Approved catalogue_

Choose the model. Keep the contract.

Provider pricing stays visible. Routing and governance stay behind one API.

ROUTABLE MODELS58 of 61 available
Z.AI

GLM 5.3

glm-5.3

Approved for governed routing through Relay.

texttoolscache
CONTEXT
1M context
INPUT
$1.4/M
OUTPUT
$4.4/M
Anthropic

Claude Opus 5

claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

texttoolscache
CONTEXT
1M context
INPUT
$5/M
OUTPUT
$25/M
Kimi AI

Kimi K3

kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

texttoolscache
CONTEXT
1M context
INPUT
$3/M
OUTPUT
$15/M
DISCOUNT30% OFF
DeepSeek

Deepseek V4 Pro

deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model aimed at advanced reasoning, coding, long-context analysis, and long-horizon agent workflows. It is suited for full-codebase analysis, multi-step automation, and large-scale information synthesis where capability and efficiency both matter.

texttoolscache
CONTEXT
1M context
INPUT
$1.26/M$1.8/M
OUTPUT
$2.52/M$3.6/M
OpenAI

GPT 5.6 Sol

gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

texttoolscache
CONTEXT
1M context
INPUT
$5/M
OUTPUT
$30/M
OpenAI

GPT 5.6 Terra

gpt-5.6-terra

GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced.

texttoolscache
CONTEXT
1M context
INPUT
$2/M
OUTPUT
$12/M
OpenAI

GPT 5.6 Luna

gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

texttoolscache
CONTEXT
1M context
INPUT
$0.2/M
OUTPUT
$1.2/M
DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-focused Mixture-of-Experts model built for fast inference, high-throughput applications, coding assistants, chat systems, and agent workflows. It keeps strong reasoning and coding performance while prioritizing responsiveness and cost efficiency.

texttools
CONTEXT
1M context
INPUT
$0.1/M
OUTPUT
$0.2/M
xAI

Grok 4.5

grok-4.5

Grok 4.5 is xAI's smartest model, with frontier performance across coding, knowledge work, and STEM.

textimagetools
CONTEXT
500K context
INPUT
$2/M
OUTPUT
$6/M
Anthropic

Claude Sonnet 5

claude-sonnet-5

Claude Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels, a 1M-token context window, and text, image, and file inputs.

textimagetools
CONTEXT
1M context
INPUT
$2/M
OUTPUT
$10/M
Google

Gemini 3.6 Flash

gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

texttoolscache
CONTEXT
1M context
INPUT
$1.5/M
OUTPUT
$7.5/M
OpenAI

GPT Image 2

gpt-image-2

OpenAI's latest image generation model. Supports high-fidelity image generation and editing via the dedicated Images API.

textimage
CONTEXT
N/A
INPUT
$8/M
OUTPUT
$30/M
OpenAI

GPT-5.4 Mini

gpt-5.4-mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

texttoolscache
CONTEXT
N/A
INPUT
$0.75/M
OUTPUT
$4.5/M
OpenAI

GPT-5

gpt-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

texttoolscache
CONTEXT
400K context
INPUT
$1.25/M
OUTPUT
$10/M
Google

Gemini 3.1 Flash-Lite Preview

gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across...

text
CONTEXT
N/A
INPUT
$0.25/M
OUTPUT
$1.5/M
Google

Gemini 3.1 Flash-Lite Image

gemini-3.1-flash-lite-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

textimage
CONTEXT
N/A
INPUT
$0.25/M
OUTPUT
$1.5/M
Google

Gemini 3.1 Flash Image Preview

gemini-3.1-flash-image-preview

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines...

textimage
CONTEXT
N/A
INPUT
$0.5/M
OUTPUT
$3/M
Google

Gemini 3 Pro Image

gemini-3-pro-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

textimage
CONTEXT
N/A
INPUT
$2/M
OUTPUT
$12/M
Google

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

text
CONTEXT
N/A
INPUT
$0.1/M
OUTPUT
$0.4/M
Google

Gemini 2.5 Flash

gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

text
CONTEXT
1M context
INPUT
$0.3/M
OUTPUT
$2.5/M
Anthropic

Claude Fable 5

claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports reasoning and long-running, complex, asynchronous tasks that previously required frequent human check-ins, with a 1M-token context window and text, image, and file inputs.

texttoolscache
CONTEXT
1M context
INPUT
$10/M
OUTPUT
$50/M
MiniMax

MiniMax M2.7

minimax-m2.7

MiniMax M2.7 is an agentic productivity model designed for autonomous task execution, planning, and continuous improvement across complex real-world environments. It focuses on multi-agent collaboration, coding, and business workflows that require sustained reasoning and iteration.

texttoolscache
CONTEXT
204.8K context
INPUT
$0.3/M
OUTPUT
$1.2/M
MiniMax

MiniMax M2.5

minimax-m2.5

MiniMax M2.5 is a productivity-focused language model for coding, office document work, multi-step planning, and agentic workflows. It extends MiniMax’s coding strengths into broader real-world digital work such as document, spreadsheet, and presentation-oriented tasks.

texttoolscache
CONTEXT
204.8K context
INPUT
$0.3/M
OUTPUT
$1.2/M
Kimi AI

Kimi K2.6

kimi-k2.6

Kimi K2.6 is Moonshot AI’s next-generation multimodal model for long-horizon coding, coding-driven UI and UX generation, and multi-agent orchestration. It targets complex end-to-end tasks across languages such as Python, Rust, and Go, including production-ready interface generation from prompts and visual inputs.

textimagetools
CONTEXT
262.1K context
INPUT
$6.5/M
OUTPUT
$27/M