PipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow livePipeLLM RelayAuto RouterNow live

PipeLLM Relay · Auto Router for production AI

Every model.
One governed path.

PipeLLM Relay automatically routes every request to an approved, healthy, cost-efficient model—behind one endpoint.

Relay
01

Every model behind one endpoint

02

Lower cost for eligible workloads

03

Fail over without rewrites

Auto Router · Cost-aware by default

Route each request to the lowest-cost approved model that still fits.

Relay evaluates prompt size, context fit, catalog discounts, policy, region, and provider health before it sends the request. Your application keeps one stable integration.

Copy your Relay endpoint

Relay / Auto Router Live

Routing decision

200 OK

Requested

Claude Sonnet 4.6

Auto-routed to

AWS Bedrock

First token

428ms

Total time

3.86s

Billed cost

$0.065530

Response streamed

Fallback ready
OpenAI
Anthropic
Google Gemini
xAI
AWS Bedrock
DeepSeek
Kimi AI
Z.AI
DeepInfra
Volcano Engine
OpenAI. Anthropic. Google Gemini. xAI. AWS Bedrock. DeepSeek. Kimi AI. Z.AI. DeepInfra. Volcano Engine.

Auto Router strategies

Decide what every route optimizes for.

Choose a strategy per environment, team, or API key. Relay Auto Router evaluates it on every request without moving the decision into application code.

relay/strategy/resilience Evaluating

AWS Bedrock

claude-sonnet-4-6

eligible · 200K ctx

GCP Vertex

claude-sonnet-4-5-20250929

eligible · 200K ctx

OpenAI

gpt-5.4

eligible · 1M ctx

Primary signal
Provider health
Resolved route
Claude Sonnet 4.6 → AWS Bedrock
Decision
Failover ready

Built for builders.
Ready for operators.

Engineering gets a stable integration. Platform teams get the controls needed to approve, route, and replace models safely.

Let Auto Router decide.

Select approved providers by health, context fit, input and output price, discounts, and availability—not hard-coded application logic.

Keep access governed.

Centralize model approvals, API keys, budgets, and provider credentials outside application source code.

Change without migration.

Keep OpenAI-, Anthropic-, and Gemini-compatible clients while Relay changes the route underneath.

Auto Router · Catalog-aware routing

Put price beside every route.

Relay Auto Router evaluates the model catalog together with policy, provider health, region, and workload fit—before a request leaves your gateway.

Cost impact · eligible traffic

$150 $25.20

10M output tokens

$150.00

$118.80

$87.60

$56.40

$25.20

0%25%50%75%100%

Share of eligible traffic using the discounted route. This compares price only; Relay moves a request only when its policy and workload requirements still match.

Savings without instability

Use the lower-cost route. Keep a healthy one ready.

Relay evaluates discounts and token economics before routing. Long-prompt workloads can favor an approved model with enough context and a lower input-token price, while healthy fallbacks keep production traffic moving.

01

Discount-aware

Apply catalog discounts before the request is sent.

02

Prompt-aware

Compare input price and context fit when prompts run long.

03

Policy-gated

Preserve model, region, tool, and workload requirements.

04

Health-routed

Keep an approved fallback ready when the primary is unhealthy.

Long-prompt economics

More input tokens make route price matter.

Relay can account for prompt size before choosing among eligible routes—not after the bill arrives.

01 · Measure

Estimate the request's input-token volume and required context window.

02 · Compare

Evaluate input price, catalog discounts, region, and policy across eligible models.

03 · Route

Select a healthy lower-cost route only when workload requirements still match.

Public routing benchmark

Input price, context, and provider—side by side.

Current catalog data. This table compares route economics and availability, not model quality.

10M+

requests routed monthly

61

active catalog models

16

provider routes

9

model publishers

Customer-verifiable evidence

Every saving leaves a request-level record.

Cost

List price, applied discount, tokens, and final request cost.

Route

Requested model, selected provider, region, and policy match.

Stability

Health signal, failover decision, latency, and response status.

Catalog snapshot: August 6, 2026. Cost trend assumes 10M output tokens and shifts only eligible traffic from a $15/M route to DeepSeek V4 Pro at the current 30% PipeLLM discount. Actual spend depends on workload, token mix, cache usage, and provider terms.

One control surface

Control before the call. Evidence after it.

Access

Keys and teams

Issue scoped credentials and keep provider secrets out of application environments.

Policy

Approved model sets

Define which models, providers, regions, and tools each workload may use.

Operations

Health and failover

React to upstream availability and rate limits using approved fallback routes.

Evidence

Request-level records

Connect routing decisions, usage, and approvals to a trace your team can inspect.

FAQ

Questions before you route.

Relay is PipeLLM's governed Auto Router. Your application sends one compatible request, and Relay automatically selects an eligible model and provider using policy, health, region, availability, and cost signals.

Relay is designed for agent platforms, multi-model products, enterprise AI teams, and evaluation pipelines that need one governed route across models and providers.

No. Relay exposes OpenAI-, Anthropic-, and Gemini-compatible paths. Most teams switch by changing the base URL and API key.

Relay can support provider-managed routing and bring-your-own-provider configurations, depending on the provider and deployment model.

Relay evaluates approved fallback routes and can move eligible requests to a healthy provider without changing the client request.

Yes. Model access, routing strategy, budgets, and credentials can be scoped by environment, team, or API key.

Each request carries route, provider, usage, and policy context that can be connected to PipeLLM Lens for review and audit.

Models change.
Your integration stays put.

Start with one endpoint today. Relay Auto Router keeps every future provider, policy, cost, and failover decision behind the same integration.