Back to approved models
xAIactiverelay/request.mjsready to route CONTEXT200K contextprompt window MAX OUTPUTN/Aper response INPUT$0.20 / Mtoken price OUTPUT$0.50 / Mtoken price
Approved model profile_
Grok 4 Fast Reasoning
[Grok 4.1 Fast is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window. Reasoning can be enabled/disabled using the `reasoning` `enabled` parameter in the API. Learn more in our docs]
grok-4-fast-reasoningconst completion = await client.chat.completions.create({
model: "grok-4-fast-reasoning",
messages: [
{ role: "user",
content: "Plan a multi-step task" }
]
});Route configuration_
Choose a supported call format.
Select a format below to view its documentation. Use the displayed converter route when calling through PipeLLM.
https://api.pipellm.aiOne base URL for OpenAI Chat Completions, Responses, Anthropic, and Gemini.SUPPORTED CALL FORMATS
OpenAI/openai/v1/chat/completionsDocs Responses/responses/v1/responsesDocs Anthropic/anthropic/v1/messagesDocs Gemini/gemini/v1beta/models/{model}:generateContentDocs Choose a call format to open its request and response documentation.grok-4-fast-reasoningUse this in every requestPROVIDER AVAILABILITY2 routes
| Provider | Context | Max output | Discount | Input / M | Output / M |
|---|---|---|---|---|---|
| X.AI | 200K context | N/A | — | $0.20 / M | $0.50 / M |
| Grok AI | N/A | N/A | — | $0.20 / M | $0.50 / M |
SUPPORTED SURFACESactive
input: textinput: imageoutput: texttool use
Relay governs model access with the same routing policy that Lens keeps visible and auditable.
