LLM Proxy

One integration. Many supported models.

Integrate once. Use the model that fits.

Connect through one authenticated API—or start with the official Go, Python, and command-line clients. Then route canonical messages to any supported text provider and model without adding another provider SDK, credential flow, or request contract to your product.

  • Keep one canonical request contract as models change
  • Start with HTTP, Go, Python, or the command line
  • Keep upstream credentials out of client requests

One endpointRoute across native and compatible provider APIs.

One credentialApplications send a tenant secret, never provider keys.

One contractConsistent request, error, usage, and timeout behavior.

One integration. Choose the exact route.

12 families · 43 exact models · 43 offerings
Your product HTTP · Go · Python · CLI
LLM Proxy Authenticate · validate · route

Choose a model family

Choose an exact model

Claude Fable1 exact model

Choose a provider offering

Provider offerings1 route

Selected route: anthropicclaude-fable-5

Integrate once

Use the API directly or start with a client.

Every integration surface targets the same canonical POST /v2 messages contract. Provider selection changes the route, not the shape your application owns.

  1. 1
    Choose an integration

    Use ordinary HTTP or an official Go, Python, or command-line client.

  2. 2
    Authenticate once

    Give the application one client key while provider credentials remain behind the proxy.

  3. 3
    Select the model

    Use centrally managed defaults or choose a supported provider and model for one request.

HTTP API

Any language. No client dependency.

Send canonical JSON over HTTPS and keep provider-specific wire formats out of application code.

Read the v2 reference
Official Go client

Typed requests and media attachments.

Construct validated messages, attach supported media, and preserve typed transport and HTTP failures.

Use the Go client
Official Python client

A small synchronous client.

Send the same canonical messages with explicit configuration and provider/model selection.

Use the Python client
Official CLI

Pipe prompts from scripts and tools.

Install the reusable command for stdin workflows, output limits, and request-level time budgets.

Install the CLI

Built for the way teams ship

One boundary. Three ways to benefit.

Start with the path closest to your work. Each path reaches the same API, client-key boundary, and supported model catalog.

AI-assisted builders

Turn generated code into a durable integration.

Give coding agents and rapid prototypes one documented request shape instead of accumulating provider SDKs and raw keys.

Start with a copyable request
Startups and product teams

Test model fit without multiplying integrations.

Compare supported providers behind the same product boundary, then change the route as quality, capability, or product needs evolve.

Compare native providers
Platform and engineering teams

Standardize access across applications.

Centralize provider credentials, routing validation, usage signals, and client access while product teams keep one contract.

Design an internal gateway

What stays stable

Change providers without rebuilding the product boundary.

LLM Proxy owns provider-specific wire contracts, credential handling, route validation, and observable failure behavior so application code can stay focused on the job it performs.

01

Switch routes, not integrations.

Select a supported provider and model per request, use managed defaults, or update an application-owned model profile without replacing the client contract.

  • Provider selection
  • Model defaults
  • Reasoning effort
02

Keep secrets server-side.

Tenant-secret authentication gives each application access without exposing upstream provider keys. Client access can be generated and rotated independently, while managed provider keys are encrypted at rest and explicitly verified before use.

Explore the security boundary
03

Keep one request and error contract.

Send canonical messages while the proxy maps provider payloads, output limits, timeouts, usage metadata, and provider failures into one documented boundary.

Inspect the API contract
04

Use capabilities the route declares.

Send images or audio, request reasoning or constrained web search, or transcribe audio only when the selected model explicitly publishes that capability.

Inspect model capabilities

Current runtime contract

Providers and model capabilities.

This matrix is generated from the same validated provider registry used by request routing. It describes proxy support—not whether a specific account has configured a provider key.

11Providers

11Publishers

24Families

66Exact models

67Offerings

Provider-independent exact models and every current provider offering.
Anthropicanthropic claude-fable-5Claude Fable · claude-fable-5
  • Anthropicanthropicanthropic_messages128000 token output
Anthropicanthropic claude-haiku-4-5Claude Haiku · claude-haiku-4-5
  • Anthropicanthropicanthropic_messages64000 token output
Anthropicanthropic claude-haiku-4-5-20251001Claude Haiku · claude-haiku-4-5-20251001
  • Anthropicanthropicanthropic_messages64000 token output
Anthropicanthropic claude-opus-4-1Claude Opus · claude-opus-4-1
  • Anthropicanthropicanthropic_messages32000 token output
Anthropicanthropic claude-opus-4-1-20250805Claude Opus · claude-opus-4-1-20250805
  • Anthropicanthropicanthropic_messages32000 token output
Anthropicanthropic claude-opus-4-8Claude Opus · claude-opus-4-8
  • Anthropicanthropicanthropic_messages128000 token output
Anthropicanthropic claude-sonnet-4-5Claude Sonnet · claude-sonnet-4-5
  • Anthropicanthropicanthropic_messages64000 token output
Anthropicanthropic claude-sonnet-4-5-20250929Claude Sonnet · claude-sonnet-4-5-20250929
  • Anthropicanthropicanthropic_messages64000 token output
Anthropicanthropic claude-sonnet-4-6Claude Sonnet · claude-sonnet-4-6
  • Anthropicanthropicanthropic_messages64000 token output
Anthropicanthropic claude-sonnet-5Claude Sonnet · claude-sonnet-5
  • Anthropicanthropicanthropic_messages128000 token output
DeepSeekdeepseek deepseek-chatDeepSeek V3 · deepseek-chat
  • DeepSeekdeepseekopenai_chat_completionsProvider-enforced output
DeepSeekdeepseek deepseek-reasonerDeepSeek R1 · deepseek-reasoner
  • DeepSeekdeepseekopenai_chat_completionsProvider-enforced output
  • SiliconFlowsiliconflowopenai_chat_completionsProvider-enforced output
DeepSeekdeepseek deepseek-v4-flashDeepSeek V4 · deepseek-v4-flash
  • DeepSeekdeepseekopenai_chat_completionsProvider-enforced output
DeepSeekdeepseek deepseek-v4-proDeepSeek V4 · deepseek-v4-pro
  • DeepSeekdeepseekopenai_chat_completionsProvider-enforced output
Googlegoogle gemini-2.5-flashGemini · gemini-2.5-flash
  • Geminigeminigemini_interactions65536 token output
Googlegoogle gemini-2.5-flash-liteGemini · gemini-2.5-flash-lite
  • Geminigeminigemini_interactions65536 token output
Googlegoogle gemini-2.5-proGemini · gemini-2.5-pro
  • Geminigeminigemini_interactions65536 token output
Googlegoogle gemini-3-flash-previewGemini · gemini-3-flash-preview
  • Geminigeminigemini_interactions65536 token output
Googlegoogle gemini-3.1-flash-liteGemini · gemini-3.1-flash-lite
  • Geminigeminigemini_interactions65536 token output
Googlegoogle gemini-3.1-pro-previewGemini · gemini-3.1-pro-preview
  • Geminigeminigemini_interactions65536 token output
Googlegoogle gemini-3.5-flashGemini · gemini-3.5-flash
  • Geminigeminigemini_interactions65536 token output
Z.AIzai glm-5.1GLM-5 · glm-5.1
  • Z.AIzaiopenai_chat_completionsProvider-enforced output
Z.AIzai glm-5.2GLM-5 · glm-5.2
  • Z.AIzaiopenai_chat_completions131072 token output
Z.AIzai glm-asr-2512GLM ASR · glm-asr-2512
  • Z.AIzaimultipart_transcription
OpenAIopenai gpt-4.1GPT-4 · gpt-4.1
  • OpenAIopenaiopenai_responsesProvider-enforced output
OpenAIopenai gpt-4oGPT-4 · gpt-4o
  • OpenAIopenaiopenai_responsesProvider-enforced output
OpenAIopenai gpt-4o-miniGPT-4 · gpt-4o-mini
  • OpenAIopenaiopenai_responsesProvider-enforced output
OpenAIopenai gpt-4o-mini-transcribeGPT Transcribe · gpt-4o-mini-transcribe
  • OpenAIopenaimultipart_transcription
OpenAIopenai gpt-4o-transcribeGPT Transcribe · gpt-4o-transcribe
  • OpenAIopenaimultipart_transcription
OpenAIopenai gpt-5GPT-5 · gpt-5
  • OpenAIopenaiopenai_responsesReasoning: minimal, low, medium, highProvider-enforced output
OpenAIopenai gpt-5-miniGPT-5 · gpt-5-mini
  • OpenAIopenaiopenai_responsesReasoning: minimal, low, medium, highProvider-enforced output
OpenAIopenai gpt-5.5GPT-5 · gpt-5.5
  • OpenAIopenaiopenai_responsesReasoning: none, low, medium, high, xhighProvider-enforced output
OpenAIopenai gpt-5.5-proGPT-5 · gpt-5.5-pro
  • OpenAIopenaiopenai_responsesReasoning: medium, high, xhighProvider-enforced output
OpenAIopenai gpt-5.6GPT-5 · gpt-5.6
  • OpenAIopenaiopenai_responsesReasoning: none, low, medium, high, xhigh, maxProvider-enforced output
OpenAIopenai gpt-5.6-lunaGPT-5 · gpt-5.6-luna
  • OpenAIopenaiopenai_responsesReasoning: none, low, medium, high, xhigh, maxProvider-enforced output
OpenAIopenai gpt-5.6-solGPT-5 · gpt-5.6-sol
  • OpenAIopenaiopenai_responsesReasoning: none, low, medium, high, xhigh, maxProvider-enforced output
OpenAIopenai gpt-5.6-terraGPT-5 · gpt-5.6-terra
  • OpenAIopenaiopenai_responsesReasoning: none, low, medium, high, xhigh, maxProvider-enforced output
xAIxai grok-4.20-0309-non-reasoningGrok · grok-4.20-0309-non-reasoning
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-4.20-0309-reasoningGrok · grok-4.20-0309-reasoning
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-4.3Grok · grok-4.3
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-4.3-latestGrok · grok-4.3-latest
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-4.5Grok · grok-4.5
  • xAIxaiopenai_responsesProvider-enforced output
xAIxai grok-build-0.1Grok Build · grok-build-0.1
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-code-fastGrok Code · grok-code-fast
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-code-fast-1Grok Code · grok-code-fast-1
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-code-fast-1-0825Grok Code · grok-code-fast-1-0825
  • xAIxaiopenai_chat_completionsProvider-enforced output
xAIxai grok-imagine-video-1.5Grok Imagine · grok-imagine-video-1.5
  • xAIxaixai_videos_generations
xAIxai grok-latestGrok · grok-latest
  • xAIxaiopenai_chat_completionsProvider-enforced output
Moonshot AImoonshot kimi-k2.6Kimi K2 · kimi-k2.6
  • Moonshotmoonshotopenai_chat_completionsProvider-enforced output
Moonshot AImoonshot kimi-k2.7-codeKimi K2 · kimi-k2.7-code
  • Moonshotmoonshotopenai_chat_completionsProvider-enforced output
Moonshot AImoonshot kimi-k2.7-code-highspeedKimi K2 · kimi-k2.7-code-highspeed
  • Moonshotmoonshotopenai_chat_completionsProvider-enforced output
Moonshot AImoonshot kimi-k3Kimi K3 · kimi-k3
  • Moonshotmoonshotopenai_chat_completionsReasoning: low, high, maxProvider-enforced output
MiniMaxminimax minimax-m2MiniMax M2 · minimax-m2
  • MiniMaxminimaxopenai_chat_completions204800 token output
MiniMaxminimax minimax-m2.1MiniMax M2 · minimax-m2.1
  • MiniMaxminimaxopenai_chat_completions204800 token output
MiniMaxminimax minimax-m2.1-highspeedMiniMax M2 · minimax-m2.1-highspeed
  • MiniMaxminimaxopenai_chat_completions204800 token output
MiniMaxminimax minimax-m2.5MiniMax M2 · minimax-m2.5
  • MiniMaxminimaxopenai_chat_completions204800 token output
MiniMaxminimax minimax-m2.5-highspeedMiniMax M2 · minimax-m2.5-highspeed
  • MiniMaxminimaxopenai_chat_completions204800 token output
MiniMaxminimax minimax-m2.7MiniMax M2 · minimax-m2.7
  • MiniMaxminimaxopenai_chat_completions204800 token output
MiniMaxminimax minimax-m2.7-highspeedMiniMax M2 · minimax-m2.7-highspeed
  • MiniMaxminimaxopenai_chat_completions204800 token output
Metameta muse-spark-1.1Muse Spark · muse-spark-1.1
  • Metametaopenai_chat_completionsProvider-enforced output
Alibabaalibaba qwen-plusQwen · qwen-plus
  • DashScopedashscopeopenai_chat_completionsProvider-enforced output
Alibabaalibaba qwen3.6-flashQwen · qwen3.6-flash
  • DashScopedashscopeopenai_chat_completions65536 token output
Alibabaalibaba qwen3.7-maxQwen · qwen3.7-max
  • DashScopedashscopeopenai_chat_completions65536 token output
Alibabaalibaba qwen3.7-plusQwen · qwen3.7-plus
  • DashScopedashscopeopenai_chat_completions65536 token output
FunAudioLLMfunaudio sensevoice-smallSenseVoice · sensevoice-small
  • SiliconFlowsiliconflowmultipart_transcription
xAIxai xai-sttxAI STT · xai-stt
  • xAIxaimultipart_transcription

4 MiBMaximum JSON request body

25 MiBMaximum input audio

3600 secondsMaximum request work budget

Current support contract

The catalog is the source of truth.

The public matrix is generated from the same validated provider registry used by request routing. It states exactly which providers, models, and capabilities the current proxy release supports.

  • Provider and model routes are explicit, validated runtime contracts.
  • Unsupported models and capabilities fail before provider dispatch.
  • The canonical client contract stays stable while provider adapters own upstream API differences.
  • Provider lifecycle, model-onboarding, and hosted uptime commitments are outside the current catalog contract pending an approved support and SLA policy.

One integration, ongoing choice

Start once. Keep choosing the model that fits.

Choose HTTP, Go, Python, or the command line, then use the current model catalog to select the route your product needs.