Every model. Every harness.
One API that speaks OpenAI and Anthropic natively — drop it into any compatible harness without rewriting a client.
/api/v1/chat/completions
- Prompt Caching
- Spend Tracking
- Never Downgraded
glm-5.3-flash$0.0300USD per 1M input tokensqwen3.8-flash$0.0300USD per 1M input tokensgpt-6-luna$0.0350USD per 1M input tokensgpt-5.6-luna$0.0400USD per 1M input tokensminimax-m3$0.0420USD per 1M input tokensdeepseek-v4-flash-0731$0.0506USD per 1M input tokensdeepseek-v4.1-flash$0.0525USD per 1M input tokensmimo-v2.6-flash$0.0560USD per 1M input tokensgemini-3.8-flash$0.0675USD per 1M input tokensqwen3.8-omni-flash$0.0750USD per 1M input tokensglm-5.3$0.0840USD per 1M input tokensqwen3.8-max-0902$0.1200USD per 1M input tokenskimi-k3$0.1500USD per 1M input tokensmimo-v2.6-pro$0.15225USD per 1M input tokensgpt-5.6-terra$0.1600USD per 1M input tokensdeepseek-v4-pro-0813$0.1980USD per 1M input tokensgpt-6-sol$0.2000USD per 1M input tokensgrok-4.6$0.2000USD per 1M input tokensgrok-4.7$0.2000USD per 1M input tokensgpt-5.6-sol$0.2500USD per 1M input tokensgpt-6-astra$0.3000USD per 1M input tokensCompare the numbers.
| Relative | |||||
|---|---|---|---|---|---|
| GLM 5.3 Flash | 0.0300 | 0.0060 | 0.1000 | 0.0382 | |
| Qwen 3.8 Flash | 0.0300 | 0.0032 | 0.0940 | 0.03474 | |
| GPT-6 Luna | 0.0350 | 0.0035 | 0.1750 | 0.0567 | |
| GPT 5.6 Luna | 0.0400 | 0.0040 | 0.2400 | 0.0748 | |
| MiniMax M3 | 0.0420 | 0.0084 | 0.1680 | 0.06048 | |
Blended assumes 70% of input served from cache and output equal to 25% of input volume — a reference mix for comparison only, not a billed rate. The gateway records usage after a successful response; rates shown are the currently configured rates.
Change one base URL.
Keep the request shape your application already understands. Change the model ID when the workload changes.
one base URL- Chat Completions
POST /chat/completions - Responses
POST /responses - Messages
POST /messages
Cache decides the bill.
Message 12 of a conversation costs 3.5× less than it would without cache — because every message resends everything before it, and we charge full price only for the part that is new.
$1.06instead of $3.70 without cache. The first message has nothing cached yet, so it is billed in full.
GLM 5.3 Flash rates, USD per 1M tokens — a cache hit is billed at 5× less than a miss. Illustration: a conversation of 12 messages, each one request that adds 10,000 input and 1,000 output tokens and resends everything before it, priced per 1,000 requests. Token counts are illustrative; the rates are published. Cache ratio differs per model.
Spend less on every token.
Pay only for what you use. Input starts at $0.0300 per 1M tokens — lower cache rates are applied automatically, per model.