Every model. Every harness.

One API that speaks OpenAI and Anthropic natively — drop it into any compatible harness without rewriting a client.

Base URL

/api/v1/chat/completions

  • Prompt Caching
  • Spend Tracking
  • Never Downgraded

From $0.0300 / 1M1,000 credits = $1

glm-5.3-flash$0.0300USD per 1M input tokensqwen3.8-flash$0.0300USD per 1M input tokensgpt-6-luna$0.0350USD per 1M input tokensgpt-5.6-luna$0.0400USD per 1M input tokensminimax-m3$0.0420USD per 1M input tokensdeepseek-v4-flash-0731$0.0506USD per 1M input tokensdeepseek-v4.1-flash$0.0525USD per 1M input tokensmimo-v2.6-flash$0.0560USD per 1M input tokensgemini-3.8-flash$0.0675USD per 1M input tokensqwen3.8-omni-flash$0.0750USD per 1M input tokensglm-5.3$0.0840USD per 1M input tokensqwen3.8-max-0902$0.1200USD per 1M input tokenskimi-k3$0.1500USD per 1M input tokensmimo-v2.6-pro$0.15225USD per 1M input tokensgpt-5.6-terra$0.1600USD per 1M input tokensdeepseek-v4-pro-0813$0.1980USD per 1M input tokensgpt-6-sol$0.2000USD per 1M input tokensgrok-4.6$0.2000USD per 1M input tokensgrok-4.7$0.2000USD per 1M input tokensgpt-5.6-sol$0.2500USD per 1M input tokensgpt-6-astra$0.3000USD per 1M input tokens
01Model catalog

Compare the numbers.

Input, cached input, and output are priced independently. No blended platform fee hidden in the rate.

USD per 1M tokens
Relative
GLM 5.3 Flash0.03000.00600.10000.0382
Qwen 3.8 Flash0.03000.00320.09400.03474
GPT-6 Luna0.03500.00350.17500.0567
GPT 5.6 Luna0.04000.00400.24000.0748
MiniMax M30.04200.00840.16800.06048

Blended assumes 70% of input served from cache and output equal to 25% of input volume — a reference mix for comparison only, not a billed rate. The gateway records usage after a successful response; rates shown are the currently configured rates.

02Compatible by design

Change one base URL.

Keep the request shape your application already understands. Change the model ID when the workload changes.

Gateway base URL
1endpointone base URL
  • Chat CompletionsPOST /chat/completions
  • ResponsesPOST /responses
  • MessagesPOST /messages
03Rate anatomy

Cache decides the bill.

Message 12 of a conversation costs 3.5× less than it would without cache — because every message resends everything before it, and we charge full price only for the part that is new.

Cost per message as the conversation growsUSD per 1,000 requests
you payalready in cache
message 1$0.40
message 12$1.06
without cache: $3.70

$1.06instead of $3.70 without cache. The first message has nothing cached yet, so it is billed in full.

GLM 5.3 Flash rates, USD per 1M tokens — a cache hit is billed at 5× less than a miss. Illustration: a conversation of 12 messages, each one request that adds 10,000 input and 1,000 output tokens and resends everything before it, priced per 1,000 requests. Token counts are illustrative; the rates are published. Cache ratio differs per model.

Spend less on every token.

Pay only for what you use. Input starts at $0.0300 per 1M tokens — lower cache rates are applied automatically, per model.

Get an API key Quick setup

Usage-based pricing · 1,000 credits = $1