LLM API pricing
134 models · 9 providers · all rates per million tokens, USD
This is not a hand-maintained marketing table. It is the exact pricing data the BurnLens proxy bills every request from, generated from burnlens/cost/pricing_data in the open-source repo. When a rate changes there, this page changes with it.
Cache columns matter more than they look. A coding agent re-sends its whole context every turn, so on Anthropic-style billing 90–99% of prompt tokens are cache reads at a tenth of the input rate — pricing a run off the input column alone overstates it by an order of magnitude.
Anthropic
anthropic · 23 models · rates verified 2026-08-10
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
claude-2.0 | $8.00 | $24.00 | — | — | — |
claude-2.1 | $8.00 | $24.00 | — | — | — |
claude-3-5-haiku-20241022 | $0.80 | $4.00 | $0.08 | $1.00 | — |
claude-3-5-sonnet-20240620 | $3.00 | $15.00 | $0.30 | $3.75 | — |
claude-3-5-sonnet-20241022 | $3.00 | $15.00 | $0.30 | $3.75 | — |
claude-3-7-sonnet-20250219 | $3.00 | $15.00 | $0.30 | $3.75 | — |
claude-3-haiku-20240307 | $0.25 | $1.25 | $0.03 | $0.30 | — |
claude-3-opus-20240229 | $15.00 | $75.00 | $1.50 | $18.75 | — |
claude-3-sonnet-20240229 | $3.00 | $15.00 | — | — | — |
claude-fable-5 | $10.00 | $50.00 | $1.00 | $12.50 | — |
claude-haiku-4-5 | $1.00 | $5.00 | $0.10 | $1.25 | — |
claude-mythos-5 | $10.00 | $50.00 | $1.00 | $12.50 | — |
claude-opus-4 | $15.00 | $75.00 | $1.50 | $18.75 | — |
claude-opus-4-1 | $15.00 | $75.00 | $1.50 | $18.75 | — |
claude-opus-4-5 | $5.00 | $25.00 | $0.50 | $6.25 | — |
claude-opus-4-6 | $5.00 | $25.00 | $0.50 | $6.25 | — |
claude-opus-4-7 | $5.00 | $25.00 | $0.50 | $6.25 | — |
claude-opus-4-8 | $5.00 | $25.00 | $0.50 | $6.25 | — |
claude-opus-5 | $5.00 | $25.00 | $0.50 | $6.25 | — |
claude-sonnet-4 | $3.00 | $15.00 | $0.30 | $3.75 | — |
claude-sonnet-4-5 | $3.00 | $15.00 | $0.30 | $3.75 | — |
claude-sonnet-4-6 | $3.00 | $15.00 | $0.30 | $3.75 | — |
claude-sonnet-5 | $2.00 | $10.00 | $0.20 | $2.50 | from 2026-09-01: $3.00 in / $15.00 out |
AWS Bedrock
bedrock · 10 models · rates verified 2026-07-18
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
anthropic.claude-fable-5 | $10.00 | $50.00 | $1.00 | $12.50 | — |
anthropic.claude-haiku-4-5 | $1.00 | $5.00 | $0.10 | $1.25 | — |
anthropic.claude-opus-4-5 | $5.00 | $25.00 | $0.50 | $6.25 | — |
anthropic.claude-opus-4-6 | $5.00 | $25.00 | $0.50 | $6.25 | — |
anthropic.claude-opus-4-7 | $5.00 | $25.00 | $0.50 | $6.25 | — |
anthropic.claude-opus-4-8 | $5.00 | $25.00 | $0.50 | $6.25 | — |
anthropic.claude-sonnet-4 | $3.00 | $15.00 | $0.30 | $3.75 | — |
anthropic.claude-sonnet-4-5 | $3.00 | $15.00 | $0.30 | $3.75 | — |
anthropic.claude-sonnet-4-6 | $3.00 | $15.00 | $0.30 | $3.75 | — |
anthropic.claude-sonnet-5 | $2.00 | $10.00 | $0.20 | $2.50 | from 2026-09-01: $3.00 in / $15.00 out |
DeepSeek
deepseek · 4 models · rates verified 2026-07-20
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
deepseek-chat | $0.14 | $0.28 | $0.0028 | — | — |
deepseek-reasoner | $0.14 | $0.28 | $0.0028 | — | — |
deepseek-v4-flash | $0.14 | $0.28 | $0.0028 | — | — |
deepseek-v4-pro | $0.435 | $0.87 | $0.0036 | — | — |
google · 13 models · rates verified 2026-07-23
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
gemini-1.0-pro | $0.50 | $1.50 | — | — | — |
gemini-1.5-flash | $0.075 | $0.30 | — | — | — |
gemini-1.5-flash-8b | $0.0375 | $0.15 | — | — | — |
gemini-1.5-pro | $1.25 | $5.00 | — | — | — |
gemini-2.0-flash | $0.10 | $0.40 | — | — | — |
gemini-2.0-flash-lite | $0.075 | $0.30 | — | — | — |
gemini-2.5-flash | $0.30 | $2.50 | — | — | — |
gemini-2.5-flash-lite | $0.10 | $0.40 | — | — | — |
gemini-2.5-pro | $1.25 | $10.00 | — | — | over 200k ctx: $2.50 in / $15.00 out |
gemini-3-flash-preview | $0.50 | $3.00 | — | — | — |
gemini-3.1-flash-lite | $0.25 | $1.50 | — | — | — |
gemini-3.1-pro-preview | $2.00 | $12.00 | — | — | over 200k ctx: $4.00 in / $18.00 out |
gemini-3.5-flash | $1.50 | $9.00 | $0.15 | — | — |
Groq
groq · 7 models · rates verified 2026-07-17
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
llama-3.1-8b-instant | $0.05 | $0.08 | — | — | — |
llama-3.3-70b-versatile | $0.59 | $0.79 | — | — | — |
moonshotai/kimi-k2-instruct | $1.00 | $3.00 | — | — | — |
openai/gpt-oss-120b | $0.15 | $0.75 | — | — | — |
openai/gpt-oss-20b | $0.075 | $0.30 | — | — | — |
qwen/qwen3-32b | $0.29 | $0.59 | — | — | — |
qwen/qwen3.6-27b | $0.60 | $3.00 | — | — | — |
Mistral
mistral · 11 models · rates verified 2026-07-17
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
codestral-latest | $0.30 | $0.90 | — | — | — |
magistral-medium-latest | $2.00 | $5.00 | — | — | — |
ministral-3b-latest | $0.04 | $0.04 | — | — | — |
ministral-8b-latest | $0.10 | $0.10 | — | — | — |
mistral-large-2512 | $0.50 | $1.50 | — | — | — |
mistral-large-latest | $0.50 | $1.50 | — | — | — |
mistral-medium-3-5 | $1.50 | $7.50 | — | — | — |
mistral-medium-latest | $1.50 | $7.50 | — | — | — |
mistral-small-2603 | $0.15 | $0.60 | — | — | — |
mistral-small-latest | $0.15 | $0.60 | — | — | — |
open-mistral-nemo | $0.15 | $0.15 | — | — | — |
OpenAI
openai · 39 models · rates verified 2026-07-17
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
gpt-3.5-turbo | $0.50 | $1.50 | — | — | — |
gpt-3.5-turbo-instruct | $1.50 | $2.00 | — | — | — |
gpt-4 | $30.00 | $60.00 | — | — | — |
gpt-4-32k | $60.00 | $120.00 | — | — | — |
gpt-4-turbo | $10.00 | $30.00 | — | — | — |
gpt-4-turbo-preview | $10.00 | $30.00 | — | — | — |
gpt-4.1 | $2.00 | $8.00 | $0.50 | — | — |
gpt-4.1-mini | $0.40 | $1.60 | $0.10 | — | — |
gpt-4.1-nano | $0.10 | $0.40 | $0.025 | — | — |
gpt-4o | $2.50 | $10.00 | $1.25 | — | — |
gpt-4o-audio-preview | $2.50 | $10.00 | — | — | audio: $40.00 in / $80.00 out |
gpt-4o-mini | $0.15 | $0.60 | $0.075 | — | — |
gpt-4o-mini-audio-preview | $0.15 | $0.60 | — | — | audio: $10.00 in / $20.00 out |
gpt-4o-realtime-preview | $5.00 | $20.00 | $2.50 | — | audio: $40.00 in / $80.00 out |
gpt-5 | $1.25 | $10.00 | $0.125 | — | — |
gpt-5-codex | $1.25 | $10.00 | $0.125 | — | — |
gpt-5-mini | $0.25 | $2.00 | $0.025 | — | — |
gpt-5-nano | $0.05 | $0.40 | $0.005 | — | — |
gpt-5.2 | $1.75 | $14.00 | $0.175 | — | — |
gpt-5.2-pro | $21.00 | $168.00 | — | — | reasoning: $168.00 |
gpt-5.4 | $2.50 | $15.00 | $0.25 | — | — |
gpt-5.4-mini | $0.75 | $4.50 | $0.075 | — | — |
gpt-5.4-nano | $0.20 | $1.25 | $0.02 | — | — |
gpt-5.4-pro | $30.00 | $180.00 | — | — | reasoning: $180.00 |
gpt-5.5 | $5.00 | $30.00 | $0.50 | — | — |
gpt-5.5-pro | $30.00 | $180.00 | — | — | reasoning: $180.00 |
gpt-5.6 | $5.00 | $30.00 | $0.50 | — | — |
gpt-5.6-luna | $1.00 | $6.00 | $0.10 | — | — |
gpt-5.6-sol | $5.00 | $30.00 | $0.50 | — | — |
gpt-5.6-terra | $2.50 | $15.00 | $0.25 | — | — |
gpt-realtime-2.1 | $4.00 | $24.00 | $0.40 | — | audio: $32.00 in / $64.00 out |
gpt-realtime-2.1-mini | $0.60 | $2.40 | $0.30 | — | audio: $10.00 in / $20.00 out |
o1 | $15.00 | $60.00 | $7.50 | — | reasoning: $60.00 |
o1-mini | $1.10 | $4.40 | $0.55 | — | reasoning: $4.40 |
o1-preview | $15.00 | $60.00 | $7.50 | — | reasoning: $60.00 |
o3 | $2.00 | $8.00 | $0.50 | — | reasoning: $8.00 |
o3-mini | $1.10 | $4.40 | $0.55 | — | reasoning: $4.40 |
o3-pro | $20.00 | $80.00 | — | — | reasoning: $80.00 |
o4-mini | $1.10 | $4.40 | $0.275 | — | reasoning: $4.40 |
Together AI
together · 26 models · rates verified 2026-07-17
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
LiquidAI/LFM2.5-8B-A1B | $0.03 | $0.12 | — | — | — |
MiniMaxAI/MiniMax-M2.7 | $0.30 | $1.20 | $0.06 | — | — |
MiniMaxAI/MiniMax-M3 | $0.30 | $1.20 | $0.06 | — | — |
Qwen/Qwen2.5-72B-Instruct-Turbo | $1.20 | $1.20 | — | — | — |
Qwen/Qwen2.5-7B-Instruct-Turbo | $0.30 | $0.30 | — | — | — |
Qwen/Qwen3.5-9B | $0.17 | $0.25 | — | — | — |
Qwen/Qwen3.6-Plus | $0.50 | $3.00 | — | — | — |
Qwen/Qwen3.7-Max | $1.25 | $3.75 | — | — | — |
Qwen/Qwen3.7-Plus | $0.32 | $1.28 | — | — | — |
deepcogito/cogito-v2-1-671b | $1.25 | $1.25 | — | — | — |
deepseek-ai/DeepSeek-R1 | $3.00 | $7.00 | — | — | — |
deepseek-ai/DeepSeek-V3 | $1.25 | $1.25 | — | — | — |
deepseek-ai/DeepSeek-V4-Pro | $1.74 | $3.48 | $0.20 | — | — |
google/gemma-3n-E4B-it | $0.06 | $0.12 | — | — | — |
google/gemma-4-31B-it | $0.39 | $0.97 | — | — | — |
meta-llama/Llama-3.3-70B-Instruct-Turbo | $1.04 | $1.04 | — | — | — |
meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo | $3.50 | $3.50 | — | — | — |
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo | $0.18 | $0.18 | — | — | — |
moonshotai/Kimi-K2.6 | $1.20 | $4.50 | $0.20 | — | — |
moonshotai/Kimi-K2.7-Code | $0.95 | $4.00 | $0.19 | — | — |
nvidia/nemotron-3-ultra-550b-a55b | $0.60 | $3.60 | $0.20 | — | — |
openai/gpt-oss-120b | $0.15 | $0.60 | — | — | — |
openai/gpt-oss-20b | $0.05 | $0.20 | — | — | — |
pearl-ai/gemma-4-31b-it | $0.28 | $0.86 | — | — | — |
thinkingmachines/Inkling | $1.00 | $4.05 | $0.17 | — | — |
zai-org/GLM-5.2 | $1.40 | $4.40 | $0.26 | — | — |
xAI
xai · 1 models · rates verified 2026-07-20
| Model | Input | Output | Cache read | Cache write | Notes |
|---|---|---|---|---|---|
grok-4.5 | $2.00 | $6.00 | — | — | — |
Reading the table
- Cache read / cache write — prompt-cache rates. An em dash means the provider does not bill caching separately.
- OpenAI and Google fold cached tokens into the input count; Anthropic reports them separately. Adding cache reads to input tokens on an OpenAI response double-counts them.
- over Nk ctx — a long-context tier. The higher rate applies to the whole request once the prompt crosses the threshold, not just the tokens above it.
- from YYYY-MM-DD — a rate change the provider has already announced. The proxy switches on that date with no upgrade.
- Bedrock keys omit the geo prefix (
us.,eu.,apac.,global.). These are global cross-region rates; regional inference runs roughly 10% higher.
Stop reading prices, start measuring them
A price list tells you the rate. It does not tell you what your agents actually spend, or which repo, PR, or developer spent it. BurnLens is a local proxy that does — pip install burnlens, one environment variable.
Read the docs · Scan existing agent logs · See the dashboard