BURNLENSDashboard

LLM API pricing

134 models · 9 providers · all rates per million tokens, USD

This is not a hand-maintained marketing table. It is the exact pricing data the BurnLens proxy bills every request from, generated from burnlens/cost/pricing_data in the open-source repo. When a rate changes there, this page changes with it.

Cache columns matter more than they look. A coding agent re-sends its whole context every turn, so on Anthropic-style billing 90–99% of prompt tokens are cache reads at a tenth of the input rate — pricing a run off the input column alone overstates it by an order of magnitude.

Anthropic

anthropic · 23 models · rates verified 2026-08-10

ModelInputOutputCache readCache writeNotes
claude-2.0$8.00$24.00———
claude-2.1$8.00$24.00———
claude-3-5-haiku-20241022$0.80$4.00$0.08$1.00—
claude-3-5-sonnet-20240620$3.00$15.00$0.30$3.75—
claude-3-5-sonnet-20241022$3.00$15.00$0.30$3.75—
claude-3-7-sonnet-20250219$3.00$15.00$0.30$3.75—
claude-3-haiku-20240307$0.25$1.25$0.03$0.30—
claude-3-opus-20240229$15.00$75.00$1.50$18.75—
claude-3-sonnet-20240229$3.00$15.00———
claude-fable-5$10.00$50.00$1.00$12.50—
claude-haiku-4-5$1.00$5.00$0.10$1.25—
claude-mythos-5$10.00$50.00$1.00$12.50—
claude-opus-4$15.00$75.00$1.50$18.75—
claude-opus-4-1$15.00$75.00$1.50$18.75—
claude-opus-4-5$5.00$25.00$0.50$6.25—
claude-opus-4-6$5.00$25.00$0.50$6.25—
claude-opus-4-7$5.00$25.00$0.50$6.25—
claude-opus-4-8$5.00$25.00$0.50$6.25—
claude-opus-5$5.00$25.00$0.50$6.25—
claude-sonnet-4$3.00$15.00$0.30$3.75—
claude-sonnet-4-5$3.00$15.00$0.30$3.75—
claude-sonnet-4-6$3.00$15.00$0.30$3.75—
claude-sonnet-5$2.00$10.00$0.20$2.50from 2026-09-01: $3.00 in / $15.00 out

AWS Bedrock

bedrock · 10 models · rates verified 2026-07-18

ModelInputOutputCache readCache writeNotes
anthropic.claude-fable-5$10.00$50.00$1.00$12.50—
anthropic.claude-haiku-4-5$1.00$5.00$0.10$1.25—
anthropic.claude-opus-4-5$5.00$25.00$0.50$6.25—
anthropic.claude-opus-4-6$5.00$25.00$0.50$6.25—
anthropic.claude-opus-4-7$5.00$25.00$0.50$6.25—
anthropic.claude-opus-4-8$5.00$25.00$0.50$6.25—
anthropic.claude-sonnet-4$3.00$15.00$0.30$3.75—
anthropic.claude-sonnet-4-5$3.00$15.00$0.30$3.75—
anthropic.claude-sonnet-4-6$3.00$15.00$0.30$3.75—
anthropic.claude-sonnet-5$2.00$10.00$0.20$2.50from 2026-09-01: $3.00 in / $15.00 out

DeepSeek

deepseek · 4 models · rates verified 2026-07-20

ModelInputOutputCache readCache writeNotes
deepseek-chat$0.14$0.28$0.0028——
deepseek-reasoner$0.14$0.28$0.0028——
deepseek-v4-flash$0.14$0.28$0.0028——
deepseek-v4-pro$0.435$0.87$0.0036——

Google

google · 13 models · rates verified 2026-07-23

ModelInputOutputCache readCache writeNotes
gemini-1.0-pro$0.50$1.50———
gemini-1.5-flash$0.075$0.30———
gemini-1.5-flash-8b$0.0375$0.15———
gemini-1.5-pro$1.25$5.00———
gemini-2.0-flash$0.10$0.40———
gemini-2.0-flash-lite$0.075$0.30———
gemini-2.5-flash$0.30$2.50———
gemini-2.5-flash-lite$0.10$0.40———
gemini-2.5-pro$1.25$10.00——over 200k ctx: $2.50 in / $15.00 out
gemini-3-flash-preview$0.50$3.00———
gemini-3.1-flash-lite$0.25$1.50———
gemini-3.1-pro-preview$2.00$12.00——over 200k ctx: $4.00 in / $18.00 out
gemini-3.5-flash$1.50$9.00$0.15——

Groq

groq · 7 models · rates verified 2026-07-17

ModelInputOutputCache readCache writeNotes
llama-3.1-8b-instant$0.05$0.08———
llama-3.3-70b-versatile$0.59$0.79———
moonshotai/kimi-k2-instruct$1.00$3.00———
openai/gpt-oss-120b$0.15$0.75———
openai/gpt-oss-20b$0.075$0.30———
qwen/qwen3-32b$0.29$0.59———
qwen/qwen3.6-27b$0.60$3.00———

Mistral

mistral · 11 models · rates verified 2026-07-17

ModelInputOutputCache readCache writeNotes
codestral-latest$0.30$0.90———
magistral-medium-latest$2.00$5.00———
ministral-3b-latest$0.04$0.04———
ministral-8b-latest$0.10$0.10———
mistral-large-2512$0.50$1.50———
mistral-large-latest$0.50$1.50———
mistral-medium-3-5$1.50$7.50———
mistral-medium-latest$1.50$7.50———
mistral-small-2603$0.15$0.60———
mistral-small-latest$0.15$0.60———
open-mistral-nemo$0.15$0.15———

OpenAI

openai · 39 models · rates verified 2026-07-17

ModelInputOutputCache readCache writeNotes
gpt-3.5-turbo$0.50$1.50———
gpt-3.5-turbo-instruct$1.50$2.00———
gpt-4$30.00$60.00———
gpt-4-32k$60.00$120.00———
gpt-4-turbo$10.00$30.00———
gpt-4-turbo-preview$10.00$30.00———
gpt-4.1$2.00$8.00$0.50——
gpt-4.1-mini$0.40$1.60$0.10——
gpt-4.1-nano$0.10$0.40$0.025——
gpt-4o$2.50$10.00$1.25——
gpt-4o-audio-preview$2.50$10.00——audio: $40.00 in / $80.00 out
gpt-4o-mini$0.15$0.60$0.075——
gpt-4o-mini-audio-preview$0.15$0.60——audio: $10.00 in / $20.00 out
gpt-4o-realtime-preview$5.00$20.00$2.50—audio: $40.00 in / $80.00 out
gpt-5$1.25$10.00$0.125——
gpt-5-codex$1.25$10.00$0.125——
gpt-5-mini$0.25$2.00$0.025——
gpt-5-nano$0.05$0.40$0.005——
gpt-5.2$1.75$14.00$0.175——
gpt-5.2-pro$21.00$168.00——reasoning: $168.00
gpt-5.4$2.50$15.00$0.25——
gpt-5.4-mini$0.75$4.50$0.075——
gpt-5.4-nano$0.20$1.25$0.02——
gpt-5.4-pro$30.00$180.00——reasoning: $180.00
gpt-5.5$5.00$30.00$0.50——
gpt-5.5-pro$30.00$180.00——reasoning: $180.00
gpt-5.6$5.00$30.00$0.50——
gpt-5.6-luna$1.00$6.00$0.10——
gpt-5.6-sol$5.00$30.00$0.50——
gpt-5.6-terra$2.50$15.00$0.25——
gpt-realtime-2.1$4.00$24.00$0.40—audio: $32.00 in / $64.00 out
gpt-realtime-2.1-mini$0.60$2.40$0.30—audio: $10.00 in / $20.00 out
o1$15.00$60.00$7.50—reasoning: $60.00
o1-mini$1.10$4.40$0.55—reasoning: $4.40
o1-preview$15.00$60.00$7.50—reasoning: $60.00
o3$2.00$8.00$0.50—reasoning: $8.00
o3-mini$1.10$4.40$0.55—reasoning: $4.40
o3-pro$20.00$80.00——reasoning: $80.00
o4-mini$1.10$4.40$0.275—reasoning: $4.40

Together AI

together · 26 models · rates verified 2026-07-17

ModelInputOutputCache readCache writeNotes
LiquidAI/LFM2.5-8B-A1B$0.03$0.12———
MiniMaxAI/MiniMax-M2.7$0.30$1.20$0.06——
MiniMaxAI/MiniMax-M3$0.30$1.20$0.06——
Qwen/Qwen2.5-72B-Instruct-Turbo$1.20$1.20———
Qwen/Qwen2.5-7B-Instruct-Turbo$0.30$0.30———
Qwen/Qwen3.5-9B$0.17$0.25———
Qwen/Qwen3.6-Plus$0.50$3.00———
Qwen/Qwen3.7-Max$1.25$3.75———
Qwen/Qwen3.7-Plus$0.32$1.28———
deepcogito/cogito-v2-1-671b$1.25$1.25———
deepseek-ai/DeepSeek-R1$3.00$7.00———
deepseek-ai/DeepSeek-V3$1.25$1.25———
deepseek-ai/DeepSeek-V4-Pro$1.74$3.48$0.20——
google/gemma-3n-E4B-it$0.06$0.12———
google/gemma-4-31B-it$0.39$0.97———
meta-llama/Llama-3.3-70B-Instruct-Turbo$1.04$1.04———
meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo$3.50$3.50———
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo$0.18$0.18———
moonshotai/Kimi-K2.6$1.20$4.50$0.20——
moonshotai/Kimi-K2.7-Code$0.95$4.00$0.19——
nvidia/nemotron-3-ultra-550b-a55b$0.60$3.60$0.20——
openai/gpt-oss-120b$0.15$0.60———
openai/gpt-oss-20b$0.05$0.20———
pearl-ai/gemma-4-31b-it$0.28$0.86———
thinkingmachines/Inkling$1.00$4.05$0.17——
zai-org/GLM-5.2$1.40$4.40$0.26——

xAI

xai · 1 models · rates verified 2026-07-20

ModelInputOutputCache readCache writeNotes
grok-4.5$2.00$6.00———

Reading the table

  • Cache read / cache write — prompt-cache rates. An em dash means the provider does not bill caching separately.
  • OpenAI and Google fold cached tokens into the input count; Anthropic reports them separately. Adding cache reads to input tokens on an OpenAI response double-counts them.
  • over Nk ctx — a long-context tier. The higher rate applies to the whole request once the prompt crosses the threshold, not just the tokens above it.
  • from YYYY-MM-DD — a rate change the provider has already announced. The proxy switches on that date with no upgrade.
  • Bedrock keys omit the geo prefix (us., eu., apac., global.). These are global cross-region rates; regional inference runs roughly 10% higher.

Stop reading prices, start measuring them

A price list tells you the rate. It does not tell you what your agents actually spend, or which repo, PR, or developer spent it. BurnLens is a local proxy that does — pip install burnlens, one environment variable.

Read the docs · Scan existing agent logs · See the dashboard