AI Economics Control Plane
BurnLens measures, explains and controls the cost of AI agents and LLM applications by repository, team, customer, feature and outcome. Start locally without sending prompts or code to BurnLens Cloud.
$6.84 per accepted PR on this repository, measured 2026-08-15. 104 merged · 2 closed unmerged · $710.85 of agent spend, with the failed attempts charged to the successes. A floor, not a ceiling — any model missing from the pricing tables counts as $ unknown. Cost is attributed per repo: agent logs record which repo a session ran in, not which branch. See the full method.
Your invoice says gpt-4o: $4,287. It doesn't say which feature, team, or customer burned it. By the time you trace the spike, it's already on next month's card.
A bad deploy, a runaway agent, or one abusive customer can trigger thousands of API calls before any dashboard turns red. You find out when you open the bill.
OpenAI's usage page. Anthropic's console. Azure Cost Management. Bedrock CloudWatch. No unified view, no way to ask which feature is your biggest AI spend across all providers.
burnlens scan reads local session logs from Claude Code, Cursor, Codex, and Gemini CLI. Cost per repo, per model — instantly. If gh is installed, it also derives cost per merged PR from your GitHub history.
BurnLens runs a local proxy on :8420. Set OPENAI_BASE_URL or ANTHROPIC_BASE_URL and your existing SDK code routes through automatically. Designed for low overhead with full streaming passthrough.
Three request headers — X-BurnLens-Tag-Feature, -Team, -Customer — attribute cost to any dimension you care about. Tags are stripped before the request reaches the AI provider. If you enable cloud sync, tag values are uploaded to your workspace alongside cost metadata — never prompt or response bodies.
Register an API key with a daily dollar limit. At 100%, BurnLens returns 429 before the call is forwarded upstream. 50% and 80% thresholds fire Slack or email alerts. This is a recorded-spend guardrail, not an absolute concurrent ceiling;burnlens controls shows the active guarantees.
Budget enforcement semantics — exact behaviour under concurrency, streaming, retries, and unpriced models, with the implementing function cited for every claim.
OpenAI, Anthropic, Google, Groq, Mistral, Together, xAI, DeepSeek, Azure OpenAI, and AWS Bedrock spend in one view. Model breakdowns, waste detection, and budget tracking using versioned provider pricing.
Claude Code, Cursor, Codex, Gemini CLI — cost per repo and developer from local logs. Merged PRs are accepted outcomes when ghis present, at repo grain, not per branch. Hard daily caps per API key stop one runaway agent from burning the team's monthly budget overnight.
Tag each request with a customer ID. See which customers drive the most cost. Alert on per-customer budget thresholds and configure cheaper-model routing.
Tag retrieval calls, tool calls, and generation separately. See whether your vector search or synthesis step is the cost driver — and whether it justifies the output quality.
Set per-team monthly budgets, get Slack alerts at 80% and 100%, and export monthly records for comparison with provider invoices.
Record accepted and rejected outcomes for any workflow. BurnLens divides total spend by accepted outcomes — failed attempts are charged to the successes, because that's what one working result actually costs.
For coding agents, merged PRs are outcomes automatically — burnlens scan derives them from GitHub via gh. No instrumentation required.
BurnLens caches repeated prompts using exact match and cosine-similarity embedding search. Identical or near-identical queries are served from cache — the upstream API call never happens and you pay nothing.
Off by default. Observation never caches. Enable with cache.enabled: true in burnlens.yaml.
burnlens recommend analyses your usage and suggests cheaper model alternatives with confidence scores and projected savings. Switch from gpt-5.6-sol to gpt-5.6-terra where output tokens stay short.
Sliding-window statistical analysis (MAD / Z-score) across 1-minute, 5-minute, 15-minute, and 1-hour windows. Cost spikes and runaway agent loops fire alerts before the damage compounds.
Opt in with routing.budget_downgrade: true. Then, when a budget you set drops below 20% remaining (or $5), BurnLens routes requests to a cheaper model instead of hard-blocking with 429. Your app stays live — it just gets more cost-efficient.
Off by default. A budget alone never changes the model. Every downgrade is logged; omit the flag or set routing.budget_downgrade: false to always receive the model you asked for.
| Capability | BurnLens approach |
|---|---|
| Coding-agent observation | Local log scanning — Claude Code, Cursor, Codex, Gemini CLI |
| Prompt handling | Prompt bodies are never uploaded to BurnLens Cloud |
| Economics | Spend + accepted outcomes (merged PR where GitHub data exists) |
| Missing pricing | Explicit $ unknown, never a silent $0 |
| Runtime enforcement | Optional local proxy; recorded-spend guardrail returns 429 before upstream |
| Budget model changes | Explicit opt-in. Cache and routing stay off by default |
| Savings | Projected and verified reported separately |
| Self-hosting | Open-source deployment available |
Dedicated comparisons come later and are sourced. Until then we describe BurnLens, not other products. Methodology and evidence →
burnlens scan then burnlens repos. Cost per repository, confidence, and accepted outcomes stay on disk. Prompts never leave.
Sync cost metadata (never prompt bodies) to a workspace. Recurring review across developers. Self-service: start 7-day free trial, card required, $29/month.
Owner plus engineering share history, permissions, and budgets. Agencies: see project economics by client.
Free forever for individual developers. Pay only when your team needs cloud sync.
Card required. Cancel anytime before day 7 — no charge.
Production APIs: burnlens start, then point OPENAI_BASE_URL at http://127.0.0.1:8420/proxy/openai. Google needs burnlens.patch.patch_google().