AI Costs

Tokenpocalypse: why the AI bill is rising even though prices are falling

August 14, 2026

The word of the summer comes from 404 Media: the “Tokenpocalypse”. A leaked internal Accenture meeting revealed that the firm is fighting soaring token spend. A large share of it arises when non-technical staff hand token-intensive tasks such as “turn this PDF into slides” to frontier models. Uber has capped the use of Claude Code and Cursor after its own CTO admitted that the entire annual AI budget was exhausted after four months — shortly after the company had encouraged staff to use the tools as much as possible.

Why consumption is exploding now

Because the nature of usage has changed. A chat prompt was one question and one answer. An agent is a loop: it plans, calls tools, reads the results, corrects itself and tries again. And at every single step it sends the entire history so far back to the model. One task therefore becomes dozens of model calls with an ever-longer run-up. Multi-agent setups, where an orchestrator dispatches several sub-agents in parallel, multiply this again.

A task completed by an agent quickly consumes ten to a hundred times what a simple chat prompt does. That is not a bug — it is the price of AI working independently rather than merely answering. But it explains why budgets calculated for chatbot usage are empty after four months in the age of agents.

Diagram: a chatbot is one model call, an agent with sub-agents is dozens
Image: own graphic — one agent multiplies the model calls. Figures illustrative.

The paradox

Prices per token are falling rapidly at the same time. OpenAI cut GPT-5.6 Luna by 80% in August to USD 0.20 per million input tokens; Google offers Gemini 3.7 Flash at half its predecessor’s introductory price. Cheaper per unit, more expensive in total: this is the Jevons effect that economists know from the steam engine. Falling costs per use lead to exploding use.

Our take

The problem is routing, not usage. Anyone converting PDFs into slides with a frontier model is paying for a sports car to fetch bread rolls. The answer is model routing by task: cheap workhorse models for everyday work and for the many intermediate steps agents take, frontier models for what justifies them. That is exactly what harness and orchestration are for — our common thread from July.

Waste has to be prevented, not usage. Uber’s swing from “use everything” to a hard limit repeats the mistake we know from shadow AI: blanket brakes push usage into uncontrolled channels. Better: budgets per use case, transparency dashboards and cheap default paths for routine tasks.

Token economics now belongs in controlling. When AI costs become consumption-based, they need the same discipline cloud costs needed ten years ago: forecasting, showback, price negotiations with volume commitments. If you are negotiating contracts in 2027, price in falling list prices explicitly.

The question is no longer what a token costs, but who in the organisation decides what tokens are spent on.

Want to know where your token spend actually goes? We are happy to look at it with you.

Lucas Roesler

Your contact

Lucas Roesler

Head of Engineering