AI Insight

Knowledge work just acquired a unit cost. Most organisations aren’t ready.

Why AI bills are rising as token prices collapse, and what leaders should measure instead of usage

For about two years, the fastest way to look forward-thinking inside a large organisation was to use a lot of AI.

That is not a figure of speech. Meta made AI-driven impact a formal expectation in every employee’s performance review from 2026. Accenture began tracking weekly logins to its internal AI platforms and told managers that progression depended on regular adoption. At several technology firms, engineers were reported to be competing on internal leaderboards for AI consumption, a practice that acquired the only name it could plausibly have been given: tokenmaxxing. Heavy usage read as capability. Light usage read as resistance.

Then the invoices arrived.

In April 2026, Uber’s chief technology officer confirmed that the company had exhausted its entire annual AI budget four months into the year, after an agentic coding tool spread across roughly 5,000 engineers faster than any finance model had anticipated. Monthly costs ran between $500 and $2,000 per engineer. What makes the episode instructive is not the overspend. It is what Uber’s president and COO, Andrew Macdonald, said when asked whether all those tokens had produced better products: “That link is not there yet.”

Uber is not a badly run company. It is technically sophisticated, financially disciplined, and unusually good at operating things that are metered. If Uber could not connect its AI consumption to its output, the problem is unlikely to be Uber.

The problem is that AI has done something to knowledge work that few leadership teams have named. It has given it a variable unit cost. Almost no organisation has a management system built for that.

The proxy that outlived its purpose

Measuring usage was not stupid. It was correct, briefly, for a specific reason. In 2023 and 2024 the dominant risk was under-adoption. Nobody knew what these tools were good for, so the sensible strategy was to get them into as many hands as possible and watch. Usage was the only observable signal available. You cannot measure the value of experiments you have not run, so leaders measured the running instead.

The trouble is what happens to a proxy when the underlying risk changes. Goodhart’s law is usually quoted as a warning about gaming: when a measure becomes a target, it stops being a good measure. The version that matters more to executives is quieter. Proxies do not announce their own expiry. They keep producing tidy numbers long after the question they answered has been replaced.

By late 2025 the risk had flipped. McKinsey’s global survey found 88% of organisations regularly using AI in at least one business function. Adoption was no longer the constraint; discrimination was, meaning the ability to tell which uses were worth what they cost. The dashboard did not change, because nothing forces a dashboard to change. So organisations spent a year rewarding the precise behaviour that was about to become their largest uncontrolled expense.

This will happen again. Every organisation running an AI programme today has at least one metric installed for a risk that has since evaporated. The fix is unglamorous and takes ten minutes: when you adopt a proxy metric, write down the condition that will retire it. Almost nobody does this. It is the cheapest governance intervention available.

Cheaper units, bigger bills

The arithmetic that has caught out so many finance teams is not complicated.

The price of AI has collapsed. Industry analyses of enterprise API traffic put the blended cost of inference at roughly $6 per million tokens in early 2026, down from around $18 a year earlier. Measured against capability rather than volume, the fall is steeper still.

Yet the FinOps Foundation’s State of FinOps 2026, drawn from 1,192 practitioners responsible for over $83 billion of annual cloud spend, found 73% of organisations exceeding their original AI cost projections. Two years ago, 31% of FinOps teams managed any AI spend at all. In 2026 it is 98%, and AI cost management is the discipline’s most requested skill.

Total spend is price multiplied by volume. The first term is falling fast. The second is growing faster.

The mechanism is architectural rather than behavioural, which is why it surprised people. A chatbot exchange is one call. An agentic workflow is a loop: the model plans, calls a tool, reads the result, revises, retries, and carries the accumulated context forward on every turn. History, tool definitions and retrieved documents are resent as input each time. Goldman Sachs analysts described it in May 2026 as taking a simple request and multiplying it many times over. Most agentic business cases were built on chatbot-era volume assumptions; production numbers came in an order of magnitude higher. Compounding this is the habit of sending every task to the most capable model regardless of need, across a price spread now running to three or four orders of magnitude.

The lesson energy managers learned a century ago

Anyone who has worked on decarbonisation will recognise this pattern, and will know it already has a name.

In 1865, the economist William Stanley Jevons observed that more efficient coal-fired engines had not reduced Britain’s coal consumption. They had increased it, because cheaper steam made viable a whole class of uses not previously worth the fuel. Energy professionals have watched the same effect since in lighting, insulation and vehicle fuel economy. It is now happening to cognition.

The International Energy Agency’s 2026 analysis makes the parallel almost uncomfortably exact. Measured per individual task, the IEA reports, the energy efficiency of AI is improving at a rate unprecedented in energy history. In the same report, it projects that total data centre electricity consumption will roughly double by 2030, with AI-focused capacity tripling. Efficiency per task improving faster than anything the energy sector has ever recorded, and total consumption doubling anyway. Both statements are true, and they are true for the same reason your AI bill is rising.

 

The valuable part of this parallel is not the diagnosis. It is the distinction energy practitioners were forced to develop once they accepted that efficiency alone would never deliver absolute reductions.

Efficiency asks how to get the same output from less input. Sufficiency asks whether the output was needed at all.

Almost every AI cost-control programme now being stood up is an efficiency programme. Route to cheaper models. Cache aggressively. Trim context windows. Cap spend per seat. All sensible, and all answers to “how do we do this more cheaply?” Very few organisations are asking the sufficiency question: which of these workflows should we stop running altogether?

That conversation is harder, because it requires admitting that work now done automatically was not worth doing manually either. A great deal of AI spend goes into artefacts nobody asked for before the tool made them cheap. Longer reports. More summaries. Meeting notes nobody reads. The cost has become visible; the fact that the activity was always optional has not.

Efficiency will take twenty or thirty per cent off the bill. Sufficiency is where the larger number lives.

Why the return looks like it isn’t there

Leaders keep saying there is no ROI, and the evidence appears to back them. McKinsey found that only 39% of organisations attribute any EBIT impact at all to AI, most of those below 5%. PwC’s 29th Global CEO Survey, published in January 2026, found 56% of chief executives reporting that AI had so far delivered neither revenue growth nor cost savings.

Something is missing. It is not value. Two different things are being confused.

The first is a measurement asymmetry. The cost of AI arrives as a single monthly invoice, precise to the cent, attributable to a named budget line. The benefits arrive as twenty minutes saved here, a better first draft there, a research task that took an afternoon instead of two days, spread thinly across hundreds of people and attributed to nobody. Compare a precise number with a diffuse one and the precise number wins. It looks like evidence; the other looks like an anecdote.

Sustainability leaders have fought this exact asymmetry for twenty years. Efficiency retrofits, waste reduction, supply chain resilience: the cost is a capital line, the benefit is spread across operations and lands in nobody’s numbers. The organisations that eventually made the business case work did not find better arguments. They changed the accounting so the benefit had somewhere to land.

The second is the harvest problem, and it matters more. Time saved is not money saved. It becomes money only if something is done with it.

Suppose AI saves each of two hundred knowledge workers thirty minutes a day. That is a hundred hours of daily capacity, and it will appear nowhere in your accounts unless it is converted into more output, fewer hours worked, or work that was being deferred. Left alone it dissipates into slack. Meetings expand. Drafts get another revision. The work fills the time, as it always has.

Which is why the strongest finding in McKinsey’s data is also the least discussed. The variable most closely correlated with EBIT impact is not spend, tool choice or model quality. It is fundamental workflow redesign. The 6% classed as high performers are around three times more likely to have rebuilt processes end to end rather than layering AI on top of existing ones.

 

 

Erik Brynjolfsson, Daniel Rock and Chad Syverson formalised this as the productivity J-curve. General purpose technologies demand large complementary investments in redesign, retraining and reorganisation, and those are expensed rather than capitalised, so measured productivity dips before it climbs. Electrification took decades to reach the statistics, not because electric motors disappointed, but because factories had to be rebuilt around them before the gains existed.

Most organisations bought the input and skipped the complementary investment. Then they measured the result and concluded the input did not work.

One observation from facilitated sessions sticks with me. Ask a room which tasks have got faster since they started using AI and everyone speaks at once. Ask the same room to name a single decision AI has changed, and it goes quiet. That gap is the ROI problem in miniature. Faster tasks do not move a P&L. Better decisions do.

Meter, ratio, harvest

If usage was the wrong metric, what replaces it? Three questions, in strict order. Taking them out of order is where most programmes fail.

The meter. Can you attribute spend to a specific workflow, team and purpose? Most cannot, which is why granular token monitoring tops the capability wish list in the FinOps Foundation’s 2026 survey. Instrument before you restrict. A cap imposed without attribution does not cut waste; it cuts waste and exploration in equal measure, and exploration is the part you cannot rebuild quickly.

The ratio. The unit of account is the completed task, never the token. What does one finished, accepted piece of work cost across a hundred real production runs rather than a demonstration? And measured against what? Not last year’s AI bill. Against the fully loaded cost of the alternative: the human hours, the agency fee, the cycle time, the thing you were not doing at all.

Then adjust for quality, using the metric almost nobody tracks and everybody could measure by Friday: the rework rate. What proportion of AI output does a human materially rewrite before it ships? That single number captures quality, trust and hidden labour cost at once. An output costing four dollars that ships is cheaper than one costing forty cents that gets rebuilt twice. Your people already know the answer. Nobody has asked them.

The harvest. For every workflow where AI has released capacity, someone must be accountable for where that capacity went. Named, with a number. Did the team take on work it had been deferring? Did cycle time fall in a way a customer noticed? Did the function absorb growth without adding headcount? If the honest answer is “people are a bit less busy”, there is no return. There is only a cost.

The harvest is the question almost universally missing, and the only one of the three that produces a number a CFO recognises.

Two ways this goes wrong

The flat-rate trap. The predictable market response to token anxiety is repackaging. Vendors will push bundled, per-seat, all-you-can-eat pricing, and finance teams will find it a relief. Deloitte notes that packaged AI solutions already abstract tokens away entirely, leaving buyers a comfortable subscription fee and no visibility into consumption efficiency. Flat rates smooth your forecast and destroy your unit economics data in the same stroke. Accept one if you must, but insist on metering anyway.

Over-correction. Gartner forecasts that more than 40% of agentic AI projects will be cancelled by 2027. Some of those cancellations will be sound judgement. Many will be organisations killing exploration because exploration cannot yet justify itself, which is a category error, since that is what exploration means. Uber’s own response is a better template: rather than a ban, per-tool monthly caps, a usage dashboard for every employee, and a route to request more. And remember that the J-curve is real. Demanding a return at month nine may mean measuring the trough and mistaking it for the destination. The discipline is overdue; the verdict can wait.

What to do differently

  1. Give every adoption metric an expiry condition. Write down now what would make “active users” the wrong thing to track, then diarise the review.
  2. Instrument before you cap. Tag every AI call by team, workflow and purpose, before imposing limits rather than after.
  3. Change the unit of account from cost per token to cost per completed, accepted task, benchmarked against the fully loaded cost of the alternative.
  4. Start measuring rework rate this quarter. The cheapest quality signal available, sitting unused in your teams’ heads.
  5. Run a sufficiency review, not only an efficiency review. Take the ten highest-spending workflows and ask what happens if you stop each one entirely.
  6. Name a harvest owner for every workflow where AI has freed capacity, accountable for converting it into output, cycle time or absorbed growth.
  7. Split the budget in two. Production, held to unit economics. Exploration, capped and deliberately unmeasured.

The larger point

The AI cost story is usually told as a cautionary tale about enthusiasm. It is more interesting than that.

For thirty years the marginal cost of one more unit of knowledge work was effectively zero. An analyst produced another report, a lawyer another memo, and the cost was already sunk in the salary. Knowledge work was managed through headcount, because headcount was the only lever that moved the number. Manufacturing, meanwhile, spent a century building a discipline around cost of goods sold: unit economics, yield, rework, scrap rates, capacity utilisation.

Metered AI has quietly dragged knowledge work across that line. It now has a marginal cost, a yield and a rework rate. The vocabulary that fits was developed on factory floors, not in professional service firms, and most executives running knowledge organisations have never had to use it.

That is the transition actually underway, and it explains the disorientation. These organisations are not failing to find the ROI. They are being asked, for the first time, to run cognitive work as a production system while still managing it as a payroll.

The ones that come through this well will not be the ones that cut hardest. They will be the ones that learned to measure the right unit, then had the discipline to ask what the output was for.