Why most organisations can no longer tell whether their AI is working
In the first half of 2025, a research group called METR ran an experiment that ought to have changed a great many board conversations, and mostly did not. Sixteen highly experienced open-source developers were given 246 real tasks in codebases they had worked in for years. Each task was randomly assigned to be completed either with AI assistance or without it.
Before starting, the developers expected AI to make them roughly 24% faster. Machine-learning experts asked to forecast the outcome predicted a speed-up of nearly 40%.
The measured result was that the AI-assisted tasks took 19% longer.
The more uncomfortable finding came afterwards. Having lived through the slowdown, the same developers still estimated that AI had made them about 20% faster. They were not exaggerating for the researchers. They simply could not distinguish between work that felt quick and work that finished quickly.

Figure 1. The gap between belief and measurement was larger than the effect being measured.
METR is careful about this study, and so should we be. Sixteen developers is a small sample, the tools were those available between February and June 2025, and the organisation now treats the finding as historical. A 2026 follow-up found some evidence of speed-up but was complicated by selection effects: developers who most valued AI were reluctant to enrol in a study that might take it away from them.
Take all of those caveats seriously. Then notice what they do not touch. Whatever the true effect was, these professionals could not detect it in themselves. That is the finding that should concern executives, because almost every AI business case now circulating in industry rests on precisely the instrument this study found to be unreliable: people’s own sense of how much faster they have become.
Adoption is near-universal. Evidence is not.
McKinsey’s late-2025 global survey found that 88% of organisations were using AI in at least one business function and 72% were using generative AI, up from 33% the year before. Yet nearly two-thirds had not begun scaling AI across the enterprise, only 39% attributed any EBIT impact at all to AI, and roughly 6% qualified as high performers attributing more than 5% of EBIT to it.

Figure 2. Adoption has become ordinary. Demonstrable impact has not.
The usual reading of this gap is an implementation-failure story: wrong use cases, poor data foundations, weak change management. All plausible, all partly true. But there is a simpler reading that almost nobody adopts, and it follows directly from METR.
Many of these organisations are not failing to create value. They are failing to know whether they created value.
Consider what that 39% figure actually is. It is a self-report, from executives, about a counterfactual almost none of them constructed. Nobody ran the business for a year without AI as a control. If a chief financial officer were handed a capital programme justified by a survey of how the users felt about it, the conversation would end quickly. AI is the one significant category of spend where that standard has been quietly accepted.
The strongest counterargument, and why it does not help
There is a serious economic case against reading any of this as failure. Erik Brynjolfsson and colleagues have documented what they call the productivity J-curve: when a genuinely general-purpose technology arrives, measured productivity tends to fall before it rises. Firms are pouring money into intangible capital, which is to say new processes, new skills, restructured work, and none of that appears as output while it is being built. Electricity took decades to show up in factory statistics, largely because factories had to be physically redesigned around it first.
On that reading, disappointing returns in 2026 are exactly what a real transformation looks like from the inside. I find this persuasive. But it cuts in the opposite direction from how it is usually deployed in strategy meetings.
If the J-curve is the right frame, the returns do not come from the tool. They come from the organisational reinvention that the tool makes worthwhile. “We have given everyone a licence and the benefits will follow” is not a J-curve argument. It is the absence of one. The J-curve says the value sits in the work you have not done yet.
A better mental model: from production to verification
Researchers at Microsoft Research and Carnegie Mellon University surveyed 319 knowledge workers about 936 real instances of using generative AI at work. Their central observation was not that AI reduced thinking, but that it relocated it: from gathering information to verifying it, from solving problems to integrating an AI’s answer, from doing the task to supervising it.
They also found a relationship worth reading twice. The more confidence someone had in the AI, the less critical thinking they applied. The more confidence they had in their own expertise, the more they applied. Expertise here is not only a knowledge asset. It is a form of permission to disagree.
Put those findings together and a usable model emerges. AI does not remove work. It moves work from production to verification.
Producing a first draft, a forecast, a summary, a competitor comparison used to be the expensive step. It is now close to free. Checking has become the expensive step, and checking has three awkward properties: it is invisible in most workflows, it is nobody’s job title, and it is the easiest thing in the world to skip when the output already looks finished.
Which leads to the question I would put to any AI deployment before it scales: who holds the verification cost, and do they have the time and the standing to pay it?
When nobody has been assigned that cost, it does not disappear. It travels downstream. Researchers at BetterUp Labs and Stanford’s Social Media Lab gave this a name in the Harvard Business Review: workslop, meaning AI-generated content that looks like completed work but does not advance the task. In their survey of 1,150 US desk workers, 41% had received some in the previous month, each instance taking almost two hours to resolve, which they costed at around $186 per affected worker per month.
The financial number is the least interesting part, and the study is a company-sponsored survey rather than peer-reviewed research, so treat it as indicative. The social finding is harder to wave away. Of those who received workslop, 42% trusted the sender less afterwards and around a third were less willing to work with them again. Unmanaged AI adoption does not only cost hours. It spends trust, which is the one resource an organisation cannot afford to run down while it is changing.

Figure 3. The same technology produces opposite results depending on whether verification is designed in or left to find its own owner.
The 1983 paper that explains 2026
In 1983, the psychologist Lisanne Bainbridge published a short paper called “Ironies of Automation”. Her observation was this: automating a process removes the operator’s routine practice while leaving them responsible for the exceptions. Over time, the operator becomes less capable of handling precisely the situations that only a human can handle. The more automated the system, the more it depends on a skill it is quietly eroding.
For four decades this was treated as an aviation and process-control problem. It is now a knowledge-work problem, and very few organisations have noticed.
If your analysts stop writing the first draft of a market assessment, then in three years they will be measurably worse at judging whether a market assessment is any good. If your managers stop drafting difficult messages, they will be worse at recognising when a drafted one will land badly. This is not an argument against the tools. It is an argument for deciding, deliberately and in advance, which forms of practice you intend to preserve.
The practical version is a single question, asked once per automated process: what judgement was this task exercising, and where will that judgement now be developed instead? In my experience running sessions with leadership teams, this question produces longer silences than any question about risk or compliance.
Regulation has just stopped doing the work for you
For most of the past two years, 2 August 2026 stood as the EU AI Act’s enforcement cliff: the date on which the high-risk regime would apply to systems used in employment, education, credit assessment and access to essential services. Compliance programmes across Europe were built around it.
That date has moved. The Digital Omnibus on AI, proposed by the European Commission in November 2025, endorsed by Parliament on 16 June 2026 and approved by the Council on 29 June, defers stand-alone high-risk obligations to 2 December 2027, and to 2 August 2028 for AI embedded in regulated products. Most transparency obligations still apply from 2 August 2026, and a new prohibition on AI-generated non-consensual intimate imagery arrives in December.
One further change has received far less attention than it deserves. Article 4, the AI literacy duty that has applied since February 2025, was softened. Providers and deployers are no longer required to ensure a sufficient level of AI literacy among their staff, only to support its development. In legal terms, an obligation of result became an obligation of effort.
Read as a compliance matter, that is a relief: sixteen additional months and a lighter standard. Read as a management matter, it is close to the opposite. The external forcing function for building genuine understanding weakened at exactly the moment internal adoption peaked. Nothing that made AI literacy urgent last year has become less true. Only the deadline changed.
Anyone who worked through the equivalent sustainability reporting delays will recognise the pattern. When those deadlines slipped, the organisations that had built a compliance project quietly stood it down and lost the capability. The organisations that had built actual carbon literacy kept it, and are now the ones able to answer customer, investor and procurement questions that never went away. The deadline was never the point. It was only the prompt.
There is at least suggestive evidence that this is not merely virtuous. McKinsey’s 2026 research on AI trust found that organisations investing most heavily in responsible AI were considerably more likely to report material EBIT impact. That is a correlation, and the heaviest investors also tend to be the largest and most operationally mature firms. But it is evidence against the assumption that governance functions primarily as a brake.
Getting the energy argument right
Environmental cost belongs in this conversation, and it is usually argued badly in both directions.
The International Energy Agency estimates that data centres consumed around 415 TWh of electricity in 2024, roughly 1.5% of the global total, and projects this will more than double to around 945 TWh by 2030, a little over Japan’s entire current consumption, with AI the single largest driver. Stated on its own, that sounds like an emergency.
Now add the context that rarely travels with it. That increase represents about 8% of the total growth in global electricity demand expected by 2030. Electric vehicles account for more. Air conditioning accounts for more. Industry accounts for far more.
So the defensible position is neither that AI is an environmental catastrophe, a claim that will not survive contact with an informed chief financial officer, nor that 3% is nothing. It is more specific and considerably more useful: in aggregate the load is manageable, but its concentration is not. Around 80% of the projected growth lands in the United States and China, on particular grids, in particular water catchments, on particular local planning timetables. Data centres are also one of a small number of sectors where emissions are expected to rise while most others fall.
For a leadership team, that converts a moral argument into an operational question: where does our inference actually run, on whose grid, and does anything in our reported footprint include it? Very few sustainability reports currently answer this. That is a governance gap rather than an ethical failing, and governance gaps are fixable.
What to do differently
None of this argues for caution as a posture. It argues for instrumentation. Six things are worth doing in the next quarter.
- Stop accepting self-reported productivity. Choose two or three live deployments and measure a real before-and-after with a comparison group. Cycle time, error rate, rework volume, escalation rate. Where a counterfactual genuinely cannot be built, say so out loud and label the business case a hypothesis rather than a result.
- Name a verification owner for every deployment. In writing, with time budgeted, visible in their objectives. If the honest answer to “who checks this?” is nobody in particular, you have not saved effort. You have exported it to colleagues who did not agree to receive it.
- Decide which practice you are keeping. For each process you automate, name the judgement it was exercising and where that judgement will now be trained. Bainbridge’s irony is avoidable, but only on purpose.
- Measure rework, not just output. Add one line to project reviews: how much time went into correcting AI-assisted work produced elsewhere in the organisation? It is the cheapest available instrument for detecting workslop before it hardens into culture.
- Treat the extra sixteen months as capability time, not slack. Re-baseline high-risk work to December 2027, then spend the interval on system inventory, human-oversight design and literacy. Article 4 still applies from 2 August 2026 even in its softened form, and a documented literacy gap is treated as an aggravating factor when regulators investigate anything else.
- Build literacy through practice, not policy. This is where facilitation earns its place. Teams cannot govern what they cannot picture. In a Climate Fresk, the shift happens not when people are told how the carbon cycle works but when they build the causal map themselves and argue about where the arrows belong. The same principle applies here. A two-hour session in which a team blind-scores AI-assisted outputs against unassisted ones will do more for AI literacy than any policy document, because it recalibrates the instrument the METR study found to be broken: their own judgement.
The capability that actually matters
An earlier version of this argument would have concluded that leaders should slow down. That is not quite right, and it is not what the evidence supports. Speed is not the problem.
Undetectable speed is the problem.
The organisations that come out of this decade well will not be the fastest adopters or the most cautious ones. They will be the ones that can tell the difference: the ones that built an honest instrument for knowing whether the thing they just deployed actually worked, and kept enough human judgement in practice to read it. It is an unglamorous capability. Judging by the evidence, it is also a rare one.