Most AI agents being switched off this year are switched off for arithmetic, not accuracy. They work well enough. What they cannot do is show which number they moved, so when the invoice lands there is nothing to weigh it against, and the cheapest decision in the room is to turn them off.
The pullback is now measurable. KPMG's Global AI Pulse for Q2 2026 found that 49% of senior leaders had scaled back AI agent deployments because operating costs outweighed the benefits, and that only 26% of US leaders had full, real-time visibility into what AI costs to run at scale. Those are the same finding stated twice. Usage-based pricing turned AI from a project you finish into a meter you read, and most teams never fitted the meter.
We built an AI design assistant for Annex, which supplies custom modular buildings to fabricators across the US. The unit of value was fixed before we wrote anything: every design the assistant touched had to come out cheaper or fit more usable space. The published result was a "10% average cost reduction or space increase" — a figure a fabricator can hold against one specific quote and see whether the tool paid for itself.
A fixed unit of value is what makes running cost decidable. When the unit is a design, a quote or an enrolment, you can put the monthly bill beside the margin on the work it touched and get an answer in an afternoon. When the unit is "productivity", there is no denominator, and the first serious cost review ends the pilot. The agents dying this year are mostly the second kind, and the teams running them cannot argue back, because nobody ever agreed what the thing was supposed to move.
The part that cost us more than expected on Annex was not the model. It was the data underneath it: cataloguing thousands of construction materials with their specifications, regional availability and volume pricing, then digitising state building codes so the assistant could reason about them. That groundwork took longer than the AI implementation itself. It is also what kept the running cost sane, because an agent forced to rediscover the same unstructured mess on every request pays for that rediscovery every single time.
Human review belongs in the cost model, not in the reassurance section. The Annex assistant proposes and the fabricator decides, and we treated that review time as a line in the cost of every run rather than a safeguard bolted on afterwards. A workflow where someone has to check output they cannot verify quickly is where the economics quietly go bad. The model bill arrives itemised every month; the hour your team spends checking it never does.
If you are running delivery and weighing whether to renew or scale an agent, the useful test is not another demo. It is whether you can state today, without commissioning a project to find out, the unit of work the agent improves and what one unit costs to run. A demonstration is not evidence of either. If you cannot compute both numbers, that is the finding, and it is worth more than the next pilot.
None of this is an argument against putting AI into production. It is an argument for knowing the price of what you already switched on.