A finance director approving an artificial intelligence programme is usually approving a number that can be defended in a committee: licences, an integrator’s fee, some training days, a contingency line of maybe fifteen per cent. Clean, bounded, capitalisable.
What turns up eighteen months later does not behave like that at all. It arrives monthly, and it moves with how many people used the thing and how verbose they were. It also keeps arriving long after the project board has been stood down. The mismatch between the shape of the approval and the shape of the spend is the single biggest reason AI programmes overrun, and it runs through John Margerison’s arguments about what enterprises actually take on when they put AI into live operations, written by the CEO of XFactorAi.
The overlooked costs are not exotic. They are ordinary operating costs that no one assigned to a cost centre, because the business case was written as though the work ended at deployment.
The unit of spend is a token, and nobody owns the meter
Consumption pricing changes who controls the bill. When a seat licence costs a fixed sum per user per year, procurement controls spend by controlling headcount. When the charge is per token, spend is controlled by whoever is writing prompts at four in the afternoon, and that person has no visibility of the rate card.
Gartner put a striking figure on where this leads in a press release published in June 2026, predicting that AI coding costs will overtake the average developer’s salary by 2028 as token consumption grows and vendors move to consumption-based licensing. Read that as a governance problem rather than a pricing one. Engineering teams optimise for the thing they are measured on, which is shipping, and token spend is invisible to them unless someone builds the dashboard and hands them the ceiling.
Most organisations have no equivalent of a phone bill review for AI usage. They will need one, and the person who runs it is a cost.
The pilots that never die become a second budget
The much-quoted MIT finding is usually read as a story about wasted investment. It deserves a harder reading. As reported by The Hill, MIT’s research found that despite thirty to forty billion dollars of enterprise investment in generative AI, ninety-five per cent of organisations were seeing zero return.
A pilot with no return does not necessarily stop costing money. It leaves behind a service account, an API key still billing, a copy of customer data sitting in a vector store that now falls inside the retention policy, and a half-finished integration that the next upgrade of the underlying system will break. Decommissioning is skilled work and nobody’s promotion depends on it, so the residue accumulates.
Any organisation that has run more than three or four proofs of concept should price the clean-up before it funds the fifth. The cost of proliferation is real even when the experiments were cheap.
The price per unit is drifting upward with the power bill
Software buyers have spent two decades assuming that the unit cost of compute falls every year. That assumption is doing quiet damage to five-year AI models.
The input side has stopped cooperating. CNBC reported in February 2026 that Goldman Sachs analysts expect household electricity prices to rise a further six per cent through 2027, with data centres accounting for around forty per cent of growth in electricity demand. Industrial and commercial buyers are exposed to the same grid.
Inference efficiency is improving, and that pushes the other way. The point for a budget holder is that the direction of travel is contested, and a multi-year plan built on assumed deflation has taken a position on an energy market it has no view of. Model the renewal at a higher rate than today’s.
Every model upgrade is a re-qualification project
This is the cost that almost never appears in a business case. When a vendor deprecates a model version, the replacement is not a drop-in. Outputs shift, tone shifts, edge cases that were tested and signed off behave differently, and anything downstream that parsed a predictable response format may quietly stop working.
For a regulated process, that means running the validation again: the test set, the reviewer sign-off, the documentation. Organisations that built evaluation harnesses early absorb this in days. Those relying on a human spot-checking a spreadsheet absorb it in weeks, several times a year, indefinitely.
Deprecation schedules are set by the vendor. That makes re-qualification a recurring cost on someone else’s calendar, which is the worst kind to leave unbudgeted.
Proving the output stands up is a permanent operating line
The comforting version of the story says human review is scaffolding that comes down once the model is trusted. Three years of enterprise deployment suggest otherwise. Where a decision affects a customer’s money, health or employment, someone competent looks at it, and that person needs to be senior enough for their judgement to mean something.
Add the audit trail. Retaining prompts and outputs so a decision can be reconstructed months later is storage, retention policy and disclosure exposure at once. It is cheap per record and substantial per year.
Budget it as an operating cost of the same order as the licence, and the business case either survives that test or was never sound.
Tighter approval will not fix any of this. What helps is a different question asked at approval: ask what the steady-state monthly run rate looks like in year three, with usage at three times pilot volume, a model migration behind you and a review function staffed. Most organisations cannot answer that, which is precisely why the answer matters. Firms that can will be the ones still expanding their AI estate when everyone else is quietly explaining a variance.
