The most cited statistic in AI in 2026 is now MIT NANDA's finding that 95% of organisations that invested in generative AI saw zero measurable return. Methodologists have argued with it. Reasonable people have pushed back. The figure is overstated, or it is roughly right, depending on whom you read. Either way, it has done what statistics do when they touch a nerve: it has forced a conversation that should have happened earlier. The conversation is not really about return. It is about measurement, and increasingly, about whether the AI consulting behind a project actually earns its keep.
Most AI initiatives are not failing because the models do not work. They are failing because nobody set up the measurement framework that would tell anyone whether they were working. The default of every AI dashboard (daily active users, requests served, demos given) measures activity. None of it tells you whether the feature is producing value. The teams getting actual returns out of AI in 2026 are the ones who decided early what they would measure, measured it honestly, and were willing to keep or kill features based on the answer.
This post is about how to do that, and it is exactly the discipline a serious AI consulting engagement should bring to the table before any spend is signed off.
The headline number, and what it actually says
The MIT NANDA State of AI in Business 2025 report, published in July of that year, was based on 150 interviews with leaders, a 350-employee survey, and analysis of 300 public AI deployments. Its central finding was that despite $30-40 billion in enterprise spending on generative AI, only about 5% of AI pilot programmes were achieving rapid revenue acceleration, while the vast majority were stalling with little to no measurable impact on P&L.
This finding has been criticised. Some of the criticism is fair: the sample is small relative to the population of enterprise AI projects, the definitions of "measurable" and "P&L impact" do real work in the result, and the figure does not capture diffuse productivity gains that are genuine but hard to attribute. UC Berkeley's executive education group made a substantive case that ROI is the wrong primary metric for AI at this stage of adoption, in the same way it was wrong for the early internet.
The criticisms are useful. They do not refute the underlying observation. McKinsey's State of AI 2025, working from a much larger sample, has consistently shown the same shape from another angle: a small number of organisations capture most of the value, and they look operationally different from the rest. The McKinsey work found that high performers were nearly three times more likely than others to have fundamentally redesigned workflows around AI, and that the practices most strongly associated with value capture were unglamorous: strong technology and data infrastructure, clear measurement, and ownership of outcomes.
A more recent set of data from DX, who instrument AI productivity at engineering organisations including Booking.com, Workhuman, and Mercari, puts a useful corrective on the productivity claims specifically. The realistic throughput gain from AI coding tools across their measured population is 5-15%, meaningful, but a long way from the 3x and 10x figures that vendor marketing has trained leaders to expect. When leaders see modest results, they assume something is wrong. The data says: calibrate against validated peer-company signals, not against hype.
Put these threads together and a coherent picture emerges. The technology can produce real value. Most organisations are not capturing it. The reasons are not technological. They are about measurement and operating discipline.
The lying-to-yourself failure modes
If you read the failed AI deployments closely, the measurement mistakes cluster into a small number of categories. Most teams have made at least three of them.
Measuring activity, not outcomes. "We had 12,000 sessions with the AI assistant this quarter" tells you nothing about whether the assistant was useful. Useful is whether someone got their question answered, or their ticket avoided, or their decision made better. Activity metrics are easy to produce and uninformative. Outcome metrics are harder to define and the only ones that matter.
Measuring deflection without measuring resolution. This one is specific to support and service use cases, and it is the most common honest mistake in the published case studies. Zendesk and others have made the point clearly: AI deflection rate counts interactions handled without human involvement even if the problem was not resolved, while AI resolution rate counts only the issues fully solved. A high deflection rate with a low resolution rate is a system that is making customers angrier, on a metric that looks like success.
Confusing pilot performance with production performance. The 5% of organisations that ship valuable AI tend to look very different in pilot than in production. The 95% that fail typically had a working pilot. The thing that breaks between pilot and production is the long tail: edge cases, real-user phrasing, data the system was never tested against. Pilot metrics that do not reflect this gap will systematically overstate the production result.
Comparing to an imaginary baseline. "Our AI feature handles 60% of incoming queries" is not informative unless you know what the previous system handled. If your existing FAQ deflected 50%, the AI is worth a 10-point bump and a probably-larger ops bill. If your previous system deflected 0%, the AI is transformative. Without a real baseline, the AI number can be made to sound like anything.
Counting savings against costs that weren't real. This is the most common ROI mistake in the published case studies. "We saved 4,000 hours by automating X" is only true if those 4,000 hours of work were actually being done before. Often the work was being done badly, partially, or not at all. The honest ROI is against the previous-state cost, not against a fully-staffed-up version of the previous state.
Hiding cost in operating budgets. Token costs, infrastructure costs, the engineering time required to keep the AI feature running, the cost of incidents when it fails publicly. These don't disappear because they are spread across teams. They just stop being attributed to the feature that caused them. An AI feature whose true cost is hidden in three different budgets cannot be assessed for ROI.
Most of these are not malicious. They are the natural failure modes of a measurement framework that was added after the feature shipped, by people who were not present when the feature was scoped. The fix is to define the measurement framework at the same time as the feature, against an honest baseline, with all costs attributed.
What to measure, when
The metrics that actually tell you whether an AI feature is producing value vary by use case, but the structure is consistent. For each AI feature, define four things, in this order.
The outcome. What changes for the user, the customer, or the business if this feature works? Not "the AI does X," but "the user gets to Y faster" or "the business avoids Z cost." If you cannot write this in one sentence in plain language, you do not yet have a feature worth building. The MIT report's most useful framing was that the divide between the 5% and the 95% is mainly about whether organisations did this work upfront.
The leading metric. A measurable signal you will see within two weeks of launch. Task success rate. Acceptance rate of AI-generated content. Human escalation rate. User-rated quality. These are not the outcome; they are the early-read on whether you are heading toward it. Leading metrics let you iterate before the lagging numbers come in.
The lagging metric. The actual outcome, on a 90-day or 180-day horizon. Cost-per-resolved-ticket. Time-to-decision. Revenue per AI-assisted interaction. Customer satisfaction at the end of a fully-resolved interaction (not at the end of a deflected one). These are the numbers that justify the spend.
The honest cost. All-in. Token cost. Infrastructure cost. Engineering maintenance time at fully-loaded compensation. Incident cost when the feature misbehaves in public. The cost of the team meeting that scoped it, if you want to be thorough. The point is not to discourage AI features. The point is to be able to compare against the outcome they produce.
For AI features that touch customer-facing service, four specific metrics tend to be the most useful in combination: AI resolution rate (fully-solved, not just-handled); human escalation rate (and its trend over time); revenue per AI-assisted interaction (for features in the conversion path); and a quality signal from real users (CSAT specifically on AI-handled interactions, segmented).
For AI features that touch internal productivity, the most useful metrics are throughput-and-quality together. Throughput-only metrics get gamed quickly. DX's data on engineering productivity is instructive here: Change Failure Rate has swung significantly across organisations adopting AI coding tools, with some seeing meaningful improvements and others seeing up to 50% more defects than before AI adoption. Faster shipping only creates value if quality holds. If you cannot measure quality alongside speed, you cannot tell which kind of team you are.
The things that fall out of doing this honestly
There are some uncomfortable consequences of taking measurement seriously, and they are worth naming because most teams encounter them and then quietly back away from the discipline.
Some features get killed. A feature that does not move its leading metric within a few weeks, and does not move its lagging metric within a quarter, is consuming engineering capacity without producing value. The honest move is to retire it. Most organisations cannot bring themselves to do this with AI features, because killing the demo is harder than killing a normal feature. The teams that do this end up with smaller AI portfolios and larger AI returns.
Some features turn out to need a different team. MIT's data, again, found that more than half of generative AI budgets in their study population were being spent on sales and marketing applications, but the highest ROI was in back-office automation: eliminating business-process outsourcing, cutting external agency costs, streamlining operations. Honest measurement reveals which side of that pattern you are on, and a team that has been pushing AI in the wrong domain for a year will find this uncomfortable. It is still the right finding.
Some costs get reattributed. Token costs, infrastructure costs, engineering time. The full cost of an AI feature, attributed properly, is often significantly larger than the team building it has been claiming. This is information leadership needs. It is also information the building team would rather not surface. The measurement framework has to be set up by people who do not depend on the answer being flattering.
The pace of AI adoption looks slower, not faster. Teams that measure honestly tend to ship fewer AI features but ship them with higher impact. Teams that do not measure honestly tend to ship more AI features and produce less value. The first looks slower in a board deck. It is the one that compounds.
The discipline behind the discipline
Underneath the specific metrics, there is a deeper operating principle that the high-ROI teams seem to share: they treat AI features the same way they treat any other product feature. They write down the outcome they are trying to create. They define how they will know whether they got it. They count the cost honestly. They are willing to retire what is not working.
This is not a sophisticated insight. It is the ordinary discipline of product development, applied to a layer of the stack that has, for reasons of novelty, hype, and the difficulty of measurement, often escaped it. The teams that get returns out of AI in 2026 are the teams that stopped exempting AI from the rules they apply to everything else they ship.
The companies separating themselves from the median in 2026 are doing two things, in this order. They are getting their data layer right: the foundation under all of it. And they are measuring honestly: what worked, what did not, what cost what. The other side of that ledger, in rands, is what AI actually costs a small business. The combination compounds. The absence of either is what 95% looks like. Getting both right is usually faster with proper AI consulting than by working it out alone.
References
- MIT NANDA Initiative: The GenAI Divide: State of AI in Business 2025 (July 2025), summarised in Fortune. https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
- McKinsey: State of AI 2025, summarised in QuickLaunch Analytics. https://quicklaunchanalytics.com/bi-blog/data-ai-readiness-enterprise/
- UC Berkeley Executive Education: Beyond ROI: Are We Using the Wrong Metric in Measuring AI Success? https://exec-ed.berkeley.edu/2025/09/beyond-roi-are-we-using-the-wrong-metric-in-measuring-ai-success/
- SR Analytics: Why 95% of AI Projects Fail and How Data Fixes It. https://sranalytics.io/blog/why-95-of-ai-projects-fail/
- The AI Consultancy (Medium): The MIT "95% of GenAI Pilots Fail" Report: What It Gets Wrong. https://medium.com/@ai_93276/the-mit-95-of-genai-pilots-fail-report-what-it-gets-wrong-and-what-leaders-should-do-instead-3a6a1bd7a3d5
- Google Cloud: KPIs for Gen AI: Measuring Your AI Success. https://cloud.google.com/transform/gen-ai-kpis-measuring-ai-success-deep-dive
- Zendesk: What is AI resolution rate and how to measure it. https://www.zendesk.com/blog/ai/productivity/ai-resolution-rate/
- Moveworks: Enterprise Automation ROI: A Guide To Measuring and Maximizing ROI With AI. https://www.moveworks.com/us/en/resources/blog/measure-and-improve-enteprise-automation-roi
- Chrono: AI Feature KPIs That Actually Tell You Something Useful. https://www.chronoinnovation.com/resources/measuring-ai-feature-success-kpis/
- DX: How to measure AI's impact on developer productivity. https://getdx.com/blog/ai-measurement-hub/
- DX: How to measure AI performance in software engineering (Q1 2026 data). https://getdx.com/blog/measure-ai-impact/
- Decagon: What is deflection rate? https://decagon.ai/glossary/deflection-rate
Written by JP, Sixees Labs. Last reviewed May 2026.