The most useful AI statistic of the past year is not about what the technology can do. It is about what companies get back. MIT researchers examined more than 300 enterprise deployments and found that 95 percent of generative AI pilots produced no measurable profit-and-loss impact. The 5 percent that paid were not using better models. They were running the work differently. This article sets out the verified numbers, the reasons pilots fail, and the sequence a mid-sized company should follow to be in the minority that gets a return.

MIT's Project NANDA published The GenAI Divide: State of AI in Business 2025 in August 2025, built on a review of more than 300 publicly disclosed AI initiatives, structured interviews with 52 organisations and survey responses from 153 senior leaders. Its headline finding: 95 percent of generative AI pilots stall, delivering little or no measurable impact on the profit and loss. About 5 percent of pilots reach production and move a number the board cares about.
McKinsey's State of AI survey, published on 25 August 2026 from 1,719 respondents in 97 countries, tells the same story from the other end. Nearly nine in ten organisations now use AI regularly in at least one function, and 44 percent report AI scaling across the enterprise. Yet only 37 percent can attribute any EBIT impact to it, essentially unchanged from a year earlier. Adoption has raced ahead of return.
BCG's Widening AI Value Gap report, published on 30 September 2025 from a survey of 1,250 senior executives, sorts companies into three groups: 5 percent are built to generate AI value at scale, 35 percent are beginning to see value, and 60 percent report minimal gains. The leaders show 1.7 times the revenue growth of the laggards and 3.6 times the three-year shareholder return. The distribution matters more than the averages: AI returns are not spread thinly across everyone, they are concentrated in a small group doing specific things.
The MIT researchers are unusually direct about the cause. The failures were not about model quality. Generic tools work well for individuals precisely because the individual bends around the tool; in a business, the tool has to bend around the workflow, and most deployments never manage it. MIT calls this the learning gap: most enterprise GenAI systems do not retain feedback, adapt to context or improve with use. A pilot that sits beside the workflow rather than inside it produces demonstrations, not savings.
Two secondary findings from the same report deserve an owner's attention. First, the money is often pointed at the wrong end of the business: more than half of generative AI budgets go to sales and marketing tools, while MIT found the largest measured returns in back-office automation, replacing outsourced processing, cutting agency spend and streamlining operations. Second, buying beats building: tools purchased from specialised vendors and delivered through partnerships succeeded roughly 67 percent of the time, while internal builds succeeded about a third as often. For a mid-sized company without a machine-learning team, this is good news. The winning move is not hiring one.
McKinsey's high performers, the roughly 6 percent of respondents attributing a meaningful share of EBIT to AI, share one habit above all: nearly three quarters report fundamentally redesigning workflows because of AI, up from 55 percent a year earlier. They did not add AI to the old process. They changed the process. BCG expresses the same point as its 10/20/70 principle: about 10 percent of the effort in a successful AI programme goes into algorithms, 20 percent into technology and data, and 70 percent into people and process change. Companies that budget the first 30 percent and ignore the 70 are the ones filling MIT's failure column.
MIT's successful buyers also behaved differently as customers. They demanded deep customisation to their own operations, and they held vendors accountable to business metrics rather than technical ones. The question they asked was not whether the tool worked, but whether the invoice queue shrank.
Not a function, a workflow: invoice matching, quote preparation, first-draft responses to enquiries, report assembly. It should be frequent, rule-rich and currently absorbing paid hours you can count. The back office is the right place to look first; the data says the returns are there.
Hours per week, error rate, turnaround time, cost. If the pilot cannot be judged against a number recorded before it started, it will be judged on enthusiasm, and enthusiasm is how 95 percent of pilots end.
A specialised vendor tool, configured to your process, carries roughly twice the success rate of building in-house. Reserve internal development for the rare case where the workflow is genuinely unlike anyone else's.
The person accountable for the pilot should be the person who owns the workflow's number. Integration decisions, exception handling and staff training are operational questions, and MIT's evidence says they are where the value is won or lost.
Ninety days is enough to know whether the baseline moved. If it did, redesign the workflow around the tool and scale it. If it did not, stop, keep the baseline, and pick the next workflow. A cheap, honest failure is a result; an indefinite pilot is a cost.
The 2026 data reads as a warning, but for an SME it is closer to an opening. The 95 percent failure rate belongs mostly to enterprises running broad, centrally approved programmes far from the work itself. The successful pattern, one workflow, one bought tool, one owner, one measured number, is easier to run in a 50-person company than in a 5,000-person one. The distance between the decision-maker and the workflow is shorter, and that distance is precisely what the failures have in common.
Offer and credential pages, Google Business Profile, consistent details across sources, content shaped to the questions customers ask, and a monthly prompt panel that records who the engines name.
MeasurementTracking from first touch to paid sale, without customer data leaving your systems, with source tagged on every sale.
SalesFor founder-led sales: discovery questions, qualification, call structure, proposal and follow-up, written from your real deals and trained until the founder runs it without us.