Written by Gowtham Raj, Director at TartLabs, who leads custom software and AI engagements for enterprise and mid-market clients.
The Short Answer
The model is rarely what sinks an AI pilot. A pilot only has to show that a model can do a task. Production has to show that a business can lean on it daily. Five gaps lie between the two: data that isn't ready, no agreed meaning of "good enough", no owner or business case, no operating layer, and no slot in the actual workflow. S&P Global Market Intelligence reported that 42% of companies dropped most of their AI initiatives in 2025, compared with 17% a year earlier. You can test for each gap before a pilot begins, and that costs far less than discovering one after the demo.
Key Takeaways
- In 2025, 42% of companies scrapped most of their AI initiatives, against 17% in 2024. The average organisation also dropped 46% of its proofs of concept before they reached production (S&P Global Market Intelligence, 1,000+ respondents in North America and Europe, published March 2025).
- McKinsey's 2026 survey has 44% of organisations scaling AI across the enterprise, up from 38%. Only 37% credit AI with any EBIT impact, about the same as a year before.
- Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, citing rising costs, unclear business value or inadequate risk controls.
- The five gaps are data, evaluation, ownership, operations and workflow. That grouping is our view, built from the sources below and our delivery work. It is not a published standard.
- Most of the fix is sequencing. Answer the five gap questions first, and treat the pilot as the opening slice of production rather than a separate experiment.
What the Numbers Say
The surveys count different things, so look at where they point, not at the exact digits.
| Source | Finding | What it measures |
|---|---|---|
| S&P Global Market Intelligence, via CIO Dive (March 2025) | 42% abandoned most AI initiatives, up from 17%; 46% of proofs of concept scrapped on average | 1,000+ respondents, North America and Europe |
| McKinsey, State of AI 2026 (fielded May to June 2026) | 44% scaling AI across the enterprise (38% a year earlier); 37% report any EBIT impact | 1,719 respondents in 97 countries |
| Gartner (25 June 2025) | Over 40% of agentic AI projects cancelled by end of 2027 | A forecast; also a January 2025 poll of 3,412 webinar attendees |
Two cautions. S&P counted companies that dropped most of their initiatives, which is not the same as every pilot, and the figure is a 2025 snapshot. McKinsey's aggregate numbers are getting better. That shows scaling can be done, but it says nothing about whether a particular pilot will manage it.
Blog posts also keep repeating "95% of pilots fail" and "88% never reach production". We left both out of our headlines. The 95% number traces to MIT NANDA's 2025 report. We could not retrieve the full report. It drew on 153 survey responses, interviews and a review of public deployments over about six months, and its authors call it a directionally accurate snapshot, not a definitive market analysis. The 88% number spreads mainly through secondary write-ups, and we found no primary page we could read. Read both as signals of direction, not as rates to plan around.
The 5 Gaps Between a Pilot and Production
Gap 1: The data was clean for the demo, not for the business
Pilots tend to run on a hand-picked extract, a few thousand tidy records somebody prepared by hand. Production brings every odd record, late feed and duplicate customer. In the S&P Global survey (as reported by CIO Dive), cost, data privacy and security risk topped the list of obstacles. Those are not data-quality findings as such, but they sit next to it. Gartner's August 2024 forecast adds that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, with poor data quality among the reasons (as reported by Intelligent CIO).
Check before the pilot: does the pilot pull its data the way production will, on the same schedule and with the same holes? If a person cleaned it by hand, the pilot result proves little. Legacy systems complicate this further, and our guide to why legacy systems block AI adoption lays out the order to fix things in.
Gap 2: Nobody defined "good enough"
A demo gets judged on feel. Production needs a number: which accuracy, which response time, which rate of unsafe answers you will accept, all measured on a fixed test set the team never tunes against. Skip that, and each release turns into a debate.
In our view, this is a common reason a pilot looks fine and then stalls in review. The sponsor asks "how do we know it works?" and all anyone can offer is the demo. Check before the pilot: settle the pass mark and the evaluation set first. If the team can't agree on them, that disagreement is the pilot's real result.
Gap 3: No owner and no business case
An innovation budget can fund a pilot without anyone being accountable for what it delivers. Production wants a named owner for the outcome, a cost line, and a baseline to beat. Among the reasons in Gartner's cancellation forecast are unclear business value and escalating costs. In our view, both only show up once someone has to defend the spend.
McKinsey's 2026 finding fits. Eight in ten respondents say AI has made them personally more productive, yet only 37% report any EBIT impact. Personal gains are easy to feel, while gains for the enterprise need someone to measure them. Check before the pilot: name the business owner, pick the baseline metric, and set the monthly run cost at which the project stops being worth it.
Gap 4: No operating layer
A demo has nobody on call. Production needs monitoring for quality drift, cost per request, access control, audit logs, a rollback path and a human fallback. This is where Gartner's third stated reason, inadequate risk controls, sits. McKinsey's 2026 survey also found that AI operating costs now limit AI use at roughly one in five organisations, so for many teams run cost arrives as a surprise on launch day.
Check before the pilot: can you say who gets paged when output quality falls, what one request costs at ten times the pilot's volume, and how to turn it off safely? Our AI coding standards template applies the same discipline.
Gap 5: It was never part of the real workflow
This gap is the quietest. The pilot sits in a standalone tool that a handful of enthusiasts open. In production, the output has to land in the system people already use, exceptions need a route, and the affected staff must have been trained and consulted. Gartner also cautions about "agent washing", where products are relabelled as agentic with no real capability behind them. In our view, a pilot that never touches a real workflow can conceal that problem for months.
Check before the pilot: which screen, which approval step and which existing report will this change? If the answer is "none, it's a side tool", you are testing a demo, not a workflow.
How the Gaps Compound
The gaps feed each other. Without an evaluation set (Gap 2) you can't tell if bad data (Gap 1) is hurting you. Without an owner (Gap 3), nobody pays for the operating layer (Gap 4). A side-tool pilot (Gap 5) leaves no baseline to measure against. In our experience, this is why a team that fixes one gap often watches the project stall at the next. Our view is to run the five as a single readiness review in one week, before any build begins. For the wider strategy around that review, see our AI strategy for software companies guide.
A Pre-Pilot Checklist You Can Run This Week
Answer yes or no for the pilot you're about to approve:
- Does the pilot use production-grade data, pulled the way production will pull it?
- Is there a written pass mark and a fixed evaluation set?
- Does one named person answer for the business outcome?
- Do we know the baseline metric and the monthly cost at full volume?
- Are monitoring, a rollback path and a human fallback planned?
- Will the output land inside a system people already use?
- Was the team that will live with the change involved?
In our view, three or more "no" answers mean you have a demo with a budget. Fix the foundation first, or cut the pilot down to a slice you can take to production within a quarter.
The Bottom Line
AI pilots stall because they are built to impress rather than to ship. The sources agree on direction: scaling is growing, but financial returns haven't caught up. Five gaps account for most of it: data, evaluation, ownership, operations and workflow.
Run the checklist on your next pilot before you approve its budget. If a pilot is stuck and you want a second opinion, get in touch, read our AI strategy for software companies guide, or see how a dedicated team can carry one bounded use case from pilot to production. If the pilot should be stopped, we will say so.




