Whitepaper 04 · The Agentic Enterprise series

Escaping Pilot Purgatory

Why enterprise AI stalls between the demo and the P&L — and what the few who escape do differently

Start reading
Data Outcome Talent Control P&L impact the 5 percent FUNDED PILOTS
Figure 1. Enterprises do not have an AI shortage. They have an attrition problem: pilots stop at gates that have nothing to do with the model.
However beautiful the strategy, you should occasionally look at the results.
— attributed to Winston Churchill
• • •
01

Pilot purgatory

The hardest problem in enterprise AI is no longer building something impressive. It is getting something impressive to matter.

Almost every large organization can now produce a compelling demonstration in a fortnight. A model summarizes claims files, drafts underwriting notes, answers questions about a policy manual, writes code against an internal API. The room is impressed. A budget line appears. And then, somewhere between that room and the operating plan, the work quietly stops mattering. It does not fail loudly. It is renewed, re-scoped, folded into another initiative, and eventually described in the past tense.

The industry has a name for this condition: pilot purgatory. It is not a technology failure and it is rarely an intelligence failure. It is the predictable result of running experiments in an organization that has not decided what it wants the experiments to change.

Figure 2 · The four gaps that define the problem
0

Of generative AI pilots produce no measurable P&L impact, in MIT's study of 300 public deployments1

0

Of enterprises say their data is completely ready for AI adoption4

0

Of CIOs name a lack of in-house talent as the top obstacle to their AI strategy5

0

Of organizations have mature governance for the AI agents they are already deploying7

Figure 2. Four independent research programs, four different questions, one shape of answer: capability is running well ahead of the conditions capability needs.

The number everyone quotes, and what it actually says

The 95 percent figure comes from MIT's State of AI in Business research — 150 leader interviews, 350 employee responses, and an analysis of 300 public deployments.1 It is worth being precise about what it measured, because the headline has been stretched well past its evidence. It did not find that 95 percent of AI systems do not work. It found that roughly 95 percent of generative AI pilots produced no measurable acceleration in the P&L. Those are very different claims. Most of these pilots worked exactly as designed. They simply were not attached to anything that shows up in a quarterly result.

The same pattern appears in data that is not about pilots at all. McKinsey's 2026 global survey of 1,719 respondents found that 44 percent of organizations are now scaling AI across the enterprise, up from 38 percent a year earlier — real, measurable progress. Over the same period, the share reporting any EBIT contribution from AI stayed flat at 37 percent, and only 6 percent qualified as high performers, meaning they attribute at least 5 percent of EBIT to AI and describe the value as significant.2 Adoption moved. Impact did not.

Figure 3 · The narrowing
Scaling AI enterprise-wide 44% of organizations, 2026 Reporting any EBIT contribution 37% unchanged year over year High performers (5%+ of EBIT) 6% the group worth studying Pilots with measurable P&L impact ~5% MIT, 300 public deployments
Figure 3. Two independent studies, measuring different populations, arrive at the same narrow end. Sources: McKinsey (2026)2 and MIT (2025).1
The gap is not between what AI can do and what it does. It is between what AI does and what the organization has decided to change.

Read that way, the numbers stop being an indictment of the technology and become a description of an operating problem — which is good news, because operating problems are the kind executives already know how to solve. The rest of this paper is about the specific shape of that problem: where programs break, why the obvious fixes make it worse, and what the organizations on the right side of the gap do differently.

02

Where programs break

Stalled AI programs fail in remarkably similar ways. In practice there are four breakpoints, and a program rarely survives more than one of them unattended.

Figure 4 · The path from pilot to production — and the four places it breaks
Ground Aim Staff Control Pilot proof of concept Production at scale decisions that run the business 01 02 03 04
01

The ground was never made ready

Siloed systems, thin metadata, no lineage. A pilot survives on a curated extract; production cannot. Only 7 percent of enterprises call their data completely AI-ready, and 56 percent name silos as the leading obstacle.4

02

The technology was chosen before the outcome

The initiative is named after a model or a vendor. No one can say which business decision gets faster, cheaper, or more defensible if it works — so no one can say when it is finished.

03

Nobody owns the system after the demo

Pilots are built by a team that does not run anything. Forty percent of CIOs name missing in-house talent as their top AI obstacle5 — and the scarcest skill is not model building, it is operating a model in production.

04

Speed and safety are pulling apart

The business adopts for speed; IT holds for risk. The gap fills with shadow AI: 49 percent of employees use unapproved tools and 51 percent have connected one to a work system without approval.6

Figure 4. Hover or focus a break to connect it to its cause. Programs stall at whichever of these four is left unattended first — which is why the failure looks different in every organization and is structurally identical in all of them.

01 · Data that only works in a demo

Every pilot begins with a curated extract. Someone pulls a clean slice of claims, tickets, or contracts, and the model performs beautifully against it. Production asks a harder question: can the system reach this data continuously, know where it came from, and be trusted to act on it when it is late, partial, or contradicted by another system of record?

For most enterprises the honest answer is not yet. In research by Cloudera and Harvard Business Review Analytic Services, just 7 percent of organizations said their data was completely ready for AI adoption; 27 percent said it was not very or not at all ready; and 56 percent named siloed data and integration difficulty as their leading obstacle.4 Generative systems do not forgive this. They amplify whatever is underneath them: a fragmented record produces confident, fluent, fragmented answers, and the fluency is what makes it dangerous.

02 · Technology chosen before the outcome

The second break is the one leaders have the most control over and address the least. A program is launched because the organization has decided it needs AI, rather than because it has decided which results it intends to change. The initiative gets the name of a technology, not a business outcome.

This is not a semantic complaint. Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, and the reasons it names are escalating costs, unclear business value, and inadequate risk controls.3 Two of those three are consequences of an unnamed outcome: without a business result to bound the work, cost has no ceiling and value has no test. The same analysis notes that only about 130 of the thousands of vendors claiming agentic capability are doing anything genuinely agentic — a market condition Gartner calls agent washing. An organization that has not defined its own outcome is precisely the buyer that cannot see through it.

03 · No owner after the demo

Pilots are usually run by innovation groups, centers of excellence, or a vendor's delivery team — none of whom will still be there when the system needs a retraining schedule, an escalation path, and a named accountable executive. The handoff is the moment most programs die, and it dies quietly because nobody has been assigned to notice.

The talent gap is real: 40 percent of CIOs cite a lack of in-house talent as the leading obstacle to executing their AI strategy.5 But it is misdiagnosed as often as it is cited. The shortage is rarely in model development. It is in the unglamorous middle: people who can run evaluation harnesses, monitor drift, hold a vendor to a service level, and translate a risk finding into a business decision. Organizations recruit for the first skill and stall for want of the second.

04 · Speed and safety pulling in opposite directions

The fourth break is cultural before it is technical. Business units adopt for speed. Security, legal, and risk hold for safety. Neither is wrong, and in the absence of a shared decision rule the friction resolves itself the way friction always does: unofficially.

In a January 2026 survey of 2,000 employees at organizations with more than 500 people, 49 percent admitted using AI tools their employer had not approved, 51 percent had connected an AI tool to a work system without IT's involvement, and 58 percent were doing it on free consumer tiers.6 Roughly a third had shared internal research or datasets. Most striking: 69 percent of C-suite respondents said they were fine with it. Shadow AI is not a compliance nuisance at the edge of the organization. It is the organization routing around a governance process that is slower than the work.

And governance is not catching up on its own. In Deloitte's 2026 survey of 3,235 leaders across 24 countries, only 21 percent reported mature governance for agentic AI, even as deployment accelerates.7 Agents are being given the ability to act before most organizations have decided what they are allowed to do, who reviews it, and what the audit trail looks like.

Shadow AI is not an employee discipline problem. It is a governance latency problem — and the fix is a faster path to yes, not a longer list of no.
03

The fixes that don't work

Organizations that recognize they are stalled usually reach for one of three remedies. Each is reasonable. Each, on its own, deepens the stall.

Buying a bigger platform

The instinct when a pilot underdelivers is to conclude the tooling was too small, and to replace it with an enterprise platform. Sometimes that is right. The evidence suggests that buying is often the better path on the technology axis: MIT found that purchased, specialized tools succeeded roughly two-thirds of the time, while internal builds succeeded about a third as often.1

But that finding is about which capability to buy, not about whether purchase substitutes for decision. A platform inherits every condition the pilot lacked — the same fragmented data, the same unnamed outcome, the same absent owner — and adds a contract, an implementation program, and a sunk cost that makes the eventual reckoning more expensive. Buying accelerates a program that knows where it is going. It does not tell a program where to go.

Running more pilots

The second remedy is volume: if one in twenty pilots produces impact, run more pilots. This is the most expensive misreading of the data available, because it treats a five percent success rate as a sampling problem rather than a structural one. The failures were not random. They clustered at four specific breakpoints, and running more experiments through the same four gates produces the same attrition at greater cost — plus an organization that is now visibly tired of AI.

There is a second-order cost too. MIT's research found more than half of generative AI budgets flowing to sales and marketing, while the strongest returns showed up in back-office work — eliminating outsourced processing, reducing agency spend, compressing operational cycles.1 A portfolio built from whichever pilot found a sponsor drifts toward the most visible functions rather than the most valuable ones. More pilots amplifies that drift.

Waiting for the models to get better

The third remedy is patience: the technology is moving quickly, so wait for the next generation to close the gap. This is the only one of the three that costs nothing today, which is exactly what makes it attractive and what makes it dangerous. The four breakpoints in Chapter 02 are not model limitations. A better model does not integrate your claims system, name your target outcome, hire your operator, or decide what an agent is permitted to do without a human in the loop.

Meanwhile the clock is real. Gartner projects that by 2028, 15 percent of day-to-day work decisions will be made autonomously — from effectively zero in 2024 — and that a third of enterprise software will include agentic capability.3 The organizations that have solved their data access, their ownership model, and their governance latency will absorb that shift as an upgrade. The organizations that waited will meet it as a migration.

Figure 5 · The reframe
Stop asking which AI projects to fund. Start asking which decisions the business needs to make faster, clearer, and more defensibly — then fund the capability those decisions require.
Skip Vanderburg
A finish line appears

A decision has an owner, a cadence, and a moment it is made. That gives the work a definition of done that a capability never has.

Value becomes measurable

Cycle time, rework, and the cost of delay on a named decision are all countable before and after — which is what a CFO means by evidence.

Governance gets a subject

Controls attach to a specific decision and its blast radius, instead of to AI in the abstract — which is why they can be written in weeks rather than quarters.

Figure 5. The single change that separates programs that scale from programs that repeat: the unit of planning stops being a technology and becomes a decision.
04

The exit: six moves

Escaping pilot purgatory is not a bigger version of what stalled. It is a change in what the portfolio is made of.

The single most useful finding in the 2026 McKinsey data is what separates the 6 percent of high performers from everyone else. It is not model choice, spend, or headcount. Nearly three-quarters of high performers — 72 percent — had fundamentally redesigned workflows around AI, against 25 percent of the rest.2 They did not add intelligence to the way the work already ran. They changed how the work runs, which means they changed how decisions get made.

Figure 6 · Two ways to hold the same budget

Switch the view. The money is the same in both. The accountability is not.

Figure 6. A pilot portfolio is a list of things being tried. A decision portfolio is a list of choices the business has agreed to make better, each with an owner, a cadence, and a measurable delta.

Converting one into the other is a sequence, and the order matters more than the speed. These are the six moves I run with leadership teams, and they map directly to the six phases of the Decision Velocity Framework.

Figure 7 · Six moves out of purgatory
01 · FRAME

Name the decisions

List the ten to fifteen recurring choices that actually move the P&L — pricing, adjudication, allocation, triage, sequencing. Write each as a question with an owner and a cadence. This list, not the technology inventory, becomes the portfolio.

02 · DIVERGE

Widen before you narrow

For each decision, generate the full set of ways it could be improved — better data, a model, a changed policy, a removed approval step. Several will not need AI at all. Finding those early is what keeps the AI work credible.

03 · SCORE

Score on one sheet

Value, confidence, effort, and risk — scored the same way across every candidate, by the people who own the outcome. The point is not the arithmetic. It is forcing four functions to argue against one set of numbers instead of four narratives.

04 · SORT

Sequence against readiness

Order the survivors by the conditions they need, not by enthusiasm. A decision whose data is three quarters away goes behind one whose data is already governed — no matter how attractive the demo was.

05 · CUT

Kill the rest, visibly

Stopping work is the move most programs skip, and it is the one that buys everything else. Announce what is being stopped and why. An organization that has never seen an AI project cancelled does not believe the criteria are real.

06 · PROVE

Instrument the decision

Baseline the cycle time, rework rate, and cost of delay before anything ships, and publish the delta on the same cadence as the decision itself. Evidence produced monthly is what converts a pilot budget into an operating line.

Framename the decision
Divergewiden the options
Scoreone sheet, one scale
Sortsequence by readiness
Cutstop the rest
Provepublish the delta
Figure 7. The six moves are deliberately unglamorous. None of them is an AI problem, which is precisely why AI programs stall without them.
A portfolio of pilots asks the organization to be patient. A portfolio of decisions asks it to be accountable. Only one of those survives a budget cycle.

Governance belongs inside this sequence, not after it. The Deloitte finding — 21 percent with mature agentic governance while deployment accelerates7 — describes organizations writing controls for AI in general, which cannot be finished. Controls written for a specific decision can: what the system may decide alone, what it must escalate, what evidence it records, and who reviews the exceptions. That is a two-week document per decision rather than a two-quarter policy program, and it is what closes the latency gap that shadow AI is exploiting.

05

What this looks like by sector

The four breakpoints are universal. Which one bites first is not — it depends on how the industry is regulated, how its data was accumulated, and where its margin actually sits.

Figure 8 · Where the stall shows up first

Healthcare

Providers & payers

The stall is almost always at Control. Clinical and privacy review is rigorous and slow by design, so ambient documentation and triage pilots pass their evaluations and then wait months for an approval pathway that was never designed for a system that learns.

Meanwhile the data breakpoint hides behind an integrated EHR: one system of record is not the same as one governed, retrievable, lineage-tracked corpus across imaging, notes, claims, and scheduling.

Opening move · Pick one decision with a bounded blast radius — documentation burden, prior-authorization triage, discharge sequencing — and write its control document before the build, not after.

Figure 8. Select a sector. The diagnosis is structural, but the order of attack is industry-specific — and starting with the wrong breakpoint is how well-run programs lose a year.
06

The first ninety days

If your organization is somewhere in the 95 percent, the way out does not begin with a new model, a new platform, or a new team. It begins with a decision inventory and a willingness to stop things.

Figure 9 · A ninety-day exit
1

Days 1–15 · Inventory both sides. Every AI initiative currently funded, with its sponsor and its stated outcome. Beside it, the ten to fifteen recurring decisions that move the P&L. The gap between the two lists is the diagnosis, and it is usually visible in a single meeting.

2

Days 16–35 · Score once, together. One scoring sheet, one scale, the four functions that will have to live with the answer in the same room. Value, confidence, effort, risk. What comes out is not a ranking so much as an agreement about what the organization believes.

3

Days 36–50 · Cut, and say so. Stop the initiatives that do not attach to a named decision. Publish the list and the reasoning. This is the move that converts a scoring exercise into a governing one, and it is the one most leadership teams flinch from.

4

Days 51–90 · Prove one decision end to end. One decision, one owner, one control document written before the build, one baseline captured before anything ships. Report the delta in cycle time and cost of delay at the decision's own cadence — weekly if it is made weekly.

Figure 9. Ninety days is enough to leave purgatory with one decision governed and proven. It is not enough to transform the enterprise — and programs that promise the second usually fail to deliver the first.

Notice what this sequence does not require. No new platform. No reorganization. No hiring plan that takes two quarters to fill. It requires an executive team willing to name a small number of decisions, agree on how they are scored, stop the work that does not serve them, and publish evidence on a cadence. The technology, for once, is the easy part.

The 5 percent are not better at AI. They are better at deciding.

That is the uncomfortable and ultimately encouraging conclusion of every study cited in this paper. The differentiator McKinsey found was workflow redesign. The differentiator MIT found was whether the system was embedded in how work actually happens. The reasons Gartner gives for cancellation — unclear value, uncontrolled cost, inadequate controls — are all failures of decision, not of engineering. None of these are things a vendor can sell you, and all of them are things a leadership team can start on this quarter.

Agentic systems raise the stakes on all of it. A generative pilot that has no owner produces a document nobody reads. An agent that has no owner takes actions nobody reviews. The organizations that spend the next year building the habit of naming decisions, governing them at the decision level, and proving the delta will be the ones able to hand real authority to a machine when it is worth handing over. The rest will still be running pilots.

Pilot purgatory is not where AI programs go to fail. It is where organizations wait until they decide what they want.

Sources

Every figure in this paper is drawn from named research published between mid-2025 and mid-2026. Where a study is widely paraphrased — the 95 percent figure above all — the original methodology is noted so the claim can be read at its actual strength.

  1. 01 MIT / Project NANDA, The GenAI Divide: State of AI in Business (2025). 150 leader interviews, 350 employee survey responses, and analysis of 300 public generative AI deployments; finds ~95% of pilots produced no measurable P&L acceleration, purchased tools succeeded ~67% of the time versus roughly a third as often for internal builds, and more than half of budgets went to sales and marketing while back-office work delivered stronger returns. Reported coverage.
  2. 02 McKinsey & Company, The State of AI: Global Survey 2026. Fielded 4 May – 8 June 2026; 1,719 respondents across 97 nations. 44% scaling AI enterprise-wide (from 38%), 37% reporting EBIT contribution (unchanged), 6% high performers at ≥5% of EBIT, 72% of high performers having fundamentally redesigned workflows against 25% of others, ~20% scaling coding agents, 32% declining a software purchase in favor of building. mckinsey.com.
  3. 03 Gartner, Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (25 June 2025). Cancellation drivers: escalating costs, unclear business value, inadequate risk controls. Also: ~130 of thousands of claimed agentic vendors assessed as genuine (“agent washing”); 15% of day-to-day work decisions made autonomously by 2028, from 0% in 2024. gartner.com.
  4. 04 Cloudera with Harvard Business Review Analytic Services (surveyed October 2025; published 5 March 2026). 230+ respondents involved in AI data decisions: 7% call their data completely ready for AI, 27% not very or not at all ready, 56% cite siloed data and integration difficulty as the top obstacle, 73% say data quality deserves more priority than it gets. cloudera.com.
  5. 05 2026 State of the CIO survey, reported by CIO (30 April 2026). 40% of respondents name a lack of in-house talent as the top challenge to implementing their AI strategy, with the scarcity concentrated in production-scale and AI-security skills rather than model development. cio.com.
  6. 06 BlackFog shadow-AI research (published 29 January 2026); 2,000 employees at organizations with 500+ staff. 49% use AI tools without employer approval, 51% have connected an AI tool to a work system without IT approval, 58% use free consumer tiers, 33% have shared internal research or datasets, and 69% of C-suite respondents accept unsanctioned use. Reported coverage.
  7. 07 Deloitte, AI agents are scaling faster than their guardrails (24 April 2026); 3,235 business and IT leaders across 24 countries. Only 21% report mature governance models for agentic AI, with decision boundaries, real-time monitoring, and audit trails identified as the most common gaps. deloitte.com.

Where this goes next

Which decisions would you fix first?

This paper is drawn from The Agentic Enterprise: From Automation to Autonomous Decision Intelligence. If your AI portfolio is full of promising work that has not reached the P&L, the first useful conversation is not about models — it is about which decisions the work was supposed to improve. That is the conversation I have most often.