Figure 1. Enterprises do not have an AI shortage. They have an attrition problem: pilots stop at gates that have nothing to do with the model.
However beautiful the strategy, you should occasionally look at the results.
— attributed to Winston Churchill
• • •
01
Pilot purgatory
The hardest problem in enterprise AI is no longer building something impressive. It is getting
something impressive to matter.
Almost every large organization can now produce a compelling demonstration in a fortnight. A model
summarizes claims files, drafts underwriting notes, answers questions about a policy manual, writes
code against an internal API. The room is impressed. A budget line appears. And then, somewhere between
that room and the operating plan, the work quietly stops mattering. It does not fail loudly. It is
renewed, re-scoped, folded into another initiative, and eventually described in the past tense.
The industry has a name for this condition: pilot purgatory. It is not a technology failure and
it is rarely an intelligence failure. It is the predictable result of running experiments in an
organization that has not decided what it wants the experiments to change.
Figure 2 · The four gaps that define the problem
0
Of generative AI pilots produce no measurable P&L impact, in MIT's study of 300 public deployments1
0
Of enterprises say their data is completely ready for AI adoption4
0
Of CIOs name a lack of in-house talent as the top obstacle to their AI strategy5
0
Of organizations have mature governance for the AI agents they are already deploying7
Figure 2. Four independent research programs, four different questions, one shape of answer: capability is running well ahead of the conditions capability needs.
The number everyone quotes, and what it actually says
The 95 percent figure comes from MIT's State of AI in Business research — 150 leader interviews,
350 employee responses, and an analysis of 300 public deployments.1
It is worth being precise about what it measured, because the headline has been stretched well past its
evidence. It did not find that 95 percent of AI systems do not work. It found that roughly 95 percent of
generative AI pilots produced no measurable acceleration in the P&L. Those are very different
claims. Most of these pilots worked exactly as designed. They simply were not attached to anything that
shows up in a quarterly result.
The same pattern appears in data that is not about pilots at all. McKinsey's 2026 global survey of 1,719
respondents found that 44 percent of organizations are now scaling AI across the enterprise, up from 38
percent a year earlier — real, measurable progress. Over the same period, the share reporting any EBIT
contribution from AI stayed flat at 37 percent, and only 6 percent qualified as high performers, meaning
they attribute at least 5 percent of EBIT to AI and describe the value as
significant.2 Adoption moved. Impact did not.
Figure 3 · The narrowing
Figure 3. Two independent studies, measuring different populations, arrive at the same narrow end. Sources: McKinsey (2026)2 and MIT (2025).1
The gap is not between what AI can do and what it does. It is between what AI does and what the
organization has decided to change.
Read that way, the numbers stop being an indictment of the technology and become a description of an
operating problem — which is good news, because operating problems are the kind executives already know
how to solve. The rest of this paper is about the specific shape of that problem: where programs break,
why the obvious fixes make it worse, and what the organizations on the right side of the gap do
differently.
02
Where programs break
Stalled AI programs fail in remarkably similar ways. In practice there are four breakpoints, and a
program rarely survives more than one of them unattended.
Figure 4 · The path from pilot to production — and the four places it breaks
01
The ground was never made ready
Siloed systems, thin metadata, no lineage. A pilot survives on a curated extract; production cannot. Only 7 percent of enterprises call their data completely AI-ready, and 56 percent name silos as the leading obstacle.4
02
The technology was chosen before the outcome
The initiative is named after a model or a vendor. No one can say which business decision gets faster, cheaper, or more defensible if it works — so no one can say when it is finished.
03
Nobody owns the system after the demo
Pilots are built by a team that does not run anything. Forty percent of CIOs name missing in-house talent as their top AI obstacle5 — and the scarcest skill is not model building, it is operating a model in production.
04
Speed and safety are pulling apart
The business adopts for speed; IT holds for risk. The gap fills with shadow AI: 49 percent of employees use unapproved tools and 51 percent have connected one to a work system without approval.6
Figure 4. Hover or focus a break to connect it to its cause. Programs stall at whichever of these four is left unattended first — which is why the failure looks different in every organization and is structurally identical in all of them.
01 · Data that only works in a demo
Every pilot begins with a curated extract. Someone pulls a clean slice of claims, tickets, or contracts,
and the model performs beautifully against it. Production asks a harder question: can the system reach
this data continuously, know where it came from, and be trusted to act on it when it is late, partial,
or contradicted by another system of record?
For most enterprises the honest answer is not yet. In research by Cloudera and Harvard Business Review
Analytic Services, just 7 percent of organizations said their data was completely ready for AI adoption;
27 percent said it was not very or not at all ready; and 56 percent named siloed data and integration
difficulty as their leading obstacle.4 Generative systems do not
forgive this. They amplify whatever is underneath them: a fragmented record produces confident,
fluent, fragmented answers, and the fluency is what makes it dangerous.
02 · Technology chosen before the outcome
The second break is the one leaders have the most control over and address the least. A program is
launched because the organization has decided it needs AI, rather than because it has decided which
results it intends to change. The initiative gets the name of a technology, not a business outcome.
This is not a semantic complaint. Gartner expects more than 40 percent of agentic AI projects to be
canceled by the end of 2027, and the reasons it names are escalating costs, unclear business value, and
inadequate risk controls.3 Two of those three are consequences
of an unnamed outcome: without a business result to bound the work, cost has no ceiling and value has no
test. The same analysis notes that only about 130 of the thousands of vendors claiming agentic
capability are doing anything genuinely agentic — a market condition Gartner calls
agent washing. An organization that has not defined its own outcome is precisely the buyer that
cannot see through it.
03 · No owner after the demo
Pilots are usually run by innovation groups, centers of excellence, or a vendor's delivery team — none
of whom will still be there when the system needs a retraining schedule, an escalation path, and a
named accountable executive. The handoff is the moment most programs die, and it dies quietly because
nobody has been assigned to notice.
The talent gap is real: 40 percent of CIOs cite a lack of in-house talent as the leading obstacle to
executing their AI strategy.5 But it is misdiagnosed as often as
it is cited. The shortage is rarely in model development. It is in the unglamorous middle: people who
can run evaluation harnesses, monitor drift, hold a vendor to a service level, and translate a risk
finding into a business decision. Organizations recruit for the first skill and stall for want of the
second.
04 · Speed and safety pulling in opposite directions
The fourth break is cultural before it is technical. Business units adopt for speed. Security, legal,
and risk hold for safety. Neither is wrong, and in the absence of a shared decision rule the friction
resolves itself the way friction always does: unofficially.
In a January 2026 survey of 2,000 employees at organizations with more than 500 people, 49 percent
admitted using AI tools their employer had not approved, 51 percent had connected an AI tool to a work
system without IT's involvement, and 58 percent were doing it on free consumer
tiers.6 Roughly a third had shared internal research or
datasets. Most striking: 69 percent of C-suite respondents said they were fine with it. Shadow AI is
not a compliance nuisance at the edge of the organization. It is the organization routing around a
governance process that is slower than the work.
And governance is not catching up on its own. In Deloitte's 2026 survey of 3,235 leaders across 24
countries, only 21 percent reported mature governance for agentic AI, even as deployment
accelerates.7 Agents are being given the ability to act before
most organizations have decided what they are allowed to do, who reviews it, and what the audit trail
looks like.
Shadow AI is not an employee discipline problem. It is a governance latency problem — and the fix is a
faster path to yes, not a longer list of no.
03
The fixes that don't work
Organizations that recognize they are stalled usually reach for one of three remedies. Each is
reasonable. Each, on its own, deepens the stall.
Buying a bigger platform
The instinct when a pilot underdelivers is to conclude the tooling was too small, and to replace it with
an enterprise platform. Sometimes that is right. The evidence suggests that buying is often the better
path on the technology axis: MIT found that purchased, specialized tools succeeded roughly two-thirds of
the time, while internal builds succeeded about a third as often.1
But that finding is about which capability to buy, not about whether purchase substitutes for
decision. A platform inherits every condition the pilot lacked — the same fragmented data, the same
unnamed outcome, the same absent owner — and adds a contract, an implementation program, and a sunk cost
that makes the eventual reckoning more expensive. Buying accelerates a program that knows where it is
going. It does not tell a program where to go.
Running more pilots
The second remedy is volume: if one in twenty pilots produces impact, run more pilots. This is the most
expensive misreading of the data available, because it treats a five percent success rate as a sampling
problem rather than a structural one. The failures were not random. They clustered at four specific
breakpoints, and running more experiments through the same four gates produces the same attrition at
greater cost — plus an organization that is now visibly tired of AI.
There is a second-order cost too. MIT's research found more than half of generative AI budgets flowing
to sales and marketing, while the strongest returns showed up in back-office work — eliminating
outsourced processing, reducing agency spend, compressing operational
cycles.1 A portfolio built from whichever pilot found a
sponsor drifts toward the most visible functions rather than the most valuable ones. More pilots
amplifies that drift.
Waiting for the models to get better
The third remedy is patience: the technology is moving quickly, so wait for the next generation to close
the gap. This is the only one of the three that costs nothing today, which is exactly what makes it
attractive and what makes it dangerous. The four breakpoints in Chapter 02 are not model limitations.
A better model does not integrate your claims system, name your target outcome, hire your operator, or
decide what an agent is permitted to do without a human in the loop.
Meanwhile the clock is real. Gartner projects that by 2028, 15 percent of day-to-day work decisions will
be made autonomously — from effectively zero in 2024 — and that a third of enterprise software will
include agentic capability.3 The organizations that have solved
their data access, their ownership model, and their governance latency will absorb that shift as an
upgrade. The organizations that waited will meet it as a migration.
Figure 5 · The reframe
Stop asking which AI projects to fund. Start asking which decisions the business needs to make
faster, clearer, and more defensibly — then fund the capability those decisions require.
Skip Vanderburg
A finish line appears
A decision has an owner, a cadence, and a moment it is made. That gives the work a definition of done that a capability never has.
Value becomes measurable
Cycle time, rework, and the cost of delay on a named decision are all countable before and after — which is what a CFO means by evidence.
Governance gets a subject
Controls attach to a specific decision and its blast radius, instead of to AI in the abstract — which is why they can be written in weeks rather than quarters.
Figure 5. The single change that separates programs that scale from programs that repeat: the unit of planning stops being a technology and becomes a decision.
04
The exit: six moves
Escaping pilot purgatory is not a bigger version of what stalled. It is a change in what the portfolio
is made of.
The single most useful finding in the 2026 McKinsey data is what separates the 6 percent of high
performers from everyone else. It is not model choice, spend, or headcount. Nearly three-quarters of
high performers — 72 percent — had fundamentally redesigned workflows around AI, against 25 percent of
the rest.2 They did not add intelligence to the way the work
already ran. They changed how the work runs, which means they changed how decisions get made.
Figure 6 · Two ways to hold the same budget
Switch the view. The money is the same in both. The accountability is not.
Figure 6. A pilot portfolio is a list of things being tried. A decision portfolio is a list of choices the business has agreed to make better, each with an owner, a cadence, and a measurable delta.
Converting one into the other is a sequence, and the order matters more than the speed. These are the
six moves I run with leadership teams, and they map directly to the six phases of the Decision Velocity
Framework.
Figure 7 · Six moves out of purgatory
01 · FRAME
Name the decisions
List the ten to fifteen recurring choices that actually move the P&L — pricing, adjudication, allocation, triage, sequencing. Write each as a question with an owner and a cadence. This list, not the technology inventory, becomes the portfolio.
02 · DIVERGE
Widen before you narrow
For each decision, generate the full set of ways it could be improved — better data, a model, a changed policy, a removed approval step. Several will not need AI at all. Finding those early is what keeps the AI work credible.
03 · SCORE
Score on one sheet
Value, confidence, effort, and risk — scored the same way across every candidate, by the people who own the outcome. The point is not the arithmetic. It is forcing four functions to argue against one set of numbers instead of four narratives.
04 · SORT
Sequence against readiness
Order the survivors by the conditions they need, not by enthusiasm. A decision whose data is three quarters away goes behind one whose data is already governed — no matter how attractive the demo was.
05 · CUT
Kill the rest, visibly
Stopping work is the move most programs skip, and it is the one that buys everything else. Announce what is being stopped and why. An organization that has never seen an AI project cancelled does not believe the criteria are real.
06 · PROVE
Instrument the decision
Baseline the cycle time, rework rate, and cost of delay before anything ships, and publish the delta on the same cadence as the decision itself. Evidence produced monthly is what converts a pilot budget into an operating line.
Framename the decision
Divergewiden the options
Scoreone sheet, one scale
Sortsequence by readiness
Cutstop the rest
Provepublish the delta
Figure 7. The six moves are deliberately unglamorous. None of them is an AI problem, which is precisely why AI programs stall without them.
A portfolio of pilots asks the organization to be patient. A portfolio of decisions asks it to be
accountable. Only one of those survives a budget cycle.
Governance belongs inside this sequence, not after it. The Deloitte finding — 21 percent with mature
agentic governance while deployment
accelerates7 — describes organizations writing controls for AI
in general, which cannot be finished. Controls written for a specific decision can: what the system may
decide alone, what it must escalate, what evidence it records, and who reviews the exceptions. That is
a two-week document per decision rather than a two-quarter policy program, and it is what closes the
latency gap that shadow AI is exploiting.
05
What this looks like by sector
The four breakpoints are universal. Which one bites first is not — it depends on how the industry is
regulated, how its data was accumulated, and where its margin actually sits.
Figure 8 · Where the stall shows up first
Healthcare
Providers & payers
The stall is almost always at Control. Clinical and privacy review is rigorous and slow by design, so ambient documentation and triage pilots pass their evaluations and then wait months for an approval pathway that was never designed for a system that learns.
Meanwhile the data breakpoint hides behind an integrated EHR: one system of record is not the same as one governed, retrievable, lineage-tracked corpus across imaging, notes, claims, and scheduling.
Opening move · Pick one decision with a bounded blast radius — documentation burden, prior-authorization triage, discharge sequencing — and write its control document before the build, not after.
Insurance
Carriers & brokers
The stall is at Aim. Underwriting, claims, and distribution each run their own experiments against their own definition of success, so nothing aggregates to a loss-ratio or expense-ratio number the CFO recognizes.
This is the sector where the decision framing pays back fastest, because the decisions are already named, already measured, and already have owners: what to bind, what to price up, what to adjudicate automatically, what to investigate.
Opening move · One prioritized roadmap across underwriting, claims, and service, expressed in leakage and cycle time rather than in models deployed.
Financial Services
Banking & capital markets
Model risk management is the strongest asset and the tightest constraint. Institutions here have decades of practice validating models — and a validation process built for statistical models with stable inputs, not for systems whose behavior changes with a prompt.
The result is a two-speed organization: heavily governed production models alongside a large, undocumented population of assistant use in the front office. That is the shadow AI pattern in its most regulated setting.
Opening move · Extend model risk management to cover generative and agentic behavior explicitly — tiered by decision impact — so the sanctioned path is faster than the unsanctioned one.
High Tech
Software & platforms
Capability is rarely the constraint; conviction is. Engineering adoption is genuine — roughly one in five organizations is scaling coding agents enterprise-wide, and 32 percent have declined a software purchase because they intend to build it instead.2
The stall shows up as productivity without portfolio: individual velocity rises, the roadmap does not change, and the gain is absorbed rather than captured. This is the Aim break wearing a friendly face.
Opening move · Convert engineering throughput into a roadmap decision — what ships sooner, what gets built that was previously out of reach — before the gain disappears into baseline expectation.
Manufacturing & Industrial
Operations & supply chain
The stall is at Ground. Plant data is abundant and almost never governed: historians, MES, quality systems, and maintenance logs that were built per site, per vendor, per decade. A pilot that works in one plant meets a different schema in the next.
Because the returns are operational rather than customer-facing, they also compete poorly for attention against front-office pilots — the exact misallocation MIT identified.1
Opening move · Prove one decision — changeover sequencing, expedite selection, maintenance triage — at two sites rather than one, so the second site pays for the data work the first exposed.
Figure 8. Select a sector. The diagnosis is structural, but the order of attack is industry-specific — and starting with the wrong breakpoint is how well-run programs lose a year.
06
The first ninety days
If your organization is somewhere in the 95 percent, the way out does not begin with a new model, a new
platform, or a new team. It begins with a decision inventory and a willingness to stop things.
Figure 9 · A ninety-day exit
1
Days 1–15 · Inventory both sides. Every AI initiative currently funded, with its sponsor and its stated outcome. Beside it, the ten to fifteen recurring decisions that move the P&L. The gap between the two lists is the diagnosis, and it is usually visible in a single meeting.
2
Days 16–35 · Score once, together. One scoring sheet, one scale, the four functions that will have to live with the answer in the same room. Value, confidence, effort, risk. What comes out is not a ranking so much as an agreement about what the organization believes.
3
Days 36–50 · Cut, and say so. Stop the initiatives that do not attach to a named decision. Publish the list and the reasoning. This is the move that converts a scoring exercise into a governing one, and it is the one most leadership teams flinch from.
4
Days 51–90 · Prove one decision end to end. One decision, one owner, one control document written before the build, one baseline captured before anything ships. Report the delta in cycle time and cost of delay at the decision's own cadence — weekly if it is made weekly.
Figure 9. Ninety days is enough to leave purgatory with one decision governed and proven. It is not enough to transform the enterprise — and programs that promise the second usually fail to deliver the first.
Notice what this sequence does not require. No new platform. No reorganization. No hiring plan that
takes two quarters to fill. It requires an executive team willing to name a small number of decisions,
agree on how they are scored, stop the work that does not serve them, and publish evidence on a
cadence. The technology, for once, is the easy part.
The 5 percent are not better at AI. They are better at deciding.
That is the uncomfortable and ultimately encouraging conclusion of every study cited in this paper. The
differentiator McKinsey found was workflow redesign. The differentiator MIT found was whether the system
was embedded in how work actually happens. The reasons Gartner gives for cancellation — unclear value,
uncontrolled cost, inadequate controls — are all failures of decision, not of engineering. None of these
are things a vendor can sell you, and all of them are things a leadership team can start on this
quarter.
Agentic systems raise the stakes on all of it. A generative pilot that has no owner produces a document
nobody reads. An agent that has no owner takes actions nobody reviews. The organizations that spend the
next year building the habit of naming decisions, governing them at the decision level, and proving the
delta will be the ones able to hand real authority to a machine when it is worth handing over. The rest
will still be running pilots.
Pilot purgatory is not where AI programs go to fail. It is where organizations wait until they decide
what they want.
◆
Sources
Every figure in this paper is drawn from named research published between mid-2025 and mid-2026.
Where a study is widely paraphrased — the 95 percent figure above all — the original methodology is
noted so the claim can be read at its actual strength.
01
MIT / Project NANDA, The GenAI Divide: State of AI in Business (2025). 150 leader interviews,
350 employee survey responses, and analysis of 300 public generative AI deployments; finds ~95% of
pilots produced no measurable P&L acceleration, purchased tools succeeded ~67% of the time versus
roughly a third as often for internal builds, and more than half of budgets went to sales and
marketing while back-office work delivered stronger returns.
Reported coverage.
02
McKinsey & Company, The State of AI: Global Survey 2026. Fielded 4 May – 8 June 2026;
1,719 respondents across 97 nations. 44% scaling AI enterprise-wide (from 38%), 37% reporting EBIT
contribution (unchanged), 6% high performers at ≥5% of EBIT, 72% of high performers having
fundamentally redesigned workflows against 25% of others, ~20% scaling coding agents, 32% declining a
software purchase in favor of building.
mckinsey.com.
03
Gartner, Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027
(25 June 2025). Cancellation drivers: escalating costs, unclear business value, inadequate risk
controls. Also: ~130 of thousands of claimed agentic vendors assessed as genuine (“agent washing”);
15% of day-to-day work decisions made autonomously by 2028, from 0% in 2024.
gartner.com.
04
Cloudera with Harvard Business Review Analytic Services (surveyed October 2025; published
5 March 2026). 230+ respondents involved in AI data decisions: 7% call their data completely ready
for AI, 27% not very or not at all ready, 56% cite siloed data and integration difficulty as the top
obstacle, 73% say data quality deserves more priority than it gets.
cloudera.com.
052026 State of the CIO survey, reported by CIO (30 April 2026). 40% of respondents name a lack
of in-house talent as the top challenge to implementing their AI strategy, with the scarcity
concentrated in production-scale and AI-security skills rather than model development.
cio.com.
06
BlackFog shadow-AI research (published 29 January 2026); 2,000 employees at organizations with 500+
staff. 49% use AI tools without employer approval, 51% have connected an AI tool to a work system
without IT approval, 58% use free consumer tiers, 33% have shared internal research or datasets, and
69% of C-suite respondents accept unsanctioned use.
Reported coverage.
07
Deloitte, AI agents are scaling faster than their guardrails (24 April 2026); 3,235 business
and IT leaders across 24 countries. Only 21% report mature governance models for agentic AI, with
decision boundaries, real-time monitoring, and audit trails identified as the most common gaps.
deloitte.com.
Where this goes next
Which decisions would you fix first?
This paper is drawn from The Agentic Enterprise: From Automation to Autonomous Decision
Intelligence. If your AI portfolio is full of promising work that has not reached the P&L, the
first useful conversation is not about models — it is about which decisions the work was supposed to
improve. That is the conversation I have most often.