Whitepaper 05 · The Agentic Enterprise series

The Future of AI

Generate, assess, score, prioritize — the four machine capabilities that decide what an enterprise does next

Start reading
01 02 03 04 Generate Assess Score Prioritize impact, cost, risk against capacity the decision a human still makes
Figure 1. Four capabilities, running continuously. The machine carries the volume; the judgment at the center stays where it belongs.
It is not the strongest of the species that survives, but the one most responsive to change.
— commonly attributed to Charles Darwin
• • •
01

The decision load

Artificial intelligence is moving from a supportive tool to an integrated system that touches every part of how a business runs. The reason is not that the models got clever. It is that the decision load outgrew the people carrying it.

Every enterprise now generates more signal than its leadership can process: telemetry from products, transcripts from service, movement in supply and price, regulatory change, competitive noise, and the internal exhaust of a dozen systems that were never designed to be read together. The volume did not arrive gradually. It arrived faster than the decision-making structures built to absorb it, which were designed for quarterly cycles and a manageable number of options on the table.

The result is familiar to any executive team. The list of things the business could do is longer than it has ever been, the evidence behind each one is thinner than anyone would like, and the meeting where it all gets decided has not grown any longer. Something has to carry the volume.

Figure 2 · What the machine is actually for
0

Of organizations are now scaling AI enterprise-wide, up from 38 percent a year earlier1

0

Report any EBIT contribution from that AI — a figure that has not moved1

0

Of the high performers fundamentally redesigned how work runs, against 25 percent of everyone else1

0

Of day-to-day work decisions will be made autonomously by 2028, from effectively zero in 20242

Figure 2. Adoption is not the constraint and capability is not the constraint. What separates the two groups is whether the machine was pointed at how decisions actually get made.

Read those four numbers together and a pattern appears. Deploying AI is now ordinary. Getting it to change a business result is not, and the organizations that manage it are the ones that redesigned the work rather than decorating it. Meanwhile the horizon is moving: Gartner expects roughly 15 percent of day-to-day work decisions to be made autonomously by 2028, from a base of essentially zero.2 The question stops being whether machines participate in decisions and becomes which parts of the decision they are given.

Four capabilities, one loop

Strip the technology conversation back to what a decision actually requires and four capabilities fall out. Options have to be generated. They have to be assessed against evidence. They have to be scored so that unlike things can be compared. And they have to be prioritized against a capacity that is always smaller than the ambition.

Humans have always done all four. What changes is the scale at which the first three can be run, and how often the fourth can be revisited. That is the whole of the shift, and the rest of this paper takes the four in turn.

Figure 3 · The four capabilities
Generateoptions at machine scale, from signal no team could read
Assessdemand, feasibility, cost, risk, expected return
Scoreone scale, so unlike things can be compared
Prioritizeagainst real capacity, and re-cut when it changes
Figure 3. The order matters. Scoring without generation narrows the field prematurely; prioritizing without scoring is just seniority in a spreadsheet.
The machine's job is to make the volume tractable. The human's job is to decide. Confusing the two is how organizations end up automating the wrong half.
02

Generate

The first capability is the one most organizations underuse: continuous, structured option generation from signal no human team could read in a quarter, let alone a week.

Consider what a system with access to the right data can now do. It reads global trend data, industry shifts, customer behavior, competitor movement, and internal performance telemetry, and proposes new products, business models, and process changes against them. Not once at the offsite — continuously, as the underlying signal moves.

The instinct is to be suspicious of this, and the suspicion is half right. Generation at scale is worth nothing if what comes out is random. The value depends entirely on the constraint the generation runs against: options have to be produced against a named decision and an explicit objective, or the organization simply drowns in a larger pile of ideas than it had before.

Figure 4 · Where options come from
SIGNAL Market & trend data Customer behaviour Product telemetry Service transcripts Supply & cost movement Regulatory change Generation bounded by objective Candidate options, each attached to a decision The unglamorous ones — retrieval, assembly The agentic ones — synthesis, monitoring An option with no decision behind it does not enter the portfolio, however impressive the demo.
Figure 4. Generation is only as good as its constraint. Widen the inputs; narrow the objective.

Done well, this surfaces the opportunities an organization would never have reached on its own — not because nobody was clever enough, but because the pattern sat across four systems and three functions that never meet. Done badly, it produces a thousand plausible sentences and a leadership team that trusts the exercise less than it did before.

Generate widely. Bound tightly. Those are not in tension — the second is what makes the first safe.

03

Assess

The second capability is evaluation: taking a long list of candidates and establishing, quickly and consistently, what each one would actually involve.

Assessment asks five questions of every option — is there demand, can we operationally do it, what does it cost, what does it risk, and what would it return. Human teams answer these well and slowly, one initiative at a time, with the depth of the answer depending on who happened to be in the room. A machine answers them consistently across hundreds of candidates in the time it takes to schedule the meeting.

Consistency is the underrated half of that sentence. The value is not only speed; it is that every option gets asked the same questions with the same rigour, which is precisely what does not happen when assessment depends on which executive is sponsoring which idea.

Figure 5 · Five questions, asked of everything
01

Demand

Is there evidence anyone wants this — measured in behaviour, not in enthusiasm from the team proposing it?

02

Feasibility

Can this organization actually run it — with the data it has, the systems it runs, and the people it employs today?

03

Cost

The full cost — build, change management, governance, and the ongoing operation that arrives after launch.

04

Risk

What breaks if this is wrong, who is exposed, and what the organization is obliged to prove afterwards.

05

Return

The expected effect on a number the business already tracks — stated before the work starts, not reconstructed after.

06

And the timing

Assessment has a shelf life. Anything evaluated against last quarter's conditions is a historical document, not a recommendation.

Figure 5. The same five questions, asked of every candidate, are worth more than a deep answer to one of them for the initiative with the loudest sponsor.
Assessment at machine speed is not about answering faster. It is about answering evenly — which is the thing human assessment almost never does.
04

Score

Assessment tells you what each option involves. Scoring is what lets you compare things that have nothing else in common.

This is the step organizations skip, and skipping it is why AI portfolios end up sorted by seniority. A claims automation project and a customer service assistant have no shared unit — different functions, different sponsors, different definitions of success — until someone imposes one. The imposition is the point. A score is not a verdict; it is a common language that makes an argument possible.

The formulation I use is Decision RICE: Reach × Impact × Confidence ÷ Effort, recast so that reach counts decisions per quarter rather than users, confidence blends evidence quality with data readiness, and effort includes the change and governance work rather than just the build. Machines can compute it continuously as inputs move; what they cannot do is choose the criteria, which remains an executive act.

Figure 6 · Score your own initiative against a live portfolio
600
Medium
70%
5 pm

Reach is decisions touched per quarter. Effort is person-months of build, change, and governance. Move a slider and watch the ranking answer back.

Figure 6. The arithmetic is deliberately simple. Its value is that four functions have to argue against one set of numbers instead of four narratives.

Two things become obvious the moment a portfolio is scored this way. High-reach, low-impact work — summarization, assembly, retrieval — often out-scores the ambitious agentic initiative everyone came in excited about, because it touches so many decisions at such low cost. And confidence does more damage than any other input: an initiative with genuinely uncertain data readiness cannot buy its way back with a strong impact estimate.

Score ranks. It does not decide. A high-reach, low-impact initiative can top the chart and still be the wrong thing to fund this year.

Which is why scoring is the third capability and not the last one.

05

Prioritize

Prioritization is where the previous three capabilities either become an operating decision or stay an analysis. It is also the only one of the four that is fundamentally about scarcity.

A ranked list is not a plan. A plan is a ranked list met with a capacity — the engineering months, the change bandwidth, the governance attention an organization actually has this quarter — and a decision about where the line falls. Everything above the line gets funded, staffed, and governed. Everything below it is not a backlog and not a phase two; it is a list of things the organization has agreed not to do while the funded work finishes.

Machines make this dynamic in a way it has never been. Capacity changes, a supplier moves, a regulation lands, a competitor ships — and the line can be redrawn against current conditions rather than the conditions that held at the last planning offsite. That is the real prize in the fourth capability: not faster ranking, but the ability to re-cut without waiting for the calendar.

Figure 7 · Where the line falls
Capacity this quarter · 18 person-months
4 funded · 2 on the avoid list
Funded — fits the capacity Avoid list — explicitly not this quarter Capacity line

Drag the capacity. Notice what has to leave the plan before anything new can enter it.

Figure 7. Prioritization is subtraction. The organizations that never practise it are the ones whose AI portfolios grow every quarter and deliver every other one.

Dynamic does not mean continuous. Re-cutting a portfolio every week produces whiplash and destroys the conditions under which anything finishes; re-cutting once a year means running the business against conditions that expired months ago. Quarterly, with a standing mechanism to force a re-cut when something material moves, is the cadence I see working.

06

Putting it to work

The four capabilities are worth nothing as a diagram. What follows is the sequence I run with leadership teams, and the order is load-bearing.

Figure 8 · Four steps, in this order
1

Connect the machine to real data. Generation and assessment are only as good as their reach into CRM, ERP, finance, service, and operations. This is the least glamorous step and the one that determines everything downstream — an assessment layer running on a curated extract will produce confident answers about a business that does not exist.

2

Name the objective before generating anything. Market expansion, cost reduction, cycle-time compression, risk posture — the objective is the constraint that turns generation from noise into options. Teams that skip this step get volume and mistake it for progress.

3

Automate assessment and scoring, not the choosing. Let the machine evaluate and rank continuously against the criteria you set. Keep the criteria, the weights, and the final cut in human hands — that is the part your board is actually holding you accountable for.

4

Make prioritization dynamic, and give it a cadence. Build the mechanism that re-cuts the portfolio when conditions move, then decide how often it is allowed to fire. Dynamic without cadence is churn; cadence without a mechanism is an annual planning cycle wearing new language.

Figure 8. Four steps. None of them is a model selection decision, which is usually the first thing an organization wants to talk about.

What this changes for the people in the room

The honest description of this shift is not that AI decides. It is that the preparation for a decision — the assembling, the comparing, the ranking, the endless reconciliation of numbers from four systems — stops consuming the week before the meeting. What arrives at the meeting is a ranked, evidenced, capacity-tested set of options with the assumptions visible.

Executives then spend their time on the part machines are worst at: deciding what the organization is for, what it will not do, and which uncomfortable trade-off it is prepared to accept. That work does not get automated. It gets clearer, and it gets faster, because the arithmetic underneath it was already done and everyone can see it.

The constraint was never intelligence. It was throughput on the way to judgment.

Where to start

Start with one decision your business makes repeatedly and expensively — claims routing, pricing moves, capacity commitments, portfolio investment, hiring sequence. Run the four capabilities against that one decision end to end: generate the options, assess them evenly, score them on one sheet, and cut against real capacity. Measure the cycle time before and after.

One decision, done completely, teaches an organization more than a dozen pilots aimed at capabilities. It also produces the thing every subsequent conversation needs: evidence, from your own business, that the loop works.

Generate widely. Assess evenly. Score honestly. Prioritize ruthlessly. Then do it again next quarter, against the conditions that actually hold.

Sources

Figures cited in this paper come from named research published between mid-2025 and mid-2026, with the methodology noted so each claim can be read at its actual strength.

  1. 01 McKinsey & Company, The State of AI: Global Survey 2026. Fielded 4 May – 8 June 2026; 1,719 respondents across 97 nations. 44% scaling AI enterprise-wide (from 38% a year earlier), 37% reporting EBIT contribution (unchanged), 6% qualifying as high performers at ≥5% of EBIT, and 72% of those high performers having fundamentally redesigned workflows against 25% of other organizations. mckinsey.com.
  2. 02 Gartner, Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (25 June 2025). Includes the projection that 15% of day-to-day work decisions will be made autonomously by 2028, from 0% in 2024, and that a third of enterprise software will include agentic capability — alongside the cancellation drivers of escalating cost, unclear value, and inadequate risk controls. gartner.com.
  3. 03 MIT / Project NANDA, The GenAI Divide: State of AI in Business (2025). 150 leader interviews, 350 employee survey responses, and 300 public generative AI deployments; the source of the widely quoted finding that roughly 95% of pilots produced no measurable P&L acceleration. Examined in depth in Whitepaper 04.

Where this goes next

Which decision would you run this against first?

The four capabilities are easy to agree with and hard to sequence. If your team is working out where generation, assessment, scoring, and prioritization belong in your operating model — and which decision to prove them against first — that is the conversation I have most often.