July 23, 2026
The Missing Unit in the New Economics of AI
Assiduity AI
McKinsey’s series on the new economics of AI makes an important shift. It moves the discussion away from the price of intelligence and toward the economics of the work that intelligence performs. Tokens determine part of the bill. They do not determine whether an AI system created value. Model prices can fall while total spending rises. Agents can take different paths, use different tools, retry actions, and consume radically different amounts of compute while attempting the same task. The meaningful unit is becoming the completed business outcome. That is the right direction.
Across the series, McKinsey makes four related arguments:
- AI should be evaluated at the level of business outcomes rather than token consumption.
- The trajectory taken by an agent may contain information that the final output conceals.
- Autonomous systems require defined mandates, budgets, owners, and stopping rules.
- The ability to govern machine work will influence operating models and the boundary between work performed inside and outside the firm.
Together, these ideas describe the beginnings of a management discipline for machine work.
But they also expose an unresolved economic question. A completed outcome can create value without having been produced in a way the enterprise can accept.
A valuable outcome is not necessarily a governed outcome
Consider an AI system preparing a risk recommendation. The final memo may be clear, accurate, and commercially useful. It may even support a profitable decision. But the analysis may have omitted a required constraint, lost a specified fact, or developed in a direction increasingly distant from the evidence it was required to preserve. The outcome may still look correct. The process may still be unacceptable.
A profitable decision does not prove that the supporting analysis was authorized or reproducible. A correct answer does not establish that the process used to produce it remained within the organization’s mandate.
Enterprises therefore need to answer two separate questions:
Did the work create value?
And:
Was the work performed in a way the organization could authorize and accept?
The first is an outcome question. The second is a governance question. The economics of enterprise AI needs both.
The trajectory is part of the product
One of the strongest observations in McKinsey’s series is that agentic systems cannot be understood by inspecting only their final outputs. Agents choose tools, construct intermediate prompts, revise plans, abandon approaches, and retry actions. Two executions of the same task may arrive at similar endpoints through materially different paths.
In McKinsey’s interview with David Tepper of Pay-i, the trajectory is described as wide while the endpoint is narrow. A failure can begin several steps before completion, and a final-output evaluator may miss it because the last step conceals the earlier deviation. That observation has consequences beyond technical evaluation. Once the trajectory matters, it becomes part of the economic object being purchased.
The enterprise is not accepting only a document, recommendation, classification, or completed transaction. It is accepting a sequence of machine choices that produced the result. Those choices affect reliability, authorization, reproducibility, verification effort, and whether the organization can assume responsibility for the work.
Most current trajectory analysis is retrospective. Organizations collect successful and unsuccessful traces, connect them to business outcomes, and use the resulting evidence to evaluate or improve future performance. That is useful.
But visibility into a completed trajectory is not the same as influencing the trajectory while it develops.
Observability tells the enterprise what happened. Evaluation tells it whether the system succeeded.
Generation-time control addresses a different question:
While the work is still developing, which available path remains most aligned with the operating mandate?
McKinsey names the mandate but leaves it unmechanized
McKinsey does not overlook the importance of mandates. Its lead article states that autonomous systems should not operate without a defined mandate, budget, and stopping rule. The unresolved issue is not whether a mandate is necessary.
It is how that mandate remains active during execution. A mandate stated at the beginning of a task can lose influence as the work becomes longer, the context becomes more complex, and intermediate decisions accumulate. A system can begin with the correct instruction and still drift from it. A required constraint may be present in the initial instruction but lose influence as the work develops.
The missing mechanism sits between design-time policy and outcome-time evaluation.
McKinsey’s sequence is broadly:
Define the mandate → run the system → measure the outcome → inspect the trajectory
The additional control problem is:
Define the mandate → measure alignment while the work develops → influence the emerging path → preserve the trajectory as evidence
That is a narrower claim than saying the series misses governance. It does not.
The series identifies the governance requirement. What remains largely unspecified is how the mandate becomes an active property of execution rather than a static instruction placed at its beginning.
Context is valuable. It is not control.
McKinsey argues that as foundation models converge, competitive differentiation will increasingly shift toward proprietary data and context. Customer signals, operating data, prior decisions, policies, and process telemetry may allow otherwise similar AI systems to produce materially different results. That is persuasive.
But possessing context does not mean that context governs the work. Most enterprises already possess extensive organizational knowledge: policies, procedures, risk limits, precedents, delegated authorities, and escalation rules. The difficult problem is not merely making those materials available to an AI system. It is converting selected elements into an operating mandate that remains measurable and influential while the system works.
Placing policies and procedures in a context window does not establish that they constrained the resulting work. Retrieval makes organizational knowledge available; it does not prove that the relevant requirement remained influential as execution developed. Context is an input. A mandate is a control structure.
The competitive advantage will not come solely from possessing more proprietary information. It will come from turning organizational context into governed execution.
The real denominator is cost per governed outcome
McKinsey’s series also correctly moves the economic discussion beyond inference expense. Agentic work includes the cost of refinement, checking, exception handling, repair, and reverification. The same task can produce materially different total costs depending on the path the system takes.
The Pay-i interview describes verification and rework as an “agency tax.” Work that appears inexpensive at the inference layer may become unattractive after human review and remediation are included. That suggests a more useful denominator. Enterprises should not optimize only for cost per token, cost per model call, or cost per completed task.
They should increasingly optimize for:
Cost per governed outcome
That includes:
- inference,
- orchestration,
- generation-time control,
- human verification,
- rework,
- exception handling,
- and the expected cost of an undetected failure.
This changes how the economics of control should be assessed.
A control mechanism is not inefficient merely because it adds computation. Its relevant economic value lies in the review, remediation, and failure exposure it removes.
The question is not:
Did control make generation more expensive?
The relevant question is:
Did control reduce the total cost of producing work the enterprise could accept?
A higher generation cost can be economically rational if it materially reduces expert review, silent failure, or the probability that completed work must be reconstructed or discarded. The correct comparison is not controlled inference against uncontrolled inference. It is governed work against apparently cheaper work that still requires substantial human effort to determine whether it can be trusted.
The structure of agentic cost is what makes this comparison favorable. McKinsey and Pay-i report that in the agentic coding workflows where it has been measured, refinement — checking, repairing, and re-verifying — accounts for on the order of 60 percent of a task’s total cost, with first-pass generation the minority. The exact share varies by workload, but the direction is consistent: verification dominates creation.
Generation-time control adds cost to the smaller pool and targets the larger one. It raises the cost of first-pass generation — in Assiduity’s case by a bounded, measured amount — in order to reduce the refinement and re-verification that follows. Because the cost it adds falls on the minority of task cost and the cost it removes falls on the majority, the trade is favorable whenever control eliminates even a modest fraction of rework. The size of that fraction is an open empirical question; the asymmetry that makes the trade worth taking is not. It follows from the cost structure itself.
What Assiduity has demonstrated
Assiduity is designed around the gap between a stated mandate and a completed outcome. The system represents selected facts, constraints, and evidence requirements as a semantic operating contract. During generation, it evaluates alternative continuations and selects the path that remains most closely aligned with the specified anchors and constraints. The broader architecture can also define conditions under which developing work should be routed for review or escalation. Those functions are distinct from the alignment results reported below.
Assiduity’s current evidence supports a bounded claim. The bounded overhead referred to above is measured. In long-form generation tests across multiple models and domains, sparse generation-time branching occurred on approximately 12 percent of steps, holding average computational overhead to roughly 1.36× — compared with the approximately 4× cost of generating four candidates at every step. That 1.36× is the cost added to first-pass generation; it is the number that must be weighed against the refinement it is intended to displace.
The controlled runs demonstrated stronger retention of specified semantic anchors, while comparison and placebo tests indicated that the result came from the selection mechanism rather than branching alone. This does not prove that every possible legal, policy, or operational requirement has been satisfied. Semantic alignment is not the same as comprehensive compliance certification.
The present evidence demonstrates something narrower but still important:
Specified elements of a mandate can remain measurable and influential while work develops, rather than being checked only after the output is complete.
The resulting trajectory can show where measured divergence emerged and which alternatives were considered, giving reviewers evidence they can use to determine whether intervention or escalation is warranted.
That is different from placing another judge at the end of the process. A final-output evaluator determines whether the completed result appears acceptable. Generation-time control helps shape the process that produces it.
From buying intelligence to buying governed work
The larger implication concerns the boundary of the firm. McKinsey argues that agentic economics will lead enterprises to reconsider which activities should remain internal and which can be sourced externally. Research, analysis, content production, and other forms of knowledge work may increasingly be delivered through agentic services. That future is plausible.
But outsourcing machine work is economically credible only when the buyer receives more than a finished output. Traditional outsourcing works because execution can be specified, tested, inspected, contractually bounded, and supported by evidence. The buyer does not need to employ every person performing the work. It still needs a basis for accepting the result.
Machine work will require a similar institutional structure. An enterprise may outsource execution. It cannot outsource accountability merely because the executor is artificial. The commercial product will therefore need to include the completed work, evidence of how specified controls influenced its production, visibility into material deviations, and a record sufficient to support review and acceptance.
This is the idea behind governed execution as an attested outcome.
Attestation here does not mean that every legal or policy requirement has been independently certified. It means that the output is accompanied by inspectable evidence about how specified elements of the operating mandate influenced its production. The product is no longer simply access to a model. It is not even access to an agent. It is machine work delivered with a basis for organizational acceptance.
The next economics of AI
McKinsey’s series correctly reframes the economic problem. Tokens are not value. Model access is not productivity. Falling inference prices do not guarantee falling operating costs. And AI cannot be managed at scale without new disciplines for allocation, measurement, sourcing, and governance.
The next step is to recognize that a completed business outcome is necessary but not sufficient. Enterprises need to know not only whether machine work produced value, but whether the organization can accept responsibility for how that value was produced. Tokens measure the consumption of intelligence. Outcomes measure the value produced by intelligence. Governed execution measures whether specified elements of the operating mandate remained active while the work developed.
The new economics of AI will not end with cost per token or return on intelligence. It will increasingly be built around the cost, reliability, and transferability of governed machine work. The missing unit is the governed outcome.