The Fifth Measure in the AI Scorecard

Assiduity AI

The Fifth Measure in the AI Scorecard

An AI system produces 1,000 reports.

Reviewers classify 900 as ready to use. But 100 of those reports quietly omitted a required issue, weakened an important constraint, or failed to provide evidence the assignment required. Did the system complete 900 successful tasks? Or did it complete 800, accompanied by 100 latent control failures? The distinction matters because the economics of enterprise AI depend on what counts as success.

OpenAI recently proposed a better scorecard for the AI age. Instead of measuring adoption through seats purchased, active users, or tokens consumed, enterprises should measure the work accomplished and the full cost of completing it. That cost includes not only model usage, but also employee time, human review, retries, corrections, and rework.

The scorecard asks four questions:

  1. Is AI completing work that matters?
  2. What does each successful task cost?
  3. Can people depend on the result?
  4. Does each AI dollar produce more value as usage grows?

This is the right economic frame. Tokens have value only when they become useful work, and a cheaper model call is not cheaper if it creates more review, correction, or repetition.

But the scorecard needs a fifth question:

Did the work remain within the operating mandate, and can the organization show how it was controlled?

That is the difference between a successful task and a governed successful task.

The denominator determines the economics

OpenAI’s framework correctly shifts attention from the cost of generating an answer to the full cost of completing the task. For consequential enterprise work, however, the denominator needs an additional condition.

Cost per governed successful task = total cost of compute, control, review, retries, and rework ÷ tasks meeting quality, mandate, evidence, and escalation standards.

The numerator may be slightly larger because control is not free. But the denominator is more honest. A task should not count as successfully completed merely because the final output appears reasonable. It should count when the work satisfies the conditions under which the organization is prepared to accept it.

Return to the 1,000 reports. If 900 appear acceptable but 100 contain unrecognized mandate failures, the system did not produce 900 governed successful tasks. It produced 800, along with 100 problems that have not yet been priced. Those problems may later surface as additional review, corrections, retries, downstream rework, missed escalations, delayed decisions, control exceptions, or failures the organization must explain after the work has already been used.

A scorecard that counts only visible corrections will overstate both dependability and economic value. This is also why governance should not automatically be treated as overhead. Control creates value when its cost is lower than the review, rework, delay, and failure cost it removes. The relevant question is not whether control makes an individual generation more expensive. It is whether it lowers the total cost of producing work the organization can responsibly accept.

Ready to use according to whom?

OpenAI recommends tracking three practical outcomes:

These categories are much more useful than model accuracy alone because they connect AI performance to the amount of human effort still required. But they assume the organization can reliably determine which category an output belongs in. That is not always easy.

Consider an illustrative contract-review workflow. The point is not that every element of this example has already been benchmarked in legal deployment. It is to show the shape of a mandate failure that can remain hidden inside an otherwise credible result. The operating mandate requires the system to identify specified categories of nonstandard clauses, cite the relevant language for each finding, and route exposures above a defined threshold to a human reviewer. The resulting memo is fluent and professionally organized. It identifies several important provisions and reaches plausible conclusions.

But it omits one required clause. Another observation lacks the required source citation. A threshold condition that should have prompted human review passes without escalation. Nothing needs to be obviously false for the output to be ungoverned. The work may still look ready to use because the failure is not a malformed sentence or an absurd conclusion. It is the quiet disappearance of something the assignment required.

That creates a fourth operational category:

Appears ready to use, but departed from mandate.

It is the most difficult category because the work does not announce its own failure. The output is fluent. The omission is plausible. The reviewer has little reason to suspect that something required has disappeared. The work ships precisely because it looks finished.

Quality and mandate adherence are not the same measure

Mandate adherence is not a complete measure of output quality. A system might preserve all specified requirements and still produce weak analysis. It might temporarily move away from one requirement and later recover. A reviewer might disagree with a conclusion even when every required issue and source is present. The operating mandate measures whether defined requirements remain represented in the work. It does not settle every question of truth, professional judgment, or usefulness. That limitation is also what makes mandate adherence valuable.

Final quality is usually judged after the answer exists. Mandate adherence can provide a control signal while the answer is still developing and alternative paths remain available. It can indicate that a required subject is receiving insufficient treatment, an exclusion is losing influence, an evidence requirement remains unsatisfied, or the developing answer is moving away from the completion standard. It is a leading indicator, not a substitute for human judgment.

The goal is not to declare the output correct before it is finished. The goal is to prevent specified requirements from quietly disappearing before a reviewer ever sees the result.

“Done” has to remain active

OpenAI advises enterprises to begin with one workflow, define what “done” means, and measure the result in the system where the work occurs.

That is the right starting point. But consequential work is rarely defined by a single output instruction. “Done” may include:

Together, these conditions form an operating mandate.

A prompt asks a system to produce something. An operating mandate defines the conditions under which the organization is willing to accept the work. The control problem is keeping those conditions active as the work develops.

Generative systems produce work through a sequence of locally plausible choices. Each continuation may read well on its own while gradually shifting attention away from an earlier requirement. By the time the final answer reaches a reviewer, the path is complete. The reviewer must detect the omission, diagnose what happened, and return the work for correction. Generation-time control moves the control point earlier.

Assiduity translates the operating mandate into a semantic contract, evaluates candidate continuations against it while the work develops, and selects the path that remains closest to that mandate. The resulting work is accompanied by evidence showing how control was applied.

The model remains the generative engine. The control layer influences which developing path becomes the finished work.

Evidence is part of the deliverable

In consequential workflows, the final document is often only part of the work product. A credit recommendation needs supporting analysis. A compliance review needs evidence. A project decision may require a record showing that approval and escalation thresholds were followed.

The deliverable therefore consists of two things:

  1. the completed work; and
  2. the evidence required to review, accept, and defend it.

AI does not eliminate this requirement. It makes it more important because execution is moving into systems while accountability remains with people and organizations.

Traditional output review provides a snapshot of the final answer. It may reveal a visible error, but it does not necessarily show which mandate governed the task, whether required elements remained active, where material drift emerged, whether control affected the path, or why an exception did—or did not—reach human review.

Useful execution evidence should not drown reviewers in technical traces. It should help the accountable person answer a practical question:

Can I accept this work without reconstructing the entire process myself?

That is where governance becomes economically relevant. When fewer required elements disappear, reviewers spend less time finding omissions. When evidence accompanies the work, they spend less time reconstructing it. When genuine exceptions are surfaced appropriately, human attention can be reserved for decisions that actually require judgment.

The value is not surveillance of the model. It is lower-cost acceptance of the work.

The fifth measure

Not every AI interaction requires this level of control. Brainstorming, stylistic exploration, casual drafting, and easily reversible tasks may not justify it.

The fifth measure matters where the work has an explicit mandate, a plausible omission would be costly, and a person or organization remains accountable for the result.

In those settings, the enterprise AI scorecard should ask:

  1. Useful work: Is AI completing work that matters?
  2. Task economics: What does each successful task actually cost?
  3. Dependability: How much of the work can people use without correction?
  4. Scale economics: Does each AI dollar produce more value as usage grows?
  5. Governed execution: Did the work remain within mandate, and is there sufficient evidence to stand behind it?

The first four measures tell an enterprise whether AI is becoming more productive. The fifth tells it whether that productivity can be responsibly absorbed into the organization.

Assiduity addresses that fifth measure by encoding the operating mandate as a semantic contract, evaluating candidate continuations against it while the work develops, and preserving evidence of how control was applied.

The relevant enterprise unit is governed successful work: work that meets the quality bar, remains within mandate, escalates appropriately, and carries the evidence required for acceptance. Useful output is not necessarily governed output.

The AI scorecard should measure both.

Assiduity AI

Move Fast. Build Reliable.

Assiduity is building runtime control infrastructure for enterprise AI systems that need to stay aligned, auditable, and reliable during generation.