A token is not a unit of work
Token consumption is an input, not an output. How to budget machine cognition by cost per accepted unit of work: define the accepted unit, price consequence with crossover arithmetic from public benchmark data, and follow one case through the routing architecture.

The CFO has the invoice. The CTO has a dashboard of model calls, token counts, and error rates. Both can say exactly what the AI consumed last month. Neither can answer the question that decides whether the investment is working: what did one accepted result cost?
A token is not a unit of work. It is a unit of consumption. On one of the largest production AI gateways in June 2026, open-weight models carried 29% of all tokens on under 4% of normalized spend. The comparison is aggregated across the gateway's traffic at public list prices, so it is a statement about price dispersion rather than matched workloads, and as a statement about price dispersion it is enough: the cost of a unit of machine cognition varies by more than an order of magnitude depending on what runs where. A consumption number that elastic cannot stand in for output. Reporting token spend to a board is reporting hours typed instead of software shipped.
The moment to fix the metric is now, because the input is no longer experimental. Token volume on that gateway grew 29% in a single month. Corporate card spending data from American businesses puts the median firm near $12 per employee per month while the top tenth spends $650: a 54x spread before the extreme tail. In August, a vendor-neutral consortium of 30 companies launched for a single purpose: standardizing how organizations measure AI cost, consumption, and value. When an input gets its own standards body, it has stopped being an experiment. Before setting a 2027 number, map the input across API accounts, per-seat subscriptions, AI embedded in vendor products, and tools employees expense on their own.
The spending spread has produced a fear with a name: compute inequality, the idea that a minimum spend on machine intelligence decides who competes in every knowledge industry. Machine cognition is a production input, like talent, software, and capital, and an unbudgeted input is an unmanaged one. What a company gets back per unit of spend, however, is a property of architecture rather than of budget size, and architecture can be designed at any budget.
That design is a cognition budget: a managed allocation of human judgment, efficient throughput, and premium reasoning. What follows is the arithmetic for building one.
The arithmetic of acceptance
Start with the number that should replace the token bill: cost per accepted unit of work. Everything the unit consumed, model calls, routing, failed runs and retries, review time, rework, divided by the outputs that cleared the quality bar. Accepted is the load-bearing word. An output is accepted when it passes the evaluation set for its task class and, where consequence demands, an accountable human's review.
A public benchmark of agent routing supplies unusually complete numbers: 145 multi-step agent tasks, run three ways, with cost and accuracy both reported. Running the whole suite on a frontier model cost $11.45 at 86% accuracy. A routed configuration, an efficient model for most calls with escalation to the frontier, cost $3.00 at 80%. The efficient model alone cost $0.72 at 77.7%.
The headline read: routing cut cost by 74%. Now price the quality in. Per accepted task, the frontier configuration cost roughly 9.2 cents, the routed one 2.6 cents, the small model 0.6 cents. On cost per accepted unit, the cheapest configuration wins by a wide margin, which raises the obvious question: why not run everything on the smallest model?
Because the metric is missing its last variable: what a failure costs when it escapes. Give that a number and charge each configuration for its failures, and the ranking flips at computable thresholds. On this workload, the small model stays cheapest only while an escaped failure costs less than roughly seventy cents. Between that and roughly a dollar, the routed configuration wins. Beyond one dollar per escaped failure, the frontier configuration is already the cheapest option in the room, sticker prices notwithstanding. (The thresholds solve cost plus failure rate times escape cost across the three published configurations; figures rounded.)
Sit with how low those crossovers are. When a task costs cents, quality becomes economical almost immediately: a failure worth even a dollar tips the arithmetic toward the most capable configuration. The zone where the cheapest model wins outright is confined to work whose failures cost practically nothing. Your workflows need their own thresholds; the method transfers even where the numbers do not.
The same benchmark also shows why routing itself must be priced, not assumed. The routed configuration's escalation judge, the component deciding which tasks deserve the frontier, consumed about a fifth of total spend while producing none of the final work. Across five identical runs, frontier traffic ranged from 4.1 to 9.1% and total cost swung by 67%, while the routed configuration's accuracy gain over the efficient model alone sat inside run-to-run variance. Routing is machinery with its own cost and its own variance, and it earns its place only when measured against your acceptance bar, on your cases, with your evaluation sets.
Define the accepted unit first
None of the arithmetic works until the accepted unit is defined, and this is the step most budgets skip because the categories look obvious until someone prices them. A draft, an extraction, a classification, and an approved decision are different products; give them one metric and the cheapest incomplete output looks the most productive.
Take a generic invoice-exception workflow. "One invoice processed" is too loose; the fields can be extracted perfectly while the purchase-order mismatch that caused the exception slips through. A unit worth paying for reads more like: one exception resolved, invoice and purchase order matched, the applicable policy identified, the discrepancy classified, the evidence linked, and the accountable approver recorded. That sentence is a contract. Engineering can validate it, finance can price it, and operations can reject work that fails it without debating whether the model sounded convincing.
Every accepted-unit contract fixes the same few decisions:
| The contract fixes | Invoice-exception example |
|---|---|
| What has to be completed | One exception resolved and recorded |
| What must be correct | Fields match source documents and policy |
| What a failure costs | Payment delay, duplicate payment, policy breach |
| What evidence travels with the answer | Invoice, purchase order, policy clause, trace |
| Who may accept it | Automatic below policy limits, named approver above |
The same discipline transfers everywhere: a software team's accepted unit is a change that passes tests, security checks, and review; a compliance team's is an assessment with source-level evidence and a named approval. Decide what done means before measuring how cheaply machines can produce it.
Budget by consequence, not by volume
With units defined, the allocation resolves into three resources at three price points, and the most expensive one is the easiest to misbudget. Accountable humans carry judgment: the approvals, the exceptions, the decisions that carry someone's name. Human attention is the scarcest line in the budget, and it is a priced input: spent deliberately where consequence demands a signature, withdrawn from work a machine can verify. Open-weight models carry recurrent throughput: the screening, extraction, classification, and assembly that repeats daily against a defined bar. Frontier models carry consequential reasoning: the analysis and security-critical review where a better answer changes the outcome, and where a model's confidence does not carry accountability, which is why the human boundary stays fixed even when the model sounds certain.
We have written about the routing architecture that makes the assignment, task by task, on quality, cost, license, and provenance. Production traffic shows spend already concentrating where consequence lives: on the gateway above, back-office agents, the workload closest to money and compliance, consumed 5% of tokens but 14% of spend. Across paterhn deployments the same pattern is deliberate: open-weight models carry at least 60% of production agent tasks, each task class behind its own evaluation set, frontier spend reserved for the work that justifies it.
The difference between instinct and discipline is the acceptance bar. The benchmark above bought its saving by letting quality float six points downward. In production we run the trade in the opposite direction: the bar per task class is fixed by the evaluation set and the reviewer, and cost falls as far as the bar allows, not further. That is the posture behind our standing result that intelligent routing can cut inference cost by 40 to 60 percent: the saving is real precisely because acceptance, not consumption, is the number under management.
"A cognition budget is not a cap on spend. It is a design for where judgment, throughput, and reasoning each earn their keep. Fix the quality bar first and let the cost fall to meet it. A budget that lets quality float is not saving money; it is deferring the invoice." paterhn
One case through the router
Watch the invoice exception move, because the mechanism is where the economics live. The case arrives with an invoice, a purchase order, supplier history, and the current payment policy; intake classifies the exception type and sets its consequence class from amount and supplier status. An efficient model extracts the fields and proposes a classification; deterministic checks reconcile totals, currencies, and line items; retrieval attaches the policy clause the classification depends on. If everything reconciles and the amount sits below the automatic-closure limit, the case closes as one accepted unit, and no frontier token was spent.
Now give the purchase order two amendments with conflicting delivery terms. The first path cannot tell which amendment governs, the evidence check fails, and the case escalates with its documents, its failed attempt, and the contradiction attached. A frontier model resolves the amendment sequence and produces a new classification with the clauses linked. The amount exceeds the closure limit, so the case lands with the named approver, who sees the resolution, the trace, and the model path. Approval closes one accepted unit with its full cost attached, including the failed first attempt. A correction records what changed and why, and becomes a case in the evaluation set. The flow never asked which model is best in the world, only what this case needed to be accepted.
A budget that compounds
That closing step separates a cognition budget from a cloud bill: managed well, it improves itself. Every run leaves a record: which model handled the task, what it cost, whether it was accepted, what the reviewer corrected. Retained in architecture you control, that record moves your crossover points from estimates to measurements: evaluation sets sharpen, routing thresholds migrate to where quality actually breaks, and frontier spend concentrates on a smaller, harder core. The learning loop we have argued for since the beginning is what turns this quarter's budget into next quarter's productivity. Left inside a vendor's platform, the same record may still compound, but on the vendor's terms rather than yours; price that difference into every renewal.
This is also the honest answer to compute inequality. At the frontier of model training, scale decides, and most operating companies are not in that race. An operating company buys cognition and converts it into accepted work, and the conversion rate is set by allocation: the 54x spending spread is real, and nothing in it says the top spender converts best.
The cognition-budget review
Structure the number as three pools rather than one line. Routine capacity funds expected volume on the lowest-cost paths that have cleared evaluation. Exception capacity funds frontier escalation and accountable review, budgeted as a tested range rather than an average; the benchmark's 67% cost swing shows how sharply escalation frequency moves the total. Learning capacity funds the evaluation sets, observability, and evidence retention that keep the other two honest; starve it and the budget goes blind.
Then one working session with the people who own the money and the people who own the architecture, before the 2027 numbers lock. Five questions on the table:
| The session asks | What the answer must cover |
|---|---|
| What are we actually spending? | API accounts, per-seat subscriptions, AI embedded in vendor products, and the tools employees expense on their own |
| Which work carries consequence? | Workflows sorted into recurrent-and-verifiable versus consequential-and-judgment-bound, with a rough escape cost per class |
| Where is the acceptance bar, per class? | An evaluation set for every recurrent task class, a named reviewer wherever accountability lives |
| What does an accepted unit cost, all in? | Model spend plus review plus rework, over accepted outputs, per workflow |
| What do the runs leave behind? | Evidence and corrections retained in architecture you control, or exhaust escaping into a vendor's platform |
The estimates can stay rough; the crossover arithmetic only needs the order of magnitude, and card data captures only the visible layer of spend. Two answers do the most work. Cost per accepted unit is the number that should fall quarter over quarter; if only the token bill moves, the budget is measuring consumption. And a task class without an acceptance bar produces opinions, not work.
Starting requires no new model spend, some engineering time, and one decision: acceptance, not consumption, becomes the number under management.
Machine cognition is a production input; the 2027 cycle is the time to budget it as one. Define the accepted unit, budget by consequence, measure by acceptance, keep the learning. Three configurations and one escape cost are enough to compute where your allocation should flip; the crossover points on your own workflows are a measurement away.
To price one accepted unit on your own workflow, Talk to an Engineer.
A token is a unit of consumption, not of work. On one of the largest production AI gateways, open-weight models carried 29% of tokens on under 4% of spend; the bill alone cannot tell management what one accepted result cost.
The managing metric is cost per accepted unit of work, and the accepted unit comes first: what must be complete, what must be correct, what evidence travels with it, and who may accept it, set by the consequence of a failure.
Public benchmark data makes it computable: roughly 9.2 cents per accepted task frontier-only, 2.6 cents routed, 0.6 cents on a small model alone, with crossover points below one dollar of failure cost. Allocation measured on your own cases decides more than budget size.
Frequently asked questions
What is a cognition budget?
The money and accountable human attention assigned to machine-assisted work, managed as a production budget. It covers model and infrastructure spend, routing and evaluation, and the review, failure, retry, and rework required to reach acceptance, allocated across humans for judgment, open-weight models for throughput, and frontier models for consequential reasoning.
How should executives measure AI productivity?
Define an accepted unit of work per workflow, then divide accepted units into their full cost: model spend, routing, failed runs, review time, and rework. If cost per accepted unit falls while the quality bar holds, the budget works. A falling token price with rising review burden is not a saving.
Should every task use the strongest available model?
Above a computable failure-cost threshold, the configuration with the highest measured acceptance rate is the cheapest option per accepted result. Below it, premium reasoning adds cost without adding acceptable quality. The threshold is found with private evaluation sets on your own cases, and high-consequence decisions keep a named human reviewer regardless of model confidence.
Related Articles

An open-weight model is the graduate, not the school
Open weights release the graduate, not the school. Kimi K3 shows why the difference from open source matters: the license activates at scale, while the architecture keeps the model swappable. Across paterhn deployments, open models carry at least 60% of production agent tasks.

The frontier model is the easy part. The learning loop is the moat.
The durable advantage is not the frontier model you rent. It is the owned loop between your people and your AI: private evals, your traces, your judgment, your evidence. Compliance is where that thesis gets tested under load. Own the loop, or let the compounding accrue somewhere else.

Agent labor is the new subscription
$800B wiped from software stocks in a week. Everyone is writing the obituary. We wrote the birth announcement. From seats to agent labor: what comes next.

Code is cheap now. Software isn't.
The barrier to writing code collapsed. The coding agent market exploded from zero to $3 billion in five years. Production teams now deliver in 12 weeks what took 6 months. This article shows how the economics changed, what we've learned shipping code agents into production, and why this moment is not the end of anything. It's a beginning.