Brian Armstrong published a good cost story this month. Coinbase cut its AI spend nearly in half while token usage kept climbing. The how is sound engineering. Smarter model routing, so each task runs on the cheapest model that can actually do it. Aggressive caching, so repeated queries stop paying for redundant outputs. A shift of routine work onto cheaper open-weight models, where the frontier adds nothing.
He goes further with a prediction. Within twelve to eighteen months, he argues, roughly 80% of workloads will run on models that cost about 99% less than today's frontier. The flagship models get reserved for the hard problems: scientific breakthroughs, complex agent orchestration. The binding constraint becomes energy and compute, not better models. The press has a name for the shift now circulating widely, the end of the tokenmaxxing era, though that headline is the coverage talking, not Armstrong.
I have no quarrel with the engineering. Coinbase was reportedly generating about 40% of its daily code with AI as of late 2025, with a stated goal to pass half, so this is not theory for them. It is shipping practice.
It does, and that is the point. Cheap inference is the most reliable trend in computing right now. Since ChatGPT shipped in November 2022, the cost of reaching a fixed level of capability has fallen on the order of 10x a year. The 2025 to 2026 price war alone knocked mainstream API prices down 60% to 80% in a single year. So 99% cheaper is a 100x drop, and the curve keeps clearing that bar in well under two years. Armstrong's twelve-to-eighteen-month window is a forecast, and an aggressive one, but it sits inside the trend, not beyond it. The claim that 80% of workloads commoditize is just the rule that yesterday's frontier becomes today's default. We watched it happen to GPT-3.5, then to GPT-4. One honest twist: the frontier itself can still get more expensive. OpenAI roughly doubled its top-line price when GPT-5.5 shipped in April 2026. That is not a counterexample, it is the rule restated. The curve collapses the cost of a fixed capability, while the new frontier, the genuinely hard 20%, stays premium.
Two honest caveats. This is mostly a software and market story, not a chip story. Algorithmic efficiency improves around 3x a year. Mixture-of-experts designs and better inference systems compound it, and an oversupply price war is doing much of the rest. Raw hardware cost per unit of compute, by contrast, improves only around 1.3x a year. And the fall is uneven. Routine work deflates fastest, which is the easy 80%. The genuine frontier stays expensive, which is exactly why the flagships get reserved for it, and why a new flagship can still cost more than the one it replaced. You can watch the split in a live market. OpenRouter routes across hundreds of models and publishes token volume by model, and outside analysts who layer in price find a familiar shape: a premium lab pulling a large share of revenue on a small share of the tokens while cheap open-weight models carry the bulk. That is the 80-20 of self-liquidation showing up as market structure.
Here is the uncomfortable part for anyone bragging about a cost cut. A deflation this predictable and this exogenous is not a moat. It arrives for every competitor on the same schedule. If your token bill falls 90% next year, so does theirs. Riding the curve well, the routing and the caching, can buy a temporary edge, but the deflation underneath is a tide, not a position you hold. What you do with the cheap tokens, the return, is the only part you durably control.
So cost is one side of the equation. It is the denominator. And almost every public story about AI spend right now leads with the denominator, which matters more than it sounds, because what leadership publishes is what the org builds dashboards around.
02 · The numeratorThe return number your AI budget never shows you
Here is the side that stays in the drawer. Only 27% of CEOs say AI has met or exceeded their ROI expectations, down from 38% the year before. That figure comes from the Oliver Wyman Forum's CEO Agenda 2026, a survey of 415 chief executives, the people who sign the checks. And 53%, an outright majority, say it is simply too early to tell. Read that slowly. The number who can point to a clear, met return is small and shrinking, and the largest group cannot yet see the return at all.
Be precise about what that proves, because the honest version is stronger than the slogan. "Has not met expectations" is not the same as "cannot see the return." Some of those CEOs measure their return cleanly and just do not like the figure. Some set the bar too high at the start. But the 53% who say it is too early to tell are describing something specific: nine months and millions of tokens in, they still cannot put a number on what the spend returned. And the distinction lets nobody off the hook, because here is the trap that catches all three groups. If you cannot see return by workflow, you cannot tell "AI genuinely underdelivered here" apart from "I just cut the one workflow that was actually working." Measurement settles that question no matter which camp you are in. That is why it comes first.
Now hold the two facts side by side. Cost is being optimized hard, in public, with case studies and threads and crisp before-and-after charts. Return is barely being measured at all. We have gotten very good at the half of the equation that tells us the least.
It is not for lack of appetite. In one widely cited industry survey, 85% of organizations want to be agentic within three years and 76% say their current operating model cannot support it (MIT Technology Review Insights, May 2026, sponsored research in partnership with Ema, with the 76% figure drawn from Celonis). Treat those as directional, not gospel, given who funded them. But the direction rings true, and notice what an operating model that "cannot support" agentic AI is actually missing. You cannot redesign around the workflows worth running if you have never measured which ones return. The readiness gap is a measurement gap wearing an org-chart costume.
The same token spend can produce a return of negative one hundred percent or positive ten thousand percent, depending entirely on what you do with it. Cost tells you nothing about which one you have. A team that cut its AI bill in half may have trimmed the workflow returning thirty times its cost while lovingly protecting the one returning nothing. The bill went down. The mistake went up. Nobody noticed, because nobody was watching the right number.
The marketing move that makes your AI spend pay for itself
Marketers solved a version of this problem a long time ago, and they gave it a name.
A self-liquidating offer is a front-end offer whose revenue covers, or liquidates, the cost of acquiring the customer. The acquisition pays for itself. The instant that happens, good marketers stop trying to minimize ad spend. They scale it. Why would you cap something that returns more than a dollar for every dollar in, plus a customer on the back end? Cost stops being a ceiling. It becomes a throttle you open.
Transplant the idea straight across. Token spend is self-liquidating when the measurable value it produces covers its fully-loaded cost. Once a workflow clears that bar, the whole optimization flips. You do not cut its token spend. You feed it, in steps. A workflow that liquidates itself at five thousand prospects a month can behave differently at fifty thousand, as rework creeps up and the list saturates, so each order of magnitude earns its own quick test before the next. The frontier-versus-cheap-model debate shrinks into a denominator sideshow. The real game is the numerator.
One honest difference, because the analogy is a tool and not a proof. Attribution is harder here than in direct response. A coupon code ties a sale to an ad with near-certainty. A workflow's contribution to revenue travels through more hands, the rep, the offer, the brand, the timing. That difficulty is not a reason to skip the measurement. It is the reason the value ladder below starts at the rung you can defend cold and only climbs when the evidence earns it.
So the operating question for any founder is not "how do I spend less on tokens?" It is "is my token spend paying for itself, and how would I know?"
Put a return number on every AI workflow you run
Start with one metric and one ratio.
Token ROI (tROI) = (value produced minus fully-loaded token cost) divided by fully-loaded token cost. Self-Liquidation Ratio (SLR) = value produced divided by fully-loaded token cost. That is the friendly form, the one you can say out loud in a meeting.
- SLR of 1.0 or higher: the spend is self-liquidating. It pays for itself. Open the throttle and scale.
- SLR below 1.0: not yet. Fix the value side, or kill it.
The word that does the work is "fully-loaded." It means API and token cost, plus caching and infrastructure, plus the human review and rework time the output actually demands. A 99%-cheaper model whose work a person has to redo is not cheaper. It is negative. The quality drop that creates rework is a real cost, and it belongs in the denominator where it can be seen.
You do not need new vocabulary for any of this. The Self-Liquidation Ratio is just the general form. Departmentalize it into the cost metric each team already trusts. MarketingOps lives it as cost of acquisition: you spend to win a customer only when that customer is worth more than the spend, which is the original self-liquidating logic. SalesOps has cost per opportunity. Support has cost per resolved ticket. Engineering has cost per shipped feature. Tokens are just another input cost flowing into each of those lines. Call the ratio whatever your team says out loud. The name is local. The discipline is universal: every dollar of added cost, token or otherwise, has to produce a valid, measurable business return, or it gets cut.
There is a deeper move here, the one marketers learned the hard way. A spend number means nothing without a return number sitting next to it. The bare token bill is a rear-view mirror. It reports what you already paid. The Self-Liquidation Ratio is a windshield. It estimates what the spend will earn back, the same way a marketer pairs the cost of winning a customer against the gross profit that customer throws off over a lifetime. Publish only the rear-view number and you are driving a fast car by watching the road you already drove. Which raises the question the windshield depends on: how do you put a number on what a workflow returns?
The value attribution ladder
People skip measuring return because "value produced" feels impossible to pin down. It is not. You attribute at the highest rung you can defend, you discount for confidence, and you never count a rung you cannot evidence.
| Rung | What it measures | How to value it |
|---|---|---|
| 1 Displaced cost | The spend replaced a known cost: an FTE-hour, an outsourced deliverable, a SaaS seat | The cost you no longer pay. Easiest, most defensible |
| 2 Throughput | Same headcount, more output or a faster cycle | Incremental margin from extra output, or the dollar value of compressed time-to-cash |
| 3 Revenue lift | The spend produced or influenced revenue directly | Attributed revenue times margin, discounted by attribution confidence |
| 4 Optionality | Capabilities, moats, proprietary data, learning | Track qualitatively with a confidence flag. Never launder it into a hard number |
Rung 1 is where you start, and it is more than enough to begin. A workflow that produces twenty-five thousand custom drawings in three minutes, replacing roughly five full-time people, has a value you can defend to a CFO without hand-waving. A support assistant that deflects forty tickets a month, against a blended loaded cost of a few hundred dollars, has a rung-1 number you can post the first week it runs. The two failure modes are mirror images. One is claiming rung 3 or 4 with no evidence. The other is refusing to measure anything because rung 3 is hard, while rung 1 sat in plain sight the whole time.
The middle rungs are just as defensible when you anchor them to a number you already keep. A contract-review workflow that clears agreements 40% faster produces no revenue you can attribute, but the compressed cycle time has a dollar value: deals signing sooner, cash landing sooner, a backlog clearing without another hire. That is a rung-2 figure a CFO will accept, and it never required a single attributed sale.
The four numbers that turn the ratio into a system
The ratio tells you where a workflow stands. Four supporting numbers tell you what to do about it. Cost per outcome, not cost per token. Quality-adjusted cost, which is token cost plus rework cost, the trap-catcher that exposes the cheap model that is actually expensive. Payback velocity, how fast an initiative crosses SLR 1.0. And the yield curve by workload, the one that changes how you run the company.
Token ROI: net return over fully-loaded cost
Self-Liquidation Ratio: value over cost
Cost per accepted unit of value
Token cost plus rework cost
How fast a workflow crosses SLR 1.0
tROI and SLR by workload
Put plainly: cost per outcome is the honest denominator, quality-adjusted cost is the lie detector, payback velocity is the clock, and the yield curve is the map. Together they turn one ratio into an operating instrument.
05 · The portfolioWhich of your workflows to scale, fix, optimize, or kill
Once you have a yield curve, you stop managing AI as a budget line and start managing it as a portfolio. Some workloads are thirty times self-liquidating. Some are negative one hundred percent. A single company-wide "AI ROI" number hides every bit of that, and so do the headline figures most companies report instead: total tokens consumed, or the bare token bill. Both are aggregates. Both are blind to which workflow earned its keep and which one quietly bled. Plot every material workload on two axes, return and cost, and the next move writes itself.
| Low cost | High cost | |
|---|---|---|
| High return | Scale Self-liquidating. Cost is a throttle. Open it. |
Optimize Proven return. Now apply routing and caching, here. |
| Low return | Fix or kill The vanity trap. Busy, cheap, pointless. |
Kill Bleeding. Kill it fast. |
Sit with the high-return boxes, because that is where the conventional wisdom breaks. Reflexive cost-cutting there is not prudent. It is wrong. Trimming the spend on a workflow that returns thirty times its cost is starving your single best asset to look disciplined on a report. The instinct that earned you applause for cutting the bill is, in this exact quadrant, the instinct quietly destroying the most value you have. The thrift that built your margins is now blinding you to the return sitting one column over. That reversal, where the strength that got you here becomes the constraint that holds you back, is the founder's oldest trap.
06 · The synthesisYour edge lives on the return side, not the cost cut
So where does the routing-and-caching playbook belong? Look at the matrix again. It belongs in the high-return, high-cost box, and earns its keep there. Proven return. That is exactly where you optimize the denominator, because you already know the numerator is real.
The mistake is cost-only, not cost-early.
Armstrong's playbook is not wrong. The trap is not running it. The trap is running it instead of measuring return, on spend you have never measured. And be careful with the strong version of this: the claim is not "never touch cost until every return is proven." Routing and caching can run in parallel with return measurement, and often should. Optimizing the cost of spend you have never measured for return is optimizing in the dark. Routing and caching make a returning workflow return more efficiently. They tell you nothing about whether it returns at all. Run them as your only discipline and you will dutifully, efficiently, scale your way deeper into the group that cannot see its return.
One fair pushback: measurement is not free either, and a founder's bandwidth is finite. So tier it. Point your instrumentation at the highest-cost workflows first, the ones where a wrong call costs the most, and let cheap heuristics ride on the small, uncertain spend until it earns a closer look. You do not have to measure everything at once. You have to stop cutting blind.
07 · The litmusCan you state your token return? A three-check test
Here is the litmus, made operational. For each material workload, can you state its SLR? If you cannot, you are measuring tokens, not return. Three flags say you are not there yet.
- Leaderboard gap. You can quote your token spend, or your cost-cut percentage, but not your token return.
- Attribution gap. None of your top three workloads has a value-attribution rung assigned to it.
- Aggregate blindness. You report one company-wide number, total tokens used or the bare token bill or a single "AI ROI" figure, with no per-workload yield curve underneath it.
How to instrument it
It is less work than it sounds, because half the rig already exists.
- The cost meter is built. The same model-routing infrastructure that produced Armstrong's cost cut emits per-call metadata. The cost meter is wired. The value meter usually is not. Tag spend at the workflow level, not the company level.
- Pick the highest defensible attribution rung per workflow, and write down the confidence.
- Emit an accepted-outcome event from every material AI workflow. Outcomes divided by cost gives you cost per outcome, for free.
- Run a monthly yield-curve review. Scale the top quartile. Fix or kill the bottom.
- Report cost per outcome and SLR to the board and the team. Not total spend. Not token volume.
A worked example, the winner. Take an AI-assisted outbound sales workflow, measured over one month.
| Cost, fully-loaded (one month) | Figure |
|---|---|
| Tokens and API: about 5,000 prospects run multi-step (research, draft, two follow-ups), roughly 100 million tokens at a blended $6 per million | $600 |
| Orchestration, enrichment data, caching | $200 |
| Human review and approval, about 13 hours of a rep's time | $700 |
| Fully-loaded cost | $1,500 |
| Value, attributed gross profit (one month) | Figure |
|---|---|
| 5,000 personalized emails, 0.8% book a meeting | 40 meetings |
| 25% become qualified opportunities | 10 opps |
| 20% of opportunities close | 2 deals |
| Gross profit per deal ($15,000 average deal at 60% margin) | $9,000 |
| Gross profit produced | $18,000 |
| Attributed to the workflow, a conservative 50%, since the rep and the offer also worked | $9,000 |
SLR = $9,000 / $1,500 = 6.0. tROI = 500%. Cost per closed deal = $750, against $9,000 of gross profit on each one. Notice the bare token bill was $600. Reported on its own, the number most teams would actually publish, it tells you nothing. The fully-loaded ratio says this workflow is not a cost to trim. It is a position to scale, and the next dollar belongs in more prospects and a better research model, not a cheaper one.
A word on that 50%. It is a placeholder for a test you have not run yet, not a number to fall in love with. The honest way to earn a real attribution figure is to run the same lead set through an AI-assisted rep and an unassisted one and watch the gap. Until you run that test, 50% is a conservative anchor, and the rule is to round down, not up. The real number will catch up with you either way.
Now run it on an AI-assisted competitive-research workflow. A strategist spends three hours a week prompting, reading, and verifying outputs that mostly confirm what the team already suspected. Token cost is modest, call it $120 a month. Loaded with the strategist's time, the cost is closer to $900. Rung-1 displacement: zero, because nothing was being outsourced before. Rung-2 throughput: marginal, because no decision moved faster. Rung-3 revenue: none anyone can trace. SLR lands well under 0.5. This is the bottom-left box. It is not self-liquidating, and no amount of cheaper tokens fixes it, because the problem was never the denominator. Redesign it toward a rung-1 displacement that actually exists, or kill it. The cheap-model reflex would have quietly kept it alive.
Here is the sharpest version of cost per outcome, and almost nobody tracks it: cost per accepted change. Not tokens spent, not loops run. The cost of the work that actually cleared for production. The threshold is intuitive once you name it. If a loop hands you ten results and you throw six away, you are doing the review work the tool was supposed to save, and below roughly a 50% accept rate it can cost more than it gives back. It started with coding agents, but it generalizes to every loop a business runs: drafts a rep sends versus drafts they rewrite, support replies that ship versus replies redone. Acceptance rate is where token cost meets human rework, which makes it the truest read on productivity, because it counts only the work that survived the bar. Fold it into the denominator. A falling accept rate is the early warning that a cheaper model is quietly costing you more. (The framing is Anatoli Kopadze's; the rough 50% line is my own working threshold, not a measured constant.)
If your CFO cannot calculate your cost of inference per completed workflow, the architecture is not yet operationalized. Tokens are cost of goods sold. Treat them like it.
Your move this quarter: publish yield, not spend
Cost discipline without return measurement is just thrift. Useful, but blind. The companies that win the next eighteen months will not be the ones that spent the least on tokens. They will be the ones who knew, per workflow, exactly what each token returned, and fed the spend that paid for itself while killing the spend that did not. Cost optimization gets you to efficient. Return measurement gets you to right. Run cost-only and you optimize the dark. Pair them, return first, and the cost work finally has a target worth aiming at.
One honest caveat, because the hard edge cuts both ways. Some workflows will never look token-return optimal, and sitting next to the ones that do, they will read as expensive. Be careful what you call a loser. The cost curve is still falling, so a workflow that is underwater today can surface in a year without changing at all, the denominator simply dropping out from under it. Some value is real but refuses to become a number, which is exactly what rung 4 is for, and you do not kill it just because it will not fit a spreadsheet. And some buyers have more money than price could ever matter to, so for them the return question was never about cost.
Underneath all of it sits a question nobody has answered yet. As inference trends toward free, does the self-liquidation test stay binding, or does almost everything eventually pay for itself, leaving judgment, taste, and trust as the only things left to compete on? I do not know. As a society, we do not know yet. That is the part worth watching.
Spend is not the score. Yield is.
Stop publishing your token-spend brags and start publishing your token-return KPIs. Cost per outcome. SLR by workload. Your yield curve. That is the post that would actually help the majority of leaders who still cannot see their return. Pick your top three AI workloads this week. For each one, answer a single question: is this spend paying for itself, and how do I know?
Sources
- Brian Armstrong, post on X, June 2026: x.com/brian_armstrong
- Tekedia: "Coinbase CEO Brian Armstrong Urges Shift to Cheaper AI Models, Signaling End of the 'Tokenmaxxing' Era" (the "end of the tokenmaxxing era" phrasing is this outlet's framing, not Armstrong's words)
- The Block: "Brian Armstrong says about 40% of Coinbase's daily code is AI-generated" (figure as of late 2025, with a stated goal to pass 50%)
- Oliver Wyman Forum, CEO Agenda 2026 (survey of 415 CEOs: 27% say AI met or exceeded ROI expectations, down from 38% the prior year, with a majority saying it is too early to assess)
- MIT Technology Review Insights, May 2026, "Rethinking organizational design in the age of agentic AI" (sponsored research, in partnership with Ema; the 85% agentic-intent and 76% operating-model figures, the latter drawn from Celonis)
- Stanford HAI, AI Index Report 2025 (GPT-3.5-level inference cost fell about 280x, November 2022 to October 2024)
- Andreessen Horowitz (Guido Appenzeller): "Welcome to LLMflation", November 2024 (about 10x per year, near 1000x in three years)
- OpenAI API pricing (2026): GPT-5.5 flagship at $5 / $30 per million input/output tokens, roughly doubling the top line at the April 23, 2026 release; mainstream API prices fell roughly 60% to 80% across 2025 to 2026 (the AI price war)
- Epoch AI: "LLM inference prices have fallen rapidly but unequally across tasks" (per-tier range roughly 9x to 900x per year, median near 50x; GPT-4-level output fell roughly 40x cumulatively off a March 2023 launch price near $60 per million; algorithmic efficiency about 3x per year versus hardware about 1.3x per year)
- OpenRouter, market-level usage across hundreds of models (token volume published officially at openrouter.ai/data; per-model revenue shares are derived by third-party analysts layering price onto that volume, not an official OpenRouter disclosure)
- Anatoli Kopadze (X, 2026): the cost-per-accepted-change framing and the rough 50% accept-rate threshold, used here as a springboard; the threshold is the author's synthesis. x.com/AnatoliKopadze
- Jeremy Kahn, Fortune (May 2026) and Azeem Azhar with Nathan Warren, Exponential View (May 2026), on the "tokenmaxxing is over" discourse
- Founder OS, building-an-exo and measuring-token-roi skills: the tokens-as-COGS framing, cost of inference per completed workflow