← Blog
June 29, 2026 · Analysis · Kent Langley

Is your AI spend paying for itself?

You can recite the token bill. You cannot see what it returned. Here is the half nobody publishes, and how to put a number on it this quarter.

Brian Armstrong published a good cost story this month. Coinbase cut its AI spend nearly in half while token usage kept climbing. The how is sound engineering. Smarter model routing, so each task runs on the cheapest model that can actually do it. Aggressive caching, so repeated queries stop paying for redundant outputs. A shift of routine work onto cheaper open-weight models, where the frontier adds nothing.

He goes further with a prediction. Within twelve to eighteen months, he argues, roughly 80% of workloads will run on models that cost about 99% less than today's frontier. The flagship models get reserved for the hard problems: scientific breakthroughs, complex agent orchestration. The binding constraint becomes energy and compute, not better models. The press has a name for the shift now circulating widely, the end of the tokenmaxxing era, though that headline is the coverage talking, not Armstrong.

I have no quarrel with the engineering. Coinbase was reportedly generating about 40% of its daily code with AI as of late 2025, with a stated goal to pass half, so this is not theory for them. It is shipping practice.

Cost curve check: does the prediction hold?

It does, and that is the point. Cheap inference is the most reliable trend in computing right now. Since ChatGPT shipped in November 2022, the cost of reaching a fixed level of capability has fallen on the order of 10x a year. The 2025 to 2026 price war alone knocked mainstream API prices down 60% to 80% in a single year. So 99% cheaper is a 100x drop, and the curve keeps clearing that bar in well under two years. Armstrong's twelve-to-eighteen-month window is a forecast, and an aggressive one, but it sits inside the trend, not beyond it. The claim that 80% of workloads commoditize is just the rule that yesterday's frontier becomes today's default. We watched it happen to GPT-3.5, then to GPT-4. One honest twist: the frontier itself can still get more expensive. OpenAI roughly doubled its top-line price when GPT-5.5 shipped in April 2026. That is not a counterexample, it is the rule restated. The curve collapses the cost of a fixed capability, while the new frontier, the genuinely hard 20%, stays premium.

~40x
GPT-4-tier output, cumulative fall ($60 to under $1.50 per million, 2023 to 2025)
60-80%
mainstream API price drop across 2025 to 2026, the price war
~10x / yr
fixed-capability decline (LLMflation); Epoch range 9x to 900x, median near 50x
frontier ↑
GPT-5.5 roughly doubled the top line, April 2026 ($2.50 to $5 in)

Two honest caveats. This is mostly a software and market story, not a chip story. Algorithmic efficiency improves around 3x a year. Mixture-of-experts designs and better inference systems compound it, and an oversupply price war is doing much of the rest. Raw hardware cost per unit of compute, by contrast, improves only around 1.3x a year. And the fall is uneven. Routine work deflates fastest, which is the easy 80%. The genuine frontier stays expensive, which is exactly why the flagships get reserved for it, and why a new flagship can still cost more than the one it replaced. You can watch the split in a live market. OpenRouter routes across hundreds of models and publishes token volume by model, and outside analysts who layer in price find a familiar shape: a premium lab pulling a large share of revenue on a small share of the tokens while cheap open-weight models carry the bulk. That is the 80-20 of self-liquidation showing up as market structure.

Here is the uncomfortable part for anyone bragging about a cost cut. A deflation this predictable and this exogenous is not a moat. It arrives for every competitor on the same schedule. If your token bill falls 90% next year, so does theirs. Riding the curve well, the routing and the caching, can buy a temporary edge, but the deflation underneath is a tide, not a position you hold. What you do with the cheap tokens, the return, is the only part you durably control.

So cost is one side of the equation. It is the denominator. And almost every public story about AI spend right now leads with the denominator, which matters more than it sounds, because what leadership publishes is what the org builds dashboards around.

The return number your AI budget never shows you

Here is the side that stays in the drawer. Only 27% of CEOs say AI has met or exceeded their ROI expectations, down from 38% the year before. That figure comes from the Oliver Wyman Forum's CEO Agenda 2026, a survey of 415 chief executives, the people who sign the checks. And 53%, an outright majority, say it is simply too early to tell. Read that slowly. The number who can point to a clear, met return is small and shrinking, and the largest group cannot yet see the return at all.

27%
CEOs say AI met or exceeded ROI (down from 38%)
53%
say it is too early to tell: the measurement gap
85%
want to be agentic in 3 yrs (sponsored survey)
76%
say their model cannot support it (Celonis)

Be precise about what that proves, because the honest version is stronger than the slogan. "Has not met expectations" is not the same as "cannot see the return." Some of those CEOs measure their return cleanly and just do not like the figure. Some set the bar too high at the start. But the 53% who say it is too early to tell are describing something specific: nine months and millions of tokens in, they still cannot put a number on what the spend returned. And the distinction lets nobody off the hook, because here is the trap that catches all three groups. If you cannot see return by workflow, you cannot tell "AI genuinely underdelivered here" apart from "I just cut the one workflow that was actually working." Measurement settles that question no matter which camp you are in. That is why it comes first.

Now hold the two facts side by side. Cost is being optimized hard, in public, with case studies and threads and crisp before-and-after charts. Return is barely being measured at all. We have gotten very good at the half of the equation that tells us the least.

It is not for lack of appetite. In one widely cited industry survey, 85% of organizations want to be agentic within three years and 76% say their current operating model cannot support it (MIT Technology Review Insights, May 2026, sponsored research in partnership with Ema, with the 76% figure drawn from Celonis). Treat those as directional, not gospel, given who funded them. But the direction rings true, and notice what an operating model that "cannot support" agentic AI is actually missing. You cannot redesign around the workflows worth running if you have never measured which ones return. The readiness gap is a measurement gap wearing an org-chart costume.

The asymmetry

The same token spend can produce a return of negative one hundred percent or positive ten thousand percent, depending entirely on what you do with it. Cost tells you nothing about which one you have. A team that cut its AI bill in half may have trimmed the workflow returning thirty times its cost while lovingly protecting the one returning nothing. The bill went down. The mistake went up. Nobody noticed, because nobody was watching the right number.

The marketing move that makes your AI spend pay for itself

Marketers solved a version of this problem a long time ago, and they gave it a name.

A self-liquidating offer is a front-end offer whose revenue covers, or liquidates, the cost of acquiring the customer. The acquisition pays for itself. The instant that happens, good marketers stop trying to minimize ad spend. They scale it. Why would you cap something that returns more than a dollar for every dollar in, plus a customer on the back end? Cost stops being a ceiling. It becomes a throttle you open.

Transplant the idea straight across. Token spend is self-liquidating when the measurable value it produces covers its fully-loaded cost. Once a workflow clears that bar, the whole optimization flips. You do not cut its token spend. You feed it, in steps. A workflow that liquidates itself at five thousand prospects a month can behave differently at fifty thousand, as rework creeps up and the list saturates, so each order of magnitude earns its own quick test before the next. The frontier-versus-cheap-model debate shrinks into a denominator sideshow. The real game is the numerator.

One honest difference, because the analogy is a tool and not a proof. Attribution is harder here than in direct response. A coupon code ties a sale to an ad with near-certainty. A workflow's contribution to revenue travels through more hands, the rep, the offer, the brand, the timing. That difficulty is not a reason to skip the measurement. It is the reason the value ladder below starts at the rung you can defend cold and only climbs when the evidence earns it.

The flip

So the operating question for any founder is not "how do I spend less on tokens?" It is "is my token spend paying for itself, and how would I know?"

Put a return number on every AI workflow you run

Start with one metric and one ratio.

Token ROI (tROI) = (value produced minus fully-loaded token cost) divided by fully-loaded token cost. Self-Liquidation Ratio (SLR) = value produced divided by fully-loaded token cost. That is the friendly form, the one you can say out loud in a meeting.

The word that does the work is "fully-loaded." It means API and token cost, plus caching and infrastructure, plus the human review and rework time the output actually demands. A 99%-cheaper model whose work a person has to redo is not cheaper. It is negative. The quality drop that creates rework is a real cost, and it belongs in the denominator where it can be seen.

Use the metric you already trust

You do not need new vocabulary for any of this. The Self-Liquidation Ratio is just the general form. Departmentalize it into the cost metric each team already trusts. MarketingOps lives it as cost of acquisition: you spend to win a customer only when that customer is worth more than the spend, which is the original self-liquidating logic. SalesOps has cost per opportunity. Support has cost per resolved ticket. Engineering has cost per shipped feature. Tokens are just another input cost flowing into each of those lines. Call the ratio whatever your team says out loud. The name is local. The discipline is universal: every dollar of added cost, token or otherwise, has to produce a valid, measurable business return, or it gets cut.

There is a deeper move here, the one marketers learned the hard way. A spend number means nothing without a return number sitting next to it. The bare token bill is a rear-view mirror. It reports what you already paid. The Self-Liquidation Ratio is a windshield. It estimates what the spend will earn back, the same way a marketer pairs the cost of winning a customer against the gross profit that customer throws off over a lifetime. Publish only the rear-view number and you are driving a fast car by watching the road you already drove. Which raises the question the windshield depends on: how do you put a number on what a workflow returns?

The value attribution ladder

People skip measuring return because "value produced" feels impossible to pin down. It is not. You attribute at the highest rung you can defend, you discount for confidence, and you never count a rung you cannot evidence.

RungWhat it measuresHow to value it
1 Displaced costThe spend replaced a known cost: an FTE-hour, an outsourced deliverable, a SaaS seatThe cost you no longer pay. Easiest, most defensible
2 ThroughputSame headcount, more output or a faster cycleIncremental margin from extra output, or the dollar value of compressed time-to-cash
3 Revenue liftThe spend produced or influenced revenue directlyAttributed revenue times margin, discounted by attribution confidence
4 OptionalityCapabilities, moats, proprietary data, learningTrack qualitatively with a confidence flag. Never launder it into a hard number

Rung 1 is where you start, and it is more than enough to begin. A workflow that produces twenty-five thousand custom drawings in three minutes, replacing roughly five full-time people, has a value you can defend to a CFO without hand-waving. A support assistant that deflects forty tickets a month, against a blended loaded cost of a few hundred dollars, has a rung-1 number you can post the first week it runs. The two failure modes are mirror images. One is claiming rung 3 or 4 with no evidence. The other is refusing to measure anything because rung 3 is hard, while rung 1 sat in plain sight the whole time.

The middle rungs are just as defensible when you anchor them to a number you already keep. A contract-review workflow that clears agreements 40% faster produces no revenue you can attribute, but the compressed cycle time has a dollar value: deals signing sooner, cash landing sooner, a backlog clearing without another hire. That is a rung-2 figure a CFO will accept, and it never required a single attributed sale.

The four numbers that turn the ratio into a system

The ratio tells you where a workflow stands. Four supporting numbers tell you what to do about it. Cost per outcome, not cost per token. Quality-adjusted cost, which is token cost plus rework cost, the trap-catcher that exposes the cheap model that is actually expensive. Payback velocity, how fast an initiative crosses SLR 1.0. And the yield curve by workload, the one that changes how you run the company.

tROI

Token ROI: net return over fully-loaded cost

SLR >= 1

Self-Liquidation Ratio: value over cost

$ / outcome

Cost per accepted unit of value

Q-adj cost

Token cost plus rework cost

Payback

How fast a workflow crosses SLR 1.0

Yield curve

tROI and SLR by workload

Put plainly: cost per outcome is the honest denominator, quality-adjusted cost is the lie detector, payback velocity is the clock, and the yield curve is the map. Together they turn one ratio into an operating instrument.

Which of your workflows to scale, fix, optimize, or kill

Once you have a yield curve, you stop managing AI as a budget line and start managing it as a portfolio. Some workloads are thirty times self-liquidating. Some are negative one hundred percent. A single company-wide "AI ROI" number hides every bit of that, and so do the headline figures most companies report instead: total tokens consumed, or the bare token bill. Both are aggregates. Both are blind to which workflow earned its keep and which one quietly bled. Plot every material workload on two axes, return and cost, and the next move writes itself.

Low costHigh cost
High return Scale
Self-liquidating. Cost is a throttle. Open it.
Optimize
Proven return. Now apply routing and caching, here.
Low return Fix or kill
The vanity trap. Busy, cheap, pointless.
Kill
Bleeding. Kill it fast.

Sit with the high-return boxes, because that is where the conventional wisdom breaks. Reflexive cost-cutting there is not prudent. It is wrong. Trimming the spend on a workflow that returns thirty times its cost is starving your single best asset to look disciplined on a report. The instinct that earned you applause for cutting the bill is, in this exact quadrant, the instinct quietly destroying the most value you have. The thrift that built your margins is now blinding you to the return sitting one column over. That reversal, where the strength that got you here becomes the constraint that holds you back, is the founder's oldest trap.

Your edge lives on the return side, not the cost cut

So where does the routing-and-caching playbook belong? Look at the matrix again. It belongs in the high-return, high-cost box, and earns its keep there. Proven return. That is exactly where you optimize the denominator, because you already know the numerator is real.

The sequencing

The mistake is cost-only, not cost-early.

Armstrong's playbook is not wrong. The trap is not running it. The trap is running it instead of measuring return, on spend you have never measured. And be careful with the strong version of this: the claim is not "never touch cost until every return is proven." Routing and caching can run in parallel with return measurement, and often should. Optimizing the cost of spend you have never measured for return is optimizing in the dark. Routing and caching make a returning workflow return more efficiently. They tell you nothing about whether it returns at all. Run them as your only discipline and you will dutifully, efficiently, scale your way deeper into the group that cannot see its return.

One fair pushback: measurement is not free either, and a founder's bandwidth is finite. So tier it. Point your instrumentation at the highest-cost workflows first, the ones where a wrong call costs the most, and let cheap heuristics ride on the small, uncertain spend until it earns a closer look. You do not have to measure everything at once. You have to stop cutting blind.

Can you state your token return? A three-check test

Here is the litmus, made operational. For each material workload, can you state its SLR? If you cannot, you are measuring tokens, not return. Three flags say you are not there yet.

How to instrument it

It is less work than it sounds, because half the rig already exists.

  1. The cost meter is built. The same model-routing infrastructure that produced Armstrong's cost cut emits per-call metadata. The cost meter is wired. The value meter usually is not. Tag spend at the workflow level, not the company level.
  2. Pick the highest defensible attribution rung per workflow, and write down the confidence.
  3. Emit an accepted-outcome event from every material AI workflow. Outcomes divided by cost gives you cost per outcome, for free.
  4. Run a monthly yield-curve review. Scale the top quartile. Fix or kill the bottom.
  5. Report cost per outcome and SLR to the board and the team. Not total spend. Not token volume.

A worked example, the winner. Take an AI-assisted outbound sales workflow, measured over one month.

Cost, fully-loaded (one month)Figure
Tokens and API: about 5,000 prospects run multi-step (research, draft, two follow-ups), roughly 100 million tokens at a blended $6 per million$600
Orchestration, enrichment data, caching$200
Human review and approval, about 13 hours of a rep's time$700
Fully-loaded cost$1,500
Value, attributed gross profit (one month)Figure
5,000 personalized emails, 0.8% book a meeting40 meetings
25% become qualified opportunities10 opps
20% of opportunities close2 deals
Gross profit per deal ($15,000 average deal at 60% margin)$9,000
Gross profit produced$18,000
Attributed to the workflow, a conservative 50%, since the rep and the offer also worked$9,000
The verdict

SLR = $9,000 / $1,500 = 6.0. tROI = 500%. Cost per closed deal = $750, against $9,000 of gross profit on each one. Notice the bare token bill was $600. Reported on its own, the number most teams would actually publish, it tells you nothing. The fully-loaded ratio says this workflow is not a cost to trim. It is a position to scale, and the next dollar belongs in more prospects and a better research model, not a cheaper one.

A word on that 50%. It is a placeholder for a test you have not run yet, not a number to fall in love with. The honest way to earn a real attribution figure is to run the same lead set through an AI-assisted rep and an unassisted one and watch the gap. Until you run that test, 50% is a conservative anchor, and the rule is to round down, not up. The real number will catch up with you either way.

The same math, a loser

Now run it on an AI-assisted competitive-research workflow. A strategist spends three hours a week prompting, reading, and verifying outputs that mostly confirm what the team already suspected. Token cost is modest, call it $120 a month. Loaded with the strategist's time, the cost is closer to $900. Rung-1 displacement: zero, because nothing was being outsourced before. Rung-2 throughput: marginal, because no decision moved faster. Rung-3 revenue: none anyone can trace. SLR lands well under 0.5. This is the bottom-left box. It is not self-liquidating, and no amount of cheaper tokens fixes it, because the problem was never the denominator. Redesign it toward a rung-1 displacement that actually exists, or kill it. The cheap-model reflex would have quietly kept it alive.

Cost per accepted change

Here is the sharpest version of cost per outcome, and almost nobody tracks it: cost per accepted change. Not tokens spent, not loops run. The cost of the work that actually cleared for production. The threshold is intuitive once you name it. If a loop hands you ten results and you throw six away, you are doing the review work the tool was supposed to save, and below roughly a 50% accept rate it can cost more than it gives back. It started with coding agents, but it generalizes to every loop a business runs: drafts a rep sends versus drafts they rewrite, support replies that ship versus replies redone. Acceptance rate is where token cost meets human rework, which makes it the truest read on productivity, because it counts only the work that survived the bar. Fold it into the denominator. A falling accept rate is the early warning that a cheaper model is quietly costing you more. (The framing is Anatoli Kopadze's; the rough 50% line is my own working threshold, not a measured constant.)

CFO test

If your CFO cannot calculate your cost of inference per completed workflow, the architecture is not yet operationalized. Tokens are cost of goods sold. Treat them like it.

Your move this quarter: publish yield, not spend

Cost discipline without return measurement is just thrift. Useful, but blind. The companies that win the next eighteen months will not be the ones that spent the least on tokens. They will be the ones who knew, per workflow, exactly what each token returned, and fed the spend that paid for itself while killing the spend that did not. Cost optimization gets you to efficient. Return measurement gets you to right. Run cost-only and you optimize the dark. Pair them, return first, and the cost work finally has a target worth aiming at.

One honest caveat, because the hard edge cuts both ways. Some workflows will never look token-return optimal, and sitting next to the ones that do, they will read as expensive. Be careful what you call a loser. The cost curve is still falling, so a workflow that is underwater today can surface in a year without changing at all, the denominator simply dropping out from under it. Some value is real but refuses to become a number, which is exactly what rung 4 is for, and you do not kill it just because it will not fit a spreadsheet. And some buyers have more money than price could ever matter to, so for them the return question was never about cost.

Underneath all of it sits a question nobody has answered yet. As inference trends toward free, does the self-liquidation test stay binding, or does almost everything eventually pay for itself, leaving judgment, taste, and trust as the only things left to compete on? I do not know. As a society, we do not know yet. That is the part worth watching.

Bottom line

Spend is not the score. Yield is.

Stop publishing your token-spend brags and start publishing your token-return KPIs. Cost per outcome. SLR by workload. Your yield curve. That is the post that would actually help the majority of leaders who still cannot see their return. Pick your top three AI workloads this week. For each one, answer a single question: is this spend paying for itself, and how do I know?

Sources

Subscribe

Notes like this one also go out through factually, my newsletter. Subscribe at news.kentlangley.com, or point your reader at the RSS feed.

Founder OS · Published 2026-06-29 · Instance: factual · Project: fos-www-blog
Skills applied: designing-fos, writing-copy, diagnose-you-get-benefit-now, analyzing-text, measuring-token-roi, analyzing-ai-costs, building-an-exo
fos.kentlangley.com