Winslow Tandler ← Issue 01


SECOND ORDER
Issue 02 · June 12, 2026
AI, Capital & Work

The Wrong Denominator

Since Issue 01 we have held that the binding constraint on AI is diffusion and cost, and this week we push that argument down to the unit of measurement. A token is an input priced by the seller, and it is a poor stand-in for the value a task delivers to the buyer.

Read the companion deck (PDF)

01 · WHAT HAPPENED

Seventy days repriced the cost of AI

For two years the AI trade ran on a convenient proxy, that more tokens meant more intelligence consumed and so more value created, and that proxy broke this spring once agents changed the unit economics of usage. An agent does not hold a single conversation, it loops and plans and retries and reviews, so a typical agent job now burns about 96,000 tokens (SemiAnalysis) where a chat exchange burned a fraction of that. The providers withdrew the subsidy in sequence, OpenAI moving Codex to usage billing on April 2, Google following with Gemini on May 19, and GitHub converting every Copilot plan to metered “AI credits” on June 1, with unlimited use surviving only for code completions. Sam Altman’s own summary this week, that token costs went from a non-issue in January to “a huge issue” and a meme, marks how fast the AI cost line moved from rounding error to board topic (Exhibit 1).

Exhibit

Amazon had run an internal leaderboard, KiroRank, that rewarded engineers for token consumption, and it removed the board once busywork flooded through agents to climb the rankings, so the response reads as governance rather than retreat. Uber’s roughly 5,000 engineers burned the company’s full-year 2026 budget for agentic coding tools in four months, with individual bills reported between $150 and $2,000 a month, and the company now caps spend at $1,500 per engineer per month, per tool, with an exception process (Bloomberg, June 2). Walmart capped its internal AI agent after demand ran too hot (Bloomberg, June 1), and Microsoft cancelled most of its Claude Code subscriptions, partly on cost (The Verge). Across these cases the buyers are metering and capping an input they intend to keep using, which is how firms treat a supply they depend on. Axios reported one enterprise spending $500 million on Claude in a single month with no usage limits, a figure that is single-sourced and unconfirmed by Anthropic, so we carry it as illustrative and not established. The Wall Street Journal’s summary was blunter, that “the free-money period for AI is definitively over.”

The tape registered the shift before the commentary did. The Silicon Data LLM Token Expenditure Index, a usage-weighted measure of what the market pays per million tokens, roughly doubled off its December base to an all-time high of 1.99 on May 20 and then rolled over, with Citadel’s June 10 note counting six straight down days through June 9, the longest such streak since January, and the index’s public page printing 1.75 on June 11, about 12% below the peak. The daily series is licensed and we do not chart it, so we work from Citadel’s published chart, the latest public print, and the reported anchors. The print is genuinely ambiguous: a falling expenditure index can mean demand rolling over, or it can mean deflation arriving in the spend line, the same work bought cheaper. Citadel’s desk leans to the second reading, treating the decline as a rotation toward cheaper models and adoption turning on “the price and scarcity of the inputs required to make AI operational at scale”, a framing we share. The source is worth noting, since Citadel is one of the two firms Zito identifies as gladly paying anything for the frontier, which makes it a price-insensitive buyer describing everyone else trading down.

02 · THE REFRAME

Tokens price throughput, and buyers pay for completed tasks

The sharpest correction came from inside the bull camp, which is why it carries weight. John Zito, co-president of Apollo, spoke at the Morgan Stanley US Financials Conference on June 10, and his words were “I think tokenmaxxing and token talk is - it’s a lot of BS, honestly. Like, if you look at per unit of knowledge and cost per unit of knowledge, prices are collapsing. Prices are collapsing per unit of IQ, if you did it that way.” The economics under the remark are the substance: a token is a unit of throughput, buyers ultimately pay for completed tasks, and the price of the two can diverge at the same moment. Zito traces the rising bills to firms pointing frontier models at work that does not justify the compute, captured in his line that “our IQs are so low that we’re actually using [AI tools] to check out the recipe for, you know, French toast.” Goldman partner Rich Privorotsky had reached the same denominator five days earlier, naming useful task completion per watt and per dollar as the metric that matters, and when two institutions of that weight converge on one framework inside a week the consensus is forming in real time.

The price data supports the denominator (Exhibit 2). In Epoch AI’s series the cheapest model clearing GPT-4-level general knowledge fell from $37.50 to $0.175 per million tokens in 23 months, GPT-3.5-level fell from $20.00 to $0.07 over the window the Stanford AI Index records as a 280x collapse in 18 months, and GPT-4-level coding fell about 100x in 16 months. Across the capability thresholds Epoch tracks the declines run from 9x to 900x per year, a median near 50x, with the fastest of them beginning after January 2024, and open-weight and Chinese models now deliver near-frontier capability at a tenth to a twenty-fifth of frontier prices (Citrini), while Cursor’s latest model is reported to match frontier coding performance at roughly a tenth of the cost per task. Priced per unit of capability, the deflation has no modern precedent, and it explains why token revenue is a poor proxy for value delivered.

Exhibit

The opposing case deserves its strongest form, because the CFOs complaining about the bills are right about their own ledgers. Gartner forecasts that inference on a trillion-parameter model will cost over 90% less by 2030 than in 2025 and still expects enterprise bills to rise, on the logic that agentic workloads burn 5 to 30 times the tokens of a chatbot exchange and that consumption is compounding faster than price is falling, and its analyst Will Sommer warns that “Chief Product Officers should not confuse the deflation of commodity tokens with the democratization of frontier reasoning”, so a unit cost down 90% multiplied by a unit count up thirtyfold is still a larger invoice. We accept the arithmetic in full, and it is consistent with the reframe, since bills can rise while cost per unit of capability collapses. Token expenditure says little about value created, but it measures the labs' revenue directly, which is why the index rollover matters even if Zito is right about everything else: it speaks to whether the revenue beneath roughly $3 trillion of AI-linked infrastructure investment is durable.

03 · THE CORPORATE RESPONSE

Enterprises ration without retreating, and the capital swap runs through hiring

The most instructive finding in the UBS enterprise checks is what enterprises are declining to do. About 60% of the IT executives polled call token costs a real issue, “this is not a made-up media story” in the bank’s own words, one described the Copilot pricing change in a single word, “chaos”, and another opened the first genuine AI bill and heard leadership say “we don’t have the money for this.” Not one check showed a company slamming the brakes, and the dominant behavior is guardrails, caps and alerts and model downshifting and pooled token budgets, with several firms explicitly refusing to throttle usage and instead cannibalizing other line items to make room, starting with external IT services, then cloud spend, and then headcount growth. That choice, to protect the AI line by cutting elsewhere, shows AI being treated as a near-inelastic input even at prices buyers openly resent. One sourcing caveat is owed, since the checks reached us through coverage of the Zito interview and not a standalone UBS report we could locate, so the survey numbers carry that flag, though the pattern they describe is corroborated by every public channel we track.

The same substitution surfaces in every adjacent survey (Exhibit 3). Gartner’s CFO work has expected headcount growth collapsing from 6% to 2% for 2026 even as 75% of CFOs raise technology budgets, 48% of them by double digits, and about 60% plan double-digit increases in AI investment specifically. The Atlanta Fed’s business survey has executives expecting +2.25% productivity from AI over three years against a 1.2% headcount reduction, with firm AI spend per employee rising from $1,358 in 2025 to $2,068 this year, roughly $280 billion in aggregate on the Fed’s flagged ballpark, while Walmart has frozen its 2.1 million-person workforce for three years as revenue grows. The trade runs through hiring plans rather than layoffs, it compounds quarter on quarter, and flat headcount paired with rising AI spend is hardening into the default planning template, a substitution of capital for labor we call the capital swap.

Exhibit

Underneath the survey noise is a structural change in the shape of the cost, as software turns from a fixed cost into a variable one that scales with use. Exponential View’s survey work found that over 70% of companies blew through their 2025 AI budgets, and SemiAnalysis’s Doug O’Laughlin argues that “automated intelligence” is a permanent new opex category, with firms swinging between overspend and underspend until each settles on its own labor-to-compute ratio. Boards will be asking for that ratio ahead of an adoption percentage by this time next year, because a variable cost that grows with usage demands a unit economics answer rather than a deployment headline.

04 · THE MACRO STAKES

The denominator decides a one-trade economy

The sums involved are large enough that we computed them from the filings this issue rather than quoting a bank deck. Cash capex at the big five hyperscalers ran $158 billion in 2022, $154 billion in 2023, $239 billion in 2024, and $412 billion in 2025, close to a tripling in two years, with 2026 company guidance summing to about $660 billion and Morgan Stanley’s lease-inclusive estimate reaching $805 billion (Exhibit 4), and Amazon alone plans roughly $200 billion and sold C$14 billion of Canadian-dollar bonds on June 8 to help fund it. Apollo’s chief economist Torsten Slok calculates that hyperscalers now devote about 60% of operating cash flow to capex and that the top ten S&P 500 names trade richer than they did at the 1999 peak. The bubble thesis and the deflation thesis now coexist inside the same firm, and the open question is whether the revenue this spending presumes will materialize.

Exhibit

We then tested the strongest macro claim on the tape directly (Exhibit 5). Morgan Stanley estimates that AI-related investment drove about 75% of Q1 2026 GDP growth, and the national accounts cannot isolate “AI” as such, though they do publish the categories where AI equipment and software sit, and information processing equipment contributed 0.87 points and software 0.49 points of the quarter’s 1.6% annualized growth, 85% of it between the two. Against their 2017 to 2024 average of 0.31 points the surge above trend is about 1.05 points, two thirds of all growth, and stripping that surge out leaves an economy growing roughly 0.5%. The 75% figure is close to what the accounts allow, provided one assumes nearly all of the IT surge is AI, which is the assumption the honest caveats then qualify in both directions. The measure is gross of imports, and the same equipment line contributed +0.89 points in Q1 2025 while GDP shrank 0.6% in the tariff quarter, it misses data centers and power, and it includes non-AI IT spend, so Goldman’s import adjustment and UBS’s shallow-adoption objection live in exactly those gaps. Our calculation bounds the claim more than it settles it.

Exhibit
05 · SPILLOVER

The narrative is firing ahead of the capability

The cost panic is leaking into the labor market through the narrative channel well ahead of the capability channel, and the distinction governs how durable the displacement turns out to be. Challenger’s May report makes AI the No. 1 stated reason for US job cuts for the third consecutive month, at 38,579 cuts citing AI, 40% of the month’s total, up from 7% in January, 25% in March, and 26% in April (Exhibit 6), and year-to-date AI-attributed cuts of 87,714 already run 1.6 times the whole of 2025, while May’s 97,006 total was the highest for the month since 2020. This is employer self-attribution, and Challenger itself flags that AI may be a socially acceptable cover for ordinary cost cuts, while HBR’s Davenport and Srinivasan document that the wave is anticipatory, with firms cutting ahead of demonstrated AI capability because investors reward the story. One CEO put the circularity more sharply, that some companies may be citing AI to cover the same cuts made to pay for AI.

Exhibit

The floor argument from our first briefing is arriving early, that substitution stops where the marginal cost of compute exceeds the marginal cost of labor for a given task. Tokenpanic made that floor visible faster than we expected, because once an agent job costs real money the instruction to “automate everything” becomes a budget line someone has to defend, and the marginal calculation reasserts itself. The displacement question remains genuinely contested in the data, with the NY Fed now attributing most of the rise in new-graduate unemployment to remote work over AI, a fight we take up directly in the next issue. Either way, the story of AI is restructuring organizations faster than the technology itself is.

06 · WHERE THE DOLLARS POOL

Good enough wins the volume, and the IPO inherits the tension

If buyers optimize cost per completed task, the demand curve migrates rather than shrinks, and the migration follows price. Simple work routes to local and open models, harder tasks go to the cloud, and the frontier is reserved for users whose return on it is provable, with Zito naming Citadel and Jane Street as the houses gladly paying anything for the frontier, and Citadel’s own note expecting frontier-intensive AI to concentrate in “a narrower set of firms with the balance sheets to absorb the compute cost.” Citrini’s rotation map points the same way, toward smart routing, observability, edge inference, and good-enough models, and Microsoft used Build this week to pitch MAI models it says burn 60% fewer tokens on coding. Volume keeps growing while the dollars and the margins pool somewhere other than where today’s valuations assume, so the token metric and the value it is taken to proxy are pulling apart.

That divergence is what makes the IPO timing awkward. Anthropic filed its confidential S-1 on June 1 at a reported $47 billion run-rate, up from about $10 billion a year earlier, following a Series H at a $965 billion post-money valuation, with press reports putting the IPO ambition near $1.8 trillion, and the filing landed the same morning flat-rate pricing ended and a day before its largest customers began rationing its product. OpenAI’s single largest customer now burns 100 billion tokens a month, per Altman, even as the broad base of buyers learns to spend less per task. The labs are monetizing token volume at the moment their customers are learning to minimize it, so underwriters will call that volume growth and buyers should ask how much of it survives the routing layer.

NEXT ISSUE · THE ENTRY-LEVEL PUZZLE

The same denominator problem runs through the entry-level labor data, where three credible sources disagree on what is being measured. Stanford’s payroll data reads AI as displacing workers aged 22 to 25 in exposed occupations, a 16% relative employment decline, the NY Fed attributes 64% of the rise in graduate unemployment to remote work, and Yale finds the occupational mix barely changing. Someone is measuring the wrong thing, and next week we work out who, and what the answer means for your hiring pipeline.

SELECTED SOURCES

Bloomberg interview with John Zito, Morgan Stanley US Financials Conference (Jun 10, 2026) · Citadel Securities, “Tokenomics,” F. Flight (Jun 10, 2026) · Citrini Research, “State of the Themes” (Jun 8, 2026) · D. Thompson with D. O’Laughlin (May 29, 2026) · Axios, “AI sticker shock” (May 28, 2026) · WSJ, “Corporate America Is Starting to Ration AI” (May 2026) · Bloomberg on Uber caps (Jun 2) and Walmart’s agent (Jun 1) · Business Insider on Amazon’s leaderboard (May 2026) · The Verge on Microsoft and Claude Code · GitHub changelog (Jun 1, 2026) · UBS enterprise checks via press coverage (Jun 2026; see caveat in 03) · Silicon Data Token Market Pulse · Epoch AI / Artificial Analysis price data (CC-BY) · Stanford AI Index 2025 · Gartner press releases (Oct 15, 2025; Feb 10 and Mar 25, 2026) · Atlanta Fed macroblog (May 6, 2026) · Challenger, Gray & Christmas (Jun 4, 2026) · Apollo 2026 Outlook (T. Slok) · Morgan Stanley estimates via press reports (May 2026) · SEC EDGAR XBRL company filings · BEA via FRED (May 28, 2026 vintage) · Fortune on the Anthropic S-1 (Jun 1, 2026) · BIS WP 1179 · G. Gopinath, Odd Lots (May 29, 2026) · HBR, Davenport & Srinivasan (Jan 2026).

Estimates are flagged where they appear; single-sourced figures carry a visible caveat. Second Order calculations are filed with data and method under data/ in the production repository. Second Order is an independent research briefing. It is provided for discussion purposes only and is not investment, legal, tax, or accounting advice. © 2026.