Over the past several months, I’ve spent a lot of time researching the rapidly emerging neocloud space. That has meant going well beyond the surface-level debate around AI infrastructure and trying to understand the entire stack: semiconductor architectures, data center economics, liquid versus air cooling, power generation and grid constraints, and what fundamentally separates neoclouds from the hyperscalers.
The deeper I went, the more puzzled I became by how differently the market seemed to view the sector.
Eventually, I realized the disconnect was not necessarily that others were missing an obvious fact. It was that understanding the economics of these businesses requires following several layers of the AI stack at once. And increasingly, one of the most important layers may be the newest and least understood of them all: the economics of tokens.
Token economics are still developing in real time. We do not yet know exactly how the long-run relationship between falling inference costs, exploding usage, model efficiency, and physical compute demand will settle.
But after months of following the data, I kept hearing the same signal.
There was something about the neocloud model, and CoreWeave in particular, that I could see in the numbers but had struggled to articulate clearly. The more I studied it, the more the pieces began to fit together.
This piece is my attempt to finally put that idea into words.
My goal is to explain why I believe the value proposition of the neoclouds extends far beyond the two narratives that dominate the debate today:
“Compute is scarce.” and “There is too much leverage.”
Both matter. Neither, in my view, captures the full economics of what is happening.
Because there is a strange divergence developing inside the AI economy.
And I’ve decided to call it The Compute Paradox.
The Compute Paradox: As the cost of intelligence falls, demand for it can rise even faster, causing total compute consumption, and potentially the value of the infrastructure supplying it, to increase rather than decline.
When Cheaper Intelligence Demands More Compute
There is a strange divergence developing inside the AI economy.
The price of intelligence is collapsing.
The price of the compute required to produce it is not.
That distinction may end up being one of the most important signals in the entire AI infrastructure trade.
Citadel Securities recently published Elastic Expectations, an excellent follow-up to its earlier work on Tokenomics. The central question is deceptively simple:
When intelligence becomes cheaper, do users spend less on AI, or do they consume so much more intelligence that aggregate spending actually increases?
The latest evidence increasingly points toward the second outcome.
Usage-weighted token prices have fallen roughly 40% since the end of June, according to Silicon Data. Yet Citadel calculates that July AI spend per employee increased approximately 49% month over month among the top 1% of enterprises, 25% among the top 10%, and 9% at the median. GPU rental prices, meanwhile, have rebounded from their June lows even as token prices continued falling.
This is not what the simple AI commoditization bear case was supposed to look like.
And for companies such as CoreWeave, it may fundamentally change the valuation framework.
The Economics of Cheaper Intelligence
The conventional AI infrastructure bear case is straightforward:
Better GPUs → greater efficiency → cheaper inference → lower compute prices → commoditization → lower infrastructure returns
There is nothing irrational about that argument.
In fact, if demand for intelligence were relatively inelastic, it would probably be correct.
But there is another possibility.
Better GPUs → cheaper intelligence → more economically viable AI workloads → exponentially more consumption → higher aggregate compute demand
This is essentially Jevons paradox applied to intelligence.
Efficiency does not necessarily destroy demand for a resource. Sometimes it makes the resource useful enough that total consumption increases.
We may now be watching that happen in real time.
The more affordable intelligence becomes, the more intelligence the economy may decide to consume.
That distinction matters enormously.
Citadel’s most important chart may be the divergence between token prices and H100 rental prices.
Earlier this summer, the two were declining together.
That was dangerous.
Falling token prices alongside falling compute prices could have indicated something much worse than technological efficiency: weakening demand across the entire AI stack.
But that relationship has since broken.
GPU rental prices measured by Ornn have rebounded while token prices have continued to decline. Citadel describes this as tentative evidence that the Jevons effect may be becoming more dominant. Lower-priced intelligence appears to be stimulating enough additional usage to support demand for the underlying physical compute.
That is an extraordinarily important development for the neoclouds.
Quantity Is Exploding
Now add the quantity side.
OpenRouter recently crossed approximately 75 trillion tokens per week, with the subsequent week pacing above 87 trillion tokens.
That represents an annualized run rate north of 4.5 quadrillion tokens.
OpenRouter’s chart is remarkable because it is presented on a logarithmic scale and still looks almost like a straight line.
From roughly:
10 billion tokens per week in January 2024
1 trillion per week by March 2025
10 trillion per week by February 2026
75 trillion-plus today.
The exact growth rate will obviously decelerate. No serious valuation model should extrapolate this curve indefinitely.
That is not the point.
The point is that the quantity response to declining intelligence costs is enormous.
Token prices can fall dramatically while the economic value of the infrastructure layer rises if quantity grows faster than price declines.
At the simplest level:
Revenue ≈ Price × Quantity
But AI infrastructure requires another layer:
Compute demand ≈ AI workloads × compute required per workload
The variable that matters for infrastructure investors is therefore not the nominal price of a token.
It is the total physical computation demanded by the economy.
The Compute Elasticity Ratio
While doing my research, I asked ChatGPT if there was a useful way to conceptualize this dynamic. What it came up with was actually pretty great.
Call it the Compute Elasticity Ratio, or CER:
CER = (1 + AI workload growth) ÷ (1 + effective compute-efficiency growth)
The intuition is simple.
CER < 1
Efficiency grows faster than workloads.
Physical compute demand eventually contracts.
Bearish neocloud regime.
CER ≈ 1
Additional workloads roughly absorb efficiency gains.
Physical compute consumption stabilizes.
Commodity regime.
CER > 1
AI workloads grow faster than compute efficiency.
Physical compute demand continues expanding despite relentless technological improvement.
Bullish infrastructure regime.
And if CER remains materially above 1?
That is where the neocloud economics become extremely interesting.
The Firms That Understand AI Are Spending More, Not Less
Citadel’s enterprise spending data make this story even stronger.
The highest-spending AI companies are not showing signs of saturation.
They are accelerating.
Since October 2023, monthly AI spending per employee among the top 1% of firms has increased by approximately $6,542, compared with only $9.63 at the median. The ratio between 90th-percentile and median spending has roughly doubled from 27x to 54x.
Think about what this implies.
The companies furthest along the AI adoption curve are not discovering that they have “enough AI.”
They are discovering more things to do with it.
Coding agents.
Reasoning models.
Long-context workloads.
Multi-agent systems.
Inference at scale.
Video generation.
Autonomous workflows.
Scientific computing.
Persistent enterprise agents.
The firms that have already learned how to consume intelligence appear to want dramatically more of it when its effective price declines.
This is exactly the demand curve an infrastructure owner would want.
Enter CoreWeave
This brings us to CoreWeave CRWV 0.00%↑.
The common bear framework for CRWV views the company as a highly levered owner of rapidly depreciating GPUs selling a product that will eventually become commoditized.
That is the risk.
But it may not be the whole story.
According to its Q2 2026 results, CoreWeave generated approximately $2.58 billion of revenue, up 112% year over year. Revenue backlog reached $104.2 billion, up from $99.4 billion one quarter earlier, before more than $25 billion of additional customer commitments secured early in Q3. Management said near-term capacity was effectively sold out and that new agreements were being signed on increasingly favorable terms.
CoreWeave subsequently raised full-year 2026 revenue expectations to approximately $12.4 billion to $13.2 billion and increased expected capital expenditures to $35 billion to $39 billion.
Those numbers matter on their own.
But they matter more when viewed alongside the token market.
The external market data are telling us:
Token prices ↓ Token consumption ↑↑↑
Enterprise AI spending ↑↑
GPU rental pricing firm
Cloud revenue ↑↑
And CoreWeave is simultaneously telling us:
Capacity sold out
Pricing improving
Backlog expanding
CapEx increasing to meet contracted demand
That is a meaningful convergence of signals.






