Intelligence got 1,000x cheaper. Your AI bill went up. An 1865 book on coal explains why
The price of a unit of AI has collapsed faster than chips or bandwidth ever did, and subscriptions still climbed from $20 to $200. That is not greed. It is a mechanism William Stanley Jevons described about coal in 1865.
"It is wholly a confusion of ideas to suppose that the economical use of fuel is equivalent to a diminished consumption. The very contrary is the truth." That could be a line from an internal AI memo. It is 160 years old, written about coal by a 29-year-old economist in a book almost nobody reads, and it explains why one of the most deflationary products in modern economic history costs you more every month.
Deflationary is not an exaggeration. Holding quality fixed, GPT-3-level intelligence fell from $60 per million tokens in late 2021 to $0.06 by 2024, a factor of 1,000 in three years, and Epoch AI measures the price of reaching a given benchmark score falling between 9x and 900x per year depending on the task. Meanwhile the tiers went $20, then $100, then $200 a month, and even at the top the quotas tightened.
So the unit cost of intelligence has never been lower, and you have never paid more for access. It is tempting to call that a scam. It is closer to a mechanism, and a famous one: on 27 January 2025 Nvidia lost $589 billion in a session on the theory that efficiency destroys demand, and Satya Nadella replied that the Jevons paradox was striking again. Only one of them can be right.
The mechanism, the evidence, and what it does to your bill
1. 1865: the efficiency that ate a country's coal
William Stanley Jevons, then 29, published The Coal Question in April 1865 into a Britain where coal did nearly everything. Every generation of steam engine since Watt burned less coal per unit of useful work, and engineers promised falling consumption would follow. Jevons noticed the opposite. Output was near 100 million tons a year and compounding at about 3.5% annually, a doubling every twenty years, roughly 40% a decade against a population growing just over 10% a decade. Cheaper motive power did not shrink the national fuel bill; it enlarged the territory of steam, into deeper mines, railways and whole new factories.
2. The line to keep: "new applications of coal are of an unlimited character"
That is Jevons's own phrasing and it is the load-bearing part of the argument. Efficient steam did not do the same work for less money, it conquered work that had been out of reach. The book sold slowly at first; recognition took about a year. John Stuart Mill cited it in the Commons on 17 April 1866, Gladstone leaned on it in that year's budget speech, and a motion for a royal commission on coal followed on 12 June 1866. His forecast, though, was wrong. He concluded Britain must "not only stop, we must go back", and instead oil, gas and electricity moved the constraint off coal.
3. The awkward part: measured rebound is usually partial, not total
Economists exhumed the paradox after the oil shocks and renamed it the rebound effect. Steve Sorrell's UKERC work puts the long-run direct rebound for the best-documented energy services at roughly 10% to 30%, not 100% and nowhere near 150%. Two reasons. Saturation: once the living room is at 20°C, cheaper heat gives you no reason to go to 35°C. And share of spend: when energy is a small slice of total cost, a price fall frees a limited sum, and much of it leaks to restaurants and phones rather than back into heat. Sorrell's published verdict on outright backfire is that the evidence is "far from conclusive".
4. Which means backfire needs three conditions, and AI is the test case
The same literature describes the rare configuration in which rebound can pass 100%. The resource has to sit at the core of the cost, demand has to be far from saturation, and each efficiency gain has to open genuinely new uses rather than discount old ones. Coal and steam in 1865 met all three. Your radiator meets none of them. Whether AI meets them decides both your subscription price and a decade of capital expenditure, so it is worth taking them one at a time instead of arguing by analogy.
5. Condition one: compute is the cost, not a line item on it
In 1865 coal supplied more than 90% of Britain's energy and entered the price of every unit of steam power produced. The AI analogue is inference. Each additional answer demands additional compute, so when the price per token falls, the marginal cost of serving that answer falls with it, and any saving reinvested in more tokens pushes total compute straight back up. Published estimates put inference at roughly 60% to 70% of AI compute demand, up from about 40% in 2024, though that split comes from secondary analysis rather than operator disclosure, so treat the exact share as an estimate rather than a fact.
6. Condition two: nobody has found the thermostat
Google is the one operator publishing a usage series, and Sundar Pichai's numbers are stark: 9.7 trillion tokens a month in 2024, 480 trillion in 2025, and more than 3.2 quadrillion by the I/O keynote of 20 May 2026. That is roughly 330x in two years and 7x in the last year alone, while the compute needed to reach a given level of capability kept falling. It aggregates models, products and modalities, so it is not a clean elasticity measurement. It does show the opposite of saturation: the unit price collapses and volumes accelerate.
7. Condition three: cheap tokens buy new behaviour, not cheaper old behaviour
Deep-research agents that read and reconcile hundreds of sources before writing a report. Reasoning models that spend tokens thinking before they answer. Coding agents that work autonomously for hours. None of that is affordable at $60 per million tokens; all of it is unremarkable at six cents. The irony of 27 January 2025 is precise: R1, the model that crashed the market in the name of efficiency, belonged to exactly the generation designed to spend more compute at answer time. Nvidia's own statement that day called it "a perfect example of Test Time Scaling".
8. Which is why your subscription was the first thing to break
If a single request can consume far more tokens than the last one, a flat unlimited plan becomes a losing bet. Cursor rewrote its $20 Pro plan on 16 June 2025; users hit surprise charges, and CEO Michael Truell apologised on 4 July and refunded, explaining that new models "spend more tokens per request on longer-horizon tasks". Anthropic added weekly caps on 28 August 2025 even for $200-a-month Max subscribers. Uber exhausted its entire 2026 AI coding-tools budget by April. Pichai, in May 2026: "many companies are already blowing through their annual token budgets, and it's only May."
9. The market's January 2025 bet has not paid off
The trade was clean: if frontier intelligence costs 20x less to produce, the GPU mountain is oversized and Stargate's $500 billion (announced 21 January 2025, $100 billion committed up front) is a monument to excess. Nvidia fell 17%, a record $589 billion of market value gone in a day, the Nasdaq about 3%, Jensen Huang roughly $20 billion on paper. Nvidia's reply was that cheaper models mean wider deployment, not less compute. The stock closed up 8.8% the next day, no hyperscaler trimmed its budget, and Morgan Stanley now models about $805 billion of hyperscaler capex for 2026.
10. The constraint does not vanish. It moves down the stack
This is where Jevons was genuinely wrong. He assumed the constraint would stay attached to coal; it moved to oil, gas and electricity, and demand kept climbing on top of it. The modern version will not be a token shortage, because tokens are what comes out of the machine. Scarcity sits in what produces them: leading-edge wafers from a handful of fabs, memory, data centres, gigawatts on grids that are already tight, and the heat you then have to get rid of. When price stops rationing demand, the queue gets paid in chips, fabs and watts.
Read as a thesis rather than a curiosity, that is uncomfortable in both directions. If intelligence keeps commoditising while the physical layers stay slow and hard to copy, value migrates toward GPUs, grids, fabs, memory and cooling: the things a click cannot duplicate. The product trends toward free, and the means of producing it become strategic. That is the trade the market got backwards in January 2025, and has kept getting backwards in slower motion since.
Two things would break the trajectory. Demand could actually saturate, capability could plateau, agents could stop multiplying, and today's hundreds of billions would become one of the great overcapacity episodes on record. Or a different substrate, reversible computing among the candidates, could route around the energy wall and let demand climb by a level instead of stalling. Until one of those happens, the planning assumption is unglamorous: your price per token keeps falling and your bill keeps rising, so budget for the workload you will actually run. Picking the right tool and the right depth for each job is the part you still control, and it is what TaskNorth is built to answer.
Questions people ask
- What is the Jevons paradox?
- It is the observation, published by William Stanley Jevons in The Coal Question in 1865, that making the use of a resource more efficient can raise rather than lower total consumption of it. Cheaper useful work makes previously uneconomic applications worthwhile, and the new demand can exceed the saving. Jevons wrote that "it is wholly a confusion of ideas to suppose that the economical use of fuel is equivalent to a diminished consumption."
- Why is my AI subscription getting more expensive if tokens are cheaper?
- Because the cheap tokens changed what a single request looks like. Reasoning models and agents spend far more compute per task than a one-shot chat answer did, so the cost of serving one subscriber can rise even while the price of each token falls. Cursor rebuilt its $20 plan in June 2025 for exactly that reason, Anthropic added weekly caps on $200 Max plans in August 2025, and enterprise buyers such as Uber exhausted a full-year AI coding budget in four months.
- Does the Jevons paradox really apply to AI?
- Not automatically. Economists measuring the rebound effect since the 1970s find it is usually partial, on the order of 10% to 30% for well-studied energy services, and Steve Sorrell describes the evidence for outright backfire as far from conclusive. Full backfire needs three conditions: the resource dominates the cost, demand is far from saturation, and efficiency unlocks new uses. AI currently appears to meet all three, which is why the pattern holds for now rather than as a law.
- Was the market right to sell Nvidia after DeepSeek R1?
- So far, no. The 27 January 2025 sell-off wiped a record $589 billion off Nvidia in one session on the premise that cheaper training means less compute demand. Nvidia argued the opposite, the stock recovered 8.8% the next day, no major cloud provider cut its capital budget, and hyperscaler capex forecasts for 2026 have since been raised to roughly $805 billion. The efficiency gain expanded the market instead of shrinking it.
Written 23 August 2026 from an owner-supplied French video essay; every figure was re-verified against primary or first-tier sources on 23 August 2026. Verified directly: Jevons's quotations and the 1865 coal figures (The Coal Question, 1865); Mill's citation in the Commons on 17 April 1866 and the royal-commission motion of 12 June 1866 (Hansard); Steve Sorrell's rebound estimates and his "far from conclusive" verdict (UKERC review; Energy Policy, 2009); the a16z LLMflation series ($60 to $0.06 per million tokens for GPT-3-level quality) and Epoch AI's 9x-900x per year benchmark-price analysis; DeepSeek's own V3 technical report; Nvidia's 27 January 2025 statement and closing prices (CNBC, Bloomberg); Nadella's post; the Stargate announcement; Cursor's June 2025 pricing post and Truell's apology; Anthropic's weekly-limit announcement; Pichai's I/O 2026 token figures and budget remark; Uber's exhausted 2026 budget (The Information, Fortune); Morgan Stanley's 2026 capex forecast. Corrections to the source material: the source states Stargate at "$5 billion", the announced figure is up to $500 billion over four years with $100 billion committed initially; Nvidia's one-day loss was $589 billion, not "about $600 billion"; the Google token series is 9.7 trillion (2024), 480 trillion (2025) and 3.2 quadrillion (May 2026) per month, not the billions the transcript gives, though the 330x-in-two-years ratio is correct; Sorrell's published wording is "far from conclusive", not "suggestive rather than definitive"; Jevons compared a population growing just over 10% a decade with coal output growing about 40% a decade, which implies closer to tenfold than the eightfold the source claims; and The Coal Question was not an instant success, it sold slowly for roughly a year before Mill and Gladstone took it up. Labelled but unconfirmed: the share of AI compute going to inference ("roughly 60% to 70%") rests on secondary analysis, not operator disclosure. Omitted as unverifiable: the claim that Jevons's 1863 gold pamphlet sold 74 copies, and the anecdote about his hoarded paper outliving him by fifty years. Disclosure: TaskNorth's knowledge base recommends Claude models, and Anthropic, whose rate limits are cited above, makes the models used to build this site.
Trying to work out which AI tools fit your task? Describe the outcome and get a Blueprint: the tools, the prompt, and the steps, with pricing we verified ourselves.