- Jul 14
JERK Report #22 AI inference is getting cheaper. So why isn't my bill?
I listened to Peter Diamandis' Moonshots podcast this week on a new Chinese open model that matches or beats the best from OpenAI and Anthropic.[^1]
The excitement was about the frontier race. Data centers in space. Who wins. Who falls behind.
My question was smaller; if inference is getting cheaper, does this mean my AI bill decreases?
Not necessarily is the short answer.
Training is when a model learns. Inference is when it answers. Every question you ask is inference.
Inference is getting cheaper. Training is getting more expensive. Frontier training cost has been rising about three and a half times a year since 2020, and doubling roughly every seven months.[^2]
So when you read that the AI companies are not profitable, and you also read that the price of a token has collapsed, both are true.
Serving an answer makes money. Building the next model loses it.
That matters because the training bill has to be recovered somewhere. There is only one place it can be recovered from: the price of an answer.
Position
The news is that AI use just got cheap.
It is true at the level people mean it. Stu Clott, an operations manager and part-time developer in San Diego, was spending about ten dollars on an hour-long coding session. The same work on DeepSeek cost him under fifty cents.[^3]
Flo Crivello moved his company Lindy off Anthropic's models onto DeepSeek and said it saved millions. Model cost had been Lindy's biggest expense. Bigger than payroll.[^4]
Crivello, speaking on the technology show MTS, said "You don't need God to write your email."[^5]
The measured collapse is real. In November 2022, GPT-3.5-level performance cost twenty dollars per million tokens. By October 2024, seven cents. A drop of more than 280 times in about two years. That is a fall of more than ninety-nine percent.[^6]
That is the price of a fixed capability.
The frontier still carries a premium. Claude Opus 4.8 lists at five dollars per million input tokens and twenty-five per million output. Fable 5, the newer frontier model, lists at ten and fifty.[^7] Reasoning models cost roughly six times their non-reasoning siblings.[^8]
The cost of LLM use has three terms, not one.
The price of a token x The number of tokens per answer x The number of answers it takes to finish the job.
Only the price of a token has fallen.
Velocity
Term one. Price per token at a fixed level of performance is falling about forty times a year. That is a halving roughly every two months. Epoch notes the rate varies by task, from about nine times a year to nine hundred.[^2]
Term two. Tokens per answer are rising. Reasoning model output lengths grow about five times a year. Non-reasoning models, 2.2 times. Reasoning models already use around eight times more tokens per answer.[^9] Switching one open model into reasoning mode raised its token use ten to fourteen times.[^10]
Two terms. Opposite signs.
Forty times down vs five times up is not a wash. For the same task, run once, the answer gets cheaper. That is why Stu Clott's bill fell. That is why Lindy saved millions.
One caution on term two. Those numbers come from benchmark questions answered at the highest available reasoning effort, not from ordinary work.[^9] Treat the direction as evidence. Do not treat the rate as a forecast for your own bill.
Term three has no public number at all. Nobody is measuring how many attempts it takes to finish a real job.
One question becomes a research process. One prompt becomes an agent. One run becomes three retries, four tool calls and a correction.
Acceleration
Epoch AI ran the token price analysis across several years and found a median decline of about fifty times a year. Then they ran it again using only data from January 2024 onward. The median went to two hundred.[^11]
That is a shorter window re-estimated, not a second derivative measured. It points at acceleration. It does not prove it.
Epoch also excludes reasoning models from the falling-price series entirely, because reasoning models generate far more tokens per answer.[^11] That is another reason the acceleration here is qualitative and not quantitative.
Term two has the same shape and the same problem. Reasoning modes are becoming defaults rather than options. Agentic runs are longer than chat turns. Both true. Neither measured as a second derivative.
Two curves that look like they are bending but no data clean enough to prove it.
Jerk
Jerk is the change in acceleration. It is where the rules underneath the trend start to move. Like acceleration, this is qualitative.
What might be driving prices up? Three candidates. The first two explain why the token price fell. The third is about how you are charged.
Chips. AI chip performance per dollar improves about thirty-seven percent a year, and has for over a decade.[^2] Steady. Jerk near zero. If chips were the reason your AI got cheap, you would have seen thirty-seven percent off, not ninety-nine.
Innovation. Hardware alone cannot explain a ninety-nine percent fall. The rest comes from some mixture of better model design, smaller models, quantization, routing, utilization and competition. Nobody has decomposed it publicly. Chinese labs, denied the best chips, got better with the ones they had, and published how. The direction is known. The contribution is not. Jerk unclear.
The business model switch from all you can eat to a meter.
GitHub announced in April 2026 that Copilot was going to usage-based billing. On June 1, 2026, every Copilot plan switched to counting tokens at listed API rates. GitHub said in writing that the change was an important step toward a sustainable Copilot business.[^12] When your credits run out, there is no longer a free fallback to a cheaper model. You pay more, or you stop.[^13]
It is not one vendor. Across AI companies surveyed by ICONIQ, consumption-based pricing rose from thirty-five percent to forty-two percent in a single survey cycle. Outcome-based pricing rose from eighteen to twenty-three. The average company now runs 1.7 pricing models, up from 1.5.[^14] Subscriptions are not being replaced. They are being layered.
Anthropic's documentation says tuning effort is often a better lever than switching models, and effort defaults to high on Opus 4.8.[^7] On Fable 5 and Mythos 5, adaptive thinking is the only thinking mode. You can still set effort. You cannot set it to off.[^7]
GitHub started counting tokens on June 1. ICONIQ shows consumption pricing rising and being layered under subscriptions. Anthropic removed the off switch on some models.
All of this happening while the unit price is still collapsing.
Redesigning billing during a price collapse suggests the vendors have already read term three.
What it means for your business
On the discount. It is real and it is widely available. Use the cheapest model and the lowest effort that clears your quality bar. But everyone receives the same discount, so collecting it is not an advantage. It is table stakes.
On the budget. Your unit cost will fall. Your invoice depends on all three terms. The flat monthly fee is not going away, but a meter is being fitted underneath it. To save money, manage the cost of a finished job, not the cost of a token.
On the moat. When thinking was expensive, the big firm held the edge. It could afford the analysts you could not. Cheap thinking closes that gap. It does not close the gaps in trust, in customer knowledge, in distribution, in execution.
The advantage is not the saving, it is what you do with thinking you could not previously afford.
The five-minute practice
Two practices, because there are two questions. What am I spending? What could I do now?
What am I spending
Pick one task you run every week. The same one, worded roughly the same way, every time.
Run it on the default model at the default effort.
Run it again on the lowest effort available.
Run it again on the cheapest model in the list.
For each of the three, write down four things.
The credits, tokens or cost, if the product shows them.
Whether it cleared your quality bar.
How much correcting it needed.
Whether it finished the job on the first attempt.
Do not pick the cheapest answer. Pick the cheapest finished job.
If the cheap model or the lower effort clears the bar, you have found the discount everyone is writing about and you had not collected.
If it does not, you have learned something more useful. You now know what you are actually buying at the frontier, and why.
What could I do now?
Open your LLM and paste this.
Based on everything you know about my business from our work together, answer three questions. Do not ask me for more information first. Make your best inference from what you have, and tell me what you are inferring.
The price of a specific answer has collapsed. Frontier-level thinking has not.
Which of my prices was protected by expert thinking being expensive? My competitors can buy that thinking now. Which price is most exposed? Name one.
Which thing do I not do, because it would take too much expert time to be worth it, that is now cheap enough to do? Pick one. Tell me what it would actually take.
What advantage do I have that gets bigger, not smaller, when thinking is nearly free?
Be specific. Do not flatter me.
Next Tuesday, July 21, the JERK session runs as part of Workshop Week, July 20 to 24. We take the derivatives on a live signal and find the decision hiding before the position catches up.
https://www.designrosetta.com/workshop-week/
Best wishes,
Rose
Sources
[^1]: Peter Diamandis, Moonshots, EP 266, June 2026. The framing that Chinese labs found a way to reason more cheaply is theirs, not a measured finding. https://open.spotify.com/episode/44BAYgPTAs9sYm6pq7YeET
[^2]: Epoch AI, AI Trends dashboard, updated February 2026. Frontier training cost rising ~3.5× per year since 2020, doubling roughly every 7 months. Inference price at a fixed level of performance falling ~40× per year, equivalent to halving every 2 months; the rate varies by performance milestone, from about 9× to 900× per year. AI chip performance per dollar has improved by 37% per year across accelerators released 2012–2025. https://epoch.ai/trends
[^3]: Low-cost Chinese AI models like DeepSeek gain traction in the U.S., Rest of World, July 2026. https://restofworld.org/2026/when-americans-choose-chinese-ai/
[^4]: Flo Crivello on Lindy's switch from Anthropic to DeepSeek. https://thenewstack.io/lindy-deepseek-anthropic-switch/
[^5]: Crivello, speaking on the technology show MTS, as reported by Rest of World. https://restofworld.org/2026/when-americans-choose-chinese-ai/
[^6]: Stanford HAI, AI Index Report 2025, Research and Development. GPT-3.5-equivalent performance on MMLU, $20.00 to $0.07 per million tokens, November 2022 to October 2024. https://hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development
[^7]: Anthropic, Choosing the right model and Introducing Claude Fable 5 and Claude Mythos 5, Claude Platform documentation. Fable 5 and Mythos 5 priced at $10 per million input tokens and $50 per million output tokens. Adaptive thinking is the only thinking mode on both models; thinking: {"type": "disabled"} is not supported. The effort parameter is used to control thinking depth. Tuning effort is often a better lever than switching models. The effort parameter defaults to high on Claude Opus 4.8 across all surfaces, including Claude Code and the Messages API. https://platform.claude.com/docs/en/about-claude/models/choosing-a-model and https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5
[^8]: Uptime Institute, 'Reasoning' will increase the infrastructure footprint of AI. OpenAI's o1 roughly six times the price of GPT-4o; DeepSeek R1 roughly six times V3. https://journal.uptimeinstitute.com/reasoning-will-increase-the-infrastructure-footprint-of-ai/
[^9]: Epoch AI, LLM responses to benchmark questions are getting longer over time, April 2025. Reasoning model output lengths growing 5× per year, non-reasoning 2.2× per year; reasoning models use around 8× more tokens on average. Measured on benchmark questions, using the highest available reasoning effort where models allow it to be set. https://epoch.ai/data-insights/output-length
[^10]: Towards a Standard, Enterprise-Relevant Agentic AI Benchmark, arXiv 2511.08042. Qwen3 models in reasoning mode showed a 10–14× token increase for accuracy gains of 10–13 points. https://arxiv.org/pdf/2511.08042
[^11]: Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks, March 2025. Epoch measure the price per token of the cheapest model reaching a given benchmark score, and exclude reasoning models from that analysis, because reasoning models generate far more tokens per answer. Published under Creative Commons Attribution. https://epoch.ai/data-insights/llm-inference-price-trends
[^12]: GitHub, GitHub Copilot is moving to usage-based billing, GitHub Blog, April 2026. All Copilot plans transitioned to usage-based billing on June 1, 2026, calculated on token consumption at listed API rates. GitHub described the change as an important step toward a sustainable, reliable Copilot business. https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/
[^13]: GitHub Community, All GitHub Copilot plans are now on usage-based billing (FAQ), effective June 1, 2026. Once the AI Credit allowance and any additional usage limit are exhausted, there is no free fallback to a cheaper model. https://github.com/orgs/community/discussions/197089
[^14]: ICONIQ Analytics + Insights, 2026 State of AI: The Builder's Economy, July 2026. Q2 2026 survey of approximately 305 executives at software companies building AI products. Consumption-based pricing rose from 35% to 42%; outcome-based from 18% to 23%; companies now blend 1.7 pricing models on average, up from 1.5. https://www.iconiq.com/growth/reports/2026-state-of-ai-bi-annual-snapshot
Check out the Jerk Report,
The JERK Report is a weekly signal read for small business owners. One signal. Four layers. A five-minute practice. Every Monday. From Rose Thun at Design Rosetta