Skip to main content

[@DwarkeshPatel] Why smarter AI models could drive up compute prices 10x

· 4 min read

@DwarkeshPatel - "Why smarter AI models could drive up compute prices 10x"

Link: https://youtu.be/oZBGAuANX6I

Duration: 11 min

Transcript: Download plain text

Short Summary

AI lab economics are diverging sharply: revenue at companies like Anthropic is 10x'ing year-over-year (from $9B last year to a projected $100-150B this year) while lab compute only 3x's, forcing the gap to be filled by higher margins, higher compute prices, or a shift toward inference. Compute growth itself decomposes into roughly 1.4x from Moore's Law, 1.2x from new fabs (bottlenecked by ASML EUV through 2030), and 1.8x from AI absorbing leading-edge wafer allocation from smartphones and PCs. Market signals confirm tightening: spot prices are 40%+ above the February trough, Anthropic's inference margins reportedly jumped from 40% to 80%, and Google is paying $900M a month to rent 110,000 GPUs from SpaceX at 2x spot.

Key Quotes

  1. "Anthropic's revenue has 10x'd year over year, and it's likely to do so again this year. They ended last year with nine billion in revenue. I think they'll probably end this year with somewhere between one hundred billion to one hundred and fifty billion dollars in revenue." (00:00:05)
  2. "One, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that labs spend on inference rather than training has to increase." (00:00:54)
  3. "If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over 250K a year." (00:04:17)
  4. "if it costs twenty dollars an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model, because it's gonna burn more tokens on your expensive compute to get the exact same result." (00:05:59)
  5. "I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration, but it seems we do." (00:11:02)

Detailed Summary

AI Lab Economics: Revenue Outrunning Compute

Anthropic's Revenue Trajectory

  • Anthropic's revenue has 10x'd year-over-year for three consecutive years, ending last year at $9 billion.
  • It is projected to reach $100-150 billion this year, and would need $1 trillion by the end of next year if the trend continues.

The Compute-vs-Revenue Gap

  • Lab compute is only 3x'ing year-over-year while revenues 10x, creating a structural gap that must be filled by one or more of: higher lab margins, higher compute prices, or a larger share of compute going to inference versus training.
  • Anthropic's inference margins reportedly went from 40% in mid-last year to upwards of 80% for Fable now.

Compute Market Signals

  • Spot prices for compute are more than 40% higher than the February trough earlier this year.
  • According to Epoch, in 2024 OpenAI was spending just a quarter of its compute on inference, and that number is now likely closer to 50% or higher.
  • Google is paying $900 million per month to rent 110,000 GPUs (a blend of GB200s and GB300s) from SpaceX, at 2x the spot price per hour.
  • At current software engineer prices, an H100 running a true human-level software engineer should rent for over $250K per year, more than 15x the current H100 spot price.

Why Compute Only 3x's

  • The 3x year-over-year lab compute scaling breaks down into roughly 1.4x from Moore's Law, 1.2x from building new fabs (bottlenecked by ASML EUV machine production through 2030), and 1.8x from AI absorbing leading-edge wafer allocation from smartphones and PCs.
  • At leading-edge N3 nodes at TSMC, AI's share of wafer allocation is expected to rise from 60% to 86% by the end of next year, after which further growth in this dimension will hit a wall.

The Alchian-Allen Effect

  • The Alchian-Allen effect applies to AI labs: if a model can achieve the same result using less compute, it effectively creates more compute, and labs can charge higher margins for more efficient frontier models that economize a scarce input.