The breakthrough of pre-training is the realization that this recipe is good. You say, “Hey, if you mix some compute with some data into a neural net of a certain size, you will get results. You will know that you’ll be better if you just scale the recipe up.” ... Companies love this because it gives you a very low-risk way of investing your resources.
—Ilya Sutskever
The two key ideas:
1. Another prominent AI scientist joining me in saying that AI progress in the current "scaling" orthodoxy is grinding to a halt.
2. The scaling orthodoxy floated most data center capex boats in recent years, and its slowing has broad consequences.
I'm surprised it didn't get more attention, but the Dwarkesh interview from which this came was long enough that I doubt many made it through.
The Gist
A new interview with former OpenAI scientist Ilya Sutskever captures, almost accidentally and in passing, something important about the AI boom. It helps answer the question everyone asks: Why are companies willing to spend so much?
The naive answer is that it is all about the perceived size of the AI opportunity. But that is uncertain, and captures only one side. What it misses is how, for a halcyon period, from 2017-2022, compute spending on AI had not only been derisked; it had turned into a predictable capability production function.
compute + data + paramaters +training = capability
This created a new kind of speculation, one that doesn't feel like speculation. Pre-training scaling "laws" created the illusion of a physics-like production function: add compute, get capability. That belief is what has been driving a trillion-dollar capex cycle with no historical parallel. And now that the curve's costs have soared and capabilities bent, we’re left with what increasingly looks the largest mispriced engineering bet in modern technology.