OpenAI has argued that increased token efficiency on its GPT 5.5 model offsets higher prices—prices for it doubled versus its predecessor—especially on longer prompts. So, is that true? There is new data from OpenRouter on the question, and the answer is a clear "No". At all prompt lengths, the increased efficiency is offset by higher prices.

This matters because it's a falsifiable claim about rapid efficiency gains that didn't survive contact with real data. Model providers routinely argue that software optimization outruns rising compute intensity. Here, it didn't. A model can score better on benchmarks and generate higher user satisfaction while still costing more at scale—and that's what appears to be happening.
The broader pattern is worth watching. AI efficiency gains are behaving less like Moore's Law and more like Jevons Paradox: lower friction expands usage more than it lowers costs. More capable models induce longer contexts, more autonomous behavior, and higher reliability expectations. The result is that inference costs stay sticky even as the models improve. If that pattern holds, the standard industry narrative—that capability and cost efficiency move together—needs revision.