You Can’t Train Your Way Past Missing Information
The modality ceiling — what Kimi K3 and the race for scale are missing
In the summer of 2026, China’s Moonshot AI shook the field with Kimi K3. Released as an open-weight model, it matched — and on some benchmarks beat — the best models out of the United States. Scientists called it a “turning point,” a “Sputnik moment.” Within two days, demand forced Moonshot to suspend new subscriptions. Silicon Valley and Washington tensed up.
The excitement is understandable. But the question here is not “who is winning.” It is whether the direction is right. Kimi K3, and every large model competing the same way, sits on one road: more data, more scale, and distillation — training on one another’s output to squeeze out performance. Does that road end at general intelligence (AGI)? Or are we knocking harder and harder on the wrong door?
Cracks in the faith in scale
The belief of the past few years was simple: scale up and intelligence follows. More data and compute, more capability — the “scaling laws.” But signs that the curve is bending have started to come from inside the very labs that wrote those laws. In late 2024, reports piled up that OpenAI (its next-generation ‘Orion’), Google, and Anthropic each failed to get the leap they expected from their next models. OpenAI co-founder Ilya Sutskever said “pre-training as we know it will end,” and that “the 2010s were the age of scaling; now we’re back in the age of discovery.”
Problems that scale simply does not solve also surfaced. On abstraction tests like ARC, models grew tens of thousands of times larger with almost no gain. We needn’t jump to “large models are a dead end.” But it is signal enough that scale alone hits a wall.
Recycling is not progress
Where Kimi K3 became both a sensation and a controversy is distillation — one model absorbing performance by training on another model’s (or its own) output. This is, at bottom, recycling information that already exists. Real frontier systems aren’t fully closed, of course; real data and human feedback are mixed in. So “they’re collapsing” would be an overstatement.
But there is a more careful claim, and a harder one to refute. Training repeatedly on model-generated data has been shown experimentally to destroy performance — the rare tails of the original distribution die first (Shumailov et al., Nature 2024). Conversely, if you keep the real data and accumulate alongside it, that collapse disappears (Gerstgrasser et al., 2024). Put the two together and one conclusion follows. Accumulation keeps performance from getting worse — it buys stability. But stability is not progress. Progress — genuinely new knowledge or invention actually coming into being — requires information freshly drawn from the world: through experiment, observation, discovery, the kind that enters only by contact with reality. Scaling and mutual distillation rearrange and amplify what is already there; they do not create this new information. At best they circle the ceiling with great precision. They do not break through it.
Two kinds of “can’t”
Here is the crux. What the race for scale misses is that it conflates two different kinds of “can’t.”
One is an epistemic limit. The information is there, but the model can’t yet extract it. This shrinks with more training. It is the zone where scale and data genuinely work.
The other is an informational limit. The information isn’t there at all. This does not shrink with training. No model, however good, can conjure information that isn’t present — because processing cannot add information. The only way past this wall is not to run the model harder but to open a different channel that actually carries the missing information.
And this informational limit shows up in two places: in space, and in time.
The ceiling in space — what isn’t there now
The first is easy. Information that a channel physically cannot capture, right now.
Tesla is trying to finish self-driving with cameras alone, no LiDAR. Cameras can take you about as far as a human. But they can’t prevent what a human can’t see. On a snowy night with no visibility, the obstacle ahead simply isn’t in the camera’s photons. The information isn’t in that channel. No neural network can make what isn’t there. Only a different channel — LiDAR — catches the obstacle. It is not an algorithm problem; it is a channel problem.
Medical imaging is the same. An early cancer too small to leave a signal in the pixels of an X-ray, CT, or MRI is not caught by training the AI harder. That threshold is crossed not by learning but by another method — a biopsy, a follow-up scan. A different channel has to bring in the missing information.
This axis is hard to argue with. If the information is physically absent, no amount of computation on top of it helps.
The ceiling in time — what isn’t there yet
Large language models look different at first. Text is not short on information; it holds nearly everything humanity has written. So “LLMs fail because the information isn’t there” is wrong. (People often say “LLMs can’t reason,” but that is contested, and it isn’t the point here.)
The real problem is elsewhere. The moment a piece of knowledge becomes knowledge is the moment it is tested against reality. But that test happens in a future that hasn’t arrived. Tomorrow’s reality doesn’t exist yet, so there is no data that contains it. Information not yet in any corpus cannot be computed out of that corpus, no matter how much you scale.
This is not a different principle from the ceiling in space. It is that ceiling extended through time. Just as LiDAR brings in information not present now, genuinely new knowledge brings in information not present yet — and only contact with reality can bring it. Scale, distillation, recycling: none of them substitute for that contact. The limit of self-driving and the limit of large models are one wall with two faces — space and time.
The line AlphaFold draws
This isn’t abstract. AlphaFold — the crown jewel of AI in science — draws the line exactly.
AlphaFold predicted protein 3D structure superbly. Why did it work? Because the information that determines that structure — traces left in the sequence by hundreds of millions of years of evolution — was already in the data. The information was there, so computation pulled it out.
And AlphaFold stops precisely where the information isn’t. It cannot predict a truly novel fold with no comparable precedent, because there is no evolutionary trace to lean on. The same holds for how proteins move (dynamics). A 2026 roadmap by dozens of researchers named the cause plainly: not that the models are weak, but that the recorded data on protein motion is orders of magnitude smaller than the data on static structure. This is not a gap a bigger model fills. It is filled only by new measured data.
AlphaFold enacts this essay exactly. Where the information is present, computation solves it; where it is absent, no scale can make it. The line always sits where information leaves the channel.
Knowing the structure doesn’t make the drug
The line is even clearer in drug development. Knowing a protein’s structure doesn’t make a drug.
The numbers say it. Since 2019, roughly $60 billion has gone into AI drug discovery and about 175 AI-found candidates have entered human trials — yet as of 2026, not one has FDA approval. AI shortened candidate design from 3–4 years to about one, but the Phase 2 success rate that decides whether a drug actually works (~40%) and the overall ~90% attrition are the same as before AI.
The reason is simple. The answer to “does this molecule cure disease in a real human body” is in no dataset — that information doesn’t exist until a trial is actually run. AI rushes you up to the wall but can’t break it. The bottleneck didn’t disappear; it moved to the back end, and the back end still belongs to researchers and experiments.
So, is it the wrong direction?
Honestly, the other side has a strong voice too. Reasoning models with long chains of thought and reinforcement learning are already lifting performance; the “bitter lesson” says general methods that ride scale beat hand-built structure; and predictions that “scale is about to hit a wall” have often been wrong before. That humility is worth keeping.
But those rebuttals don’t touch the core. What reasoning models add is more processing over the same material — not new information the world hasn’t yet produced. The “bitter lesson” holds when the information is already there. Where the information isn’t there yet, more scale is just a more expensive mirror.
This is not to say large models’ achievements are fake. They are a powerful tool. But no matter how sharply you hone a knife, where the knife can’t reach you need a different tool. Today’s race looks like an endless sharpening of one knife.
So return to Kimi K3. However much you scale it and cross-distill it, it ends up in the same place as those drugs stalled at the clinical wall. The front end gets dazzlingly fast, but the place where genuinely new knowledge is born — the place where the answer comes only from contact with reality — stays put. If crossing that place by scale is where this race is headed, the question we should ask is not about more scale, but about the nature of the threshold we cannot cross. On the evidence so far, there is no guarantee that general intelligence waits at the end of endlessly sharpening one knife.
So the question changes. If there is a threshold you can’t train your way past, two things matter: knowing when you’ve reached it, and what you do when you’re there. The race for scale doesn’t even ask — because it pretends the threshold isn’t there.
I believe there is another way to meet this threshold head-on. But that is not this essay’s job. This essay stops at opening a door — at asking which door to knock on, instead of knocking harder on the wrong one.