You Can’t Train Your Way Past Missing Information

The modality ceiling — what Kimi K3 and the race for scale are missing

CAIDE Systems, Inc. — AI Debate · a problem, not a solution

In the summer of 2026, China’s Moonshot AI shook the field with Kimi K3. Released as an open-weight model, it matched — and on some benchmarks beat — the best models out of the United States. Scientists called it a “turning point,” a “Sputnik moment.” Within two days, demand forced Moonshot to suspend new subscriptions. Silicon Valley and Washington tensed up.

The excitement is understandable. But the question here is not “who is winning.” It is whether the direction is right. Kimi K3, and every large model competing the same way, sits on one road: more data, more scale, and distillation — training on one another’s output to squeeze out performance. Does that road end at general intelligence (AGI)? Or are we knocking harder and harder on the wrong door?

Cracks in the faith in scale

The belief of the past few years was simple: scale up and intelligence follows. More data and compute, more capability — the “scaling laws.” But signs that the curve is bending have started to come from inside the very labs that wrote those laws. In late 2024, reports piled up that OpenAI (its next-generation ‘Orion’), Google, and Anthropic each failed to get the leap they expected from their next models. OpenAI co-founder Ilya Sutskever said “pre-training as we know it will end,” and that “the 2010s were the age of scaling; now we’re back in the age of discovery.”

Problems that scale simply does not solve also surfaced. On abstraction tests like ARC, models grew tens of thousands of times larger with almost no gain. We needn’t jump to “large models are a dead end.” But it is signal enough that scale alone hits a wall.

Recycling is not progress

Where Kimi K3 became both a sensation and a controversy is distillation — one model absorbing performance by training on another model’s (or its own) output. This is, at bottom, recycling information that already exists. Real frontier systems aren’t fully closed, of course; real data and human feedback are mixed in. So “they’re collapsing” would be an overstatement.

But there is a more careful claim, and a harder one to refute. Training repeatedly on model-generated data has been shown experimentally to destroy performance — the rare tails of the original distribution die first (Shumailov et al., Nature 2024). Conversely, if you keep the real data and accumulate alongside it, that collapse disappears (Gerstgrasser et al., 2024). Put the two together and one conclusion follows. Accumulation keeps performance from getting worse — it buys stability. But stability is not progress. Progress — genuinely new knowledge or invention actually coming into being — requires information freshly drawn from the world: through experiment, observation, discovery, the kind that enters only by contact with reality. Scaling and mutual distillation rearrange and amplify what is already there; they do not create this new information. At best they circle the ceiling with great precision. They do not break through it.

Two kinds of “can’t”

Here is the crux. What the race for scale misses is that it conflates two different kinds of “can’t.”

One is an epistemic limit. The information is there, but the model can’t yet extract it. This shrinks with more training. It is the zone where scale and data genuinely work.

The other is an informational limit. The information isn’t there at all. This does not shrink with training. No model, however good, can conjure information that isn’t present — because processing cannot add information. The only way past this wall is not to run the model harder but to open a different channel that actually carries the missing information.

And this informational limit shows up in two places: in space, and in time.

The ceiling in space — what isn’t there now

The first is easy. Information that a channel physically cannot capture, right now.

Tesla is trying to finish self-driving with cameras alone, no LiDAR. Cameras can take you about as far as a human. But they can’t prevent what a human can’t see. On a snowy night with no visibility, the obstacle ahead simply isn’t in the camera’s photons. The information isn’t in that channel. No neural network can make what isn’t there. Only a different channel — LiDAR — catches the obstacle. It is not an algorithm problem; it is a channel problem.

Medical imaging is the same. An early cancer too small to leave a signal in the pixels of an X-ray, CT, or MRI is not caught by training the AI harder. That threshold is crossed not by learning but by another method — a biopsy, a follow-up scan. A different channel has to bring in the missing information.

This axis is hard to argue with. If the information is physically absent, no amount of computation on top of it helps.

The ceiling in time — what isn’t there yet

Large language models look different at first. Text is not short on information; it holds nearly everything humanity has written. So “LLMs fail because the information isn’t there” is wrong. (People often say “LLMs can’t reason,” but that is contested, and it isn’t the point here.)

The real problem is elsewhere. The moment a piece of knowledge becomes knowledge is the moment it is tested against reality. But that test happens in a future that hasn’t arrived. Tomorrow’s reality doesn’t exist yet, so there is no data that contains it. Information not yet in any corpus cannot be computed out of that corpus, no matter how much you scale.

This is not a different principle from the ceiling in space. It is that ceiling extended through time. Just as LiDAR brings in information not present now, genuinely new knowledge brings in information not present yet — and only contact with reality can bring it. Scale, distillation, recycling: none of them substitute for that contact. The limit of self-driving and the limit of large models are one wall with two faces — space and time.

The line AlphaFold draws

This isn’t abstract. AlphaFold — the crown jewel of AI in science — draws the line exactly.

AlphaFold predicted protein 3D structure superbly. Why did it work? Because the information that determines that structure — traces left in the sequence by hundreds of millions of years of evolution — was already in the data. The information was there, so computation pulled it out.

And AlphaFold stops precisely where the information isn’t. It cannot predict a truly novel fold with no comparable precedent, because there is no evolutionary trace to lean on. The same holds for how proteins move (dynamics). A 2026 roadmap by dozens of researchers named the cause plainly: not that the models are weak, but that the recorded data on protein motion is orders of magnitude smaller than the data on static structure. This is not a gap a bigger model fills. It is filled only by new measured data.

AlphaFold enacts this essay exactly. Where the information is present, computation solves it; where it is absent, no scale can make it. The line always sits where information leaves the channel.

Knowing the structure doesn’t make the drug

The line is even clearer in drug development. Knowing a protein’s structure doesn’t make a drug.

The numbers say it. Since 2019, roughly $60 billion has gone into AI drug discovery and about 175 AI-found candidates have entered human trials — yet as of 2026, not one has FDA approval. AI shortened candidate design from 3–4 years to about one, but the Phase 2 success rate that decides whether a drug actually works (~40%) and the overall ~90% attrition are the same as before AI.

The reason is simple. The answer to “does this molecule cure disease in a real human body” is in no dataset — that information doesn’t exist until a trial is actually run. AI rushes you up to the wall but can’t break it. The bottleneck didn’t disappear; it moved to the back end, and the back end still belongs to researchers and experiments.

So, is it the wrong direction?

Honestly, the other side has a strong voice too. Reasoning models with long chains of thought and reinforcement learning are already lifting performance; the “bitter lesson” says general methods that ride scale beat hand-built structure; and predictions that “scale is about to hit a wall” have often been wrong before. That humility is worth keeping.

But those rebuttals don’t touch the core. What reasoning models add is more processing over the same material — not new information the world hasn’t yet produced. The “bitter lesson” holds when the information is already there. Where the information isn’t there yet, more scale is just a more expensive mirror.

This is not to say large models’ achievements are fake. They are a powerful tool. But no matter how sharply you hone a knife, where the knife can’t reach you need a different tool. Today’s race looks like an endless sharpening of one knife.

So return to Kimi K3. However much you scale it and cross-distill it, it ends up in the same place as those drugs stalled at the clinical wall. The front end gets dazzlingly fast, but the place where genuinely new knowledge is born — the place where the answer comes only from contact with reality — stays put. If crossing that place by scale is where this race is headed, the question we should ask is not about more scale, but about the nature of the threshold we cannot cross. On the evidence so far, there is no guarantee that general intelligence waits at the end of endlessly sharpening one knife.

So the question changes. If there is a threshold you can’t train your way past, two things matter: knowing when you’ve reached it, and what you do when you’re there. The race for scale doesn’t even ask — because it pretends the threshold isn’t there.

I believe there is another way to meet this threshold head-on. But that is not this essay’s job. This essay stops at opening a door — at asking which door to knock on, instead of knocking harder on the wrong one.


없는 정보는 학습으로 채워지지 않는다

키미 K3와 거대 모델 경쟁이 놓치고 있는 것

CAIDE Systems, Inc. — AI Debate · 문제 제기

2026년 여름, 중국 Moonshot AI의 Kimi K3가 세계 AI 판을 흔들었다. 오픈 웨이트로 공개된 이 모델이 미국 최전선 모델과 대등하거나 일부에서는 앞선다는 평가가 나왔다. 과학자들은 “전환점”, “스푸트니크 순간”이라고 불렀고, 출시 이틀 만에 수요가 몰려 신규 구독이 중단됐다. 실리콘밸리와 워싱턴은 긴장했다.

흥분은 이해할 만하다. 하지만 이 글의 질문은 “누가 이기고 있는가”가 아니다. 그 방향이 옳은가이다. Kimi K3도, 같은 방식으로 경쟁하는 거대 모델들도 같은 길 위에 있다. 더 많은 데이터, 더 큰 규모, 그리고 서로의 출력을 증류(distillation)해 성능을 끌어올리는 길이다. 이 길의 끝에 정말 범용지능(AGI)이 있을까? 아니면 잘못된 문을 점점 더 세게 두드리고 있는 걸까?

스케일이라는 믿음에 생긴 균열

지난 몇 년의 믿음은 단순했다. 규모를 키우면 지능이 따라온다. 데이터와 연산을 늘리면 능력이 좋아진다는 “스케일링 법칙”이다. 그런데 이 곡선이 꺾이고 있다는 신호가, 다름 아닌 그 법칙을 만든 랩 안에서 나오기 시작했다. 2024년 말, OpenAI(차세대 ‘Orion’)·구글·앤트로픽이 나란히 다음 세대 모델에서 기대한 만큼의 도약을 얻지 못했다는 보도가 이어졌다. OpenAI 공동창업자 일리야 수츠케버는 “지금 방식의 사전학습은 끝날 것”이라며 “2010년대가 스케일링의 시대였다면 이제는 다시 발견의 시대”라고 말했다.

규모를 아무리 키워도 잘 안 풀리는 문제도 드러났다. 추상적 추론을 재는 ARC 같은 과제에서는 모델을 수만 배 키워도 성능이 거의 오르지 않았다. 여기서 “거대 모델은 막다른 길”이라고 단정할 필요는 없다. 다만 규모 하나만으로는 넘지 못하는 벽이 있다는 신호로는 충분하다.

재활용은 진보가 아니다

Kimi K3가 화제이자 논란이 된 지점이 증류다. 한 모델이 다른 모델(또는 자기 자신)의 출력을 학습 재료로 삼아 성능을 빨아들이는 방식이다. 본질적으로 이미 있는 정보를 재활용하는 것이다. 물론 실제 최전선 시스템이 완전히 닫힌 것은 아니다. 실제 데이터와 사람의 피드백이 함께 섞인다. 그러니 “그들이 무너지고 있다”고 말하면 과장이다.

하지만 더 조심스럽고, 그래서 반박하기 어려운 사실이 있다. 모델이 만든 데이터를 계속 다시 학습시키면 원래 분포의 드문 부분부터 사라지며 성능이 무너진다는 것이 실험으로 확인됐다(Shumailov 외, Nature 2024). 반대로 실제 데이터를 버리지 않고 함께 쌓으면 이 붕괴가 사라진다는 것도 확인됐다(Gerstgrasser 외, 2024). 두 결과를 합치면 결론은 하나다. 실제 데이터를 쌓는 ‘축적’은 성능이 더 나빠지지 않게 지켜주기는 한다. 그러나 안정은 진보가 아니다. 진보 — 전에 없던 지식이나 발명이 실제로 새로 생기는 것 — 에는 세상에서 새로 얻은 정보가 필요하다. 실험, 관찰, 발견처럼 현실과 부딪혀야만 들어오는 정보다. 규모를 키우고 서로를 증류하는 일은 이미 가진 정보를 재배열하고 부풀릴 뿐, 이 새 정보를 만들지 못한다. 잘해야 천장을 정교하게 맴돌 뿐, 뚫지는 못한다.

두 종류의 “못한다”

여기서 핵심에 이른다. 거대 모델 경쟁이 놓치는 것은 두 종류의 “못한다”를 구분하지 못한다는 점이다.

하나는 인식적 한계다. 정보는 분명히 있는데 모델이 아직 못 꺼내는 경우다. 이건 학습을 더 하면 줄어든다. 규모와 데이터가 실제로 효과를 내는 영역이다.

다른 하나는 정보적 한계다. 정보가 애초에 없는 경우다. 이건 학습으로 줄지 않는다. 아무리 좋은 모델도 없는 정보를 만들어낼 수는 없다. 정보는 처리한다고 늘어나지 않기 때문이다. 이 벽을 넘는 길은 모델을 더 굴리는 것이 아니라, 그 정보를 실제로 담고 있는 다른 통로를 여는 것뿐이다.

그리고 이 정보적 한계는 두 곳에서 나타난다. 공간에서, 그리고 시간에서.

공간의 천장 — 지금 없는 것

첫째는 쉽다. 지금 이 순간, 어떤 통로가 아예 담지 못하는 정보다.

테슬라는 라이다 없이 카메라만으로 자율주행을 완성하려 한다. 카메라만으로도 사람만큼은 갈 수 있다. 하지만 사람이 못 보는 상황까지 막을 수는 없다. 눈이 내려 앞이 안 보이는 밤, 장애물은 카메라에 잡히지 않는다. 그 정보가 카메라라는 통로에 없기 때문이다. 어떤 신경망도 없는 것을 만들 수 없다. 라이다라는 다른 통로만이 그 장애물을 잡는다. 알고리즘의 문제가 아니라 통로의 문제다.

의료 영상도 같다. 너무 작아서 X-ray나 CT, MRI에 신호가 거의 안 남는 초기 암은, AI를 더 학습시켜서 잡는 것이 아니다. 그 문턱은 학습으로 넘는 것이 아니라 다른 방법 — 조직검사, 추적검사 — 으로 넘어야 한다. 없는 정보를 다른 통로가 가져와야 한다.

이 축은 반박이 어렵다. 정보가 물리적으로 없으면, 그 위에서 아무리 계산해도 소용없다.

시간의 천장 — 아직 없는 것

거대 언어모델은 얼핏 사정이 다르다. 텍스트에는 정보가 없지 않다. 오히려 인류가 쓴 거의 모든 것이 들어 있다. 그래서 “LLM은 정보가 없어서 못 한다”는 말은 틀렸다. (사람들은 흔히 “LLM은 추론을 못 한다”고도 하지만, 그것은 논쟁적이고 이 글의 요점도 아니다.)

진짜 문제는 다른 데 있다. 어떤 지식이 지식이 되는 순간은 현실과 부딪혀 검증되는 순간이다. 그런데 그 검증은 아직 오지 않은 미래에 일어난다. 내일의 현실은 아직 존재하지 않으니, 그것을 담은 데이터도 없다. 어떤 데이터에도 아직 없는 정보는, 그 위에서 아무리 계산해도 나오지 않는다.

이건 공간의 천장과 다른 원리가 아니다. 그것을 시간으로 늘린 것이다. 라이다가 지금 없는 정보를 가져오듯, 새로운 지식은 아직 없는 정보를 오직 현실과의 새 접촉으로만 가져온다. 규모도, 증류도, 재활용도 그 접촉을 대신하지 못한다. 자율주행과 거대 모델의 한계는 공간과 시간이라는 두 얼굴을 한 같은 벽이다.

알파폴드가 그은 선

이건 추상적인 이야기가 아니다. AI 과학의 가장 빛나는 성취인 알파폴드가 이 선을 그대로 보여준다.

알파폴드는 단백질의 3차 구조를 훌륭히 예측했다. 왜 됐을까. 그 구조를 결정하는 정보 — 수억 년의 진화가 서열에 남긴 흔적 — 가 이미 데이터에 있었기 때문이다. 정보가 있으니 계산이 꺼냈다.

그리고 알파폴드는 정보가 없는 곳에서 정확히 멈춘다. 참고할 사례가 하나도 없는 완전히 새로운 형태의 단백질은 예측하지 못한다. 진화의 흔적이 없기 때문이다. 단백질이 어떻게 움직이는지도 마찬가지다. 2026년 연구자 수십 명이 낸 보고서는 그 원인을 분명히 짚었다. 모델이 부족해서가 아니라, 단백질의 움직임을 기록한 데이터 자체가 정적 구조 데이터보다 수천 배 이상 적기 때문이다. 이건 더 큰 모델로 메울 공백이 아니다. 실제로 측정한 새 데이터가 있어야 메워진다.

알파폴드는 이 글을 그대로 실연한다. 정보가 있으면 계산이 풀고, 없으면 규모로도 못 만든다. 선은 늘 정보가 통로를 떠나는 자리에 있다.

구조를 알아도 약이 되지는 않는다

이 선은 신약에서 더 분명해진다. 구조를 안다고 약이 나오는 건 아니다.

숫자가 말한다. 2019년 이후 AI 신약에 약 600억 달러가 들어갔고 후보 175개가 임상에 올랐지만, 2026년 현재 FDA 승인은 하나도 없다. AI는 후보 설계를 3~4년에서 1년으로 줄였을 뿐, 약이 진짜 듣는지 가리는 임상 2상 성공률(약 40%)도 전체 탈락률(90%)도 AI 이전과 똑같다.

이유는 간단하다. “이 물질이 사람 몸에서 병을 낫게 하는가”의 답은 어떤 데이터에도 없다 — 실제로 임상을 돌리기 전까지 그 정보는 세상에 존재하지 않기 때문이다. AI는 벽 앞까지 데려다줄 뿐 벽은 못 뚫는다. 병목은 사라진 게 아니라 뒷단으로 옮겨갔고, 그 뒷단은 여전히 연구자와 실험의 몫이다.

그래서 잘못 가고 있는가

반대편 목소리도 강하다. 긴 사고사슬과 강화학습을 붙인 추론 모델이 이미 성능을 끌어올리고 있고, 손으로 짠 구조보다 규모로 미는 일반적 방법이 결국 이긴다는 “쓴 교훈”도 있다. “스케일이 곧 벽에 부딪힌다”는 예측이 과거에 자주 틀렸다는 것도 사실이다. 이 겸손은 지켜야 한다.

그러나 그 반박들은 핵심을 건드리지 못한다. 추론 모델이 늘리는 것은 같은 재료 위에서의 처리일 뿐, 아직 세상에 없는 정보를 새로 들여오는 일이 아니다. “쓴 교훈”은 정보가 이미 있을 때 참이다. 정보가 아직 없는 곳에서는, 규모가 클수록 더 비싼 거울일 뿐이다.

거대 모델의 성취가 가짜라는 말이 아니다. 그것은 강력한 도구다. 다만 칼을 아무리 갈아도 칼이 닿지 않는 곳은 다른 도구가 필요하다. 지금의 경쟁은 하나의 칼을 무한히 가는 경주처럼 보인다.

그래서 Kimi K3로 돌아가자. 그 모델을 아무리 키우고 서로 증류해도, 결국 임상의 벽 앞에서 멈춘 그 신약들과 같은 자리에 선다. 앞단은 눈부시게 빨라지지만, 진짜 새로운 지식이 태어나는 자리 — 현실과 부딪혀야만 답이 나오는 자리 — 는 그대로 남는다. 그 자리를 규모로 넘으려는 것이 지금 경쟁의 방향이라면, 우리는 더 큰 규모가 아니라 넘지 못하는 문턱의 정체를 물어야 한다. 적어도 지금까지의 증거로는, 하나의 칼을 무한히 가는 경주의 끝에 범용지능이 기다린다는 보장은 어디에도 없다.

그러면 질문이 바뀐다. 학습으로 넘을 수 없는 문턱이 있다면, 정작 중요한 것은 두 가지다. 언제 그 문턱에 닿았는지를 아는 것, 그리고 그 앞에서 무엇을 하는지. 규모를 키우는 경주는 이 질문을 던지지도 않는다. 문턱이 없는 척하기 때문이다.

나는 이 문턱을 정면으로 다루는 다른 길이 있다고 생각한다. 그러나 그것은 이 글의 몫이 아니다. 이 글은 문을 여는 데서 멈춘다 — 잘못된 문을 더 세게 두드리는 대신, 두드려야 할 문이 어디인지 묻는 자리에서.

References: Shumailov et al., “AI models collapse when trained on recursively generated data,” Nature 631 (2024). · Gerstgrasser et al., “Is Model Collapse Inevitable?” arXiv:2404.01413 (2024). · Sutskever, remarks on the end of pre-training scaling (2024). · Public reporting on AI drug discovery outcomes, 2026.