World Models

LLMs are ngmi. Next-token prediction understands nothing. A truly intelligent agent needs a world model, spatial intelligence, and signals beyond text.

The Thesis

This is Kevin's strongest technical opinion. Large language models are not a path to general intelligence. They are capable pattern matchers doing next-token prediction, and next-token prediction is fundamentally insufficient for genuine understanding. The industry treats LLMs as the foundation for AGI. Kevin thinks they're a tool - an impressive, useful tool - but a dead end for the goal they're being aimed at. Source: User, 2026-05-17

The Biological Argument

The evidence is in how humans actually acquire intelligence, and it starts before birth.

Babies at three years old can understand whole sentences. They can parse grammar, follow narratives, respond to questions. But they lack object permanence, a sense of time, understanding of sequences and patterns. Language comes first; comprehension of the physical world comes much later and with enormous effort. The acquisition process itself reveals how much more there is to intelligence beyond language. Source: User, 2026-05-17

Even more telling: as fetuses, babies already pick up on patterns of cadence and tone variation in their mother's speech. This is why children acquire their mother's language faster and more naturally - they've been pre-training on its acoustic structure for months before birth. The biological learning process starts with non-linguistic signal processing. Rhythm, tone, variation. Not tokens. Not text. Sound patterns that encode structure without encoding meaning. Source: User, 2026-05-17

Young children spend enormous brain power on language acquisition. It's one of the most computationally expensive things the human brain ever does. Our biology is incredibly capable but also incredibly complex - and language is just one subsystem of the full cognitive architecture. An LLM gets the language part (sort of) and nothing else.

What's Missing

A truly capable and intelligent agent must have:

  • A world model - an internal representation of how the physical world works, not just how text patterns flow. Object permanence. Cause and effect. Spatial relationships.
  • Spatial intelligence - understanding of 3D space, navigation, embodiment. Knowing that a cup on a table can fall. Knowing that behind a door is a room.
  • Multi-modal signal processing - responding to signals beyond text tokens. Sound, vision, touch, proprioception. The fetus processes acoustic patterns before it processes language. Intelligence starts below the level of words.
  • Multi-modal output - producing signals beyond text. Actions, movements, manipulations. Understanding the world means being able to act on it, not just describe it.

LLMs have none of this. They predict the next token in a sequence. They produce text that reads like understanding because the training data was written by beings who actually understand. The fluency is borrowed. Source: User, 2026-05-17

The Research Gap

It will take vastly more research into how skill acquisition actually works - how humans go from pre-linguistic pattern recognition to embodied intelligence - to determine what a truly intelligent model even looks like. The current approach of scaling next-token prediction is optimizing the wrong objective. More parameters, more data, more compute - none of it addresses the fundamental gap between text fluency and world understanding. Source: User, 2026-05-17

This connects to Kevin's research direction with Professor Danqi Chen on agent orchestration. The question isn't "how do we make LLMs smarter?" It's "what are the other components that a capable agent needs beyond language, and how do we orchestrate them?" The language model is one module, not the whole system. Source: User, 2026-05-17

Connections

Connection Implication
Services-as-Software LLMs are useful now for services-as-software (selling outcomes via text-based automation). But that's a business thesis about current tools, not a claim about intelligence
The Brain-Agent Loop The wiki's brain-agent-loop uses LLMs as a knowledge interface - reading and writing text. It works precisely because the task is text manipulation, not world understanding

Timeline

  • 2026-05-17 | Page created. Kevin's strongest technical opinion: LLMs are ngmi for genuine intelligence. The biological argument from language acquisition - babies process acoustic patterns before birth, acquire language before object permanence - reveals how much more than text prediction is needed. Connects to research with Professor Danqi Chen. Source: User, 2026-05-17