Introduction
In recent years, remarkable advances have been made in artificial intelligence, particularly in deep learning and large language models. The success of models such as GPT, computer vision systems, and many reinforcement learning algorithms shows that neural-network-based methods can solve a wide range of problems with considerable accuracy. Despite these achievements, a fundamental question remains: do such systems genuinely possess intelligence, or can their behavior be understood primarily as the result of highly complex statistical learning?
Yann LeCun raises this question as part of his broader critique of the current direction of artificial intelligence. From his perspective, the issue is not simply processing power, model size, or the amount of training data; it also concerns the learning paradigm itself. In other words, LeCun argues that the way today's machines learn differs in important respects from how humans and other intelligent animals learn about the world. The question is therefore not merely an engineering problem, but also one connected to cognition, learning, and the nature of intelligence.
The Dominant Paradigm in Today's AI
A large part of modern AI relies on gradient-based optimization: the parameters of a model are updated according to an objective or cost function using gradient information.
The learning objectives of language models, computer vision systems, and reinforcement learning agents are not identical, but many of these systems share a general pattern. A model produces an output or behavior according to its current state, this behavior is evaluated against an objective or feedback signal, and the model's parameters are then updated to improve performance.
📐 From a mathematical perspective, the core idea can be summarized as
"Change the model's parameters so that the objective or cost is reduced."
This process can be repeated many times until the model reaches the desired level of performance.
Reinforcement Learning: Experience Instead of Correct Answers
In reinforcement learning, the correct answer is not specified in advance for every state. Instead, an agent interacts with an environment and changes its behavior according to feedback. The agent selects an action, observes its consequence, receives a reward or penalty, and then updates its policy based on that feedback.
🔄 The Reinforcement Learning Cycle
Action → Environment → Reward → Update
In many reinforcement-learning settings, the agent must interact with the environment to discover the consequences of its actions. This dependency, particularly in model-free reinforcement learning, can result in a very large number of trials and interactions. LeCun explicitly describes model-free RL as extremely sample-inefficient when compared with human and animal learning.
From LeCun's perspective, the problem is that when learning depends heavily on direct trials and limited feedback, an agent may need to experience the consequences of many actions in the real world. He therefore argues for learning mechanisms that can acquire much of the necessary knowledge by observing and modeling the world, rather than by repeatedly testing every possibility through physical interaction.
Critique One: Dependency on Direct Interaction with the World
LeCun's first criticism concerns the strong dependence of many learning systems, especially in model-free settings, on interaction with the environment. In such systems, the agent acquires part of its knowledge by performing actions and observing their consequences.
LeCun's argument, paraphrased: Interaction with the real world is costly and can be dangerous in some applications. An intelligent agent should therefore be able to learn as much as possible by observation and internal modeling, reducing the need for costly direct trials.
— Based on LeCun's discussion in A Path Towards Autonomous Machine Intelligence
LeCun views this as an important scalability problem. In many real-world applications, direct experience is not merely time-consuming; it may also be expensive or potentially dangerous. Learning through observation and prediction can therefore help reduce the number of interactions that an agent needs in order to acquire a useful skill.
Critique Two: The Lack of a World Model
One of the most fundamental concepts in LeCun's view is the "World Model." In his proposal, an intelligent agent should have some internal and usable representation of how the world changes over time.
The point is not that every contemporary system necessarily lacks every form of internal representation. Rather, LeCun's criticism is directed at approaches that do not make explicit, predictive modeling of the world's dynamics a central part of learning. A system may learn highly complex statistical relationships between observations, states, and actions without necessarily having a world model that can be effectively used to simulate future states.
LeCun's idea, paraphrased: One path toward machine intelligence may lie in the ability of humans and animals to learn world models—internal models that describe how the world works.
— Yann LeCun, A Path Towards Autonomous Machine Intelligence
From this perspective, what we call "common sense" depends substantially on internal knowledge about how objects, events, and processes behave. Such knowledge can help an agent make better predictions about what is likely, what is possible, and what is unlikely to occur.
Critique Three: Inability to Mentally Simulate the Future
Another important difference emphasized by LeCun is the ability to predict the consequences of actions before actually executing them. Humans can imagine several possible scenarios, evaluate their likely outcomes, and then choose an action.
LeCun argues that an autonomous intelligent agent should be able to use its world model in a similar way. Instead of physically testing every alternative, the agent should be able to predict possible future states internally and use those predictions to guide its decision.
In his proposed architecture, the world model can be used to predict possible future states and provide a basis for planning and searching through alternative action sequences. In this sense, the system can evaluate consequences inside an internal predictive model before committing to an action in the real world.
Critique Four: Low Sample Efficiency
One direct consequence of heavy dependence on experience is the potentially large amount of data and environmental interaction required for learning. LeCun specifically argues that model-free reinforcement learning is extremely sample-inefficient compared with human and animal learning.
In many real-world applications, this amount of trial and error is not practical. A robot cannot repeatedly fail millions of times to learn a physical skill, and an autonomous vehicle cannot discover dangerous consequences by repeatedly experiencing real accidents.
📊 Sample Efficiency
From LeCun's perspective, a major motivation for world-model-based architectures is to learn a large amount of knowledge through observation and prediction, thereby reducing the number of direct interactions required to acquire skills.
This makes the gap between model-free reinforcement learning and the sample efficiency associated with human and animal learning an important research question on the path toward autonomous intelligence.
Critique Five: Over-Reliance on Reward Signals
In many reinforcement learning algorithms, reward plays an important role in guiding the learning process. The agent modifies its policy according to positive or negative feedback received from the environment.
LeCun argues that such a signal is not sufficient by itself to provide all the information needed to learn the complex structure required by an intelligent agent. In his discussion, scalar rewards provide low-information feedback and therefore cannot replace extensive learning about the structure of the world.
LeCun's argument, paraphrased: Reward is a low-information signal, so it should not be expected to teach a complex system all the knowledge required for intelligent behavior by itself.
— Based on the section "Reward is not enough" in A Path Towards Autonomous Machine Intelligence
In LeCun's proposal, world-model learning through prediction can train a much larger portion of the system, while reward or intrinsic cost can be used primarily to guide behavior, specify objectives, and evaluate possible future trajectories.
Critique Six: Limitations in Reasoning and Planning
LeCun's final major criticism concerns the limited forms of reasoning and planning available in many current systems. He argues that some existing models primarily establish mappings from observations to outputs or actions, leaving relatively limited room for richer search over alternative interpretations and possible courses of action toward a goal.
In contrast, the kind of agent described by LeCun should be able to operate with goals at multiple levels, explore alternative action sequences, predict their consequences using a world model, and then select an appropriate sequence of actions.
In LeCun's paper, this form of reasoning is connected to search through energy minimization or constraint satisfaction, where the actor searches for suitable combinations of actions and latent variables. He also points out that the absence of abstract latent variables can limit the ability of a model to explore multiple interpretations of a percept and search for appropriate courses of action toward a goal.
Conclusion
An analysis of Yann LeCun's perspective shows that his criticism is not simply a rejection of the engineering achievements of deep learning. Rather, it concerns the limitations of particular learning paradigms and cognitive architectures. From his point of view, the central challenge is not merely to make models more accurate, but to build systems that can observe the world, learn an internal model of it, and use that model for prediction and decision-making.
Within this framework, heavy dependence on direct interaction, low sample efficiency in model-free reinforcement learning, strong reliance on reward signals, the absence of a useful world model, limited future simulation, and difficulties in reasoning and planning all point toward a deeper question: can learning be organized so that a large portion of the knowledge required by an intelligent agent is acquired through observation and prediction before all possible actions must be tested in the real world?
🎯 LeCun's Key Message
The next generation of AI, in LeCun's vision, should not rely only on learning statistical mappings. It should also be able to learn the structure of the world, predict future states, explore alternative scenarios inside an internal model, and evaluate the consequences of decisions before acting in the environment.
In this perspective, the world model is not merely an auxiliary component. It is a central part of the architecture required for efficient learning, planning, and autonomous intelligence.