Introduction: From Simple Mapping to an Integrated System
So far, in previous articles, we have discussed LeCun's philosophy, his critiques of current systems, and the core idea of "reasoning as energy optimization." But perhaps the most important part of LeCun's paper is where he moves from theory to practice and proposes a concrete cognitive architecture for an autonomous intelligent agent.
When I first read this part of the paper, a sentence from LeCun caught my attention. At the beginning, he emphasizes that this writing is not a description of a ready-made system:
"This paper is not a description of a system, but a proposal for an architecture and a research agenda."
In other words: "This paper is not a description of a complete system; it is a proposal for an architecture and a research agenda."
LeCun does not claim to have solved everything. He is drawing a roadmap—one that, by his own admission, still has much room for improvement. But what distinguishes this roadmap from other proposals is the integration and coordination of its components.
To understand this architecture in relation to the broader ideas in LeCun's work, it is useful to read this article alongside LeCun's philosophy of World Models and the relationship between EBM, JEM and JEPA.
1. Fundamental Difference from Current Systems
A significant portion of current AI approaches, especially model-free approaches in reinforcement learning, attempt to directly map perception to action. In these approaches, a function is learned that directly converts observation to action:
This mapping has been successful in many applications. But LeCun believes that for truly autonomous intelligence, this approach is insufficient. Because between observation and action, there is a complete cognitive world that is ignored in this simple mapping: an internal model of the world, memory, goals, and mechanisms for planning and reasoning.
For this reason, LeCun proposes a multi-module architecture where each module has a specific role and all interact with each other.
2. Components of the Intelligent Agent Architecture
LeCun's proposed architecture includes eight main components:
- Configurator
- Perception
- World Model
- Short-Term Memory
- Intrinsic Cost
- Critic
- Cost
- Actor
The relationship between these components is not a simple linear chain. It is a dynamic system in which:
- Actor proposes action sequences to the World Model.
- World Model predicts possible consequences of those imagined actions.
- Cost / Critic estimate future costs.
- Actor optimizes the action sequence based on these evaluations.
This cycle can be schematically represented in the figure below.
3. Configurator: The System-Wide Regulator
It may seem surprising, but LeCun does not start his architecture with the Configurator. However, conceptually, this module plays a key role in the system.
The Configurator's job is to configure the other modules for the task the agent is currently performing. In complex planning, it may also need to determine sequences of subgoals.
The Configurator can be conceptually compared to some ideas related to Executive Control in cognitive science. But this is my interpretive comparison, not LeCun's explicit equivalence.
LeCun describes this module as one of the least understood parts of the proposed architecture:
"Of all the least understood aspects of the current proposal, the configurator module is the most mysterious."
In other words, the Configurator determines what matters for the task at hand and modulates the other modules accordingly.
4. Perception: Estimating the State of the World
After the Configurator sets the priorities, it is time for Perception.
Perception in LeCun's architecture is not meant merely to pass raw sensory data. Its job is to estimate an abstract representation of the current state of the world—a representation usable by the World Model and other parts of the system.
More precisely, Perception should be able to:
- Extract a compressed representation from sensory data such as images, sound, or text.
- Include information relevant to future prediction and decision-making.
- Ignore details that are irrelevant to the current task.
5. World Model: The Internal Model of the World
And now we come to the most important part of the architecture.
In LeCun's formulation, the World Model is the most complex component of the architecture. Its role is not simply to store information, but to model the relevant aspects of the world and predict possible future states.
What exactly is the World Model?
The World Model is an internal, predictive model of how the state of the world may evolve. It can estimate missing information and predict plausible future states, including outcomes resulting from imagined action sequences.
In this sense, the World Model behaves like a simulator of the relevant aspects of the world. It does not need to reconstruct every detail; the relevant representation depends on the task.
Key Features of the World Model
1. Prediction in Abstract Space
The World Model is not meant to reconstruct all details of the world at the pixel level. Predictions can instead be performed in an abstract representation space, potentially across multiple levels of abstraction and different time horizons.
2. Representing Multiple Futures
The world is not fully predictable. Therefore, the World Model should be able to represent multiple plausible futures, including futures parameterized by latent variables that represent uncertainty about the world state.
This point also connects directly to the discussion of Latent Variables in LeCun's architecture.
3. Ability to Learn from Observation
The World Model should be capable of learning from passive observation, without requiring an agent to interact with the environment for every aspect of its learning.
Capabilities Derived from This Architecture
Several important capabilities can be derived from LeCun's architecture:
- Prediction: Predicting the consequences of actions.
- Planning: Finding a sequence of actions to achieve a goal.
- Reasoning: Using internal models and optimization to solve problems.
- Imagination: Simulating different possible future scenarios.
- Hierarchical Representation: Representing the world at different levels of abstraction.
These are capabilities that can be derived from the architecture, rather than an official list of five World Model capabilities given by LeCun.
6. Short-Term Memory
In LeCun's architecture, Short-Term Memory is an independent component. It is responsible for keeping track of:
- The current state of the world.
- States predicted by the World Model.
- Intrinsic costs associated with those states.
Therefore, the cognitive cycle is not simply Perception → World Model → Actor. Memory plays an important role in maintaining context and connecting the current state with predicted future states.
7. Cost: A Function for Measuring Energy / Discomfort
When the World Model has constructed several possible futures, the agent needs a mechanism for evaluating them. This is where the Cost Module comes in.
In LeCun's architecture, Cost is composed of two sub-modules:
The Cost Module produces a scalar output called energy, interpreted as a measure of the agent's level of discomfort.
1. Intrinsic Cost
This part is immutable and non-trainable. It represents intrinsic drives and immediate costs associated with the current state.
LeCun describes examples such as pain, pleasure, and hunger as possible manifestations of intrinsic cost.
2. Critic
Critic is a trainable module that predicts future values of the Intrinsic Cost based on the agent's experience.
Therefore, the architecture does not reduce evaluation to an externally supplied reward signal. Instead, future consequences are evaluated through the internal cost mechanism formed by Intrinsic Cost and Critic.
8. Actor: Proposer and Optimizer of Action Sequences
Finally, we come to the Actor. Rather than simply being a conventional decision layer, the Actor has two important roles:
1. Proposer
The Actor proposes action sequences.
2. Optimizer
Using evaluations produced by the World Model and Critic, the Actor searches for an action sequence that minimizes estimated future cost.
LeCun discusses several possible methods for finding an optimal action sequence, including:
- Gradient-Based Optimization
- Graph Search
- Dynamic Programming
- A* Search
- Monte Carlo Tree Search (MCTS)
Therefore, the architecture should not be presented as if the Actor necessarily relies only on Gradient Descent.
Conceptually, the Actor seeks an action sequence that minimizes predicted future energy or cost:
where H is the planning horizon.
9. Differentiability: A Key Feature of the Architecture
One of the important features of LeCun's architecture is differentiability.
The trainable modules are designed so that gradient estimates of the scalar cost can propagate through the architecture. This creates the possibility of training interconnected modules through gradient-based methods.
However, an important distinction must be maintained: Intrinsic Cost is defined as immutable and non-trainable. Therefore, not every component of the architecture is trainable in the same way.
10. Conclusion: A Multi-Module Architecture for Autonomous Intelligence
If I were to summarize LeCun's architecture in a few sentences:
Summary of LeCun's Architecture
Instead of proposing a single network, LeCun proposes a multi-module architecture for autonomous intelligence; an architecture where perception, memory, world modeling, cost evaluation, and action selection interact within a coordinated system.
The architecture consists of eight central components:
- Configurator: Configures the rest of the system for the current task.
- Perception: Estimates the current state of the world.
- World Model: Models and predicts relevant aspects of the world.
- Short-Term Memory: Maintains current and predicted states.
- Intrinsic Cost: Immutable, non-trainable intrinsic cost.
- Critic: Predicts future intrinsic costs.
- Cost: Produces the scalar energy used for evaluating outcomes.
- Actor: Proposes and optimizes action sequences.
Perhaps the most important lesson I take from this architecture is that LeCun presents it as a research direction, not as a finished technical blueprint. Its value is therefore not simply in the individual modules, but in the way the modules are assembled into a unified framework for perception, prediction, evaluation, planning, and action.
"This document is ... expressing my vision for a path towards intelligent machines."
And this is ultimately what makes the architecture interesting to me: it tries to move the discussion of intelligence away from a single input-output mapping and toward an integrated system that can represent the world, imagine possible futures, evaluate them, and use those evaluations to guide action.
Frequently Asked Questions About LeCun's Autonomous Intelligence Architecture
What is Yann LeCun's autonomous intelligence architecture?
It is a proposed multi-module architecture for autonomous intelligent agents that combines perception, world modeling, short-term memory, intrinsic cost, a critic, cost evaluation, an actor, and a configurator.
What is the role of the World Model?
The World Model estimates missing information about the state of the world and predicts plausible future states, including outcomes associated with imagined action sequences.
What does the Configurator do?
The Configurator configures the other modules for the task at hand. LeCun also describes it as one of the least understood and most mysterious parts of the proposed architecture.
What is the difference between Intrinsic Cost and Critic?
Intrinsic Cost is defined as immutable and non-trainable, representing immediate intrinsic costs. The Critic is trainable and predicts future values of intrinsic cost.
Is the Actor simply a decision-making neural network?
Not exactly. In LeCun's formulation, the Actor proposes action sequences and searches for sequences that minimize estimated future cost. Several optimization and search methods can potentially be used.
How does this architecture relate to Energy-Based Models and JEPA?
Energy-based formulations are relevant to the Cost component and to LeCun's broader approach to reasoning and decision-making, while JEPA is connected to predictive representation learning and World Model construction. The concepts are related but should not be treated as identical architectures.
References & Further Reading
- LeCun, Y. (2022). A Path Towards Autonomous Machine Intelligence . OpenReview. https://openreview.net/pdf?id=BZ5a1r-kVsf
- Assran, M., et al. (2023). Self-Supervised Learning From Images With a Joint-Embedding Predictive Architecture . CVPR.
- Dawid, A., & LeCun, Y. (2024). Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence . Journal of Statistical Mechanics: Theory and Experiment, 2024, 104011.