Artificial Intelligence · Deep Learning

What Is EBM? Energy-Based Models, Energy Functions, and JEPA

Fundamental Concepts of Energy-Based Models in AI Architectures
Energy-Based Model EBM Energy-Based Learning Latent Variable LV-EBM JEPA EB-JEPA Yann LeCun World Model Self-Supervised Learning

1. Introduction: Energy, But Not What You Think

When we first encounter the word Energy in machine learning literature, we naturally think of physics: conservation laws, joules, and calories. But in Energy-Based Models (EBMs), that interpretation is not only unnecessary; it can be misleading.

Key Point

In an EBM, Energy is a mathematical compatibility score, not a physical quantity.

Yann LeCun, a Turing Award winner and one of the pioneers of deep learning, presents the Energy Function as the fundamental object of the model in his influential paper "A Path Towards Autonomous Machine Intelligence" (2022). He also emphasizes that an EBM does not have to be interpreted as a conventional probabilistic model.

"The energy function is viewed as the fundamental object and is not assumed to implicitly represent the unnormalized logarithm of a probability distribution."

In simple terms, an EBM gives the system a way to ask: how compatible is this candidate state, answer, representation, or prediction with the input and the constraints?

Energy-Based Model as a compatibility function in artificial intelligence
Figure 1: Energy as a compatibility function in an Energy-Based Model (EBM)

2. What Is EBM? What Does an Energy-Based Model Actually Model?

In the context of artificial intelligence, EBM usually refers to Energy-Based Model or Energy-Based Models. Instead of requiring a normalized probability for every possible output, an EBM defines an Energy Function that scores configurations according to their compatibility.

Conceptual Map of EBM

EBM → Energy Function → Energy Landscape → Energy Minimization → Inference → Latent Variable → Joint Embedding → JEPA

This chain shows that EBM is not merely a method for computing a number. It is a framework for defining compatibility, performing Inference and, in some architectures, predicting in a Latent Space and supporting planning.

Term Role in EBM Related Concepts
Energy-Based Model General modeling framework based on an energy function EBM, Energy-Based Models, Deep EBM
Energy Function Scores the compatibility of a configuration Energy Function, Energy-Based Learning
Energy Landscape Describes how energy varies across candidate states Energy Landscape, Energy Minimization
Latent Variable EBM Introduces hidden variables that explain unobserved structure LV-EBM, Latent Variable Energy-Based Model
JEPA Predicts representations rather than reconstructing raw inputs Joint Embedding Predictive Architecture, JEPA

One important distinction is that Energy is not physical energy. Another is that an EBM should not automatically be treated as a probability model. Some formulations connect energy to probability distributions, but in LeCun's compatibility-based view the Energy Function itself can be the primary object.

3. Mathematical Definition of an Energy-Based Model

An Energy-Based Model can be expressed as a function:

Fw : 𝒳 × 𝒴 → ℝ

where:

  • 𝒳: input space, such as images, text, or sensory observations
  • 𝒴: output space or candidate states, answers, labels, or predictions
  • w ∈ ℝd: model parameter vector

For each pair (x, y), the Energy Function returns a scalar:

E(x, y) = Fw(x, y)

The training objective is to shape the function so that compatible configurations receive lower energy than incompatible ones:

Fw(x, ytrue) ≪ Fw(x, yfalse)

4. Latent Variables, Free Energy, and the Connection Between EBM and JEPA

So far we have assumed that (x, y) directly describes the state of interest. But many real problems contain hidden factors that are not directly observed. This is where the idea of a Latent Variable becomes important.

A latent variable z can represent hidden factors or structure that are not directly visible in the input but are useful for explaining the relationship between x and y.

In a Latent Variable Energy-Based Model (LV-EBM) , energy can be extended to include the latent state:

E(x, y, z)

The observable pair (x, y) can then be evaluated through possible latent configurations. This leads to the related concept of Free Energy, which summarizes the contribution of latent explanations.

F(x, y) = −(1/β) log ∫ exp(−βE(x, y, z)) dz

In this view, Latent Space is not merely a compression space. It can be a space in which the model represents structure that matters for prediction, inference, and world modeling.

EBM → Latent Variable → JEPA

In this view, EBM provides a compatibility-based language, the Latent Variable provides hidden explanatory structure, and JEPA uses representation-space prediction to model what is predictable without requiring full input-space reconstruction.

In the framework presented by Dawid and LeCun, the advantages of Energy-Based Models and Latent Variable Models are combined in the building blocks of Hierarchical JEPA (H-JEPA).

The central distinction is useful: a conventional generative model may attempt to predict the raw pixels of a future image, while JEPA aims to predict a representation of the future. This shifts the objective toward meaningful, predictable structure in representation space rather than reconstructing every low-level detail.

Recent work makes this connection even more explicit. In 2026, the EB-JEPA open-source project and its corresponding paper introduced a library for Energy-Based Joint-Embedding Predictive Architectures , with examples spanning image representation learning, video prediction, and action-conditioned world-model planning.

5. Energy Landscape

The set of energy values over all candidate y's forms an Energy Landscape:

ℰ(x) = { Fw(x, y) | y ∈ 𝒴 }
Energy Landscape with low-energy valleys and high-energy peaks
Figure 2: Schematic Energy Landscape showing low-energy valleys and high-energy regions
  • Valleys → configurations consistent with the input, data, or constraints
  • Peaks → configurations that are less compatible

6. Energy in Discrete and Continuous Spaces

Important Note

Continuity of Energy is not required in order to define an EBM.

An EBM can be defined on a discrete space. For a classification problem:

E(x, k) = Fw(x, k)   where   k ∈ {1, 2, ..., K}

If Energy is differentiable and defined on a continuous space, gradients can be used to search for low-energy configurations:

∇y E(x, y)

7. Gradient, Energy Minimization, and Inference-as-Optimization

If the Energy Function is defined on a continuous space and is differentiable:

∇y E(x, y) = [ ∂E/∂y1, ∂E/∂y2, ..., ∂E/∂yd ]T

To move toward low-energy points:

yt+1 = yt − η ∇y E(x, yt)

Therefore:

y* = arg min y ∈ 𝒴 E(x, y)

The inference perspective can also be connected to reasoning as optimization. For further discussion, see Reasoning and Gradient-Based Learning .

"Reasoning can be formulated as energy minimization under constraints."

8. Energy Collapse and Representation Collapse

If the model cannot preserve sufficient differences between compatible and incompatible configurations, the Energy Landscape can become excessively flat:

E(x, y) ≈ Constant

This is related to Energy Collapse. A different but related issue in representation learning is Representation Collapse, where different inputs are mapped to overly similar representations.

Architecture Collapse Risk? Explanation
Deterministic Predictive Architecture Not in the specific formulation discussed by LeCun In the formulation under discussion, direct predictive matching avoids the collapse mechanism being considered.
Latent-Variable Generative Architecture Yes Degenerate solutions may emerge when latent-variable structure permits them.
Autoencoder Possible Representation quality depends on the training objective and constraints.
Joint-Embedding Architecture Possible Encoders may ignore meaningful input variation without appropriate constraints.

9. Main Approaches for Training EBMs

Training Energy-Based Models is one of the technically demanding parts of the field. The broader EBM literature includes Maximum Likelihood, MCMC, Langevin Dynamics, Score Matching, and Noise Contrastive Estimation (NCE).

9.1. Contrastive Methods

The basic idea is to lower the energy of positive samples and raise the energy of contrastive or negative samples.

ℒ(w, x, y, ŷ) = [ E(x, y) − E(x, ŷ) + m(y, ŷ) ] +

where:

  • y: positive or compatible sample
  • ŷ: contrastive sample
  • m(y, ŷ): margin function
  • [a]+ = max(0, a): positive-part operator

A well-known related objective in representation learning is InfoNCE.

9.2. Regularized Methods

Regularized methods lower the energy of compatible samples while also constraining the geometry or volume of the low-energy region.

ℒ(w, x, y) = E(x, y) + λ · ℛ(w)

Here ℛ(w) represents a regularization term designed to prevent degenerate solutions and shape the energy landscape.

Why Regularization Matters

In LeCun's 2022 discussion, regularized methods are presented as a promising way to avoid the dimensionality problems associated with purely contrastive approaches:

"Regularized methods are much more promising in the long run than contrastive methods because they can eschew the curse of dimensionality that plagues contrastive methods."

10. Energy as a Compatibility Metric in a World Model

In LeCun's broader proposal for Autonomous Machine Intelligence, an intelligent agent needs internal representations, prediction, and the ability to evaluate possible future states.

  1. Observe the world.
  2. Build an internal representation.
  3. Predict future states with a World Model.
  4. Evaluate possible consequences of actions.
  5. Select or search for actions under constraints and objectives.

For a broader discussion of LeCun's view of world models, see LeCun's Philosophy of World Models .

E(st, at, st+1) ∝ −Compatibility(st, at, st+1)

A simplified planning objective can be written as:

at* = arg min at Σt' E( st', at', st'+1 )

In other words: find a sequence of actions that minimizes the energy or cost of the trajectory.

11. Energy vs Reward vs Cost

Concept Energy Reward
Role Compatibility, inference, prediction, or planning objective Feedback signal used to evaluate behavior, commonly in reinforcement learning
Definition Usually learned or defined as an Energy Function over configurations Specified by the environment, task, or reward design
Optimization Direction Often minimized Usually maximized
Domain Can be continuous or discrete Can be continuous or discrete

In some designs, a task-specific transformation can make:

Reward ≈ −Energy

but this does not establish a general equivalence between the two concepts.

"Contrary to the title of a recent position paper by Silver et al. (2021), the reward plays a relatively minor role in this scenario."

Cost is another related term. It often denotes an objective that should be minimized, while the precise interpretation depends on the model and its semantics.

12. Conclusion: EBM, Latent Space, and JEPA

In this article we saw that EBM (Energy-Based Model) offers a different way to build intelligent systems. Instead of requiring the model to directly generate every answer, we can learn an Energy Function that scores compatibility and then perform Energy Minimization to find a suitable configuration.

Once we introduce Latent Variables and Latent Space, energy can also represent hidden structure that is not directly observable. This connects to Joint Embedding and JEPA, where the system predicts in representation space rather than reconstructing every detail of the raw input.

Conceptual Map

Energy-Based Model → Energy Function → Energy Landscape → Inference as Optimization → Latent Variable / Latent Space → Joint Embedding → JEPA → World Model → Planning

World 1: Reasoning World 2: Learning
"If... then..." constraints Gradient-based optimization
Discrete structure Continuous neural representations
Constraints Loss functions and regularizers

Key Message

Energy provides a common mathematical language for compatibility, inference, prediction, and planning—especially when these operations are carried out in a learned latent or representation space.

13. Frequently Asked Questions About EBM, Latent Variables, and JEPA

What is EBM?

EBM stands for Energy-Based Model. It assigns energy values to configurations or input-output pairs so that compatible states can be represented by lower energy than incompatible ones.

Is an EBM the same as a probabilistic model?

Not necessarily. Some EBMs can be connected to probability distributions, but in the compatibility-based formulation discussed by LeCun, the Energy Function can be the primary object without directly representing a normalized probability distribution.

What does a Latent Variable do in an EBM?

A Latent Variable represents factors that are not directly observed. In a Latent Variable Energy-Based Model (LV-EBM), latent variables can explain hidden structure and lead to related concepts such as Free Energy.

How is JEPA related to EBM?

In Dawid and LeCun's framework, Energy-Based Models and Latent Variable Models contribute building blocks to Hierarchical JEPA (H-JEPA). JEPA predicts in representation space, which creates a conceptual link among EBM, latent representations, self-supervised learning, and prediction.

What is EB-JEPA?

EB-JEPA is the name used by a 2026 paper and open-source library for Energy-Based Joint-Embedding Predictive Architectures. The project includes examples for image representation learning, video prediction, and action-conditioned world-model planning.

What is the difference between Representation Collapse and Energy Collapse?

Representation Collapse means that different inputs are mapped to overly similar or nearly identical representations. Energy Collapse means that the energy landscape loses meaningful discrimination between candidate configurations. They can be related, but they are not identical.

References & Further Reading

  1. LeCun, Y. (2022). "A Path Towards Autonomous Machine Intelligence" . OpenReview. https://openreview.net/pdf?id=BZ5a1r-kVsf
  2. Dawid, A., & LeCun, Y. (2023). "Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence" . arXiv. https://arxiv.org/abs/2306.02572
  3. LeCun, Y., et al. (2006). "A Tutorial on Energy-Based Learning" . In Predicting Structured Data.
  4. Grill, J. B., et al. (2020). "Bootstrap Your Own Latent" . NeurIPS.
  5. Chen, T., et al. (2020). "A Simple Framework for Contrastive Learning of Visual Representations" . ICML.
  6. Zbontar, J., et al. (2021). "Barlow Twins: Self-Supervised Learning via Redundancy Reduction" . ICML.
  7. Silver, D., et al. (2021). "Reward is Enough" . Artificial Intelligence, 299, 103535.
  8. Terver, B., et al. (2026). "A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures" . arXiv:2602.03604. https://arxiv.org/abs/2602.03604