Witex

Former NVIDIA AI Director Is Taking Aim at the Transformer With a 5-Trillion-Token Vision of Physical AI

NEWS

A new startup says it is building AI models that do not merely generate language or images, but directly model the four-dimensional dynamics of the physical world. If its claims hold up, the implications could be enormous.

The next frontier of artificial intelligence may not be language.

It may be the physical universe itself.

Large language models predict sequences of words. Video-generation systems attempt to predict how pixels evolve over time. But a new generation of so-called physical AI models is pursuing something more fundamental: predicting the underlying state of physical systems across both space and time.

At the center of this effort is Accelerated Understanding, a startup founded by Caltech professor and former NVIDIA AI research director Anima Anandkumar and AI infrastructure engineer Benedikt Jenik.

The company is reportedly pursuing an architecture based on neural operators, rather than relying primarily on the Transformer architecture that has dominated modern generative AI.

And the numbers being associated with the project are extraordinary.

According to information circulating around the company, its models have reached 1 trillion parameters, with scaling experiments extending as far as 35 trillion parameters. Its reported training context window has reached 1 trillion tokens, while inference-time context is said to exceed 5 trillion tokens.

If accurate, that would represent an entirely different scale of computation from today’s mainstream language models.

The more important question, however, is not how large the context window is.

It is what the model is using that context to simulate.

Beyond Language, Beyond Pixels

Today’s AI systems largely operate through representations designed for human communication or perception.

An LLM predicts language.

A video model predicts visual patterns.

Accelerated Understanding is attempting to model something different: the evolution of physical fields through space and time.

That means phenomena such as fluid dynamics, turbulence, heat transfer, electromagnetic fields and plasma behavior could potentially be represented directly as computational objects rather than inferred indirectly from language or images.

In other words, the goal is not to make a simulation look physically plausible.

The goal is to learn the underlying mapping that governs how a physical system evolves.

This distinction is critical.

A video-generation model may produce a convincing image of a ball falling under gravity while still failing to maintain physically consistent trajectories over long periods. A model trained directly on physical states, by contrast, could potentially learn relationships that remain consistent across an entire simulation.

The company describes an approach in which a model can generate an entire four-dimensional trajectory — three dimensions of space plus time — in a single inference pass, rather than predicting one frame after another.

It is an ambitious departure from conventional autoregressive generation.

Why Neural Operators Matter

The architectural idea behind the system is the neural operator.

Unlike conventional neural networks, which typically learn mappings between finite-dimensional vectors, neural operators are designed to learn mappings between functions or continuous fields.

That makes them particularly attractive for scientific computing.

A physical system can be described as a field: temperature distributed throughout a material, pressure and velocity across a fluid, or plasma density evolving inside a fusion reactor.

Instead of reducing these phenomena to a sequence of tokens or pixels, neural-operator methods attempt to learn the transformation from one physical state to another.

This has been an active area of research for years, including work associated with Anandkumar and her collaborators.

The potential applications are broad.

Weather forecasting.

Computational fluid dynamics.

Materials science.

Semiconductor thermal modeling.

Fusion research.

Aerodynamics.

Climate simulation.

Drug and molecular discovery.

The underlying proposition is that many of these problems, despite their radically different physical manifestations, can be treated as instances of a more general mathematical problem:

Given the current state of a physical system, what happens next?

If a sufficiently capable model can learn that mapping across many domains, the result could resemble a general-purpose simulator for physical reality.

From Simulation to a “Reality Validator”

The more interesting possibility is not simply faster simulation.

It is the creation of a feedback loop:

simulate → evaluate → improve → simulate again.

That changes the role of AI.

Instead of generating an answer and waiting for a human to judge it, a physical model could potentially generate a hypothesis, test its consequences inside a simulation, measure the result against known constraints, and use that feedback to improve its next prediction.

In this framework, AI becomes less like a chatbot and more like an experimental scientist.

The system could propose a material configuration, simulate its thermal behavior, evaluate the result and iterate.

It could design an aerodynamic shape, simulate airflow around it and refine the geometry.

It could explore thousands of potential physical configurations without requiring each one to be tested in the real world.

That is where the concept of physical AI becomes particularly powerful.

The ultimate objective is not merely to generate information.

It is to optimize the physical world.

The 5-Trillion-Token Question

The reported context length is perhaps the most eye-catching part of the story.

Accelerated Understanding is said to have pushed training contexts to approximately 1 trillion tokens, with inference contexts exceeding 5 trillion tokens.

Those figures are difficult to conceptualize.

A context window of that magnitude could theoretically contain an enormous amount of information about a physical system — potentially representing extensive spatial and temporal states within a single computation.

But there is an important distinction between a large context window and genuine physical understanding.

More tokens do not automatically produce better reasoning.

The real technical challenge is determining whether a model can maintain coherent physical relationships across such a vast representation while remaining computationally efficient.

That is precisely where the neural-operator approach becomes interesting.

Rather than treating every physical observation as an independent token, the architecture can exploit the mathematical structure of continuous physical fields.

If successful, that could make extremely large-scale physical reasoning tractable in ways that conventional token-based architectures struggle to achieve.

Is the Transformer Really on the Way Out?

Probably not — at least not yet.

The Transformer remains the dominant architecture behind modern language models and many generative systems.

But the rise of physical AI raises a more nuanced question:

Does every form of intelligence need to be built around language?

Anandkumar has argued that an AI paradigm centered on language risks being inherently human-centric.

Language is our interface to the world.

It is not necessarily the world’s native representation.

From that perspective, physics may provide a more fundamental substrate for intelligence.

The universe does not communicate in English, Chinese or Python.

It evolves according to mathematical relationships.

A sufficiently general physical AI system would therefore need to represent those relationships directly rather than translating everything into language first.

That is the philosophical bet behind this approach.

From NVIDIA Research to an Independent Startup

Anandkumar’s background makes the story particularly notable.

She joined NVIDIA in 2018 as its AI research director, where she led research exploring how GPU computing could accelerate scientific and machine-learning workloads.

One of the areas associated with her research was AI-powered weather forecasting.

The idea was straightforward but transformative: use machine learning to approximate computationally expensive physical simulations while dramatically reducing the time required to generate predictions.

The broader vision was that AI could eventually learn the operators underlying complex physical systems.

At an NVIDIA GTC event in 2021, Anandkumar presented research related to neural operators and AI-based scientific computing.

NVIDIA CEO Jensen Huang was reportedly enthusiastic about the potential.

Anandkumar has recalled joking that AI might eventually “take the lunch” of theoretical physicists.

Huang’s response, according to her account, was even more ambitious: he wanted AI to take all of their lunch.

The exchange captures the scale of the idea.

For NVIDIA, AI is no longer simply about language models.

It is increasingly about using accelerated computation to model the real world.

The Bezos Connection

The story becomes even more intriguing when viewed alongside Jeff Bezos’s recent interest in AI-driven physical systems.

According to reports surrounding the company’s formation, investor and biotech entrepreneur Vik Bajaj approached Anandkumar and Jenik with an opportunity connected to Project Prometheus, the AI and engineering venture backed by Bezos.

The reported offer was unusually large.

Anandkumar was reportedly invited to serve as a public-facing scientific leader and board member, while the founders were offered as much as 35% of the company.

The package reportedly included a base salary of $1 million per year, potentially rising to $2 million, as well as a financing commitment of more than $2 billion before a Series B round.

The founders nevertheless declined.

Their reasoning, according to the account, was that the opportunity they were pursuing required independence.

Rather than building AI primarily for industrial automation, they wanted to develop a general-purpose foundation for modeling physical systems.

That decision ultimately led to Accelerated Understanding.

Is NVIDIA Behind It?

This is where the story moves from documented history into speculation.

Accelerated Understanding has reportedly indicated that it has partnerships with computing providers that supply the hardware infrastructure required to train and operate its models.

Given the computational requirements implied by trillion-parameter models and multi-trillion-token contexts, NVIDIA would be an obvious candidate.

Anandkumar’s previous relationship with NVIDIA makes the possibility even more interesting.

But there is an important caveat:

NVIDIA has not publicly confirmed that it is an investor or undisclosed backer of Accelerated Understanding.

Any suggestion that Huang or NVIDIA is secretly financing the startup should therefore be treated as speculation unless and until the companies disclose such a relationship.

What is beyond dispute is that NVIDIA has spent years investing in the intersection of GPUs, AI and scientific computing — precisely the technological territory in which physical AI is now emerging.

The Bigger Shift: From Generating Information to Modeling Reality

The most important development here may not be the five-trillion-token number.

It may be the change in what we expect AI to understand.

The first generation of generative AI learned to generate information.

LLMs generate text.

Image models generate pictures.

Video models generate moving images.

The next generation could attempt to generate states of the physical world.

That would represent a profound shift.

Instead of asking an AI system:

“What should happen?”

we could ask:

“Given these physical conditions, what will happen?”

And eventually:

“What should we change to make the outcome we want happen?”

That is the transition from generation to optimization.

From content creation to scientific discovery.

From digital information to physical reality.

The Road Ahead

It would be premature to declare that the Transformer era is over, or that Accelerated Understanding has already built a universal simulator of the universe.

The claims surrounding the startup still need to be independently validated, particularly around model scale, context length, generalization across physical domains and real-world accuracy.

Those are extraordinarily difficult engineering problems.

But the direction is unmistakable.

For years, the AI industry has focused on making machines better at understanding what humans say and see.

Physical AI asks a more fundamental question:

Can machines learn how the world itself behaves?

If the answer eventually becomes yes, the implications could extend far beyond chatbots and generative media.

AI could become a tool for discovering new materials, designing machines, forecasting weather, optimizing energy systems, accelerating fusion research and exploring physical configurations that would be impossible to test individually in the real world.

The ultimate ambition is even larger.

Not an AI that merely describes the universe.

Not an AI that creates a convincing picture of it.

But an AI that can model the dynamics of reality itself.

And if that vision succeeds, the next great AI breakthrough may not be another chatbot.

It may be a machine that learns the rules of the physical world — and then uses those rules to help us change it.