Computer scientist Louis Castricato was dedicating his eighth year to studying large language models, the artificial intelligence technology behind chatbots like ChatGPT and Claude, when he began to feel like he was at a dead end.
“We’ve basically passed the point of doing real LLM fundamental research,” Castricato said. “Now, it’s just application.”
The researcher finished his doctoral studies at Brown University and founded a new company, Overworld. Its ambition is clear in the name: to create an artificial intelligence capable of understanding and navigating a physical world, not just words.
While there is still a lot of money to be made from AI chatbots, investors are betting on it, committing trillions of dollars to leading developers like Anthropic and OpenAI, a growing number of artificial intelligence entrepreneurs are dedicating themselves to what they see as the next frontier: “models of the world.” These systems aim to teach artificial intelligence, and sometimes robots, how to react in a physical environment.
Among those involved are some of the most prominent scientists in the field, such as the “godmother of AI”, Fei-Fei Li, who describes the concept of world model as “one of the most important and most overloaded terms in AI today”.
What are world models and why does AI need them?
At the heart of world model research is the premise that artificial intelligence cannot be truly intelligent if it just “reads a book.” She also needs to “read the environment.”
More on this story: World models in AI: the future of physical simulation and its challenges
“While language models learn the statistical structure of text, world models learn the statistical structure of space and time: how light falls on a surface, what a garden looks like from an angle no camera has captured, how objects respond to force and follow the laws of physics,” wrote Li, founder of San Francisco startup World Labs, in an essay published this month.
Another enthusiast is AI pioneer Yann LeCun, who left his role as chief AI scientist at Meta last year to found Paris-based Advanced Machine Intelligence Labs.
“Model of the world is quickly becoming a buzzword,” LeCun said on a recent “Unsupervised Learning” podcast. He sees the concept as something that allows an AI agent to “predict the consequences of its own actions.”
There are several ways to define world models, often based on the technologies that are expected to be built with them, whether robots or a more interactive video game.
Robots and the limitation of language models in real environments
Training on all of humanity’s books, news articles, and visual media, as done by AI language models, has resulted in intelligent assistants that transform the nature of office work and some creative fields. However, some advocates see limitations in generative AI models that work by repeatedly predicting the next word or pixel to produce new dialogue, images or lines of code.
Chatbots can’t pick up a mug of coffee, notes Martial Hebert, dean of computer science at Carnegie Mellon University.
“There’s all the geometry in the world, the dynamics of how I move my hand, the physical interaction of contact with the cup,” Hebert said. “This is much more complex than just predicting the next word in a sentence.”
For scientists like Hebert, who has devoted more than four decades to robotics research, the most useful application for world models is as a faster, more affordable path to “physical AI,” another term on the rise in the technology industry.
“Some may have different definitions, but physical and embodied AI are kind of an evolution of what we used to call robotics,” Hebert explained in an interview. He added that some of the AI advances that have made chatbots so useful can also be applied to building artificial intelligence with environmental awareness broad enough to function like a robot’s brain.
Learn more: Researchers use AI to discover hidden laws in charged particle systems
“In your body and spinal cord, you have a very general model of how to balance, how to walk, and you can adapt to a bad knee in the morning, so you walk a little bit differently,” he said. “You don’t need to think about it. You have a general blueprint somewhere in your nervous system and brain that allows your body to adapt very quickly.”
Simulated worlds arouse the interest of investors in the sector
Smarter robots are not the only goal of world models. Castricato founded Overworld last year, and the small Rhode Island startup is now building video game worlds where a scene, say a spooky forest, can adapt as a virtual character moves and interacts with the objects in it.
“There’s no other world model where you can just walk through doors or interact with a detailed environment like this,” he said in an interview. “We optimize for interaction above anything else.”
While near-term applications aren’t as immediately apparent as AI coding tools, world model builders are attracting interest from venture capitalists like Steve Jang, co-founder and managing partner of Kindred Ventures.
The company is investing in Overworld and other companies focused on world models, including Causal Labs, which develops AI models for weather forecasting, and Extropic, which builds specialized computer chips for such models.
“I believe the future will be many different types of models with many different philosophies and architectures,” said Jang. “I don’t think it’s going to be one big, dense model to govern them all.”
In his recent essay, Li sought to create a “taxonomy of world models” to help organize the confusion over competing views.
“A video model that produces magnificent but physically impossible flames, a language model that improvises a playable game, and a physics engine that faithfully simulates combustion all bear the same name,” she wrote.
On the same topic: Nvidia advances in Physical AI and ensures that robots can now develop fast reasoning with current data, says executive
She divided the world models into three categories. The most commercially viable today are “renderers,” which prioritize the visual fidelity of the virtual worlds they create, but cannot be completely trusted to teach robots much.
Next come “simulators”, which create virtual training grounds that faithfully represent the physical structure of a world; and “planners,” who try to predict what an AI agent or robot should do in an unstructured world.
“A robot that can plan is a robot that can work, and the entire industry is racing to be the first to get there,” she concluded.

