World Models and the Robot Race: The Case for a Measurement Strategy

As China rapidly advances in using world models to bypass data bottlenecks and train physical AI, the US must establish rigorous federal benchmarking and secure data commons to maintain its leadership and defend against dual-use risks.

First imagine a highly intelligent, AI-powered chatbot, like ChatGPT. Such agents have significantly shaped the daily lives of people around the world since they became public in 2022. 

Now imagine an AI chatbot that also has arms and legs and can interact with humans in the physical world. This is physical AI – and this is the next step in the evolution of AI.

To support physical AI, engineers from companies like NVidia are developing a technology with similarly revolutionary potential: the world model.

Instead of painstakingly duplicating human work, which relies on a human action database, robots are now able to manufacture experience in a simulated world that allows them to practice in advance of having to do the real-world job. The basic idea is to create an environment that replicates the way things change when something acts on them, ultimately giving robots much less task-specific data to work with. 

As it happens, researchers in China are making major advances in developing world models.

Beijing recognizes the importance of world models, as made clear by the establishment of a World Model Expert Committee in August during the World Robot Conference as part of China's Five-Year Plan. Washington has not yet recognized that the robot race now runs through world models. The United States retains a strong robotics lead, but now it must set comprehensive policy for measurement and validation, as well as dual-use governance, to protect its edge over China and realize the full promise of physical AI.

Why Data Is the Bottleneck

It is well known in the new world of physical AI that the need for data can be a breaking point for most systems. While language models can feed off the existing web, no comparable archive of machines doing physical work exists, so every example for physical AI must be produced on purpose. Beijing’s own telecommunications research institute has said that the high cost and low efficiency of collecting physical-world data is the core bottleneck for the field. 

A particularly ambitious Chinese firm developing physical intelligence, AgiBot, tried to solve the data problem by building a facility more than 4,000 square meters in area and filling it with over 3,000 objects for the robots to interact with. It then ran the robots through one-million demonstrations. 

What does any of this have to do with world models? Well, everything. 

AgiBot recently published a paper on Genie Envisioner, a world model trained on roughly 3,000 hours of its own robot video, which then adapted that data to a robot it had never seen before using only one hour of teleoperated demonstrations. It turns out their world models could be trained to navigate a facility and deal with a simulated equivalent of ordinary occurrences in the world, and that can allow standard computer vision methods to more quickly collect from the real world situations that often cause confusion for models. The model then works out of the box, producing a significant upgrade to the output of the teleoperation, reducing the need for brute force data collection of real normal operations. 

Collectively, the community is now working on building generalized models of the physical world that will make this approach easier. The best known so far, Nvidia’s Cosmos 3, trained on a vast library of 348 million videos of real world dynamics, is openly published. As these models improve, the biggest means to increase the efficiency of our physical intelligence systems will come from being able to generate new scenes to test robots and other machines against, add in new unfamiliar objects, change conditions like lighting or humidity, and introduce the sort of rare accidents that we normally need to wait a long time to catch. Experience may no longer be a function of floor space, hardware, and human operators, but computing power alone.

The Dual-Use Problem

World models are dual-use by nature. China’s military press recently noted that simulations will play a key role in the PLA’s “modernization” and “intelligentization” agenda, and the sorts of world models developed at AgiBot or in any similar Chinese project could help compress the information needed to train the main functions of a battlefield into quicker learning experiences for perception, situation assessment, and planning. The real complexity of training combat agents is getting good information into these cheaper training models, given the expense of simulating the necessary inputs and outputs. 

At the same time, simulators that train machines can also become targets. In February, researchers showed that quietly corrupting the physical assumptions inside a driving world model degraded the systems trained on its output, worsening their planning by roughly 20 percent, even though the generated video still looked convincing. Separately, poisoning 0.31 percent of training episodes was enough to plant a hidden behavior that worked almost every time. If any hallucination happens on robots, it could cause them to derail or collide with humans and properties. Not understanding the full scope of what is at stake in this dual-use field and pushing some systems into use is going to lead to extortionate costs.

Not the Only Bet

World models, however, are not the only option and golden standard for fostering robotics real-world applications. Teleoperation has not disappeared; it has changed jobs, from bulk data collection to correction. Human-carried camera rigs are closing on teleoperation for some precision tasks, though real robot data still gets mixed back in for deployment-grade reliability. And in strategic industrial work, firms skip generic simulation altogether. Path Robotics, which builds industrial welding robots in Ohio for the shipbuilding sector, trained its model on “tens of millions of welded inches” drawn from eight years of real welds, because the physics of an arc and a molten pool, and the variation between supposedly identical parts, are not things the current world models have paid attention to. Path’s approach was rewarded this past August with a $900 million contract to automate welding and finishing on Navy carriers and submarines. 

China is also setting up nearly 30 state-backed “training grounds” in an effort to give robots the sort of real operating experience they expect to be difficult to replicate. Given the push to get robots into state-owned factories, some of these are simply going to generate records of real machines failing at their work to refine the software models greased into operation.

What Washington Should Do

Congress should give the National Institute of Standards and Technology (NIST) a key role in creating benchmarks for independent validation of how well simulation can be trusted to create world models, starting with measures of physical fidelity rather than visual realism. The results of these benchmarks should be published, together with standards of minimum evidence for claims about simulated performance in the real world. Those claims are currently thin: robot policies scoring about 95 percent on a standard test can fall below 30 percent once objects and viewpoints shift. Any federal procurement of autonomous systems should require independent tests against these benchmarks and disclosure of the environments in which the systems were validated. The US should work with its allies, especially the EU and East Asia to promote standards for independent validation against these benchmarks.

In areas where high-quality simulation is unavailable, the federal government should support the creation of data commons and physical AI data centers with different application scenarios to empower data collection, testing, and benchmarking. Simulators are only as reliable as the real-world data against which they are validated. The US federal government should itself encourage the industry and allies to share a large body of teleoperation data—including touch, force, and human behavior data—collected in facilities and relevant contexts, including laboratories, ports, and depots, to stay ahead of the PRC in the robotics race. 

The Pentagon must continually assess how adversaries may exploit simulation capabilities, from creating and testing dangerous taunts of the simulation to using teleoperation interfaces. Therefore, its purchases of autonomous systems should come with contracts guaranteeing data-provenance and anti-poisoning capabilities, as well as controls on who can operate and update software on deployed systems. 

The United States cannot afford to wait. This is the moment to act—and the United States needs a comprehensive robotics strategy to lead in one of the most defining technologies of our time.