Physical AI
Simulation
Which Parts of the Sim-to-Real Gap Can Simulation Actually Close?
The sim-to-real gap is four separate gaps. Which ones close with measurement, which ones no simulation budget closes, and how to tell them apart.
Physical AI
Simulation
The sim-to-real gap is four separate gaps. Which ones close with measurement, which ones no simulation budget closes, and how to tell them apart.
Written by Sneha Tanna, Content Strategist, FutureBeeAI · Reviewed by Jesal Thakkar, Founder, FutureBeeAI
Last updated: September 28, 2026
Author block
Sneha Tanna | Content Strategist, FutureBeeAI.
Sneha Tanna is Content Strategist at FutureBeeAI, where she has run the content operation since June 2025 under the guidance of the founders. She came to AI data from hands-on annotation, transcription and localization QC work, and she writes on training data quality, speech and multimodal datasets, low-resource language collection and FutureBeeAI's Physical AI research pillar.
Reviewer: Jesal Thakkar | Founder, FutureBeeAI.
Jesal Thakkar co-founded FutureBeeAI in 2020 and leads it as CEO. The company builds production training datasets for speech AI, LLMs, computer vision and multimodal systems, collected in 100+ languages across 70+ countries through a contributor community of 20,000+. He reviews FutureBeeAI's content for technical accuracy and category positioning.
Move a robot's camera a few centimeters, raise the table it works on, and a policy that looks finished can stop working, even though nothing about the model or the task has changed. The sim-to-real gap is the performance a policy loses when it moves from the simulator it was trained into to the physical world it was built for, and a 2025 NVIDIA benchmarking study found that changes in lighting and camera pose alone degrade model success by 30 to 50 percent. A 2026 empirical study of vision-language-action models goes further and concludes that spatial features speak louder than appearance: augmenting spatial factors such as table height and camera pose produced larger improvements than the appearance changes most pipelines focus on.
That finding should worry anyone currently funding photorealism. The factors that dominate transfer are spatial, while the factors most randomization pipelines spend their effort on are visual, and the budget tends to follow the effort rather than the evidence. Before any of that can be corrected, the gap has to stop being treated as one thing.

A 2026 peer-reviewed review in the Annual Review of Control, Robotics, and Autonomous Systems breaks the reality gap down by the structure of the control problem instead of treating it as one blurry discrepancy, and that breakdown is the most useful planning tool in the literature. The first is a dynamics gap, which covers what the simulator gets wrong about how the world evolves, including rigid-body assumptions and contact models simplified to point contacts and linearized friction cones. The second is a perception gap, where the review notes that real sensor noise is non-Gaussian, state-dependent, temporally correlated and influenced by motion, temperature, lighting and surface properties, rather than the simple additive noise most pipelines model.
The third is an actuation gap, where real actuators behave as higher-order systems that introduce phase lag and controller latency. The fourth is a system-design gap, which covers communication delay, packet loss, safety limits that exist on the real robot but not in simulation, and reward or termination conditions that depend on information the real system cannot observe.
Four gaps means four interventions and four prices, which is why "closing the gap" is not a deliverable you can put on a quote. A vendor improving your rendering is working on one part of one gap, and on current evidence not the expensive part. The same breakdown explains why the same system fails in a different building even when nothing about the model has changed. Knowing there are four, though, does not tell you which one is draining your success rate.

The evidence on visual policies points in one direction: the spatial relationship between the sensor and the workspace carries more transfer risk than most of what the renderer touches. The NVIDIA study places camera pose alongside lighting as the largest source of degradation, and the vision-language-action study found that training against spatial variation was worth more than training against appearance variation, which is the opposite of where most simulation budgets go.
Compare that with how simulation budgets are actually spent and the standard priority list turns upside down. A program that funds photorealism while mounting its camera on a tripod that gets bumped between sessions has paid for the cheaper problem and left the expensive one open.
The cheapest intervention in this category is one nobody sells: fixing and documenting camera extrinsic as part of the collection brief, and then randomizing around measured tolerances instead of guessed ones. That is really an argument that randomization ranges are claims about the deployment environment, and those claims inherit every representativeness problem that comes with one. On FutureBeeAI programs where pose tolerance is specified numerically at capture time, the conversation after deployment is about the policy. When it was left to whoever set up the rig, the first question was always whether the data and the simulator were ever looking at the same scene.

This is the distinction that decides where the budget goes. A parametric gap is a quantity the simulator represents but gets wrong, such as a friction coefficient, a mass, a link length or a latency. A structural gap is a phenomenon the simulator does not represent at all, such as cloth self-penetration or a truly chaotic contact event. The first kind closes with measurement, and the second kind does not close at any price most programs are willing to pay.
A 2026 study on how to spend a transfer budget tested the allocation directly by splitting a fixed measurement budget between identifying system parameters and widening domain randomization. Its finding runs against standard practice: a small number of identification rollouts closed most of the transfer gap, and once any real measurement existed, policies performed best when trained at the estimated parameters rather than across a widened band. Broad randomization that contained the true system still did not substitute for measuring it, which puts system identification before randomization rather than after.
That study ran on a two-parameter pendulum, so treat the ordering as directional rather than as a calibrated recipe, and note that the authors flag the boundary themselves: under structural model mismatch, the simulator cannot represent the target system through parameter fitting alone. That gives a simple allocation rule. Measure wherever the gap is parametric, and collect real data wherever it is structural, because no simulation budget converges on a phenomenon the simulator never represented. If you are about to renew a simulation contract, run that classification first.
Deformable objects are the clearest structural case, and the evidence is specific. A 2025 garment manipulation benchmark reports that prevailing simulators rely on simplified position-based dynamics that produce distortion, stretching and frequent self-penetration failures. It also reports that one widely used engine fails to generate valid grasping visuals at higher mesh resolutions, while another becomes unstable on gripper contact and produces exaggerated twisting or inflation.
Tactile and force-rich contact behaves the same way. A dedicated 2026 method for closing the tactile gap still leaves 14.7 to 18.5 percent error in deformation depth after its fix is applied, and the stated reason for stopping there is that finite-element contact modeling is computationally prohibitive at reinforcement-learning scale. The review adds two categories that are structural by nature: chaotic phenomena, which it describes as inherently non-reproducible, and human behavior, which it characterizes as complex, context-dependent and often irrational.
Now the counterexample, because "contact-rich manipulation never transfers" is a convenient belief for a company that sells real data, and it is false. IndustReal, a simulation-trained assembly system covering peg insertion, gear meshing and connector insertion, reported 83 to 99 percent success over 600 real trials. What it took was signed-distance-field rewards, simulation-aware policy updates and a sampling-based curriculum engineered for that task family.
So structural gaps are not closed once and for all; they are closed one task family at a time by whoever pays for the research. If your task class already has a published transfer result, simulation is a live option and you are building on somebody else's research. If it does not, you are funding that research yourself, which is a different line item and a different timeline.
The classification is a working session, not a research project, and for most tasks it takes about an afternoon.
1. Name the failure against the four gaps: Describe where the policy breaks on the real system and assign it to dynamics, perception, actuation or system design, because a failure that spans two gaps needs two answers.
2. Check whether the simulator has a parameter for it: If the phenomenon exists in the simulator with the wrong value, such as friction, mass, latency or camera pose, the gap is parametric. If the simulator has no representation of it at all, such as cloth self-contact or soft tactile deformation, it is structural.
3. Run a short identification pass on parametric gaps: A handful of real rollouts measuring the suspect parameters, including camera extrinsic, usually tells you more than another round of widened randomization.
4. Search for a published transfer result on structural gaps: If one exists for your task class, price the engineering to reproduce it. If none exists, price real-world collection, or price the research explicitly and give it its own timeline.
5. Write the result into the collection brief: Measured tolerances, camera pose specifications and the list of structural phenomena become requirements that both the simulation team and the capture crew work against.
Simulation's clearest advantage is generating diversity, and no collection budget comes close. Legged locomotion policies have been trained to walk in under four minutes on flat terrain and in around twenty minutes on uneven terrain using thousands of parallel simulated robots on a single workstation GPU, then transferred to real hardware. The 2025 MuJoCo Playground report documents a quadruped's flat-ground locomotion stage trained within five minutes on two consumer GPUs, with its joystick, handstand, foot stand and fall-recovery policies all transferring to the real robot without additional fine-tuning.
Those numbers matter because generalization in imitation learning depends heavily on the diversity of environments and objects a policy has seen. In the RoboCasa household benchmark, real-robot success rose from 13.6 percent with real data alone to 24.4 percent when the same real data was trained alongside simulated data, and in simulation its scaling study moved from 28.8 percent on fifty human demonstrations to 47.6 percent on a fully generated dataset of three thousand.
The case for real collection was never that simulation fails. It is that simulation covers an enormous space, and real data is what tells you where inside that space your actual deployment sits.
Generative video arrives with stronger claims than either simulation or real data and deserves the same scrutiny. A 2025 study built the Physics-IQ benchmark to test physical understanding in video generative models and found it severely limited across leading systems and unrelated to visual realism. Against a real-video ceiling scored at 100, the best model reached 29.5, the model rated most visually realistic in the set scored 10.0, and the correlation between realism and physical understanding was not statistically significant.
The payoff figures point the same way. When NVIDIA's GR00T N1 humanoid foundation model expanded 88 hours of real teleoperation into 827 hours of generated video at roughly 105,000 GPU-hours, the ablation credited that data with an 8.8 percent average improvement in simulation at the 100-demonstration setting and 5.8 percent across eight real-world tasks in the low-data setting. That gain is useful and real, and still a long way from replacing collection.
A model that cannot represent physics inherits every structural limit described above, one step further removed from reality. Treat generated video as a coverage amplifier sitting beneath simulation, never as a substitute for either input, and never as an evaluator of the policies it helped train.
Red flags:
The whole gap is sold as one deliverable: The proposal promises to "close the sim-to-real gap" without naming which of the four gaps it addresses.
Randomization ranges with no measurement pass: Ranges are chosen from judgment instead of surveyed tolerances on the real system.
Photorealism presented as the main lever: Appearance is treated as the priority even though the evidence ranks spatial factors such as camera pose above it.
No concession anywhere: The proposal never names a task class where simulation is the wrong tool.
Green flags:
Gap classification done before the quote: Each task is classified as parametric or structural before any price is attached.
Identification priced as its own line item: Measurement is visible in the quote instead of folded invisibly into simulation setup.
Randomization tied to surveyed tolerances: Ranges are measured on the real system, not assumed.
Camera pose specified numerically: The spec includes a stated tolerance that holds between sessions.
FutureBeeAI runs that classification before scoping a collection, because it decides whether the next budget belongs to a simulation engineer or a capture crew, and getting it backwards is expensive in both directions.
Most teams arrive at the sim-to-real gap as a performance problem and start spending on whichever side of it they already understand. Simulation teams widen randomization and collection teams add episodes, and both are reasonable responses to a question nobody has actually asked yet, which is what kind of gap this is. An afternoon per task settles it, and no amount of spending on either side recovers a wrong answer.
If you are deciding where the next phase of the budget goes, talk to the FutureBeeAI team about running the gap classification and the collection specification together, before the first episode is recorded. If you are earlier than that, the companion piece on building a sim-to-real data strategy covers what to do with the mixture once the classification exists, the library of 2000+ ethically sourced datasets shows what production-grade collection looks like, and our case studies show how collection specifications are written in practice.
A. The sim-to-real gap is the drop in performance a robot policy suffers when it moves from the simulator it was trained into the real world. It is made up of four separate gaps (dynamics, perception, actuation and system design), and each one needs a different fix.
A. Classify the gap first. If the simulator represents the phenomenon but has the number wrong, such as friction, mass, link length or latency, fix the simulator, because identification is cheap and one study found that widening randomization around the true value still did not substitute for measuring it. If the simulator does not represent the phenomenon at all, such as cloth self-penetration or chaotic contact, collect real data.
A. The published record suggests so, though no source states it as a consensus. Locomotion papers report whole policy suites transferring zero-shot after minutes of training on one or two GPUs. Contact-rich manipulation results exist and are strong, including 83 to 99 percent success on assembly tasks over 600 real trials, but they required task-specific reward shaping, simulation-aware policy updates and a curriculum built for that task family. Read the difference as a difference in the cost of transfer, not in whether transfer is possible.
A. A parametric gap is a quantity your simulator represents but gets wrong, and it closes with measurement. A structural gap is a phenomenon your simulator does not represent at all, and it closes only when somebody funds the research to represent it, one task family at a time. The classification takes about an afternoon per task and decides whether your next spend belongs in measurement or in collection.
Acquiring high-quality AI datasets has never been easier!!!
Get in touch with our AI data expert now!
