SimEX: Simulation-Integrated Robotics AutoResearch Coding agents that learn to control real robots by experimenting in simulation first

Coding agents are good at reasoning about goals in the digital world, but bringing them onto robots has been hard. Writing a policy in one shot fails because the model does not understand the robot or the physics well enough, and iterating directly on hardware is slow, costly and unsafe. SimEX uses simulation as the laboratory. In Stage 1 a coding agent runs open-ended probe-and-optimize cycles in a simulated sandbox, inventing its own tasks and building a reusable robot toolbox. In Stage 2 it adapts that toolbox to the real robot with only five physical trials, replaying each trial in simulation to diagnose the failure and screening several candidate repairs before spending the next trial. No demonstrations, no hand-written skills, and about ten minutes of robot time per task.

SimEX: discover robot skills in simulation, build a reusable toolbox of code and usage knowledge, and adapt in the real world through real trials and simulation tests.

Real-world qualitative results

Task completion on the physical robot

Real-world rollouts of the adapted toolbox on the physical dual-arm YAM robot. All clips are shown at 8× speed.

Barcode scanning

real robot · 8×

Barcode scanning · Rollout 1

real robot · 8×

Barcode scanning · Rollout 2

Plate to tote

real robot · 8×

Plate to tote · Rollout 1

real robot · 8×

Plate to tote · Rollout 2

Towel folding

real robot · 8×

Two-fold towel task · Rollout 1

real robot · 8×

Two-fold towel task · Rollout 2


Real-world qualitative results

Emergent behaviors

None of the behaviors below were specified. They emerged from the agent's own probe-and-optimize cycles in simulation and its repairs after real trials.

real robot · 8×

Zero-shot generalization. The scanning skill generalizes to unseen objects in a zero-shot setting.

real robot · 8×

Bottle at the tote edge. The robot figured out how to pick up objects at the edge of the tote by aligning its gripper with the tote’s orientation and gently nudging the tote.

real robot · 8×

Flipping a box. When the label is on the other side, the agent learned to flip the object in the tote and scan it again.

real robot · 8×

Recovery during scan and transfer. The agent is robust to failures and can recover from them.

real robot · 8×
real robot · 8×
The left and right videos show the agent grasping the plate from different sides based on its distance from the tote, avoiding collisions with the tote or wall.

Real-world qualitative results

Instruction following and multi-task capabilities

The same toolbox, a new instruction. Here we change the task at deployment time and let the coding agent write a fresh program against the toolbox it built.

real robot · 8×

“Scan two objects and place them in the destination tote.”

real robot · 8×

“Scan one object and place it in the destination tote.”

real robot · 8×

“Leave one plate on the table.”

real robot · 8×

“Place the red plate in the tote.”

real robot · 8×

“Fold the towel in half.”


Simulation

Sim-to-sim behaviors

The controlled sim-to-sim experiments in the paper use an independently implemented simulator as the stand-in for the robot: Isaac Sim for barcode scanning and plate to tote, MJWarp Flex for towel folding. Below, two successful rollouts of the adapted toolbox on the held-out tasks of each family.