Figure evaluated the model in 30 homes across the San Francisco Bay Area using three tasks: tidying living rooms, folding towels and making beds. No training data was collected in the evaluation homes, and none of the toys, towels or bedding used in the tests appeared in the task-specific training data.
Each task used the same model checkpoint across all 30 homes. The robot worked with the existing beds, couches and folding surfaces and received no fine-tuning or adaptation after entering a property. Figure says the behaviors combine autonomous locomotion, manipulation, bimanual coordination, perception and whole-body control.
Helix 2.5 was pretrained on Index, Figure’s dataset of human behavior. In a controlled comparison, Figure trained two policies using identical task-specific data and architecture. The model trained from scratch completed 9% of zero-shot trials, while the Index-pretrained model completed 56%. Success required completing the entire task, with no partial credit. These results are company-reported and have not been independently validated.
Figure also reports that Helix 2.5 matched the success rate of a representative Helix 02 behavior using half as much task-specific robot data, while operating across 30 unseen environments rather than the environment in which its training data was collected.
The company tested four versions of the model across an eightfold increase in Index pretraining data and reported that robot-action prediction improved consistently as the dataset increased. Figure says its largest training run’s loss could be forecast from the smaller runs with an error equivalent to 0.54% of the variation across the full 8× data range.
Index is currently collecting approximately 35 minutes of human experience every second. Figure has separately committed $3.5 billion to compute infrastructure for training future Helix models.



