The training took place in NVIDIA Isaac Sim. During the self-play phase, the robot competed against recent versions of its own policy. As its performance improved, its opponents improved with it. The objective was to score goals rather than separately reward behaviors such as dribbling, shielding the ball or tackling. Skild says those behaviors emerged because they improved the policy’s ability to score.
Before self-play, S1 was trained on individual soccer skills using human references and skill-specific rewards. The self-play stage then combined and adapted those capabilities through competition. The model outputs target angles for the humanoid’s joints at 50 Hz, according to additional technical reporting on the release.
Skild reports that the policy initially struggled to walk in simulation but eventually learned to recover after falling, maneuver around defenders, protect possession and tackle opponents. After more than 140 simulated years, the company transferred the policy to a physical humanoid and demonstrated it playing against humans and other robots. The 140-year figure represents accumulated simulated experience, not elapsed training time. Skild has not disclosed the compute budget, number of parallel simulations or quantitative real-world success rates.
The work is a separate post-training development from S1’s previously announced ability to learn tasks from a single video demonstration. Skild says it is now extending self-play to other robotics domains and plans to publish further work involving larger teams, collaborative manipulation and navigation.



