Instead of retraining the model for each new task, S1 receives a video demonstration as part of its context and uses it to determine what actions to take. Skild says the same model can then execute tasks lasting up to 10 minutes and involving dozens of manipulation steps.
The company demonstrated S1 on four tasks that it says were absent from its training data: making pour-over coffee, potting a plant, assembling a kit and cooking pancakes. In each case, the robot was given one visual demonstration before attempting the task itself.
Skild says S1 can also adapt when the physical environment differs from the demonstration rather than simply replaying the observed movements. In one test, the demonstration used a watering can while the robot was given a cup instead. The model completed the task using the available object. It has also shown the ability to retry actions after making mistakes.
In Skild’s internal testing, S1 reached 66% per-step success on unseen tasks after 100,000 hours of pre-training, compared with 9% for a language-prompted VLA trained with the same data and compute. The company says the advantage of in-context learning increased as more pre-training data was added.
S1 is already being deployed with a limited number of Skild’s industrial partners, with a broader rollout to customers planned over the coming months.


