Robotics inherited a belief from language modeling: scale wins. Add more data, add more parameters, and capability follows. It's a good prior, it built the models that made this series worth writing, and for a while it drove VLA progress too. But the most careful recent work in the 1,228-paper corpus we analyzed has started to find its edge. On the margin that matters, more robot data is no longer the reliable lever the field assumed, and the reason ties every earlier part of this series together.
Here is a statistic that says more than its size suggests: when we split the 1,228 VLA papers into ten chunks for parallel analysis, one chunk of 123 papers contained zero tactile-centric work. Not a little. None. Across the full corpus, audio is thinner still: roughly three efforts in 1,228 papers touch it meaningfully.
The single most exciting number in the VLA corpus is also the most misleading. HumanNet assembles roughly one million hours of egocentric human video (arXiv:2605.06747). Being-H0.5 works from 35,000 hours more (arXiv:2601.12993). Set against the thousands of hours of teleoperated robot data that represent the field's hard-won total, these corpora look like the end of the data bottleneck. Humans manipulate objects all day, on camera, for free. Why teleoperate a robot when YouTube has already recorded the demonstrations?
If you've read the first four parts of this series, you might expect this one to argue against synthetic data. It doesn't. The synthetic-data boom in robotics is real, it's accelerating, and on the evidence of the 1,228 VLA papers we analyzed, it is one of the most promising answers to the data bottleneck the field faces. Consequently, fighting it would be at worst, a mistake and at best, a bit of a waste of time.
You fine-tune a Vision-Language-Action model, run it on LIBERO, and post a 95% success rate. The number goes in the paper, the demo video, the investor update. Then you deploy the same model on a real robot in a slightly different room, and it faceplants.
Every robot data collection session produces two datasets, and you keep only one.
Take a state-of-the-art Vision-Language-Action model, one posting 90%+ success rates on a standard benchmark, and corrupt its language input. Swap "put the mug in the sink" for a paraphrase. Or for an instruction about a different object. Or for meaningless tokens.
A synthesis of the data problem across three years of Vision-Language-Action research, and what it says about where the robotics field goes next.
CEO & Co-Founder
Subscribe for news.