The five axes of egocentric dataset diversity, what the published datasets actually cover, and how to specify coverage before collection starts.
Where teleoperation episodes are lost across collection, annotation, and curation, and how to budget a campaign by yield rather than by session count.
What the measurements show about automated judges on creative output, where a judge earns its place, and where it cannot settle the question at all.
What aesthetic data is, why the published category stays thin where it matters, and the four decisions that determine whether a dataset is usable.
A test that tells you whether your task needs expert judgement or simply a better written specification, before you commit an annotation budget.
Why preference tuning narrows image model output, what the published measurements show, and how to collect preference data that keeps range.
How to separate a real design quality regression from a drifting standard, using a frozen reference set, per-criterion scores, and agreement.
The calibration, timing, and annotation failures that produce a valid but unusable teleoperation dataset, and what to verify before teardown.
The measured ceiling on public human text, why generated data only partly answers it, and the three mechanisms that produce net-new human data.
How to write and audit an annotation specification for a domain nobody on your team understands, using seeded gold items and agreement patterns.
On some tasks the judgment is the label, and no guideline document transfers it. Here is how to tell which tasks those are and how to run expert annotation well.
Crowdsourced annotation fails in ways throughput dashboards are not built to detect. Five failure modes, how to test for each, and where the model stops being appropriate.
The final few percent of cases resists the methods that got you the first 95%, because rarity is a property of the distribution you are sampling from.
Internet video is abundant and free, and it records what happened rather than what was commanded. That missing action channel is the constraint that shapes world model training.
Robot foundation models are not short on trajectories. They are short on diversity, grounded language, failure coverage, and modalities, and more of the wrong data makes them…
Visual realism and physical understanding are measurably different capabilities. Here is what a world model needs in its training data to learn the second one.
Contact is where manipulation succeeds or fails, and it is the signal robot datasets are least likely to contain. Here is what makes it hard to capture and what it costs to fix.
Embodied AI runs on data that has to be produced under a protocol rather than collected from the web, which turns model quality into an operations problem.
Two manipulation datasets of the same size can differ completely in what they teach a policy. Five design decisions, made before collection, account for most of the difference.
Simulation solves cost and volume for robot training data, but four classes of signal stay out of reach at any fidelity, and each one has to be captured in the real world.