Simulation scales. Fleets scale. The engineer who has to watch the footage does not. Every rollout, in the lab and in the field, waits on a person to say what happened before it means anything. We take that person out of the loop and leave them the five runs that actually needed them.
Different machines, different signals, different phase vocabularies. Same record, same three questions asked of it. Note the humanoid: it completed the task and still should not have shipped.
Evaluation is not the first step. Before anything can be judged, what happened has to be represented: entities, contacts, behavioural phases, events and outcomes, each carrying where it came from and how confidently it is known. Build that representation once and evaluation becomes a layer on top of it, changeable at will. The representation is the product. Everything else is a question asked of it.
Phases are cut where the physics changes: a force step, a velocity inflection, a current spike, found by deterministic extractors. The model supplies the words. No timestamp or magnitude in your output originates in a language model.
What happened is stored as structure, not as a caption. Change what counts as success and nothing re-reads your footage. Judge the same hundred thousand executions against completion, safety, quality or your own criteria.
A human evaluator samples. A machine-readable record lets you reason across every execution you generate. That is where the rare failure modes live, and where the regressions that aggregate metrics hide actually turn up.
We are not another simulator and not another storage format. Point us at executions wherever they come from. We normalise them into one representation, and everything downstream works the same way.
Atomic Skills Industries is the team formerly known as DiffuseDrive, whose synthetic-data platform for autonomous systems is used by teams at the companies above. Perception and data infrastructure for physical AI, out of automotive and autonomous driving, now pointed at machine behaviour.
Ten executions is enough. We break each one down, judge what happened, and mark exactly where confidence drops. Results in under thirty minutes.