Built by a team that has shipped data infrastructure to:

Anduril Lockheed Martin Palantir Denso Continental John Deere Blue River Technology Auterion Quantum Systems TurbineOne Atoms and Pronto
Previously DiffuseDrive
Behavioral observability for autonomous systems

You can do thousands
of policy rollouts.
You cannot possibly watch all of them.

Simulation scales. Fleets scale. The engineer who has to watch the footage does not. Every rollout, in the lab and in the field, waits on a person to say what happened before it means anything. We take that person out of the loop and leave them the five runs that actually needed them.

View A · workspace
Evaluate against frames re-read  0
Evidence
TSignalMeasurementConsequence
Switching the rubric re-runs the judgement, not the perception. The record of what happened was built once.
‹ illustrative · sample executions ›

Different machines, different signals, different phase vocabularies. Same record, same three questions asked of it. Note the humanoid: it completed the task and still should not have shipped.

Evaluation is not the first step. Before anything can be judged, what happened has to be represented: entities, contacts, behavioural phases, events and outcomes, each carrying where it came from and how confidently it is known. Build that representation once and evaluation becomes a layer on top of it, changeable at will. The representation is the product. Everything else is a question asked of it.

01

Every boundary has a source

Phases are cut where the physics changes: a force step, a velocity inflection, a current spike, found by deterministic extractors. The model supplies the words. No timestamp or magnitude in your output originates in a language model.

02

Understand once, evaluate forever

What happened is stored as structure, not as a caption. Change what counts as success and nothing re-reads your footage. Judge the same hundred thousand executions against completion, safety, quality or your own criteria.

03

Coverage, not sampling

A human evaluator samples. A machine-readable record lets you reason across every execution you generate. That is where the rare failure modes live, and where the regressions that aggregate metrics hide actually turn up.

Connects to what you already run.

We are not another simulator and not another storage format. Point us at executions wherever they come from. We normalise them into one representation, and everything downstream works the same way.

Simulators

  • Isaac Sim
  • MuJoCo
  • Gazebo
  • your own

Benchmarks & harnesses

  • LIBERO
  • RoboCasa
  • VLA eval harnesses
  • custom scenarios

Logs & datasets

  • ROS bags · MCAP
  • RLDS · TFDS
  • HDF5
  • DROID · Open X-E

Real hardware

  • lab test rigs
  • teleop recordings
  • deployed fleets
  • onboard telemetry

Atomic Skills Industries is the team formerly known as DiffuseDrive, whose synthetic-data platform for autonomous systems is used by teams at the companies above. Perception and data infrastructure for physical AI, out of automotive and autonomous driving, now pointed at machine behaviour.

Test your own rollouts.

Ten executions is enough. We break each one down, judge what happened, and mark exactly where confidence drops. Results in under thirty minutes.