Back to core deliverables
Client Use Case Matrix

Simulation Requirements Matrixfor Algorithm Research

The same simulation foundation plays a different role in each algorithm pipeline. This page cites public research for every path and states the limits of that evidence, so research teams can identify environment, data, ground-truth, evaluation and real-world feedback needs. It is not a named-customer delivery report.

Coverage

6 algorithm research scenarios

From offline data production to control-loop evaluation, every path feeds a data flywheel driven by real-world failure samples.

COMMON FOUNDATION

Shared simulation foundation

BASE 01

Platform and sensor digital twins

Capture hardware structure, dynamics, actuators, extrinsics, intrinsics, clocks and noise models into a versioned experimental baseline.

BASE 02

Authorable scenes and conditions

Turn nominal, boundary, failure and long-tail conditions into scene definitions that can be regenerated, with random seeds and dependency versions recorded.

BASE 03

A shared data and ground-truth protocol

Align images, point clouds, state, actions, trajectories and labels on one timeline so they can be used for training, replay and comparison.

BASE 04

Closed-loop evaluation and failure feedback

Compare algorithm versions against the same evaluation baseline, and turn real-world failures back into reproducible experiments and the next data task.

VLAUSE CASE 01

Vision-language-action models

From a natural-language task to verifiable embodied execution

Train the model to relate language instructions, visual state and action sequences, and to complete interpretable, reproducible tasks under physical constraints.

Research Basis

arXiv · 2026

AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild

The paper trains an end-to-end UAV VLA on a joint simulation and real-data pipeline, and validates it in both simulated and real environments. That supports unifying the task, observation, action and data protocol first, then moving to hardware validation.

Scope: The study focuses on autonomous UAV navigation. Moving it to a manipulator or another platform still requires a new action space, task boundary and safety constraints.

PAIN 01

Real demonstration capture is expensive; failure trajectories and safety-boundary data are especially scarce.

PAIN 02

Language, object relations, visual observations and robot actions are hard to keep aligned.

PAIN 03

Offline instruction-matching metrics do not prove that the model can execute a continuous task reliably.

Simulation needs in the pipeline

  1. 01

    Task definition

    Split the language goal into objects, relations, constraints, success conditions and forbidden behaviour.

  2. 02

    World construction

    Generate a physically stable, semantically queryable scene and retain object state and a scene graph.

  3. 03

    Trajectory production

    Collect expert, exploration, recovery and failure trajectories on a shared observation and action timeline.

  4. 04

    Closed-loop validation

    Replay the instruction across scene variants and check success, timeout, collision and safety violations.

Environment, platform and sensors

  • Interactive objects, joints, grasp points, occlusions and material properties.
  • Camera, depth, tactile or force feedback, synchronised with robot state.
  • Controllable randomisation of object pose, relations and initial state.

Data, ground truth and labels

  • Language–observation–action triples and step-wise task state.
  • Scene graph, object pose, contact state, and success or failure reasons.
  • Expert trajectories, recovery trajectories and safety-constraint labels.

Closed-loop evaluation and regression

  • Task completion, instruction following, path efficiency and safety constraints.
  • Generalisation regressions under object, phrasing and scene perturbations.
  • Failed-step localisation and version comparison from the same initial condition.

Versioned deliverables

  • Versioned task-scene packs and language templates.
  • Multimodal trajectory data and a failure-sample index.
  • Closed-loop task evaluation reports and reproducible configurations.

What the customer provides

  • Target platform, action space and control interface.
  • Business task definition, allowed behaviour and safety bounds.
  • A private language vocabulary or task-planning constraints.
PERCEPTIONUSE CASE 02

Multimodal perception

Cover the long tail that real capture cannot label stably, with explainable ground truth

Produce controllable data for detection, segmentation, occupancy, depth, optical flow and multi-sensor fusion, and test how stable the model stays under environment change and sensor degradation.

Research Basis

ICCV · 2025

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

The paper uses native 3DGS scene data directly for vision-language pretraining. That shows a queryable spatial representation, scene semantics and multi-view observations can form a perception-pretraining base, and supports a pipeline that starts from scene assets and opens into a data and label protocol.

Scope: The paper supports the scene-semantics and pretraining-data path. It does not replace calibration or noise modelling for a specific camera, LiDAR or radar.

PAIN 01

Real capture rarely yields complete, precise and time-synchronised multimodal ground truth at once.

PAIN 02

Extreme weather, occlusion, rare targets and fault states are under-covered.

PAIN 03

After the data distribution shifts, a fixed replay is a poor way to locate why the model degraded.

Simulation needs in the pipeline

  1. 01

    Data specification

    Define the vocabulary, modalities, projection model, coordinate frames, frame rate and label precision.

  2. 02

    Scene production

    Author target density, occlusion, lighting, weather, materials and dynamic interaction.

  3. 03

    Sensor simulation

    Configure camera, LiDAR and radar extrinsics, intrinsics, noise, latency and dropped frames.

  4. 04

    Distribution regression

    Map real false positives and misses onto scene factors, then top up the training and regression sets.

Environment, platform and sensors

  • Controllable lighting, weather, materials, occlusion, target density and motion.
  • Synchronisation and calibration across multi-camera, depth, LiDAR and radar.
  • Lens distortion, exposure, noise, returns, latency and fault models.

Data, ground truth and labels

  • RGB, depth, normals, semantic/instance masks, optical flow, 2D/3D boxes and occupancy ground truth.
  • Object attributes, visibility, occlusion rate, sensor pose and scene factors.
  • A shared coordinate frame, timestamps and cross-modal correspondences.

Closed-loop evaluation and regression

  • Slice evaluation by scene factor, range, scale, occlusion and weather.
  • Version regression on detection, segmentation, occupancy and depth metrics.
  • Sim-to-real distribution gap and failure-sample coverage reports.

Versioned deliverables

  • Versioned multimodal datasets and a label specification.
  • Scene and sensor configuration snapshots plus quality-check results.
  • Model-failure slices and a targeted data-top-up list.

What the customer provides

  • Sensor models, mounting positions, calibration results and data formats.
  • Business vocabulary, target classes and the evaluation protocol.
  • A small set of real samples that can be used for distribution calibration.
SLAMUSE CASE 03

Localisation and mapping

Turn trajectories, clocks, motion distortion and environmental degradation into repeatable experiments

Validate visual, inertial and point-cloud localisation and mapping on continuous motion, dynamic interference and sensor error, covering trajectory accuracy, map consistency and recovery.

Research Basis

ICCV · 2025

EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile Device

The paper reconstructs a real deployment environment as Gaussian Splats, then trains a navigation policy in Habitat-Sim. That validates a real-to-sim-to-real path connecting field capture, simulation compilation, training and hardware validation.

Scope: The object of study is a navigation policy, not a SLAM benchmark. Time sync, pose ground truth, ATE/RPE and map consistency remain the SLAM engineering requirements stated on this page.

PAIN 01

Real trajectory ground truth is expensive and site-limited, so systematic degradation is hard to cover.

PAIN 02

Time sync, extrinsic drift and motion distortion are difficult to isolate as variables.

PAIN 03

Dynamic objects, weak texture, repetitive structure and transparent or reflective scenes are hard to reproduce stably.

Simulation needs in the pipeline

  1. 01

    Trajectory design

    Define speed, acceleration, turns, vibration, loop closures and lost-lock recovery segments.

  2. 02

    Timed capture

    Generate camera, IMU, LiDAR and wheel-odometry data on a shared simulation clock.

  3. 03

    Degradation injection

    Control noise, bias, latency, dropped frames, extrinsic change and dynamic interference.

  4. 04

    Map regression

    Compare pose, map and recovery behaviour on the same world and trajectory.

Environment, platform and sensors

  • Indoor and outdoor continuous space, loop closures, ramps, narrow passages, weak texture and repetitive structure.
  • Clock and coordinate relationships among camera, IMU, LiDAR, wheel odometry and GNSS.
  • Motion distortion, sensor bias, dynamic occlusion and positioning-signal degradation.

Data, ground truth and labels

  • High-rate pose, velocity, acceleration, depth, point-cloud and map ground truth.
  • Per-sensor timestamps, extrinsics, biases and fault events.
  • Static/dynamic regions, loop-closure locations and observability labels.

Closed-loop evaluation and regression

  • Absolute and relative trajectory error, loop-closure consistency and map overlap.
  • Lost-lock rate, recovery time and stability under degradation.
  • Version-diff localisation on a fixed trajectory and random seed.

Versioned deliverables

  • A repeatable trajectory library and multi-sensor sequences.
  • Pose and map ground truth plus a clock-calibration checklist.
  • Degraded-condition regression reports and a failed-segment index.

What the customer provides

  • Sensor suite, rates, synchronisation method and coordinate conventions.
  • Motion-platform dynamics envelope and typical routes.
  • Algorithm I/O interfaces and map format.
RL CONTROLUSE CASE 04

Reinforcement learning and motion control

Explore and regress policies inside a safe, reproducible physics loop

Build a parallelisable dynamics environment for robots, quadrupeds, UAVs and other intelligent hardware, and test how stable a policy stays under disturbance, latency and faults.

Research Basis

arXiv · 2025

GRaD-Nav: Efficiently Learning Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics

The paper puts a 3DGS visual environment, differentiable UAV dynamics and deep RL into one training system, then validates sim-to-real on a real UAV. That supports designing the environment, dynamics, policy training and deployment regression together.

Scope: The study focuses on visual UAV navigation. Other robot platforms still need real test data to recalibrate dynamics, contact and actuator models.

PAIN 01

Real exploration risks hardware damage and is rate-limited by wall-clock physics.

PAIN 02

A mismatch in simulated dynamics, actuators or control-chain latency widens the sim-to-real gap.

PAIN 03

Reward changes, randomisation ranges and training versions rarely form an auditable experiment.

Simulation needs in the pipeline

  1. 01

    Dynamics modelling

    Configure mass, inertia, joints, actuators, contact, aerodynamics or wheel-ground interaction.

  2. 02

    Task and reward

    Define the goal, reward, constraints, termination conditions and safety guards.

  3. 03

    Parallel training

    Support resettable episodes, randomisation ranges, curriculum learning and seed records.

  4. 04

    Policy regression

    Compare policies under nominal conditions, disturbances, faults and hardware-calibrated parameters.

Environment, platform and sensors

  • Platform dynamics, actuator saturation, control rate, latency and contact models.
  • Randomisation of terrain, payload, wind, friction, external force and structural parameters.
  • State estimation, IMU, encoder, vision or point-cloud observation interfaces.

Data, ground truth and labels

  • State, action, reward, constraints, contact forces and termination reasons.
  • Dynamics parameters, random seeds, environment versions and policy versions.
  • Expert trajectories, failed episodes and safety events.

Closed-loop evaluation and regression

  • Task return, success, stability, energy use and constraint violations.
  • Robustness under parameter disturbance, observation noise, latency and faults.
  • Policy-version comparison on a fixed seed and a hardware-calibrated domain.

Versioned deliverables

  • Versioned training environments, tasks and reward configurations.
  • Episode trajectories, control logs and failure samples.
  • A policy regression matrix and a hardware-calibration parameter list.

What the customer provides

  • Platform CAD/URDF, mass and inertia, actuators and the control interface.
  • Task goals, safety constraints and the real control rate.
  • Test data that can be used to calibrate the dynamics.
WORLD MODELUSE CASE 05

World-model training

Produce action-conditioned, multimodal, traceable long-horizon world data

Train the model to predict future state, video, occupancy or a latent from past observations and actions, and evaluate long-horizon rollout consistency and out-of-distribution degradation.

Research Basis

arXiv · 2025

Learning to Drive from a World Model

The paper builds an on-policy simulator from real driving data, then trains and closed-loop evaluates a driving policy inside a learned world-model simulation. That supports a world model that does more than emit future observations: it also has to host action-conditioned rollouts and policy feedback.

Scope: The paper validates an autonomous-driving world model. A robot or general physics world model still needs its own state, action, event and long-horizon consistency metrics.

PAIN 01

Actions, environment changes and future outcomes in real logs are hard to control as isolated variables.

PAIN 02

Long-horizon data easily picks up timeline breaks, label drift and sparse key events.

PAIN 03

Open-loop prediction metrics do not capture accumulated rollout error or controllability.

Simulation needs in the pipeline

  1. 01

    World-state definition

    Pin down observable state, hidden state, actions, events and the future prediction targets.

  2. 02

    Sequence generation

    Run long episodes from a policy or a script and keep replayable world snapshots.

  3. 03

    Multimodal alignment

    Synchronise vision, point clouds, occupancy, state, actions, events and scene parameters.

  4. 04

    Rollout evaluation

    Check short-horizon accuracy, long-horizon stability and action controllability in and out of distribution.

Environment, platform and sensors

  • Dynamic agents, interaction events, causal variables and a world state that can keep evolving.
  • A shared timeline for cameras, point clouds, occupancy, maps, state and actions.
  • Controllable policy distributions, environment changes and out-of-distribution conditions.

Data, ground truth and labels

  • Long-horizon multimodal observations, actions and next-state ground truth.
  • World-state snapshots, event boundaries, object trajectories and scene factors.
  • Branched rollouts, counterfactual scenes and replayable seeds.

Closed-loop evaluation and regression

  • Single-step and multi-step prediction error, plus rollout stability.
  • Action-conditioned consistency, event prediction and out-of-distribution degradation.
  • Behaviour comparison on a fixed world branch before and after a model update.

Versioned deliverables

  • Versioned long-horizon data packs and a world-state protocol.
  • Action-conditioned data, an event index and branched rollouts.
  • Long-horizon prediction and out-of-distribution regression reports.

What the customer provides

  • The target world representation, action space and prediction modalities.
  • Sequence length, sampling rate and the event vocabulary.
  • Real log formats and the out-of-distribution problems that matter most.
FSD / E2EUSE CASE 06

Full-stack autonomous driving and end-to-end models

From traffic-world generation to scored perception–prediction–planning–control loops

Validate the closed-loop behaviour of a modular or end-to-end driving stack under controllable traffic interaction, covering long-tail events, policy games, system coupling and version regression.

Research Basis

ICCV · 2025

DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving

The paper combines a generative world model with a traffic manager into a closed-loop simulation arena, so a driving agent can run, be tested and be developed inside ongoing traffic interaction. That supports moving FSD evaluation from fixed replay to a closed-loop environment that can respond to the ego vehicle.

Scope: A generative closed-loop environment still has to be calibrated against the company ODD, vehicle dynamics, safety rules and real events. It cannot replace road testing on its own.

PAIN 01

Fixed-log replay cannot show how other traffic participants respond after the ego action changes.

PAIN 02

A local improvement in perception, prediction, planning or control can still degrade the system.

PAIN 03

Dangerous and rare events are hard to capture often, or retest, on real roads.

Simulation needs in the pipeline

  1. 01

    ODD and event definition

    Structure the roads, traffic flow, rules, weather and key interaction events.

  2. 02

    Full-stack integration

    Connect perception, prediction, planning, control or an end-to-end model, plus vehicle dynamics.

  3. 03

    Closed-loop run

    Keep the ego vehicle and other participants interacting, and record the causal chain and intermediate state.

  4. 04

    Event regression

    Grow scene variants around collision, takeover, rule-break, discomfort and failure samples.

Environment, platform and sensors

  • Road topology, traffic rules, signals, traffic flow and multi-agent interaction.
  • Vehicle dynamics plus camera, LiDAR, radar, localisation and map interfaces.
  • Weather, lighting, construction, occlusion, irregular targets and sudden events.

Data, ground truth and labels

  • Sensor data, vehicle state, object trajectories, maps and traffic-light ground truth.
  • Module intermediates or end-to-end actions, event chains and takeover reasons.
  • Scene parameters, participant policies, random seeds and software versions.

Closed-loop evaluation and regression

  • Collision, takeover, rule following, comfort and task completion.
  • Key-event recall, interaction plausibility and closed-loop pass results.
  • End-to-end models versus a modular baseline on the same scene set.

Versioned deliverables

  • An ODD/event scene library and traffic-flow configurations.
  • Full-stack closed-loop logs, event slices and failure replay.
  • Version regression reports and a targeted retest list.

What the customer provides

  • Vehicle and hardware configuration, software interfaces and the target ODD.
  • Safety rules, takeover definitions and the company evaluation standard.
  • Anonymised real events, logs and a list of the problems that matter most.

BEFORE IMPLEMENTATION

Scoping checklist

These details set the boundary of the simulation work, the form of delivery and the evaluation protocol. Gaps should be closed during scoping, not filled with generic data or default parameters.

  1. 01Target algorithm, I/O interfaces and current research stage
  2. 02Platform structure, dynamics, sensor configuration and calibration materials
  3. 03Operational scenes, long-tail problems, failure samples and safety bounds
  4. 04Data format, label vocabulary, coordinate frames and version rules
  5. 05Training framework, deployment environment and closed-loop evaluation protocol
  6. 06Anonymised real data and logs that can be used for calibration

Imitatio Artis embeds through forward-deployed engineering, turning non-standard platforms, data, scenes and evaluation protocols into modules that keep evolving.