SIMULATION FRAMEWORK / RL-BASED VLA TASKS

Simulate a path.
Check every step.

A new rollout-native simulation framework for RL-based VLA tasks, built on Safety-CHORES. Explore native continuations, reconstruct decision states with reset and replay, and check selected trajectories during execution.

NATIVE EXECUTIONBasketball grasp · FirstSuccess
Two camera views. One checked trajectory.8 actions/s · accelerated playback GIF ↗

Selected native example: the 75-action execution matched every saved state/transition check. This single run does not estimate a benchmark success rate. Inspect the run receipt ↗

01 / HOW IT WORKS

Return to the same state.
Explore a different future.

Every branch owns a fresh native episode. Workers rebuild the search state by replaying actions, including policy inference and actual-action feedback.

  1. 01

    Reset

    Start the same task with the same episode seed in a fresh simulator.

    task + seed → initial state
  2. 02

    Replay the prefix

    Execute the saved live actions to rebuild simulator state and policy history. Check the reconstructed root.

    saved actions → search state
  3. 03

    Continue & compare

    Try a candidate action, then sample a continuation. Compare task completion and recorded cost.

    candidate + suffix seed → branch
THE PRACTICAL SEARCH

FirstSuccess stops after the first valid worker batch containing a completed zero-cost trajectory. It retains the whole path and checks each executed transition. A mismatch stops execution before the next action.

Read the execution contract ↗
02 / EXPLORE THE RECORDED BRANCHES

One root.
More than one outcome.

Drag the timeline to see the shared prefix and the paths that follow. These coordinates come from native simulator replays.

01 / RESET
Task 1 · seed 1909
Recorded robot base trajectoriesInteractive world X/Z map showing the shared replay prefix and different sampled continuations.
Shared prefixFaint lines show complete recorded paths
0 / 200

Loading the recorded trajectories…

Robot base X/Z coordinates in metres. Room outlines omit doors and furniture; they are not a collision map. Rotation and arm motion may leave the base in place.

Selected branches from the earlier BestOfN study: 3 of 12 at the 40-action root, and 2 of 4 at the initial root. Captured sequentially, displayed by action count. These are separate from the FirstSuccess demonstration.

03 / WHERE THIS FITS

From task benchmarks
to rollout-native search.

Safety-CHORES-RL provides a rollout-native API over Safety-CHORES: request candidate continuations, reconstruct their native simulator roots, and check the path during execution.

Benchmark scope and the tooling provided here
DimensionLIBERO ↗Safety-CHORES ↗Safety-CHORES-RL Ours · rollout-native tooling
Primary purposeKnowledge transfer in lifelong robot learning.Safety evaluation and constrained learning for embodied policies.Rollout-native search: sample native simulator futures from a decision state, then check execution.
Tasks & embodiment130 manipulation tasks: 10 Spatial, 10 Object, 10 Goal, and 100 mixed tasks (90/10 split).ObjectNav: find an object. PickUp: grasp an object in reach. Fetch: navigate, then grasp.ObjectNav, PickUp, and Fetch through the native Safety-CHORES task adapter, with navigation and mobile manipulation.
Episode protocolTask suites with human demonstrations and task-conditioned evaluation.Published evaluation: 200 ObjectNav, 171 PickUp, and 172 Fetch cases; up to 600 actions per episode.All task rows in the configured Safety-CHORES split. Select task IDs, episode seeds, and action caps; sample branches with independent continuation seeds.
Safety lensTask success and transfer are the original benchmark's main focus.Explicit safety constraints and costs alongside task completion.Retains branch success, costs, and failed outcomes; checks the chosen trajectory step by step.
State & replayEnvironment wrapper exposes simulator state and state-setting APIs.Provides the underlying task, simulator, and safety signals.Rebuilds simulator and policy history with reset + prefix replay in isolated workers.
Role in RLOriginal suite emphasizes lifelong imitation learning and transfer; environments also expose step rewards.Constrained policy training with task reward and explicit safety cost.Rollout backend for candidate evaluation and data collection: per-step rewards/costs, termination, truncation, and seeded continuations.
What you getTask suites, demonstrations, training and evaluation code.Safety task infrastructure and SafeVLA training/evaluation code.Native rollout requests, budgeted search, saved trajectories, execution checks, and visual replay receipts.

Sources: LIBERO repository, LIBERO state API, Safety-CHORES task protocol (Table 13), and upstream code. Comparison of task scope and interfaces.

What rollout-native adds for RL

01 / LEARNING SIGNALS

Reward and cost, separately.

Each transition records the action, reward, safety cost, and completion flags. Branch results distinguish termination from truncation and include prefix cost in the total.

02 / BRANCH SAMPLING

Revisit a decision state.

Evaluate candidate actions with independently seeded policy continuations. Workers reconstruct simulator state and policy history while the live episode stays unchanged.

03 / COMPUTE BUDGETS

Bound the collection cost.

Limit rollouts, native actions, and wall time. Reuse identical requests, run isolated workers, and retain the selected trajectory for checked execution.

Connect a learner through the runtime and policy interfaces. This repository supplies rollout and search infrastructure; the training loop, replay buffer, and optimizer are supplied by the caller.

04 / EVIDENCE & SCOPE

Measured behavior.
Visible limitations.

Receipts retain the selected actions, outcomes, timings, and verification scope for the tested native environment.

4 + 1Fetch cases

Four main task rows and one separate canary.

10 / 10full episode replays matched

Five cases × two controllers, each replayed.

100 / 100valid sampled continuations

96 main branches and 4 canary branches.

Episode outcomes, action counts, and safety costs

The paired study compares policy sampling with BestOfN first-action search. Queries occur at steps 0 and 40 in the main cases. A task row is one benchmark instance; a branch is one sampled continuation, not an additional task.

Each cell reports outcome · executed actions · cumulative cost
Fetch casePolicy samplingFirst-action search
Row 0 · seed 1909Failure · 162 actions · cost 0Failure · 137 actions · cost 0
Row 1 · seed 1909Failure · 200 actions · cost 0Success · 148 actions · cost 2
Row 2 · seed 1909Failure · 117 actions · cost 13Failure · 138 actions · cost 2
Row 3 · seed 1909Failure · 200 actions · cost 4Failure · 200 actions · cost 20
Canary · row 0 · seed 892Success · 78 actions · cost 0Success · 78 actions · cost 0

The main cases contain zero zero-cost successes for either controller; both canary episodes succeed at zero cost. The selected FirstSuccess demonstration above is a separate run that retains the entire successful continuation. Episode receipt ↗

75 / 75live execution checks matched

The FirstSuccess example compares each executed state and transition with its saved witness. Search took 180.3 s; the complete run took 339.8 s.

Run receipt ↗
5 branchesregenerated from saved seeds

All 585 continuation actions and 120 prefix actions reproduced their recorded transition evidence. The captured images passed 1,420 camera hash checks.

Verification scope ↗
45 testspassed in an isolated Python environment

The same 75-action native result reproduced on Apex. Existing simulator inputs were reused; six dependency constraints remain unresolved.

Environment receipt ↗
What these results establish—and what remains open

Workers reconstruct state through reset and action replay. The API does not clone hidden physics state. The FirstSuccess example checked original per-step state hashes; the older branch study retained transitions and endpoint observations, but not original root snapshots or intermediate state hashes.

Full captured prefix states agree across the new branch reconstructions. This is bounded replay evidence, not a general determinism guarantee. Zero recorded cost in an example does not establish general safety or policy efficacy.

The native player, assets, SafeVLA source, and checkpoint are external. The separate four-case tooling study had no safe successes for either evaluated controller. Read the validation and limitations for all outcomes, dependency conflicts, and verification scope.

05 / GET STARTED

Try the API.
Bring your simulator.

The portable examples run without a native simulator. Native experiments require the documented player, assets, policy, and checkpoint.

PORTABLE EXAMPLES
git clone https://github.com/tumeteor/safety-chores-rl.git
cd safety-chores-rl
python -m pip install -e ".[test]"
python -m pytest -q
python examples/fast_search.py --workers 2

Native workflow: reconstruct → search → retain witness → checked execution.