← Project website

TECHNICAL NOTE · NATIVE EXECUTION

Native rollouts

Reconstruct a native decision state, sample policy continuations, and check the selected trajectory during execution. This note covers the tested Safety-CHORES environment and the commands behind the recorded demonstrations.

Setup

Use Linux with NVIDIA/Vulkan and activate your existing SafeVLA Python environment. The setup helper locates the installed sources, modified player, task/house data, assets, and checkpoint, then generates the configs and launcher:

python -m pip install -e ".[native]" pillow
python tools/setup_native.py
bash runs/native-setup/run_demo.sh --save-frames

It prints the full number of task rows in your configured benchmark. The default asset preflight checks task 0; it does not restrict the available episode set. The simulator and Python environment must already be installed. See the tested environment and dependency constraints.

Discovery, checkpoint download, and generated files
# Use a different installation root, or reuse a previously pinned config.
python tools/setup_native.py --installation-root ../SafeVLA
python tools/setup_native.py --from-config runs/native.json

# If the task's official checkpoint is not installed, download it (~2 GB).
python tools/setup_native.py --download-checkpoint

# Check assets for more task IDs without changing the benchmark file.
python tools/setup_native.py --task fetch --task-ids 0 1 2 3 --output runs/fetch-setup

These are alternative setup commands. Discovery checks standard directories under the installation root, sibling SafeVLA directories, and existing SAFEVLA_REPO, SAFEVLA_CHECKPOINT, OBJAVERSE_*, PYTHONPATH, and SAFETY_CHORES_PLAYER settings. SAFETY_CHORES_NATIVE_CONFIG can identify an existing config. Ambiguous or missing inputs produce an explicit error and an override flag; run python tools/setup_native.py --help for all options.

The helper writes native.json, policy.json, env.sh, run_demo.sh, and a setup.json receipt under runs/native-setup. It hashes the checkpoint, verifies the player's artifact manifest, checks assets for the requested task IDs, and preserves existing source and data files. Repeating an identical setup is safe; use a new output directory when changing inputs. For the other recipes in this note, first run source runs/native-setup/env.sh.

Optional downloads use the official SafeVLA weights and verify their SHA256: safe_fetch.pt, safe_pickup.pt, or safe_objnav.pt. The Fetch checkpoint is the same artifact used in the recorded demonstration. Use --checkpoint for a custom policy. The policy loader still checks model compatibility; test-time image augmentation is disabled.

Keep generated files local because they contain installation paths. Artifact checks configure the installation; native execution is checked when you run the demo. This helper does not install the simulator or rebuild the Python environment. Proxy settings are inherited from the calling shell.

First-success demonstration

The selected Fetch case asks the robot to navigate to a basketball and grasp it. FirstSuccess evaluates candidates until a valid worker batch contains a completed zero-cost continuation, then retains and executes the entire path.

Task / seed Search Checked execution Recorded cost
0 / 1909 2 of 8 allowed rollouts 75 of 75 actions matched 0

On one shared H100 with one worker, search took 180.3 s and the complete run 339.8 s. Candidate 2 failed after 85 actions; candidate 1 succeeded after 75. Both used continuation seed 4100. Run receipt · Watch the demonstration.

bash runs/native-setup/run_demo.sh \
  --out runs/first-success --task-id 0 --seed 1909 \
  --vulkan-device 0 --workers 1 --candidates 2 \
  --sample-seeds 4100 4101 4102 4103 --episode-limit 200 --save-frames

Use a new output directory and the physical Vulkan GPU index; CUDA visibility alone does not select it. xvfb-run provides the display required by the upstream controller. The task and seed were selected using earlier experiments.

Execution contract, budgets, and outputs

The native signatures check camera images, exposed physics/cost state, policy history, and source identity. Per-step execution comparisons exclude the policy sampling RNG deliberately forked for proposals; full root validation includes it. See the checked-execution API.

Other evaluation recipes: paired study and integration probe

The paired study compares ordinary policy sampling with first-action search at declared query steps. Both full live episodes are replayed; failures, unsafe outcomes, and invalid branches remain in the report.

PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 xvfb-run -a \
  python examples/native_task_study.py \
  --config runs/native-setup/native.json --policy-config runs/native-setup/policy.json \
  --out runs/task-study --task-id 0 --seed 1909 --episode-limit 200 \
  --vulkan-device 0 --workers 2 --query-steps 0 40 \
  --candidates 3 --sample-seeds 4100 4101 4102 4103

A smaller integration probe uses uniform-random continuations and executes two live actions after a fixed prefix:

PYTHONHASHSEED=0 xvfb-run -a python examples/native_search.py \
  --config runs/native-setup/native.json --out runs/native-search \
  --vulkan-device 0 --episode-limit 150 --samples 2 --live-steps 2 \
  --wall-seconds 600 --prefix 0 0 0 0 0 0 0 0

native_search_validated=true means the requested rounds reconstructed their roots and produced valid branches before execution. Read task success, cost, and length separately. This probe supplies no live safety filter, and its action IDs depend on the pinned vocabulary. --planner halving selects successive halving; --policy-factory module:Class --policy-config policy.json supplies a custom continuation policy. Both runners retain diagnostics and need an outer timeout for unattended operation.

Trajectory visualization

The original report contains state hashes and camera frames. Forced replay with the same runtime, configuration, and checkpoint recovers base positions from cached metadata without extra render or movement calls. For the selected run, both paths matched their original signatures and all 152 successful-path camera images matched the saved live frames.

Recover coordinates and render the animation
PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 xvfb-run -a \
  python examples/native_trajectory_replay.py \
  --config runs/native-setup/native.json --policy-config runs/native-setup/policy.json \
  --report runs/first-success/report.json --out runs/trajectory-replay \
  --vulkan-device 0

python -m pip install matplotlib pillow
python tools/plot_trajectories.py \
  --trajectories runs/trajectory-replay/trajectories.json \
  --out runs/first-success/visuals
python tools/render_first_success.py \
  --report runs/first-success/report.json --frames runs/first-success/frames \
  --trajectories runs/trajectory-replay/trajectories.json \
  --out runs/first-success/visuals

Omit --trajectories from the renderer for a camera-only animation. Replay is additional work, excluded from the reported search time. Every replayed camera hash must match the original live frame before it is used in the animated map.

Maps show base centres and headings in Unity world X/Z metres. Room outlines omit doors and furniture; rotation and arm motion may leave the base stationary. Playback runs at eight actions per second with a final hold, rather than wall time.

Branch comparison animations

The two comparisons use selected continuations from the BestOfN study. Action counts below exclude the shared prefix.

Case / root Candidate / continuation seeds Outcomes: actions, cost
Task 0, seed 892 / initial state 1 / 9002, 9001 Success: 74, 0; limit: 150, 0
Task 1, seed 1909 / 40 actions 0 / 4102, 4100, 4101 Success: 98, 0; success: 103, 3; limit: 160, 0

Each capture resets the episode, replays the live prefix and candidate, then regenerates the continuation from its saved seed. Checks cover every action, reward, cost, termination, root candidate probabilities, and final observation, including both camera hashes. Full reset/prefix states agree across the new reconstructions. The historical study did not save original root snapshots or intermediate state hashes, so the receipts distinguish these checks from historical per-step state equality.

Capture and render the recorded branches

Use the exact sources, configuration, and checkpoint from the original report. These clips used frozen runtime sources corresponding to d0de2d9, including the original file bytes and line endings that contribute to source identity.

PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 xvfb-run -a \
  python examples/native_branch_replay.py \
  --config runs/native-setup/native.json --policy-config runs/native-setup/policy.json \
  --report runs/canary-task0-seed892/report.json --query-step 0 \
  --selection 1:9002 1:9001 --out runs/branches-seed892 --vulkan-device 0

PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 xvfb-run -a \
  python examples/native_branch_replay.py \
  --config runs/native-setup/native.json --policy-config runs/native-setup/policy.json \
  --report runs/task1-seed1909/report.json --query-step 40 \
  --selection 0:4102 0:4100 0:4101 --out runs/branches-prefix40 --vulkan-device 0

python tools/render_branch_replay.py \
  --capture runs/branches-seed892/branches.json --out runs/branches-seed892.gif
python tools/render_branch_replay.py \
  --capture runs/branches-prefix40/branches.json --out runs/branches-prefix40.gif

The renderer verifies source-image hashes, preserves outcome/cost milestones, and labels reset, replay, and continuation. It keeps every second action by default, with holds at reset, the root, and the outcome. Columns align by action count; captures were collected sequentially. Navigation images keep their resolution; manipulation images are reduced to fit. Cyan marks the shared prefix.

For complete episode outcomes, environment checks, and operational limits, see validation and limitations. For interfaces and return types, see the API reference.