Closed-loop driving safety benchmark

NavSafe-∞: Benchmarking Closed-Loop Driving Safety
in Photorealistic Environments

Yuxin Bao1,* , Hongwei Ruan2,* , Luobin Wang2,† , Seth Z. Zhao1,† , Ziyang Leng1 , Zihan Zhang2 , Yu Zeng3 , Rowan McAllister3 , Henrik Christensen2 , Bolei Zhou1

1UCLA  ·  2UCSD  ·  3Toyota Research Institute

*Equal contribution; order by last name   †Corresponding authors

Intersection traversalTraffic crashes
Pedestrian on crosswalkVulnerable road users
Cones in the drivable pathTraffic incidents
Signalized intersectionTraffic violations
Car-park maneuverTraffic crashes

Photorealistic reconstructions of nuPlan scenes, front camera.

280scenarios
28safety event types
20E2E policies evaluated
4safety categories
Open-loop policies diverge under closed-loop execution
Comparable open-loop scores can lead to different state sequences and safety outcomes once policies act in the environment.
01 · Benchmark

NavSafe Benchmark

Event-based evaluation in photorealistic closed-loop environments.

NavSafe evaluates 280 event-based scenarios across 28 event types, adapted from RoadSafe365. Each bounded scenario isolates a safety capability while preserving the feedback between a policy and its environment.

TC

Traffic crashes

Respond to vehicle conflicts and collision hazards.

TI

Traffic incidents

Resolve work zones, obstructions, and unexpected conflicts.

TV

Traffic violations

Follow traffic controls and event-specific rules.

VRUC

Vulnerable road users

Yield to pedestrians and other vulnerable road users.

Representative scenarios from the four NavSafe safety categories
Representative safety events in reconstructed, photorealistic environments.

Closed-Loop Simulation

The renderer produces camera observations from the current ego and actor states. The policy proposes a plan, the simulator advances, and the next observation reflects the consequences. Reconstructed 3D scenes support configurable camera rigs and controlled edits to traffic elements and actor behavior.

NavSafe closed-loop simulation process
An event-specific reconstruction connects camera observations, policy planning, and simulation feedback.
Rendering fidelity over a long-horizon scenario
Rendering comparison at 0, 5, 10 and 15 seconds
The paper compares first-person views and bird’s-eye views across DriveArena, DreamStream, and NavSafe; NavSafe maintains temporal consistency in this example.

Evaluation Metrics

Driving score (DS)Event progress multiplied by infraction penalties, scaled to 0–100.

Success rate (SR)Events completed within the time budget without a disallowed infraction.

Driving efficiency (DE)Ego speed relative to surrounding traffic.

ComfortCompliance with acceleration, jerk, yaw-rate, and yaw-acceleration bounds.

Category ability scores are the unweighted mean of event-type success rates within TC, VRUC, TV, and TI.

02 · Results

Leaderboard

Strong open-loop performance does not guarantee closed-loop safety.

The best learned policy reaches 61.79% success, compared with 81.31% for the privileged PDM-Closed reference. No learned policy leads all four safety categories.

Section 4.1 main-table results · 280 scenarios · 3 seeds (0, 1, 1024). Values are mean ± sample standard deviation. Select a metric heading to sort; the human reference remains first, followed by the privileged reference and learned policies.

Policy Venue
Human Driver ReferenceReference -- 100.00± 0.00 100.00± 0.00 316.52± 0.42 75.15± 1.89 100.00± 0.00 100.00± 0.00 100.00± 0.00 100.00± 0.00
PDM-ClosedPrivileged reference CoRL 2023 90.17± 1.00 81.31± 1.03 305.27± 8.57 82.49± 1.39 88.67± 1.53 72.50± 0.00 84.24± 1.39 57.78± 1.92
GTRS-Dense-V2-99 (SimScale)Passive Demonstration Perturbation Methods CVPR 2026 73.54± 0.84 61.79± 1.56 141.50± 1.30 87.69± 1.66 69.33± 0.58 50.00± 2.50 64.24± 1.89 43.33± 3.33
SparseDriveV2IL-based Methods ECCV 2026 65.26± 1.72 48.81± 3.04 167.99± 1.01 93.41± 0.43 54.00± 0.00 43.33± 8.04 48.18± 4.17 41.11± 3.85
DrivoRIL-based Methods CVPR 2026 64.74± 1.60 50.71± 0.71 188.02± 1.07 87.92± 1.25 58.67± 2.52 31.67± 5.77 52.42± 2.62 43.33± 0.00
GTRS-Dense-V2-99IL-based Methods arXiv 2025 64.27± 0.76 49.52± 1.15 163.88± 0.87 95.90± 0.63 52.00± 1.73 40.00± 2.50 54.24± 0.52 36.67± 0.00
DrivoR (SimScale)Passive Demonstration Perturbation Methods CVPR 2026 57.66± 1.32 42.50± 0.36 176.73± 0.92 85.88± 0.39 44.00± 1.00 17.50± 2.50 48.18± 2.41 50.00± 0.00
SimWAMWorld-Model-based Methods arXiv 2026 57.53± 2.72 37.50± 1.56 160.54± 1.64 97.19± 0.27 41.33± 0.58 26.67± 7.22 40.61± 0.52 27.78± 5.09
MTDrive-mtGRPORLFT-based Methods arXiv 2026 55.86± 0.50 40.60± 0.90 277.62± 3.17 93.94± 0.41 40.33± 2.89 16.67± 1.44 46.36± 0.91 52.22± 3.85
ReCogDrive-2B-RLRLFT-based Methods ICLR 2026 53.64± 1.42 30.95± 2.06 173.96± 1.90 85.10± 0.49 35.33± 2.31 18.33± 6.29 32.42± 1.39 27.78± 5.09
LTFIL-based Methods NeurIPS 2024 53.04± 1.92 27.74± 1.76 163.71± 0.11 97.81± 0.15 31.67± 0.58 25.00± 7.50 28.48± 1.89 15.56± 1.92
LTF (SimScale)Passive Demonstration Perturbation Methods CVPR 2026 52.65± 0.88 34.05± 0.21 192.09± 0.85 97.41± 0.48 35.00± 1.73 17.50± 2.50 38.48± 0.52 36.67± 0.00
DiffusionDrive (SimScale)Passive Demonstration Perturbation Methods CVPR 2026 52.19± 2.05 35.24± 2.38 189.66± 4.28 96.92± 0.34 38.33± 1.15 18.33± 7.22 36.67± 5.01 42.22± 3.85
AutoVLARLFT-based Methods NeurIPS 2025 51.05± 0.80 29.52± 2.43 186.02± 0.58 65.37± 0.47 33.00± 2.65 15.00± 4.33 29.70± 2.92 36.67± 8.82
ReCogDrive-2B-ILIL-based Methods ICLR 2026 48.98± 0.95 28.21± 1.43 166.94± 3.15 97.58± 0.11 30.00± 1.00 16.67± 3.82 31.21± 0.52 26.67± 3.33
MTDrive-SFTIL-based Methods arXiv 2026 48.78± 0.88 31.79± 1.43 173.77± 0.76 94.51± 1.25 30.33± 1.15 15.00± 6.61 35.76± 1.39 44.44± 1.92
RAPIL-based Methods ICLR 2026 47.70± 0.66 27.26± 0.41 220.78± 3.77 48.38± 0.60 32.33± 1.15 14.17± 6.29 28.79± 0.52 22.22± 3.85
DiffusionDrive (BeyondDrive)Passive Demonstration Perturbation Methods ECCV 2026 46.61± 0.07 21.79± 0.94 172.79± 1.26 97.24± 0.19 22.67± 1.53 10.00± 0.00 25.15± 1.39 22.22± 1.92
DiffusionDriveIL-based Methods CVPR 2025 44.77± 0.75 20.95± 0.55 165.08± 1.84 97.68± 0.56 21.33± 1.15 11.67± 1.44 23.64± 0.00 22.22± 1.92

19 policies and references

Human Driver Reference is the paper’s hybrid upper reference using front-camera observations. PDM-Closed is privileged and non-learned; neither is ranked against learned policies. DE is a relative-speed measure and can exceed 100.

03 · Analysis

Experiments

We investigate how driving policies fail in closed-loop execution and why common training remedies do not consistently improve safety.

Failure Investigations

Planning · LTF

Hazard-conditioned planning

LTF detects crossing traffic at 29 m but continues accelerating and brakes too late. The observed plan remains consistent with routine car following despite the crossing conflict.

Scoring · DrivoR

Trajectory scoring

DrivoR executes its top-ranked trajectory and collides with the lead vehicle. A safe stopping trajectory exists in the candidate set, but is ranked 43rd.

LTF planning failure and DrivoR trajectory scoring failure
A planning failure and a scoring failure expose different points of breakdown.
Prediction · SimWAM

Safety-critical actor prediction

SimWAM omits an oncoming vehicle from its future prediction and plans through the space it occupies, removing a safety-critical constraint.

Reasoning · AutoVLA

Reasoning–action alignment

At the same instant in two policy-induced states, AutoVLA identifies crossing pedestrians and calls for stopping. Only one decoded trajectory realizes that decision.

SimWAM prediction failure and AutoVLA reasoning-action mismatch
Prediction errors and state-dependent reasoning–action alignment can lead to collisions.

Passive Demonstration Perturbation

Passive perturbation pairs perturbed ego states with expert recovery trajectories. Its effectiveness depends on whether closed-loop rollouts remain covered by the perturbed state distribution. Because these demonstrations are generated independently of the policy, they may miss the states visited during execution.

Anchor-based policies (DiffusionDrive and GTRS) retain stable, expert-like candidates; training mainly improves scoring and keeps rollouts closer to the perturbed distribution. Query-based policies (DrivoR and LTF) improve open-loop scores but still drift beyond that distribution. The figure diagnoses this gap through state shifts and Monte Carlo returns, complementing the benchmark success rates below.

DiffusionDrive + SimScaleSuccess rate: 20.95% → 35.24%▲ 14.29

DrivoR + SimScaleSuccess rate: 50.71% → 42.50%▼ 8.21

LTF + SimScaleSuccess rate: 27.74% → 34.05%▲ 6.31

GTRS-Dense-V2-99 + SimScaleSuccess rate: 49.52% → 61.79%▲ 12.27

State distribution shifts and Monte Carlo returns under demonstration perturbation
State-distribution shifts and Monte Carlo returns reveal how trajectory proposals and scoring affect the transfer of perturbation training to closed-loop execution.

Reinforcement Learning Fine-Tuning

RLFT scores trajectories against a frozen logged future. Because the proxy rewards progress more directly than safety margin, open-loop reward gains can come at the expense of safety. Repeated closed-loop replanning compounds this risk, with collisions ending safety-critical rollouts early.

ReCogDrive largely preserves safety-related scores, allowing gains to transfer. MTDrive gains progress at the cost of safety, reducing driving scores in the safety-critical diagnostics shown below. These diagnostics are distinct from the aggregate leaderboard.

Open-loop and closed-loop RLFT reward exploitation diagnostics
Open-loop reward exploitation trades safety margin for progress. The trade-off persists in nominal closed-loop planning, while collision-induced early termination amplifies its cost in safety-critical rollouts.