Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization

1University of Toronto 2Vector Institute 3NVIDIA Research 4ELLIS Institute Tübingen 5KE:SAI
*Work partly done while at NVIDIA Research.
Base VaVAM against VaVAM with ECO in HUGSIM

The same HUGSIM scene driven twice. The sampled base plan (red) has lateral inconsistencies in its interior; ECO keeps the endpoint, anchors the executed state and repairs the waypoints between them (blue). Base VaVAM later collides, VaVAM + ECO completes the route.

Abstract

End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or easy for the controller to track. We observe that these inconsistencies concentrate primarily at intermediate waypoints, while the predicted endpoint remains comparatively reliable. Based on this observation, we introduce Endpoint-Constrained Optimization (ECO), a lightweight post-processing layer that anchors the trajectory to the vehicle's executed history, preserves the policy's predicted endpoint, and reshapes the intermediate waypoints to improve feasibility. ECO requires no map, privileged simulator state, or additional training, and can be inserted between a broad range of waypoint-emitting policies and their controllers. Across two closed-loop simulators, it improves the aggregate closed-loop score of all six evaluated generative and regression-based driving policies, and the gains tend to increase with how often the base plans violate motion limits. On HUGSIM, ECO improves VaVAM from 18.1 to 31.0 HD-Score (+71%), achieving 1st place on the HUGSIM Closed-Loop Driving Challenge. Similarly, on AlpaSim, ECO increases the scene scores of VaVAM and DiffusionDrive by 123% and 22%, respectively. These results show that for a broad collection of end-to-end driving models, repairing the intermediate geometry of predicted trajectories without changing the policy's predicted endpoint can substantially improve closed-loop performance.

Video

Where the open-loop/closed-loop gap comes from

Displacement metrics such as ADE and FDE score each waypoint in isolation, so a plan can agree with the logged human drive and still ask the controller for accelerations and jerk the vehicle cannot deliver. Across the nuScenes episodes of HUGSIM, VaVAM's plans from collision episodes are closer to the human path (median 1.2 m) than plans from completed episodes (2.0 m). In closed loop, both collisions and motion-limit violations concentrate at the intermediate waypoints; the predicted endpoint remains comparatively reliable.

Collisions concentrate at intermediate waypoints
Open-loop metrics see neither the failure nor the repair

Endpoint-Constrained Optimization

ECO pipeline

ECO sits between the policy and its controller. It prepends the two most recently executed poses and the current pose to the predicted plan, holds those and the final predicted waypoint fixed, and optimizes the intermediate waypoints with an objective that balances smoothness, turn sharpness and fidelity to the predicted plan. The problem is solved with L-BFGS-B in a median of 20.7 ms on one CPU core. ECO consumes and returns the same fixed-period waypoint representation, needs no map, no agent states, no privileged simulator state and no training, and uses the same three weights for every policy and both simulators.

Trajectory-shaping operators on a sampled plan

Closed-loop rollouts

The same scene driven twice from the same start. Left, base VaVAM executes its sampled plan (red); right, VaVAM with ECO executes the repaired plan (blue). The bird's-eye replay below each pair shows both driven paths. HUGSIM clips play at 2× real time.

HUGSIM, nuScenes scene-0038 (medium). Base VaVAM collides; VaVAM + ECO completes the route.

HUGSIM, nuScenes scene-0041 (hard). Base VaVAM leaves the route; VaVAM + ECO completes it.

AlpaSim, NuRec validation clip, nonlinear MPC at 10 Hz. The base policy drifts out of its lane and collides with a car ahead at 11.0 s; with ECO the vehicle holds its lane.

Results

Every evaluated policy improves in aggregate closed-loop score in both simulators. HD-Score on the 185-scene HUGSIM test set (mean over passes) and scene score on the 441-scene AlpaSim NuRec validation set; scores in percent.

HUGSIM, 185 scenesBase+ ECOΔ
VaVAM18.131.0+12.9
UniAD26.832.4+5.6
VAD15.717.4+1.7
DrivoR (w/ SimScale)34.636.2+1.6
Latent TransFuser17.920.7+2.8
AlpaSim, 441 scenesBase+ ECOΔ
VaVAM, nonlinear MPC11.024.5+13.5
VaVAM, linear MPC32.839.2+6.4
DiffusionDrive, nonlinear MPC28.935.2+6.3
Results by difficulty tier

(a) Every scene of the 185-scene HUGSIM test set: HD-Score without and with ECO. (b) The gain by difficulty tier.

Gain against base-plan deficit

Paired gain against the fraction of base-plan waypoints that violate a motion limit, over policies and sampler settings.

71% of the test scenes improve, and the gain increases with how often a policy's base plans violate motion limits (Spearman ρ = 0.64 over 10 policy and sampler variants). The gain is not only comfort: VaVAM's route completion rises from 40.9 to 45.2, and AlpaSim has no comfort term at all.

BibTeX

@article{zhang2026eco,
  title   = {Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization},
  author  = {Brayden Zhang and Mahsa Golchoubian and Igor Gilitschenski and Boris Ivanovic and Kashyap Chitta},
  journal = {arXiv preprint arXiv:2609.31383},
  year    = {2026}
}