top of page

End-to-End Autonomous Driving

beyond-sight-arch-v1.png

BeyondSight: Object Permanence for End-to-End Autonomous Driving

Sandro Papais, Letian Wang, Mudit Jain, Behnaz Rezaei, Steven L. Waslander

European Conference on Computer Vision (ECCV), 2026

We introduce BeyondSight, a permanence-aware end-to-end driving framework that decouples actor existence from observability by maintaining persistent actor hypotheses over time. BeyondSight propagates actor queries temporally and updates them with observation-conditioned evidence, enabling joint perception, prediction, and planning to reason about actors even when they are temporarily unobservable. To enable principled training and evaluation of persistence-aware models, we further introduce nuScenes-Permanence, an extension of nuScenes that provides supervision and observability-conditioned evaluation for unobservable actors.

LMDrive.jpg

LMDrive: Closed-Loop End-to-End Driving with Large Language Models

Hao Shao, Yuxuan Hu, Letian Wang, Steven L. Waslander, Yu Liu, Hongsheng Li

Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), 2024

We propose a novel end-to-end, closed-loop, language-based autonomous driving framework, LMDrive, which interacts with the dynamic environment via multi-modal multi-view sensor data and natural language instructions. We also present LangAuto, a new benchmark for evaluating the autonomous agents that take language instructions as navigation inputs, which include misleading/long instructions and challenging adversarial driving scenarios.

Paper       Video       Dataset

©2026 Toronto Robotics and AI Laboratory

bottom of page