{
  "meta": {
    "generated_at": "2026-08-19T01:15:51Z",
    "window_days": 180,
    "total_papers": 1238,
    "repository_url": "https://github.com/alaliqing/AlphaAD",
    "taxonomy_version": "2.0",
    "categories": [
      {
        "name": "Perception & Sensor Fusion",
        "slug": "perception-sensor-fusion",
        "count": 197,
        "latest_published": "2026-08-17"
      },
      {
        "name": "Prediction & World Models",
        "slug": "prediction-world-models",
        "count": 86,
        "latest_published": "2026-08-18"
      },
      {
        "name": "Planning & Decision-Making",
        "slug": "planning-decision-making",
        "count": 79,
        "latest_published": "2026-08-14"
      },
      {
        "name": "Control & Vehicle Dynamics",
        "slug": "control-vehicle-dynamics",
        "count": 51,
        "latest_published": "2026-08-18"
      },
      {
        "name": "Mapping & Localization",
        "slug": "mapping-localization",
        "count": 48,
        "latest_published": "2026-08-12"
      },
      {
        "name": "End-to-End & VLA",
        "slug": "end-to-end-vla",
        "count": 108,
        "latest_published": "2026-08-16"
      },
      {
        "name": "Safety, Security & Verification",
        "slug": "safety-security-verification",
        "count": 225,
        "latest_published": "2026-08-17"
      },
      {
        "name": "Systems, Deployment & Connectivity",
        "slug": "systems-deployment-connectivity",
        "count": 152,
        "latest_published": "2026-08-17"
      },
      {
        "name": "Human Factors & Policy",
        "slug": "human-factors-policy",
        "count": 27,
        "latest_published": "2026-08-17"
      },
      {
        "name": "Cross-cutting / Other",
        "slug": "cross-cutting-other",
        "count": 265,
        "latest_published": "2026-08-18"
      }
    ],
    "tags": [
      {
        "name": "Dataset",
        "slug": "dataset",
        "count": 109
      },
      {
        "name": "Benchmark",
        "slug": "benchmark",
        "count": 88
      },
      {
        "name": "Simulation",
        "slug": "simulation",
        "count": 154
      },
      {
        "name": "Synthetic Data",
        "slug": "synthetic-data",
        "count": 40
      },
      {
        "name": "Cooperative / V2X",
        "slug": "cooperative-v2x",
        "count": 114
      },
      {
        "name": "VLA",
        "slug": "vla",
        "count": 64
      },
      {
        "name": "World Model",
        "slug": "world-model",
        "count": 80
      },
      {
        "name": "Reinforcement Learning",
        "slug": "reinforcement-learning",
        "count": 117
      },
      {
        "name": "Imitation Learning",
        "slug": "imitation-learning",
        "count": 17
      },
      {
        "name": "Hardware / Real-Time",
        "slug": "hardware-real-time",
        "count": 93
      },
      {
        "name": "Survey",
        "slug": "survey",
        "count": 17
      },
      {
        "name": "Explainability",
        "slug": "explainability",
        "count": 110
      },
      {
        "name": "Human Interaction",
        "slug": "human-interaction",
        "count": 26
      }
    ],
    "classification_confidence": {
      "high": 447,
      "medium": 357,
      "low": 434
    }
  },
  "papers": [
    {
      "id": "2608.17882",
      "anchor": "paper-2608-17882",
      "title": "ControlledShifts: Towards Standardizing Robustness Evaluation in Trajectory Prediction Under Distribution Shifts",
      "authors": [
        "Ingrid navarro",
        "Pablo Ortega-Kral",
        "Yutong Duan",
        "Jonathan Francis",
        "Jean Oh"
      ],
      "abstract": "Trajectory prediction is central to safety in autonomous driving, yet learning-based predictors tend to degrade sharply when encountering scenarios poorly represented by their training data. Many methods attempt to mitigate distribution shift degradation through data-centric or test-time adaptation approaches; however, they are typically validated along fragmented axes of generalization, leaving the field without a standardized way to compare robustness across shifts a model may encounter. To address this, we introduce ControlledShifts, a framework and benchmark suite that systematically re-splits existing trajectory datasets into in-distribution (seen) and out-of-distribution (unseen) partitions, via a shared characterization-and-splitting formulation, in which a characterization function fixes the axis of variation a benchmark probes and a splitting function fixes how the tail of that axis is withheld. The suite comprises three benchmarks targeting key topological and behavioral distribution shifts. Furthermore, to aggregate multi-dimensional performance metrics across these benchmarks, we propose a unified robustness score that evaluates models along two complementary dimensions: prediction quality (relative performance gain) and prediction stability (performance preservation under shift). We showcase ControlledShifts by benchmarking prominent transformer-based architectures, exposing critical differences in how models of varying capacities handle latent relevance and environmental structure.",
      "short_abstract": "Trajectory prediction is central to safety in autonomous driving, yet learning-based predictors tend to degrade sharply when encountering scenarios poorly represented by their training data. Many methods attempt to mitigate distribution shift degradation...",
      "published": "2026-08-18",
      "updated": "2026-08-18",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 16,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "new",
      "age_days": 1,
      "arxiv_url": "https://arxiv.org/abs/2608.17882",
      "pdf_url": "https://arxiv.org/pdf/2608.17882.pdf"
    },
    {
      "id": "2608.17779",
      "anchor": "paper-2608-17779",
      "title": "Stability Control for Real World Testing in Autonomous Racing",
      "authors": [
        "Phillip Pitschi",
        "Simon Sagmeister",
        "Frederik Werner",
        "Markus Lienkamp",
        "Boris Lohmann"
      ],
      "abstract": "Controlling an autonomous vehicle at the limits of handling is a challenging task. Due to external influences, such as road conditions or weather, a vehicle can easily become unstable. Since most control algorithms assume stable vehicle behavior, they might fail in these situations. Especially when operating expensive vehicles without a safety driver on board, as in autonomous racing, this poses a significant challenge. To enable safe operation at the vehicle's dynamic limits, we present a comprehensive stability control system that safeguards motion control algorithms in autonomous driving. The proposed system consists of an electronic stability control (ESC), a slip control (SC), and a countersteer system (CS), which collectively adapt steering and brake commands from the motion controller to maintain vehicle stability. We validate our approach through both simulation and experiments on a real-world, full-scale vehicle. The results show that the stability control system maintains vehicle stability in critical situations and extends the operational feasible region. To simplify integration, we provide an open-source implementation at github.com/TUMFTM/tam-stability-control.",
      "short_abstract": "Controlling an autonomous vehicle at the limits of handling is a challenging task. Due to external influences, such as road conditions or weather, a vehicle can easily become unstable. Since most control algorithms assume stable vehicle behavior, they might...",
      "published": "2026-08-18",
      "updated": "2026-08-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 13,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: controller",
          "Control & Vehicle Dynamics: title: control",
          "Control & Vehicle Dynamics: abstract: control",
          "Control & Vehicle Dynamics: abstract: steering or braking",
          "Safety, Security & Verification: abstract: safety",
          "Safety, Security & Verification: title: validation or testing"
        ]
      },
      "recency": "new",
      "age_days": 1,
      "arxiv_url": "https://arxiv.org/abs/2608.17779",
      "pdf_url": "https://arxiv.org/pdf/2608.17779.pdf"
    },
    {
      "id": "2608.17420",
      "anchor": "paper-2608-17420",
      "title": "SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering",
      "authors": [
        "Gen Li",
        "Shu Han",
        "Yun Xi Qiao",
        "Hua Chen",
        "Xuyang Dai",
        "Bohan Li",
        "Hao Zhao",
        "Chaojian Li"
      ],
      "abstract": "Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry structures, temporal flicker, and foreground-background misalignment. Existing refinement methods are commonly designed for a specific setting, such as image-level novel-view repair or object-editing correction. In this paper, we introduce SPVC, a structured and panoptic video fixing framework for cross-dataset driving scene rendering. The name summarizes four design principles. (1) Structured fixing denotes the use of explicit spatial conditions, including camera pose, 3D bounding boxes, and HD maps, to guide the repair process and reduce uncontrolled hallucination. (2) Panoptic fixing refers to correcting both background rendering artifacts, such as distorted roads, buildings, and lanes, and foreground vehicle artifacts introduced by scene editing, such as inconsistent object appearance. (3) Video fixing means that the model operates on driving sequences rather than isolated frames, allowing temporal cues to be used during artifact correction. (4) Cross-dataset fixing means that a single shared network is trained and applied across multiple driving datasets, reducing the need for dataset-specific or scene-specific fixers. Concretely, we construct paired degraded-clean training data by simulating under-constrained 3DGS rendering and foreground vehicle insertion artifacts, and train a two-stage controllable video diffusion model that first addresses video-level appearance and then refines scene layout with structured controls.",
      "short_abstract": "Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry...",
      "published": "2026-08-18",
      "updated": "2026-08-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Mapping & Localization: abstract: HD map",
          "Perception & Sensor Fusion: abstract: camera"
        ]
      },
      "recency": "new",
      "age_days": 1,
      "arxiv_url": "https://arxiv.org/abs/2608.17420",
      "pdf_url": "https://arxiv.org/pdf/2608.17420.pdf"
    },
    {
      "id": "2608.17567",
      "anchor": "paper-2608-17567",
      "title": "Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries",
      "authors": [
        "Henrik Wille",
        "Luis-Finley Schütz",
        "Felix Strieth-Kalthoff"
      ],
      "abstract": "Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.",
      "short_abstract": "Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we...",
      "published": "2026-08-18",
      "updated": "2026-08-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Planning & Decision-Making: abstract: decision-making",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "new",
      "age_days": 1,
      "arxiv_url": "https://arxiv.org/abs/2608.17567",
      "pdf_url": "https://arxiv.org/pdf/2608.17567.pdf"
    },
    {
      "id": "2608.17902",
      "anchor": "paper-2608-17902",
      "title": "Adaptive Model Predictive Control for Ground Vehicles: Review and Demonstrative Implementation",
      "authors": [
        "Chetana Gadgil",
        "Mahendra Singh Tomar"
      ],
      "abstract": "This paper reviews Adaptive Model Predictive Control (AMPC) methods for Autonomous Vehicles (AVs), focusing on control strategies that dynamically adapt to uncertainties and changing conditions in real-time. The critical role of Adaptive Model Predictive Control (AMPC) in addressing the challenges of autonomous vehicle control are discussed. For the scope of this paper, AMPC is defined as a class of Model Predictive Control (MPC) techniques that modify the system model, cost function, constraints, or prediction horizon, based on real-time data. Traditional MPC, while effective for constrained optimization, struggles with model inaccuracies, computational demands, and dynamic environments, necessitating AMPC methods. The review covers existing literature on Gain scheduled MPC, Online Model Estimation MPC, Weight Adaptive MPC, Horizon Adaptive MPC, Learning Based MPC, and Hybrid MPC that combines MPC with other control methods. In addition to the survey, a demonstrative simulation of an adaptive MPC controller is presented that illustrates practical aspects of weight and speed adaptation in trajectory tracking.",
      "short_abstract": "This paper reviews Adaptive Model Predictive Control (AMPC) methods for Autonomous Vehicles (AVs), focusing on control strategies that dynamically adapt to uncertainties and changing conditions in real-time. The critical role of Adaptive Model Predictive...",
      "published": "2026-08-18",
      "updated": "2026-08-18",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Survey"
      ],
      "classification": {
        "confidence": "high",
        "score": 39,
        "margin": 36,
        "evidence": [
          "title: model predictive control",
          "abstract: model predictive control",
          "abstract: vehicle control",
          "abstract: controller",
          "title: control",
          "abstract: control",
          "abstract: trajectory tracking"
        ]
      },
      "recency": "new",
      "age_days": 1,
      "arxiv_url": "https://arxiv.org/abs/2608.17902",
      "pdf_url": "https://arxiv.org/pdf/2608.17902.pdf"
    },
    {
      "id": "2608.17739",
      "anchor": "paper-2608-17739",
      "title": "Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control",
      "authors": [
        "Lu Liu",
        "Chi Xie",
        "Xi Xiong"
      ],
      "abstract": "This study investigates cooperative control of connected and automated vehicles (CAVs) at partially observable highway bottlenecks in mixed traffic, aiming to mitigate congestion without relying on complete global traffic states or online trial-and-error. We propose a physics-informed world model-based offline multi-agent reinforcement learning framework that reconstructs a physically interpretable global traffic state from local CAV observation-action histories, with coupled macroscopic-microscopic traffic dynamics providing physics-based supervision. A probabilistic ensemble world model learns traffic-state transitions and system rewards, while model disagreement quantifies epistemic uncertainty. Multi-step imagined rollouts with pessimistic rewards and uncertainty-driven truncation are then used for offline policy learning. Experiments in a SUMO-based on-ramp bottleneck using approximately $1\\times10^6$ offline transitions show that physics supervision improves state reconstruction and world-model prediction accuracy.",
      "short_abstract": "This study investigates cooperative control of connected and automated vehicles (CAVs) at partially observable highway bottlenecks in mixed traffic, aiming to mitigate congestion without relying on complete global traffic states or online trial-and-error. We...",
      "published": "2026-08-18",
      "updated": "2026-08-18",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Cooperative / V2X",
        "World Model",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 18,
        "margin": 2,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "new",
      "age_days": 1,
      "arxiv_url": "https://arxiv.org/abs/2608.17739",
      "pdf_url": "https://arxiv.org/pdf/2608.17739.pdf"
    },
    {
      "id": "2608.17178",
      "anchor": "paper-2608-17178",
      "title": "Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving",
      "authors": [
        "Christopher Lang",
        "Alexander Braun",
        "Abhinav Valada"
      ],
      "abstract": "Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes regions irrespective of their semantic or temporal relevance. In ego-centric driving videos, this can weaken the pretext signal since safety-critical cues such as pedestrians, vehicles, lane boundaries, and dynamic interactions often occupy only a small portion of the frame, yet are central to downstream perception. We introduce V-JEPA4A, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy. It accounts for semantically and temporally relevant context. The proposed policy preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction. We evaluate the resulting encoders on four driving benchmarks spanning tracking, semantic segmentation, and depth estimation. The results demonstrate that V-JEPA4A reduces identity switches on BDD100k MOT by 25% over V-JEPA with random masking, achieves 73.2 mIoU on Cityscapes, and 3.75 RMSE on KITTI-2015 depth, while incurring only ~14% additional pre-training iteration overhead.",
      "short_abstract": "Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 6,
        "evidence": [
          "abstract: perception",
          "abstract: scene segmentation",
          "abstract: depth estimation"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.17178",
      "pdf_url": "https://arxiv.org/pdf/2608.17178.pdf"
    },
    {
      "id": "2608.16490",
      "anchor": "paper-2608-16490",
      "title": "Towards Real-Time and Adaptable LiDAR Scene Completion",
      "authors": [
        "Azhar Hussian",
        "Martin Vossiek",
        "Vasileios Belagiannis"
      ],
      "abstract": "LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed in real time to be usable in downstream tasks. Existing approaches typically follow an initialize-and-refine paradigm, in which a coarse initialization of the scene is first constructed, then refined into complete 3D geometry. Generative models are slower because they iteratively refine random Gaussian noise into the scene, while non-generative methods perturb the partial scene with a fixed noise scale, which limits coverage of large gaps and occluded regions and requires manual recalibration for each new sensor configuration. We present RapidLiDAR, a LiDAR scene completion method that treats the initialization itself as a learned, data-driven component. We propose an adaptive initialization module that predicts a spatially varying displacement for each partial input point, expanding the partial observations into a coarse scene initialization adapted to the local geometry, without requiring manual noise tuning. To refine this coarse initialization into a complete and coherent scene, we additionally propose a multi-scale reconstruction module that further refines point positions by querying multi-scale 3D voxel and 2D BEV feature maps constructed from the input scan. By replacing point-neighborhood operators such as farthest point sampling and $k$-nearest neighbor search with voxel- and BEV-based feature extraction, our architecture is faster and can handle different input resolutions by design. Experiments on SemanticKITTI and KITTI-360 show that our method achieves completion performance on par with the state of the art while completing a full scene in 0.1 seconds, which is 2.3 times faster than the fastest prior method. This matches the 10 Hz acquisition rate of typical automotive LiDAR sensors, taking a step toward real-time LiDAR scene completion.",
      "short_abstract": "LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed in real time to be usable in downstream tasks. Existing approaches typically follow an initialize-and-refine paradigm, in which a coarse...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 5,
        "evidence": [
          "abstract: perception",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.16490",
      "pdf_url": "https://arxiv.org/pdf/2608.16490.pdf"
    },
    {
      "id": "2608.16354",
      "anchor": "paper-2608-16354",
      "title": "DriveCache: Action-Aware Caching for Driving World Model Inference",
      "authors": [
        "Jianchun Yang",
        "Jian Liang",
        "Xianda Guo",
        "Pinhan Fu",
        "Yanlun Peng",
        "Conglang Zhang",
        "Wenke Huang",
        "Mang Ye"
      ],
      "abstract": "Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Diffusion-based driving generators repeatedly evaluate large backbones across denoising steps, which limits generation throughput. Existing diffusion acceleration methods reduce this cost, but general-purpose designs omit driving signals available before generation, such as ego speed and planned trajectories. Experiments across driving motions show that cache tolerance varies with ego translation and rotation, denoising progress, and consecutive reuse length. We propose DriveCache, a training-free, action-aware controller that uses planned motion to allocate reuse across scenes and dynamic programming to place it across denoising steps under a calibrated response budget. A causal drift check refreshes features and replans the remaining schedule when generation departs from calibration. Across three generator configurations, DriveCache improves the overall fidelity-efficiency trade-off over evaluated cache methods. Our code will be publicly available.",
      "short_abstract": "Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Diffusion-based driving generators repeatedly evaluate large backbones across...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "title: world model"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.16354",
      "pdf_url": "https://arxiv.org/pdf/2608.16354.pdf"
    },
    {
      "id": "2608.16285",
      "anchor": "paper-2608-16285",
      "title": "Audio-Visual Segmentation via Depth-Guided Collaborative Modeling",
      "authors": [
        "Zhaojin Fu",
        "Yuyang Hong",
        "Qi Yang",
        "Zili Wang",
        "Kun Ding",
        "Shiming Xiang",
        "Bin Fan"
      ],
      "abstract": "Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues. It has broad applications in video understanding, human-computer interaction, and autonomous driving. However, most existing AVS methods do not explicitly model geometric cues such as relative distance and occlusion, thereby limiting the robustness of cross-modal alignment. In human perception, spatial structure is naturally integrated with audio-visual evidence to accurately localize sounding objects. Motivated by this, we incorporate estimated depth as a spatial structural cue for AVS and propose DGCM-AVS, a tri-modal framework that jointly models audio, visual, and depth information. Specifically, we design a Depth-Aware Dynamic Modulator to improve the separation of adjacent objects while preserving intra-object feature consistency. Furthermore, we propose Depth-Guided Progressive Fusion, which uses depth as an intermediate bridge to progressively align audio cues with visual features. Compared to state-of-the-art methods, DGCM-AVS achieves relative improvements of 10.2 percent in M_J and 8.7 percent in M_F on the AVSS dataset. We believe our study highlights depth as a promising yet underexplored modality for AVS and may encourage further research in this direction.",
      "short_abstract": "Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues. It has broad applications in video understanding, human-computer...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 6,
        "evidence": [
          "title: cooperative systems"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.16285",
      "pdf_url": "https://arxiv.org/pdf/2608.16285.pdf"
    },
    {
      "id": "2608.17009",
      "anchor": "paper-2608-17009",
      "title": "PowderLine: a programmatic powder diffraction analysis application",
      "authors": [
        "Adam A. Corrao",
        "Jennifer A. Perez",
        "John D. Langhout",
        "Megan M. Butal",
        "Thomas A. Caswell",
        "Daniel Olds"
      ],
      "abstract": "Whole-pattern fitting methods, such as Rietveld refinement, excel at extracting detailed structural, chemical, and microstructural information from powder diffraction data. Obtaining reliable results requires both considerable expertise and software-specific knowledge, and applying these methods at scale typically relies on custom scripts written for each application. High-throughput experiments and autonomous self-driving laboratories increasingly utilize powder diffraction analysis to proceed programmatically and to return structured, machine-readable results. Here, we introduce PowderLine, a Python application that encapsulates a complete refinement into a single declarative recipe, validates that recipe against a versioned schema, and executes it through refinement software to return structured results. The refinement recipe is an all-inclusive, machine-readable and -writable description of either Rietveld or single peak analysis that users, scripts, and automated agents can specify and run in the same way. As a result of PowderLine's composability, it naturally fits into interactive, scripted, and autonomous workflows alike.",
      "short_abstract": "Whole-pattern fitting methods, such as Rietveld refinement, excel at extracting detailed structural, chemical, and microstructural information from powder diffraction data. Obtaining reliable results requires both considerable expertise and software-specific...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.17009",
      "pdf_url": "https://arxiv.org/pdf/2608.17009.pdf"
    },
    {
      "id": "2608.17092",
      "anchor": "paper-2608-17092",
      "title": "Structured Driving-State Narratives for Small Language Model-Based GNSS Spoofing Detection",
      "authors": [
        "Abyad Enan",
        "Sagar Dasgupta",
        "Mizanur Rahman",
        "Mashrur Chowdhury"
      ],
      "abstract": "Autonomous vehicles (AVs) depend on reliable Global Navigation Satellite System (GNSS) positioning. However, spoofed GNSS signals can induce plausible but incorrect vehicle states. This study develops a small language model (SLM)-based framework for detecting and classifying GNSS spoofing attacks by comparing vehicle behaviors independently derived from GNSS and other sensing sources. The framework converts independent driving states from GNSS and other sensing sources into structured semantic narratives that are provided to an SLM for spoofing detection and attack classification. The performance of the SLM-based framework is compared with large language models (LLMs) fine-tuned on identical training data and evaluated on the same test set. The evaluation considers five classes: no attack, overshoot attack, stopped attack, turn-by-turn attack, and wrong-turn attack. The framework is also evaluated with geographically unseen field data collected in Clemson, South Carolina, United States. Experimental results indicate that the evaluated SLMs achieve performance similar to the LLMs, achieving an average accuracy of 96.99%, precision of 99.05%, recall of 95.59%, and F1-score of 97.18%. In terms of computational efficiency and resource utilization, the SLMs demonstrate advantages over the LLMs by requiring lower inference latency and less GPU memory during both fine-tuning and inference. Evaluation using field data collected in a geographically distinct location further demonstrated its efficacy. The presented framework can detect and classify GNSS spoofing attacks in real-time while requiring relatively low computational and memory resources, and is therefore suitable for deployment on resource-constrained vehicular computing platforms.",
      "short_abstract": "Autonomous vehicles (AVs) depend on reliable Global Navigation Satellite System (GNSS) positioning. However, spoofed GNSS signals can induce plausible but incorrect vehicle states. This study develops a small language model (SLM)-based framework for detecting...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.17092",
      "pdf_url": "https://arxiv.org/pdf/2608.17092.pdf"
    },
    {
      "id": "2608.16710",
      "anchor": "paper-2608-16710",
      "title": "The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback",
      "authors": [
        "Thomas Mbrice",
        "Ammar Ali",
        "Sami Mian",
        "Khai Hern Low",
        "Eric Chen",
        "Arshia Aghajani",
        "Wolf Schäfer",
        "Amin Shirangi"
      ],
      "abstract": "As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight. This paper details the Ethical Decision Head (EDH), a deep re- inforcement learning (RL) framework that encodes ethical reasoning as a differentiable reward signal, enabling a pol- icy gradient agent to learn morally-aligned driving behavior in scenarios whose state representation is aligned with the CARLA simulation environment [Dosovitskiy et al., 2017]. Two normative frameworks are instantiated and evaluated: a Utilitarian framework minimizing total casualties and a Kan- tian framework enforcing course maintenance as a categori- cal imperative. The EDH is trained via Proximal Policy Op- timization (PPO) [Schulman et al., 2017] against a Bradley- Terry reward model [Bradley and Terry, 1952] learned from pairwise human preference annotations over 200 collision- imminent scenarios. Results reveal an asymmetry in the learnability of normative ethical frameworks under human su- pervision. The Kantian condition, which reduces to a con- stant prediction task under the codebook, serves as a pipeline control: it confirms training stability and rules out infrastruc- ture failure as an explanation for the utilitarian result. The Utilitarian agent learned something more unsettling: human raters rewarded self-sacrifice over casualty minimization, and the model learned that preference faithfully. This divergence between what humans prescribe in theory and what they re- ward in practice suggests that RLHF does not learn ethics as philosophers define it, but as humans live it.",
      "short_abstract": "As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight. This paper details the Ethical...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 14,
        "evidence": [
          "title: ethics",
          "abstract: ethics",
          "abstract: law or policy"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.16710",
      "pdf_url": "https://arxiv.org/pdf/2608.16710.pdf"
    },
    {
      "id": "2608.16696",
      "anchor": "paper-2608-16696",
      "title": "UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures",
      "authors": [
        "Homa Esfahanizadeh",
        "Matin Mortaheb",
        "Jinfeng Du",
        "Harish Viswanathan"
      ],
      "abstract": "Physical AI systems such as autonomous vehicles and robots rely on timely exchange of high-dimensional sensory signals under tight bandwidth, latency, and energy budgets. Because the task driving downstream decisions evolves over time, a task-specific codec is brittle and retraining one per task is infeasible in the field. We propose UniTAC, a single learned image codec spanning universal (task-agnostic) to task-specialized operation, re-targeted at runtime without retraining. The task is abstracted as a per-component importance vector, derived, e.g., from gradient attribution of any downstream model, and transmitted as low-overhead side information that conditions both encoder and decoder. Trained once over a broad, randomized family of such vectors against weighted-reconstruction distortion, UniTAC keeps a fixed backbone and a single human-viewable reconstruction whose fidelity is steered to the active task by swapping the injected vector. We analyze the underlying weighted rate-distortion problem, characterizing when a diagonal weighted distortion is task-consistent and how weights relate to task sensitivity. Guided by this, we design a Vision Transformer (ViT) codec whose token-level conditioning natively realizes this weight-driven code. On a localized task at 0.034 bpp, a single UniTAC model reaches 91.4% accuracy, only 1.9% below a task-based codec (93.3%) and above universal codecs (76.9%).",
      "short_abstract": "Physical AI systems such as autonomous vehicles and robots rely on timely exchange of high-dimensional sensory signals under tight bandwidth, latency, and energy budgets. Because the task driving downstream decisions evolves over time, a task-specific codec...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.16696",
      "pdf_url": "https://arxiv.org/pdf/2608.16696.pdf"
    },
    {
      "id": "2608.17258",
      "anchor": "paper-2608-17258",
      "title": "A Hybrid End-to-End and Modular Control Architecture Toward Safe Vehicle Lateral Control: Combining Soft Actor-Critic with Model Predictive Control",
      "authors": [
        "Farzaneh Tatari"
      ],
      "abstract": "Connected and automated vehicles demand lateral controllers that are simultaneously accurate, low-effort, and safe under model error and sensor noise. Modular controllers such as model predictive control (MPC) are interpretable and constraint-aware but rely on accurate models and hand-tuned weights. End-to-end learned policies, in particular continuous-action deep reinforcement learning, are adaptable and require no hand-designed control law, but offer no intrinsic safety guarantees and limited interpretability. This paper presents a hybrid architecture that combines an end-to-end Soft Actor-Critic (SAC) policy with a constrained linear MPC into a single steering command, using the MPC's first-step optimum as the model-based anchor and a single monotone blending coefficient that interpolates between the two paradigms. The architecture is evaluated on a linearized lateral bicycle model against a PID baseline, a tuned linear MPC, and a stand-alone SAC policy, across nominal, single-axis robustness, and multi-initial-condition ensemble experiments. The hybrid retains the tracking quality of stand-alone SAC while remaining inside the MPC's actuator envelope and preserving a deterministic, model-based contribution to every steering command. The architecture provides an actuator-envelope guarantee by construction but does not establish recursive feasibility or terminal invariance, and the closed-form blend does not prevent all corner-case divergences at the boundary of the training distribution. A corner-case analysis shows that the blend attenuates but cannot prevent failure under distribution shift, motivating a connectivity-aware extension in which the blending coefficient is scheduled by vehicle-to-everything (V2X) signals to restore model-based authority. Limitations and a path toward a constrained-QP predictive safety filter are discussed.",
      "short_abstract": "Connected and automated vehicles demand lateral controllers that are simultaneously accurate, low-effort, and safe under model error and sensor noise. Modular controllers such as model predictive control (MPC) are interpretable and constraint-aware but rely...",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Cooperative / V2X",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 45,
        "margin": 33,
        "evidence": [
          "title: model predictive control",
          "abstract: model predictive control",
          "title: vehicle control",
          "abstract: controller",
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.17258",
      "pdf_url": "https://arxiv.org/pdf/2608.17258.pdf"
    },
    {
      "id": "2608.17093",
      "anchor": "paper-2608-17093",
      "title": "Digital Twin-Based Intrusion Detection for Vehicle Powertrain CAN Bus Systems",
      "authors": [
        "Araf Rahman",
        "M Sabbir Salek",
        "Mashrur Chowdhury"
      ],
      "abstract": "Existing automotive intrusion detection systems (IDSs) for the Controller Area Network (CAN) largely target discrepancies in message timing, frequency, or sequencing and cannot detect attacks that preserve these properties while manipulating the payload. Digital twins (DTs) have been used to emulate CAN traffic and generate attack scenarios for IDS evaluation, but their use for intrusion detection remains unexplored. This study develops a DT-based IDS that jointly models physical relationships among decoded powertrain signals and identifies attacks through residuals between predicted and observed behavior. A shared-encoder LSTM DT was trained on 17 decoded signals from a real Hyundai/Kia CAN log to jointly predict seven numeric and two categorical gear signals over a 24-step window. A timestep is flagged when a residual exceeds a calibrated threshold, while adaptive rollout protects the twin's input history from sustained contamination. Four attacks (plateau, continuous drift, masquerade, and gear masquerade) were evaluated against the twin and a range-and-plausibility baseline. The DT outperformed the baseline across all attacks, achieving detection rates of 94.6% for continuous drift and 89.2% for masquerade, while the baseline detected almost none of the fabricated payload attacks. These results demonstrate that learning coupled vehicle dynamics enables detection of stealthy payload manipulations that preserve normal CAN communication patterns. False positive rates reached 39.6%, highlighting the need for improved robustness under sustained attacks. The DT-based IDS shows promise for detecting stealthy payload-level CAN attacks that preserve normal communication patterns, supporting behavior-based cybersecurity for connected and automated vehicles.",
      "short_abstract": "Existing automotive intrusion detection systems (IDSs) for the Controller Area Network (CAN) largely target discrepancies in message timing, frequency, or sequencing and cannot detect attacks that preserve these properties while manipulating the payload....",
      "published": "2026-08-17",
      "updated": "2026-08-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 3,
        "evidence": [
          "abstract: security",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "new",
      "age_days": 2,
      "arxiv_url": "https://arxiv.org/abs/2608.17093",
      "pdf_url": "https://arxiv.org/pdf/2608.17093.pdf"
    },
    {
      "id": "2608.16031",
      "anchor": "paper-2608-16031",
      "title": "AdROD: HyperNetwork-based Adversarially Robust Object Detection for Autonomous Driving",
      "authors": [
        "Yuting Wu",
        "Dongfang Guo",
        "Xiangzhong Luo",
        "Qun Song",
        "Rui Tan"
      ],
      "abstract": "Camera-based object detectors are vulnerable to physical adversarial attacks designed to suppress detections. While adversarial training and input purification offer some protection, they often overfit to specific attack distributions and fail on adaptive adversaries. This paper presents AdROD, an embedded, stochastic ensemble defense software designed for autonomous driving. AdROD employs {\\em low-rank HyperNetworks}, which require only 1.6\\% of the parameter footprint of standard HyperNetworks, to generate diverse detectors at a per-frame rate, making it impractical for attackers to obtain the deployed detectors in time. To further improve adversarial robustness, AdROD incorporates a novel \\emph{functional diversity} mechanism, which couples stochastic weight updates with unique input-space transformations. We design two serving modes of AdROD that strike different trade-offs between robustness and runtime overhead: AdROD-I, a continuous protection mode for maximum resilience that leverages inter-detector disagreement to recover compromised detections, and AdROD-II, an on-demand mode triggered by kinematic discontinuities in object tracking. Through comprehensive evaluation with synthetic benchmarks, physically deployed adversarial patches, and end-to-end safety tests in the OpenCDA co-simulator, AdROD outperforms five baseline defenses and exhibits superior generalizability compared with the evaluated adversarial-training baselines, while maintaining real-time performance for safely stopping the vehicle at a stop sign instrumented with adversarial patches.",
      "short_abstract": "Camera-based object detectors are vulnerable to physical adversarial attacks designed to suppress detections. While adversarial training and input purification offer some protection, they often overfit to specific attack distributions and fail on adaptive...",
      "published": "2026-08-16",
      "updated": "2026-08-16",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 2,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat",
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "new",
      "age_days": 3,
      "arxiv_url": "https://arxiv.org/abs/2608.16031",
      "pdf_url": "https://arxiv.org/pdf/2608.16031.pdf"
    },
    {
      "id": "2608.15731",
      "anchor": "paper-2608-15731",
      "title": "Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification",
      "authors": [
        "Haadia Amjad",
        "Ronald Tetzlaff"
      ],
      "abstract": "Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also understandable to ensure user trust. In real-world computer vision tasks, these models often operate on complex images containing background noise and are heavily annotated. To make such models explainable, Concept-based Explainable AI (CXAI) methods need to be assessed for their applicability and problem-solving capacity. In this work, we explore CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset (Microsoft Common Objects in Context). We apply two CXAI methods, CRP (Concept Relevance Propagation) and CRAFT (Concept Recursive Activation FacTorization), to generate concept-level explanations and investigate the overall evaluations. Our analysis reveals three key findings: (1) CXAI highlights learning weaknesses in DNNs, (2) higher concept distinctiveness reduces label and concept confusion, and (3) environmental concepts expose dataset-induced biases. Our results demonstrate the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.",
      "short_abstract": "Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also understandable to ensure user trust. In real-world computer vision tasks, these models often operate on complex images...",
      "published": "2026-08-16",
      "updated": "2026-08-16",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: explainability",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "new",
      "age_days": 3,
      "arxiv_url": "https://arxiv.org/abs/2608.15731",
      "pdf_url": "https://arxiv.org/pdf/2608.15731.pdf"
    },
    {
      "id": "2608.15573",
      "anchor": "paper-2608-15573",
      "title": "Not All History Helps: Velocity-Aware Selective Memory for Long-Horizon End-to-End Autonomous Driving",
      "authors": [
        "Yuchen Liu",
        "Ziying Song",
        "Shengkai Zhang",
        "Jiannan Chen",
        "Peiliang Wu",
        "Lei Yang",
        "Bin Sun",
        "Yan Gong",
        "Li Wang"
      ],
      "abstract": "Reliable long-horizon planning remains a key challenge in end-to-end autonomous driving. By accounting for future motion evolution and potential consequences, it provides forward-looking guidance for safe and consistent driving in evolving traffic environments. Existing methods use historical planning states as temporal context. Self-generated history may become stale or conflict with the current motion stage, introducing unreliable priors. We propose StableDrive to address cross-cycle historical reliability and within-horizon motion-stage evolution. Selective Momentum Memory (SMM), implemented with a Mamba selective state-space operator, controls the influence of the preceding self-predicted planning state on the current cycle. Motion-Stage Training Scaffold (MSTS) uses motion-stage, long-horizon trajectory, and longitudinal-motion supervision to guide stage-aware future motion learning and is removed before inference. A fixed parameter midpoint between two architecture-aligned endpoints yields a single deployable SMM planner without model ensembling or extra inference-time computation. On nuScenes under the MomAD evaluation protocol, StableDrive achieves SOTA performance across all reported planning metrics from 1 to 6 s, reducing average collision rate by 23.3%, TPC by 30.9%, and L2 by 11.8% over the best previously reported value for each metric. On the curated Longitudinal-Transition nuScenes (LT-nuScenes), StableDrive reduces 6-s collision rate by 23.81%, TPC by 10.90%, and L2 by 6.37%. On NAVSIM v1 and v2, StableDrive achieves the highest PDMS/EPDMS in all three reported settings, including a 5.7-point EPDMS gain on v2 navhard over the previous best.",
      "short_abstract": "Reliable long-horizon planning remains a key challenge in end-to-end autonomous driving. By accounting for future motion evolution and potential consequences, it provides forward-looking guidance for safe and consistent driving in evolving traffic...",
      "published": "2026-08-16",
      "updated": "2026-08-16",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 33,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "new",
      "age_days": 3,
      "arxiv_url": "https://arxiv.org/abs/2608.15573",
      "pdf_url": "https://arxiv.org/pdf/2608.15573.pdf"
    },
    {
      "id": "2608.15539",
      "anchor": "paper-2608-15539",
      "title": "CrossView: Can Vision-Language Models Reason Across Cameras?",
      "authors": [
        "Sahil Shah",
        "S P Sharan",
        "Harsh Goel",
        "Manvik Pasula",
        "Adithya Hebbalae",
        "Minkyu Choi",
        "Sandeep P. Chinchali"
      ],
      "abstract": "Video understanding benchmarks have long centered on single-camera settings, where modern multi-modal language models achieve strong performance across image and video tasks. Yet, the real world runs on multi-camera networks: autonomous vehicles, security systems, and robots all gather data across many simultaneous views. We argue that this is not simply \"more\" of the single-camera problem; it is fundamentally different. Multi-camera reasoning requires handling context that scales with the number of views, resolving occlusions visible from only a subset of cameras, judging which views matter, and integrating evidence across perspectives that may overlap or diverge. Current models struggle with exactly these challenges, yet no benchmark systematically targets them. We introduce CrossView, a multi-camera video question-answering benchmark spanning autonomous driving, security surveillance, egocentric/exocentric video, and robotics. Evaluation of proprietary models, such as GPT-5.2, and open-source models, like Qwen3-VL, reveals consistently low accuracy, with open-source models trailing by a wide margin. Performance scales strongly with a model's ability to jointly process multiple viewpoints, positioning CrossView as a rigorous benchmark for multi-camera video. We open-source our code and dataset at https://utaustin-swarmlab.github.io/CrossView.",
      "short_abstract": "Video understanding benchmarks have long centered on single-camera settings, where modern multi-modal language models achieve strong performance across image and video tasks. Yet, the real world runs on multi-camera networks: autonomous vehicles, security...",
      "published": "2026-08-16",
      "updated": "2026-08-16",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 3,
        "evidence": [
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "new",
      "age_days": 3,
      "arxiv_url": "https://arxiv.org/abs/2608.15539",
      "pdf_url": "https://arxiv.org/pdf/2608.15539.pdf"
    },
    {
      "id": "2608.15929",
      "anchor": "paper-2608-15929",
      "title": "Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning",
      "authors": [
        "Šimon Sukup",
        "Ariyan Bighashdel",
        "Pavol Jancura"
      ],
      "abstract": "Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous studies explored different learning-task formulations for pedestrian path prediction and compared these formulations using shallow neural networks, but did not extend this analysis to more complex deep-learning models. This paper adapts the Spatial-Temporal Graph Attention Network (STGAT) to a unified pedestrian path prediction framework and introduces state and action definitions specific to STGAT. The resulting formulations support deterministic and stochastic policies, one-time and sequential decision-making, and reinforcement-learning algorithms including REINFORCE and proximal policy optimization. The proposed learning-task formulations improve prediction performance across the selected benchmark datasets compared with the standard supervised-learning formulation. These results demonstrate that reformulating the decision process and training objective can improve an advanced pedestrian trajectory prediction architecture and may provide a path toward improving other graph-based prediction models.",
      "short_abstract": "Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous studies explored different learning-task formulations for pedestrian path prediction and compared these formulations using...",
      "published": "2026-08-16",
      "updated": "2026-08-16",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Benchmark",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 9,
        "evidence": [
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "new",
      "age_days": 3,
      "arxiv_url": "https://arxiv.org/abs/2608.15929",
      "pdf_url": "https://arxiv.org/pdf/2608.15929.pdf"
    },
    {
      "id": "2608.15279",
      "anchor": "paper-2608-15279",
      "title": "Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting",
      "authors": [
        "Sitao Chen",
        "Zhuangwei Zhuang",
        "Hui Luo",
        "Qingyao Wu",
        "Mingkui Tan"
      ],
      "abstract": "4D occupancy forecasting models the spatio-temporal evolution of 3D scenes and is crucial for autonomous driving, especially for corner-case simulation. Existing methods often rely on discrete tokenization followed by autoregressive prediction, yet struggle with geometric distortion in static structures and inconsistent temporal coherence over the forecasting horizon. In this work, we propose a Geometry-Aware Spatio-Temporal context modeling method (GAST) for 4D occupancy forecasting, built upon progressive explicit-implicit generation and dual-path spatio-temporal modeling. Specifically, the generation module produces per-frame occupancy with high geometric fidelity and semantic plausibility through pose-driven warping, motion-aware feature modulation, and attention-based feature refinement. Subsequently, the spatio-temporal module enhances spatial consistency through global context aggregation while capturing scene evolution through temporal dynamics extraction. This unified design enables joint optimization of historical reconstruction and future forecasting in an end-to-end manner. Extensive experiments on Occ3D-nuScenes demonstrate the superiority of our method, outperforming the state-of-the-art by 7.67% in mIoU and 6.44% in IoU with a 2.84x speedup, while maintaining strong performance in long-term forecasting.",
      "short_abstract": "4D occupancy forecasting models the spatio-temporal evolution of 3D scenes and is crucial for autonomous driving, especially for corner-case simulation. Existing methods often rely on discrete tokenization followed by autoregressive prediction, yet struggle...",
      "published": "2026-08-15",
      "updated": "2026-08-15",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 6,
        "evidence": [
          "title: forecasting",
          "abstract: forecasting",
          "abstract: prediction"
        ]
      },
      "recency": "new",
      "age_days": 4,
      "arxiv_url": "https://arxiv.org/abs/2608.15279",
      "pdf_url": "https://arxiv.org/pdf/2608.15279.pdf"
    },
    {
      "id": "2608.15230",
      "anchor": "paper-2608-15230",
      "title": "PersonaDrive: Controllable Trajectory Prediction with Multi-Dimensional Driving Personas",
      "authors": [
        "Chan Lee",
        "Kimin Yun",
        "Yuseok Bae",
        "Seong Tae Kim",
        "Jung Uk Kim"
      ],
      "abstract": "Although recent trajectory prediction and end-to-end autonomous driving methods improve robustness in urban environments, they still lack meaningful controllability. Existing benchmarks either provide no persona-conditioned annotations or support only a single urgency spectrum (i.e., emergency, normal, relaxed), which cannot distinguish personas that share the same urgency level but require different driving dynamics. To address this, we propose (i) the Persona-Conditioned Trajectory (PCT) dataset, which decomposes driving personas along two axes, Temporal Urgency and Ride Comfort, and combines three levels of each to form a grid of nine personas, each paired with natural-language descriptions and trajectories, and (ii) PersonaDrive, a framework that can learn driving personas from language and can generate persona-specific trajectories. PersonaDrive incorporates Persona-Conditioned Anchor Transform (PCAT), which hierarchically reshapes anchors along both axes, and Persona-Conditioned Multi-Modal Fusion (PCMF) for BEV-level persona fusion. Training is supervised by a Hierarchical Guide Loss enforcing axis-aligned physical orderings and an Axis-Decomposed Diversity Loss preventing diagonal mode collapse. Experimental results show that PersonaDrive consistently improves over the compared baselines across multi-dimensional scenarios. The code and PCT dataset are available at https://github.com/VisualAIKHU/PersonaDrive",
      "short_abstract": "Although recent trajectory prediction and end-to-end autonomous driving methods improve robustness in urban environments, they still lack meaningful controllability. Existing benchmarks either provide no persona-conditioned annotations or support only a...",
      "published": "2026-08-15",
      "updated": "2026-08-15",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 19,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "new",
      "age_days": 4,
      "arxiv_url": "https://arxiv.org/abs/2608.15230",
      "pdf_url": "https://arxiv.org/pdf/2608.15230.pdf"
    },
    {
      "id": "2608.15048",
      "anchor": "paper-2608-15048",
      "title": "Beyond Overt Reactions: Analyzing Subtle User Emotional Response to Unexpected In-Vehicle System Behavior",
      "authors": [
        "Huy Quyen Ngo",
        "Suresh Kumaar Jayaraman",
        "Brian Mok",
        "Ken Friedl",
        "Oliver Krause",
        "Aaron Steinfeld",
        "Nikolas Martelaro"
      ],
      "abstract": "Modern vehicles, with advanced AI voice and autonomous navigation features, extend beyond traditional driving but, like any autonomous system, can potentially make mistakes or behave in ways unexpected by users. Although providing real-time explanations can alleviate some confusion, constant information can overwhelm users and potentially cause unnecessary distractions. Some situations may require explanations or corrective vehicle behavior, and thus, recognizing user response to unexpected vehicle behavior is critical. To investigate such user responses, our study focused on collecting and analyzing user behavioral responses to unexpected events while interacting with a fully autonomous vehicle in a driving simulator. We also aimed to address the lack of datasets capturing subtle user responses (facial, spoken language, physiological signals) to in-vehicle events, as existing datasets primarily focus on strong emotional signals in conventional human-driven cars and user response to external road and traffic conditions. Users were exposed to stimuli designed to induce surprise, confusion, and frustration while performing a secondary task on a tablet and interacting with the vehicle through voice commands and in-vehicle displays. We collected a multi-modal dataset with video, audio, and heart rate data and gained insights into subtle user responses that underscored the need for further investigation of nuanced user behaviors. These observations highlight the importance of designing vehicles that recognize and adapt to occupants' behavior, potentially improving their experience.",
      "short_abstract": "Modern vehicles, with advanced AI voice and autonomous navigation features, extend beyond traditional driving but, like any autonomous system, can potentially make mistakes or behave in ways unexpected by users. Although providing real-time explanations can...",
      "published": "2026-08-15",
      "updated": "2026-08-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Planning & Decision-Making: abstract: navigation"
        ]
      },
      "recency": "new",
      "age_days": 4,
      "arxiv_url": "https://arxiv.org/abs/2608.15048",
      "pdf_url": "https://arxiv.org/pdf/2608.15048.pdf"
    },
    {
      "id": "2608.14991",
      "anchor": "paper-2608-14991",
      "title": "Risk-Adaptive Edge--Cloud Visual Reasoning for Communication-Efficient Autonomous Driving",
      "authors": [
        "Meng Ma",
        "Shuyang Li",
        "Naigang Wang",
        "Ruimin Ke"
      ],
      "abstract": "Cloud-hosted vision-language models (VLMs) offer greater contextual reasoning capabilities than smaller onboard models, but frequent visual uploads increase communication overhead and add network and inference latency to tactical decisions. We present a risk-adaptive edge-cloud architecture in which onboard traffic assessment determines when cloud reasoning is requested. An onboard VLM and a lightweight detector capture temporal traffic conditions and path-relative hazards for conservative local response and selective cloud access. The cloud model provides tactical advice, while validation, vehicle control, and automatic emergency braking remain local. In CARLA experiments, our method matched the task success rate of periodic cloud access while reducing cloud requests by 54.1% and recording fewer automatic emergency braking (AEB) activations. In a delayed-roadwork ablation, semantic events triggered requests before the next scheduled audit. Across three emulated network profiles, the method continued to reduce cloud traffic, although lane changes took longer than with periodic access. Onboard traffic assessment therefore served as a practical trigger for selective VLM inference in these experiments.",
      "short_abstract": "Cloud-hosted vision-language models (VLMs) offer greater contextual reasoning capabilities than smaller onboard models, but frequent visual uploads increase communication overhead and add network and inference latency to tactical decisions. We present a...",
      "published": "2026-08-14",
      "updated": "2026-08-14",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 10,
        "margin": 2,
        "evidence": [
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "new",
      "age_days": 5,
      "arxiv_url": "https://arxiv.org/abs/2608.14991",
      "pdf_url": "https://arxiv.org/pdf/2608.14991.pdf"
    },
    {
      "id": "2608.14767",
      "anchor": "paper-2608-14767",
      "title": "NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving",
      "authors": [
        "Ashkan Yousefi Zadeh",
        "Zishuo Zhu",
        "Xiaomeng Li",
        "Andry Rakotonirainy",
        "Sebastien Glaser",
        "Ronald Schroeter",
        "Patricia Delhomme",
        "Zahra Mehraban"
      ],
      "abstract": "Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving.",
      "short_abstract": "Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than...",
      "published": "2026-08-14",
      "updated": "2026-08-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR"
        ]
      },
      "recency": "new",
      "age_days": 5,
      "arxiv_url": "https://arxiv.org/abs/2608.14767",
      "pdf_url": "https://arxiv.org/pdf/2608.14767.pdf"
    },
    {
      "id": "2608.14428",
      "anchor": "paper-2608-14428",
      "title": "GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure",
      "authors": [
        "Mohamed Abdelsamad",
        "Bin Yang",
        "Michael Ulrich",
        "Miao Zhang",
        "Yakov Miron",
        "Alexandru Paul Condurache",
        "Abhinav Valada"
      ],
      "abstract": "3D object detection from LiDAR point clouds is a core problem in autonomous driving. Recent advances in self-supervised learning (SSL) enable scalable pretraining and transfers well to per-point tasks such as semantic and panoptic segmentation, but transfer to 3D detection remains weaker. We analyze recent SSL methods and find that most objectives are defined only on measured LiDAR returns from visible surfaces, leaving occluded and unobserved regions unconstrained. This visible-surface bias can be sufficient for point-wise prediction, but 3D detection requires robustness to missing structure. To address this gap, we propose GhostPoint, an SSL framework that hallucinates latent features in local neighborhoods around discovered instances, generated via a novel instance voxel dilation. In GhostPoint, an encoder processes observed returns, and an additional predictor infers neighborhood representations from observed context. In addition to standard encoder-level supervision, we introduce a predictor-level supervision scheme on sampled voxels from generated neighborhoods. Specifically, observed (visible/masked) voxels match teacher-encoder targets, while unobserved voxels match teacher-predictor hallucinations. This design encourages the learned representation to explicitly model structure beyond observed returns. Extensive evaluations on nuScenes and Waymo demonstrate that our method achieves state-of-the-art performance, consistently improving downstream 3D detection, especially under sparse scans and limited labels.",
      "short_abstract": "3D object detection from LiDAR point clouds is a core problem in autonomous driving. Recent advances in self-supervised learning (SSL) enable scalable pretraining and transfers well to per-point tasks such as semantic and panoptic segmentation, but transfer...",
      "published": "2026-08-14",
      "updated": "2026-08-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 21,
        "evidence": [
          "abstract: object detection",
          "abstract: scene segmentation",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "new",
      "age_days": 5,
      "arxiv_url": "https://arxiv.org/abs/2608.14428",
      "pdf_url": "https://arxiv.org/pdf/2608.14428.pdf"
    },
    {
      "id": "2608.14394",
      "anchor": "paper-2608-14394",
      "title": "IRGNN: Efficient Invariant Radar Graph Neural Network for Radar Point Cloud Object Detection",
      "authors": [
        "Xiao Guo",
        "Wanke Xia",
        "Lili Yang",
        "Caicong Wu"
      ],
      "abstract": "Perception is a fundamental component of autonomous driving systems. While LiDAR-based methods have achieved remarkable progress in object detection, their reliability can degrade under adverse weather conditions. Radar point clouds provide a robust alternative due to their resilience to bad weather and low-illumination scenarios. However, radar point clouds are typically sparse, unordered, and less informative than LiDAR data, making it challenging to directly apply existing LiDAR-based perception methods. To address these challenges, we propose IRGNN, an Invariant Radar Graph Neural Network for radar point cloud object detection. IRGNN first reconstructs radar point clouds into graph representations using translation- and rotation-invariant feature designs, enabling robust modeling of sparse radar measurements. It then employs an improved message passing neural network (MPNN) with residual connections and a virtual node layer to enhance local feature propagation and global context modeling. Finally, task-specific heads are applied to the learned graph representations for object classification and bounding box prediction. Experimental results on the RadarScenes dataset show that IRGNN outperforms existing radar-based object detection methods and achieves competitive performance. In addition, IRGNN significantly reduces computational cost and memory usage during inference, demonstrating its effectiveness and practical potential for efficient radar-based perception in autonomous driving.",
      "short_abstract": "Perception is a fundamental component of autonomous driving systems. While LiDAR-based methods have achieved remarkable progress in object detection, their reliability can degrade under adverse weather conditions. Radar point clouds provide a robust...",
      "published": "2026-08-14",
      "updated": "2026-08-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 46,
        "margin": 38,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "title: point cloud",
          "abstract: point cloud",
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar"
        ]
      },
      "recency": "new",
      "age_days": 5,
      "arxiv_url": "https://arxiv.org/abs/2608.14394",
      "pdf_url": "https://arxiv.org/pdf/2608.14394.pdf"
    },
    {
      "id": "2608.14024",
      "anchor": "paper-2608-14024",
      "title": "SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models",
      "authors": [
        "Haojie Feng",
        "Peizhi Zhang",
        "Xinrui Zhang",
        "Zhuoren Li",
        "Junpeng Huang",
        "Xiurong Wang",
        "Dongxiao Yin",
        "Yuxiang Zhang",
        "Junfan Zhu",
        "Lu Xiong"
      ],
      "abstract": "Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.",
      "short_abstract": "Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured...",
      "published": "2026-08-14",
      "updated": "2026-08-14",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation",
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 20,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "new",
      "age_days": 5,
      "arxiv_url": "https://arxiv.org/abs/2608.14024",
      "pdf_url": "https://arxiv.org/pdf/2608.14024.pdf"
    },
    {
      "id": "2608.14407",
      "anchor": "paper-2608-14407",
      "title": "The Past and Future of AI Scientists",
      "authors": [
        "Ross D. King"
      ],
      "abstract": "We present a survey of the past and future of AI Scientists: machines capable of automating science. AI Scientists can originate hypotheses, deduce their consequences, design and execute experiments, interpret their results, and revise their beliefs. Such systems are integrated scientific agents, connected to the literature, formal knowledge, mathematical models, simulations, data-analysis systems and physical laboratories. Adam was the first machine to make novel scientific discoveries through cycles of hypothesis formation and physical experimentation. Eve established the architecture of the modern self-driving laboratory. Foundation models, autonomous agents and laboratory robotics now make it possible to build systems far more general than either Adam or Eve. The central problem is no longer whether individual components of science can be automated. They can. The problem is integration. AI Scientists must combine neural learning with logic, probability, mathematics, causal reasoning, simulation, experimental design, robotics and formal scientific records. AI Scientists have the potential to transform science: to make science faster, cheaper, more systematic and more reproducible. AI Scientists could investigate systems too complicated for unaided human science, and enable thousands of AI scientists to work together on single problems. The Nobel Turing Challenge sets the goal of developing by 2050 AI systems capable of automating Nobel-quality discoveries. Progress is ahead of schedule. When we succeed it will create a new form of science and transform the world.",
      "short_abstract": "We present a survey of the past and future of AI Scientists: machines capable of automating science. AI Scientists can originate hypotheses, deduce their consequences, design and execute experiments, interpret their results, and revise their beliefs. Such...",
      "published": "2026-08-14",
      "updated": "2026-08-14",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "new",
      "age_days": 5,
      "arxiv_url": "https://arxiv.org/abs/2608.14407",
      "pdf_url": "https://arxiv.org/pdf/2608.14407.pdf"
    },
    {
      "id": "2608.14448",
      "anchor": "paper-2608-14448",
      "title": "Control-Informed Constraint Adaptation in Minimum-Time Trajectory Planning for Autonomous Racing",
      "authors": [
        "Ann-Kathrin Schwehn",
        "Alexander Langmann",
        "Mattia Piccinini",
        "Johannes Betz"
      ],
      "abstract": "Autonomous racecars operate at the limits of vehicle dynamics, where small control errors translate into safety-critical behavior and lost performance. Trajectory planners assume perfect tracking and remain blind to execution errors. To guarantee safety, trajectory planners therefore restrict themselves to conservative spatial margins, leaving usable track space untapped. To overcome these issues, we introduce a control-informed online trajectory planning framework that learns from its own execution errors. By measuring systematic tracking deviations during runtime, we dynamically adapt spatial track constraints and iteratively expand the free-space planning area. The planner remains time-optimal while compensating for accumulated execution errors. This method was analyzed in a high-fidelity closed-loop simulation environment with autonomous racecars. The results demonstrate that our approach reduces lap time by 1.8\\,s without increasing computational burden, maintaining a median runtime of 25 ms. Our finding indicates that feeding control-induced deviations back into the planning layer unlocks performance previously inaccessible to modular architectures and enables autonomous vehicles to exploit track limits systematically.",
      "short_abstract": "Autonomous racecars operate at the limits of vehicle dynamics, where small control errors translate into safety-critical behavior and lost performance. Trajectory planners assume perfect tracking and remain blind to execution errors. To guarantee safety,...",
      "published": "2026-08-14",
      "updated": "2026-08-14",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 19,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "new",
      "age_days": 5,
      "arxiv_url": "https://arxiv.org/abs/2608.14448",
      "pdf_url": "https://arxiv.org/pdf/2608.14448.pdf"
    },
    {
      "id": "2608.13719",
      "anchor": "paper-2608-13719",
      "title": "Coverage Aware Active Evaluation for Failure Discovery with Paired Systems",
      "authors": [
        "Anjali Parashar",
        "Rachel Luo",
        "Apoorva Sharma",
        "Sushant Veer",
        "Edward Schmerling",
        "Carson Sobolewski",
        "Mingxin Yu",
        "Chuchu Fan",
        "Marco Pavone"
      ],
      "abstract": "Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult under limited testing budgets. Although cheaper proxies such as simulators, lower-fidelity systems, or related policies can be sampled extensively to find failures, proxy failures often do not transfer to the real world due to sim-to-real and system-to-system gaps. The key challenge is therefore to effectively leverage proxy system information for accurate prediction of severe target system failures. We propose an adaptive failure discovery method that combines proxy evaluations with limited target system results to guide scenario selection for target system testing. Our method learns a local predictor of target risk by correcting proxy failure signals using control-variate-inspired residual modeling. To find failures that are both likely and diverse, we combine this predictor with a support-aware mutual-information objective that favors realistic, well-supported regions while expanding coverage across failure modes. Across autonomous driving, manipulation, and quadruped velocity-tracking tasks, our method discovers up to 2$\\times$ as many failures as random sampling and active-learning baselines, including severe and diverse failures missed by competing methods.",
      "short_abstract": "Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult under limited testing budgets. Although cheaper proxies such as simulators, lower-fidelity systems, or related policies can be sampled extensively to...",
      "published": "2026-08-13",
      "updated": "2026-08-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "abstract: validation or testing",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "new",
      "age_days": 6,
      "arxiv_url": "https://arxiv.org/abs/2608.13719",
      "pdf_url": "https://arxiv.org/pdf/2608.13719.pdf"
    },
    {
      "id": "2608.13450",
      "anchor": "paper-2608-13450",
      "title": "LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles",
      "authors": [
        "Md Wasiul Haque",
        "Sagar Dasgupta",
        "Mizanur Rahman",
        "Md Rayhanur Rahman"
      ],
      "abstract": "Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming exploitability requires executable test artifacts that are difficult to construct manually. We investigate whether large language models (LLMs) can automate this process for Autoware, an open-source autonomous-driving stack. We perform compiler-precise static analysis across 185 packages, identifying 1,375 decision rules, 2,274 validation checks, and 482 input-to-safety-output flows, from which we derive a weakness taxonomy and sample 740 reachable sites. Two local open-weight LLMs, a no-static-context ablation, and a naive-template baseline generate 3,700 artifact sets, which are compiled against the real build under sanitizers, repaired through compiler-in-the-loop feedback, and fuzzed when executable. The main result is a build-integration failure taxonomy showing that 80% of first-shot compilation failures arise from dependency wiring rather than program logic. The reasoning model compiled 64% of harnesses on the first attempt, compared with 6% for the code-specialized model. Repair achieved full object-compileability for the reasoning model only through extensive stubbing; fewer than half of its harnesses reached the fuzzer, and all 37 observed crashes originated in stubbed code rather than Autoware. No candidate weakness was dynamically confirmed within budget. These results show that build integration, not candidate generation or fuzzing, is the primary barrier to reliable LLM-assisted dynamic analysis of full autonomous-vehicle software stacks.",
      "short_abstract": "Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming...",
      "published": "2026-08-13",
      "updated": "2026-08-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 21,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "new",
      "age_days": 6,
      "arxiv_url": "https://arxiv.org/abs/2608.13450",
      "pdf_url": "https://arxiv.org/pdf/2608.13450.pdf"
    },
    {
      "id": "2608.13395",
      "anchor": "paper-2608-13395",
      "title": "FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving",
      "authors": [
        "Hao Dou"
      ],
      "abstract": "Reinforcement learning improves autonomous-driving vision-language-action (VLA) models by evaluating trajectories sampled from the current policy. Group relative policy optimization (GRPO) learns from reward differences within each rollout group. When all sampled trajectories are poor, this relative signal can rank failures without identifying behavior outside the failed region. We introduce FIRE-VLA, a failure-informed self-evolution framework that converts such unresolved failures into privileged supervision for the next policy. Low-reward, low-diversity groups trigger self-distillation from a frozen round-start copy of the same model. Teacher and student have the same parameter scale, but only the teacher observes the hidden future trajectory. Supervision follows the student's generated prefix and is restricted to answer tokens, while GRPO remains active for every group. The updated policy supplies the teacher for the next round, allowing the routed failure distribution to change with the policy without requiring a larger external teacher. Starting from the same Qwen2.5-VL-3B SFT checkpoint, the comparison matches student rollout and policy-update counts. On 6,019 examples from 150 held-out nuScenes scenes, FIRE-VLA retains comparable single-sample planning, reduces G=4 mean L2 from 1.848 to 1.500 m, and lowers evaluation-persistent failure prevalence from 13.03% to 11.20%. The reduction in mean error arises mainly from rare severe rollouts rather than uniform improvement across ordinary trajectories.",
      "short_abstract": "Reinforcement learning improves autonomous-driving vision-language-action (VLA) models by evaluating trajectories sampled from the current policy. Group relative policy optimization (GRPO) learns from reward differences within each rollout group. When all...",
      "published": "2026-08-13",
      "updated": "2026-08-13",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 16,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "new",
      "age_days": 6,
      "arxiv_url": "https://arxiv.org/abs/2608.13395",
      "pdf_url": "https://arxiv.org/pdf/2608.13395.pdf"
    },
    {
      "id": "2608.13147",
      "anchor": "paper-2608-13147",
      "title": "Geometry-Grounded Unified 3D Perception for Autonomous Driving",
      "authors": [
        "Longfei Xu",
        "Xiaohui Wang",
        "Zehao Huang",
        "Han Li",
        "Ya Yang",
        "Naiyan Wang",
        "Si Liu"
      ],
      "abstract": "Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and introduce 3D geometry through downstream task-specific modules. As a result, their shared representations may fail to preserve explicit metric geometry and consistent 3D scene structure. In this paper, we present a Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes. GeoUP factorizes cross-image interaction into self, temporal, and view attention to capture structurally distinct temporal and cross-view correspondences. It further injects calibration-aware raymap encodings to provide metric scale and camera geometry. The resulting geometry-grounded latent is decoded for metric depth estimation, 3D object detection, and semantic occupancy prediction, corresponding to surface-, instance-, and volume-level readouts of the same 3D scene. Through joint multi-task and multi-dataset training, GeoUP effectively leverages heterogeneous annotations and generalizes across diverse sensor configurations and perception ranges. Extensive experiments on nuScenes, Argoverse 2, Waymo, KITTI, and DDAD demonstrate that GeoUP achieves SOTA performance across detection, occupancy, and depth estimation. These results validate the effectiveness of geometry-grounded representations for unified 3D driving perception.",
      "short_abstract": "Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and...",
      "published": "2026-08-13",
      "updated": "2026-08-13",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 21,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: object detection",
          "abstract: occupancy",
          "abstract: camera",
          "abstract: depth estimation"
        ]
      },
      "recency": "new",
      "age_days": 6,
      "arxiv_url": "https://arxiv.org/abs/2608.13147",
      "pdf_url": "https://arxiv.org/pdf/2608.13147.pdf"
    },
    {
      "id": "2608.12932",
      "anchor": "paper-2608-12932",
      "title": "FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving",
      "authors": [
        "Zekai Li",
        "Yihao Liang",
        "Hongfei Zhang",
        "Jian Chen",
        "Yesheng Liang",
        "Zhijian Liu"
      ],
      "abstract": "Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.",
      "short_abstract": "Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade...",
      "published": "2026-08-13",
      "updated": "2026-08-13",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 28,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "new",
      "age_days": 6,
      "arxiv_url": "https://arxiv.org/abs/2608.12932",
      "pdf_url": "https://arxiv.org/pdf/2608.12932.pdf"
    },
    {
      "id": "2608.12854",
      "anchor": "paper-2608-12854",
      "title": "BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving",
      "authors": [
        "Bing Zhan",
        "Shuyao Shang",
        "Jiahao Gu",
        "Shuo Lu",
        "Yuan Xu",
        "Zhao Wang",
        "Yida Wang",
        "Xueyang Zhang",
        "Kun Zhan",
        "Lue Fan",
        "Zhaoxiang Zhang"
      ],
      "abstract": "Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find that a naive combination through joint token-level attention suffers from an attention-allocation mismatch, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics. Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. We further introduce an asynchronous rectified-flow inference strategy with decoupled video and action denoising, which shortens inference latency while preserving planning-relevant predictive context. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.",
      "short_abstract": "Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for...",
      "published": "2026-08-13",
      "updated": "2026-08-13",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "VLA",
        "World Model",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 18,
        "margin": 3,
        "evidence": [
          "abstract: world model",
          "title: predictive dynamics",
          "abstract: predictive dynamics",
          "abstract: prediction"
        ]
      },
      "recency": "new",
      "age_days": 6,
      "arxiv_url": "https://arxiv.org/abs/2608.12854",
      "pdf_url": "https://arxiv.org/pdf/2608.12854.pdf"
    },
    {
      "id": "2608.13878",
      "anchor": "paper-2608-13878",
      "title": "Knowledge-Data-Dual-Driven Reinforcement Learning for Autonomous Vehicle Control in Mixed Traffic",
      "authors": [
        "Jie Fang",
        "Wei Zheng",
        "Mengyun Xu",
        "Eui-Jin Kim"
      ],
      "abstract": "In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, physics-based priors incorporated into reinforcement learning (RL) models fail to capture latent interactive vehicle intentions and diverse driver behaviors, limiting the proactive reasoning capabilities. Second, abrupt maneuvers by surrounding vehicles cause non-stationarity, leaving long-tail safety events under-explored. Third, hybrid action spaces destabilize unified RL training due to the different temporal scales of continuous car-following and discrete lane-changing maneuvers. To address these issues, we propose Knowledge-Data Dual-driven Reinforcement Learning (KDDRL). First, a conditional deep generative model synthesizes intention-aware future trajectories, converting passive perception into proactive predictive states. Second, a knowledge-data dual-driven paradigm operates on these predictive states, fusing probabilistic data-driven insights with physical constraints to guide safe exploration through safety-critical scenarios. Third, a coupling module compresses both intention-aware trajectories and physical constraints into compact shared embeddings. This unified representation enables asynchronous multi-timescale optimization of continuous car-following and discrete lane-changing while preserving mutual information. Evaluations on dataset-calibrated simulations demonstrate that KDDRL effectively handles intention uncertainty, accelerates training convergence, and outperforms conventional baseline methods in terms of safety, efficiency, and comfort.",
      "short_abstract": "In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, physics-based priors incorporated into reinforcement learning (RL) models fail to capture latent interactive vehicle intentions and diverse driver...",
      "published": "2026-08-13",
      "updated": "2026-08-13",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 12,
        "evidence": [
          "title: vehicle control",
          "title: control"
        ]
      },
      "recency": "new",
      "age_days": 6,
      "arxiv_url": "https://arxiv.org/abs/2608.13878",
      "pdf_url": "https://arxiv.org/pdf/2608.13878.pdf"
    },
    {
      "id": "2608.14725",
      "anchor": "paper-2608-14725",
      "title": "Spatial Attention Noise Masking for Causally Sufficient Interpretability",
      "authors": [
        "Benjamin Formby",
        "Kuang-Ching Wang",
        "D Hudson Smith"
      ],
      "abstract": "We present a novel causal approach to interpretability for computer vision models that dynamically masks the input image prior to classification. The interpretability of deep learning predictions is critical in high-stakes fields such as medical imaging, security, and autonomous driving. Most interpretability methods are applied passively to already trained models, which typically result in correlational rather than causal explanations. Existing causal interpretability methods are limited to post hoc analysis, weakening the causal claims. Additionally, existing active methods generally lack explanations that explicitly assign responsibility to input features. This work proposes a spatial attention noise masking framework that provides causal explanations about the features sufficient for the prediction. The proposed framework consists of: 1) a UNet-style mask generator, and 2) a Resnet18 encoder and linear classifier that classifies both masked and unmasked versions of an input image. The generated masks are regularized to be sparse and spatially smooth, while masked image embeddings are constrained to remain consistent with embeddings from the corresponding unmasked images. The resulting masks can be interpreted as feature attribution maps that are competitive with related interpretability methods while additionally providing strong causal explanations of model predictions. Quantitative evaluations demonstrate mask faithfulness, near-baseline classification performance across five classification tasks despite substantial masking of image information, and robustness to distribution shifts such as background swapping and natural adversarial examples. Qualitative comparisons further demonstrate mask behavior and competitive interpretability relative to state-of-the-art feature attribution methods.",
      "short_abstract": "We present a novel causal approach to interpretability for computer vision models that dynamically masks the input image prior to classification. The interpretability of deep learning predictions is critical in high-stakes fields such as medical imaging,...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "abstract: security",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.14725",
      "pdf_url": "https://arxiv.org/pdf/2608.14725.pdf"
    },
    {
      "id": "2608.14724",
      "anchor": "paper-2608-14724",
      "title": "Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering",
      "authors": [
        "Mohammed Abdul Al Arafat Tanzin",
        "Rudzidatul Akmam Dziyauddin"
      ],
      "abstract": "The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. However, curating high-fidelity video imagery in complex tropical urban environments---specifically Kuala Lumpur, Malaysia---presents severe challenges for Personally Identifiable Information (PII) anonymization due to high motorcycle density, dark acrylic license plates, dynamic camera tilt, and extreme tropical glare. We propose an automated anonymization framework tailored for the Kuala Lumpur Road Dataset, captured via a mobile cycling platform at 2 FPS. We document how legacy Haar cascades and YOLOv8 fail under these conditions---generating false positives on background elements while missing rotated or occluded targets. Our architecture resolves this by integrating Grounding DINO---a zero-shot open-set vision-language transformer---with a novel Spatial Vehicle Region of Interest (ROI) Containment Engine. By requiring license plate centroids to reside within validated vehicle boundaries, the pipeline suppresses environmental false positives while automatically obfuscating faces, heads, and license plates. An initial evaluation on 1,266 frames demonstrates a $\\sim$95\\% success rate, with remaining failures restricted to small, heavily occluded, oblique, or ambiguous targets. Coupled with temporal persistence mechanisms and an automated quality-control auditor, the framework minimizes privacy-related false negatives while preserving scene context for downstream vision tasks. While formal legal compliance depends on broader governance procedures, this publicly available pipeline and demonstration notebook provide an auditable preprocessing stage for privacy-aware dataset curation.",
      "short_abstract": "The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. However, curating high-fidelity video imagery in complex tropical urban environments---specifically Kuala Lumpur,...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.14724",
      "pdf_url": "https://arxiv.org/pdf/2608.14724.pdf"
    },
    {
      "id": "2608.12600",
      "anchor": "paper-2608-12600",
      "title": "PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping",
      "authors": [
        "Chikao Tsuchiya",
        "Dhaval Bhanderi",
        "David Ilstrup",
        "Hsinmin Cheng",
        "Christopher Ostafew"
      ],
      "abstract": "A critical challenge in deploying online HD map construction systems to real-world scenarios is the scarcity of labeled training data, which limits model generalization in diverse environments. To address this limitation, we propose a teacher-student semi-supervised learning (SSL) framework that generates high-quality pseudo-labels from unlabeled data through confidence-aware map refinement. Our approach first trains a teacher model on limited labeled data, then leverages Beta-distribution-based confidence maps to assess the reliability of predicted map elements across temporal observations. Unlike conventional filtering methods that discard entire elements, we introduce a spatial clipping technique that selectively preserves high-confidence regions while removing unreliable segments. The refined map elements serve as map priors that improve the teacher model's prediction accuracy on unlabeled data in a second pass. These enhanced predictions become pseudo-labels for training a student model from scratch, followed by fine-tuning on the original labeled data. Experimental results on the nuScenes dataset demonstrate that our teacher-student framework with refined pseudo-labels improves performance by +6.1 mAP under a low-label regime compared to training on labeled data alone, offering a practical solution to the labeled data scarcity problem in online HD map construction.",
      "short_abstract": "A critical challenge in deploying online HD map construction systems to real-world scenarios is the scarcity of labeled training data, which limits model generalization in diverse environments. To address this limitation, we propose a teacher-student...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 19,
        "evidence": [
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.12600",
      "pdf_url": "https://arxiv.org/pdf/2608.12600.pdf"
    },
    {
      "id": "2608.12051",
      "anchor": "paper-2608-12051",
      "title": "Do Not Forget the Obvious - RISC: A Risk-Informed Slice-Coverage Protocol for Safe Autonomous Driving",
      "authors": [
        "Fabian Hüger"
      ],
      "abstract": "Aggregate metrics may not fully reflect performance in insufficiently examined high-risk driving conditions. We propose RISC (Risk-Informed Slice Coverage), a practical protocol for risk-guided stress testing and coverage-qualified evaluation. Risk-guided stress testing directs a finite audit budget toward risk-relevant sub-datasets, called risk slices, while coverage-qualified evaluation reports results together with explicit statements about which slices are sufficiently or insufficiently covered. The protocol translates safety concerns into machine-readable risk slices, uses lightweight signals to tag candidate data, selects a compact audit set by risk, and qualifies the results using coverage evidence. An LLM can optionally support this process by surfacing relevant but potentially overlooked conditions during test planning, thereby helping engineers not to forget the obvious. RISC is model-agnostic and can be applied to perception modules, driving models, and other autonomous-driving subsystems. We instantiate the protocol for monocular pedestrian perception using 1,000 frames from the Zenseact Open Dataset, image statistics, and a YOLO-based detector proxy. In this proof-of-concept study, risk-guided selection increases critical failure discovery from 34.0% under random sampling to 98.5%. RISC provides a lightweight, assurance-oriented evaluation layer that complements scenario categorization, coverage assessment, and broader testing-and-verification workflows.",
      "short_abstract": "Aggregate metrics may not fully reflect performance in insufficiently examined high-risk driving conditions. We propose RISC (Risk-Informed Slice Coverage), a practical protocol for risk-guided stress testing and coverage-qualified evaluation. Risk-guided...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 11,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.12051",
      "pdf_url": "https://arxiv.org/pdf/2608.12051.pdf"
    },
    {
      "id": "2608.11790",
      "anchor": "paper-2608-11790",
      "title": "High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving",
      "authors": [
        "Muhammad Ayub Sabir",
        "Junbiao Pang",
        "Fatima Ashraf"
      ],
      "abstract": "Accurate Global Navigation Satellite System (GNSS)-based localization is essential for safe and reliable autonomous driving. However, spoofing attacks can manipulate vehicle position estimates. Continuous and subtle attacks are particularly difficult to detect because individual GNSS observations may remain plausible while the inconsistency between GNSS-implied displacement and onboard vehicle motion gradually increases. Existing methods often rely on static vehicle-behavior features or a single residual signal and do not explicitly model this evolution. To address this problem, we propose a causal high-order liquid evidence framework for GNSS spoofing detection. The method first constructs a physics-guided GNSS--motion inconsistency residual by comparing GNSS-implied displacement with onboard-motion-derived displacement. It then forms separate evidence streams for the residual level and its first- and second-order discrete variations, with relevant contextual cues selected according to the evidence order. Each stream is processed by a separate adaptive liquid encoder, and the resulting temporal states are hierarchically coupled to predict spoofing at the window endpoint using only current and past observations. Experiments on three subsets of the real-world AV-GPS dataset show that the proposed method achieves the highest F1-scores among the evaluated temporal models on Dataset~1 and Dataset~3, reaching 0.9535 and 0.9777, respectively. On Dataset~3, it detects both labeled normal-to-attack transitions within four sampling steps. Code and datasets are publicly available at: https://github.com/pangjunbiao/GNSS_Spoofing.git.",
      "short_abstract": "Accurate Global Navigation Satellite System (GNSS)-based localization is essential for safe and reliable autonomous driving. However, spoofing attacks can manipulate vehicle position estimates. Continuous and subtle attacks are particularly difficult to...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 3,
        "evidence": [
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.11790",
      "pdf_url": "https://arxiv.org/pdf/2608.11790.pdf"
    },
    {
      "id": "2608.11770",
      "anchor": "paper-2608-11770",
      "title": "Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines",
      "authors": [
        "Vaishnav Raju"
      ],
      "abstract": "Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis. Running all models on the GPU creates a serial bottleneck that limits real-time throughput as pipeline stages grow. Modern edge SoCs pair GPUs with dedicated neural accelerators (NPUs, DLAs) capable of concurrent execution, yet deploying custom models on these accelerators remains impractical due to strict operator constraints, quantization incompatibilities, and an undocumented end-to-end pipeline. We target NVIDIA Jetson DLA cores as the representative platform. We present a five-step methodology for zero GPU fallback DLA INT8 deployment of classification backbones, comprising architecture adaptation, manual dynamic range workaround to rescue TensorRT's implicit quantization (recovering 94.0% accuracy from implicit quantization's 75%) for rapid pipeline validation before explicit quantization, quantization-aware training, ONNX graph surgery for DLA compilation, and a concurrent GPU-detection/DLA-classification inference pipeline. We document nine engineering constraints with root-cause analysis and generalizable solutions. Validation on a dual-head person attribute classifier running on DLA alongside a GPU object detector on a Jetson Orin NX demonstrates near-zero pipeline overhead (12.5 vs. 13.3~FPS detector-only at 1080p), with dual-DLA scaling at no additional cost. The methodology is backbone-agnostic and generalizes to any detection-classification edge pipeline.",
      "short_abstract": "Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 12,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: hardware"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.11770",
      "pdf_url": "https://arxiv.org/pdf/2608.11770.pdf"
    },
    {
      "id": "2608.12520",
      "anchor": "paper-2608-12520",
      "title": "Weight Certificates for Convex Multi-Objective MPC: Geometric Characterization, $\\ell^1$ Construction, and $\\ell^2$ Foreclosure",
      "authors": [
        "Hadi Hajieghrary",
        "Benedikt Walter",
        "Chaitanya Shinde",
        "Miguel Hurtadoand Jerry Lopez"
      ],
      "abstract": "Automated-driving rulebooks rank rule violations lexicographically, and model predictive control enforces that ranking either exactly, through $L{+}1$ sequential programs per tick, or approximately, through a weighted sum tuned by the separation heuristic $w_1\\gg w_2\\gg\\cdots\\gg w_L$. We show the heuristic answers the wrong question. For a convex priority-ordered program, a weighted sum reproduces the lexicographic optimum precisely when its weight, augmented by a unit performance coefficient, supports the upper image of the achievement map at the lexicographic point; the admissible weights form the unit-performance slice of an outward normal cone. Under hinge penalties this slice is a polyhedron obtained by projecting a scaled-KKT system, and a linear program returns an interior weight with a certified margin; under squared-hinge penalties no finite weight is exact whenever the limiting multiplier is nonzero, with violation along the local minimizer branch decaying as $O(1/w)$. Calibrated on held-out logs, the resulting weights have near-equal tier components in nine of eleven calibration-eligible scenario classes and roughly double legal-tier event precision against a matched heuristic weight in closed-loop nuPlan experiments on a 25-rule rulebook. The certificate is, however, pointwise: no single weight is valid across the sampled ticks of an episode, the median lifetime is one sampling interval (zero subsequent ticks at the native rate), and persistence tracks active-set stability. These findings motivate monitored weighted solves with selective cascade fallback, although the compliance-pattern monitor detects only a subset of measured lapses.",
      "short_abstract": "Automated-driving rulebooks rank rule violations lexicographically, and model predictive control enforces that ranking either exactly, through $L{+}1$ sequential programs per tick, or approximately, through a weighted sum tuned by the separation heuristic...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 18,
        "evidence": [
          "title: model predictive control",
          "abstract: model predictive control",
          "abstract: control"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.12520",
      "pdf_url": "https://arxiv.org/pdf/2608.12520.pdf"
    },
    {
      "id": "2608.12198",
      "anchor": "paper-2608-12198",
      "title": "Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment",
      "authors": [
        "Jean-Pierre Busch",
        "Guido Linden",
        "Jan Bergmann",
        "Lutz Eckstein"
      ],
      "abstract": "Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of transparency can hinder explainability and trustworthiness and complicate safety assurance. Motivated by these challenges, we propose a hybrid planning architecture that combines the advantages of machine learning with the verifiability and the determinism of classical approaches. Specifically, we developed a deep neural network to interpret complex traffic scenes and propose driving behavior, while an optimization-based supervision layer validates this proposal and enforces explicit drivability and safety constraints. We evaluate the learned planner's driving behavior in open-loop studies on real-world urban data, discuss system integration aspects for stable closed-loop operation, and report results from real-world deployment on our research vehicle karl..",
      "short_abstract": "Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of transparency...",
      "published": "2026-08-12",
      "updated": "2026-08-12",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 18,
        "evidence": [
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "title: behavior planning"
        ]
      },
      "recency": "new",
      "age_days": 7,
      "arxiv_url": "https://arxiv.org/abs/2608.12198",
      "pdf_url": "https://arxiv.org/pdf/2608.12198.pdf"
    },
    {
      "id": "2608.11580",
      "anchor": "paper-2608-11580",
      "title": "RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation",
      "authors": [
        "Yueyuan Li",
        "Zexi Chen",
        "Weijie Xi",
        "Mingyang Jiang",
        "Songan Zhang",
        "Hanyang Zhuang",
        "Ming Yang"
      ],
      "abstract": "Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road networks. Existing approaches either rely on handcrafted or reconstructed real-world maps, which limits scalability, or generate only local road structures rather than complete HD maps. We present RoadWeaver, a coarse-to-fine framework for from-scratch generation of diverse, large-scale HD maps. RoadWeaver first synthesizes a global road layout, expands it into a connected road network, and then constructs lane-level geometry with topologically consistent lane connectivity. Experimental results show that RoadWeaver achieves a 99.8\\% reachability, a 10.7\\% dead-end ratio, and an endpoint alignment error of 0.24 m. Compared with SOTA generation methods, it reduces endpoint alignment error by 94.4\\% while generating complete HD maps in 1.39--3.50 s. The generated maps can be directly deployed in driving simulators, providing scalable simulation environments for future closed-loop evaluation of autonomous driving systems. The training code and an out-of-the-box implementation of RoadWeaver will be released upon acceptance.",
      "short_abstract": "Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road networks. Existing approaches either rely on handcrafted or reconstructed real-world maps, which limits scalability, or...",
      "published": "2026-08-11",
      "updated": "2026-08-11",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 18,
        "evidence": [
          "title: HD map",
          "abstract: HD map"
        ]
      },
      "recency": "recent",
      "age_days": 8,
      "arxiv_url": "https://arxiv.org/abs/2608.11580",
      "pdf_url": "https://arxiv.org/pdf/2608.11580.pdf"
    },
    {
      "id": "2608.11451",
      "anchor": "paper-2608-11451",
      "title": "Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards",
      "authors": [
        "Simón Patiño Idarraga",
        "Erick Silva",
        "Rehana Yasmin",
        "Ali Shoker"
      ],
      "abstract": "Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statistical patterns rather than the physical conditions that guarantee safe driving, leaving their decision-making process opaque and safety constraints unenforced. We introduce a neuro-symbolic safety guard, a lightweight module that attaches to the final command interface of an already-trained agent. Immediately before a command reaches the vehicle, it checks the command against explicit safety rules and, only when necessary, replaces it with the nearest safe alternative. Each intervention is directly executable and traceable to the rule that triggered it, while the guard itself requires no retraining and adds no learned component. Evaluated on the long-tail benchmarks Fail2Drive and Bench2Drive using the state-of-the-art TransFuser v6 (TFv6) as a case study, the guard improves Success Rate by 15% and reduces safety-critical collisions by up to 53%, while preserving the original Driving Score.",
      "short_abstract": "Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statistical patterns rather than the physical conditions that guarantee safe...",
      "published": "2026-08-11",
      "updated": "2026-08-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 20,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 8,
      "arxiv_url": "https://arxiv.org/abs/2608.11451",
      "pdf_url": "https://arxiv.org/pdf/2608.11451.pdf"
    },
    {
      "id": "2608.10976",
      "anchor": "paper-2608-10976",
      "title": "XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving",
      "authors": [
        "Foundation Model Team",
        "XPeng Inc"
      ],
      "abstract": "Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time control because it is open-ended, costly to decode, and difficult to optimize as an action-facing representation. We propose XCoT-VLA, which replaces descriptive rationales with compact executable CoT tokens learned from automatically constructed Reason-Action supervision. Logged trajectories provide action evidence, while scene context supplies causal semantics. The predicted XCoT sequence remains in context and conditions fixed trajectory queries through shared multimodal self-attention. Deterministic token-function routing applies the Reason FFN to XCoT tokens and the Control FFN to trajectory queries for flow-matching trajectory generation. We further introduce XCoT Policy Optimization (XCPO) as an optional refinement extension in the same executable token space. XCoT-VLA reduces longitudinal ADE from 1.645 to 1.323 on a general-distribution set and lateral FDE from 1.616 to 0.648 in lane-change scenarios. By representing driving-oriented reasoning with only 2-6 executable XCoT tokens, our method substantially reduces autoregressive reasoning overhead and remains within the real-time planning budget. These results demonstrate that driving-oriented reasoning can be compact, executable, and directly connected to trajectory generation.",
      "short_abstract": "Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time control because it is open-ended,...",
      "published": "2026-08-11",
      "updated": "2026-08-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 18,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "recent",
      "age_days": 8,
      "arxiv_url": "https://arxiv.org/abs/2608.10976",
      "pdf_url": "https://arxiv.org/pdf/2608.10976.pdf"
    },
    {
      "id": "2608.10864",
      "anchor": "paper-2608-10864",
      "title": "Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models",
      "authors": [
        "Kiet T. Nguyen",
        "Hanbo Shim",
        "Jinwoo Kim",
        "Seunghoon Hong"
      ],
      "abstract": "Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning needed for embodied AI, robotics, and autonomous driving. Prior approaches to geometry grounding either fine-tune VLMs on spatial question answering, which can perpetuate spurious visual representations, or fuse features from large geometry-grounded vision models, which substantially increases model size at inference. Knowledge distillation from geometry-grounded vision models offers an alternative, but directly matching multi-view teacher features can disrupt the pretrained alignment between visual and textual representations, degrading object- and language-semantic capabilities. We propose multi-view relational distillation (MVRD), which distills patch-wise cosine similarities across views instead of the teacher features themselves. These relations encode geometric correspondences adequate for spatial understanding, while leaving the student representation underdetermined, allowing it to remain close to its pretrained vision- language space. Across representative VLMs, MVRD improves visual-spatial reasoning, outperforming supervised fine-tuning and feature distillation while approaching feature fusion methods with considerably fewer added parameters and lower latency. We show that MVRD makes visual representations more geometric while retaining language alignment, and generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering.",
      "short_abstract": "Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning needed for embodied AI, robotics, and autonomous driving. Prior...",
      "published": "2026-08-11",
      "updated": "2026-08-11",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: scene understanding",
          "Safety, Security & Verification: abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 8,
      "arxiv_url": "https://arxiv.org/abs/2608.10864",
      "pdf_url": "https://arxiv.org/pdf/2608.10864.pdf"
    },
    {
      "id": "2608.10660",
      "anchor": "paper-2608-10660",
      "title": "Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving",
      "authors": [
        "Jiaping Wang",
        "Shaobo Li",
        "Zhen Wang"
      ],
      "abstract": "Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satellite maps, providing complementary localization cues for pipelines that depend on Global Navigation Satellite System (GNSS) signals and high-definition (HD) maps. Most existing cross-view visual localization methods process each frame independently, leaving temporal information underused and limiting accuracy under dynamic occlusion, illumination variation, and repetitive textures. This study proposes a temporal-context-enhanced framework for cross-view sequence visual localization. The proposed recurrent cross-frame module aggregates historical context from the previous state to enhance the coarse ground feature of each current frame. These enhanced features facilitate satellite candidate-region classification, while hierarchical fine-grained features enable precise local offset estimation. On the CVIS dataset, the proposed method reduces mean localization error from 3.80 m to 1.57 m and increases R@1 m from 8.14% to 40.22%. Direct transfer to KITTI-CVL achieves a mean error of 2.61 m, with target-domain fine-tuning further reducing the mean error to 2.27 m. Zero-shot field experiments on a real-world vehicle achieve a mean error of 2.84 m and R@5 m of 96.86%. These results demonstrate that temporal context enhancement significantly improves cross-view localization accuracy and supports robust deployment on public benchmarks and real-world roads.",
      "short_abstract": "Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satellite maps, providing complementary localization cues for pipelines that depend on Global Navigation Satellite System...",
      "published": "2026-08-11",
      "updated": "2026-08-11",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 19,
        "evidence": [
          "title: localization",
          "abstract: localization",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "recent",
      "age_days": 8,
      "arxiv_url": "https://arxiv.org/abs/2608.10660",
      "pdf_url": "https://arxiv.org/pdf/2608.10660.pdf"
    },
    {
      "id": "2608.11407",
      "anchor": "paper-2608-11407",
      "title": "Top-down Traffic Scenario Generation via Joint Initial-Goal Diffusion and Trajectory Infilling",
      "authors": [
        "Da Saem Lee",
        "Yash Vardhan Pant",
        "Sebastian Fischmeister"
      ],
      "abstract": "Robust traffic simulators are crucial for developing and testing autonomous vehicles to reduce the costly, labor-intensive real-world data collection process and the need for physical presence on the road. However, existing simulators require agents' initial states to generate trajectories, which limits scalability and diversity due to restrictions on the given initial states. While data-driven agent initialization has been widely studied, the generated initial states are not interpretable in terms of why the agents are initialized at those specific locations. Given known initial states, trajectory generation is also a challenging problem, as the model must learn the variability of the destination and how agents should reach it over time. In this paper, we propose TrafficDiffuser, a top-down traffic scenario generation framework that generates high-level traffic scenarios, defined by initial and goal state pairs, by jointly modeling them. The high-level scenario generation makes initial states better interpretable and reduces trajectory generation into as simple as an infilling problem. We demonstrate how the generated high-level traffic scenarios can be used, including constraining based on different trajectory modes and integrating them with existing trajectory generation models. We conduct extensive experiments on the Argoverse 2 motion prediction dataset to evaluate how well the generated outputs capture real-world distributions. In addition to generating goal states, TrafficDiffuser outperforms the next-best approach for agent initialization, reducing speed distribution distance by 55.3% and the off-road rate by 2.8%.",
      "short_abstract": "Robust traffic simulators are crucial for developing and testing autonomous vehicles to reduce the costly, labor-intensive real-world data collection process and the need for physical presence on the road. However, existing simulators require agents' initial...",
      "published": "2026-08-11",
      "updated": "2026-08-11",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 2,
        "evidence": [
          "abstract: motion forecasting",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 8,
      "arxiv_url": "https://arxiv.org/abs/2608.11407",
      "pdf_url": "https://arxiv.org/pdf/2608.11407.pdf"
    },
    {
      "id": "2608.11592",
      "anchor": "paper-2608-11592",
      "title": "Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems",
      "authors": [
        "Mengjie Tian",
        "Xinrui Zhang",
        "Tianyu Li",
        "Peizhi Zhang",
        "Guirong Zhou",
        "Haojie Feng",
        "Junpeng Huang",
        "Qixiang Zhang",
        "Lu Xiong"
      ],
      "abstract": "Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical scenarios, by enabling reproducible evaluation under controlled conditions. However, most existing approaches still rely on standardized protocols or predefined trajectories, leading to overly scripted interactions and limited ability to reproduce the natural complexity of public-road traffic. To address this limitation, we propose Video2Track, a framework that transfers real-world interactive driving scenarios from videos into steerable adversarial closed-track testing. The framework consists of two tightly coupled modules. The first is a scenario semantic mapping module, which extracts structured semantics from driving videos using a vision-language model and grounds them onto a closed-track topology library via retrieval-augmented generation, thereby identifying compatible map segments and interaction anchors. The second is a dynamic interactive testing module, which conditions on the grounded topology and anchors to generate diverse multi-agent trajectories through a conditional diffusion model, while regulating interaction intensity via a Stackelberg game with a parameterized adversarial objective. Closed-track experiments demonstrate that the proposed framework can faithfully reproduce representative real-world interaction scenarios and generate executable scenario variants with controllable risk levels and interaction styles, providing a scalable approach for realistic and steerable ADS validation.",
      "short_abstract": "Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical scenarios, by enabling reproducible evaluation under controlled conditions. However, most existing approaches...",
      "published": "2026-08-11",
      "updated": "2026-08-11",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 37,
        "margin": 33,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "title: validation or testing",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "recent",
      "age_days": 8,
      "arxiv_url": "https://arxiv.org/abs/2608.11592",
      "pdf_url": "https://arxiv.org/pdf/2608.11592.pdf"
    },
    {
      "id": "2608.10413",
      "anchor": "paper-2608-10413",
      "title": "DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving",
      "authors": [
        "Zebin Xing",
        "Yupeng Zheng",
        "Qiang Chen",
        "Linbo Wang",
        "Yichen Zhang",
        "Pengxuan Yang",
        "Junli Wang",
        "Deheng Qian",
        "Xiaoqing Ye",
        "Junyu Han",
        "Yifeng Pan",
        "Qichao Zhang",
        "Dongbin Zhao"
      ],
      "abstract": "Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and planning. However, existing approaches lack mechanisms to exploit past failures or adapt to distribution shifts, causing the model to persistently underperform on similar scenarios where it has previously failed. In this paper, we propose DriveVLA-M0, a retrieval-augmented VLA with failure-aware latent memory. We construct a latent memory pool that stores failure cases along with their structure scene representations and expert trajectory labels, and design a dedicated Retrieve Model that decouples static road structure and dynamic agent interactions to enable structurally grounded retrieval. At inference time, retrieved cases are injected into the model via a lightweight decoupled LoRA-based test-time training (TTT) mechanism, allowing targeted and scenario-specific correction without modifying the backbone. Extensive experiments on NAVSIMv1 and NAVSIMv2 benchmark demonstrate that our approach consistently outperforms prior methods, achieving 94.1 PDMS on Navtest and 47.0 EPDMS on Navhard with only 26.44 ms TTT backward latency overhead. Furthermore, we show that DriveVLA-M0 scales effectively with additional memory, enabling training-free performance gains through memory expansion. The code is available at https://github.com/ZebinX/DriveVLA-M0.",
      "short_abstract": "Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and planning. However, existing approaches lack mechanisms to exploit past failures...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 7,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.10413",
      "pdf_url": "https://arxiv.org/pdf/2608.10413.pdf"
    },
    {
      "id": "2608.10403",
      "anchor": "paper-2608-10403",
      "title": "Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning",
      "authors": [
        "Xincong Hu",
        "Lei Ou",
        "Maosen Li",
        "Jingtao Zhang",
        "Liguo Hou",
        "Zongzhang Zhang"
      ],
      "abstract": "Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challenging scenes or adversarial situations. However, these approaches typically optimize scene generation objectives separately from the evolving policy, without explicitly modeling how generated perturbations relate to the current policy's weaknesses and learning needs. In this paper, we propose Threat-guided Policy-aware Scene Perturbation (TPSP) for safe autonomous driving with online RL. TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy. Based on this representation, TPSP selectively perturbs critical objects rather than applying uniform modifications across the scene. Furthermore, we develop a threat-guided optimization strategy that evaluates perturbed scenes through threat-level differences between policy rollouts on original and perturbed scenes, guiding the generation of safety-critical scenes with higher training value. Comprehensive experiments demonstrate that TPSP improves safety learning efficiency, achieving strong safety performance on NAVSIM v2 with approximately 4 million kilometers of simulated driving data. Ablation studies verify that policy-aware targeted perturbations provide more informative safety-critical experiences than random or policy-unaware strategies, enabling safer driving under limited interaction budgets.",
      "short_abstract": "Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Synthetic Data",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 22,
        "margin": 6,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.10403",
      "pdf_url": "https://arxiv.org/pdf/2608.10403.pdf"
    },
    {
      "id": "2608.10386",
      "anchor": "paper-2608-10386",
      "title": "Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving",
      "authors": [
        "Jiazhuo Li",
        "Linjiang Cao",
        "Qi Liu",
        "Xi Xiong"
      ],
      "abstract": "Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world models reduce the reliance on costly environment interactions, policy optimization over learned dynamics remains sensitive to prediction errors. This paper proposes the Dreamer-SAC framework, which integrates a recurrent state-space world model with an off-policy soft actor-critic algorithm trained directly in latent space. The framework uses a combination of real interactions and short-horizon generated trajectories with n-step target estimation and multi-objective supervision. Evaluated in autonomous driving scenarios with objectives encompassing driving efficiency and safety, the proposed framework consistently outperforms representative reinforcement learning baselines, including DreamerV3, SAC, and PPO, while achieving improved performance with substantially fewer real environment interactions. Experiments reveal an inverted-U relationship between rollout horizon and policy performance, where short-horizon latent rollouts achieve the best trade-off between additional training signals and accumulated model bias. Furthermore, n-step target estimation demonstrates more effectiveness over one-step temporal-difference targets in exploiting predicted experience for value learning.",
      "short_abstract": "Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world models reduce the reliance on costly environment interactions, policy optimization over learned dynamics...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 18,
        "margin": 2,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.10386",
      "pdf_url": "https://arxiv.org/pdf/2608.10386.pdf"
    },
    {
      "id": "2608.10278",
      "anchor": "paper-2608-10278",
      "title": "Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models",
      "authors": [
        "Hunter Schofield",
        "Mohammed Elmahgiubi",
        "Mohammad Mahdavian",
        "Richard Shi",
        "Jinjun Shan",
        "Amir Rasouli",
        "Dongfeng Bai"
      ],
      "abstract": "Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent vision-language models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, state-of-the-art approaches typically rely on additional spatial encoders or architectural modifications during inference, increasing computational cost. We introduce Space Tokens, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules. By distilling scene-level 3D geometry and object-centric spatial attributes into continuous latent tokens, our method enables these modalities to be directly incorporated into a chain-of-thought reasoning process, thereby improving the VLM's spatial reasoning capabilities. At the same time, the learned representations can be explicitly decoded to verify that they encode meaningful geometric information, while the unified token interface remains extensible to additional modalities. Experiments on VSI-Bench improve Qwen3-VL-8B by 4.3% and SenseNova-SI-1.3 by 1.3%, while achieving state-of-the-art performance on object size (79.2%) and room size estimation (75.7%). These results demonstrate that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism for integrating geometric reasoning into large vision-language models.",
      "short_abstract": "Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent vision-language models (VLMs) have achieved impressive performance on spatial...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Planning & Decision-Making: abstract: navigation"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.10278",
      "pdf_url": "https://arxiv.org/pdf/2608.10278.pdf"
    },
    {
      "id": "2608.10107",
      "anchor": "paper-2608-10107",
      "title": "4D-WAM: 4D Consistent World Modeling for Autonomous Driving",
      "authors": [
        "Jiacheng Fu",
        "Yibo Yuan",
        "Meng Tian",
        "Yue Li",
        "Jiangtong Zhu",
        "Jianhua Han",
        "Yueyi Zhang",
        "Jianwu Fang",
        "Jianru Xue",
        "Hang Xu",
        "Zhiwei Xiong"
      ],
      "abstract": "Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D projections of the underlying 4D driving scene. Consequently, WAMs fail to understand and capture the structure of 4D scenes and thus generate visually plausible yet 4D inconsistent future predictions that mislead downstream planning. To alleviate this issue, we present 4D-WAM, a model that leverages geometric foundation models for training-time supervision to enable 4D consistent world modeling. Specifically, we feed WAM-predicted future frames into a geometric foundation model, and use 4D-aware responses to define a 4D consistency loss. This loss encourages the model to understand, represent, and predict physically consistent 4D scenes during training, without additional inference cost. Moreover, we identify an early-decision phenomenon in WAMs and propose a decision-oriented timestep sampling strategy that emphasizes supervision at early, high-noise stages, where driving decisions are primarily formed. By propagating 4D supervision to this critical decision-formation phase, the proposed strategy further improves trajectory planning. Extensive experiments demonstrate that 4D-WAM effectively models 4D consistent scene evolution and achieves state-of-the-art performance on challenging NAVSIM-v1 and NAVSIM-v2 benchmarks.",
      "short_abstract": "Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "abstract: motion planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.10107",
      "pdf_url": "https://arxiv.org/pdf/2608.10107.pdf"
    },
    {
      "id": "2608.09591",
      "anchor": "paper-2608-09591",
      "title": "FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving",
      "authors": [
        "Guolei Huang",
        "Tengfei She",
        "Yuxuan Lu",
        "Yao Huang",
        "Yuqi Ye",
        "Yongjun Shen"
      ],
      "abstract": "Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of scene-specific planning demands. Furthermore, reasoning-path optimization for higher planning quality remains largely unexplored in autonomous-driving post-training. To address these limitations, we propose FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs). We first perform large-scale driving-domain instruction tuning to establish foundational driving knowledge. Building on this foundation, we construct PCF-CoT, a chain-of-thought (CoT) dataset that grounds planning reasoning in trajectory-relevant spatial-physical evidence and organizes reasoning around scene-specific PCFs, enabling the composition and depth of reasoning paths to adapt to different planning demands. We further introduce Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which guides Monte Carlo Tree Search (MCTS) with trajectory-level planning rewards to discover reasoning paths with higher planning quality and uses the resulting responses to optimize the policy through GRPO, thereby improving trajectory planning performance. Extensive experiments on both open-loop (nuScenes) and closed-loop-oriented (NAVSIM) benchmarks demonstrate that FactorDrive achieves state-of-the-art planning performance.",
      "short_abstract": "Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 19,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.09591",
      "pdf_url": "https://arxiv.org/pdf/2608.09591.pdf"
    },
    {
      "id": "2608.09493",
      "anchor": "paper-2608-09493",
      "title": "GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction",
      "authors": [
        "Khang Minh Le",
        "Hieu Dinh Trung Pham",
        "Luu Thanh Danh",
        "Nam-Tien Le",
        "Hieu Anh Ngo",
        "Phuong Huu Vu Tran",
        "Son Nguyen Minh Le",
        "Nguyen Trong Nghia",
        "Tu Tran Thi Cam",
        "Huy Minh Nhat Nguyen",
        "Cuong Tuan Nguyen"
      ],
      "abstract": "Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.",
      "short_abstract": "Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 22,
        "evidence": [
          "title: future-frame prediction",
          "abstract: future-frame prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.09493",
      "pdf_url": "https://arxiv.org/pdf/2608.09493.pdf"
    },
    {
      "id": "2608.09402",
      "anchor": "paper-2608-09402",
      "title": "Beyond the Plane: Coupling Planar Vehicle Dynamics with Three-Dimensional Road Geometry",
      "authors": [
        "Simon Sagmeister",
        "Phillip Pitschi",
        "Nico Haja",
        "Markus Lienkamp"
      ],
      "abstract": "Simulation is crucial for developing and testing autonomous driving systems. In particular, the development of localization and control algorithms relies on an accurate vehicle dynamics simulation. However, most vehicle dynamics models are two-dimensional while real-world roads are three-dimensional. For example, effects from the three-dimensional road geometry on the Las Vegas Motor Speedway can increase the normal forces on the tires by more than 66% compared to the nominal load at standstill. As a result, even highly detailed planar vehicle dynamics models struggle to accurately reproduce the real vehicle's behavior. While solutions for three-dimensional vehicle dynamics exist, they are rarely adopted, computationally expensive, and complex. To address this issue, we present a novel method to couple planar vehicle dynamics models with real-world three-dimensional road geometry. We transform the planar vehicle state from the vehicle model's two-dimensional plane to its corresponding representation in three-dimensional space. Additionally, we calculate road-geometry-induced forces and moments and apply them to the planar vehicle model. We validate our approach using high-speed data recorded with a full-scale race car on the banked Las Vegas Motor Speedway. Furthermore, on synthetic tracks, we show that our method yields accurate results even in edge cases. Together, our results demonstrate that the gap between planar simulation and real-world three-dimensional roads can be closed without abandoning simpler planar models. To simplify adoption of our method, we provide the implementation as open-source software on github.com/TUMFTM/3d-road-geometry-coupling.",
      "short_abstract": "Simulation is crucial for developing and testing autonomous driving systems. In particular, the development of localization and control algorithms relies on an accurate vehicle dynamics simulation. However, most vehicle dynamics models are two-dimensional...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 17,
        "evidence": [
          "title: vehicle dynamics",
          "abstract: vehicle dynamics",
          "abstract: control"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.09402",
      "pdf_url": "https://arxiv.org/pdf/2608.09402.pdf"
    },
    {
      "id": "2608.09333",
      "anchor": "paper-2608-09333",
      "title": "DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving",
      "authors": [
        "Ziyi Song",
        "Chen Xia",
        "Hang Yu",
        "Sheng Zhou",
        "Zhisheng Niu"
      ],
      "abstract": "Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated vehicles, limited sensing range and occlusions restrict reliable decision-making, and the substantial computational and latency overhead makes on-board deployment impractical. Cooperative driving provides a potential solution by leveraging external agents for information exchange, but existing methods remain limited in semantic reasoning capability under practical constraints. To address these challenges, we propose DH-VLM, a dual-horizon cooperative latent reasoning framework that enables asymmetric semantic cooperation between the infrastructure and ego vehicle. The infrastructure aggregates multi-layer hidden states to form a global-reasoning horizon latent guidance, which is integrated into the ego model through an Infrastructure-Driven Latent Evolution mechanism for conditional latent refinement. This enables the ego vehicle to leverage long-range contextual understanding while preserving autonomous decision-making within its local planning horizon. Furthermore, we construct a cooperation-oriented question-answer (QA) dataset covering fundamental scene understanding and ego-personalized comprehension to support counterfactual and safety-aware reasoning. Extensive experiments demonstrate that DH-VLM achieves state-of-the-art planning performance, outperforming the previous state of the art by 14.6% in L2 error and 26.9% in collision rate. Compared with query-based end-to-end cooperative driving methods, our approach reduces the communication cost by 57.3% and GPU memory usage by 25.5%, while maintaining strong robustness against infrastructure guidance errors, providing a practical and robust paradigm for cooperative autonomous driving.",
      "short_abstract": "Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated vehicles, limited sensing range and occlusions restrict reliable decision-making, and the substantial...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Dataset",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 12,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.09333",
      "pdf_url": "https://arxiv.org/pdf/2608.09333.pdf"
    },
    {
      "id": "2608.09202",
      "anchor": "paper-2608-09202",
      "title": "CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving",
      "authors": [
        "Junyao Wang",
        "Yulin Xu",
        "Yu Li",
        "Pramod Khargonekar",
        "Mohammad Abdullah Al Faruque"
      ],
      "abstract": "Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor varies significantly across diverse real-world driving conditions, including poor visibility and adverse weather. While uncertainty quantification (UQ) mitigates this issue by allowing models to prioritize reliable signals, existing uncertainty-aware fusion methods typically rely on simple feature-level uncertainty estimates and thus often fail to generalize effectively in complex, out-of-distribution scenarios. To address this limitation, we propose CRUISE, a novel uncertainty-aware cross-modal sensor fusion framework. CRUISE integrates a vision-language model (VLM)-guided UQ module that generates fine-grained, pixel-level uncertainty estimates. By leveraging the VLM's rich prior knowledge and superior contextual reasoning, our approach provides a highly informative guide for the fusion process. Furthermore, we introduce a dynamic adaptive mechanism that explicitly models and captures cross-modal dependencies, ensuring the framework fully exploits the inherent complementary nature of multi-sensor inputs.",
      "short_abstract": "Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor varies...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 11,
        "evidence": [
          "abstract: perception",
          "title: sensor fusion",
          "abstract: sensor fusion",
          "abstract: LiDAR",
          "abstract: radar",
          "abstract: camera"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.09202",
      "pdf_url": "https://arxiv.org/pdf/2608.09202.pdf"
    },
    {
      "id": "2608.09144",
      "anchor": "paper-2608-09144",
      "title": "How Roadside Units Enhance Intersection Safety? Cooperative Autonomous Driving System Design and A Proof of Concept",
      "authors": [
        "Taoyuan Yu",
        "Kui Wang",
        "Zongdian Li",
        "Tao Yu",
        "Walid Saad",
        "Kei Sakaguchi"
      ],
      "abstract": "Intersections remain one of the most hazardous locations in urban road networks, where heterogeneous traffic participants and limited visibility frequently lead to severe traffic conflicts. In this paper, a vehicle-to-infrastructure-to-vehicle (V2I2V) cooperative system is proposed for improving road safety and traffic efficiency by using digital twins (DTs) deployed on roadside units (RSUs) to eliminate blind spots and centrally coordinate connected and automated vehicles (CAVs) in smart intersections. The proposed system integrates cloud-based global DTs for macroscopic guidance and RSU-based local DTs for real-time operations. Within this architecture, a hierarchical reinforcement learning (HRL) framework combines offline pre-training with online fine-tuning to achieve robust cooperative control. Experimental results show that the proposed system achieves substantial improvements in safety and efficiency in simulation experiments and real-world proof-of-concept (PoC) trials. In simulations, our system ensures high safety, efficiency, and smoothness under realistic communications and traffic constraints. In PoC trials, the RSU-centric control loop achieves a decision-making latency of approximately 42 ms and maintains a safe stopping distance of 8.5 m for pedestrians, while also shortening stop duration and overall traversal time. These results indicate that the proposed system provides robust and scalable performance at smart intersections.",
      "short_abstract": "Intersections remain one of the most hazardous locations in urban road networks, where heterogeneous traffic participants and limited visibility frequently lead to severe traffic conflicts. In this paper, a vehicle-to-infrastructure-to-vehicle (V2I2V)...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 23,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints",
          "title: roadside infrastructure"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.09144",
      "pdf_url": "https://arxiv.org/pdf/2608.09144.pdf"
    },
    {
      "id": "2608.09541",
      "anchor": "paper-2608-09541",
      "title": "Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives",
      "authors": [
        "Lei Wan",
        "Hannan Ejaz Keen",
        "Alexey Vinel"
      ],
      "abstract": "Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions. We present a conceptual framework for Collaborative Joint Perception and Prediction (Co-P&P) that improves motion prediction of surrounding road users, thereby enhancing situational awareness in complex and dynamic traffic environments. Building upon our preliminary study, this extended version compares the performance of different fusion strategies and establishes baseline performance for a modular design of perception and prediction. Experimental results show that prediction-level fusion leads to a decline in overall system performance compared to detection-level or tracking-level fusion. We further implement a minimal end-to-end Co-P&P prototype that couples collaborative point-cloud sharing via the RENO neural codec with joint detection-forecasting via FutureDet, showing that collaboration improves forecasting accuracy while neural compression preserves this benefit at roughly 34x lower communication bandwidth.",
      "short_abstract": "Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work...",
      "published": "2026-08-10",
      "updated": "2026-08-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 19,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "title: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 9,
      "arxiv_url": "https://arxiv.org/abs/2608.09541",
      "pdf_url": "https://arxiv.org/pdf/2608.09541.pdf"
    },
    {
      "id": "2608.09098",
      "anchor": "paper-2608-09098",
      "title": "UnsDrive: Towards Robust End-to-End Autonomous Driving in Unstructured Scenes",
      "authors": [
        "Nanxin Zeng",
        "Ruiqi Song",
        "Xiangyu Guo",
        "Baiyong Ding",
        "Yunfeng Ai"
      ],
      "abstract": "End-to-end planning has shown strong promise for autonomous driving, but most existing methods are designed for structured urban roads and generalize poorly to unstructured mining environments. In such settings, weak road structure, terrain-induced occlusions, degraded visibility, and large unobserved regions make safe planning particularly challenging. To address these challenges, we propose UnsDrive, an end-to-end planner designed for unstructured mining scenes. UnsDrive builds an unknown-aware occupancy representation that explicitly models occupied, free, and unknown space using multi-frame visibility cues, and conditions a flow-matching planner on this representation to generate multimodal future trajectories. To improve safety under partial observability, we further introduce an occupancy trajectory consistency loss and an uncertainty-aware trajectory scorer that penalize trajectories entering non-traversable or unobserved regions. We also present MineLoop, a mining-oriented closed-loop simulator for evaluating autonomous driving under irregular road geometry, degraded visibility, heavy-vehicle interactions, and mining-specific operational constraints. Experiments in both open-loop and closed-loop settings show that UnsDrive consistently outperforms strong baselines in trajectory accuracy, collision avoidance, and long-horizon driving robustness. These results demonstrate the value of explicit unknown-space reasoning for autonomous driving in unstructured mining environments.",
      "short_abstract": "End-to-end planning has shown strong promise for autonomous driving, but most existing methods are designed for structured urban roads and generalize poorly to unstructured mining environments. In such settings, weak road structure, terrain-induced...",
      "published": "2026-08-09",
      "updated": "2026-08-09",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 16,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 10,
      "arxiv_url": "https://arxiv.org/abs/2608.09098",
      "pdf_url": "https://arxiv.org/pdf/2608.09098.pdf"
    },
    {
      "id": "2608.08947",
      "anchor": "paper-2608-08947",
      "title": "Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis",
      "authors": [
        "Lennox Anderson",
        "Ahmed Boutar",
        "Jonah Mulcrone",
        "Tal Erez"
      ],
      "abstract": "Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition. We investigate whether human gaze patterns, captured via webcam-based eye tracking (WebGazer.js), can serve as privileged information to constrain mesa-objective formation. We collected 137,663 frame-level gaze samples synchronized with hazard annotations across 388 real dashcam clips, then test this hypothesis across two calibration protocols (9-point/45-click and 11-point/440-click), two model architectures (Random Forest and causal Transformer), and five random seeds per experiment with paired t-tests. No experiment yields a statistically significant improvement from gaze (p = 0.919, 0.578, and 0.667 respectively). A geometric analysis reveals the root cause: WebGazer's reported error (~130-257 px depending on configuration) exceeds 93% of detected hazard object sizes (median 36 px), rendering object-level gaze attribution physically impossible at this instrument precision.",
      "short_abstract": "Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition. We investigate whether human gaze patterns,...",
      "published": "2026-08-09",
      "updated": "2026-08-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "recent",
      "age_days": 10,
      "arxiv_url": "https://arxiv.org/abs/2608.08947",
      "pdf_url": "https://arxiv.org/pdf/2608.08947.pdf"
    },
    {
      "id": "2608.08941",
      "anchor": "paper-2608-08941",
      "title": "From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving",
      "authors": [
        "Chaitanya Shinde",
        "Hadi Hajieghrary",
        "Miguel Hurtado"
      ],
      "abstract": "Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain. This gap between operating condition specification and behavioral validation represents a critical unresolved challenge in ADS safety assurance. This paper presents a structured, standards-grounded taxonomy of 21 behavioral competencies organized across three operational domains-Highway (HWY), Urban (URB), and Hub (HUB)-derived systematically from the PEGASUS six-layer model-based ODD. Each behavior is decomposed along longitudinal and lateral control axes and characterized against a four-property framework: Safety (gap maintenance, conflict avoidance, kinematic stability), Compliance (legal rules and behavioral norms), Comfort (rider dynamics and trust), and Efficiency (mission completion and product-level metrics). We further demonstrate that the crossing of ODD layer parameterizations with behavioral competency specifications yields concrete scenario families suitable for systematic behavioral testing and SOTIF coverage evidence. The taxonomy is grounded in AVSC00008202111, SAE J3237, and SAE J3016, and is validated as an operational specification layer through its deployment in a rule-enforced trajectory optimization system. The Hub domain is identified as a structurally distinct, underspecified domain warranting dedicated research attention.",
      "short_abstract": "Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain. This gap between operating condition...",
      "published": "2026-08-09",
      "updated": "2026-08-09",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: SOTIF"
        ]
      },
      "recency": "recent",
      "age_days": 10,
      "arxiv_url": "https://arxiv.org/abs/2608.08941",
      "pdf_url": "https://arxiv.org/pdf/2608.08941.pdf"
    },
    {
      "id": "2608.08476",
      "anchor": "paper-2608-08476",
      "title": "RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion",
      "authors": [
        "Meng Wang",
        "Hongxia Yu",
        "Wenzhe He",
        "Xingdong Song",
        "Huilong Pi",
        "Jiapeng Zhang",
        "Ruihui Li"
      ],
      "abstract": "Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel representations. To address this issue, we propose RayLift, a framework that uses stereo geometry as a metric reference while incorporating complementary ray evidence to recover reliable 3D structures adaptively. RayLift first employs a Complementary Context Encoder that extracts geometry-aware priors from a frozen 3D vision foundation model, thereby enriching the scene context. It then introduces a Depth Ray Evidence Lifter module that jointly models geometric dissimilarity, depth confidence, and spatial uncertainty to adaptively sample and weight candidate surface locations along each camera ray. Finally, a Semantic-Aware Voxel Integrator injects the resulting ray evidence into voxel features by explicitly modeling their spatial support. Extensive experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that RayLift achieves competitive performance and consistently outperforms existing methods.",
      "short_abstract": "Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and...",
      "published": "2026-08-09",
      "updated": "2026-08-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera",
          "Perception & Sensor Fusion: abstract: scene understanding",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "recent",
      "age_days": 10,
      "arxiv_url": "https://arxiv.org/abs/2608.08476",
      "pdf_url": "https://arxiv.org/pdf/2608.08476.pdf"
    },
    {
      "id": "2608.10025",
      "anchor": "paper-2608-10025",
      "title": "The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software",
      "authors": [
        "Kizito Salako",
        "Rabiu Tsoho Muhammad"
      ],
      "abstract": "For safety-critical software, data from the software's operational past (e.g. a sequence of success and failure events experienced by the software) can provide strong statistical support for reliability claims about the software. However, such data might not describe past software failure events in sufficient detail, and this might leave a reliability assessment (based on this data) unable to account for important features of past software failures. In this paper, by extending conservative Bayesian inference (CBI) techniques used in reliability assessment, we illustrate a principled statistical approach for checking the robustness of reliability claims derived from insufficiently detailed operational data. We demonstrate the extent to which insufficient detail in operational data can undermine software reliability claims in autonomous vehicle (AV) safety assessment scenarios. Reliability claims derived from insufficiently fine-grained data might be dangerously optimistic, despite a concerted effort by an assessor to use such data conservatively during the assessment. While these findings are consistent with previous work on the impact of statistical model fidelity in Bayesian software reliability assessments, our work clarifies why attempts to use low-fidelity data conservatively can be naive, and we give the first conservative estimates of the impact of data fidelity on assessments.",
      "short_abstract": "For safety-critical software, data from the software's operational past (e.g. a sequence of success and failure events experienced by the software) can provide strong statistical support for reliability claims about the software. However, such data might not...",
      "published": "2026-08-09",
      "updated": "2026-08-09",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 20,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 10,
      "arxiv_url": "https://arxiv.org/abs/2608.10025",
      "pdf_url": "https://arxiv.org/pdf/2608.10025.pdf"
    },
    {
      "id": "2608.08815",
      "anchor": "paper-2608-08815",
      "title": "Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles",
      "authors": [
        "Pedram MohajerAnsari",
        "Amir Salarpour",
        "Mert D. Pesé"
      ],
      "abstract": "Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches. Existing defenses often improve robustness against one attack type while degrading performance on others, and can reduce clean accuracy. We propose LAMDA (Language-Anchored Model for Direction Alignment), a training framework that transfers language-grounded structure into TSR models without using adversarial examples or adding inference-time overhead. LAMDA builds two fixed prototype banks from VLM-generated sign descriptions and class names using a frozen OpenCLIP text encoder, and uses them to supervise visual features through two complementary auxiliary losses during training. At inference, the adapter and prototype banks are discarded, leaving a standard backbone and classifier. Evaluated on GTSRB and LISA across four backbones and three physical attack types, LAMDA is the only method among ten evaluated that consistently improves robustness across all attack-backbone-dataset combinations, with gains of up to +12.5 pp under shadow attacks and +13.2 pp under natural-light attacks, while preserving or improving clean accuracy in nearly all cases.",
      "short_abstract": "Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches....",
      "published": "2026-08-09",
      "updated": "2026-08-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: title: perception",
          "Perception & Sensor Fusion: abstract: traffic sign recognition",
          "Safety, Security & Verification: abstract: attack or threat",
          "Safety, Security & Verification: title: robustness",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 10,
      "arxiv_url": "https://arxiv.org/abs/2608.08815",
      "pdf_url": "https://arxiv.org/pdf/2608.08815.pdf"
    },
    {
      "id": "2608.08701",
      "anchor": "paper-2608-08701",
      "title": "Anchor-Based AI Approach for Pre-Crash Object Detection Utilizing Micro-Doppler Signatures in Automotive Radar",
      "authors": [
        "Patrick Zaumseil",
        "Rainer Engert",
        "Dagmar Steinhauser",
        "Jonathan Wache",
        "Soumya Dewangan",
        "Thomas Brandmeier"
      ],
      "abstract": "Advanced automated driving presents significant potential to improve modern automotive safety systems, but it depends highly on the reliable activation of restraint systems. Forward-looking sensors are crucial for immediate and precise object detection. Recent developments in automotive radar technology enable detailed environment detection and the recognition of high-resolution features, such as micro-Doppler signatures. Combined with advanced AI techniques, these features significantly enhance object detection and improve the accuracy of kinematic parameter estimation. This is essential for the early and reliable activation of irreversible safety systems, such as smart airbags and adaptive seat belts. Therefore, an anchor-based AI model is presented, designed to process high-resolution radar data with an explicit focus on micro-Doppler signatures to improve pre-crash object detection. Furthermore, these signatures can improve the accuracy of kinematic object parameter estimation and reduce false negatives, especially in the critical near-field. To address the challenges of sparse and fluctuating radar point clouds, an innovative radar-image dilation technique on the feature input channels was developed to amplify local radar patterns, like micro-Doppler features. Therefore, this approach increases the system's reliability and increases its ability to detect objects in pre-crash scenarios despite radar multipath reflections and ghost objects. In order to investigate the applicability and compare the model's performance with advanced automotive radar tracking methods, a radar data set using series sensors and pre-crash relevant scenarios was recorded. The results demonstrate the advantages of the anchor-based AI model over established tracking approaches. It excels at estimating object parameters in dynamic scenarios and underscores its ability to process different data sets effectively.",
      "short_abstract": "Advanced automated driving presents significant potential to improve modern automotive safety systems, but it depends highly on the reliable activation of restraint systems. Forward-looking sensors are crucial for immediate and precise object detection....",
      "published": "2026-08-09",
      "updated": "2026-08-09",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 27,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: point cloud",
          "title: radar",
          "abstract: radar"
        ]
      },
      "recency": "recent",
      "age_days": 10,
      "arxiv_url": "https://arxiv.org/abs/2608.08701",
      "pdf_url": "https://arxiv.org/pdf/2608.08701.pdf"
    },
    {
      "id": "2608.08293",
      "anchor": "paper-2608-08293",
      "title": "Scalable High-Speed Lateral Control for Single-Body and Articulated Autonomous Vehicles",
      "authors": [
        "Aashish Shaju",
        "Steve Southward",
        "Mehdi Ahmadian"
      ],
      "abstract": "This paper presents a scalable lateral control framework for robust path tracking of single-body and articulated autonomous vehicles at high speeds. A clothoid-based controller is extended with three key adaptations: 1) an integrated tangential check and Frechet distance method for optimized lookahead and oscillation mitigation; 2) real-time trajectory segment classification for dynamic adjustment of the lookahead search range; and 3) a dual-adaptive, rate-controlled lookahead mechanism responsive to cross-track error and upcoming path geometry. To support different vehicle configurations, the framework also incorporates flexible tracking point selection and curvature-to-steering lookup tables, enabling control of points such as the tractor rear axle, hitch point, or trailer-related locations without changing the core control architecture. The controller is evaluated using high-fidelity TruckSim and Simulink simulations for both standalone and articulated vehicles across dual lane changes and winding-road scenarios. Results demonstrate stable high-speed path tracking, reduced oscillations, effective lookahead adaptation, low cross-track errors, and acceptable lateral acceleration across different vehicle configurations. The results support the use of a unified lateral control architecture that can scale from single-body to articulated autonomous vehicles.",
      "short_abstract": "This paper presents a scalable lateral control framework for robust path tracking of single-body and articulated autonomous vehicles at high speeds. A clothoid-based controller is extended with three key adaptations: 1) an integrated tangential check and...",
      "published": "2026-08-08",
      "updated": "2026-08-08",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 26,
        "evidence": [
          "title: vehicle control",
          "abstract: vehicle control",
          "abstract: controller",
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "recent",
      "age_days": 11,
      "arxiv_url": "https://arxiv.org/abs/2608.08293",
      "pdf_url": "https://arxiv.org/pdf/2608.08293.pdf"
    },
    {
      "id": "2608.08045",
      "anchor": "paper-2608-08045",
      "title": "Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities",
      "authors": [
        "Xiaohe Li",
        "Yiru Wang",
        "Junhao Fan",
        "Mingyuan Liu",
        "Jie Huang",
        "Kaixin Zhang",
        "Jiahao Li",
        "Chen Qian",
        "Zide Fan"
      ],
      "abstract": "Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Existing platforms nevertheless isolate different embodiments and decouple them from task design and evaluation. We present \\textbf{Lingjing}, a simulation platform for heterogeneous multi-agent embodied intelligence in open-ended urban environments. Lingjing reconstructs and renders evolving cities from geographic data, synchronizes multiple physics engines, and exposes shared physical and structured urban state to agents. Its Gym-like interface supports user-defined ReAct agents and single- or multi-agent natural-language missions with configurable star or broadcast communication and resource constraints. Each episode becomes an attribution-ready replay that links agent trajectories and communication to relation-graph changes, resource consumption, and engine-based evaluations for systematic diagnosis. We evaluate twelve vision-language models on nine urban tasks under a shared engine-in-the-loop protocol. Controlled studies further examine communication, scalability, robustness, and failure provenance. Results expose persistent bottlenecks in grounding and long-horizon execution. They also show task-dependent coordination trade-offs and diminishing returns from added capacity, while heavier workloads further reduce success. Lingjing provides a unified testbed that enables reproducible end-to-end evaluation and systematic failure diagnosis in urban multi-agent embodied intelligence.",
      "short_abstract": "Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Existing...",
      "published": "2026-08-08",
      "updated": "2026-08-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness",
          "Safety, Security & Verification: abstract: failure",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 11,
      "arxiv_url": "https://arxiv.org/abs/2608.08045",
      "pdf_url": "https://arxiv.org/pdf/2608.08045.pdf"
    },
    {
      "id": "2608.08035",
      "anchor": "paper-2608-08035",
      "title": "Decentralized Multi-Agent Urban Traffic Management via Spatio-Temporal Mobility Profile Planning",
      "authors": [
        "Lorenzo Mario Amorosa",
        "Lorenzo Farina",
        "Vittorio Todisco",
        "Alessandro Bazzi"
      ],
      "abstract": "As modern cities face increasingly severe traffic congestion, connected and autonomous vehicles (CAVs) have emerged as a crucial enabling technology for next-generation intelligent traffic management. However, fully realizing this potential is hindered by the limitations of current paradigms. Existing approaches typically optimize localized interactions rather than system-wide efficiency, incur severe communication overhead, or lack the deterministic guarantees required for safe kinematic execution. Furthermore, current multi-agent adaptations are frequently restricted to small predefined scenarios, failing to scale across large and complex urban networks. To bridge this gap, this paper introduces VeloCity, a decentralized multi-agent spatio-temporal mobility profile planning framework designed for CAVs operating in arbitrary urban areas. To minimize vehicles' travel times, VeloCity distributes mobility profile optimization directly to individual CAVs. Vehicles query a localized traffic coordinator for a reservation table, independently compute their fastest conflict-free mobility profile, and reserve their requested space-time slots back with the coordinator. By natively adapting to any arbitrary road topology, the framework manages highly irregular urban areas without requiring scenario-specific tuning, all while guaranteeing collision-free and physically executable vehicle trajectories. Extensive simulations across four large-scale real-world urban maps (Tokyo, Manhattan, Rome, and Bologna) demonstrate the framework's scalability. Compared to established state-of-the-art models, VeloCity yields drastically lower travel times, tightly bounds delay variance, and successfully prevents congestion gridlocks even under extremely high vehicular densities.",
      "short_abstract": "As modern cities face increasingly severe traffic congestion, connected and autonomous vehicles (CAVs) have emerged as a crucial enabling technology for next-generation intelligent traffic management. However, fully realizing this potential is hindered by the...",
      "published": "2026-08-08",
      "updated": "2026-08-08",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "recent",
      "age_days": 11,
      "arxiv_url": "https://arxiv.org/abs/2608.08035",
      "pdf_url": "https://arxiv.org/pdf/2608.08035.pdf"
    },
    {
      "id": "2608.07750",
      "anchor": "paper-2608-07750",
      "title": "Multi-Task Consistency-based Detection of Adversarial Attacks",
      "authors": [
        "Cong Chen",
        "Jean-Philippe Monteuuis",
        "Jonathan Petit"
      ],
      "abstract": "Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to adversarial attacks has prompted concerns regarding their practical applications, specifically in the context of autonomous driving. Existing defenses often suffer from cost inefficiency, rendering their deployment impractical for resource-constrained applications. In this work, we propose an efficient and effective adversarial attack detection scheme leveraging the multi-task perception within a complex vision system. Adversarial perturbations are detected by the inconsistencies between the inference outputs of multiple vision tasks, e.g., object detection and instance segmentation. To this end, we developed a consistency score metric to measure the inconsistency between vision tasks. Next, we designed an approach to select the best model pairs for detecting inconsistencies effectively. Finally, we evaluated our defense against PGD attacks across multiple vision models on the BDD100k validation dataset. The experimental results demonstrated that our defense achieved a ROC-AUC performance of 99.9% detection within the considered attacker model.",
      "short_abstract": "Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to adversarial attacks has prompted concerns regarding their practical applications, specifically in the context of autonomous...",
      "published": "2026-08-07",
      "updated": "2026-08-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 8,
        "evidence": [
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "recent",
      "age_days": 12,
      "arxiv_url": "https://arxiv.org/abs/2608.07750",
      "pdf_url": "https://arxiv.org/pdf/2608.07750.pdf"
    },
    {
      "id": "2608.07621",
      "anchor": "paper-2608-07621",
      "title": "CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models",
      "authors": [
        "Hsu-kuang Chiu",
        "Stephen F. Smith"
      ],
      "abstract": "Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.",
      "short_abstract": "Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative...",
      "published": "2026-08-07",
      "updated": "2026-08-07",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Benchmark",
        "Cooperative / V2X",
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 45,
        "margin": 3,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "title: unified driving",
          "abstract: unified driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 12,
      "arxiv_url": "https://arxiv.org/abs/2608.07621",
      "pdf_url": "https://arxiv.org/pdf/2608.07621.pdf"
    },
    {
      "id": "2608.07468",
      "anchor": "paper-2608-07468",
      "title": "SimWAM: A Simple World Action Model for End-to-End Autonomous Driving",
      "authors": [
        "Zongchuang Zhao",
        "Xin Zhou",
        "Tianyang Xu",
        "Zhengyang Sun",
        "Kaixuan Zhou",
        "Honglin Li",
        "Dingkang Liang",
        "Xiang Bai"
      ],
      "abstract": "World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods incur costly test-time future imagination. We present SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without explicit future-frame generation at inference. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.",
      "short_abstract": "World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods incur costly test-time future imagination. We present SimWAM, a simple yet effective WAM that leverages...",
      "published": "2026-08-07",
      "updated": "2026-08-07",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 13,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 12,
      "arxiv_url": "https://arxiv.org/abs/2608.07468",
      "pdf_url": "https://arxiv.org/pdf/2608.07468.pdf"
    },
    {
      "id": "2608.07361",
      "anchor": "paper-2608-07361",
      "title": "Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model",
      "authors": [
        "Harisankar Babu",
        "Benjamin Coors",
        "Christopher Lang",
        "Hendrik Berkemeyer",
        "Tamim Asfour",
        "Simon Foell"
      ],
      "abstract": "Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by a single planning token that a generative planner decodes into a trajectory. Borrowing the planner as a trajectory-space logit lens, we decode the planning token from every one of the 32 decoder layers and measure two signals: the linear decodability of the navigation command and trajectory compatibility with the frozen native planner. Our diagnostic shows that semantic intent is linearly decodable early: command-probe accuracy reaches 97.7\\% after the first decoder layer, compared with 16.7\\% chance. In contrast, compatibility with the frozen native planner improves gradually across depth, with open-loop Avg-L2 reaching its minimum of 2.11\\,m only at the final layer. Learned readouts from the first layer recover much of this gap, indicating that planning information is already present early but is not yet represented in the format expected by the deployed planner. Ranking decoder layers by the angular deviation they induce in the planning token permits removal of 8 of 32 layers within an approximately 5\\% relative open-loop error increase and yields a measured 1.33$\\times$ decoder speedup. At the evaluated sample size, no family-specific degradation is statistically resolved. These findings are limited to the evaluated ORION checkpoint and Bench2Drive setup.",
      "short_abstract": "Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by a single planning token that...",
      "published": "2026-08-07",
      "updated": "2026-08-07",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 24,
        "margin": 6,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "recent",
      "age_days": 12,
      "arxiv_url": "https://arxiv.org/abs/2608.07361",
      "pdf_url": "https://arxiv.org/pdf/2608.07361.pdf"
    },
    {
      "id": "2608.07740",
      "anchor": "paper-2608-07740",
      "title": "Enhancing Autonomous Vehicle Navigation with a Clothoid-Based Lateral Controller",
      "authors": [
        "Aashish Shaju",
        "Steve Southward",
        "Mehdi Ahmadian"
      ],
      "abstract": "This study introduces an advanced lateral control strategy for autonomous vehicles using a clothoid-based approach integrated with an adaptive lookahead mechanism. The primary focus is on enhancing lateral stability and path-tracking accuracy through the application of Euler spirals for smooth curvature transitions, thereby reducing passenger discomfort and the risk of vehicle rollover. An innovative aspect of our work is the adaptive adjustment of lookahead distance based on real-time vehicle dynamics and road geometry, which ensures optimal path following under varying conditions. A quasi-feedback control algorithm constructs optimal clothoids at each time step, generating the appropriate steering input. A lead filter compensates for the vehicle's lateral dynamics lag, improving control responsiveness and stability. The effectiveness of the proposed controller is validated through a comprehensive co-simulation using TruckSim and Simulink, demonstrating significant improvements in lateral control performance across diverse driving scenarios. Future directions include scaling the controller for higher-speed applications and further optimization to minimize off-track errors, particularly for articulated vehicles.",
      "short_abstract": "This study introduces an advanced lateral control strategy for autonomous vehicles using a clothoid-based approach integrated with an adaptive lookahead mechanism. The primary focus is on enhancing lateral stability and path-tracking accuracy through the...",
      "published": "2026-08-07",
      "updated": "2026-08-07",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 23,
        "evidence": [
          "abstract: vehicle dynamics",
          "abstract: vehicle control",
          "title: controller",
          "abstract: controller",
          "abstract: control",
          "abstract: steering or braking",
          "abstract: trajectory tracking"
        ]
      },
      "recency": "recent",
      "age_days": 12,
      "arxiv_url": "https://arxiv.org/abs/2608.07740",
      "pdf_url": "https://arxiv.org/pdf/2608.07740.pdf"
    },
    {
      "id": "2608.07061",
      "anchor": "paper-2608-07061",
      "title": "Automated vehicles: Challenging transition from Minimal Risk Condition",
      "authors": [
        "Wolfgang Kröger",
        "Leon Johann Brettin"
      ],
      "abstract": "If an automated vehicle is no longer capable of handling the dynamic driving tasks, regulations require that a minimal risk manouvre is initiated aimed at achieving minimum risk condition. Details on measures and actions to resolve this standstill position are missing. This article provides clues for closing this gap and addresses various states in which the human intervention operator must understand the current situation and decide whether it is more acceptable to maintain the stay, e.g. on road-shoulder, or to move the vehicle back into traffic flow or 'limp it' to another position, depending on the outcome of a health check. Our investigations, based on a list of nearly 30 reasons and a subsequent development of 'scenario trees' prove the complexity of the situation and the heavy difficulties faced by the operator to handle in a correct manner. We also discuss how the technical automated driving system can be enhanced so that it stronger supports the operator or allows to even dismiss human intervention. This deems technically feasible but limits became obvious, so that assistance from out side appears indispensable at least for the ramp-up phase of respective automated vehicle. The non-trivial problem of transitioning the 'stranded' vehicle needs to be addressed by further regulation for which this paper may provide useful input.",
      "short_abstract": "If an automated vehicle is no longer capable of handling the dynamic driving tasks, regulations require that a minimal risk manouvre is initiated aimed at achieving minimum risk condition. Details on measures and actions to resolve this standstill position...",
      "published": "2026-08-07",
      "updated": "2026-08-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy"
        ]
      },
      "recency": "recent",
      "age_days": 12,
      "arxiv_url": "https://arxiv.org/abs/2608.07061",
      "pdf_url": "https://arxiv.org/pdf/2608.07061.pdf"
    },
    {
      "id": "2608.06535",
      "anchor": "paper-2608-06535",
      "title": "Flaky Test Recognition when Testing CPSs Using Hybrid Models",
      "authors": [
        "Zahra Sadri-Moshkenani",
        "Justin Bradley",
        "Gregg Rothermel"
      ],
      "abstract": "Cyber-Physical Systems (CPSs) have many applications, ranging from simple thermostat systems to autonomous driving systems and medical devices. Like all systems, they need to function correctly. To this end, validation techniques such as testing that can effectively reveal faults are required. Also, CPSs usually operate in uncertain environments where various expected/unexpected events can affect their behaviors. These events together with timing and synchronization incidents may result in various/different CPS behaviors and may cause a CPS to pass a test under some conditions and fail it under others. Such test cases are called ``flaky'' test cases and we call the conditions that are responsible for them ``flaky conditions'' and call these behaviors ``flaky behaviors''. When test cases are flaky, testing results are unreliable. To achieve more reliable test results, engineers may attempt to recognize flaky test cases and remove them. In this work, beginning with a test case generation and execution technique called HyTest that we created previously, we integrate TReVa, a new technique that validates testing results provided by HyTest and FlaRe, a new technique that recognizes flaky test cases and flaky conditions using hybrid models during the early stages of CPS development in just one additional round of testing. We present the results of an empirical study evaluating the effectiveness of our new approach (which we call HyTestTF). Our results show that correctly differentiate flaky test cases from non-flaky test cases",
      "short_abstract": "Cyber-Physical Systems (CPSs) have many applications, ranging from simple thermostat systems to autonomous driving systems and medical devices. Like all systems, they need to function correctly. To this end, validation techniques such as testing that can...",
      "published": "2026-08-06",
      "updated": "2026-08-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "recent",
      "age_days": 13,
      "arxiv_url": "https://arxiv.org/abs/2608.06535",
      "pdf_url": "https://arxiv.org/pdf/2608.06535.pdf"
    },
    {
      "id": "2608.06008",
      "anchor": "paper-2608-06008",
      "title": "Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features",
      "authors": [
        "Sining Ang",
        "Yuguang Yang",
        "Yan Wang"
      ],
      "abstract": "Large video diffusion models provide rich spatiotemporal priors for autonomous driving, but existing world-action models often inherit the cost of iterative future-video generation even though deployment only requires an ego trajectory. We ask a more basic question: how much of a video diffusion model must be executed to make a reliable driving decision? Through a controlled study of video denoising timesteps and Diffusion Transformer (DiT) depth, we find that planning performance is largely insensitive to the tested video-noise levels, whereas strong trajectories can already be decoded from intermediate layers. Based on this observation, we introduce Adaptive-WAM, a quality-aware multi-exit planner built on a Wan2.2-5B backbone. Trajectory diffusion heads are attached to selected DiT blocks, and a lightweight trajectory-quality scorer terminates inference once the best trajectory decoded so far satisfies a quality threshold; otherwise, computation continues from the cached hidden state to a deeper exit. The deployed planner therefore avoids the iterative classifier-free denoising loop and VAE decoding required for future-video synthesis, while dynamically allocating backbone depth according to trajectory quality. On NAVSIM, the adaptive single-trajectory planner achieves 90.8 PDMS; a separate fixed-exit variant reaches 92.6 PDMS with 64 proposals. It further obtains 89.9 EPDMS on NAVSIM v2, yielding the best reported results among the compared front-view video world-model planners. Without target-domain fine-tuning, Adaptive-WAM transfers to nuScenes with 0.88 m average L2 error and a 0.08\\% collision rate. On an A100, adaptive routing improves PDMS from 90.62 to 90.79 while averaging 170 ms end-to-end planning latency, approximately 10\\% below the 190 ms fixed block-15 planner and 47\\% below the 320 ms fixed full-depth planner. Code will be released.",
      "short_abstract": "Large video diffusion models provide rich spatiotemporal priors for autonomous driving, but existing world-action models often inherit the cost of iterative future-video generation even though deployment only requires an ego trajectory. We ask a more basic...",
      "published": "2026-08-06",
      "updated": "2026-08-06",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 11,
        "evidence": [
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "recent",
      "age_days": 13,
      "arxiv_url": "https://arxiv.org/abs/2608.06008",
      "pdf_url": "https://arxiv.org/pdf/2608.06008.pdf"
    },
    {
      "id": "2608.05634",
      "anchor": "paper-2608-05634",
      "title": "MEC-Patch: Visible-Infrared Cross-Modal Adversarial Attack Driven by Intrinsic Material Emissivity Laws",
      "authors": [
        "Zhixiang Huang",
        "Xinbo Nie",
        "Wenxuan Wang",
        "Lu Yang",
        "Xin Li",
        "Xuelin Qian",
        "Peng Wang"
      ],
      "abstract": "With the widespread deployment of visible-infrared multimodal perception systems in safety-critical domains such as autonomous driving, evaluating their cross-modal adversarial robustness has become increasingly vital. However, existing approaches exhibit significant limitations in approximating the intrinsic laws of imaging. Most studies either focus on a single modality, failing to bypass cross-modal verification, or simplify infrared modeling into heuristic pixel-intensity distributions, neglecting the impact of ambient temperature fluctuations on adversarial stability. To bridge this gap, this paper proposes MEC-Patch, a cross-modal adversarial attack framework driven by intrinsic physical laws. By leveraging the Stefan-Boltzmann Law, we establish a physics-grounded cross-spectral mapping that explicitly links material emissivity to thermal radiation. Building on this formulation, we reveal that, under a fixed emissivity distribution, ambient temperature variations induce consistent global scaling while preserving relative emissivity-induced contrast. We exploit this property to construct temperature-robust adversarial perturbations whose discriminative patterns remain stable in the infrared modality, thereby fundamentally mitigating environmental sensitivity. Furthermore, we employ the physics-constrained NSGA-II algorithm to synergistically optimize the material-distribution-based patch parameters effective across both modalities, while enhancing generalization through a Dynamic Adversarial Resampling (DAR) strategy. Experimental results demonstrate that MEC-Patch effectively deceives state-of-the-art multimodal detectors and exhibits high robustness within high-fidelity, physically-consistent, and multi-scene simulation environments. This research provides a physical-law-driven perspective for the security assessment of multimodal perception systems.",
      "short_abstract": "With the widespread deployment of visible-infrared multimodal perception systems in safety-critical domains such as autonomous driving, evaluating their cross-modal adversarial robustness has become increasingly vital. However, existing approaches exhibit...",
      "published": "2026-08-06",
      "updated": "2026-08-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 28,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: verification",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 13,
      "arxiv_url": "https://arxiv.org/abs/2608.05634",
      "pdf_url": "https://arxiv.org/pdf/2608.05634.pdf"
    },
    {
      "id": "2608.06445",
      "anchor": "paper-2608-06445",
      "title": "Game-Theoretic Inverse Reinforcement Learning for Modeling Competitive Human Driving: A Cut-in Prediction Study",
      "authors": [
        "Yu Song"
      ],
      "abstract": "Capturing the strategic decision-making inherent in competitive human driving is critical for autonomous vehicle safety and traffic simulation. This study demonstrates that game-theoretic Inverse Reinforcement Learning (IRL) provides a robust framework for this challenge. We present a comprehensive analysis comparing data-driven IRL models against an established physics-based game-theoretic approach for predicting aggressive, safety-critical cut-in lane changes. Using the high-fidelity highD dataset, we systematically develop and evaluate a series of IRL models with increasing feature complexity. Our results reveal significant advantages: the best-performing IRL models achieve an overall prediction accuracy exceeding 75 percent while maintaining a Cut-In precision up to 51 percent and recall up to 49 percent. This represents a significant improvement over the established physics-based benchmark, which achieved only 4.4 percent precision in these high-stakes scenarios. The analysis reveals a clear trade-off: incorporating granular, instantaneous features yields higher precision, while adding temporal consistency features maximizes recall. These findings suggest that IRL-based models can effectively bridge the gap between microscopic driver intent and macroscopic safety outcomes, providing a more reliable foundation for modeling interactions in mixed-autonomy environments.",
      "short_abstract": "Capturing the strategic decision-making inherent in competitive human driving is critical for autonomous vehicle safety and traffic simulation. This study demonstrates that game-theoretic Inverse Reinforcement Learning (IRL) provides a robust framework for...",
      "published": "2026-08-06",
      "updated": "2026-08-06",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Benchmark",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 2,
        "evidence": [
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 13,
      "arxiv_url": "https://arxiv.org/abs/2608.06445",
      "pdf_url": "https://arxiv.org/pdf/2608.06445.pdf"
    },
    {
      "id": "2608.06021",
      "anchor": "paper-2608-06021",
      "title": "Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models",
      "authors": [
        "Eulogio Quemada-Torres",
        "Alberto Jaenal",
        "Francisco-Angel Moreno",
        "Javier Gonzalez-Jimenez"
      ],
      "abstract": "Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition (VPR) is able to successfully meet the first two requirements, but its low metric accuracy makes it less suitable than standard VL approaches based on local features or neural representations. This limitation can be overcome by integrating VPR with the accurate local trajectory estimates produced by feed-forward neural 3D geometry (FF3D) models. In this paper, we address sequential appearance-based localization through a topometric framework that iteratively combines probabilistic VPR with FF3D metric pose estimation in controlled image sets. Our approach proposes an automatic offline mapping tool that models the topometric pose-appearance interaction in the different parts of the scene. This map is later employed by an online particle filter that estimates the pose from odometry and belief over places for FF3D inference, successfully incorporating neural metric estimation into probabilistic appearance-based localization. We extensively evaluate the framework on three known benchmarks, demonstrating substantial improvements over existing appearance-based methods. The modularity of our approach allows the descriptor extractor and FF3D model to remain interchangeable, and a focused analysis further shows that sequential belief can mitigate severe failures under perceptual aliasing.",
      "short_abstract": "Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition...",
      "published": "2026-08-06",
      "updated": "2026-08-06",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 28,
        "evidence": [
          "title: localization",
          "abstract: localization",
          "abstract: mapping",
          "abstract: place recognition",
          "abstract: pose estimation"
        ]
      },
      "recency": "recent",
      "age_days": 13,
      "arxiv_url": "https://arxiv.org/abs/2608.06021",
      "pdf_url": "https://arxiv.org/pdf/2608.06021.pdf"
    },
    {
      "id": "2608.09987",
      "anchor": "paper-2608-09987",
      "title": "The Impact of Cut-ins on Mixed-Autonomy Traffic Flow: Bridging Microscopic Game-Theoretic Model and Macroscopic Flow Analysis",
      "authors": [
        "Yu Song"
      ],
      "abstract": "The transition to automated transportation introduces mixed-autonomy traffic where human drivers may strategically exploit the risk-averse behavior of Connected and Automated Vehicles (CAVs). While microscopic models capture these dyadic interactions, their aggregate impact on network-level stability remains unquantified due to the scale gap between agent-based and continuum models. This study bridges this analytical divide by integrating a game-theoretic friction term directly into the macroscopic kinematic wave framework. We model the cut-in maneuver as a Stackelberg game, identifying a distinct \"exploitation window\" where human drivers leverage CAV defensiveness to execute aggressive merges. By deriving a closed-form micro-macro bridge, we translate these discrete strategic outcomes into a continuous friction parameter that endogenously modifies the traffic conservation law. Theoretical analysis and numerical simulations confirm that this behavioral asymmetry functions as a deterministic destabilizer, generating perturbation source terms that trigger phantom jams and strictly reduce road capacity. Crucially, we reveal a convex relationship between CAV penetration and system efficiency, identifying a critical instability regime at intermediate penetration rates (approximately 45 percent) where the frequency of exploitable interactions is maximized. These findings demonstrate that without socially aware control policies, the defensive nature of early-deployment CAVs may paradoxically degrade traffic flow stability.",
      "short_abstract": "The transition to automated transportation introduces mixed-autonomy traffic where human drivers may strategically exploit the risk-averse behavior of Connected and Automated Vehicles (CAVs). While microscopic models capture these dyadic interactions, their...",
      "published": "2026-08-06",
      "updated": "2026-08-06",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 6,
        "evidence": [
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 13,
      "arxiv_url": "https://arxiv.org/abs/2608.09987",
      "pdf_url": "https://arxiv.org/pdf/2608.09987.pdf"
    },
    {
      "id": "2608.05560",
      "anchor": "paper-2608-05560",
      "title": "From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs",
      "authors": [
        "Jiawei Qiu",
        "Yichen Xu",
        "Jianzhe Ma",
        "Mingyang Yu",
        "Wenbin Zhu",
        "Yang Han",
        "Pinzheng Lv",
        "Wenxuan Wang"
      ],
      "abstract": "Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a well-suited testbed: accident causes span diverse injury dimensions and pre-accident spatiotemporal cues draw on reasoning capabilities shared with broader safety domains such as autonomous driving and fall detection. We introduce SPRINT (Sports Proactive Risk INference Testbed), a benchmark of 2,888 real-world sports videos (2,440 accident, 448 safe controls) spanning 14 sports and 3 environmental settings. Accident videos feature fine-grained annotations of early hazard cues, accident timing, and hierarchical causes; safe videos are manually verified as accident-free and serve to diagnose prompt-induced false alarms. Evaluating state-of-the-art MLLMs under diverse prompts and temporal windows reveals a sharp gap between hazard sensitivity and understanding: the best model exceeds 95% in signaling hazards yet falls below 50% in identifying their causes. Diagnostic experiments further show that explicit danger queries trigger severe false alarms even on hazard-free videos. These findings indicate that current MLLMs exhibit only superficial proactive safety, lacking stable, cause-grounded early warning, and underscore the need for reliable proactive safety in dynamic physical environments. Data and code will be open-sourced upon acceptance.",
      "short_abstract": "Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a well-suited testbed: accident...",
      "published": "2026-08-05",
      "updated": "2026-08-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 14,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "recent",
      "age_days": 14,
      "arxiv_url": "https://arxiv.org/abs/2608.05560",
      "pdf_url": "https://arxiv.org/pdf/2608.05560.pdf"
    },
    {
      "id": "2608.05356",
      "anchor": "paper-2608-05356",
      "title": "LoDA: A Level of Detection Aware Method and a Multimodal Sensing Benchmark for Object Level Change Detection",
      "authors": [
        "Haitian Wang",
        "Xinyu Wang",
        "Sheldon Fung",
        "Xian Zhang",
        "Zichen Geng"
      ],
      "abstract": "High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in multi-temporal urban LiDAR to keep digital maps aligned with the physical world. Existing approaches from raster height differencing to depth image and point-cloud networks often remain tile-based and threshold-driven, yielding per-point scores without explicit detection limits or consistent object-level labels. We propose an object-level 3D change-detection pipeline that integrates detection-limit-aware registration, geometry-driven object proxies with rule-based semantic and instance segmentation, and displacement cues in height, volume, and surface-normal direction to assign five change labels with confidence. By decoupling registration, geometry, and semantics, the pipeline propagates pose uncertainty into spatially varying detection limits, stabilizes cross-epoch correspondences, and suppresses false changes caused by residual misalignment and density variation. We also present LoDA, a level-of-detection (LoD) aware benchmark for the Subiaco district with fused multi-temporal vehicle-LiDAR maps constructed with LiDAR, GNSS, and IMU support, semantic instances, and object-level annotations. On this benchmark, our method achieves 95.0% accuracy, 90.8% macro F1, and 83.0% macro IoU, exceeding the best baseline by 8.7 IoU points and 4.4 F1 points. On the public Urb3DCD-V2 benchmark evaluated under the official point-wise protocol, it reaches 96.81% mean accuracy and 89.52% mean change IoU, improving over the strongest reported baselines by 1.36 points in mAcc and 3.18 points in mIoUch.",
      "short_abstract": "High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in multi-temporal urban LiDAR to keep digital maps aligned with the physical world. Existing approaches from...",
      "published": "2026-08-05",
      "updated": "2026-08-05",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 5,
        "evidence": [
          "abstract: scene segmentation",
          "abstract: LiDAR"
        ]
      },
      "recency": "recent",
      "age_days": 14,
      "arxiv_url": "https://arxiv.org/abs/2608.05356",
      "pdf_url": "https://arxiv.org/pdf/2608.05356.pdf"
    },
    {
      "id": "2608.04852",
      "anchor": "paper-2608-04852",
      "title": "Toward Blockage-Resilient 6G-V2X Connectivity: Semi-Distributed Bandit with Dynamic Arm Set for mmWave HetNets",
      "authors": [
        "Weiqi Chi",
        "Bo Qian",
        "Hanlin Wu",
        "Donghui Li",
        "Haibo Zhou",
        "Manabu Tsukada"
      ],
      "abstract": "The vision for 6G vehicle-to-everything (V2X) communications demands reliable, adaptive connectivity for fully autonomous driving across complex dynamic environments. Millimeter-wave (mmWave) user association (UA) in heterogeneous vehicular networks presents a particularly demanding instance of this problem, where dynamic blockages and rapid channel variations continuously undermine the stationary reward assumptions of traditional multi-armed bandit (MAB) frameworks. This paper proposes a fully distributed blockage-aware non-stationary dynamic bandit algorithm (BAND) and its semi-distributed extension S-BAND for cooperative learning across vehicles. Blockage prediction is incorporated into the change-detection (CD) mechanism to suppress false alarms, while a dynamic base station (BS) set management scheme balances exploration and exploitation across large-scale BS deployments without requiring centralized channel state information (CSI) acquisition or offline training. In S-BAND, vehicles accumulate BS reward estimates as local knowledge and periodically upload them to the macro base station (MBS), which aggregates them into cluster-based central knowledge. A trajectory-aligned knowledge (TAK) region is proposed to capture the spatial correlation of mmWave channel characteristics. A knowledge inheritance fidelity (KIF) metric is introduced to quantify knowledge transfer quality. Simulation results on a realistic urban topology show that BAND and S-BAND achieve 34.9% and 59.4% regret reduction relative to a centralized MAB baseline, with performance gains sustained across blockage rates ranging from 10% to 50%. The proposed TAK region consistently outperforms the traditional K-means clustering scheme under both fidelity criteria.",
      "short_abstract": "The vision for 6G vehicle-to-everything (V2X) communications demands reliable, adaptive connectivity for fully autonomous driving across complex dynamic environments. Millimeter-wave (mmWave) user association (UA) in heterogeneous vehicular networks presents...",
      "published": "2026-08-05",
      "updated": "2026-08-05",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 33,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: cooperative systems",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 14,
      "arxiv_url": "https://arxiv.org/abs/2608.04852",
      "pdf_url": "https://arxiv.org/pdf/2608.04852.pdf"
    },
    {
      "id": "2608.04776",
      "anchor": "paper-2608-04776",
      "title": "NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment",
      "authors": [
        "Yu Zhao",
        "Jiangyu Pan",
        "Tao Hu",
        "Ming Yin",
        "Fan Yang",
        "Jiangfan Liu",
        "Xiubo Liang"
      ],
      "abstract": "The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research has made progress in collision prediction, accurately quantifying risk levels from monocular vision inputs remains challenging due to the complex dynamics of multi-agent interactions and the inherent uncertainty in real-world environments. To address these challenges, we present NSF-HRPT, a novel framework that combines learning-based perception with structured reasoning for quantitative risk assessment. Our approach features a Neural Semantic Field (NSF) that learns to model scene semantics, trajectory predictions, and probabilistic Time-to-Collision (TTC) distributions from simulation data. During inference, the pre-trained NSF serves as a prior for our Hierarchical Risk Perception Tree (HRPT), which enables efficient parallel computation and spatial reasoning about multi-agent risks. Additionally, we introduce a Sim2Real enhancement strategy that improves real-world applicability without retraining by incorporating priors from foundation models. Extensive evaluations demonstrate that our framework achieves state-of-the-art performance on synthetic benchmarks and delivers competitive, near-state-of-the-art results on real-world datasets for both TTC estimation accuracy and risk localization precision. The proposed method provides an effective solution for real-time risk awareness from monocular camera inputs.",
      "short_abstract": "The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research has made progress in collision prediction, accurately quantifying risk levels from monocular vision inputs...",
      "published": "2026-08-05",
      "updated": "2026-08-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 8,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: risk assessment",
          "abstract: uncertainty"
        ]
      },
      "recency": "recent",
      "age_days": 14,
      "arxiv_url": "https://arxiv.org/abs/2608.04776",
      "pdf_url": "https://arxiv.org/pdf/2608.04776.pdf"
    },
    {
      "id": "2608.04568",
      "anchor": "paper-2608-04568",
      "title": "Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching",
      "authors": [
        "Runwei Guan",
        "Di Tian",
        "Ningwei Ouyang",
        "Ruixiao Zhang",
        "Shaofeng Liang",
        "Haocheng Zhao",
        "Lianqing Zheng",
        "Xiaokai Bai",
        "Guotao Wang",
        "Daizong Liu",
        "Henghui Ding",
        "Hui Xiong"
      ],
      "abstract": "As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor extensions largely rely on monocular images alone. Both settings fall short of real-world outdoor perception, where heterogeneous sensors capture complementary yet distinct physical properties, such as visual texture, 3D geometry, and object kinematics, that are indispensable for flexible and robust query-adaptive grounding but remain under-exploited. To bridge this gap, we introduce Talk2Sensors, the first multi-sensor 3D visual grounding dataset built upon camera, LiDAR, and 4D radar. It contains 8,682 language instructions and 20,558 referred objects, with diverse prompts explicitly aligned with sensor-specific physical cues. Furthermore, we propose TSFormer, a unified Transformer-based framework for language-guided 3D visual grounding in autonomous driving. TSFormer adopts a coarse-to-fine property-aware fusion strategy: the Language-Routed Property Sampler first performs coarse text-conditioned feature retrieval by modulating sensor sampling weights with query-level linguistic cues, while the subsequent Sparse-Preserving Modality Arbiter module conducts fine-grained modality arbitration and text-guided refinement to determine the precise referred spatial location. This design enables dynamic routing of appearance, geometry, and motion cues according to the semantic requirements of each prompt, preventing dense modalities from overwhelming sparse but critical sensor signals. Extensive experiments demonstrate that TSFormer achieves state-of-the-art performance across multiple benchmarks: it improves over the strongest baseline by 8.05 mAP on Talk2Sensors, and transfers to the monocular Mono3DRefer benchmark with 53.05\\% Acc@0.5.",
      "short_abstract": "As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor extensions largely rely on monocular images alone. Both settings fall short of...",
      "published": "2026-08-05",
      "updated": "2026-08-05",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: radar",
          "abstract: camera"
        ]
      },
      "recency": "recent",
      "age_days": 14,
      "arxiv_url": "https://arxiv.org/abs/2608.04568",
      "pdf_url": "https://arxiv.org/pdf/2608.04568.pdf"
    },
    {
      "id": "2608.04453",
      "anchor": "paper-2608-04453",
      "title": "TwinIR: Coordinated Invisible Dual-Point Attacks on Online HD Map Construction",
      "authors": [
        "Haibo Hu",
        "Jianghuai Deng",
        "Chen Tang",
        "Yang Lou",
        "Qian Xu",
        "Jianping Wang"
      ],
      "abstract": "Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against online map construction are limited by a cross-boundary compensation effect: after the target boundary is perturbed, another visible boundary may retain sufficient geometric cues for the model to recover the original road geometry. Based on this observation, we propose TwinIR, a new mechanism-guided physical attack methodology for online map construction. TwinIR jointly optimizes attack effectiveness and point sparsity, seeking the minimum number of attack points needed to suppress compensating geometric cues from surrounding boundaries. To reduce the perceptibility of multi-point attacks, TwinIR models camera responses to near-infrared illumination and maps optimized attack points to feasible physical placements, producing camera-visible interference with minimal visible-spectrum changes. Experiments on nuScenes across state-of-the-art online map construction models show that TwinIR reduces mAP by 8.18-8.96 percentage points under RSA and 2.84-5.62 points under ETA, while increasing the unreachable-goal rate by 25-28 points and the unsafe-planned-trajectory rate by 19-20 points over clean inputs. These attacks are also validated on a real-world testbed AV, where TwinIR successfully induces both road straightening and early-turn deformations while remaining inconspicuous in full-color views.",
      "short_abstract": "Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against online map construction are limited by a cross-boundary compensation effect: after the target boundary is perturbed,...",
      "published": "2026-08-05",
      "updated": "2026-08-05",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 20,
        "evidence": [
          "title: HD map",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "recent",
      "age_days": 14,
      "arxiv_url": "https://arxiv.org/abs/2608.04453",
      "pdf_url": "https://arxiv.org/pdf/2608.04453.pdf"
    },
    {
      "id": "2608.07577",
      "anchor": "paper-2608-07577",
      "title": "Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects",
      "authors": [
        "Felix Schaller"
      ],
      "abstract": "A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-drawn carriage, road debris, livestock on a rural road) it can only force a confident but wrong specific label or drop the object. Prior work in this series replaced the flat label set with a hierarchical taxonomy and a runtime abstraction rule, but evaluated it only on the boxes a closed detector already produces. This paper takes the layer open-world: we place taxonomic abstraction on top of class-agnostic region proposals so objects the closed detector never boxes can still be classified or flagged; we report a feasibility study of three open-world signals (class-agnostic segmentation, appearance-based out-of-distribution scoring, monocular depth) that shows why no single 2D cue suffices and how they compose; and we run the evaluation the earlier papers could not, a ground-truth leave-classes-out benchmark on real annotated objects. Holding out seven COCO classes and classifying their 235 ground-truth crops, a flat closed head emits a confident wrong specific label 100% of the time (37% of them in the wrong super-category, e.g. an animal named as a vehicle), whereas the hierarchical layer emits zero confident wrong specific labels and safely handles 94% of the objects (a correct super-category, or an explicit UNKNOWN OBSTACLE). We are explicit that this is a safety result, not a specificity one: the correct super-category is recovered only 26% of the time and the remaining 69% are conservatively flagged unknown. The contribution is an open-world perception layer that never makes a confident categorical mistake on an out-of-vocabulary object, together with an honest account of its cost.",
      "short_abstract": "A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-drawn carriage, road debris, livestock on a rural road) it can only force a confident but wrong specific label or drop...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: perception",
          "abstract: perception"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.07577",
      "pdf_url": "https://arxiv.org/pdf/2608.07577.pdf"
    },
    {
      "id": "2608.04412",
      "anchor": "paper-2608-04412",
      "title": "muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards",
      "authors": [
        "Yang Chen",
        "Yicheng Zhu",
        "Tao Li",
        "Zilin Bian"
      ],
      "abstract": "High-quality driving data are essential for autonomous-driving systems and generative world models. However, rare and safety-critical scenarios involving adverse weather, braking under low tire--road friction, and uneven road geometry are costly and risky to collect at scale. Existing video-generation and 3D Gaussian editing methods can modify weather appearance or road geometry, but typically do not couple these edits with tire--road interaction and vehicle dynamics. As a result, an edited video may retain its original trajectory even when the modified road condition should alter braking, wheel slip, load transfer, and ego-camera motion. We present muSync-GS, a physics-synchronized framework for driving video synthesis under adverse-weather and road-elevation hazards. A precipitation-derived road-surface condition jointly controls road appearance and tire friction, while a shared road-elevation profile drives both visible road-geometry editing and axle excitation. A calibrated vehicle model predicts speed, slip ratio, normal loads, and pitch for constructing the ego-camera trajectory and synchronized physical annotations. On 12 held-out CarSim cases spanning precipitation levels, brake inputs, and road-profile parameters, the model achieves mean case-wise RMSEs of 0.0273 m/s for speed, 0.0590 degrees for pitch, 0.0101 for slip ratio, and 26.61 N for per-wheel normal load. Together with the reconstructed-scene experiments, these results show that muSync-GS accurately reproduces vehicle responses under held-out controls while synchronizing them with controllable scene edits and ego-camera motion.",
      "short_abstract": "High-quality driving data are essential for autonomous-driving systems and generative world models. However, rare and safety-critical scenarios involving adverse weather, braking under low tire--road friction, and uneven road geometry are costly and risky to...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 3,
        "evidence": [
          "abstract: vehicle dynamics",
          "abstract: steering or braking"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.04412",
      "pdf_url": "https://arxiv.org/pdf/2608.04412.pdf"
    },
    {
      "id": "2608.04130",
      "anchor": "paper-2608-04130",
      "title": "Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models",
      "authors": [
        "Jiaju Han",
        "Xuemeng Sun",
        "Qike Zhang",
        "Xiang Chen",
        "Luwei Yang",
        "Jiahuan Long",
        "Yiwei Wei",
        "Jiujiang Guo",
        "Chengyin Hu"
      ],
      "abstract": "Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distribution, motion state, collision risk, semantic category, and radial velocity. Radar4D-VLM combines proposal-grounded temporal object tokenization, global scene context, and explicit kinematic tokens within a unified frozen-backbone interface. On sequence-isolated K-Radar development validation, its Top-64 proposal recall reaches 98.13% at 4 m, exceeding fixed-lattice and uniform-random controls by 6.40 and 22.83 percentage points, respectively. We further evaluate 24 matched runs spanning eight frozen Qwen, Phi, Mistral, Llama, and Gemma backbones under an identical adaptation budget. The radar-token interface remains compatible across all five language-model families, while matched aligned, permuted, and no-language controls show sensor dependence but no stable direct-head gain from aligned language supervision. These results establish a reproducible foundation for radar-only multimodal scene and motion reasoning while separating interface compatibility from the benefit of language supervision.",
      "short_abstract": "Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 12,
        "evidence": [
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar",
          "abstract: camera"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.04130",
      "pdf_url": "https://arxiv.org/pdf/2608.04130.pdf"
    },
    {
      "id": "2608.03521",
      "anchor": "paper-2608-03521",
      "title": "Pivot-Centric Trajectory Prediction: Bridging Long Horizons via Dynamical Guidance",
      "authors": [
        "Xiucong Zhao",
        "Jindong Tian",
        "Hao Miao"
      ],
      "abstract": "Forecasting precise future motion of surrounding agents is essential for reliable autonomous vehicles. However, as the demand for longer prediction horizons increases, existing endpoint-completion or iterative-refine methods increasingly struggle with weak guidance and compounding errors. To tackle the long-horizon prediction challenge, we propose Pivot-Centric Trajectory Prediction (PCTP). By introducing ``pivots'' and focusing on predicting pivot points along extended trajectories, we divide the long-term prediction task into short-term sub-tasks at various scales. Specifically, PCTP decouples the long-term trajectory predicting process into two processes: pivot prediction and pivot-based trajectory refinement. The pivot prediction process aims to utilize global map context and agent-to-agent interactions to identify these ``pivot points'', while the pivot-based trajectory refinement process focuses on local map details and refines the short-term trajectory based on predicted ``pivot points''. Compared with existing methods, PCTP provides more intermediate guidance while reducing compounding errors. Moreover, PCTP is a flexible approach that can be integrated into most state-of-the-art trajectory prediction models. Experimental results show that PCTP improves the prediction accuracy of leading models on both Argoverse I and Argoverse II datasets with minimal impact on model size. Specifically, PCTP combined with QCNet outperforms all published ensemble-free methods on the Argoverse II leaderboard at submission.",
      "short_abstract": "Forecasting precise future motion of surrounding agents is essential for reliable autonomous vehicles. However, as the demand for longer prediction horizons increases, existing endpoint-completion or iterative-refine methods increasingly struggle with weak...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 31,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.03521",
      "pdf_url": "https://arxiv.org/pdf/2608.03521.pdf"
    },
    {
      "id": "2608.03490",
      "anchor": "paper-2608-03490",
      "title": "Lightweight 3D Object Detection via Mamba-Based Knowledge Distillation",
      "authors": [
        "Quoc Cuong Ninh",
        "Huy Xuan Pham",
        "Anh Tung Nguyen",
        "Dinh Hoan Trinh"
      ],
      "abstract": "3D object detection using light detection and ranging (LiDAR) sensors requires a balance between accuracy and computational efficiency for onboard perception in autonomous driving and robotic navigation. Many existing LiDAR-based detection methods employ complex architectures to extract features, integrating large amounts of contextual information to enhance accuracy. This often results in significant computational costs, leading to suboptimal performance on resource-constrained embedded devices. In this study, we propose a knowledge distillation framework that transfers object-level voxel representations from a strong teacher model to lightweight student models through selective voxel-space feature alignment. Taking advantage of the linear-time sequence model with selective state spaces (Mamba), we design a multi-branch Mamba teacher backbone and a box-aware feature transfer mechanism that aligns spatially corresponding voxel features between teacher and student networks through a Mamba-based projection module. Experimental results on both a public dataset and real-world data show that our approach significantly reduces computational load while maintaining competitive accuracy compared with state-of-the-art methods.",
      "short_abstract": "3D object detection using light detection and ranging (LiDAR) sensors requires a balance between accuracy and computational efficiency for onboard perception in autonomous driving and robotic navigation. Many existing LiDAR-based detection methods employ...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 20,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "abstract: LiDAR"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.03490",
      "pdf_url": "https://arxiv.org/pdf/2608.03490.pdf"
    },
    {
      "id": "2608.03330",
      "anchor": "paper-2608-03330",
      "title": "Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving",
      "authors": [
        "Yue Yao"
      ],
      "abstract": "This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient models based on polynomial representations. While conventional sequence-based representations often struggle with noise and generalization, this work demonstrates that polynomial representations offer significant advantages in computational efficiency, generalization, and prediction plausibility. Through theoretical analysis and empirical validation, this thesis demonstrates that moderate-degree polynomials capture real-world motion dynamics with high fidelity without constraining predictive performance. Building on this foundation, a prediction model representing both trajectories and map geometry with polynomial representations achieves near state-of-the-art accuracy on standard benchmarks while substantially improving generalization under distribution shift. Extending this concept, a diffusion- based generative framework enables multi-agent scene generation, producing traffic continuations that are more plausible and kinematically consistent than those generated by conventional baselines. Evaluations on the Argoverse 2 and Waymo Open datasets confirm that polynomial representations reduce computational cost, enhance cross-dataset generalization, and yield smoother trajectories and higher behavioral plausibility. The findings reveal that standard in-distribution evaluation and regression-based metrics may fail to reflect true model generalization and prediction plausibility. By providing theoretical justification and empirical validation, this dissertation estab- lishes polynomial trajectory representations as an efficient, expressive, and generalizable foundation for traffic scene prediction in safety critical autonomous driving.",
      "short_abstract": "This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient models based on polynomial representations. While conventional sequence-based representations often struggle...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 11,
        "evidence": [
          "title: scene prediction",
          "abstract: scene prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.03330",
      "pdf_url": "https://arxiv.org/pdf/2608.03330.pdf"
    },
    {
      "id": "2608.03590",
      "anchor": "paper-2608-03590",
      "title": "Secure Long-Range Autonomous Valet Parking: A Reservation Scheme With Three-Factor Authentication and Key Agreement",
      "authors": [
        "Di Wang",
        "Yue Cao",
        "Fei Yan",
        "Yining Liu",
        "Daxin Tian",
        "Yuan Zhuang"
      ],
      "abstract": "Long-range autonomous valet parking (LAVP) is increasingly adopted to alleviate traffic congestion and parking difficulties. For large-scale parking demand, reservation can improve parking management. However, existing schemes mainly focus on parking request verification and parking check-in, and do not adequately protect identity legitimacy and communication security during passenger drop-off and pick-up. To address this problem, we propose SecLAVP, a provably secure three-factor authentication and key agreement protocol for LAVP reservation services. SecLAVP combines passwords, biometrics, and smart cards. With assistance from the drop-off/pick-up point (DP), the passenger and the autonomous vehicle (AV) achieve mutual authentication and establish a session key for secure communication. In the Real-Or-Random (ROR) model, we formally prove that SecLAVP provides session-key security. AVISPA simulations show that SecLAVP resists man-in-the-middle attacks, while informal analysis demonstrates that it satisfies 15 defined security goals. Finally, performance evaluation in terms of communication overhead, computational overhead, and scheduling shows that SecLAVP is feasible for practical deployment.",
      "short_abstract": "Long-range autonomous valet parking (LAVP) is increasingly adopted to alleviate traffic congestion and parking difficulties. For large-scale parking demand, reservation can improve parking management. However, existing schemes mainly focus on parking request...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 7,
        "evidence": [
          "abstract: security",
          "abstract: verification",
          "abstract: attack or threat"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.03590",
      "pdf_url": "https://arxiv.org/pdf/2608.03590.pdf"
    },
    {
      "id": "2608.03384",
      "anchor": "paper-2608-03384",
      "title": "Reusing Operational Evidence After Context Changes: A Conservative Bayesian Framework for Autonomous Vehicle Safety",
      "authors": [
        "Robab Aghazadeh Chakherlou",
        "Siddartha Khastgir",
        "Xingyu Zhao"
      ],
      "abstract": "Operational evidence, i.e., evidence of operation without failure is an important component of confidence in the safety or reliability of a system in service, but it is costly to collect. When the context of operation changes, the relevance of previously collected operational evidence becomes unclear. This problem arises, for example, when an autonomous vehicle that has operated safely in one operational environment (the Target Operational Domain, TOD) is deployed in a different but related TOD. Existing practice often treats such evidence in an all-or-nothing manner: either it is fully reused, or it is discarded. Neither position is satisfactory when there are good reasons to believe that the new context is no worse than the previous one, but that belief is itself uncertain. This paper studies how to make post-change reliability claims by combining evidence from two TODs using Conservative Bayesian Inference (CBI), which combines the evidence from the new context with a weighted amount of previous context and partial prior knowledge through constraints on a set of admissible priors. This yields conservative posterior bounds on quantities of interest. A numerical example illustrates how pre-existing evidence from a previous TOD can be transferred, conservatively and transparently, to support reliability claims in the changed TOD.",
      "short_abstract": "Operational evidence, i.e., evidence of operation without failure is an important component of confidence in the safety or reliability of a system in service, but it is costly to collect. When the context of operation changes, the relevance of previously...",
      "published": "2026-08-04",
      "updated": "2026-08-04",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 18,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 15,
      "arxiv_url": "https://arxiv.org/abs/2608.03384",
      "pdf_url": "https://arxiv.org/pdf/2608.03384.pdf"
    },
    {
      "id": "2608.02953",
      "anchor": "paper-2608-02953",
      "title": "RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models",
      "authors": [
        "Yuwei Ning",
        "Liangzhi Wang",
        "Yi Xiao",
        "Zhenhua Wu",
        "Yun Pang",
        "Mingkun Chang",
        "Jichang Li",
        "Guanbin Li"
      ],
      "abstract": "Realistic weather translation is valuable for developing and evaluating autonomous driving systems, yet collecting paired videos of the same scenes under different weather conditions at scale is impractical. Existing methods therefore rely on synthetic data, 3D weather editing, or geometry-conditioned generation, often compromising weather realism or scene fidelity. We propose RealWeather, a driving world model for both realistic and scene-faithful weather translation. Our key idea is to learn authentic weather dynamics directly from real-world videos. Specifically, RealWeather employs Progressive Realism Bootstrapping, an iterative data-refinement strategy. Assisted by an auxiliary Pseudo-Clear Generation pipeline, training initially starts with pseudo-style conditioning videos. As training proceeds, these inputs are progressively replaced with increasingly realistic videos generated by the model itself. This strategy bridges the pseudo-to-real domain gap, allowing the model to adapt seamlessly to real-world input distributions and naturally support bidirectional clear adverse translation. Furthermore, to strictly enforce structural integrity and suppress hallucinations, we introduce Scene-Fidelity RL Optimization, a reward-driven policy optimization strategy that explicitly penalizes alterations to safety-critical driving elements. Extensive experiments demonstrate that RealWeather significantly outperforms existing methods in visual realism and structural preservation, while enabling robust long-tail weather scenario generation and strong zero-shot out-of-distribution generalization. Our video demos can be found at https://hust-umi.github.io/RealWeather/.",
      "short_abstract": "Realistic weather translation is valuable for developing and evaluating autonomous driving systems, yet collecting paired videos of the same scenes under different weather conditions at scale is impractical. Existing methods therefore rely on synthetic data,...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Synthetic Data",
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 10,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.02953",
      "pdf_url": "https://arxiv.org/pdf/2608.02953.pdf"
    },
    {
      "id": "2608.02449",
      "anchor": "paper-2608-02449",
      "title": "MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving",
      "authors": [
        "Ambarish Govindarajulu Kaliamurthi",
        "Kaikai Liu"
      ],
      "abstract": "Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact model size and reliable metric grounding. We present MoRAL (Multimodal Reasoning for Autonomous Language Models), a two-stage fine-tuning pipeline that teaches Cosmos-Reason2-2B to first read a physics-encoded Bird's Eye View (BEV) representation and then reason over it for driving decisions. The BEV image encodes LiDAR metric distance as color bands, object class as cluster morphology, and radar Doppler velocity as directional wedge overlays, externalizing spatial perception into the input image so that no learned 3D backbone is required at inference. Stage 1 fine-tunes the vision encoder on 60,000 grounding records; zero-shot baselines produce no parseable BEV outputs, confirming the vocabulary requires explicit training. Stage 2 fine-tunes the full model (52M parameters, 2.4% of total) on 57,696 chain-of-thought records generated by Cosmos-Reason2-8B as teacher, spanning eight driving question types. On 2,304 held-out nuScenes frames evaluated by Gemma 4 (31B) calibrated against human review, MoRAL wins seven of eight question types over a zero-shot 8B baseline despite using four times fewer parameters, with the largest margins on question types requiring structured multi-step physics reasoning. Emergency braking recall improves from 10.8% to 47.8%, output degeneration falls from 94.1% to 20.8%, and the full pipeline fits a consumer 8 GB GPU at 42 tok/s without quantization. These results establish a reproducible foundation for compact, physics-grounded VLM reasoning on mobile edge platforms.",
      "short_abstract": "Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact model size and reliable metric grounding. We present MoRAL (Multimodal Reasoning for Autonomous Language...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 13,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: radar",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.02449",
      "pdf_url": "https://arxiv.org/pdf/2608.02449.pdf"
    },
    {
      "id": "2608.02200",
      "anchor": "paper-2608-02200",
      "title": "RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures",
      "authors": [
        "Cheng Li",
        "Renjun Gao",
        "Boyi Fu"
      ],
      "abstract": "Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous full-frame video, remain stable around transitional arm motion, and avoid over-trusting corrupted pose measurements. This study presents RSC-GestureNet, a reliability-aware selective causal recognizer, for Chinese traffic police gestures. The model treats pose confidence as a first-class signal: unreliable joints are down weighted during graph reasoning, temporal evidence is aggregated causally, and calibrated predictions are selectively emitted through a reliability-aware inference rule. We further introduce CTPGesture-C, a reproducible feature-level corruption benchmark with seven pose/RGB degradation families, and an RGB-level diagnostic in which corrupted frames are reprocessed by MediaPipe before recognition. On the complete official CTPGesture v1 split (134,424 labeled frames and 33,451 causal windows), RSC-GestureNet achieves 93.33+-0.24% accuracy, 91.71+-0.27% macro-F1, 91.69+-0.29% online macro-F1, 98.80+-0.07% Early@10, 0.153+-0.013 s TTC, and the best robust macro-F1 among evaluated methods. Under the same split and causal protocol, it exceeds reproduced traffic-specific MD-GCN and HLP-GCN baselines by 3.23-4.11 macro-F1 points and 2.15-3.07 online-F1 points. These results, together with calibration, selective-risk, statistical, adaptive-branching, and image-level re-extraction analyses, indicate that explicit pose-reliability modeling improves early, stable, and robust traffic-command recognition.",
      "short_abstract": "Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous full-frame video, remain stable around transitional arm motion, and avoid over-trusting corrupted pose...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.02200",
      "pdf_url": "https://arxiv.org/pdf/2608.02200.pdf"
    },
    {
      "id": "2608.02191",
      "anchor": "paper-2608-02191",
      "title": "DerainSplat: Feed-Forward Clean 3D Gaussian Splatting from Sparse Rainy Views",
      "authors": [
        "Fuzhen Jiang",
        "Changyue Shi",
        "Chuxiao Yang",
        "Xinyuan Hu",
        "Wenjie Ye",
        "Minghao Chen"
      ],
      "abstract": "Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such as embodied AI and autonomous driving continue to emerge, reconstructing clean 3D scenes from sparse rainy views in a feed-forward manner becomes increasingly important. Existing feed-forward 3D Gaussian Splatting (3DGS) methods often assume clean inputs and collapse under rainy conditions. To this end, we present \\textbf{\\textit{DerainSplat}}, a feed-forward framework that reconstructs clean 3D scenes from only a few rainy views. To support this task, we build a large-scale multi-view derain dataset through a four-stage synthesis pipeline that sequentially models overcast illumination, depth-dependent haze, rain streaks, and lens raindrops, producing privileged weather factors. We introduce a weather net that predicts the weather factors from rainy context and yields two support maps. Scene support modulates cross-view cost-volume matching, while radiance support drives depth-aligned appearance fusion to fill corrupted pixels. The derived geometry evidence further attenuates Gaussian opacity to reduce spurious structures. A rainy cycle consistency re-renders clean views using the predicted factors and aligns them with rainy inputs. Extensive experiments show that \\textbf{\\textit{DerainSplat}} outperforms existing methods on various datasets, including RealEstate10K, ACID, Mip-NeRF360, and real-world rainy scenes, with strong cross-dataset generalization.",
      "short_abstract": "Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such as embodied AI and autonomous driving continue to emerge, reconstructing clean 3D scenes from sparse rainy...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.02191",
      "pdf_url": "https://arxiv.org/pdf/2608.02191.pdf"
    },
    {
      "id": "2608.02177",
      "anchor": "paper-2608-02177",
      "title": "GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes",
      "authors": [
        "Fanyu Wang",
        "Longgao Zhang",
        "Junyi Chen"
      ],
      "abstract": "Existing rainfall simulation methods for autonomous driving remain limited in physical controllability and multi-view consistency. This paper presents GSRAIN, a high-/low-frequency rainfall synthesis method for 3D Gaussian Splatting (3DGS) driving scenes. GSRAIN constructs a high-frequency raindrop model from measured rainfall data and generates low-frequency rainy appearance using a geometry-aware single-step diffusion model. The two effects are then fused in a unified 3DGS scene, enabling rainfall-intensity control over the range of 0--13~mm/h. The proposed method achieves a Fréchet Inception Distance (FID) of 149.09, outperforming CycleGAN-Turbo (155.71) and WeatherEdit (157.94). Object-detection and closed-loop driving experiments further show that the generated scenes expose scene-dependent performance changes of the evaluated algorithms under controllable rainfall. These results indicate that GSRAIN provides an effective approach for constructing physically controllable, repeatable, and closed-loop-compatible rainy-weather test scenes for autonomous driving.",
      "short_abstract": "Existing rainfall simulation methods for autonomous driving remain limited in physical controllability and multi-view consistency. This paper presents GSRAIN, a high-/low-frequency rainfall synthesis method for 3D Gaussian Splatting (3DGS) driving scenes....",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.02177",
      "pdf_url": "https://arxiv.org/pdf/2608.02177.pdf"
    },
    {
      "id": "2608.01998",
      "anchor": "paper-2608-01998",
      "title": "TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving",
      "authors": [
        "Mengmeng Zhu",
        "Yuxuan Sun",
        "Wei Chen",
        "Bo Ai"
      ],
      "abstract": "The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and reasoning capabilities. Infrastructure-assisted AD alleviates this resource constraint by enabling collaboration with large VLMs (LVLMs) at edge servers. However, in dynamic vehicular environments, the utility of sensory data for downstream tasks decays rapidly, making timeliness of information a critical concern. To balance the accuracy gains of LVLMs with their latency-induced timeliness degradation, we develop a Timeliness-Aware Large-Small VLM Collaboration (TALSC) framework. Specifically, we first model the Age of Information (AoI) evolution for VLM inference and characterize the coupling among AoI, token length, and task performance to formulate a general timeliness metric. Building on this, we propose the TALSC online scheduling algorithm. Since scheduling decisions have a delayed impact on future timeliness metric and the output token number is unknown at scheduling time, we design a Lyapunov drift-plus-estimated-penalty algorithm and provides a guaranteed performance. In simulation, we first conduct a case study to derive a fitted timeliness metric based on nuScenes dataset, and further show that TALSC outperforms baselines under various communication and computing settings, achieving up to a 12.6\\% normalized improvement in Micro-F1 score compared with the best-performing baseline.",
      "short_abstract": "The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and reasoning capabilities. Infrastructure-assisted AD alleviates...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 16,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints",
          "title: roadside infrastructure",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.01998",
      "pdf_url": "https://arxiv.org/pdf/2608.01998.pdf"
    },
    {
      "id": "2608.01761",
      "anchor": "paper-2608-01761",
      "title": "DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing",
      "authors": [
        "Siying Li",
        "Ying Ni",
        "Jie Sun",
        "Jian Sun",
        "Haotian Shi"
      ],
      "abstract": "End-to-end (E2E) autonomous driving algorithms require rigorous closed-loop validation in simulation environments offering high visual fidelity, strong interactivity, and real-time performance. Existing approaches, from game engines to static neural rendering, inherently trade off these requirements and struggle with the dynamic scene composition essential for E2E testing. To bridge this gap, we propose a novel decoupled 3D Gaussian Splatting (3DGS) framework tailored for large-scale E2E evaluation. We fundamentally decompose scenes into a high-fidelity static background and manipulable dynamic agents using an object-centric canonical representation. To resolve resulting representational conflicts, we introduce three targeted modules: (1) asset compression via perceptual pruning and vector quantization for real-time traffic rendering; (2) map-guided geometric registration leveraging semantic topology to strictly align trajectories; and (3) proxy-based relighting transferring ambient illumination for seamless photometric integration. Extensive experiments demonstrate that DecoupleGS achieves a balanced fidelity-efficiency trade-off, improves metric and photometric consistency, and provides a practical closed-loop sensor simulation platform for E2E autonomous driving evaluation.",
      "short_abstract": "End-to-end (E2E) autonomous driving algorithms require rigorous closed-loop validation in simulation environments offering high visual fidelity, strong interactivity, and real-time performance. Existing approaches, from game engines to static neural...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 18,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.01761",
      "pdf_url": "https://arxiv.org/pdf/2608.01761.pdf"
    },
    {
      "id": "2608.01755",
      "anchor": "paper-2608-01755",
      "title": "Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs",
      "authors": [
        "Zixuan Huang",
        "Yang Zhou",
        "Kaixuan Wang",
        "Guli Zhang",
        "Hongyan Xie",
        "Yakun Zhu",
        "Hao Geng",
        "Xiaozhi Chen",
        "Yikun Ban",
        "Deqing Wang"
      ],
      "abstract": "Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.",
      "short_abstract": "Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 4,
        "evidence": [
          "abstract: planning",
          "abstract: decision-making",
          "abstract: trajectory generation"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.01755",
      "pdf_url": "https://arxiv.org/pdf/2608.01755.pdf"
    },
    {
      "id": "2608.02308",
      "anchor": "paper-2608-02308",
      "title": "A General Set-Based Framework for Cognitive State Estimation: Theory and Application to Conditionally Automated Driving",
      "authors": [
        "Sibibalan Jeevanandam",
        "Neera Jain"
      ],
      "abstract": "We present a set-based framework for estimating human cognitive states in human-automation interaction (HAI) contexts. Unlike probabilistic approaches dominant in the HAI literature, our framework treats process and measurement uncertainties as unknown but bounded, avoiding the need for large structured datasets or distributional assumptions on noise. We demonstrate the framework in the context of conditionally automated (SAE Level 3) driving, where we estimate three cognitive states (trust, perceived risk, and workload) that influence human reliance on the automation during a continuous, non-trial-based interaction. We leverage a hybrid dynamical modeling framework to identify individual-specific process and measurement models, systematically estimate noise bounds by enforcing reachset conformance, and identify the subset of cognitive states that influence each individual's reliance on the automation. Set-valued estimates of those states are then produced by fusing binary reliance observations and intermittent, quantized self-reports. The framework is evaluated through an in-person experiment in a medium-fidelity driving simulator with 20 participants. The set-valued estimator achieves at least 75% consistency for most participants during testing, and multi-step-ahead reliance predictions derived from the estimates outperform both an open-loop baseline and a particle filter across all choices of prediction horizons (15, 30, 45, 60 time steps), with the performance gap more evident at longer horizons. The proposed estimation framework can enable automation systems that are continuously aware of, and responsive to, the human driver's state.",
      "short_abstract": "We present a set-based framework for estimating human cognitive states in human-automation interaction (HAI) contexts. Unlike probabilistic approaches dominant in the HAI literature, our framework treats process and measurement uncertainties as unknown but...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing",
          "Safety, Security & Verification: abstract: uncertainty",
          "Human Factors & Policy: abstract: human driver"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.02308",
      "pdf_url": "https://arxiv.org/pdf/2608.02308.pdf"
    },
    {
      "id": "2608.02289",
      "anchor": "paper-2608-02289",
      "title": "Extended Field of View Analysis for VideoGAN-based Trajectory Generation",
      "authors": [
        "Annajoyce Mariani",
        "Kira Maag",
        "Hanno Gottschalk"
      ],
      "abstract": "Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation. While rule-based and classical learning-based methods may struggle to capture the complexity of traffic behavior, generative models have already demonstrated in other fields that they can handle a comparable level of complexity. In this paper, we build upon previous work on generative adversarial network (GAN)-based semantic bird's-eye-view traffic generation and extend the proposed framework in several key aspects. We improve the semantic representation, replace the trajectory extraction procedure with a graph-based association method, and systematically investigate increasingly larger fields of view. In addition, we introduce a quantitative evaluation framework to assess hallucinations and object permanence in generated videos. Our experiments demonstrate that the framework generalizes to larger and more complex traffic scenes while maintaining statistically realistic trajectories and coherent spatial relationships between traffic participants. Within 150GPU hours of training and with inference times below 20ms for scenes of up to 20s, our results demonstrate that video-based GANs remain an efficient and scalable approach for realistic trajectory generation, even in substantially larger traffic scenes, making them well suited for downstream tasks such as prediction, planning, and simulation in automated driving.",
      "short_abstract": "Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation. While rule-based and classical learning-based methods may struggle to capture the complexity of traffic behavior, generative models have already...",
      "published": "2026-08-03",
      "updated": "2026-08-03",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 11,
        "evidence": [
          "abstract: planning",
          "title: trajectory generation",
          "abstract: trajectory generation"
        ]
      },
      "recency": "recent",
      "age_days": 16,
      "arxiv_url": "https://arxiv.org/abs/2608.02289",
      "pdf_url": "https://arxiv.org/pdf/2608.02289.pdf"
    },
    {
      "id": "2608.01572",
      "anchor": "paper-2608-01572",
      "title": "Enhancing Visual Perception in Foggy Conditions via Multiclass Fog Density Modeling",
      "authors": [
        "Mohamad Mofeed Chaar",
        "Galia Weidl"
      ],
      "abstract": "Autonomous driving (AD) systems have advanced rapidly over the past decade; however, robust perception under adverse weather conditions remains a major challenge, particularly in dense fog. In this work, we investigate fog-aware perception using synthetically generated fog data derived from the Waymo dataset. To support fog simulation, depth images are generated using an iterative learning approach. We consider five fog-density levels: clear, light fog, moderate fog, heavy fog, and very heavy fog. Instead of training a single unified model across all conditions, we train separate perception models for each fog-density level. Experimental results show that density-specific training improves performance in severe fog conditions. In particular, for the very heavy fog class, recall improves from 0.076 to 0.232, corresponding to an absolute gain of 15.6 percentage points. These findings suggest that deploying multiple specialized models, rather than a single general-purpose model, can improve perception robustness for autonomous vehicles under challenging visibility conditions. Future work will extend this strategy to additional sensing modalities, including LiDAR and radar, and evaluate generalization across diverse weather scenarios.",
      "short_abstract": "Autonomous driving (AD) systems have advanced rapidly over the past decade; however, robust perception under adverse weather conditions remains a major challenge, particularly in dense fog. In this work, we investigate fog-aware perception using synthetically...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 16,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: radar"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01572",
      "pdf_url": "https://arxiv.org/pdf/2608.01572.pdf"
    },
    {
      "id": "2608.01535",
      "anchor": "paper-2608-01535",
      "title": "STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision",
      "authors": [
        "Pou-Chun Kung",
        "Aryaman Rao",
        "Utkrisht Sahai",
        "Hemanth Murali",
        "Yi Liu",
        "Rui-Yu Lin",
        "Katherine A. Skinner"
      ],
      "abstract": "Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autonomous driving. However, existing approaches for improving spatiotemporal reasoning in VLMs often rely on complex preprocessing pipelines, expensive human annotations, or synthetic data, which limit scalability and introduce potential sim-to-real gaps. Moreover, although these methods have improved spatiotemporal understanding, they still lack strong metric reasoning capabilities for dynamic scenes, such as estimating object motion in real-world units. Prior work has explored LiDAR-based metric depth supervision to enhance spatial perception, but it does not directly address temporal reasoning. We introduce STAR-VLM, an automotive radar-supervised framework that enhances spatiotemporal VLMs with motion reasoning and metric velocity estimation for autonomous driving. Automotive radar is a low-cost and widely deployed sensor that provides complementary spatiotemporal supervision through range and Doppler measurements. By leveraging these measurements as label-free ground truth during training, STAR-VLM improves the metric spatiotemporal reasoning ability of VLMs. Through experiments on driving scenarios, we show that STAR-VLM achieves state-of-the-art performance on both motion classification and metric velocity estimation, outperforming even task-specific methods designed for each task. These results highlight automotive radar as a scalable and cost-effective source of supervision for building metric-aware spatiotemporal VLMs for real-world autonomous driving.",
      "short_abstract": "Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autonomous driving. However, existing approaches for improving spatiotemporal reasoning in VLMs often rely on...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01535",
      "pdf_url": "https://arxiv.org/pdf/2608.01535.pdf"
    },
    {
      "id": "2608.01338",
      "anchor": "paper-2608-01338",
      "title": "Driver2Map: Imitating Human Driving for Online High-Definition Map Construction",
      "authors": [
        "Pan Yin",
        "Runtian Xia",
        "Weisong Kuang",
        "Kaiyu Li",
        "Cong Zhao",
        "Xiangyong Cao"
      ],
      "abstract": "High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images provide crucial information. However, due to the modality and perspective differences among these data sources, existing methods often struggle to effectively align and fuse them, making online HD map construction still challenging. To address these issues, we propose Driver2Map, an online HD map construction model inspired by human drivers. Unlike existing HD map construction models that utilize only two modalities, our Driver2Map can simultaneously exploit three modalities. Specifically, we propose a \"two-stage alignment\" strategy to reduce spatial misalignment across different modalities. Additionally, we introduce \"Pose-Guided BEV Fusion\", a BEV (bird's-eye-view) generation module that leverages camera pose information to adaptively weight multi-view features, thereby effectively suppressing cross-view feature overlap during BEV generation. Also, we design a \"Pretrained Prior for Map Refinement\" module to refine the initial prediction by learning map structure priors, thus improving the HD map prediction under dynamic occlusions. Extensive experiments demonstrate that Driver2Map outperforms existing methods on both IoU and AP metrics.",
      "short_abstract": "High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images provide crucial information. However, due to the modality and perspective...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 32,
        "evidence": [
          "title: HD map",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01338",
      "pdf_url": "https://arxiv.org/pdf/2608.01338.pdf"
    },
    {
      "id": "2608.01336",
      "anchor": "paper-2608-01336",
      "title": "Asleep at the Wheel: JEPA's Limitations in Evaluating Novel Driving Data",
      "authors": [
        "Advait Pavuluri",
        "Shamik Karkhanis",
        "Uzma Mushtaque"
      ],
      "abstract": "Modern autonomous-driving fleets record far more video than human reviewers can inspect. This motivates the need for an automatic clip triage mechanism, to surface rare and review-worthy clips, so that driving models can be fine-tuned to better handle unideal circumstances. We test a label-free approach that scores clips by the prediction-error \"novelty\" of a self-supervised joint-embedding predictive architecture (JEPA); a frozen V-JEPA video encoder is paired with a lightweight predictor head to reconstruct masked clip embeddings, and clips whose embeddings are hard to predict are flagged as interesting. Evaluated under a realistic protocol that trains on one dataset and tests against footage from others, this approach appears highly effective. We show that this apparent success is actually a domain-shift consequence: on a fair benchmark drawn from a single dataset, this mechanism collapses to chance and is on par with simple no-training baselines. A lightly supervised probe on the same frozen embeddings results in almost double the average precision, indicating that the bottleneck is indeed the self-supervised objective, rather than the representation. We present this as a study for evaluating the effectiveness of self-supervised learning, where cross-dataset protocols can silently reward domain separation over novelty.",
      "short_abstract": "Modern autonomous-driving fleets record far more video than human reviewers can inspect. This motivates the need for an automatic clip triage mechanism, to surface rare and review-worthy clips, so that driving models can be fine-tuned to better handle unideal...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01336",
      "pdf_url": "https://arxiv.org/pdf/2608.01336.pdf"
    },
    {
      "id": "2608.01201",
      "anchor": "paper-2608-01201",
      "title": "PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning",
      "authors": [
        "Volodymyr Havrylov",
        "Faris Janjoš",
        "Andreas Look",
        "Jürgen Mathes",
        "Andreas Geiger"
      ],
      "abstract": "End-to-end autonomous driving (E2E AD) systems integrate perception, prediction, and planning into a single differentiable architecture. While these models show great promise, their standard training often relies on output-only supervision, which can lead to weak gradients for the hidden layers of increasingly complex models. Recent works have integrated vision-language model (VLM) supervision for latent features to address this, yielding substantial empirical gains, yet leaving the underlying theoretical mechanisms poorly understood. Our investigation into this methodology reveals that the resulting performance gains stem not from VLM reasoning capabilities, as previously assumed, but rather from the latent connections forged between the E2E AD model and ground-truth (GT) data during training. Building on this insight, we propose a probabilistic deep supervision framework that regularizes intermediate latent representations directly from GT data. By treating model latents as reparameterizable distributions, we optimize the architecture via the Evidence Lower Bound (ELBO). Our evaluations conducted on the nuScenes dataset demonstrate that supervising trajectory-related latents with future GT paths consistently improves planning performance. Using identical training data and E2E architectures, our method achieves an 8% reduction in planning L2 error and a 3% decrease in collision rates compared to competitive vectorized baselines, all while incurring negligible computational overhead.",
      "short_abstract": "End-to-end autonomous driving (E2E AD) systems integrate perception, prediction, and planning into a single differentiable architecture. While these models show great promise, their standard training often relies on output-only supervision, which can lead to...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 9,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01201",
      "pdf_url": "https://arxiv.org/pdf/2608.01201.pdf"
    },
    {
      "id": "2608.01133",
      "anchor": "paper-2608-01133",
      "title": "Policy Optimality Measurement for Multi-Vehicle Decision-Making: From Extrinsic Indicators to Intrinsic Quality",
      "authors": [
        "Ye Han",
        "Lijun Zhang",
        "Dejian Meng"
      ],
      "abstract": "Evaluating Multi-Agent Reinforcement Learning (MARL) policies in autonomous driving fundamentally relies on extrinsic statistical indicators (e.g., reward curves and success rates), which often mask intrinsic policy degradation and algorithmic blind spots. To break this black-box evaluation, this letter proposes a novel information-theoretic diagnostic framework. By leveraging a fully converged Monte Carlo Tree Search (MCTS) as an asymptotic oracle, we establish a theoretical ground-truth baseline distribution. We formulate a bounded policy optimality score ($\\mathcal{M}_{opt}$) using the forward KL divergence to rigorously penalize fatal collaborative omissions. Crucially, we semantically decouple this metric into lateral and longitudinal dimensions, creating a granular \"semantic microscope\". Extensive spatial and temporal diagnostics on state-of-the-art MARL architectures and exploration mechanisms demonstrate that our framework conclusively exposes hidden directional biases, identifies temporal average-policy traps, and transforms heuristic hyperparameter tuning into a visually trackable trajectory optimization. This framework establishes a rigorous, model-agnostic standard for benchmarking intrinsic multi-agent policy quality.",
      "short_abstract": "Evaluating Multi-Agent Reinforcement Learning (MARL) policies in autonomous driving fundamentally relies on extrinsic statistical indicators (e.g., reward curves and success rates), which often mask intrinsic policy degradation and algorithmic blind spots. To...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: law or policy",
          "abstract: law or policy"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01133",
      "pdf_url": "https://arxiv.org/pdf/2608.01133.pdf"
    },
    {
      "id": "2608.01035",
      "anchor": "paper-2608-01035",
      "title": "WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA",
      "authors": [
        "Zhihao Zhu",
        "Hanlin Shang",
        "Mingwang Xu",
        "Feipeng Cai",
        "Zhuolin He",
        "Yaoyi Li",
        "Jianhua Han",
        "Hang Xu",
        "Siyu Zhu"
      ],
      "abstract": "Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential autoregressive decoding. Conversely, while specialized diffusion policies enable low-latency, parallel execution, training them from scratch typically yields narrow, single-task architectures that lack holistic visual-linguistic reasoning. Successfully transforming pre-trained autoregressive generalists into parallel diffusion models could combine multi-task cognitive intelligence with execution efficiency, yet this transition presents a formidable architectural challenge due to mismatched attention patterns (causal versus bidirectional) and divergent optimization objectives. To bridge this divide, we introduce WAM-Diff2, a multi-task discrete diffusion VLA framework powered by a three-stage hierarchical distillation strategy. By structuring the architectural shift through progressive block-wise adaptation, block-wise distillation, and model-wise cross-scale distillation, WAM-Diff2 preserves the underlying semantic foundations of the base model while accelerating inference. Extensive evaluations across driving understanding, perception, and planning benchmarks demonstrate that WAM-Diff2 effectively mitigates exposure bias and achieves performance parity with autoregressive baselines. Crucially, the autoregressive-to-diffusion transition yields a 2.8x decoding speedup, which scales to an ultimate 15.1x acceleration when combined with system-level optimizations including FlashInfer and CUDA Graphs.",
      "short_abstract": "Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential autoregressive...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 28,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01035",
      "pdf_url": "https://arxiv.org/pdf/2608.01035.pdf"
    },
    {
      "id": "2608.01448",
      "anchor": "paper-2608-01448",
      "title": "Latent Two-Sample Testing for Fair Autonomous Vehicle Road Evaluation",
      "authors": [
        "Qiujing Lu",
        "Xuanhan Wang",
        "Guanghong Jia",
        "Zhenxia Zhu",
        "Huijuan Lin",
        "Kehua Sheng",
        "Zhichao Hou",
        "Shuo Feng"
      ],
      "abstract": "With the rapid advancement of autonomous vehicle (AV) systems, fast and reliable iteration through road testing has become increasingly critical. However, changes in testing environments make it difficult to disentangle true performance differences between AV versions from extraneous environmental variations, undermining fair and reliable evaluation. This challenge is further compounded by the high-dimensional and unstructured nature of large-scale road testing data, for which effective analysis and comparison methods remain limited. In this work, we address these challenges by introducing a principled framework for distribution-aware AV evaluation. We first learn structured latent representations that map high-dimensional, unstructured road testing data into a compact latent space, enabling effective characterization of scenario distributions. Building on this representation, we propose a two-stage two-sample testing framework that (i) detects and localizes distributional shifts between testing datasets and (ii) calibrates these shifts via importance sampling to reduce evaluation bias and metric estimation error. Experiments on synthetic and real-world road testing data demonstrate the effectiveness of the proposed method.",
      "short_abstract": "With the rapid advancement of autonomous vehicle (AV) systems, fast and reliable iteration through road testing has become increasingly critical. However, changes in testing environments make it difficult to disentangle true performance differences between AV...",
      "published": "2026-08-02",
      "updated": "2026-08-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "recent",
      "age_days": 17,
      "arxiv_url": "https://arxiv.org/abs/2608.01448",
      "pdf_url": "https://arxiv.org/pdf/2608.01448.pdf"
    },
    {
      "id": "2608.00878",
      "anchor": "paper-2608-00878",
      "title": "Goal-Oriented Logic-based Semantic Communication for Neuro-Symbolic Reasoning with Applications onto Autonomous Driving",
      "authors": [
        "Ahmet Faruk Saz",
        "Duo Xu",
        "Faramarz Fekri"
      ],
      "abstract": "We consider First-Order Logic (FOL)-based semantic communication for neuro-symbolic decision-making in collaborative environments such as autonomous driving networks. Each connected autonomous vehicle (CAV) converts its partial sensor observations into a natural-language scene description and corresponding grounded FOL evidence. Under an uplink budget, a semantic encoder at each car selects the observations most informative for evaluating traffic rules and transmit to a Road Side Unit (RSU). The RSU fuses all received evidence, evaluates collaborative rules, performs logical deduction for vehicle-specific safety and right-of-way information for constrained downlink transmission. Each CAV combines the received deductions with its local description, enabling a local LLM agent to select a high-level driving action. We develop a principled, verifiable semantic communication method using a random-support Dirichlet--Categorical model of inductive logical probability, providing a modern statistical reinterpretation of Carnap's and Hintikka's systems. From this model, we derive a goal-oriented semantic information-bottleneck formulation that prioritizes evidence transmission by its reduction of uncertainty over task goals. Using 152 traffic rules extracted from the California Driver Handbook, we evaluate the framework on MDrive simulator in CARLA. Under identical communication budgets, semantic evidence selection completes every scenario without safety hazards, whereas uniform evidence selection produces collisions, showcasing semantic communication's superiority.",
      "short_abstract": "We consider First-Order Logic (FOL)-based semantic communication for neuro-symbolic decision-making in collaborative environments such as autonomous driving networks. Each connected autonomous vehicle (CAV) converts its partial sensor observations into a...",
      "published": "2026-08-01",
      "updated": "2026-08-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 9,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 18,
      "arxiv_url": "https://arxiv.org/abs/2608.00878",
      "pdf_url": "https://arxiv.org/pdf/2608.00878.pdf"
    },
    {
      "id": "2608.00779",
      "anchor": "paper-2608-00779",
      "title": "SIPTraj: Map-Free End-to-End Trajectory Prediction via Physics-Guided Scene Interaction",
      "authors": [
        "Feifei Liu",
        "Zejun Wei",
        "Haozhe Wang",
        "Yazhi Ye",
        "Yuying Zhang",
        "Jintao Cheng",
        "Chi Man Vong",
        "Xieyuanli Chen",
        "Xiaoyu Tang"
      ],
      "abstract": "Trajectory prediction of surrounding agents is a prerequisite for safe planning and decision making in autonomous driving. Without high-definition (HD) maps, sensor-derived bird's-eye-view (BEV) features provide no explicit lane topology or drivable-area priors, making it inherently difficult to ground each agent in its surrounding scene context. Moreover, physical feasibility remains difficult to capture through data-driven learning alone, as kinematic constraints on agent motion cannot be explicitly encoded without structured supervision. Existing map-free predictors extract scene context in an agent-agnostic manner through a single fusion step and treat physical constraints only as output-level penalties, leaving both challenges unaddressed. We propose SIPTraj, a map-free trajectory prediction framework that jointly addresses scene grounding and physical feasibility. SIPTraj introduces a Hierarchical Agent-Scene Encoder (HASE) progressively grounding each agent in agent-guided scene evidence and refining inter-agent relations within the scene-grounded space. To tackle physical infeasibility in predicted trajectories, we develop a Physics-Guided Iterative Decoder (PGID). It conditions decoding on instantaneous kinematic states, propagating physical supervision into internal representations rather than output trajectories alone. Extensive experiments on nuScenes and Argoverse 2 Sensor show that SIPTraj surpasses prior map-free predictors and strong map-based baselines without any HD map at inference. Our code will be released as open-source.",
      "short_abstract": "Trajectory prediction of surrounding agents is a prerequisite for safe planning and decision making in autonomous driving. Without high-definition (HD) maps, sensor-derived bird's-eye-view (BEV) features provide no explicit lane topology or drivable-area...",
      "published": "2026-08-01",
      "updated": "2026-08-01",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 19,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 18,
      "arxiv_url": "https://arxiv.org/abs/2608.00779",
      "pdf_url": "https://arxiv.org/pdf/2608.00779.pdf"
    },
    {
      "id": "2608.00690",
      "anchor": "paper-2608-00690",
      "title": "LLM-Assisted Coalition Formation for Cooperative Perception in Autonomous Driving",
      "authors": [
        "Ahmad Sarlak",
        "Hao Wang",
        "Rahul Amin",
        "Abolfazl Razi"
      ],
      "abstract": "Cooperative perception (CP) enables connected autonomous vehicles (CAVs) to share complementary observations for safer navigation, but practical deployment is limited by bandwidth constraints, unreliable links, and redundant information exchange. Existing CP methods often assume predefined participants and merely focus on collective perception. Likewise, recent LLM-based cooperative driving frameworks facilitate multi-vehicle reasoning but do not regulate participation criteria to select more beneficial vehicles. To bridge this gap, we propose an LLM-assisted coalition formation framework that selects the most informative helper vehicles before LLM reasoning. The approach jointly optimizes perceptual diversity using a determinantal point process (DPP) over multimodal vehicle embeddings and communication-aware reliability. This leads to a joint coalition selection and power allocation problem, which we solve efficiently via a relaxed convex reformulation and an ADMM-based optimization strategy that decouples diversity-aware selection from network-aware resource allocation. The selected coalition is then summarized and provided with an LLM reasoning module for efficient and less redundant multi-vehicle decision support. Experimental results show that our approach outperforms other baselines in overall coalition value, while maintaining high diversity and improved networking efficiency. The framework achieves a better balance between task performance and safety across OPV2V and V2V4Real datasets, demonstrating its effectiveness for cooperative autonomous driving with communication constraints.",
      "short_abstract": "Cooperative perception (CP) enables connected autonomous vehicles (CAVs) to share complementary observations for safer navigation, but practical deployment is limited by bandwidth constraints, unreliable links, and redundant information exchange. Existing CP...",
      "published": "2026-08-01",
      "updated": "2026-08-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 11,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 18,
      "arxiv_url": "https://arxiv.org/abs/2608.00690",
      "pdf_url": "https://arxiv.org/pdf/2608.00690.pdf"
    },
    {
      "id": "2608.00625",
      "anchor": "paper-2608-00625",
      "title": "Learning-Based Motion Planning for Dynamic Environments: From Foundational Algorithms to Emerging Paradigms",
      "authors": [
        "Zongyuan Shen",
        "Shalabh Gupta",
        "Shancheng Zhao",
        "Dehua Zhou",
        "Gao Wang",
        "Rui Cheng",
        "Yaming Ou",
        "Zhongqiang Ren",
        "Yikui Zhai",
        "C. L. Philip Chen"
      ],
      "abstract": "Motion planning in dynamic environments is a fundamental problem in robotics, aiming to generate safe and efficient paths, trajectories, or control actions in the presence of moving obstacles, uncertain predictions, and multi-agent interactions. It has broad applications in autonomous driving, service robotics, warehouse logistics, human-robot collaboration, crowd navigation, and multi-robot systems. This survey reviews representative works published primarily between 2015 and 2025, with a particular focus on how recent learning-based advances extend, complement, or interact with classical planning foundations. We first revisit classical planning methods as algorithmic foundations and reference frameworks for learning-based extensions. We then propose a role-of-learning taxonomy that categorizes existing methods according to how learning participates in the planning pipeline, including direct policy learning, learning-augmented classical planning, hybrid planning, and training enhancement methods. For each category, we summarize the main problem settings, representative algorithms, key ideas, integration mechanisms, strengths, and limitations. We further analyze how observation representations, prediction uncertainty, interaction modeling, planner integration, safety constraints, and training strategies shape learning-based motion planning in dynamic environments. Finally, we discuss open challenges and future directions, including sim-to-real gap, safe and certifiable planning, dense crowd navigation, perception-planning coupling, and embodied AI.",
      "short_abstract": "Motion planning in dynamic environments is a fundamental problem in robotics, aiming to generate safe and efficient paths, trajectories, or control actions in the presence of moving obstacles, uncertain predictions, and multi-agent interactions. It has broad...",
      "published": "2026-08-01",
      "updated": "2026-08-01",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Survey"
      ],
      "classification": {
        "confidence": "high",
        "score": 34,
        "margin": 28,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: navigation"
        ]
      },
      "recency": "recent",
      "age_days": 18,
      "arxiv_url": "https://arxiv.org/abs/2608.00625",
      "pdf_url": "https://arxiv.org/pdf/2608.00625.pdf"
    },
    {
      "id": "2608.00524",
      "anchor": "paper-2608-00524",
      "title": "Cooperative Platoon Routing and Dispatching via Edge-Assisted Hybrid Quantum Optimization",
      "authors": [
        "Talha Azfar",
        "Ruimin Ke"
      ],
      "abstract": "Cooperative platooning can reduce the energy use of Connected and Autonomous Vehicle (CAV) fleets, but the routing problem becomes difficult when vehicles must meet on the same road segments at compatible times while moving through unstable urban traffic. This paper develops an edge-assisted, closed-loop evaluation pipeline for platooning-aware vehicle routing. Roadside Units estimate local traffic kinematics from video, classify segment-level flow stability, and activate platooning rewards only on road segments where close-gap coordination is physically appropriate. The resulting multi-vehicle routing problem is written directly as a Quadratic Unconstrained Binary Optimization (QUBO) model, so pairwise platooning interactions are represented as native quadratic Ising terms instead of requiring auxiliary MILP linearization variables. We evaluate the framework using a 24-hour microscopic SUMO simulation of Troy, NY, together with localized IBM Quantum hardware benchmarks. The SUMO study shows an $18.5\\%$ reduction in fleet tractive-energy demand relative to a non-cooperative baseline. On 25-active-qubit benchmark instances executed on $\\texttt{ibm_boston}$, Linear-Chain QAOA reduces two-qubit CNOT depth by $66.7\\%$ compared with dense QAOA and samples the exact classical ground state with $P_{\\text{feas}} = 38.6\\%$ and $P_{\\text{opt}} = 14.2\\%$ at $p=2$. These results suggest that edge perception and shallow quantum optimization can work together as a useful component of closed-loop CAV platoon dispatching.",
      "short_abstract": "Cooperative platooning can reduce the energy use of Connected and Autonomous Vehicle (CAV) fleets, but the routing problem becomes difficult when vehicles must meet on the same road segments at compatible times while moving through unstable urban traffic....",
      "published": "2026-08-01",
      "updated": "2026-08-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 19,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: hardware",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "recent",
      "age_days": 18,
      "arxiv_url": "https://arxiv.org/abs/2608.00524",
      "pdf_url": "https://arxiv.org/pdf/2608.00524.pdf"
    },
    {
      "id": "2608.00298",
      "anchor": "paper-2608-00298",
      "title": "WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation",
      "authors": [
        "Jianxun Cui",
        "Ping Wu",
        "Stanisa Peric",
        "Marko Milojkovic",
        "Vladan Devedzic"
      ],
      "abstract": "World models and generative simulators are emerging as interactive testing infrastructure for autonomous driving because they can react to the ego planner and produce counterfactual, rare, and safety-critical rollouts. This changes a test scenario from a fixed replayed trajectory into an interactive scenario family whose realized evolution depends on the planner under test. The unresolved question is therefore not only whether dangerous rollouts can be generated, but what valid closed-loop evidence is enough to support a specified testing intent and stopping decision. This paper formulates interactive world-model-style testing adequacy and introduces WM-Cov, a provider-agnostic evaluation layer that converts raw provider outputs into requested, realized, and valid evidence. WM-Cov reports adequacy through coverage growth, valid-failure discovery, failure-mode diversity, realism, artifact suppression, duplicate accounting, and valid-evidence precision. Studies on executed TeraSim/SUMO events, WM-like mixed trace pools, and a real DriveArena TrafficManager--WorldDreamer matrix show that dangerous-looking events can include valid ADS failures, duplicates, partial realizations, and artifacts. The DriveArena matrix evaluates two planners, two horizons, six prompt conditions, and 360 ego-route requests; 304 attempts become fully realized evidence and 56 remain partial. A disjoint 80-request route-slice check yields 74 fully realized and 6 partial attempts. The results support evaluating world-model-style testing by convergence of valid interactive evidence under budget, rather than by raw generated failures or prompt coverage alone.",
      "short_abstract": "World models and generative simulators are emerging as interactive testing infrastructure for autonomous driving because they can react to the ego planner and produce counterfactual, rare, and safety-critical rollouts. This changes a test scenario from a...",
      "published": "2026-07-31",
      "updated": "2026-07-31",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 19,
      "arxiv_url": "https://arxiv.org/abs/2608.00298",
      "pdf_url": "https://arxiv.org/pdf/2608.00298.pdf"
    },
    {
      "id": "2608.00237",
      "anchor": "paper-2608-00237",
      "title": "Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving",
      "authors": [
        "Meibo Hu",
        "Jiamian Wang",
        "Pichao Wang",
        "Zhiqiang Tao"
      ],
      "abstract": "Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs and high-level navigation instructions directly to executable trajectories. However, in practice, these models exhibit a persistent command-following gap: predicted trajectories often show weak sensitivity to navigation commands, resulting in incorrect behavior at critical decision points. We identify this issue as a form of conditional policy collapse, where regression-based training under multimodal trajectory distributions encourages the model to rely on dominant visual priors while marginalizing the language-conditioned signal. To address this issue, we introduce a principled formulation of classifier-free guidance (CFG) for regression-based vision-language driving. We show that CFG can be interpreted as isolating the instruction-induced residual in the action space by contrasting conditional and unconditional predictions, thereby explicitly amplifying the effect of the navigation command at inference time. However, a standard two-pass CFG introduces prohibitive latency for real-time control and produces noisy instance-level guidance directions. Building on a mean-shift interpretation of CFG, we propose Latent-Centroid Steering (LCS), a single-pass guidance mechanism that replaces instance-level residuals with class-level latent shifts. By projecting conditional representations toward precomputed command-specific centroids, LCS performs class-level latent steering based on cluster geometry that is both more stable and computationally efficient. We demonstrate that LCS reduces inference latency by approximately 50% while achieving stronger command adherence and improved driving performance on both closed-loop (Bench2Drive) and open-loop (nuScenes) benchmarks. Code will be released.",
      "short_abstract": "Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs and high-level navigation instructions directly to executable trajectories. However, in practice, these...",
      "published": "2026-07-31",
      "updated": "2026-07-31",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 10,
        "margin": 1,
        "evidence": [
          "abstract: control",
          "title: steering or braking",
          "abstract: steering or braking"
        ]
      },
      "recency": "recent",
      "age_days": 19,
      "arxiv_url": "https://arxiv.org/abs/2608.00237",
      "pdf_url": "https://arxiv.org/pdf/2608.00237.pdf"
    },
    {
      "id": "2607.29517",
      "anchor": "paper-2607-29517",
      "title": "STAGE: STyle-controllable Action GEneration for personalized autonomous driving",
      "authors": [
        "Zihao Liu",
        "Xing Liu",
        "Yizhai Zhang",
        "Panfeng Huang"
      ],
      "abstract": "Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varying levels of aggressiveness. If humans choose to use autonomous driving systems, they would expect the driving style of the systems to closely resemble their own habit. However, this is challenging for current industrial autonomous driving systems. To address this, we developed a style controllable action generation method, STAGE, for driving tasks. Its training process is based on imitation learning, incorporating both style value and latent value action modality encoding. Preference learning is then used to identify the user's driving style as a continuous, monotonic style value. And to reduce the cost of human involvement in the preference training process, we also developed a set of rules to compare driving style in data pairs. Then, during inference, the user inputs the style value to control the generated action patterns, dynamically meeting the user's expectations. Using the STAGE method, we verified that the style-controlled action generation results in several typical road scenarios significantly align with human expectations. Furthermore, through comparisons between the STAGE method and various other approaches, we reveal the unique functionalities of STAGE, including its style controllability, style continuity, driving style alignment capability and driving safety. The code for this work is available at: https://github.com/CarlDegio/STAGE",
      "short_abstract": "Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varying levels of aggressiveness. If humans choose to use autonomous driving...",
      "published": "2026-07-31",
      "updated": "2026-07-31",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "recent",
      "age_days": 19,
      "arxiv_url": "https://arxiv.org/abs/2607.29517",
      "pdf_url": "https://arxiv.org/pdf/2607.29517.pdf"
    },
    {
      "id": "2607.29052",
      "anchor": "paper-2607-29052",
      "title": "Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving",
      "authors": [
        "Zeyu Dong",
        "Yimin Zhu",
        "Yu Wu",
        "Yu Sun"
      ],
      "abstract": "End-to-end (E2E) autonomous driving aims to learn a direct mapping from visual observations to control actions. However, these E2E models often act as black boxes and struggle with complex scenarios. To address this, recent works incorporate Vision-Language Models (VLMs) to provide explicit reasoning, enhancing both interpretability and driving robustness. These approaches typically rely on pre-generated annotations, which suffer from potentially flawed labels and require costly human labor. In this work, we propose a new framework that integrates structured reasoning and geometric precision through a teacher-student architecture. The teacher model introduces reflective reasoning, where the VLM generates logical explanations and then reflectively refines the reasoning under the supervision of ground-truth action. This enhances zero-shot generalization without intermediate labels. The student model distills the teacher's reasoning capabilities via supervised fine-tuning. We also design a separate waypoint decoder that interprets textual reasoning into continuous trajectories. Our proposed solution integrates two goals: providing explicit reasoning for interpretability and delivering robust and accurate driving performance. It leverages the synergy between these two objectives within a staged inference engine to enhance driving performance and explicitly uses the reasoning to guide driving prediction. Evaluated on Waymo benchmarks, our framework outperforms classical reasoning-based baselines in zero-shot reasoning, waypoint accuracy, and inference efficiency. Our experiments validate this design, demonstrating that the reasoning text makes a significant contribution to driving inference, resulting in around a 24% improvement in performance compared to an identical model that lacks reasoning. Our work advances reasoning-driven autonomous driving toward interpretable and deployable systems.",
      "short_abstract": "End-to-end (E2E) autonomous driving aims to learn a direct mapping from visual observations to control actions. However, these E2E models often act as black boxes and struggle with complex scenarios. To address this, recent works incorporate Vision-Language...",
      "published": "2026-07-31",
      "updated": "2026-07-31",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: mapping",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 19,
      "arxiv_url": "https://arxiv.org/abs/2607.29052",
      "pdf_url": "https://arxiv.org/pdf/2607.29052.pdf"
    },
    {
      "id": "2607.29031",
      "anchor": "paper-2607-29031",
      "title": "Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving",
      "authors": [
        "Jiwei Yang",
        "Zhengxian Chen",
        "Chaosheng Huang",
        "Jun Li"
      ],
      "abstract": "Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion. We argue that planning need not reconstruct the complete future world, but only focus on scene features that affect future ego action. Based on this perspective, we propose Auto-JEPA, an action-oriented latent world model that learns continuous future driving intent through joint-embedding prediction. Given visual observations, egomotion history, and navigation commands, Auto-JEPA predicts an intent embedding aligned with the latent representation of the future ego trajectory. The predicted intent retrieves executable trajectories from a fixed trajectory memory, which are then ranked by a scene-conditioned candidate selection module. Auto-JEPA keeps the visual encoder frozen, requires no explicit perception annotations, and uses no learned trajectory generator. By optimizing only task-specific modules for trajectory representation, intent prediction, and candidate selection, Auto-JEPA achieves 91.3 PDMS on NAVSIM v1 and 89.1 EPDMS on NAVSIM v2. Semantic occlusion experiments show that masking dynamic-agent regions induces an average intent change 2.97x that of equal-area random masking. Moreover, occluding vehicles that affect future driving substantially changes the predicted intent and selected trajectory, whereas both remain essentially unchanged when non-influential vehicles are occluded. These results show that future-intent prediction encourages the model to focus on planning-relevant visual features and supports high-quality planning without dense future-world modeling.",
      "short_abstract": "Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion. We argue that planning need not reconstruct the complete future world, but only focus on scene features that...",
      "published": "2026-07-31",
      "updated": "2026-07-31",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 27,
        "margin": 5,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 19,
      "arxiv_url": "https://arxiv.org/abs/2607.29031",
      "pdf_url": "https://arxiv.org/pdf/2607.29031.pdf"
    },
    {
      "id": "2608.02638",
      "anchor": "paper-2608-02638",
      "title": "Studying, Identifying, and Fixing Hidden Technical Debt in AI-Intensive Cyber-Physical Systems",
      "authors": [
        "Beena"
      ],
      "abstract": "Artificial Intelligence (AI) components are increasingly pervasive in several software systems, including Cyber-Physical Systems (CPSs). AI-CPS are used in several domains, including autonomous vehicles, industry, home automation, robotics, and healthcare. Being composed of hardware, AI components, and conventional modules, AI-CPS can exhibit technical debt (TD) that is peculiar and potentially more challenging than that of conventional systems. This thesis aims to characterize AI-CPS TD and propose approaches for its identification and repair. In a first phase, we characterize AI-CPS TD by analyzing AI ecosystems and AI-CPS repositories, as well as interviewing developers. Based on the acquired knowledge, we define approaches to identify and mitigate such TD. Finally, we plan to develop and validate an automated tool that supports agentic AI solutions to monitor, govern, and repay AI-CPS TD.",
      "short_abstract": "Artificial Intelligence (AI) components are increasingly pervasive in several software systems, including Cyber-Physical Systems (CPSs). AI-CPS are used in several domains, including autonomous vehicles, industry, home automation, robotics, and healthcare....",
      "published": "2026-07-31",
      "updated": "2026-07-31",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: hardware"
        ]
      },
      "recency": "recent",
      "age_days": 19,
      "arxiv_url": "https://arxiv.org/abs/2608.02638",
      "pdf_url": "https://arxiv.org/pdf/2608.02638.pdf"
    },
    {
      "id": "2608.00113",
      "anchor": "paper-2608-00113",
      "title": "Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning",
      "authors": [
        "Sheng Zhao",
        "Bolin Zhao",
        "Xiaodong Wu",
        "Chen Lv"
      ],
      "abstract": "In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break traction to drift on loose surfaces. This maneuver rapidly aligns the vehicle for corner exits, ultimately reducing lap time. Autonomously executing such maneuvers formulates a complex dual-objective control problem: stabilizing highly nonlinear drift dynamics while strictly minimizing lap time. Addressing this challenge motivates the development of advanced Minimum-Lap-Time (MLT) drift control architectures. This paper proposes a planning-control framework specifically designed for MLT drifting scenario. First, we formulate an optimal control problem to generate a MLT drift planning trajectory, which is used as prior data to train a deep reinforcement learning drift controller. Given that drifting involves extremely large sideslip angles and is therefore challenging to learn directly, a Track-guided Reinforcement Learning (TgRL) drift control method is proposed to enable progressive training in a step-by-step manner, from drift control policy, to drift corner policy, and finally to a comprehensive drift race policy. The reward function incorporates both an instant reward term and an end reward term derived from the Minimum-Lap-Time objective. Simulation results demonstrate that the proposed framework enables the agent to learn a drift racing policy that not only ensures vehicle motion control performance but also effectively reduces lap time.",
      "short_abstract": "In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break traction to drift on loose surfaces. This maneuver rapidly aligns the vehicle for corner exits, ultimately...",
      "published": "2026-07-31",
      "updated": "2026-07-31",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "recent",
      "age_days": 19,
      "arxiv_url": "https://arxiv.org/abs/2608.00113",
      "pdf_url": "https://arxiv.org/pdf/2608.00113.pdf"
    },
    {
      "id": "2607.28483",
      "anchor": "paper-2607-28483",
      "title": "Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles",
      "authors": [
        "Luca de Martino",
        "Federico Aromolo",
        "Federico Nesti",
        "Giorgio Buttazzo"
      ],
      "abstract": "Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational cost limits their deployment on embedded hardware. This work presents an efficient and accelerated pipeline designed for both embedded and desktop platforms, targeting the autonomous driving and railway domains. The proposed approach reformulates the Neyman-Pearson scoring stage of PixOOD, a state-of-the-art out-of-distribution detection method, and deploys the full pipeline through hardware-optimized TensorRT compilation, reaching up to 182 FPS on a desktop NVIDIA RTX 4060 GPU and 75 FPS on the NVIDIA Jetson AGX Orin embedded platform, respectively 20x and 18x faster than the original baseline. The achieved results demonstrate that advanced anomaly segmentation can be efficiently deployed for onboard processing in autonomous driving and railway applications.",
      "short_abstract": "Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational cost limits their deployment on embedded hardware. This work presents an efficient and accelerated pipeline...",
      "published": "2026-07-30",
      "updated": "2026-07-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 11,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: hardware"
        ]
      },
      "recency": "recent",
      "age_days": 20,
      "arxiv_url": "https://arxiv.org/abs/2607.28483",
      "pdf_url": "https://arxiv.org/pdf/2607.28483.pdf"
    },
    {
      "id": "2607.28799",
      "anchor": "paper-2607-28799",
      "title": "FedQML-Edge: Compact Quantum Feature Sketches for Communication-Constrained Roadside Federated Learning",
      "authors": [
        "Talha Azfar",
        "Ruimin Ke"
      ],
      "abstract": "Roadside units (RSUs) supporting connected and autonomous vehicle corridors need compact models to decide when cooperative maneuvers should be rewarded, deferred, or disabled. Raw sensor streams and neural network weight checkpoints are poorly suited to bandwidth-limited, privacy-sensitive roadside learning. This paper presents $\\texttt{FedQML-Edge}$, a federated quantum feature-sketching pipeline for traffic-stability gating. Each RSU constructs a traffic-state summary and sends circuit inputs to a quantum computer; Pauli expectations form a nonlinear sketch processed by a logistic classifier. Only classifier updates are shared with an aggregator, whose head supports reward gating. Raw observations, vehicle records, event traces, and quantum sketches remain private. We evaluate the method using NGSIM trajectories, SUMO predictive gating with sensing noise, and IBM Quantum hardware. On NGSIM, the Pauli sketch reduces test log loss by $14.4\\%$ relative to the strongest matched classical sketch. On SUMO, it approaches larger MLPs in stable-window recall while using $7-28$ times less communication per round.",
      "short_abstract": "Roadside units (RSUs) supporting connected and autonomous vehicle corridors need compact models to decide when cooperative maneuvers should be rewarded, deferred, or disabled. Raw sensor streams and neural network weight checkpoints are poorly suited to...",
      "published": "2026-07-30",
      "updated": "2026-07-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 32,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: hardware",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints",
          "title: roadside infrastructure",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "recent",
      "age_days": 20,
      "arxiv_url": "https://arxiv.org/abs/2607.28799",
      "pdf_url": "https://arxiv.org/pdf/2607.28799.pdf"
    },
    {
      "id": "2607.28083",
      "anchor": "paper-2607-28083",
      "title": "Layered Architecture for Mobile Intelligence",
      "authors": [
        "Qingwen Liu",
        "Mingqing Liu"
      ],
      "abstract": "Artificial intelligence (AI) is rapidly evolving from a centralized computing capability into a pervasive infrastructure that interacts directly with the physical world. While recent perspectives highlight the roles of energy, chips, infrastructure, models, and applications in enabling large-scale AI systems, these frameworks primarily assume a static, cloud-centric computing paradigm. However, emerging intelligent applications, including autonomous vehicles, drones, robots, and wearable systems, require AI to operate in highly dynamic and mobile environments. This shift introduces mobility as a fundamental constraint across the entire AI ecosystem, affecting energy supply, computation, and intelligence deployment. In this article, we introduce the concept of the Mobile AI Stack, a mobility-aware architectural framework that integrates five tightly coupled layers: mobile energy networks, energy-efficient AI chips, cloud-edge-mobile infrastructure, distributed AI models, and embodied AI applications. The proposed framework provides a systematic perspective for understanding how energy delivery, computing architectures, communication networks, and AI algorithms must co-evolve to support large-scale mobile intelligence. We further discuss key research challenges and future directions toward building scalable, reliable, and energy-efficient mobile AI systems. Mobile AI Stack offers a conceptual blueprint of the next-generation infrastructure which deeply integrates the networks of computation, energy, and communications for mobile intelligence.",
      "short_abstract": "Artificial intelligence (AI) is rapidly evolving from a centralized computing capability into a pervasive infrastructure that interacts directly with the physical world. While recent perspectives highlight the roles of energy, chips, infrastructure, models,...",
      "published": "2026-07-30",
      "updated": "2026-07-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 20,
      "arxiv_url": "https://arxiv.org/abs/2607.28083",
      "pdf_url": "https://arxiv.org/pdf/2607.28083.pdf"
    },
    {
      "id": "2607.28953",
      "anchor": "paper-2607-28953",
      "title": "A Cooperative Implementation of Mesh Stability in Vehicular Platoons",
      "authors": [
        "Meng Qiu",
        "Di Liu",
        "He Wang",
        "Wenwu Yu",
        "Simone Baldi"
      ],
      "abstract": "This work studies the problem of mesh stability in connected and automated vehicles. Mesh stability, also known as 2D string stability, refers to studying how disturbances propagate in vehicular platoons in both longitudinal and lateral direction. As opposed to available decentralized results only relying on on-board sensing, the distinguishing feature of this work is a cooperative version of mesh stability, where onboard sensing is augmented by vehicle-to-vehicle communication. This cooperative version dramatically improves the state-of-the-art decentralized performance: for longitudinal control, a new cooperative non-identical protocol (i.e. with non-identical control gains) is proposed that improves the state-of-the-art decentralized non-identical protocol in terms of scalability of the control gains and strong notion of string stability. For lateral control, after showing that the non-identical approach is not necessary, a cooperative non-identical protocol is proposed to achieve another strong notion of string stability. Robustness of the proposed implementation against vehicle-to-vehicle communication delays and actuation time lags is considered. Numerical experiments, also performed with the vehicle simulator CarSim, validate the robustness and effectiveness of the proposed protocol.",
      "short_abstract": "This work studies the problem of mesh stability in connected and automated vehicles. Mesh stability, also known as 2D string stability, refers to studying how disturbances propagate in vehicular platoons in both longitudinal and lateral direction. As opposed...",
      "published": "2026-07-30",
      "updated": "2026-07-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 18,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 20,
      "arxiv_url": "https://arxiv.org/abs/2607.28953",
      "pdf_url": "https://arxiv.org/pdf/2607.28953.pdf"
    },
    {
      "id": "2607.27371",
      "anchor": "paper-2607-27371",
      "title": "Dynamics-matched Physical Reservoir Computing for Undersensed Traffic Prediction",
      "authors": [
        "Michael McCreesh",
        "Rohit Gupta",
        "Stephen L. Smith"
      ],
      "abstract": "Machine learning methods are increasingly used for traffic prediction in applications such as autonomous driving. Such predictions must be both highly accurate and immediately available, making methods with low computational costs and fast training times of interest. One such method is reservoir computing, in which the rich dynamics of a nonlinear system serves as a computational substrate and only a linear readout vector is trained. In this work we use a traffic network as the reservoir for predicting the behavior of an undersensed traffic network. This matching of the highly nonlinear dynamics allows for similar encoding between the behaviors of the reservoir and target network, enabling a more direct prediction. We show that a reservoir governed by the Improved Intelligent Driver Model (IIDM) satisfies the echo state property for a class of slowly-varying inputs. Through simulations we show that the echo state property likely holds for a larger class of inputs, and that the IIDM reservoir computer (IIDM-RC) accurately predicts an undersensed vehicle network governed by varying car-following models. We also compare with echo state networks (ESNs) and Long Short-Term Memory (LSTM) networks, finding improvements using IIDM-RC in both prediction accuracy and training time.",
      "short_abstract": "Machine learning methods are increasingly used for traffic prediction in applications such as autonomous driving. Such predictions must be both highly accurate and immediately available, making methods with low computational costs and fast training times of...",
      "published": "2026-07-29",
      "updated": "2026-07-29",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 6,
        "evidence": [
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "recent",
      "age_days": 21,
      "arxiv_url": "https://arxiv.org/abs/2607.27371",
      "pdf_url": "https://arxiv.org/pdf/2607.27371.pdf"
    },
    {
      "id": "2607.27058",
      "anchor": "paper-2607-27058",
      "title": "Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation",
      "authors": [
        "Danning Zhu",
        "Ziyan Lin",
        "Jing Wu"
      ],
      "abstract": "Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios. Our dataset combines real-world images captured in Weishi County, Henan Province, with parameterized virtual scenes generated via Unreal Engine. To accurately reflect the unique realities of rural traffic, we define a comprehensive 14-category object system encompassing region-specific elements such as electric tricycles, low-speed vehicles (LSVs), and roadside stalls. Under a unified training protocol, we systematically evaluate 13 mainstream detectors -- spanning the YOLOv5, YOLOv8, YOLO11, and YOLO26 series, as well as RT-DETR-L -- across three data configurations: an all-real baseline, a 1:0.5 real-to-virtual mix, and a 1:1 mix. Experimental results demonstrate that a moderate injection of synthetic data (1:0.5 ratio) effectively enhances detection performance, with YOLO11m achieving the highest mAP@0.5 of 0.758. However, a higher proportion of synthetic data (1:1) introduces domain shifts that offset the benefits of data scaling. While most models reliably identify distinct local vehicles, significant perceptual bottlenecks remain for long-tail, non-standard objects like stalls and railings. This research provides crucial empirical evidence and novel insights for model selection and synthetic data strategies, facilitating the practical deployment of autonomous driving perception systems in rural areas.",
      "short_abstract": "Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection...",
      "published": "2026-07-29",
      "updated": "2026-07-29",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 13,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "recent",
      "age_days": 21,
      "arxiv_url": "https://arxiv.org/abs/2607.27058",
      "pdf_url": "https://arxiv.org/pdf/2607.27058.pdf"
    },
    {
      "id": "2607.27036",
      "anchor": "paper-2607-27036",
      "title": "Mitigating Compounding Error via Video Representation Regularization",
      "authors": [
        "Taiye Chen",
        "Qi Zhang",
        "Yisen Wang"
      ],
      "abstract": "Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional collapse of hidden representations. Specifically, the effective rank of model representations sharply decreases at the onset of generation drift, revealing a strong connection between representational degradation and long-term rollout instability. Furthermore, we find that pure training data scaling fails to boost model resistance to error drift, contradicting mainstream scaling paradigms. To address this problem, we propose video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation. Compared with Diffusion Forcing, our method achieves improvements from 38.65 to 55.56 and from 44.37 to 72.08 on the Aesthetic Quality and Imaging Quality metrics of VBench. Our work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.",
      "short_abstract": "Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time....",
      "published": "2026-07-29",
      "updated": "2026-07-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Prediction & World Models: abstract: world model",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 21,
      "arxiv_url": "https://arxiv.org/abs/2607.27036",
      "pdf_url": "https://arxiv.org/pdf/2607.27036.pdf"
    },
    {
      "id": "2607.26802",
      "anchor": "paper-2607-26802",
      "title": "Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment",
      "authors": [
        "Marc Kaufeld",
        "Dian Zhuang",
        "Johannes Betz"
      ],
      "abstract": "This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The proposed approach combines RBFN-based candidate trajectory generation with an analytic collision probability assessment and optimization-based trajectory refinement. The network learns jerk-minimal trajectories, enabling the MPC to operate within a reduced and dynamically consistent search space. Candidate motion primitives are selected based on an accurate probabilistic risk measure. This design decreases solver complexity while preserving safety and constraint satisfaction. The framework is evaluated in numerous urban driving scenarios. Results demonstrate improved risk awareness and fewer vehicle-limit violations compared to benchmark methods. The proposed approach integrates learning-based trajectories into optimization-based motion planning, thereby ensuring safety and interpretability.",
      "short_abstract": "This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The proposed approach combines RBFN-based candidate trajectory generation with an analytic collision probability...",
      "published": "2026-07-29",
      "updated": "2026-07-29",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 7,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: trajectory generation"
        ]
      },
      "recency": "recent",
      "age_days": 21,
      "arxiv_url": "https://arxiv.org/abs/2607.26802",
      "pdf_url": "https://arxiv.org/pdf/2607.26802.pdf"
    },
    {
      "id": "2607.26651",
      "anchor": "paper-2607-26651",
      "title": "Physically Real-time Infrared Attack against Optical Flow Estimation Networks",
      "authors": [
        "Shen You",
        "Wei Jiang",
        "Jiarui Liu",
        "Yijian Ye",
        "Qiuzhen Lin",
        "Xiangtao Li",
        "Ka-Chun Wong"
      ],
      "abstract": "With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving and motion detection have become increasingly mature and relevant to human lives. In particular, Optical Flow Estimation Networks (OFENs), as upstream models, play a critical role in different domains. Its outputs are heavily assumed and adopted for different downstream tasks, and it is essential to test its robustness to prevent safety accidents. We present an approach for real-time attacks on OFENs in the physical world, leveraging infrared lights for their stealthiness. By generating a large number of Adversarial Examples in advance, our approach computes AEs in real time and dynamically displays them, which allows our method to facilitate precise and targeted attacks without modifying the victim system. Unlike previous digital-to-physical attack techniques, our method directly attacks victim models within the physical world, thereby overcoming the limitations associated with the ineffectiveness of AEs. Experimental results demonstrate the efficacy of our approach in compromising OFENs across diverse lighting conditions, varying object motion velocities, and different object placements, ultimately impairing the network's ability to accurately estimate optical flow.",
      "short_abstract": "With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving and motion detection have become increasingly mature and relevant to human lives. In particular, Optical Flow Estimation...",
      "published": "2026-07-29",
      "updated": "2026-07-29",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 22,
        "margin": 2,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 21,
      "arxiv_url": "https://arxiv.org/abs/2607.26651",
      "pdf_url": "https://arxiv.org/pdf/2607.26651.pdf"
    },
    {
      "id": "2607.27085",
      "anchor": "paper-2607-27085",
      "title": "Controlled Experiments on Lane Changing by Transitional Autonomous Vehicle: Dataset and Behavioral Insights",
      "authors": [
        "Abhinav Sharma",
        "Md Abdullah Al Hasan",
        "Danjue Chen",
        "George F. List"
      ],
      "abstract": "This paper presents the North Carolina Transitional Autonomous Vehicle Lane-Changing (NC-tALC) dataset and uses it to characterize mandatory lane-changing behavior of transitional automated vehicles (tAVs). It quantifies the evolution of lead--lag gaps throughout the lane-change process and examines how potential collision risk develops during the maneuver. A controlled field experiment comprising 78 mandatory lane-change trials was conducted on a public roadway in Apex, North Carolina. Four instrumented vehicles created repeatable traffic conditions while varying the lane changer's initial position within the candidate target gap. High-resolution RTK-GNSS/INS trajectories were processed to identify key timestamps, calculate lead, lag, and lane-change gaps, and estimate interactions using time-gap- and speed-based surrogate safety measures. Despite substantial differences in initial conditions, lead and lag gaps consistently converged toward a relatively narrow range near lane crossing. Potential collision risk increased as the maneuver progressed, peaked near physical lane entry, and was dominated by interactions with the target-lane leader. Lane-change completion did not necessarily coincide with the disappearance of collision risk. This study provides one of the first controlled empirical characterizations of the complete mandatory lane-change process of tAVs using repeatable public-road experiments. The NC-tALC dataset supports analysis of behavioral and safety evolution throughout the maneuver rather than only at the gap-acceptance instant. The dataset and findings provide empirical benchmarks for evaluating automated lane-changing behavior, calibrating behavioral models, and validating simulation and safety assessment methods for mandatory lane-change scenarios.",
      "short_abstract": "This paper presents the North Carolina Transitional Autonomous Vehicle Lane-Changing (NC-tALC) dataset and uses it to characterize mandatory lane-changing behavior of transitional automated vehicles (tAVs). It quantifies the evolution of lead--lag gaps...",
      "published": "2026-07-29",
      "updated": "2026-07-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Mapping & Localization: abstract: GNSS or GPS"
        ]
      },
      "recency": "recent",
      "age_days": 21,
      "arxiv_url": "https://arxiv.org/abs/2607.27085",
      "pdf_url": "https://arxiv.org/pdf/2607.27085.pdf"
    },
    {
      "id": "2607.26165",
      "anchor": "paper-2607-26165",
      "title": "DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving",
      "authors": [
        "Yung-Hsu Yang",
        "Luigi Piccinelli",
        "Siyuan Li",
        "Mattia Segu",
        "Lei Ke",
        "Martin Danelljan",
        "Yuqian Fu",
        "Zuria Bauer",
        "Fisher Yu",
        "Hermann Blum",
        "Marc Pollefeys"
      ],
      "abstract": "Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies these tasks, existing approaches often rely on computationally expensive, multi-stage pipelines or offline tracking, rendering them unsuitable for real-time decision-making. To address this, we propose DVPSFormer, a unified online architecture designed for efficient 4D scene understanding. Central to our approach is explicit scene discretization (ESD), a novel mechanism that leverages segmentation queries to represent foreground and background regions, enabling a discrete-to-continuous (D2C) depth head to decode metric depth in a single pass. This tightly couples semantic and geometric learning while significantly reducing latency. Furthermore, we propose an online majority voting (OMV) mechanism that exploits temporal consistency to refine classification during instance tracking. DVPSFormer establishes a new state-of-the-art on the Cityscapes-DVPS and SemKITTI-DVPS benchmarks, offering a streamlined solution for online robotic perception. Code and models are available at https://royyang0714.github.io/DVPSFormer.",
      "short_abstract": "Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies...",
      "published": "2026-07-28",
      "updated": "2026-07-28",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 16,
        "evidence": [
          "abstract: perception",
          "title: scene segmentation",
          "abstract: scene segmentation",
          "abstract: scene understanding"
        ]
      },
      "recency": "recent",
      "age_days": 22,
      "arxiv_url": "https://arxiv.org/abs/2607.26165",
      "pdf_url": "https://arxiv.org/pdf/2607.26165.pdf"
    },
    {
      "id": "2607.26121",
      "anchor": "paper-2607-26121",
      "title": "Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels",
      "authors": [
        "Xinyu Yang",
        "Tianxing Chen",
        "Honghao Su",
        "Minxuan Wang",
        "Chenze Yu",
        "Zhangzheng Tu",
        "Yue Chen",
        "Yuxiao Huo",
        "Lingfeng Zhang",
        "Yan Huang",
        "Yan Qin",
        "Shaolong Zhu",
        "Qiwei Liang",
        "Hekun Tian",
        "Shujia Liu",
        "Guangyu Chen",
        "Junhao Gong",
        "Zixuan Li",
        "Wenwei Lin",
        "Zijian Lin",
        "Wenxuan Zhu",
        "Eric J Chen",
        "Yue Yuan",
        "Qize Yu",
        "Jiaqi Liang"
      ],
      "abstract": "Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bounds. We term this objective sustained safe success. Its supporting mechanisms are organized into four interdependent layers. The model layer generates task-competent action proposals with calibrated uncertainty and explicit safety preferences. The system layer realizes authorized actions dependably through integrated sensing, computation, control, hardware safeguards, fault containment, and fallback. The evidence layer substantiates bounded claims through evaluation, verification, validation, traceability, and structured assurance arguments. The deployment layer maintains claim validity through runtime monitoring, authority management, intervention, incident response, and controlled updates. Because assumptions and failures propagate across these layers, neither model capability, isolated safeguards, nor benchmark performance alone can establish end-to-end trustworthiness. Drawing on embodied AI, robotics, control, dependable computing, distributed systems, and autonomous driving, we further propose a non-normative hierarchy of trustworthiness levels. This hierarchy grades the strength of bounded deployment claims across task capability, safety, system assurance, operational governance, and supporting evidence, providing a basis for bounded deployment, comparative evaluation, research prioritization, and future standardization.",
      "short_abstract": "Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness....",
      "published": "2026-07-28",
      "updated": "2026-07-28",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 10,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "abstract: validation or testing",
          "abstract: uncertainty",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 22,
      "arxiv_url": "https://arxiv.org/abs/2607.26121",
      "pdf_url": "https://arxiv.org/pdf/2607.26121.pdf"
    },
    {
      "id": "2607.25989",
      "anchor": "paper-2607-25989",
      "title": "Untangling Co-Drift: Proactive Multi-Intent Failure Prediction and Root-Cause Disambiguation for Self-Driving Networks",
      "authors": [
        "Md. Kamrul Hossain",
        "Walid Aljoby"
      ],
      "abstract": "The vision of self-driving networks that monitor, reason, and act upon themselves with minimal human intervention relies on tightly coupled monitoring, analytics, and actuation functions. In this work, we treat these functions as three operational macro-intents: continuous telemetry, real-time analytics, and programmatic actuation, and formalize the health of each function as an intent that the network must continuously satisfy. A critical, yet underexplored, challenge stems from the causal coupling among these intents, where a singular fault within one macro-intent propagates as a co-drift and subsequently triggers cascading, symptomatic anomalies across the remaining intents. This ambiguity makes it exceedingly difficult for existing, reactive approaches to distinguish the true root-cause intent from symptomatic victim intents, and their reliance on threshold-crossing detection leaves insufficient time for proactive remediation. We introduce MILD, a novel framework that reformulates intent assurance from reactive drift detection to proactive failure prediction. Grounded in our three-macro-intent formulation of the self-driving control loop, MILD employs a teacher-augmented Mixture-of-Experts architecture with a hybrid objective that jointly optimizes intent failure prediction and root-cause attribution. MILD enables KPI-level diagnostics via SHAP explainability and dynamic intent failure urgency estimation via multi-horizon modeling. Our extensive evaluation of MILD across three environments of increasing realism, from a controlled statistical benchmark, to a microservices application, to an SDN-based edge-to-cloud testbed, demonstrates that MILD achieves high failure detection rates, strong remediation lead times, and accurate intent-level root-cause disambiguation. This positions MILD as a practical enabler of closed-loop assurance in next-generation autonomous networks.",
      "short_abstract": "The vision of self-driving networks that monitor, reason, and act upon themselves with minimal human intervention relies on tightly coupled monitoring, analytics, and actuation functions. In this work, we treat these functions as three operational...",
      "published": "2026-07-28",
      "updated": "2026-07-28",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 3,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 22,
      "arxiv_url": "https://arxiv.org/abs/2607.25989",
      "pdf_url": "https://arxiv.org/pdf/2607.25989.pdf"
    },
    {
      "id": "2607.25612",
      "anchor": "paper-2607-25612",
      "title": "Multi-Sensor Alignment for Weather Simulations",
      "authors": [
        "Samsad Alam",
        "Devyani Lambhate",
        "Aditya Mohan",
        "Vishal Kumar",
        "Vaibhav Katewa"
      ],
      "abstract": "Perception tasks for autonomous vehicles need to work satisfactorily in adverse weather conditions. Due to lack of real-world weather datasets, weather simulations are a promising alternative. To ensure simulations closely mirror real-world weather data, it's crucial that they represent the same weather characteristics, including severity and particle positioning, across different sensors. To achieve this, we propose the Reference Dataset Alignment Method (ReDAM) for weather intensity alignment in fog and Unified-weather-edit (inspired by Weather-edit[1]) for particle positioning alignment in rain and snow. We validate both alignment methods using statistical and geometrical tests, respectively. We find that 3D detection models for non-aligned versions tend to be overly optimistic as compared to aligned versions. We also show the aligned-multi-sensor simulation's effectiveness for achieving robustness for 3D object detection task by finetuning existing sensor fusion models on it.",
      "short_abstract": "Perception tasks for autonomous vehicles need to work satisfactorily in adverse weather conditions. Due to lack of real-world weather datasets, weather simulations are a promising alternative. To ensure simulations closely mirror real-world weather data, it's...",
      "published": "2026-07-28",
      "updated": "2026-07-28",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "abstract: object detection",
          "abstract: sensor fusion"
        ]
      },
      "recency": "recent",
      "age_days": 22,
      "arxiv_url": "https://arxiv.org/abs/2607.25612",
      "pdf_url": "https://arxiv.org/pdf/2607.25612.pdf"
    },
    {
      "id": "2607.25570",
      "anchor": "paper-2607-25570",
      "title": "The LAIA Dataset: Labelled Attention for Intelligent Automobiles",
      "authors": [
        "A. Contreras",
        "D. Porres",
        "R. Abad",
        "P. Cano",
        "A. Levy",
        "G. Villalonga",
        "A. M. López",
        "A. Hernández-Sabaté"
      ],
      "abstract": "The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large volumes of sensor data with ground-truth annotations. While modular architectures are widely used, end-to-end driving paradigms offer a promising alternative by directly mapping sensor inputs to control actions. However, their adoption is limited by challenges in interpretability and explainability. To address this, we present LAIA (Labelled Attention for Intelligent Automobiles), a novel synthetic dataset designed to enrich end-to-end driving research with human attention data. Collected using the CARLA simulator in closed-loop environments, LAIA comprises over 15 hours of driving from 44 participants across carefully crafted scenarios designed to evoke natural responses. Each sequence includes RGB images under six weather conditions, semantic and instance segmentation, depth, optical flow, CAN bus signals, and synchronized eye-tracking data. LAIA enables applications including training attention-aware end-to-end AI drivers, predicting driver behavior, developing methods to detect anomalous driver-attention patterns, and improving model explainability. In this work, we use LAIA to compare human attention with the perceptual attention emerging in our end-to-end driving models, thereby providing insight into their behavior.",
      "short_abstract": "The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large volumes of sensor data with ground-truth annotations. While modular architectures are widely used, end-to-end driving...",
      "published": "2026-07-28",
      "updated": "2026-07-28",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "Simulation",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 22,
      "arxiv_url": "https://arxiv.org/abs/2607.25570",
      "pdf_url": "https://arxiv.org/pdf/2607.25570.pdf"
    },
    {
      "id": "2607.25736",
      "anchor": "paper-2607-25736",
      "title": "Image Quality Dependent Degradation for AI Systems",
      "authors": [
        "Yannick Kees",
        "Elena Hoemann",
        "Frank Köster",
        "Sven Hallerbach"
      ],
      "abstract": "Perception is one of the primary applications where neural networks outperform conventional algorithms. One example is AI systems for automated driving, which can detect pedestrians based on image data and avoid them accordingly. A substantial challenge with these AI systems is that their output depends heavily on the quality of the input images. For example, if an image is of inferior quality due to heavy contamination, such as noise or darkness, accurate predictions are hardly feasible. Additionally, various types of errors can occur, each with varying relevance to the trustworthiness of the underlying AI system. In particular, it may be more critical not to detect an existing person than to detect a person where there is none. Therefore, we want to show that we can still avoid the most critical errors in situations of inferior image quality. To achieve this, we aim to establish a fail-degraded system by lowering the network's confidence threshold based on the estimated image quality, enabling it to detect objects more cautiously in uncertain situations. Additionally, we present a novel method for estimating the quality of incoming images by comparing them to the training data using normalizing flows. We will also conduct experiments applying our method to state-of-the-art object detection. In summary, we will present a design strategy for AI-based systems in automated driving that can deal with poor-quality input data without resorting to fallback solutions. Such measures enhance trust in AI-based systems and lead to an increased provision of the AI component.",
      "short_abstract": "Perception is one of the primary applications where neural networks outperform conventional algorithms. One example is AI systems for automated driving, which can detect pedestrians based on image data and avoid them accordingly. A substantial challenge with...",
      "published": "2026-07-28",
      "updated": "2026-07-28",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 5,
        "evidence": [
          "abstract: perception",
          "abstract: object detection"
        ]
      },
      "recency": "recent",
      "age_days": 22,
      "arxiv_url": "https://arxiv.org/abs/2607.25736",
      "pdf_url": "https://arxiv.org/pdf/2607.25736.pdf"
    },
    {
      "id": "2607.24692",
      "anchor": "paper-2607-24692",
      "title": "Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines",
      "authors": [
        "Jhonatan Tavori",
        "Gur-Eyal Sela",
        "Ion Stoica",
        "Gil Zussman"
      ],
      "abstract": "Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predictions. Across several application domains, we abstract this inference architecture as a fast path, a slow path, and a coordination layer with two functions: a router that invokes the slow path and a merger that decides whether to incorporate its returned predictions. In this work, we show that this new coordination layer exposes a new attack surface: shaped workload attacks, e.g., Yo-Yo bursts, can exploit contention at shared resources along the slow path to push benign users' slow-path predictions past their latency deadlines. The merger then discards those predictions, while the fast path continues to return timely outputs. We refer to the resulting loss of slow-path accuracy benefits as accuracy collapse. We demonstrate accuracy collapse in a two-tier edge-cloud multi-object tracking pipeline in autonomous driving. In simulation, approximately 4,000 burst-shaped requests increase benign p99 latency from 92ms to 2s, nearly eliminating the benefit of the slow path's cloud inference, reducing object tracking quality by 7.0 HOTA points on average. We further find that accuracy degradation can significantly vary (2.0-18.7 HOTA points), depending on the video intervals that are targeted in the attack, and that certain rare classes (e.g., stop signs) lose nearly half of their pre-attack prediction accuracy. These results show that workload attacks can degrade prediction quality without needing either access to model weights or victim data, and motivate research on attacks and defenses for routing, merging, scheduling, and resource isolation in these emerging inference pipeline architectures.",
      "short_abstract": "Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on...",
      "published": "2026-07-27",
      "updated": "2026-07-27",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: hardware",
          "title: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 23,
      "arxiv_url": "https://arxiv.org/abs/2607.24692",
      "pdf_url": "https://arxiv.org/pdf/2607.24692.pdf"
    },
    {
      "id": "2607.24577",
      "anchor": "paper-2607-24577",
      "title": "Evaluating Fuzz Testing for Reinforcement Learning Agents",
      "authors": [
        "Zhibin Kang",
        "Hanmo You",
        "Dong Wang",
        "Haiming Zheng",
        "Junjie Chen"
      ],
      "abstract": "Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences. Fuzz testing has recently emerged as a promising method for exploring the vast state spaces of RL agents and exposing crashes. Although numerous RL fuzzing methods have been proposed, existing studies often differ in evaluation settings, baselines, and metrics, making it difficult to draw reliable conclusions about their relative effectiveness and practical usefulness. To address this gap, we present the first comprehensive empirical study that systematically evaluates RL fuzzing methods from four complementary perspectives: effectiveness, diversity, efficiency, and practical utility. We benchmark five state-of-the-art methods alongside random testing under unified configurations across three environments of increasing complexity (MountainCar, BipedalWalker, and CARLA), and further assess the downstream usefulness of detected crashes for agent robustness improvement and safety monitoring. Our results reveal several key insights. For instance,throughput-oriented methods like MDPFuzz demonstrate superior effectiveness and efficiency in crash discovery, while methods explicitly designed to encourage exploration like SeqDivFuzz excel at uncovering diverse crash behaviors. We also show that fuzzing-generated crashes can meaningfully improve agent robustness and enable accurate safety monitoring with strong cross-method generalization. Beyond these empirical findings, we distill actionable guidance for both researchers and practitioners, highlighting the benefits of combining complementary fuzzing strategies and adopting multi-level diversity analysis to achieve more comprehensive and practical RL testing.",
      "short_abstract": "Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences. Fuzz testing has recently emerged as a...",
      "published": "2026-07-27",
      "updated": "2026-07-27",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 16,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 23,
      "arxiv_url": "https://arxiv.org/abs/2607.24577",
      "pdf_url": "https://arxiv.org/pdf/2607.24577.pdf"
    },
    {
      "id": "2607.24224",
      "anchor": "paper-2607-24224",
      "title": "MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving",
      "authors": [
        "Junchen Huo",
        "Wanming Hao",
        "Song Wang",
        "Enqing Chen",
        "Shouyi Yang",
        "Guanghui Wang"
      ],
      "abstract": "Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks. However, such a single feature map hardly carries sufficient information to simultaneously meet the requirements of various perception tasks, leading to a very limited perception performance. To mitigate this limitation, this paper proposes MATS, a novel multi-modality multi-task learning approach with modality-adaptive BEV fusion and task-specific Mixture-of-Experts (MoE) for 3D perception. Specifically, a simple modality-adaptive BEV fusion module is designed to adaptively recalibrate the BEV features by modeling the global cross-modality dependencies, generating diverse BEV feature maps for various perception tasks. For joint multi-task learning, this paper proposes a task-specific MoE module to decouple the tasks and enable the network to automatically choose the appropriate BEV feature candidates for each specific task. To validate the effectiveness of the proposed approach, we conduct extensive experiments on the large-scale benchmark nuScenes. With the camera- and LiDAR-modality input data, the proposed approach outperforms the state-of-the-art (SOTA) by a significant margin. Furthermore, the experimental results on the single tasks show that the proposed approach significantly outperforms the baselines. The code and trained models will be available upon publication.",
      "short_abstract": "Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to...",
      "published": "2026-07-27",
      "updated": "2026-07-27",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 17,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "recent",
      "age_days": 23,
      "arxiv_url": "https://arxiv.org/abs/2607.24224",
      "pdf_url": "https://arxiv.org/pdf/2607.24224.pdf"
    },
    {
      "id": "2607.24199",
      "anchor": "paper-2607-24199",
      "title": "Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding",
      "authors": [
        "Yueru Luo",
        "Xu Yan",
        "Changqing Zhou",
        "Yiming Yang",
        "Chao Zhan",
        "Shuqi Mei",
        "Chao Zheng",
        "Zhen Li"
      ],
      "abstract": "Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversity and context dependence of traffic signage. Importantly, regulation understanding is not a simple recognition task, but a reasoning problem: whether a rule applies depends on interpreting the sign in relation to the spatial layout of lanes and scene context. To support such reasoning, MapDR provide fine-grained annotations that link each traffic sign's regulatory rules to the specific lanes they govern. Existing methods, however, largely treat this as direct sequence prediction, ignoring the underlying reasoning that connects sign semantics and map structure. To address this limitation, we explicitly incorporate reasoning into this task and propose a framework that equips vision-language models (VLMs) with chain-of-thought (CoT) capabilities. We first design a scalable CoT curation pipeline that bootstraps rationales from a strong LLM through a two-round strategy and employs a VLM-based verifier to filter out incorrect cases, yielding a high-quality set of (CoT, answer) pairs. Building on this foundation, we adopt a two-stage training scheme: supervised fine-tuning (SFT) to teach rationale-to-answer generation, followed by GRPO reinforcement learning with answer-grounded, fine-grained rewards to further improve final answer accuracy. Extensive experiments on MapDR show that our approach significantly improves both interpretability and accuracy, establishing the first reasoning-based framework for regulation-aware autonomous driving.",
      "short_abstract": "Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversity and context dependence of traffic signage. Importantly, regulation understanding is not a simple...",
      "published": "2026-07-27",
      "updated": "2026-07-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "recent",
      "age_days": 23,
      "arxiv_url": "https://arxiv.org/abs/2607.24199",
      "pdf_url": "https://arxiv.org/pdf/2607.24199.pdf"
    },
    {
      "id": "2607.23979",
      "anchor": "paper-2607-23979",
      "title": "Mutual Modality Trust with Lightweight Reconstruction Regularization for Fine-grained Tire Pattern Recognition",
      "authors": [
        "Jianning Yang",
        "Jie Fang",
        "Xinda Ma",
        "Zirui Song",
        "Dianwei Wang",
        "Nan Wang"
      ],
      "abstract": "Visual tire recognition serves as a core supporting technique for vehicle safety monitoring, autonomous driving perception and automated automotive maintenance. Existing fine-grained tire recognition techniques suffer from three prominent limitations. They tend to depend on only one visual source, lack the capacity to jointly model spatial and frequency cues for minute tread texture extraction, and suffer severe overfitting given limited annotated tire imagery. This paper proposes a lightweight fine-grained tire pattern recognition method incorporating dual-branch independent inference and enhanced feature fusion to boost recognition performance. The framework employs two task-specialized branches dedicated to tire surface and tread indentation, respectively, to extract modality-specific discriminative features. Each branch conducts independent prediction, while cross-branch feature fusion exploits Mutual Modality Trust (M$^2$T) to realize complementary feature enhancement across two modalities. Besides, a frequency-domain hierarchical guidance module is devised, which leverages bandpass filters to decompose feature maps into high- and low-frequency components and enables fine-grained cross-layer feature modulation. Furthermore, a Lightweight Reconstruction Regularization (LR$^2$) is introduced to retain abundant intrinsic information within feature embeddings, substantially improving feature stability and recognition robustness under limited labeled training data. In addition, we establish a surface-indentation multi-source dataset namely MTire299 for fine-grained tire tread recognition, which covers 299 categories with a total of 14795 paired image samples. Extensive experiments conducted on two public tire datasets validate the superiority and efficacy of the proposed algorithm.",
      "short_abstract": "Visual tire recognition serves as a core supporting technique for vehicle safety monitoring, autonomous driving perception and automated automotive maintenance. Existing fine-grained tire recognition techniques suffer from three prominent limitations. They...",
      "published": "2026-07-27",
      "updated": "2026-07-27",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 23,
      "arxiv_url": "https://arxiv.org/abs/2607.23979",
      "pdf_url": "https://arxiv.org/pdf/2607.23979.pdf"
    },
    {
      "id": "2607.24431",
      "anchor": "paper-2607-24431",
      "title": "InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting",
      "authors": [
        "Qi Zhang",
        "Xinquan Yu",
        "Kaiyi Zhang",
        "Hui Huang"
      ],
      "abstract": "Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong spatial-temporal modeling between the input multi-view frames is still underexplored, which limits the performance of those methods in future 4D forecasting. To address this gap, we introduce a novel framework, InterOCF, for 4D occupancy forecasting that jointly models temporal dynamics in both 3D voxel-based representations and multi-view segmentation sequences, while explicitly incorporating feature interaction between the 2D and 3D branches. Our framework incorporates three core components: 1) A 3D Spatio-Temporal (3DST) module that learns volumetric dynamics from historical voxel states to predict future voxel states; 2) A 2D Spatio-Temporal (2DST) module employing an auxiliary multi-view temporal segmentation forecasting task to enhance temporal semantic dynamics; 3) A Spatio-Temporal Interaction Modeling (STIM) module that enables feature interaction between 2D and 3D representations. Experiments on the nuScenes, Lyft-Level5, and nuScenes-Occupancy datasets show that InterOCF consistently outperforms existing baseline approaches.",
      "short_abstract": "Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong...",
      "published": "2026-07-27",
      "updated": "2026-07-27",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: occupancy",
          "abstract: occupancy",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "recent",
      "age_days": 23,
      "arxiv_url": "https://arxiv.org/abs/2607.24431",
      "pdf_url": "https://arxiv.org/pdf/2607.24431.pdf"
    },
    {
      "id": "2607.24336",
      "anchor": "paper-2607-24336",
      "title": "Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination",
      "authors": [
        "Nicole Hu",
        "Mingtao Zhang",
        "Haoyang LI",
        "Chen Jason Zhang",
        "Li Qing"
      ],
      "abstract": "City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed across trips and regions. We conduct a distributional audit on three real-city road-network and taxi-demand datasets from Manhattan, Chicago, and San Francisco. The audit reveals pervasive trip-length inequity whose direction depends on the city and coordinator. After accounting for trip length, spatial inequity becomes more pronounced as demand grows and is consistently stronger when trips are grouped by origin rather than destination. These findings motivate SPatially Aware RErouting (SPARE), a budgeted online coordination framework that assigns limited replanning capacity to delayed vehicles and redirects them using recently observed waiting pressure. SPARE provides a per-review decision guarantee and explicitly bounds online route updates. Experiments on all three datasets against six representative baselines show that SPARE delivers the strongest joint efficiency-fairness performance while retaining city-scale scalability. The results demonstrate that bounded congestion-responsive rerouting improves performance and equity without full-fleet replanning.",
      "short_abstract": "City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed across trips and regions. We conduct a distributional audit on three real-city road-network and taxi-demand...",
      "published": "2026-07-27",
      "updated": "2026-07-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 23,
      "arxiv_url": "https://arxiv.org/abs/2607.24336",
      "pdf_url": "https://arxiv.org/pdf/2607.24336.pdf"
    },
    {
      "id": "2607.24863",
      "anchor": "paper-2607-24863",
      "title": "Steeringless Drifting: Differential-Torque Control of a Four-Wheel Independently Driven Vehicle",
      "authors": [
        "Sheng Zhao",
        "Zexin Wu",
        "Dongyang Zhou",
        "Bolin Zhao",
        "Xiaodong Wu"
      ],
      "abstract": "Control methods for emerging vehicle chassis architectures are important for autonomous driving near handling limits. Unlike conventional drift control, which relies on mechanical steering and rear-tire saturation, a steering-free four-wheel independently driven (4WID) vehicle can generate direct yaw moment through differential wheel torques. This paper proposes a differential-torque drift control method for such a vehicle. A double-track vehicle model incorporating four-wheel differential actuation is established, based on which a drift-equilibrium calculation method and a closed-loop drift controller are developed. The proposed approach is validated through simulations and experiments on a 1:10-scale vehicle. The results show that the vehicle can achieve steady circular drifting with a sideslip angle of approximately 20$^\\circ$ and perform figure-eight drift tracking. This study demonstrates the feasibility of drift control using only differential wheel torques and provides a new perspective on near-limit control for steering-free vehicle architectures.",
      "short_abstract": "Control methods for emerging vehicle chassis architectures are important for autonomous driving near handling limits. Unlike conventional drift control, which relies on mechanical steering and rear-tire saturation, a steering-free four-wheel independently...",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 13,
        "evidence": [
          "abstract: controller",
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.24863",
      "pdf_url": "https://arxiv.org/pdf/2607.24863.pdf"
    },
    {
      "id": "2607.23758",
      "anchor": "paper-2607-23758",
      "title": "RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction",
      "authors": [
        "Han Jiao",
        "Chen Liu",
        "Jiakai Sun",
        "Zhanjie Zhang",
        "Mengyuan Yang",
        "Yimeng Li",
        "Mofan Zhou",
        "Kun Zhan",
        "Lei Zhao"
      ],
      "abstract": "Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require per-scene training and scene-dependent coverage design around the driving trajectory, limiting scalable reconstruction over newly collected roads. To address these limitations, we introduce RoadVGGT, a road-structure-aware feed-forward framework that reconstructs compact Gaussian road surfaces without test-time per-scene optimization. RoadVGGT uses a geometric foundation model to exploit multi-view images together with provided pose and depth observations, and predicts dense pixel-aligned Gaussian attributes through a learned Gaussian head. To make these dense predictions usable for large road surfaces, we align them into a consistent metric world coordinate system and fuse redundant Gaussians on the road-aligned XY plane through confidence-weighted grid fusion. Category-aware grouping and road--sidewalk junction protection further control fusion around vulnerable road structures. The resulting representation supports RGB and semantic bird's-eye-view maps, elevation estimation, and novel view synthesis. RoadVGGT eliminates the need for per-scene optimization in prior methods, reconstructs complete road surfaces with a compact Gaussian representation, and improves image quality, semantic mapping, and elevation accuracy. Extensive experiments demonstrate the potential of geometric foundation models for scalable feed-forward road surface reconstruction.",
      "short_abstract": "Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require...",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: mapping",
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.23758",
      "pdf_url": "https://arxiv.org/pdf/2607.23758.pdf"
    },
    {
      "id": "2607.23537",
      "anchor": "paper-2607-23537",
      "title": "ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness",
      "authors": [
        "Qiao Yan",
        "Yihan Wang",
        "Zhenghao Xing",
        "Jiaqi Xu",
        "Pheng-Ann Heng"
      ],
      "abstract": "Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models behave under real-world adverse weather with multi-modal inputs. We argue that a key difficulty lies in degraded environmental observability: under fog, rain, snow, and low illumination, multi-modal observations become unreliable and cross-modally inconsistent, posing challenges to scene understanding, and subsequent decision-making. To study this, we introduce \\textbf{ObsDriveBench}, a real-world multi-modal benchmark for adverse-weather autonomous driving. Our benchmark is designed with three capability dimensions: \\textbf{observability awareness}, \\textbf{spatial reliability}, and \\textbf{risk-aware decision-making}, enabling fine-grained diagnosis of model behavior under degraded observations. We construct the benchmark through observability meta-annotation, scene description, and capability oriented multiple-choice tasks over synchronized camera, LiDAR, and radar inputs, forming a benchmark with over 14k training and 13k test questions. Experiments reveal consistent performance degradation of existing vision-language models. We further introduce \\textbf{ObsDrive} model with normal-weather supervised fine-tuning and adverse-weather reinforcement learning, improving robustness across all three capabilities. The dataset and evaluation code will be released at \\href{https://github.com/russellyq/ObsDriveBench}{\\texttt{ObsDriveBench}}.",
      "short_abstract": "Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models...",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 5,
        "evidence": [
          "abstract: LiDAR",
          "abstract: radar",
          "abstract: camera",
          "abstract: scene understanding"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.23537",
      "pdf_url": "https://arxiv.org/pdf/2607.23537.pdf"
    },
    {
      "id": "2607.23511",
      "anchor": "paper-2607-23511",
      "title": "MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving",
      "authors": [
        "Zhijing Cheng",
        "Xuancheng Zhang",
        "Donglin Di",
        "Lei Fan",
        "Baorui Ma",
        "Hao Li",
        "Xun Yang"
      ],
      "abstract": "End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories conditioned on this context. We argue that this one-way perception-to-planning interface forces sensor inputs into a compact representation, losing the fine-grained details critical for planning. Moreover, by constraining the planner to this compressed context, it is difficult to leverage the rich representations offered by modern vision foundation models. To address these issues, we propose MOJITO, a unified sensor-to-action framework for end-to-end autonomous driving built on modal joint learning. MOJITO removes the cascaded interface and instead performs block-wise Modal Joint Attention that simultaneously updates action, image, and LiDAR features, allowing the planner to directly access multi-modal features during action generation. MOJITO achieves 88.9 PDMS on the NAVSIM v1 dataset and 88.4 EPDMS on the more challenging NAVSIM v2 dataset, setting a new state-of-the-art. Extensive experiments further demonstrate strong scalability, instruction following, and diverse trajectory generation. Code and models are available at https://github.com/mumucc01/MOJITO.",
      "short_abstract": "End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories conditioned on this context. We argue that...",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 30,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.23511",
      "pdf_url": "https://arxiv.org/pdf/2607.23511.pdf"
    },
    {
      "id": "2607.23962",
      "anchor": "paper-2607-23962",
      "title": "Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation",
      "authors": [
        "Mohammed Aldeen",
        "Muhammad Sami Irfan",
        "Sagar Dasgupta",
        "Long Cheng",
        "Mizanur Rahman",
        "Mashrur Chowdhury"
      ],
      "abstract": "Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attacks that can covertly redirect vehicles or induce unsafe maneuvers. In this paper, we develop the first Vision-Language Model (VLM)-based framework for GNSS spoofing detection for autonomous vehicles by fusing front-camera visual data with in-vehicle sensor readings (e.g., speed, acceleration, yaw rate) against GNSS-derived maneuvers. Our approach introduces a three-stage fine-tuning process that first grounds visual cues, and then calibrates sensor data within a shared semantic space to detect discrepancies between predicted and GNSS-derived maneuvers across three attack scenarios. We also generated an independent real-world dataset by driving an instrumented vehicle on public roads in Tuscaloosa, Alabama, equipped with time-synchronized GNSS, IMU, and camera logs to validate cross-regional generalization of our fine-tuned model on unseen data from training data. On this dataset, we then generated intelligent spoofing attacks, including trajectory mirroring with road-network snapping for wrong-turn attacks, position freezing for overshoot scenarios, and drift generation for stop attacks. On this validation dataset, the zero-shot VLMs baseline F1-score ranges from 23% to 32%, whereas our fine-tuned model achieves an F1-score ranging from 94% to 95%. Results show that our VLM-based approach correctly classified every wrong-turn and stop attacks, and attains 88%-93% accuracy for overshoot attacks. Furthermore, we introduce an adaptive inference policy that reduces VLM invocations to 14% (~86% computational reduction) and yields 65ms-73ms per 4s window. These results point to a practical, on-road layer of defense that complements signal-level integrity checks with the use of VLMs.",
      "short_abstract": "Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attacks that can covertly redirect vehicles or induce unsafe maneuvers. In this paper, we develop the first...",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 19,
        "margin": 6,
        "evidence": [
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.23962",
      "pdf_url": "https://arxiv.org/pdf/2607.23962.pdf"
    },
    {
      "id": "2607.23910",
      "anchor": "paper-2607-23910",
      "title": "SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception",
      "authors": [
        "Goodarz Mehr",
        "Sepideh Gohari",
        "Montasir Abbas",
        "Azim Eskandarian"
      ],
      "abstract": "Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range. However, the development of robust V2X algorithms, particularly those relying on unified spatial representations like bird's-eye view (BEV) representation, is hampered by the lack of large-scale, multi-modal, multi-task datasets. Moreover, collecting and annotating a large set of synchronized, real-world multi-agent data is prohibitively expensive. This has resulted in a landscape where existing V2X datasets are notably limited in both size and scope. To overcome this, we introduce SimBEV2X, an advanced synthetic data generation tool built on the CARLA simulator. SimBEV2X automatically creates randomized driving scenarios to collect multi-modal sensor data alongside various types of ground truth including 3D bounding boxes with unique track IDs, HD map information, BEV segmentation maps, and semantic occupancy voxel grids from both vehicles and RSUs. We also present the SimBEV2X dataset, the largest V2X perception dataset to date. The dataset comprises 258 scenes, each involving up to 8 connected vehicles and up to 4 RSUs across a variety of road networks. The SimBEV2X dataset is an order of magnitude larger than existing V2X datasets and contains 102,200 frames, 588,520 lidar point clouds, more than 3 million images, over 27 million bounding boxes, and a comprehensive set of other annotations. Finally, we establish a strong baseline on the SimBEV2X dataset using CoopDet3D and propose CoBEVFusion, a novel architecture that combines CoopDet3D with fused axial attention (FAX) for context-aware multi-agent feature aggregation, resulting in superior performance. SimBEV2X, the SimBEV2X dataset, and CoBEVFusion are available at https://simbev2x.org and https://github.com/GoodarzMehr/SimBEV2X.",
      "short_abstract": "Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range. However, the development of robust V2X algorithms,...",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Dataset",
        "Simulation",
        "Synthetic Data",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 20,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.23910",
      "pdf_url": "https://arxiv.org/pdf/2607.23910.pdf"
    },
    {
      "id": "2607.23792",
      "anchor": "paper-2607-23792",
      "title": "GNN-based Multi-Agent Control of Traffic Shockwaves in Sparse Vehicular Ad-hoc Networks",
      "authors": [
        "Prachi Nandi",
        "Madhuri Malakar",
        "Sonakshi Satpathy",
        "Pabitra Mohan Khilar"
      ],
      "abstract": "Traffic shockwaves are stop-and-go waves that propagate upstream through the streams of vehicles and are one of the major causes of traffic congestion, fuel inefficiency, and increased accident rates in modern transportation systems. Although Connected and Autonomous Vehicles (CAVs) offer a promising opportunity to mitigate such shockwaves, most existing control strategies rely on global traffic state information, making them impractical for early-stage deployment of Vehicular Ad-hoc Networks (VANETs). In this paper, we propose a decentralized Multi-Agent Reinforcement Learning (MARL) framework that integrates a Graph Neural Network (GNN) to enhance the control architecture of connected and autonomous vehicles. The proposed approach enables vehicles to learn cooperative control policies using locally available information and interaction with neighboring vehicles. The effectiveness of the proposed scheme is evaluated using a scalable simulation environment under realistic highway traffic conditions. Simulation results show that the proposed GNN-based MARL framework can reduce the propagation of traffic shockwaves by up to 80%, even when only 10% of the vehicles are connected.",
      "short_abstract": "Traffic shockwaves are stop-and-go waves that propagate upstream through the streams of vehicles and are one of the major causes of traffic congestion, fuel inefficiency, and increased accident rates in modern transportation systems. Although Connected and...",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 10,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.23792",
      "pdf_url": "https://arxiv.org/pdf/2607.23792.pdf"
    },
    {
      "id": "2607.23755",
      "anchor": "paper-2607-23755",
      "title": "DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation",
      "authors": [
        "Jianhan Lin",
        "Yuchu Qin",
        "Jiateng Yuan",
        "Wenbo Zhang",
        "Shuai Gao"
      ],
      "abstract": "Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.",
      "short_abstract": "Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation....",
      "published": "2026-07-26",
      "updated": "2026-07-26",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 18,
        "margin": 6,
        "evidence": [
          "title: pose estimation",
          "abstract: pose estimation",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "recent",
      "age_days": 24,
      "arxiv_url": "https://arxiv.org/abs/2607.23755",
      "pdf_url": "https://arxiv.org/pdf/2607.23755.pdf"
    },
    {
      "id": "2607.23365",
      "anchor": "paper-2607-23365",
      "title": "On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems",
      "authors": [
        "Muhammad Tukur",
        "Hayatullahi B. Adeyemo",
        "Tao Chen",
        "Nour Ali",
        "Anis Zarrad",
        "Rick Kazman",
        "Marco Agus",
        "Rami Bahsoon"
      ],
      "abstract": "Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education. While these systems offer powerful data-driven and adaptive capabilities, their complexity, rapid evolution, and dependence on dynamic data pipelines introduce new forms of engineering liability collectively referred to as AI Technical Debts (AITDs). AITDs arise from root causes spanning data governance, model implementation, algorithm design, architectural decisions, operational processes, documentation practices, and testing adequacy. Unlike conventional technical debt, many AITDs are latent and propagate across tightly coupled AI pipelines, leading to maintenance challenges, reliability degradation, and heightened safety or security risks. Guided by the principles of AI Trust, Risk, and Security Management (AI TRiSM), this study reinterprets technical debt through the interconnected dimensions of trustworthiness, focusing on AI safety and security technical debts. We conduct a systematic review of 60 primary studies and identify 31 distinct types of AITD, which are organized into a root-cause-oriented taxonomy comprising seven classes. The analysis examines how these debts map to 18 trust-related concerns, including 6 safety hazards and 12 security vulnerabilities. To support mitigation, the review synthesizes 34 actionable guidelines (8 safety and 26 security) targeting the prevention, detection, and reduction of AITDs across the AI lifecycle. Building on these findings, we introduce AITD-MAP, an integrated framework that connects the AITD taxonomy, quality and risk impacts, and mitigation strategies into a unified structure for risk-aware AI engineering. The framework aims to assist AI software engineers in making AI safety and security technical debts visible, understanding their root causes, and mitigating their presence.",
      "short_abstract": "Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education. While these systems offer powerful data-driven and adaptive capabilities, their complexity, rapid evolution,...",
      "published": "2026-07-25",
      "updated": "2026-07-25",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 43,
        "margin": 43,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: security",
          "abstract: security",
          "abstract: validation or testing",
          "abstract: risk assessment"
        ]
      },
      "recency": "recent",
      "age_days": 25,
      "arxiv_url": "https://arxiv.org/abs/2607.23365",
      "pdf_url": "https://arxiv.org/pdf/2607.23365.pdf"
    },
    {
      "id": "2607.23059",
      "anchor": "paper-2607-23059",
      "title": "GLST: Defending Confidence-Driven V2X Collaborative Perception Against Stealthy Multi-Attacker Feature Injection",
      "authors": [
        "Ji He",
        "Ying Wang",
        "Lijie Zheng",
        "Xinghui Zhu",
        "Yulong Shen",
        "Xiaohong Jiang"
      ],
      "abstract": "Collaborative perception (CP) improves autonomous-driving perception by enabling connected vehicles to exchange intermediate features via V2X. Confidence-driven sparse communication reduces bandwidth by transmitting only perception-critical spatial regions, but creates a security risk: once a collaborator is compromised, malicious features in high-confidence or ego-uncertain regions may be preferentially selected and amplified during fusion. Using Where2comm as a representative framework, we show that the proposed Pretend Benign attack exploits its spatial-confidence mechanism by injecting stealthy perturbations into uncertain yet perception-critical regions, substantially degrading 3D object detection while preserving benign-like feature characteristics. Beyond this attack-framework pair, we identify a broader weakness of existing trust-based defenses: their reliance primarily on a single consistency signal leaves them vulnerable when multiple attackers form a pseudo-consensus that biases trust estimation. We therefore propose Global-Local Structural Trust (GLST), a lightweight defense that assesses collaborator reliability through three complementary perspectives: global feature consistency, multi-scale local residual consistency, and structural consistency with ego-side semantic topology. The resulting trust scores guide feature fusion to suppress unreliable collaborators. Experiments on OPV2V show that GLST achieves competitive performance against single-attacker Pretend Benign attacks and substantially stronger robustness in multi-attacker settings. Under a four-attacker Pretend Benign attack, GLST maintains 0.69 AP@0.3, whereas existing single-signal defenses degrade severely. GLST also remains effective against gradient-based attacks such as PGD, indicating that multi-level trust modeling is essential for securing confidence-driven CP.",
      "short_abstract": "Collaborative perception (CP) improves autonomous-driving perception by enabling connected vehicles to exchange intermediate features via V2X. Confidence-driven sparse communication reduces bandwidth by transmitting only perception-critical spatial regions,...",
      "published": "2026-07-25",
      "updated": "2026-07-25",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 44,
        "margin": 28,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 25,
      "arxiv_url": "https://arxiv.org/abs/2607.23059",
      "pdf_url": "https://arxiv.org/pdf/2607.23059.pdf"
    },
    {
      "id": "2607.23404",
      "anchor": "paper-2607-23404",
      "title": "Transfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization",
      "authors": [
        "Jaewook Lee",
        "Ethan Errington",
        "Christian D. Lorenz",
        "Miao Guo"
      ],
      "abstract": "Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core. Gaussian processes (GPs) are the default choice, but they scale poorly as data accumulate and assume a smooth landscape that molecular and materials search spaces routinely violate. Transfer learning offers an alternative suited to this regime: it learns a representation from abundant cheap data and adapts it to sparse expensive data. Despite its use in property prediction, transfer learning has not been tested as the engine of a closed-loop optimization. Here we benchmark eleven transfer-learning surrogates against four GP methods under an identical selection rule, fidelity budget, and model size, across nine tasks spanning synthetic functions to real chemistry and materials problems. GPs win on smooth, low-dimensional functions but perform worst on molecular and materials problems, where transfer-learning surrogates reach substantially better solutions using far less computation. Because acquisition policy is held fixed across surrogates, this advantage is attributable to the surrogate itself. Uncertainty-driven exploration is not reliably beneficial, and calibration does not predict optimization performance, so greedy exploitation of the transfer-learned mean is the more robust default. Transfer learning is therefore the surrogate of choice for molecular and materials MFBO.",
      "short_abstract": "Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core. Gaussian processes (GPs) are the default choice, but...",
      "published": "2026-07-25",
      "updated": "2026-07-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Safety, Security & Verification: abstract: robustness",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "recent",
      "age_days": 25,
      "arxiv_url": "https://arxiv.org/abs/2607.23404",
      "pdf_url": "https://arxiv.org/pdf/2607.23404.pdf"
    },
    {
      "id": "2607.23132",
      "anchor": "paper-2607-23132",
      "title": "DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video",
      "authors": [
        "Muhammad Sulthan Adhipradhana",
        "Ehsan Javanmardi",
        "Naren Bao",
        "Manabu Tsukada"
      ],
      "abstract": "Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a pedestrian accident is a fatal issue that can lead to death. Recently, Vision-Language Models (VLMs) have been a promising tool for accident reasoning, yet many VLMs are not grounded in real-life accident response protocols, making them not usable in accident severity assessment off-the-shelf. We introduced DispatchRAG, an accident assessor and dispatcher framework grounded in real-life Japanese traffic-accident response protocols, designed to enhance VLMs to generate an appropriate emergency response during an emergency scenario. Utilizing a RAG-based retrieval mechanism to retrieve the most relevant accident protocol and an LLM-powered reasoner to suggest the most proper response. To support evaluation, we introduce Accident Dispatch Dataset, a comprehensive dataset of accident assessment and emergency response according to Japanese accident response protocols adapted from the MM-AU dataset. We validate our framework on the Accident Dispatch Dataset, showing strong performance across various accident scenarios compared to the baseline VLM, pointing toward integration in autonomous vehicles that can automatically report both their own and nearby accidents.",
      "short_abstract": "Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a pedestrian accident is a fatal issue that can lead to death. Recently, Vision-Language Models (VLMs) have been...",
      "published": "2026-07-25",
      "updated": "2026-07-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "recent",
      "age_days": 25,
      "arxiv_url": "https://arxiv.org/abs/2607.23132",
      "pdf_url": "https://arxiv.org/pdf/2607.23132.pdf"
    },
    {
      "id": "2607.22494",
      "anchor": "paper-2607-22494",
      "title": "CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipation",
      "authors": [
        "Zhishan Tao",
        "Ruoyu Wang",
        "Yucheng Wu",
        "Enjun Du",
        "Yilei Yuan",
        "Sherwin Ho",
        "Yue Su",
        "Jinbo Su",
        "Yi Hong"
      ],
      "abstract": "Collision anticipation in autonomous driving requires not only accurate early warnings but also interpretable reasoning about what risk factors are being tracked and how risk evolves over time. Existing methods fall short in this regard: feature-driven models are opaque, post-hoc explanations often lack fidelity, and concept-based methods are mostly designed for static recognition rather than dynamic driving scenes. We propose CARA (Concept-Aware Risk Attention), an intrinsically interpretable spatio-temporal framework for collision anticipation. CARA derives domain-grounded risk concepts from accident narratives, aligns them with video frames via vision-language similarity, and organizes them into evolving concept trajectories. These trajectories provide explicit risk evidence that guides spatial attention, temporal attention, and anticipation, allowing semantic concepts to directly influence both where the model attends and how it predicts risk over time. By treating semantic risk factors as dynamic intermediate evidence rather than auxiliary post-hoc explanations, CARA tightly couples interpretability with the predictive process. Extensive experiments on three benchmarks show that CARA consistently improves anticipation accuracy and warning earliness over strong baselines, while providing sparse and semantically grounded concept evidence.",
      "short_abstract": "Collision anticipation in autonomous driving requires not only accurate early warnings but also interpretable reasoning about what risk factors are being tracked and how risk evolves over time. Existing methods fall short in this regard: feature-driven models...",
      "published": "2026-07-24",
      "updated": "2026-07-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "recent",
      "age_days": 26,
      "arxiv_url": "https://arxiv.org/abs/2607.22494",
      "pdf_url": "https://arxiv.org/pdf/2607.22494.pdf"
    },
    {
      "id": "2607.22123",
      "anchor": "paper-2607-22123",
      "title": "DB-VIO: Dual-Branch Visual Inertial Odometry with Enhanced Visual-Inertial Representation",
      "authors": [
        "Ziyu Wan",
        "Lin Zhao"
      ],
      "abstract": "Visual inertial odometry (VIO) is essential for accurate 6-DoF motion estimation in mobile robotic systems. Recent learning-based VIO methods have shown promising progress, but they often rely on unified visual--inertial representations and a single temporal model for full-pose estimation, limiting their ability to capture the heterogeneous dynamics of rotation and translation. Moreover, monocular visual features often lack explicit geometric structure, while raw inertial encoding leaves the underlying rotational kinematics implicit, weakening the rotation-related cues in IMU features. To address these issues, we propose DB-VIO, a dual-branch visual inertial odometry framework with enhanced visual--inertial representation. DB-VIO incorporates depth cues to improve monocular visual perception, injects an explicit integrated-attitude prior to strengthen rotation-aware inertial representation, and decouples pose estimation into dedicated rotational and translational branches for motion-specific temporal modeling. Experiments on autonomous driving and aerial robot benchmarks show that DB-VIO achieves state-of-the-art performance, improving the corresponding baselines by 20\\% on KITTI and 33\\% on EuRoC. Notably, under the more agile motion patterns of EuRoC, DB-VIO improves the rotational metric by 65.7\\% over prior methods. These results demonstrate the effectiveness and generalization of DB-VIO across different platforms and motion scenarios.",
      "short_abstract": "Visual inertial odometry (VIO) is essential for accurate 6-DoF motion estimation in mobile robotic systems. Recent learning-based VIO methods have shown promising progress, but they often rely on unified visual--inertial representations and a single temporal...",
      "published": "2026-07-24",
      "updated": "2026-07-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: pose estimation",
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "recent",
      "age_days": 26,
      "arxiv_url": "https://arxiv.org/abs/2607.22123",
      "pdf_url": "https://arxiv.org/pdf/2607.22123.pdf"
    },
    {
      "id": "2607.22078",
      "anchor": "paper-2607-22078",
      "title": "CommandLM: Data driven behavior level descriptor for ego vehicles",
      "authors": [
        "Boris Tokic",
        "Constantin Selzer",
        "Fabian B. Flohr"
      ],
      "abstract": "As autonomous driving systems move toward real-world deployment, interpretable, behavior-level decision-making is essential for safety, trust, and regulation. We introduce CommandLM, a multimodal large language model that generates concise, human-readable behavior descriptions for ego vehicles from fused multi-sensor data. Our model processes temporally fused bird's-eye view representations from LiDAR and multi-camera inputs via a Q-Former adapter connected to a quantized, LoRA-fine-tuned large language model. Trained on our CommandLM-nuScenes dataset, CommandLM produces intent-aware, interpretable captions suitable for planner supervision and safety auditing. Experiments demonstrate strong linguistic and behavioral alignment, achieving CIDEr 0.67, and BERT-F1 0.88, substantially outperforming the BLIP-2 baseline (CIDEr 0.52, BERT-F1 0.86). In human evaluation, 58% of the generated descriptions were rated accurate, efficient and rule-compliant, confirming their real-world plausibility. While the remaining descriptions may not always select the most efficient, goal-oriented behavior, CommandLM's interpretable outputs enable downstream validation systems to identify and correct such cases, making it an effective tool for transparent behavior auditing. These results show that integrating multimodal fusion with language reasoning yields efficient and transparent behavior-level understanding for autonomous driving. We release our code and dataset at: https://github.com/b-tok/CommandLM",
      "short_abstract": "As autonomous driving systems move toward real-world deployment, interpretable, behavior-level decision-making is essential for safety, trust, and regulation. We introduce CommandLM, a multimodal large language model that generates concise, human-readable...",
      "published": "2026-07-24",
      "updated": "2026-07-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Perception & Sensor Fusion: abstract: camera",
          "Perception & Sensor Fusion: abstract: bird's-eye view",
          "Safety, Security & Verification: abstract: safety",
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "recent",
      "age_days": 26,
      "arxiv_url": "https://arxiv.org/abs/2607.22078",
      "pdf_url": "https://arxiv.org/pdf/2607.22078.pdf"
    },
    {
      "id": "2607.22038",
      "anchor": "paper-2607-22038",
      "title": "Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators",
      "authors": [
        "Afzal Ahmad",
        "Gaoyu Mao",
        "Shoubo Hu",
        "Hui-Ling Zhen",
        "Mingxuan Yuan",
        "Xinyu Chen",
        "Wei Zhang"
      ],
      "abstract": "Multi-task inference models share a single backbone across diverse tasks, yet execute identical computation regardless of which task is active - wasting energy and cycles on task-irrelevant operations. We observe that the task command, typically available before inference begins, provides a free signal that can be exploited to skip unnecessary computation at the hardware level. We present a HW/SW co-designed approach in which a lightweight gating network, trained jointly with the backbone, predicts per-tile binary execution masks conditioned on the task input. Each tile corresponds to a fixed group of output channels (the native scheduling granularity of the accelerator), enabling masked tiles to be skipped with zero overhead. This yields a task-dependent reduction in compute, where each command activates only the subset of the network it requires, without changes to the model architecture or inference pipeline. We co-design the full system stack: a command-conditioned training procedure that learns hardware-aligned tile masks under a sparsity objective; an instruction set architecture whose instructions carry per-tile bitmask fields, allowing the hardware to skip masked tiles without software intervention; and a tiled inference accelerator with configurable parallelism, double-buffered memory, and INT8 datapath that natively supports sparse tile execution. We prototype on an AMD/Xilinx Alveo U50 FPGA and evaluate on a closed-loop visuomotor driving task in CARLA autonomous driving simulator. Task-conditional sparsity reduces FLOPs by 66-76% while maintaining driving quality. On-device latency decreases by 51-59%, from 9.12 ms to 3.74-4.44 ms (2.1-2.4x speedup), with energy per inference dropping from 263 to 108-128mJ.",
      "short_abstract": "Multi-task inference models share a single backbone across diverse tasks, yet execute identical computation regardless of which task is active - wasting energy and cycles on task-irrelevant operations. We observe that the task command, typically available...",
      "published": "2026-07-24",
      "updated": "2026-07-24",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 16,
        "evidence": [
          "title: hardware",
          "abstract: hardware",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 26,
      "arxiv_url": "https://arxiv.org/abs/2607.22038",
      "pdf_url": "https://arxiv.org/pdf/2607.22038.pdf"
    },
    {
      "id": "2607.22975",
      "anchor": "paper-2607-22975",
      "title": "AEcroscopyWave: Towards Self-Driving Characterization Platforms for Agentic AI",
      "authors": [
        "Yongtao Liu",
        "Jawad Chowdhury",
        "Ganesh Narasimha",
        "Ralph Bulanadi",
        "Liam Collins",
        "Ruben Millan Solsona",
        "Marti Checa",
        "Asraful Haque",
        "Sumner B. Harris",
        "Stephen Jesse",
        "Rama Vasudevan"
      ],
      "abstract": "The characterization of electronic materials has traditionally been stratified into two distinct regimens: industry-scale automated systems to inspect materials for defects and ensure quality (such as in the semiconductor industry), and highly customized, operator-driven systems requiring human experts. The former offers high throughput but limited flexibility, whereas the latter is heavily bandwidth-limited but provides research-grade discovery capabilities. Recent advances in \"self-driving\" characterization tools offer the potential to bridge the two stratified regimes, by the creation of application program interfaces (APIs) that can control hardware, and the integration of AI methods to incorporate autonomy into the process. Here, we discuss our latest developments in AEcroscopyWave, a custom-built characterization platform for the agentic-AI era, that provides unified control of scanning probe microscopes with programmable peripheral instrumentation, highlighting the design choices that are necessary for maximizing the capability of the system and the ease of use for both human and AI agents. The benefits of making heterogeneous scientific instruments accessible, composable and usable by agents is demonstrated by test cases.",
      "short_abstract": "The characterization of electronic materials has traditionally been stratified into two distinct regimens: industry-scale automated systems to inspect materials for defects and ensure quality (such as in the semiconductor industry), and highly customized,...",
      "published": "2026-07-24",
      "updated": "2026-07-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: hardware",
          "Systems, Deployment & Connectivity: abstract: systems constraints",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "recent",
      "age_days": 26,
      "arxiv_url": "https://arxiv.org/abs/2607.22975",
      "pdf_url": "https://arxiv.org/pdf/2607.22975.pdf"
    },
    {
      "id": "2607.28665",
      "anchor": "paper-2607-28665",
      "title": "Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems",
      "authors": [
        "Bidhya Shrestha",
        "Christos Papadopoulos"
      ],
      "abstract": "Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both native and aftermarket such as Comma.ai's Openpilot. Monitoring systems to independently verify which automated driving system is active are important for safety monitoring, regulatory compliance, insurance assessment, and anomaly detection. In this paper, we first evaluate the effectiveness of three sequence-based classification models: Gated Recurrent Units (GRU), Long Short-Term Memory (LSTM) networks, and a Transformer encoder model for identifying Level 2 automated driving systems using vehicle telematics data alone: Comma Openpilot, Tesla Autopilot, and Cadillac Super Cruise, along with manual driving. All three models achieve strong clean-data performance with macro F1-scores of 0.92 (GRU), 0.90 (LSTM), and 0.93 (Transformer encoder model) when trained on clean data; threat-matched training yields 0.904-0.916 macro F1 with only a modest clean-data penalty. Second, we introduce a modular robustness evaluation framework that simulates realistic telematics degradation through five corruption families at five severity levels (L1-L5). Continuous channels are perturbed using additive white Gaussian noise with cumulative drift, correlated cross-channel noise, and temporal jitter. Binary event signals are subjected to burst loss, delayed transitions, spurious toggles and cross-feature inconsistencies inspired by communication errors. Robustness is measured using macro-F1, which gives equal weight to each class and is suitable for imbalanced multiclass evaluation. Our evaluation reveals a sharp failure-mode split: event-level corruptions reduce macro-F1 only slightly (greater than equal to 0.87 at L5), while temporal jitter collapses macro-F1 to 0.44-0.50 across GRU, LSTM, and Transformer encoder model.",
      "short_abstract": "Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both native and aftermarket such as Comma.ai's Openpilot. Monitoring systems to independently verify which automated driving...",
      "published": "2026-07-24",
      "updated": "2026-07-24",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 10,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 26,
      "arxiv_url": "https://arxiv.org/abs/2607.28665",
      "pdf_url": "https://arxiv.org/pdf/2607.28665.pdf"
    },
    {
      "id": "2607.21526",
      "anchor": "paper-2607-21526",
      "title": "Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving",
      "authors": [
        "Mengshi Qi",
        "Xiaoyang Bi",
        "Xianlin Zhang",
        "Huadong Ma"
      ],
      "abstract": "Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise from two main aspects. Firstly, adverse conditions can distort pixel correspondences and violate the assumptions embedded in the self-supervised loss function, leading to erroneous depth predictions. Secondly, while radar is a widely adopted sensor in adverse weather conditions, the sparse distribution of radar points in the Point of View (POV) poses challenges for self-supervised fusion. To address these issues, we introduce a novel self-training pipeline using unpaired real all-weather data through multi-teacher distillation and robust radar fusion. We propose the Uncertainty-Aware Multi-Teacher Distillation method to generate diverse teacher models with different adverse condition inputs, and then employ uncertainty modeling to weigh the knowledge distillation loss. Additionally, we design the POV-BEV Radar Fusion approach, which leverages camera-pixel ray constraints to establish connections between the camera's Point of View (POV) and the radar's Bird's-Eye View (BEV). This approach enables the utilization of denser radar points, effectively capturing the complementary perspectives of both POV and BEV. Extensive quantitative and qualitative experiments demonstrate the robustness of our proposed method on all-weather datasets, achieving state-of-the-art performance. Our code and models are available at https://github.com/MICLAB-BUPT/RobustDepth.",
      "short_abstract": "Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise from two main aspects. Firstly, adverse conditions can distort pixel...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 12,
        "evidence": [
          "abstract: perception",
          "abstract: radar",
          "abstract: camera",
          "title: depth estimation",
          "abstract: depth estimation",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.21526",
      "pdf_url": "https://arxiv.org/pdf/2607.21526.pdf"
    },
    {
      "id": "2607.21281",
      "anchor": "paper-2607-21281",
      "title": "HGeo-TopoMap: Boosting Topological Mapping with Hierarchical Geometric Priors",
      "authors": [
        "Siyu Li",
        "Kunyu Peng",
        "Di Wen",
        "Beiping Hou",
        "Zhiyong Li",
        "Kailun Yang"
      ],
      "abstract": "Topological maps are key outputs of autonomous driving perception systems, delivering essential road information for path planning. They identify instances such as centerlines and traffic signs, along with their connectivity relationships. Due to the lack of explicit markings for centerlines in real-world environments, the detection of centerline instances remains a significant challenge. To tackle this problem, we propose HGeo-TopoMap, which leverages an explicit prior map and implicit spatial relations to hierarchically boost topological mapping. First, a geometric adaptive learning module is designed for the road structure map obtained via inverse perspective mapping. This module discretely encodes semantic and spatial features from the map, followed by a prior-mask attention mechanism that selectively focuses on informative regions. Then, a geometric consistency learning module is devised, which leverages the geometric properties and spatial relationships of centerlines. Built on the geometry-aware decoder, it enforces spatial consistency by aligning features of centerline instances with identical geometric orientations. The proposed method is evaluated on the OpenLane-V2 dataset across the centerline, lane segment, and robustness benchmarks. Beyond substantial improvements in topological mapping accuracy, the proposed method offers the benefit of enhanced robustness, consistently outperforming baselines under both standard and challenging conditions. The source code and model weights will be made publicly available at https://github.com/lynn-yu/HGeo-TopoMap.",
      "short_abstract": "Topological maps are key outputs of autonomous driving perception systems, delivering essential road information for path planning. They identify instances such as centerlines and traffic signs, along with their connectivity relationships. Due to the lack of...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.21281",
      "pdf_url": "https://arxiv.org/pdf/2607.21281.pdf"
    },
    {
      "id": "2607.21043",
      "anchor": "paper-2607-21043",
      "title": "A Real-Time Generalized Nash Equilibrium Framework for Interaction-Aware Autonomous Driving in Mixed Traffic",
      "authors": [
        "Nouhed Naidja",
        "Mohamed-Cherif Rahal",
        "Steve Pechberti",
        "Stéphane Font",
        "Guillaume Sandou",
        "Marc Revilloud"
      ],
      "abstract": "Safe and efficient navigation in mixed-traffic environments remains a critical challenge for Autonomous Vehicles (AVs), primarily due to the complex interdependence between the AV's decisions and the unpredictable reactions of human drivers. This paper introduces a comprehensive decision-making framework that formulates the driving interaction as a Generalized Nash Equilibrium Problem (GNEP). Unlike decoupled optimization approaches, this framework explicitly models shared safety and geometric constraints, ensuring that the feasibility of the AV's strategy is dynamically linked to the opponent's actions. To solve this non-convex problem in real-time, we propose a dedicated solver based on Particle Swarm Optimization (PSO). The complete architecture was validated on a test track using a real autonomous Renault Zoé interacting with a human driver. Experimental results demonstrate the system's ability to handle critical scenarios by generating comfortable, human-like trajectories. Benchmarks confirm the solver's operational feasibility, achieving convergence in under 50 ms.",
      "short_abstract": "Safe and efficient navigation in mixed-traffic environments remains a critical challenge for Autonomous Vehicles (AVs), primarily due to the complex interdependence between the AV's decisions and the unpredictable reactions of human drivers. This paper...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.21043",
      "pdf_url": "https://arxiv.org/pdf/2607.21043.pdf"
    },
    {
      "id": "2607.20988",
      "anchor": "paper-2607-20988",
      "title": "HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving",
      "authors": [
        "Quanfu Yu",
        "Xian Wu",
        "Hao Xu",
        "Liulong Ma"
      ],
      "abstract": "Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning. In the pre-training stage, HyWorldVLA predicts video latents encoded by a pre-trained video VAE, while simultaneously reconstructing video frames to provide precise pixel-level grounding. During the subsequent co-fine-tuning phase, the model exclusively predicts latent features, which are fed into an action expert to generate trajectories. Extensive experiments on NAVSIM v1 and v2 benchmarks demonstrate that HyWorldVLA significantly outperforms both pixel-based and latent-based world model baselines. Notably, we present the first comprehensive qualitative and quantitative analysis of world model noise robustness in autonomous driving, establishing a new benchmark for evaluating future architectures.",
      "short_abstract": "Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Benchmark",
        "VLA",
        "World Model",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 24,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.20988",
      "pdf_url": "https://arxiv.org/pdf/2607.20988.pdf"
    },
    {
      "id": "2607.20967",
      "anchor": "paper-2607-20967",
      "title": "Update the Unseen Only: Minimizing AoI for Collaborative Perception through Online Learning",
      "authors": [
        "Yanan Ma",
        "Zhuoyi Zhao",
        "Zhengru Fang",
        "Haonan An",
        "Xianhao Chen",
        "Yuguang Fang"
      ],
      "abstract": "While collaborative perception (CP) enhances the safety of autonomous driving, limited bandwidth can cause severe shared data staleness in CP systems. Existing age-of-information (AoI) minimization policies are not well-suited for CP, as they overlook the fact that a vehicle's AoI decreases not only through updates from the source (i.e., a base station) but also through the vehicle's local sensing. To address this issue, we propose a mobility-aware AoI minimization framework for CP that explicitly accounts for vehicles' dynamic sensing ranges. We first derive a closed-form expression for the long-term time average sum AoI within a considered region, accommodating an ever-changing vehicle population and their dynamic sensed areas. Based on this characterization, we develop Local-sensing-aware Max-Weight Scheduling (LocMW), an online learning algorithm designed for sensor information broadcast from a source to vehicles under unknown environmental statistics and delayed observations. We provide performance guarantees demonstrating that LocMW achieves a sublinear cumulative excess AoI compared to the optimal stationary randomized benchmark. Extensive simulations using vehicular trajectory datasets and 3D perception tasks demonstrate that our LocMW policy substantially outperforms competing baselines, reducing the time-averaged sum AoI by up to 31.6% and improving mAP detection accuracy by up to 16.3%.",
      "short_abstract": "While collaborative perception (CP) enhances the safety of autonomous driving, limited bandwidth can cause severe shared data staleness in CP systems. Existing age-of-information (AoI) minimization policies are not well-suited for CP, as they overlook the...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 14,
        "margin": 2,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.20967",
      "pdf_url": "https://arxiv.org/pdf/2607.20967.pdf"
    },
    {
      "id": "2607.20947",
      "anchor": "paper-2607-20947",
      "title": "RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection",
      "authors": [
        "Junsheng Du",
        "Zhaocheng He",
        "Yuhuan Lu"
      ],
      "abstract": "In intelligent transportation systems, roadside 3D object detection provides wide-area perception crucial for traffic understanding, cooperative early warning, and safe autonomous driving. However, existing methods suffer from high sensitivity to camera extrinsics; even slight deviations (whether manifesting as transient jitter or persistent drift) can be significantly amplified by projective geometry. This cascade results in severe feature misalignment and degraded localization. To mitigate this limitation, we propose RECO, a region-aware extrinsic compensation framework that corrects extrinsics using piecewise 6-DoF pose offsets. RECO predicts a learnable range boundary to partition the scene into near and far regions, estimating region-specific pose corrections. A differentiable sigmoid gate then smoothly blends the two compensated geometries to preserve continuous BEV sampling and facilitate stable optimization. To supervise the refinement of extrinsics, we introduce an auxiliary reprojection loss that compares 2D bounding boxes projected from 3D ground truth against 2D annotations, optimizing it jointly with the standard detection objective. Extensive experiments on the DAIR-V2X-I and Rope3D benchmarks under extrinsic perturbations demonstrate consistent improvements over state-of-the-art baselines across both yaw and $z$-axis deviations. RECO also generalizes from transient perturbations to persistent shifts, maintaining highly competitive performance under strict calibration uncertainty.",
      "short_abstract": "In intelligent transportation systems, roadside 3D object detection provides wide-area perception crucial for traffic understanding, cooperative early warning, and safe autonomous driving. However, existing methods suffer from high sensitivity to camera...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 10,
        "evidence": [
          "abstract: V2X",
          "abstract: cooperative systems",
          "title: roadside infrastructure",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.20947",
      "pdf_url": "https://arxiv.org/pdf/2607.20947.pdf"
    },
    {
      "id": "2607.21488",
      "anchor": "paper-2607-21488",
      "title": "Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections",
      "authors": [
        "Gil Lifshits",
        "Igal Bilik",
        "Gilad Katz"
      ],
      "abstract": "Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) systems, which typically struggle with combinatorial action spaces, reliance on privileged information, or rigid agent designs. We propose Master-Agent Proto-plan System (MAPS), a hierarchical deep reinforcement learning (DRL) architecture in which a centralized Master agent generates a compact, continuous embedding, denoted as proto-plan, that encodes a global coordination strategy. Decentralized Worker agents integrate this embedding with local observations to execute vehicle-specific control, decoupling strategic intent from tactical execution and enabling independent optimization of each module. As a proof-of-concept evaluation of this coordination mechanism, we test MAPS across 72 intersection configurations in HighwayEnv. MAPS achieves collision-free navigation while significantly reducing average travel time, outperforming state-of-the-art baselines. The learned proto-plans further exhibit robust generalization: a system trained with three agents achieves a 94% success rate when deployed zero-shot to five-agent scenarios, confirming that proto-plan-based hierarchical learning provides a promising framework for multi-vehicle coordination.",
      "short_abstract": "Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) systems, which typically struggle with combinatorial action spaces, reliance on privileged information, or rigid agent...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control",
          "Planning & Decision-Making: abstract: navigation"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.21488",
      "pdf_url": "https://arxiv.org/pdf/2607.21488.pdf"
    },
    {
      "id": "2607.21243",
      "anchor": "paper-2607-21243",
      "title": "Detectors Learn the Wrong Thing: Shortcut-Resistant Adversarial Training Against Physically Realizable Attacks",
      "authors": [
        "Yuanhao Huang",
        "Yilong Ren",
        "Jinlei Wang",
        "Xuesong Bai",
        "Zheng Zhang",
        "Haiyang Yu"
      ],
      "abstract": "AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure and autonomous vehicle related applications. However, physically realizable adversarial appearances pose a significant reliability challenge for these safety-critical systems. Adversarial training is effective, but repeated co-occurrence between adversarial texture and positive person instances can cause detectors to treat the texture itself as evidence of object presence, forming a patch texture shortcut. The detector may then treat texture as evidence for the target, causing false detections on texture-only inputs and weakening cross attack generalisation. We propose InsCAT, an instance-level contrastive adversarial training framework that prevents detectors from using adversarial texture as an independent decision cue. SICA aligns adversarial person features with matched clean features and separates them from texture-only negatives, while ROPO and Guard maintain online attack pressure and coordinate training. We evaluate eight independently generated attack textures on rendered nuScenes, INRIAPerson, printed garments, and three detector families. InsCAT achieves an average attack AP of 82.3% on rendered nuScenes, exceeding the strongest baseline by 11.1 points.Relative to AT-Mix, texture FPR decreases from 46.9% to 7.3%. Physical tests yield an F1 score of 96.6% and an FPR of 1.8%. Consistent gains across separately trained detectors demonstrate applicability across architectures with direct inference. The findings show that robust physical detection depends on preserving target related evidence while preventing adversarial texture from becoming an independent decision cu",
      "short_abstract": "AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure and autonomous vehicle related applications. However, physically realizable adversarial appearances pose a significant reliability challenge for these...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 19,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.21243",
      "pdf_url": "https://arxiv.org/pdf/2607.21243.pdf"
    },
    {
      "id": "2607.21819",
      "anchor": "paper-2607-21819",
      "title": "Adaptive Driving Style for SAE Level-2 Driving Automation: Minimizing Preference Mismatch",
      "authors": [
        "Kumar Akash",
        "Zhaobo Zheng",
        "Teruhisa Misu",
        "Vidya Krishnamoorthy",
        "Mia Dong",
        "Yuni Lee",
        "Gaojian Huang"
      ],
      "abstract": "Driving style is a key factor in the comfort and acceptance of automated vehicle (AV) features. In SAE Level-2 automation, where the driver must supervise the system and remain ready to intervene, mismatches between the automation's driving style and the driver's preference can reduce trust and trigger takeovers. This paper proposes an adaptive driving-style control framework that minimizes such preference mismatch. In a driving-simulator study, we compare fixed, trust-based, and preference-based adaptation heuristics and analyze their effects on preference mismatch and trust. We then train a driving-preference prediction model and use it in an implicit adaptation policy that selects among bounded driving styles for upcoming events. A validation study shows that the predictive policy achieves equal or lower preference mismatch than comparison baselines, particularly when starting from a less defensive style, while also yielding higher average trust. The results provide a step toward developing human-aware driving automation that can implicitly adapt its driving style to the driver's preferences.",
      "short_abstract": "Driving style is a key factor in the comfort and acceptance of automated vehicle (AV) features. In SAE Level-2 automation, where the driver must supervise the system and remain ready to intervene, mismatches between the automation's driving style and the...",
      "published": "2026-07-23",
      "updated": "2026-07-23",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "recent",
      "age_days": 27,
      "arxiv_url": "https://arxiv.org/abs/2607.21819",
      "pdf_url": "https://arxiv.org/pdf/2607.21819.pdf"
    },
    {
      "id": "2607.20175",
      "anchor": "paper-2607-20175",
      "title": "PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving",
      "authors": [
        "Yushan Liu",
        "Tianxiong Lv",
        "Bohua Wang",
        "Hangqi Fan",
        "Chenxu Zhao",
        "He Zheng",
        "Xuchang Zhong",
        "Yifan Xie",
        "Congyang Zhao",
        "Zhihao Liao",
        "Leigang Luo",
        "Yang Cai",
        "Xiao-Ping Zhang",
        "Wenbo Ding"
      ],
      "abstract": "Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as the prior-to-plan transfer problem and introduce PerceptDrive, a perception prior world-action modeling framework with adaptive expert routing. PerceptDrive feeds teacher-distilled priors from a frozen, driving-adapted provider and dense observation latents from a frozen self-supervised video encoder into a trainable expert-routed world-action model. Expert-specific query branches process these signals, while a prior-retention objective anchors each branch to its prior. A router predicts soft gates from a shared scene representation and combines the expert conditions before trajectory generation. During training, privileged rule-based sub-metric estimates for branch-specific trajectory drafts provide soft-gate distillation targets. The predicted action-free future latent conditions a flow-matching actor. At inference, privileged components are absent; with one front-facing camera, PerceptDrive generates one trajectory per planning step without test-time scoring, reranking, or search. Experiments show that PerceptDrive achieves state-of-the-art performance with 90.4 PDMS on NAVSIM v1 and 90.2 EPDMS on NAVSIM v2, outperforming existing methods. Ablations confirm complementary gains from prior retention and scene-conditioned routing, alongside differential reliance on the three priors. These results demonstrate that preserving and adaptively routing perception priors improves direct planning without test-time candidate selection.",
      "short_abstract": "Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as...",
      "published": "2026-07-22",
      "updated": "2026-07-22",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 13,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 28,
      "arxiv_url": "https://arxiv.org/abs/2607.20175",
      "pdf_url": "https://arxiv.org/pdf/2607.20175.pdf"
    },
    {
      "id": "2607.20071",
      "anchor": "paper-2607-20071",
      "title": "GaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy Prediction",
      "authors": [
        "Xinzhuo Li",
        "Xianghui Pan",
        "Jiayuan Du",
        "Wei Wei",
        "Liuyi Wang",
        "Chengju Liu",
        "Qijun Chen"
      ],
      "abstract": "Vision-centric 3D occupancy prediction provides dense scene representations essential for autonomous driving and robotic navigation, yet existing methods struggle to scale to high voxel resolutions due to prohibitive computational costs. To address this, we introduce GaussianSeed, a progressive multi-scale Gaussian occupancy prediction framework that organizes primitives into a coarse-to-fine hierarchy. Benefiting from this hierarchical design, GaussianSeed effectively circumvents the memory bottlenecks inherent in dense representations, successfully scaling to a $0.1\\text{m}$ spatial resolution while maintaining real-time inference capabilities. To comprehensively evaluate high-resolution geometric perception, we further construct TJScenes, a panoramic six-camera occupancy dataset with highly detailed $0.1\\text{m}$ annotations. Extensive experiments on Occ3D-nuScenes and TJScenes demonstrate that GaussianSeed delivers the lowest latency among all evaluated methods while maintaining highly competitive accuracy, advancing the efficiency-quality frontier of high-resolution 3D occupancy prediction. Codes are available at https://github.com/Athameral/GUSD",
      "short_abstract": "Vision-centric 3D occupancy prediction provides dense scene representations essential for autonomous driving and robotic navigation, yet existing methods struggle to scale to high voxel resolutions due to prohibitive computational costs. To address this, we...",
      "published": "2026-07-22",
      "updated": "2026-07-22",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 5,
        "evidence": [
          "abstract: perception",
          "title: occupancy",
          "abstract: occupancy",
          "abstract: camera"
        ]
      },
      "recency": "recent",
      "age_days": 28,
      "arxiv_url": "https://arxiv.org/abs/2607.20071",
      "pdf_url": "https://arxiv.org/pdf/2607.20071.pdf"
    },
    {
      "id": "2607.19911",
      "anchor": "paper-2607-19911",
      "title": "LoRFT: Benchmarking Long-Range Vehicle Trajectory Reconstruction from Fixed Highway Cameras",
      "authors": [
        "Yufan Zhu",
        "Kefu Yi",
        "Xueju Zhang",
        "Yunyang Tian",
        "Long Chen",
        "Zixuan Xiao"
      ],
      "abstract": "Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autonomous driving evaluation, and data-driven traffic management, yet continuously recovering them from fixed highway cameras remains difficult. As vehicles recede into distant road regions, perspective compression and scale decay often fragment or prematurely terminate automatic tracklets, even when their continuation remains identifiable from motion consistency across neighboring frames. We formulate this problem as recovering the far-range continuation of a vehicle trajectory from a reliable near-field tracklet. We introduce LoRFT, to our knowledge the first open benchmark dedicated to long-range vehicle trajectory reconstruction from fixed highway cameras. LoRFT comprises 22 expressway surveillance scenes, 366,109 video frames, 6,601 manually verified trajectories, 2,694,889 bounding boxes, road-geometry annotations, scene-level splits, and evaluation scripts. We further propose Map-RSTNet, a map-aware residual sequence-to-sequence model that reconstructs distant trajectories in a road-geometry-aligned state space and dynamically refreshes local road geometry during decoding. On LoRFT, Map-RSTNet reduces ADE, FDE, and 5-second RMSE by 11.0%, 15.4%, and 10.5%, respectively, relative to the strongest baseline. These results demonstrate that road-geometry-aware reconstruction can extend usable trajectory records from existing fixed-camera infrastructure. LoRFT provides a reproducible testbed for long-range vehicle trajectory reconstruction.",
      "short_abstract": "Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autonomous driving evaluation, and data-driven traffic management, yet continuously recovering them from fixed highway cameras remains difficult. As...",
      "published": "2026-07-22",
      "updated": "2026-07-22",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "recent",
      "age_days": 28,
      "arxiv_url": "https://arxiv.org/abs/2607.19911",
      "pdf_url": "https://arxiv.org/pdf/2607.19911.pdf"
    },
    {
      "id": "2607.19781",
      "anchor": "paper-2607-19781",
      "title": "WASABI: Whole-graph Assignment-based Stabilizer for lAne topology By Inter-frame tracking",
      "authors": [
        "Tetsuhiro Uchida",
        "Myu Sasaki",
        "Kensho Nakajima",
        "Yasuhiro Shimada",
        "Toru Saito"
      ],
      "abstract": "Autonomous driving requires understanding the road as a graph of drivable lanes and their connectivity, beyond the ego lane alone, to follow routes through intersections and reason about cross- and merging-traffic. Recent perception models infer such lane topology, i.e., lane segments together with their inter-lane connectivity (LCLC), from onboard sensors over a 360-degree BEV view. Due to neural perception's imperfections, their outputs retain structural instabilities such as missed detections, lost or incorrect LCLC, over-detection, and label flicker. This paper presents WASABI, a real-time post-processing pipeline that stabilizes lane topology outputs both within and across frames by treating lane segments and their LCLC connectivity as joint tracking targets, under onboard real-time constraints (10 Hz / 20 ms / up to 200 input lanes). The pipeline integrates segment tracking with connectivity, noise-robust topology-aware refinement, and a resource-constrained real-time design. On internal validation data (16 sequences), WASABI improves LCLC detection F1 from 0.834 to 0.948 (+0.114, +13.6%) and reduces centerline lateral error from 2.50 m to 0.95 m, while reducing detection false-positives by 24.6%. Temporal-stability metrics on the same data show LCLC toggle rate reduced by 63.3% and boundary-label flicker rate by 30.2%, confirming across-frame stabilization beyond per-frame accuracy.",
      "short_abstract": "Autonomous driving requires understanding the road as a graph of drivable lanes and their connectivity, beyond the ego lane alone, to follow routes through intersections and reason about cross- and merging-traffic. Recent perception models infer such lane...",
      "published": "2026-07-22",
      "updated": "2026-07-22",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: bird's-eye view",
          "Safety, Security & Verification: abstract: validation or testing",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 28,
      "arxiv_url": "https://arxiv.org/abs/2607.19781",
      "pdf_url": "https://arxiv.org/pdf/2607.19781.pdf"
    },
    {
      "id": "2607.19774",
      "anchor": "paper-2607-19774",
      "title": "Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving",
      "authors": [
        "Nuoran Li",
        "Zhang Zhang",
        "Yueran Zhao",
        "Tianze Wang",
        "Chao Sun"
      ],
      "abstract": "Vehicle-to-everything-aided autonomous driving (V2X-AD) significantly enhances driving performance through information sharing. However, existing collaborative perception methods only optimize module-level perception capabilities and fail to effectively serve the ultimate planning and control tasks. We propose an end-to-end collaborative driving system that directly optimizes planning task performance. The system employs MotionNetwork to fuse historical temporal information, utilizes attention mechanisms to efficiently compress spatial features into compact tokens, and adaptively fuses multi-agent features through an autoregressive decoder. Additionally, we introduce Mixture-of-Experts (MoE) architecture to enhance the model's representation capacity for heterogeneous features. Experiments demonstrate that our method achieves a driving score of 79.72, surpassing the state-of-the-art CoDriving baseline (77.15) by 3.33% in closed-loop evaluation while maintaining communication efficiency.",
      "short_abstract": "Vehicle-to-everything-aided autonomous driving (V2X-AD) significantly enhances driving performance through information sharing. However, existing collaborative perception methods only optimize module-level perception capabilities and fail to effectively serve...",
      "published": "2026-07-22",
      "updated": "2026-07-22",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 17,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "recent",
      "age_days": 28,
      "arxiv_url": "https://arxiv.org/abs/2607.19774",
      "pdf_url": "https://arxiv.org/pdf/2607.19774.pdf"
    },
    {
      "id": "2607.20880",
      "anchor": "paper-2607-20880",
      "title": "MatDiffract: A Material-Informed Automated Analysis Platform for X-ray Powder Diffraction",
      "authors": [
        "Hongqing Wang",
        "Mingwei Chen",
        "Hongjie Luo",
        "Wen Yin",
        "Fazhi Qi",
        "Xuqing Chai",
        "Miao Liu",
        "Fengyao Hou"
      ],
      "abstract": "High-throughput experimentation and self-driving laboratories are drastically accelerating materials discovery, yet automated interpretation of X-ray powder diffraction (XRPD) data remains a critical rate-limiting step. Conventional search-match workflows rely heavily on expert manual intervention, while pure data-driven machine learning approaches suffer from limited generalizability across chemical systems and lack rigorous crystallographic interpretability. Here we present MatDiffract, a material-informed automated analysis platform for high-throughput XRPD characterization. Built on a first-principles density functional theory (DFT)-derived inorganic crystal structure database, Atomly, MatDiffract constructs a perturbation-augmented simulated diffraction database, embeds multi-scale diffraction features into indexable vectors, and integrates hierarchical vector retrieval with full-pattern fitting Rietveld refinement and quantitative phase fitting. Benchmarked on 875 single-phase experimental patterns, the platform achieves 91.3% Top-1 and 97.2% Top-10 identification accuracy after automated refinement. For binary and ternary multiphase mixtures, it delivers 85.0% and 70.0% Top-1 accuracy with mass fraction mean absolute errors as low as 1.2% and 1.8%, respectively. Beyond mere phase labeling, MatDiffract outputs full crystallographic results including refined structural models, fitted profiles, and quantitative compositions within tens of seconds per sample. Its modular vector-based architecture supports seamless incremental expansion to new material systems, providing an end-to-end solution to close the characterization throughput gap for autonomous materials discovery and high-throughput materials development.",
      "short_abstract": "High-throughput experimentation and self-driving laboratories are drastically accelerating materials discovery, yet automated interpretation of X-ray powder diffraction (XRPD) data remains a critical rate-limiting step. Conventional search-match workflows...",
      "published": "2026-07-22",
      "updated": "2026-07-22",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 28,
      "arxiv_url": "https://arxiv.org/abs/2607.20880",
      "pdf_url": "https://arxiv.org/pdf/2607.20880.pdf"
    },
    {
      "id": "2607.19701",
      "anchor": "paper-2607-19701",
      "title": "SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving",
      "authors": [
        "Jiangfan Liu",
        "Zexuan Cui",
        "Tianyuan Zhang",
        "Zonglei Jing",
        "Zonghao Ying",
        "Yaoyuan Zhang",
        "Jiakai Wang",
        "Xiaoqi Jiang",
        "Aishan Liu",
        "Xianglong Liu"
      ],
      "abstract": "VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However, existing safety-critical scenario generation methods predominantly rely on simulator-based pipelines, which suffer from a substantial sim-to-real gap and often fail to capture realistic, diverse, and unforeseen human-vehicle interaction dynamics. We present SafeGen, a goal-conditioned diffusion framework for safety-critical scenario generation in VLMADs. Our key insight is to formulate scenario generation as a goal-conditioned diffusion process, where a predefined catastrophic end-state serves as a strong supervisory signal, guiding the generation of temporally coherent video trajectories that naturally evolve toward safety-critical outcomes. Building on this formulation, we introduce Context Grounded End State Reasoning, which leverages VLMs to analyze benign driving contexts and infer latent vulnerabilities in human-vehicle interactions, producing structured end-state specifications that induce high-risk scenarios. Conditioned on these targets, we further propose End State Conditioned Video Evolution, which grounds semantic threats into physically plausible visual dynamics. Specifically, we instantiate high-risk agents within the scene via depth-aware geometric projection, followed by boundary-conditioned diffusion to generate intermediate frames with consistent motion patterns and temporal coherence. Extensive experiments across 3 VLMADs demonstrate that SafeGen increases the Judge Overall Score, a metric using a VLM judge to evaluate VLMADs' understanding and decision-making, by 24.25% on average compared to SoTA baselines. Furthermore, fine-tuning a VLMAD improves performance in real-world driving scenes by an average of 15.9%.",
      "short_abstract": "VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However,...",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 18,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.19701",
      "pdf_url": "https://arxiv.org/pdf/2607.19701.pdf"
    },
    {
      "id": "2607.19528",
      "anchor": "paper-2607-19528",
      "title": "D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models",
      "authors": [
        "Heesang Han",
        "A. Lynn Abbott",
        "Abhijit Sarkar"
      ],
      "abstract": "Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid structure for the data. For similar reasons, fusion of camera and LiDAR data within an MLLM pipeline is also uncommon. However, most autonomous systems rely on LiDAR-based sensing, and incorporating 3D data has been proven to improve performance in traditional 3D scene perception tasks. This paper presents D3VL, a novel MLLM framework that integrates 2D and 3D time-series data in a single but simple architecture. The model aims to answer questions involving traffic scene understanding and safety. D3VL shows an 11% improvement in the KITTI Question-Answering (QA) dataset compared to baseline methods in processing 2D and 3D time-series data. This paper further introduces the Waymo QA dataset extension, which assesses models' capabilities in processing 3D and time-series data under diverse driving conditions. D3VL implementation code and WaymoQA extension can be found on our supplemental website: https://automotivesafety-lvlm.github.io",
      "short_abstract": "Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM...",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: scene understanding"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.19528",
      "pdf_url": "https://arxiv.org/pdf/2607.19528.pdf"
    },
    {
      "id": "2607.19284",
      "anchor": "paper-2607-19284",
      "title": "Stochastic Multi-Objective Kinodynamic Planning Against Adversaries",
      "authors": [
        "Thomas Marshall Vielmetti",
        "Daniel Cherenson",
        "Dimitra Panagou"
      ],
      "abstract": "This paper addresses multi-objective kinodynamic planning in environments with stochastic hybrid adversaries that probabilistically transition to adversarial modes based on the ego state. The goal is to construct the Pareto-front of paths that trade off execution cost and the probability of safety constraint violation (risk). Existing chance-constrained planners evaluate risk over open-loop trajectories, yielding overly conservative solutions that fail to account for ego-agent reactivity. To address this limitation, we shift the planning space to sequences of closed-loop policies, and integrate sample-based risk evaluation directly into tree construction via Monte-Carlo particle rollouts. We first introduce Stochastic Multi-Objective RRT (SMO-RRT), for which we prove probabilistic completeness, followed by Stochastic Multi-Objective Stable Sparse RRT (SMO-SST), which leverages selective pruning to improve numerical performance at the cost of completeness. For both algorithms, we derive a finite-sample bound on the probability of chance constraint violation for systems with non-Gaussian, state-dependent uncertainty, enabling probabilistically safe planning in a broad class of environments applicable to multi-agent systems, social navigation, and autonomous driving.",
      "short_abstract": "This paper addresses multi-objective kinodynamic planning in environments with stochastic hybrid adversaries that probabilistically transition to adversarial modes based on the ego state. The goal is to construct the Pareto-front of paths that trade off...",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 4,
        "evidence": [
          "title: planning",
          "abstract: planning",
          "abstract: navigation"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.19284",
      "pdf_url": "https://arxiv.org/pdf/2607.19284.pdf"
    },
    {
      "id": "2607.19194",
      "anchor": "paper-2607-19194",
      "title": "Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency",
      "authors": [
        "Zhongyao Yang",
        "Haoyu Li",
        "Yu Yan",
        "Zhuangxuan Yu",
        "Jiangfeng Nan",
        "Jinrui Nan"
      ],
      "abstract": "High-level planning for autonomous driving is a knowledge-intensive engineering decision task that requires accurate scene understanding, timely inference, and internally consistent action selection. Vision-language models (VLMs) can make intermediate reasoning explicit, but their use in deployed planners is constrained by costly structured supervision, unnecessary reasoning in routine scenes, and possible inconsistencies between generated rationales and driving actions. We present a cognitive dual-process planning framework that represents planning-relevant scene knowledge in a machine-parsable structured chain-of-thought (S-CoT) schema. An automated data engine integrates perception foundation models, critical-path filtering, and an expert VLM to generate S-CoT supervision without manual annotation of individual rationales. A lightweight visual Arbiter estimates scene complexity from multilevel vision-encoder features before language decoding and routes each input to either fast meta-action prediction or slow structured reasoning. For slow-path outputs, a deterministic rule-based validator checks whether the parsed S-CoT fields are consistent with the final meta-action and provides verifiable rewards for Group Relative Policy Optimization (GRPO). In a 195-scene manual audit, the generated annotations achieve 91.8\\% CoT accuracy and a 98.5\\% Logical Consistency Score (LCS). On 574 manually verified NAVSIM test samples, the planner achieves 80.14\\% planning accuracy and 97.20\\% LCS while reducing average latency by 17.39\\% relative to applying slow reasoning to every scene. Evaluation on external long-tail subsets further identifies conditions under which routing and planning performance degrade. Together, these results show how explicit scene knowledge can be operationalized through adaptive reasoning and rule-based verification to support high-level VLM planning decisions.",
      "short_abstract": "High-level planning for autonomous driving is a knowledge-intensive engineering decision task that requires accurate scene understanding, timely inference, and internally consistent action selection. Vision-language models (VLMs) can make intermediate...",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.19194",
      "pdf_url": "https://arxiv.org/pdf/2607.19194.pdf"
    },
    {
      "id": "2607.22714",
      "anchor": "paper-2607-22714",
      "title": "Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems",
      "authors": [
        "Sai Sidharth D"
      ],
      "abstract": "Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power. This paper presents an optimized semantic segmentation architecture derived from the RetinaNet detection framework, adapted for dense pixel-wise prediction and tailored for deployment on resource-constrained embedded hardware. The proposed architecture, termed Opt-RetinaSeg, replaces the standard ResNet-50 backbone with a hybrid lightweight feature extractor, restructures the Feature Pyramid Network (FPN) to reduce redundant multi-scale computation, and introduces a compact segmentation head guided by focal-loss-inspired class balancing to address the severe foreground-background imbalance common in road scenes. We further apply a three-stage optimization pipeline consisting of structured channel pruning, post-training INT8 quantization, and knowledge distillation from a high-capacity teacher network. Evaluated on the Cityscapes and BDD100K datasets and deployed on an NVIDIA Jetson Xavier NX and a Qualcomm QCS610 automotive SoC, the proposed model achieves 73.9% mIoU at 70.4 FPS, representing a 7.4x inference speedup and a 4x reduction in model size relative to the ResNet-50 baseline, with less than 3% accuracy degradation. These results indicate that RetinaNet-derived architectures, when systematically optimized, are viable candidates for real-time semantic segmentation in embedded automotive perception pipelines",
      "short_abstract": "Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power. This paper presents an optimized semantic...",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 19,
        "margin": 2,
        "evidence": [
          "abstract: perception",
          "title: scene segmentation",
          "abstract: scene segmentation"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.22714",
      "pdf_url": "https://arxiv.org/pdf/2607.22714.pdf"
    },
    {
      "id": "2607.19661",
      "anchor": "paper-2607-19661",
      "title": "Leveraging ECRAM for Edge Continual Learning",
      "authors": [
        "Nabila Tasnim",
        "Haoran Liu",
        "Qing Cao",
        "Saugata Ghose"
      ],
      "abstract": "Several edge computing platforms, such as autonomous vehicles and smart sensing devices, need to adapt to dynamic environments in real time by learning from new data in the field. Continual learning has emerged as a promising solution for edge training, by incorporating techniques that successfully combine a highly summarized version of previously trained data (to avoid catastrophic forgetting) with recently sensed data. However, as is the case with other ML algorithms, continual learning generates significant data movement between general-purpose CPUs/GPUs and memory, impacting the suitability of continual learning for edge platforms. In-memory computing (IMC; also known as processing-using-memory) can curtail this waste and make continual learning feasible at the edge, but it faces two unique challenges: (1) IMC architectures make use of noisy computation operations that significantly harm training accuracy; and (2) IMC architectures have poor and often incomplete support for resource-efficient training. To address these challenges, we propose CLASP (the Continual Learning Acceleration System Platform), which to our knowledge is the first end-to-end system with IMC acceleration for continual learning. The hardware and software of CLASP are co-designed to support a wide range of continual learning algorithms, through software-visible assembly-level instructions that can be incorporated without constraints into ML-based algorithms. CLASP is designed around a back-end-of-line (BEOL) compatible ECRAM device that we fabricate, which can overcome the challenges of IMC-based training using other emerging memory devices. We show that CLASP with ECRAM approaches the accuracy of in-GPU training, while delivering a speedup of 67x and energy savings of 132x for learning without forgetting and experience replay using MNIST.",
      "short_abstract": "Several edge computing platforms, such as autonomous vehicles and smart sensing devices, need to adapt to dynamic environments in real time by learning from new data in the field. Continual learning has emerged as a promising solution for edge training, by...",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 7,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware",
          "abstract: edge or cloud"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.19661",
      "pdf_url": "https://arxiv.org/pdf/2607.19661.pdf"
    },
    {
      "id": "2607.19617",
      "anchor": "paper-2607-19617",
      "title": "EGRNet: A Lightweight Semantic Segmentation Network with Edge-Gated Refinement and Adversarial Sensing",
      "authors": [
        "Bareera Qaseem",
        "Mohsin Kamal",
        "Muhammad Naveed Aman"
      ],
      "abstract": "As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understanding becomes increasingly critical. Semantic segmentation plays a key role in enabling autonomous vehicles to comprehend complex urban environments. However, achieving high accuracy with minimal computational cost remains a significant challenge. In this paper, we present Edge-Gated Refinement Network (EGRNet), a lightweight and efficient deep learning model designed for real-time semantic segmentation in urban scenarios. The model incorporates depthwise separable convolutions to reduce computational complexity and dilated residual blocks for capturing rich multi-scale contextual information. Additionally, we introduce a novel Edge-Gated Refinement (EGR) module, which adaptively fuses original and refined features through a learnable gating mechanism, enhancing boundary preservation and edge-sensitive regions. To further improve feature representation, Squeeze-and-Excitation (SE) attention is applied across the network. With only 0.46M parameters, EGRNet achieves state-of-the-art performance while maintaining low computational overhead. When evaluated on the Cityscapes dataset, the model attains a mean Intersection over Union (mIoU) of 65.28%, demonstrating strong accuracy with minimal resource consumption. Moreover, we introduce a lightweight adversarial attack detection strategy, ensuring robustness against adversarial inputs without compromising real-time performance. By combining efficiency, accuracy, and resilience, EGRNet is well-suited for deployment on edge devices in safety-critical real-time applications.",
      "short_abstract": "As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understanding becomes increasingly critical. Semantic segmentation plays a key role in enabling autonomous vehicles to comprehend complex urban environments....",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 22,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.19617",
      "pdf_url": "https://arxiv.org/pdf/2607.19617.pdf"
    },
    {
      "id": "2607.19146",
      "anchor": "paper-2607-19146",
      "title": "Sarus: Privacy-Preserving Multi-Vendor Perception Fusion via Homomorphic Encryption",
      "authors": [
        "Munawar Hasan",
        "Apostol Vassilev"
      ],
      "abstract": "Cooperative perception enables autonomous vehicles (AVs) to improve situational awareness by aggregating detection outputs from multiple agents and sensing platforms, often via a shared fusion service in multi-vendor deployments. However, sharing such outputs at inference time exposes proprietary model behavior and sensitive environmental information, creating significant privacy and security concerns. In this paper, we present Sarus, a privacy-preserving framework for multi-vendor perception fusion via homomorphic encryption (HE), enabling aggregation without revealing individual vendor outputs. Each vendor encodes detections as compact Gaussian moment vectors over a shared spatial lattice and transmits encrypted payloads to a fusion server, which aggregates them directly in the encrypted domain. The fused result is then decrypted and reconstructed into final detections through class-wise bin merging. We analyze the computational complexity, showing linear scaling for vendor payload construction and $O(BV)$ server-side fusion with the number of occupied bins $B$ and vendors $V$, while postprocessing scales as $O(B + \\sum_{c\\in \\mathcal{C}} B_c^2)$, where $\\mathcal{C}$ denotes the set of object classes and $B_c$ is the number of occupied bins for class $c$. Experiments demonstrate linear scaling in practice with only a bounded constant-factor overhead from HE, with decryption dominating postprocessing cost. Experiments on the KITTI dataset using camera (YOLOv8) and LiDAR (PointPillars, PV-RCNN) detectors show that Sarus improves scene-level coverage by effectively aggregating complementary detections, particularly in distance-dependent regimes where individual modalities degrade. These results indicate that privacy-preserving multi-vendor perception fusion is feasible for real-time deployment when statistical compression and spatial sparsity are jointly exploited.",
      "short_abstract": "Cooperative perception enables autonomous vehicles (AVs) to improve situational awareness by aggregating detection outputs from multiple agents and sensing platforms, often via a shared fusion service in multi-vendor deployments. However, sharing such outputs...",
      "published": "2026-07-21",
      "updated": "2026-07-21",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Cooperative / V2X",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 11,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "recent",
      "age_days": 29,
      "arxiv_url": "https://arxiv.org/abs/2607.19146",
      "pdf_url": "https://arxiv.org/pdf/2607.19146.pdf"
    },
    {
      "id": "2607.18663",
      "anchor": "paper-2607-18663",
      "title": "How defensive driving enhances driving safety: A driving simulator study on drivers' defensive driving behaviors",
      "authors": [
        "Xinzheng Wu",
        "Junyi Chen",
        "Shaolingfeng Ye",
        "Yong Shen"
      ],
      "abstract": "Defensive driving is widely recognized as an advanced driving skill. However, whether and how defensive driving affects driving safety remains insufficiently investigated. This study examines the behavioral characteristics of defensive driving, its impact on driving safety, and the underlying mechanisms. First, defensive driving is defined regarding operational timing and application scenario. Then, 82 participants are recruited for driving simulator experiments, with their behavioral and eye movement data being collected. Following the experiments, participants are categorized into groups based on the frequency of defensive driving behaviors exhibited. Finally, both inter-group and inter trial comparisons are performed on the experimental data. Experimental results demonstrate that in the inter-group comparison, the high defensive driving capability group exhibits higher acceleration and deceleration magnitudes, lower average speeds, and larger average absolute yaw angles compared to the low capability group, alongside shorter fixation durations and reduced fixation frequencies. Moreover, we observe that these participants tend to initiate defensive or evasive actions earlier, resulting in lower scenario risk. Regarding the inter-trial comparison, we observe similar trends exclusively in the low capability group, whereas most metrics show no significant differences between Trial 1 and Trial 2 in the high capability group. These results reveal that drivers possessing defensive driving capabilities tend to execute more intense driving maneuvers and identify risks and take action earlier, thereby enhancing driving safety. Findings support the promotion of defensive driving and provide a basis for relevant training programs. Meanwhile, they offer insights for the training of autonomous driving algorithms with defensive driving capabilities.",
      "short_abstract": "Defensive driving is widely recognized as an advanced driving skill. However, whether and how defensive driving affects driving safety remains insufficiently investigated. This study examines the behavioral characteristics of defensive driving, its impact on...",
      "published": "2026-07-20",
      "updated": "2026-07-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 14,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "recent",
      "age_days": 30,
      "arxiv_url": "https://arxiv.org/abs/2607.18663",
      "pdf_url": "https://arxiv.org/pdf/2607.18663.pdf"
    },
    {
      "id": "2607.18637",
      "anchor": "paper-2607-18637",
      "title": "End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation",
      "authors": [
        "Jingzheng Li",
        "Yufei Ge",
        "Zhijun Chen",
        "Qianren Mao",
        "Zizhe Wang",
        "Binhang Qi",
        "Bing Li",
        "Keyu Chen",
        "Baochang Zhang",
        "Xianglong Liu",
        "Philip S Yu"
      ],
      "abstract": "Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control over traffic behavior or controllable scenarios at the expense of behavioral plausibility. This paper presents E2E-CDiff, an end-to-end conditional diffusion framework for controllable and realistic scenario generation. Conditioned on front-view visual observations, E2E-CDiff jointly denoises future motion states and executable low-level controls for route-interacting background vehicles. This unified state-action generation mitigates the planning-control mismatch in conventional two-stage trajectory-then-controller pipelines. Differentiable guidance further regulates speed, enforces drivable-area compliance, and supports collision-avoidance or collision-seeking behaviors, enabling both naturalistic and safety-critical scenario generation. Experiments on Bench2Drive show that E2E-CDiff achieves a favorable controllability-realism trade-off compared with representative reinforcement- and imitation-learning baselines, while its collision-guided variant induces challenging interactions across multiple autonomous driving systems. E2E-CDiff also performs competitively as a learning-based ego planner, demonstrating the generality of end-to-end state-action diffusion.",
      "short_abstract": "Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance...",
      "published": "2026-07-20",
      "updated": "2026-07-20",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "recent",
      "age_days": 30,
      "arxiv_url": "https://arxiv.org/abs/2607.18637",
      "pdf_url": "https://arxiv.org/pdf/2607.18637.pdf"
    },
    {
      "id": "2607.18115",
      "anchor": "paper-2607-18115",
      "title": "Compositional Semantic Communication for Physical AI: Category Theory Meets Game Theory",
      "authors": [
        "Christo Kurisummoottil Thomas",
        "Walid Saad",
        "Emilio Calvanese Strinati"
      ],
      "abstract": "Physical artificial intelligence (AI) systems involve distributed sensing agents with embedded AI models that must coordinate to perceive, reason, and act in networked environments. Transmitting raw sensor data incurs significant communication overhead, latency, and redundancy. While semantic communication (SC) mitigates these challenges by transmitting task-relevant information, existing deep learning-based joint source-channel coding approaches exhibit limited adaptability, poor out-of-distribution generalization, and scalability challenges. To address these limitations, this paper proposes a framework for compositional semantic communication (CSC), enabling heterogeneous physical AI sources to transmit semantic representations (SRs) that compose meaningfully at a base station (BS) or edge server for remote inference. First, a category-theoretic measure of compositional semantics is developed to quantify each device's contribution to inference tasks beyond mutual information. Second, Grothendieck topologies and presheaves formalize semantic composition across devices, ensuring consistency and task relevance. Building on these foundations, multi-device coordination is formulated as a Stackelberg game in which devices commit to encoding strategies and the BS optimally composes received SRs. An ADMM-based algorithm computes equilibrium signaling strategies. Equilibrium existence is established under mild conditions and is Pareto optimal when compositional information yields increasing collective benefit. Simulation results demonstrate that the proposed approach achieves up to 17% bandwidth reduction and 53% lower end-to-end latency than cooperative multi-agent, distributed gradient descent, and uniform-selection CSC baselines while maintaining 85% inference accuracy across diverse autonomous driving scenarios.",
      "short_abstract": "Physical artificial intelligence (AI) systems involve distributed sensing agents with embedded AI models that must coordinate to perceive, reason, and act in networked environments. Transmitting raw sensor data incurs significant communication overhead,...",
      "published": "2026-07-20",
      "updated": "2026-07-20",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 10,
        "evidence": [
          "abstract: cooperative systems",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "recent",
      "age_days": 30,
      "arxiv_url": "https://arxiv.org/abs/2607.18115",
      "pdf_url": "https://arxiv.org/pdf/2607.18115.pdf"
    },
    {
      "id": "2607.17813",
      "anchor": "paper-2607-17813",
      "title": "A2RL V\\textsubscript{max}: The A2RL autonomous racing dataset for long-range, high-speed perception and multi-vehicle interaction",
      "authors": [
        "Marvin Klemp",
        "Dominic Ebner",
        "Cornelius Schröder",
        "Davide Malvezzi",
        "László Turányi",
        "Riccardo Donati",
        "Ilia Schminik",
        "Xia Ning",
        "Yanxin Zhou",
        "Matthew Flagg",
        "Christoph Stiller",
        "Markus Lienkamp",
        "Marko Bertogna",
        "Gergely Bári",
        "Andreas Birk",
        "Ren Jin",
        "Chen Lv",
        "Johannes Betz"
      ],
      "abstract": "In autonomous driving development, a perception dataset is crucial, as it provides fundamental data for training, testing, and validating algorithms for an autonomous vehicle's multimodal perception systems. So far, most research has concentrated on providing datasets for well-structured urban environments. This work introduces the A2RL V\\textsubscript{max} open-source dataset, specifically designed for perception tasks in high-speed autonomous driving and multi-vehicle interaction. The dataset was captured during the 2024 Abu Dhabi Autonomous Racing League (A2RL), held at the Yas Marina F1 Circuit, with participation from all competing teams. It contains diverse scenarios, including single-vehicle data at varying speeds, multi-vehicle sessions, and the full final four-vehicle race. The dataset contains almost 30,000 professionally annotated LiDAR point clouds, along with RADAR point clouds. In particular, it is the first large-scale dataset in autonomous racing to feature professionally annotated LiDAR point clouds, enabling deep learning-based perception research. The data is provided in a developer-friendly format, enabling easy implementation and evaluation in future research. We provide implementation and evaluation for off-the-shelf 3D detection and tracking methods. Although baseline methods show promising results for both 3D detection and tracking, specialized methods are required to address the unique challenges of high-speed autonomous driving. For a detailed description of the dataset, please visit the \\href{https://tum-avs.github.io/A2RL_Dataset_website/}{A2RL V\\textsubscript{max} Dataset Website}",
      "short_abstract": "In autonomous driving development, a perception dataset is crucial, as it provides fundamental data for training, testing, and validating algorithms for an autonomous vehicle's multimodal perception systems. So far, most research has concentrated on providing...",
      "published": "2026-07-20",
      "updated": "2026-07-20",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 18,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: point cloud",
          "abstract: LiDAR",
          "abstract: radar"
        ]
      },
      "recency": "recent",
      "age_days": 30,
      "arxiv_url": "https://arxiv.org/abs/2607.17813",
      "pdf_url": "https://arxiv.org/pdf/2607.17813.pdf"
    },
    {
      "id": "2607.17767",
      "anchor": "paper-2607-17767",
      "title": "VLN-AVP: Zero-Shot Vision-Language Navigation with Hybrid Long-Short-Term Memory for Autonomous Valet Parking",
      "authors": [
        "Yijian Li",
        "Xiangru Mu",
        "Changze Li",
        "Hantian Shi",
        "Jiyuan Cai",
        "Jia Cai",
        "Xiaoxue Liu",
        "Yajing Sun",
        "Ming Yang",
        "Tong Qin"
      ],
      "abstract": "Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts their scalability to unseen environments and open-vocabulary targets. Inspired by the application of Vision-Language Models (VLMs) in Vision-Language Navigation (VLN) tasks, we propose VLN-AVP, a zero-shot navigation framework for AVP tasks. By combining the precise spatial perception of a Bird's-Eye-View (BEV) model with the general intelligence of VLMs, our framework 1) eliminates the dependency on pre-built maps, 2) interprets semantic environmental contexts in parking scenarios, and 3) enables intuitive navigation following natural language instructions. Specifically, we introduce a hybrid memory system: a short-term perception memory tracks semantic visual cues to address the limitations of VLM's single-frame reasoning in existing methods, while a long-term topological memory facilitates stable policy learning from past experiences. To bridge the gap in existing benchmarks, we also present the VLN-AVP dataset and benchmark. Featuring 10 high-fidelity parking scenes and over 1,000 navigation episodes, it has the largest number of garage scenes to date and is the first VLN benchmark for underground parking. Extensive experiments demonstrate that in simulation, our method achieves an over 25% improvement in success rate compared to VLN methods and an over 15% improvement compared to other autonomous driving methods. Furthermore, it attains a leading success rate in real-world vehicle experiments, proving its practical feasibility.",
      "short_abstract": "Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts their scalability to unseen environments and open-vocabulary targets. Inspired by the application of Vision-Language Models (VLMs) in Vision-Language...",
      "published": "2026-07-20",
      "updated": "2026-07-20",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 3,
        "evidence": [
          "title: navigation",
          "abstract: navigation"
        ]
      },
      "recency": "recent",
      "age_days": 30,
      "arxiv_url": "https://arxiv.org/abs/2607.17767",
      "pdf_url": "https://arxiv.org/pdf/2607.17767.pdf"
    },
    {
      "id": "2607.18548",
      "anchor": "paper-2607-18548",
      "title": "Engineering Trustworthy Agentic AI for Critical Systems",
      "authors": [
        "Omar Al-Refai",
        "Ibrahim Shahbaz",
        "Adam Ali Husseinat",
        "Michael Mandulak",
        "Jaewon Kim",
        "Eman Hammad"
      ],
      "abstract": "Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in current literature by treating trustworthiness, whether agentic behavior can be verified, audited, and trusted under the constraints that engineering practice actually requires, as a first-class engineering property, rather than evaluating agentic AI by task capability alone. The study adopts a trustworthiness model organized around five cross-cutting dimensions: safety and constraint satisfaction; robustness and reliability; transparency and interpretability; accountability and auditability; and privacy and security. This is mapped onto an agentic assurance workflow spanning perception through audit. Building on this foundation, agentic systems architectures, threats, concrete trust mechanisms, and quantitative metrics are surveyed for direct application in agentic systems development and evaluation. These principles are then examined across four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks, identifying recurring design patterns, shared failure modes, and domain-specific gaps. Synthesizing across those domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.",
      "short_abstract": "Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey...",
      "published": "2026-07-20",
      "updated": "2026-07-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Survey",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 14,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: attack or threat",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 30,
      "arxiv_url": "https://arxiv.org/abs/2607.18548",
      "pdf_url": "https://arxiv.org/pdf/2607.18548.pdf"
    },
    {
      "id": "2607.18106",
      "anchor": "paper-2607-18106",
      "title": "Importance Sampling and PCA for Finding Failures in Commercial Autonomous Vehicles",
      "authors": [
        "Hailey Warner",
        "Duncan Eddy",
        "Shreya Parjan",
        "Caroline Cahilly",
        "Harrison Delecki",
        "Matthias Kleinstauber",
        "Chaitanya Shinde",
        "Jerry Lopez",
        "Mykel J. Kochenderfer"
      ],
      "abstract": "Methods for discovering rare failures in autonomous systems have so far been demonstrated almost exclusively in simulations with simple, academic driving stacks, leaving open whether they generalize to the more robust planners used in commercial systems. We address this gap by applying two rare-event discovery algorithms to a commercial autonomous trucking stack. Adaptive stress testing (AST) uses reinforcement learning to search for the most likely noise trajectories leading to a simulated collision, while diffusion-based failure sampling (DiFS) trains a denoising diffusion model to sample a diverse set of failures. We show that both algorithms find simulated collisions during merge and cut-in maneuvers where traditional Monte Carlo simulation does not. To make these failures actionable, we introduce a statistical analysis based on principal component analysis (PCA) that classifies failures into common modes and identifies the timesteps that most influence the outcome. We cluster the principal components and invert the PCA transform to recover generalized noise trajectories, and show that these trajectories reproduce failures in identical and similar scenarios. This provides a path from failure discovery to systematic diagnosis of perception-level flaws.",
      "short_abstract": "Methods for discovering rare failures in autonomous systems have so far been demonstrated almost exclusively in simulations with simple, academic driving stacks, leaving open whether they generalize to the more robust planners used in commercial systems. We...",
      "published": "2026-07-20",
      "updated": "2026-07-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 10,
        "evidence": [
          "abstract: validation or testing",
          "abstract: robustness",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "recent",
      "age_days": 30,
      "arxiv_url": "https://arxiv.org/abs/2607.18106",
      "pdf_url": "https://arxiv.org/pdf/2607.18106.pdf"
    },
    {
      "id": "2607.17521",
      "anchor": "paper-2607-17521",
      "title": "GeoWorldAD: Geometry World Action Model for Autonomous Driving",
      "authors": [
        "Songyan Zhang",
        "Jinyuan Tian",
        "Hanbing Li",
        "Daqi Liu",
        "Hao Chen",
        "Wenhui Huang",
        "Fang Li",
        "Guang Chen",
        "Hangjun Ye",
        "Long Chen",
        "Kuiyuan Yang",
        "Chen Lv"
      ],
      "abstract": "Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale training data, they often lack explicit geometric grounding and future-aware spatial guidance, limiting their ability to balance collision avoidance and driving progress. In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Present geometry provides essential spatial constraints for safe planning, while future geometry reveals how surrounding agents and ego-centric free space may evolve, reducing overly conservative decisions without sacrificing safety. To efficiently exploit these geometric cues, GeoWorldAD progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.",
      "short_abstract": "Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale...",
      "published": "2026-07-19",
      "updated": "2026-07-19",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 31,
      "arxiv_url": "https://arxiv.org/abs/2607.17521",
      "pdf_url": "https://arxiv.org/pdf/2607.17521.pdf"
    },
    {
      "id": "2607.17351",
      "anchor": "paper-2607-17351",
      "title": "DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception",
      "authors": [
        "Eli Goldenshluger",
        "Barak Pinkovich",
        "Chaim Baskin"
      ],
      "abstract": "DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-modal 3D detection for autonomous mobility by learning a sparse acquisition pattern end-to-end with the fusion model. A learnable MIMO design module is trained end-to-end within a fusion network that operates directly on raw radar ADC data together with camera images and LiDAR point clouds. During training, the design module is supervised by the other sensors, enabling the system to learn both which receiver antennas to activate and the effective number of them. At deployment, the design module is removed and replaced by the learned sparse subsampling mask, leaving the downstream model architecture unchanged. Evaluated on the RADIal dataset, DeeperRadar discovers sparse, task-aware radar configurations that match or exceed full-array baselines while using fewer receivers, potentially reducing radar cost and integration complexity. These results show that learned optimal MIMO radar design depends on the fusion stack and the downstream perception task.",
      "short_abstract": "DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-modal 3D detection for autonomous mobility by learning a sparse acquisition pattern end-to-end with the fusion model. A learnable MIMO design module is...",
      "published": "2026-07-19",
      "updated": "2026-07-19",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 20,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: point cloud",
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 31,
      "arxiv_url": "https://arxiv.org/abs/2607.17351",
      "pdf_url": "https://arxiv.org/pdf/2607.17351.pdf"
    },
    {
      "id": "2607.22698",
      "anchor": "paper-2607-22698",
      "title": "FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog",
      "authors": [
        "Vansh Panwar"
      ],
      "abstract": "Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal alignments needed to evaluate robust sensor fusion. Real-world weather datasets suffer from uncontrolled collection and single-level, uncalibrated conditions, while synthetic alternatives either target camera-only restoration or lack the paired clean-and-foggy structure needed to benchmark \"defog-then-detect\" pipelines. We present FogDrive, a rigorously calibrated, multi-modal autonomous-driving dataset bridging data-centric engineering and robust machine learning. Built with the CARLA simulator, FogDrive contains 660 scenes (~133k fully annotated frames, 50:50 day/night) across four synchronized cameras (RGB, depth, semantic segmentation), a LiDAR and semantic-LiDAR pair, and front radar. Physically consistent fog is modeled independently on camera channels (Koschmieder model) and LiDAR channels (Beer-Lambert law) at three calibrated visibility densities (160m, 100m, 50m). Every scene ships in four matched variants (clean plus three graded fog levels) with cross-calibrated 2D and 3D bounding boxes. A semantic-segmentation-based quality audit over 8k images validates annotations at 95.1% precision and over 99% recall for vehicles within 40m. We establish baseline benchmarks with state-of-the-art architectures (TransFusion, BEVFusion, YOLOv8-m) across two paradigms: 3D multi-modal fusion and 2D image restoration. These yield critical data-centric insights: mixing multi-density fog during training tightens 3D bounding-box geometry without added data-scaling cost, while in 2D pipelines image-quality metrics (PSNR, SSIM) prove poor predictors of downstream detection performance. FogDrive will be fully open-sourced alongside our data-generation framework to accelerate robust, multi-modal research.",
      "short_abstract": "Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal alignments needed to evaluate robust sensor fusion. Real-world weather datasets suffer from uncontrolled...",
      "published": "2026-07-18",
      "updated": "2026-07-18",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark",
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 26,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: scene segmentation",
          "abstract: sensor fusion",
          "abstract: LiDAR",
          "abstract: radar",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 32,
      "arxiv_url": "https://arxiv.org/abs/2607.22698",
      "pdf_url": "https://arxiv.org/pdf/2607.22698.pdf"
    },
    {
      "id": "2607.16943",
      "anchor": "paper-2607-16943",
      "title": "SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections",
      "authors": [
        "Yunwei Li",
        "Shengjie Fu",
        "Chunrong Chen",
        "Chengxiang Zhao",
        "Yuchen Fan",
        "Mingyu Zhu",
        "Yanchao Xu",
        "Jiahui Xu",
        "Anran Wang",
        "Huanan Wang",
        "Yuxin Zhang",
        "Lan Yang",
        "Chuzhao Li",
        "Jie Ji",
        "Yi He",
        "Abhijit Sarkar",
        "Akash Sonth",
        "Hong Wang",
        "Jun Li"
      ],
      "abstract": "Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous driving systems (ADS), as these scenarios involve dense heterogeneous traffic, contested right of way, and long-tail safety-critical interactions, posing significant challenges to the Safety of the Intended Functionality (SOTIF). Existing naturalistic driving datasets often suffer from geographical homogeneity, sparsity of safety-critical events, and lack of semantic risk annotations, which limit the evaluation of algorithmic generalizability and targeted SOTIF verification. To address these gaps, this paper introduces SinD 2.0, a large-scale drone-based intersection dataset dedicated to cross-domain ADS safety analysis. The main contributions of SinD 2.0 are: (1) Cross-domain diversity: It covers six signalized intersections across four Chinese cities, capturing distinct intersection topologies and regional driving behavior characteristics; (2) High-density risk interactions: A total of 32,682 safety-critical events are extracted via surrogate safety measures, significantly enriching the density of boundary test scenarios; (3) Hierarchical semantic annotations: Besides integration with high-definition (HD) maps and Signal Phase and Timing (SPaT) data, it provides multi-dimensional semantic labels including traffic violations, high-risk interactions, visual shielding, and narrow feasible areas; (4) Full-stack testing toolchain: It supports automated scenario extraction, prediction-only evaluation, open-loop replay, reactive closed-loop testing, and photorealistic rendering. Benchmark experiments demonstrate that SinD 2.0 exhibits significant domain shifts across cities, and the semantic risk subsets can effectively expose the performance limitations of ADS algorithms. The dataset, annotations, and testing toolchain are available at https://github.com/SOTIF-AVLab/SinD/tree/main.",
      "short_abstract": "Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous driving systems (ADS), as these scenarios involve dense heterogeneous traffic, contested right of way, and long-tail safety-critical interactions,...",
      "published": "2026-07-18",
      "updated": "2026-07-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 57,
        "margin": 54,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: verification",
          "title: validation or testing",
          "abstract: validation or testing",
          "title: SOTIF",
          "abstract: SOTIF"
        ]
      },
      "recency": "fresh",
      "age_days": 32,
      "arxiv_url": "https://arxiv.org/abs/2607.16943",
      "pdf_url": "https://arxiv.org/pdf/2607.16943.pdf"
    },
    {
      "id": "2607.16938",
      "anchor": "paper-2607-16938",
      "title": "What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning",
      "authors": [
        "Kalpana Panda",
        "Wesley Maia",
        "Vinti Agarwal",
        "Ross Greer"
      ],
      "abstract": "End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems remains largely opaque, due to the complexity of traffic scenes. We propose a counterfactual ablation framework called Counterfactual Vision Action Analysis (CVAA) that systematically removes individual detected objects from front-camera images using photorealistic generative inpainting to prepare counterfactual sets to evaluate the difference in the model's response. This isolates the causal effect of each object's presence on the model's planning behaviour. Applied to the Alpamayo 1 trajectory predictor across 210 nuScenes driving scenes, we create a dataset Counter -nuScenes, using which we see that vehicles and pedestrians within the model's 'path' dominate causal influence as expected, while traffic lights, as expected, exert disproportionate effect relative to their image footprint. However, we also find cases where the model responds strongly to objects a human driver would consider irrelevant. This brings forth a deeper question: does the model itself view the scene as a sum of individual objects influencing the outcome, or does it encode an entirely different set of internal features that do not correspond to human-legible scene elements? To further understand this, we compare intermediate representations of original and inpainted image pairs using mechanistic interpretability techniques and examine the effect of the removal through the various model layers. Together, these two stages offer a path from behavioral auditing to representational understanding, creating explainable driving systems and solidifying human-AI trust.",
      "short_abstract": "End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these...",
      "published": "2026-07-18",
      "updated": "2026-07-18",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 21,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 32,
      "arxiv_url": "https://arxiv.org/abs/2607.16938",
      "pdf_url": "https://arxiv.org/pdf/2607.16938.pdf"
    },
    {
      "id": "2607.16758",
      "anchor": "paper-2607-16758",
      "title": "Hybrid Machine Learning for Articulation Angle Estimation of Truck-Semitrailer Combinations",
      "authors": [
        "Qixuan Zhang",
        "Jonas Boettcher",
        "Simon F. G. Ehlers",
        "Marvin Stuede"
      ],
      "abstract": "Accurate articulation angle estimation of trucks with trailers is critical for autonomous driving and advanced driver assistance system (ADAS). Existing methods either require manual initialization, additional sensors, or prior knowledge and signals from trailers, or they lack real-world validation, limiting practical deployment. This paper presents multiple learning-based models to directly estimate articulation angles from visual and kinematic inputs, eliminating the need for dedicated driving maneuvers for initialization, bounding box annotations, trailer-mounted sensor signals, or prior knowledge of trailer parameters. Two learning-based models are integrated with a kinematic model within an extended Kalman filter (EKF) framework, and an adaptive weighting scheme based on uncertainty quantification is applied for measurements involving visual input. Extensive real-world experiments with different trailer types demonstrate the approaches' robustness and generalization under out-of-domain conditions, including new trailers, varying colors, and lighting conditions. Results show that the hybrid method achieves accurate and reliable articulation angle estimation while maintaining reduced implementation requirements and practical deployment advantages.",
      "short_abstract": "Accurate articulation angle estimation of trucks with trailers is critical for autonomous driving and advanced driver assistance system (ADAS). Existing methods either require manual initialization, additional sensors, or prior knowledge and signals from...",
      "published": "2026-07-18",
      "updated": "2026-07-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 4,
        "evidence": [
          "abstract: validation or testing",
          "abstract: robustness",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 32,
      "arxiv_url": "https://arxiv.org/abs/2607.16758",
      "pdf_url": "https://arxiv.org/pdf/2607.16758.pdf"
    },
    {
      "id": "2607.16922",
      "anchor": "paper-2607-16922",
      "title": "Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing",
      "authors": [
        "Taorui Huang",
        "Namita Gaidhani",
        "Ritvik Bansal",
        "S M Jubaer",
        "Regina Lim",
        "Rhett Zhao",
        "Gavin Rafael Selin",
        "Sunnie Deng Gao",
        "Hasnain N Syed"
      ],
      "abstract": "In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive, Blind, Flock, Jaywalker, Elderly, Kid, Eventful, and Parked Pedestrian. These archetypes were introduced to move beyond single behavior labels and provide a more natural way to describe how dangerous pedestrians actually behave progressively in real-world traffic scenarios. However, upon further annotation of YouTube dash-cam videos, we identified 7 additional pedestrian archetypes with observable and significant behavioral differences from the previously proposed ones. These new archetypes capture pedestrian behavior patterns that could not be fully explained by the original taxonomy. In this pre-print, we introduce each new archetype, define its essential and optional behaviors, explain how it differs from previously proposed archetypes, and provide video-frame evidence showing the archetype in action.",
      "short_abstract": "In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash,...",
      "published": "2026-07-18",
      "updated": "2026-07-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 21,
        "evidence": [
          "title: safety",
          "title: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 32,
      "arxiv_url": "https://arxiv.org/abs/2607.16922",
      "pdf_url": "https://arxiv.org/pdf/2607.16922.pdf"
    },
    {
      "id": "2607.16847",
      "anchor": "paper-2607-16847",
      "title": "Driver Behavior Under Traffic Complexity: Variance Decomposition and Cross-Driver Generalization for Driver Monitoring",
      "authors": [
        "Lukas Köning",
        "Nataša Miličić",
        "Klaus Bogenberger"
      ],
      "abstract": "Understanding why a traffic situation is demanding for the human driver is central to safe and comfortable partially automated driving. Existing complexity metrics characterize only external environmental factors, while driver monitoring systems detect only endogenous states such as distraction. Neither side captures how external demands translate into driver-perceived situational load. Driver behavior serves as the natural bridge between both, and this paper investigates the inverse inference problem of estimating traffic complexity from behavioral signals. A systematic screening of 175 behavioral features across five domains (gaze, head pose, longitudinal control, guiding fixation, scanning strategy) is applied to data from 20 drivers in real urban traffic. Of these, 31 features exhibit statistically confirmed complexity effects, while 140 are confirmed as null effects through equivalence testing. Mixed-effects variance decomposition reveals that complexity explains only 1.5% of behavioral variance, whereas driver identity accounts for 23% and residual variance for 75%. This unfavorable ratio explains both the failure of all eight evaluated feature-level personalization strategies and the convergence of four classification architectures at F1 around 0.45 under leave-one-subject-out cross-validation. Guiding fixation rate emerges as the single most deployment-ready feature, combining speed-robustness, universality across drivers, and minimal inter-driver variation in complexity sensitivity. The results define three deployment regimes for complexity-adaptive advanced driver assistance systems and establish the variance structure as the primary bottleneck for complexity estimation.",
      "short_abstract": "Understanding why a traffic situation is demanding for the human driver is central to safe and comfortable partially automated driving. Existing complexity metrics characterize only external environmental factors, while driver monitoring systems detect only...",
      "published": "2026-07-18",
      "updated": "2026-07-18",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 16,
        "evidence": [
          "title: driver monitoring",
          "abstract: driver monitoring",
          "abstract: human driver"
        ]
      },
      "recency": "fresh",
      "age_days": 32,
      "arxiv_url": "https://arxiv.org/abs/2607.16847",
      "pdf_url": "https://arxiv.org/pdf/2607.16847.pdf"
    },
    {
      "id": "2607.22697",
      "anchor": "paper-2607-22697",
      "title": "Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning",
      "authors": [
        "Nadine Chang",
        "Maying Shen",
        "Shizhe Diao",
        "Jialiang Wang",
        "Jingde Chen",
        "Thomas Breuel",
        "Pavlo Molchanov",
        "Rafid Mahmood",
        "Jose M. Alvarez"
      ],
      "abstract": "Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods score training-side criteria rather than directly optimizing deployment match. We introduce TTCov (Test-Time Coverage), a data-level test-conditioned curation method that uses test-side information before training instead of updating model weights at inference. TTCov decomposes deployment-conditioned curation into coverage and distribution. To represent coverage, it builds a task Atlas, a collection of LLM-based atomic propositions (APs) describing deployment-relevant concepts, seeded from open task knowledge and expanded with unmatched APs extracted from unlabeled deployment samples. To represent distribution, it instantiates the matched deployment APs with their frequencies, yielding a Knowledge Atlas (K-Atlas) that operationalizes the deployment distribution as a curation target. TTCov then selects a budgeted training set whose deployment APs distribution approximates this target. We apply TTCov towards autonomous driving (AD), keeping adaptation off the inference path while selecting data with greater deployment-relevant coverage, closer K-Atlas matching, and stronger downstream end-to-end driving performance than data-curation baselines, including seamless adaptability to novel domains via city-to-city expansion.",
      "short_abstract": "Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods score training-side criteria rather than directly optimizing deployment match....",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 3,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.22697",
      "pdf_url": "https://arxiv.org/pdf/2607.22697.pdf"
    },
    {
      "id": "2607.16560",
      "anchor": "paper-2607-16560",
      "title": "From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence",
      "authors": [
        "Nadine Chang",
        "Maying Shen",
        "Shizhe Diao",
        "Jialiang Wang",
        "Jingde Chen",
        "Thomas Breuel",
        "Pavlo Molchanov",
        "Rafid Mahmood",
        "Jose M. Alvarez"
      ],
      "abstract": "We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene. A global semantic codebook unifies these into a shared vocabulary of canonical atomic propositions, placing every modality and observation into one interpretable space that spans fine grained facts to high level concepts and composes into richer ones. This brings interpretability with reasoning, cross-modal understanding and retrieval, and compositionality that enables complex multimodal understanding, rich data curation and complex structured retrieval. We demonstrate the framework on autonomous driving and open-world data.",
      "short_abstract": "We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene. A global semantic codebook...",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.16560",
      "pdf_url": "https://arxiv.org/pdf/2607.16560.pdf"
    },
    {
      "id": "2607.16181",
      "anchor": "paper-2607-16181",
      "title": "Vision-Language Assistant for Emotional Reactions to Risky Driving",
      "authors": [
        "Harine Choi",
        "Eun Hak Lee",
        "Zhengzhong Tu"
      ],
      "abstract": "This study introduces a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Although vision-language models have advanced perception and reasoning in autonomous driving, existing systems rarely consider the emotional dimension or real-world user experience. Keep Yelling Assistant (KYA) detects high-risk driving maneuvers in real time, such as sudden cut-ins. It then produces emotional responses through a large language model tailored to driver preferences. The framework comprises two core modules. The vision module uses YOLOv8 variants to detect nearby vehicles and identify risky behaviors such as sudden cut-ins. Key driving metrics, including relative distance, speed, and projected reach time, are extracted and normalized to produce a structured behavior log. The language module processes this log with user-defined emotional tone settings, such as neutral, humorous, and analytical, and generates verbal reactions using state-of-the-art large language models, including ChatGPT-4o, Claude 3, Gemini 2.5, and Copilot. We evaluated the proposed system using dashcam videos containing risky driving behaviors and a user study involving 108 participants. Participants selected preferred response styles, and the large language models were evaluated based on emotional alignment. All models received favorable ratings, although preferences varied across personas. Notably, the combination of YOLOv8s and ChatGPT-4o achieved the highest score of 4.29 out of 5.00. By integrating real-world perception with emotionally adaptive dialogue, KYA introduces a new paradigm for emotionally intelligent in-vehicle artificial intelligence. It offers promising directions for improving safety, trust, and emotional well-being in both conventional and autonomous vehicles.",
      "short_abstract": "This study introduces a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Although vision-language models have advanced perception and reasoning in autonomous...",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.16181",
      "pdf_url": "https://arxiv.org/pdf/2607.16181.pdf"
    },
    {
      "id": "2607.15888",
      "anchor": "paper-2607-15888",
      "title": "Red Light, Grey Zone: A Multi-Perspective Interactive Narrative for Autonomous Driving Ethics",
      "authors": [
        "Mengyi Wei",
        "Nianhua Liu",
        "Chenyu Zuo",
        "Liqiu Meng"
      ],
      "abstract": "Autonomous driving ethics is not only an expert concern, but also a public issue involving risk, responsibility, and governance. However, non-experts often struggle to interpret these issues in concrete incidents, especially when responsibility is distributed across multiple stakeholders. This paper investigates interactive narrative as a public-facing method for eliciting situated ethical reflection on autonomous driving. We present Red Light, Grey Zone, a web-based, multi-perspective interactive narrative prototype inspired by a real-world autonomous-driving incident. The prototype invites participants to compare stakeholder perspectives, examine scene materials, and make responsibility judgments in the face of ethical ambiguity. We report an exploratory user study (N=12) examining how differently non-experts responded to the prototype. Our analysis focuses on three dimensions of reflection: ethical cognition, responsibility-focused critical thinking, and multi-perspective reasoning. Exploratory pre-post results showed the strongest self-reported shift in responsibility-focused critical thinking among participants who completed the intended stakeholder-comparison process, while ethical cognition and multi-perspective reasoning showed positive directional trends. Qualitative findings further show how participants reflected on safety and market trade-offs, responsibility ambiguity, transparency and privacy, and governance gaps. Participants also used stakeholder comparison to corroborate evidence and, in many cases, broaden responsibility judgments from single-actor blame toward more distributed interpretations of accountability. Overall, the study suggests that multi-perspective interactive narratives may support non-expert reflection on accountability, evidence, and governance in AI-enabled systems.",
      "short_abstract": "Autonomous driving ethics is not only an expert concern, but also a public issue involving risk, responsibility, and governance. However, non-experts often struggle to interpret these issues in concrete incidents, especially when responsibility is distributed...",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "title: ethics",
          "abstract: ethics"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.15888",
      "pdf_url": "https://arxiv.org/pdf/2607.15888.pdf"
    },
    {
      "id": "2607.15820",
      "anchor": "paper-2607-15820",
      "title": "In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing",
      "authors": [
        "Qunying Song",
        "Yuan Gao",
        "Johannes Betz",
        "Dietmar Pfahl",
        "Mohammad Reza Mousavi",
        "Federica Sarro"
      ],
      "abstract": "Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks well-established standards for scenario selection, performance evaluation, and acceptance criteria. To better understand current ADS testing practices and challenges, we conducted an interview study with experts working on ADS development and testing in nine companies from six different countries. Through thematic analysis, we synthesized industrial testing practices, challenges, potential solutions, future trends, and proposed an evidence-centered closed-loop testing framework for ADS testing. Our findings show that current practices primarily focus on scenario-based and X-in-the-loop testing approaches, supported by diverse tools, metrics, benchmarks, and testing strategies. The participants highlighted major challenges related to scenario realism, scenario coverage, simulation fidelity, and acceptance criteria, while also discussing potential solutions such as the use of AI, world models, and end-to-end approaches. Furthermore, participants envisioned future ADS testing to become more automated, data-driven, and transparent across the industry. Overall, this study provides a comprehensive industry-grounded overview of ADS testing, proposes an evidence-centered closed-loop testing framework to provide actionable guidance for ADS testing, and outlines important directions for future research and practice.",
      "short_abstract": "Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks...",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.15820",
      "pdf_url": "https://arxiv.org/pdf/2607.15820.pdf"
    },
    {
      "id": "2607.15713",
      "anchor": "paper-2607-15713",
      "title": "Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization",
      "authors": [
        "Yong Chu",
        "Xun Zhou",
        "Zenglin Xu",
        "Hui Wang",
        "Yue Yu"
      ],
      "abstract": "Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel \"map-as-prompt\" framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.",
      "short_abstract": "Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments...",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 18,
        "evidence": [
          "title: localization",
          "abstract: localization"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.15713",
      "pdf_url": "https://arxiv.org/pdf/2607.15713.pdf"
    },
    {
      "id": "2607.15697",
      "anchor": "paper-2607-15697",
      "title": "SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models",
      "authors": [
        "Jinwen Xin",
        "Xixiang Lv"
      ],
      "abstract": "Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for autonomous driving, the presence of backdoor attacks introduces substantial security risks. This study focuses on implementing backdoor defense measures for speech recognition models in run-time, taking into account the characteristics of audio signals. We propose SpeechGuard, the first online backdoor defense pipeline designed to identify and purify poisoned audio samples. Specifically, we improve STRIP method to perform adaptive perturbation injection to detect and filter poisoned samples, named as S-STRIP. More importantly, we further consider the purification of poisoned samples. We utilize time-frequency (T-F) masking to suppress the expression of trigger signals and autonomously generate masks based on an autoencoder. The two-stage processing prevents the backdoor in the model from being triggered, and even input speech carrying triggers can be accurately predicted. Extensive experimental demonstrate that SpeechGuard can accurately filter out poisoned samples. Through purification, it can significantly mitigate the backdoor threat while maintaining a certain prediction accuracy.",
      "short_abstract": "Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for...",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 19,
        "evidence": [
          "abstract: security",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.15697",
      "pdf_url": "https://arxiv.org/pdf/2607.15697.pdf"
    },
    {
      "id": "2607.16175",
      "anchor": "paper-2607-16175",
      "title": "Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities",
      "authors": [
        "Md Erfan",
        "Ahmed Ryan",
        "Md Kamal Hossain Chowdhury",
        "Md Rayhanur Rahman"
      ],
      "abstract": "Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Units, in-vehicle infotainment systems, and telematics units, where vulnerabilities can compromise assets, users, and vehicle operations. These vulnerabilities are commonly documented as plain text in the Common Vulnerabilities and Exposures (CVE) database; however, security practitioners require structured information about affected assets, types of weaknesses, and attack behaviors to effectively mitigate the risks from these vulnerabilities. To this end, we evaluate open-weight Large Language Models (LLMs) for generating Structured Threat Information Expression (STIX), a well-known structured format for representing threat information, for CAV-related CVEs. We construct a dataset called CAV-STIXGen that maps CAV vulnerability descriptions to STIX domain objects (SDO), STIX relationship objects (SRO), Common Weakness Enumeration (CWE), and MITRE ATT&CK techniques mappings. Using this dataset, we evaluated 11 open-weight LLMs (4B to 120B parameters) across various prompting strategies and temperatures. Single-model configurations achieve F1 scores of 0.94 for SDO, 0.63 for SRO, and 0.99 for CWE mapping, while complete MITRE ATT&CK mapping remains challenging. In a multi-agent setup, Gemma-4-31B and Codestral-22B achieve F1 scores of 0.91 for SDOs and 0.43 for SROs, respectively. Lastly, we analyze CWE and MITRE ATT&CK co-occurrences to identify recurring threat patterns in the CAV domain, demonstrating how AI-assisted vulnerability-to-STIX translation can automate threat intelligence and prioritize defense in transportation security.",
      "short_abstract": "Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Units, in-vehicle infotainment systems, and telematics units, where vulnerabilities can compromise assets, users, and...",
      "published": "2026-07-17",
      "updated": "2026-07-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 14,
        "evidence": [
          "abstract: security",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 33,
      "arxiv_url": "https://arxiv.org/abs/2607.16175",
      "pdf_url": "https://arxiv.org/pdf/2607.16175.pdf"
    },
    {
      "id": "2607.15048",
      "anchor": "paper-2607-15048",
      "title": "RoGS: Adaptive Meshgrid Gaussian for Large-Scale Road Surface Mapping",
      "authors": [
        "Tianchen Deng",
        "Zhiheng Feng",
        "Wenhua Wu",
        "Ziming Li",
        "Siting Zhu",
        "Hesheng Wang"
      ],
      "abstract": "Road surface mapping plays a crucial role in autonomous driving, supporting high-definition map generation, lane-level perception, and automatic road annotation. Recent mesh-based road surface reconstruction methods have shown promising results, but they still suffer from limited reconstruction quality and high optimization cost, especially in large-scale driving scenarios. To address these limitations, we propose ROADGS-T, a robust and efficient large-scale road surface mapping framework based on adaptive meshgrid Gaussian representation. Specifically, we model the road surface by placing 2D Gaussian surfels on a meshgrid, where each surfel explicitly stores color, semantic, and geometric information. Compared with conventional mesh-based representations and 3D Gaussian primitives, the proposed meshgrid Gaussian representation better matches the thin-surface property of roads while significantly reducing redundant primitives and overlap during optimization. To further improve representation efficiency and structural fidelity, we introduce a road-structure-aware adaptive meshgrid strategy, which allocates denser Gaussian surfels to geometrically or semantically complex regions, such as lane markings, road boundaries, and height discontinuities, while maintaining a compact representation in flat road areas. Moreover, instead of relying on a single nearest vehicle pose, we design a trajectory-consistency-guided pose-robust refinement strategy, which estimates local surface priors from multiple neighboring poses and adaptively weights pose-guided height regularization according to their geometric consistency.",
      "short_abstract": "Road surface mapping plays a crucial role in autonomous driving, supporting high-definition map generation, lane-level perception, and automatic road annotation. Recent mesh-based road surface reconstruction methods have shown promising results, but they...",
      "published": "2026-07-16",
      "updated": "2026-07-16",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 18,
        "evidence": [
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 34,
      "arxiv_url": "https://arxiv.org/abs/2607.15048",
      "pdf_url": "https://arxiv.org/pdf/2607.15048.pdf"
    },
    {
      "id": "2607.14727",
      "anchor": "paper-2607-14727",
      "title": "WorkDrive: Roadwork Chain of Causation for Autonomous Driving",
      "authors": [
        "Tianyi Jiang",
        "Wen Zhang",
        "Sihan Yang",
        "Ming Lu",
        "Wentao Zhang"
      ],
      "abstract": "Autonomous driving vision-language models (VLMs) struggle in roadwork zones, where familiar visual cues such as lane markings and permanent signs are altered or absent, and temporary devices such as cones and barriers redefine the drivable corridor. VLMs can detect these objects, but without explicit guidance they anchor their reasoning on familiar elements from pre-training and fail to connect work-zone observations to correct planning decisions. We propose WorkDrive, a framework that constructs perception-grounded causal reasoning for work zones and aligns it with trajectory prediction. An automated multitask perception pipeline extracts structured scene facts and injects them into a Chain-of-Causation (CoC) annotation pipeline, redirecting the annotator's attention to domain-specific elements. The resulting reasoning labels are used for supervised fine-tuning, followed by reinforcement learning with a single reward: consistency between lateral meta-actions and the predicted trajectory. On ROADWork, the largest public work-zone dataset, the proposed roadwork CoC reduces trajectory average displacement error (ADE) by 9.0\\%, and consistency-based GRPO yields a further 3.0\\%, achieving progressive improvement over the trajectory-only baseline. Code and data will be publicly released.",
      "short_abstract": "Autonomous driving vision-language models (VLMs) struggle in roadwork zones, where familiar visual cues such as lane markings and permanent signs are altered or absent, and temporary devices such as cones and barriers redefine the drivable corridor. VLMs can...",
      "published": "2026-07-16",
      "updated": "2026-07-16",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 4,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 34,
      "arxiv_url": "https://arxiv.org/abs/2607.14727",
      "pdf_url": "https://arxiv.org/pdf/2607.14727.pdf"
    },
    {
      "id": "2607.14710",
      "anchor": "paper-2607-14710",
      "title": "Variational Inference for Bird's Eye View Segmentation in Autonomous Driving",
      "authors": [
        "Jingyue Shi",
        "Huaicheng Li",
        "Junhui Zhao",
        "Yanxiang Jiang"
      ],
      "abstract": "The bird's eye view (BEV) has emerged as a pivotal approach for environmental perception in autonomous driving, providing a unified spatial representation for vehicles. Nevertheless, despite BEV's significance in addressing the challenges inherent to autonomous driving, effectively fusing data from multiple camera sensors and operating in complex external driving environments remains a considerable challenge. To mitigate this issue, we recast the BEV segmentation problem within a variational inference framework. In this paper, we propose a novel transformer-based variational flow transformation network for BEV segmentation, denoted as TVB. Our architecture implicitly learns the mapping from multiple camera views to a unified canonical BEV map during training by exploiting posterior BEV supervision. TVB employs a conditional variational auto encoder (CVAE) as its backbone and produces multiple BEV map candidates. To augment the realism of the generated BEV maps, we integrate normalizing flows into the map generation process, enabling the construction of more complex and expressive probability distributions. Furthermore, we design a BEV-attention fusion (BAF) module that harnesses attention mechanisms to adaptively integrate the multiple candidate BEV maps. Experimental results, evaluated on both the nuScenes and OPV2Vdatasets, demonstrate that our proposed method achieves superior performance in multi-camera view BEV segmentation and lane environment perception.",
      "short_abstract": "The bird's eye view (BEV) has emerged as a pivotal approach for environmental perception in autonomous driving, providing a unified spatial representation for vehicles. Nevertheless, despite BEV's significance in addressing the challenges inherent to...",
      "published": "2026-07-16",
      "updated": "2026-07-16",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "abstract: camera",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 34,
      "arxiv_url": "https://arxiv.org/abs/2607.14710",
      "pdf_url": "https://arxiv.org/pdf/2607.14710.pdf"
    },
    {
      "id": "2607.15491",
      "anchor": "paper-2607-15491",
      "title": "Trajectory-aware Cross-view Geo-localization with Sequential Observations",
      "authors": [
        "Tianyi Gao",
        "Jiayu Lin",
        "Danielle Beaulieu",
        "Nathan Jacobs"
      ],
      "abstract": "Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet they overlook a complementary sequential modality: route descriptions -- which capture the same trajectory at a higher level of abstraction and are often the only input available (e.g., a user directing an autonomous vehicle to a pickup point). To bridge this gap, we introduce SeqGeo-VL, a dataset of $\\sim$39K video-text-satellite triplets, and TrajLoc, a unified framework capable of processing both video clips and route descriptions. By leveraging both dense visual and abstract linguistic semantics, TrajLoc enables these modalities to mutually reinforce cross-view matching. We further propose TrajMod, a lightweight module that conditions query embeddings on trajectory geometry, yielding spatially-aware representations. Experiments show that TrajLoc achieves substantial gains over state-of-the-art methods on both video and text geo-localization. The project page is available at https://humblegamer.github.io/trajloc/.",
      "short_abstract": "Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet they overlook a complementary...",
      "published": "2026-07-16",
      "updated": "2026-07-16",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 20,
        "evidence": [
          "title: localization",
          "abstract: localization"
        ]
      },
      "recency": "fresh",
      "age_days": 34,
      "arxiv_url": "https://arxiv.org/abs/2607.15491",
      "pdf_url": "https://arxiv.org/pdf/2607.15491.pdf"
    },
    {
      "id": "2607.15117",
      "anchor": "paper-2607-15117",
      "title": "A Model Predictive Control Framework for Assisted Vehicle Drifting",
      "authors": [
        "Marco Cortese",
        "Antonio Gallina",
        "Matteo Grandin",
        "Giovanni Righetti",
        "Mattia Bruschetta",
        "Basilio Lenzo"
      ],
      "abstract": "Model Predictive Control (MPC) has been widely applied to autonomous vehicle drifting. Assisted drifting, that is where the driver remains in the loop, is still comparatively underexplored. Existing approaches often rely on restrictive assumptions, such as precomputed drift equilibria, full actuation authority, or prior path knowledge, which limit applicability to expert drivers. This paper proposes a nonlinear model predictive control (NMPC) framework for assisted drifting on a rear-wheel-drive vehicle. Through steer-by-wire and drive-by-wire interfaces, the controller decouples driver commands from direct actuator inputs, allowing the driver to regulate the desired sideslip through the steering wheel while the NMPC maintains vehicle stability. A dedicated activation logic ensures that the controller engages only under deliberate driver intent. High-fidelity simulations show that the proposed architecture can stabilize drifting maneuvers using a simple single-track prediction model with basic tire dynamics, even when the sideslip reference is continuously varied by the driver.",
      "short_abstract": "Model Predictive Control (MPC) has been widely applied to autonomous vehicle drifting. Assisted drifting, that is where the driver remains in the loop, is still comparatively underexplored. Existing approaches often rely on restrictive assumptions, such as...",
      "published": "2026-07-16",
      "updated": "2026-07-16",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 31,
        "evidence": [
          "title: model predictive control",
          "abstract: model predictive control",
          "abstract: controller",
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "fresh",
      "age_days": 34,
      "arxiv_url": "https://arxiv.org/abs/2607.15117",
      "pdf_url": "https://arxiv.org/pdf/2607.15117.pdf"
    },
    {
      "id": "2607.14688",
      "anchor": "paper-2607-14688",
      "title": "MIND-CAVs: Multi-Intelligence Negotiation and Decision System for CAVs based on Intent-Driven Autonomy",
      "authors": [
        "Mainak Mondal",
        "Yihang Feng",
        "Yangchao Luo",
        "Song Han"
      ],
      "abstract": "Modern autonomous vehicles largely operate as isolated agents: they rely on on-board perception and decision modules and broadcast Basic Safety Messages (BSMs) that expose only low-level kinematic state. While existing cooperative driving frameworks enable limited sensor sharing, they rarely communicate high-level maneuver intentions, and edge computing is primarily used for content delivery rather than decision arbitration. As a result, current connected autonomy lacks a principled mechanism for making globally consistent, intent-aware coordination decisions across vehicles. To address this gap, we propose MIND-CAVs, a Multi-Intelligence Negotiation and Decision framework for connected autonomous vehicles (CAVs) based on intent-driven autonomy. Each vehicle abstracts raw sensor observations into structured intent representations, exchanges them over V2X links, and receives globally consistent coordination plans from roadside edge servers. Edge agents combine learned and rule-based arbitration mechanisms to negotiate conflicting intents among vehicles, while a cloud platform records decisions for auditing and continual retraining. We implement MIND-CAVs in a CARLA-based AI-in-the-loop platform and evaluate it in multi-lane highway scenarios involving conflicting maneuvers and route-constrained exits. Experimental results show improved maneuver completion time and reduced unsafe proximity and unnecessary braking compared with isolated autonomy, first-come-first-served arbitration, and multi-agent reinforcement learning baselines.",
      "short_abstract": "Modern autonomous vehicles largely operate as isolated agents: they rely on on-board perception and decision modules and broadcast Basic Safety Messages (BSMs) that expose only low-level kinematic state. While existing cooperative driving frameworks enable...",
      "published": "2026-07-16",
      "updated": "2026-07-16",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 16,
        "evidence": [
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: edge or cloud",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "fresh",
      "age_days": 34,
      "arxiv_url": "https://arxiv.org/abs/2607.14688",
      "pdf_url": "https://arxiv.org/pdf/2607.14688.pdf"
    },
    {
      "id": "2607.14387",
      "anchor": "paper-2607-14387",
      "title": "Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving",
      "authors": [
        "Yuan Gao",
        "Wenting Miao",
        "Mattia Piccinini",
        "Haoyu Wang",
        "Qunying Song",
        "Johannes Betz"
      ],
      "abstract": "Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are defined as executable scripts. Yet automatically generating such scripts from regulatory descriptions remains an open challenge, and existing approaches face fundamental trade-offs. Retrieval-assemble methods achieve reasonable compilation rates but lack scalability, whereas retrieval-based full-script generation suffers from low compilation success rates. We present Chat2Scenic, the first iterative retrieval-augmented framework to generate scenario scripts in Domain Specific Language (DSL). Specifically, Chat2Scenic provides a chatbot interface that supports interactive scenario refinement and integrates Retrieval-augmented Generation (RAG) to ground scenario generation in regulatory knowledge and DSL syntax. Furthermore, we propose an open benchmark for scenario generation comprising 123 scenarios from various regulations, including NHTSA and United Nations Vehicle Regulations, as well as other sources. Extensive evaluation with State-of-the-Art (SOTA) Large Language Models (LLMs) demonstrates that Chat2Scenic achieves 76.42% Compilation Success Rate (CSR) and 58.17% Framework Accuracy (FA), outperforming existing methods (Retrieval Assemble with 30.08% CSR, 11.03% FA and Retrieval full script generation with 16.26% CSR, 10.86% FA). To facilitate future research, we release our code as open source at https://github.com/TUM-AVS/chat2scenic.",
      "short_abstract": "Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are defined as executable scripts. Yet automatically generating such scripts from regulatory descriptions remains an open...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.14387",
      "pdf_url": "https://arxiv.org/pdf/2607.14387.pdf"
    },
    {
      "id": "2607.14248",
      "anchor": "paper-2607-14248",
      "title": "3D Lane Detection with Odometry for High-Speed Vehicle Racing",
      "authors": [
        "Omoruyi Atekha",
        "John Subosits",
        "Marcus Greiff"
      ],
      "abstract": "Lane boundary detection is a critical component in autonomous driving systems and has been rigorously studied in regular driving scenarios. However, it is less explored in vehicle racing, where the car moves at higher speeds across more extreme road geometries. To study this problem, we introduce a new dataset for 3D lane detection in racing, featuring >$250$k images from multiple camera feeds and inertial measurements taken with a Lexus LC 500 driving on a closed circuit. With this dataset, we compare various approaches to 3D lane detection and propose modifications that permit frames to be processed at rates of almost 300Hz while retaining high predictive performance in the racing application. This facilitates a multi-camera ensemble approach that is validated on hardware. We show that sensing modalities such as inertial measurements can be leveraged for pre-integration to regress road geometries over both cameras and time, yielding improvements in key metrics. Compared to methods such as BevLaneDet, adding odometry and ensemble predictions improves the F1 score by 3 points and reduces near-vehicle mean absolute errors (MAEs) by $>30 \\%$. We show F1 scores $>$0.9 and lateral MAEs of $<$0.18m in vehicle deployments.",
      "short_abstract": "Lane boundary detection is a critical component in autonomous driving systems and has been rigorously studied in regular driving scenarios. However, it is less explored in vehicle racing, where the car moves at higher speeds across more extreme road...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 11,
        "evidence": [
          "abstract: camera",
          "title: lane detection",
          "abstract: lane detection"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.14248",
      "pdf_url": "https://arxiv.org/pdf/2607.14248.pdf"
    },
    {
      "id": "2607.14203",
      "anchor": "paper-2607-14203",
      "title": "Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation",
      "authors": [
        "NVIDIA",
        ":",
        "Jiahui Huang",
        "Jiawei Ren",
        "Michal Tyszkiewicz",
        "Bjoern Haefner",
        "Michael Shelley",
        "Xin Kang",
        "Seung Wook Kim",
        "Ning Xu",
        "Qi Wu",
        "Janick Martinez Esturo",
        "Shengyu Huang",
        "Nick Schneider",
        "Laura Leal-Taixe",
        "Zan Gojcic",
        "Sanja Fidler"
      ],
      "abstract": "3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development costs and improving safety. In recent years, neural simulation has become predominant, with methods such as NuRec playing a central role; however, these methods remain relatively slow and typically require per-scene tuning. In this work, we present Instant NuRec, a feed-forward neural reconstruction model that turns a short multi-view driving log into a fully simulatable 3D Gaussian Splatting (3DGS) world in a single forward pass. The model accepts multi-view input from a calibrated camera rig and emits a layered output consisting of static and dynamic 3DGS layers, a sky cubemap, and per-camera ISP corrections, while providing native support for non-pinhole camera models via 3DGUT. It reconstructs a 10-20-second multi-camera scene in roughly 1.5 seconds and achieves a PSNR on the Waymo Open Dataset that is 2.01 dB above the strongest evaluated baseline. Instant NuRec is deeply integrated into NuRec and is compatible with AlpaSim for closed-loop simulation.",
      "short_abstract": "3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development costs and improving safety. In recent years, neural simulation has become predominant, with methods such as NuRec...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.14203",
      "pdf_url": "https://arxiv.org/pdf/2607.14203.pdf"
    },
    {
      "id": "2607.14005",
      "anchor": "paper-2607-14005",
      "title": "M$^\\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming",
      "authors": [
        "Ke Cheng",
        "Hanqiao Ye",
        "Lei Shi",
        "Yahui Liu",
        "Yunhan Shen",
        "Jingtao Dong",
        "Zhenke Wang",
        "Wenxuan Ao",
        "Weixiang Xu",
        "Kaining Huang",
        "Shuhan Shen"
      ],
      "abstract": "Driving-world generation has emerged as a core capability for scalable autonomous-driving simulation, yet existing methods remain limited in object-level controllability and long-horizon stability. We present M$^\\text{4}$World, a Multi-view and Multimodal generative driving world model that synthesizes future surround-view video streams and synchronized LiDAR scans while supporting interactive object Manipulation and stable Minute-long streaming. Fine-grained object manipulation is realized through a flexible conditioning interface that supports explicit control over both the spatial layout and visual appearance of individual objects. Stable minute-long streaming, on the other hand, is achieved through a multi-stage training framework that enables online causal generation in only four denoising steps while maintaining coherent world dynamics throughout extended rollouts. Building on these components, we introduce an efficient few-clip post-training as well as a suite of visual reference-conditioned generation models, preserving general generation ability while allowing rare-case customization for long-tail controllability. To assess controllability beyond realism, we further introduce an automated VLM-based judging pipeline that evaluates scene-level condition adherence, view-wise object controllability, and cross-view object consistency. Comprehensive experiments show that M$^\\text{4}$World consistently delivers high generation quality, precise controllability, and stable minute-long streaming. Together with downstream long-tail augmentation and scene editing, these results demonstrate the potential of M$^\\text{4}$World for controllable, scalable driving simulation.",
      "short_abstract": "Driving-world generation has emerged as a core capability for scalable autonomous-driving simulation, yet existing methods remain limited in object-level controllability and long-horizon stability. We present M$^\\text{4}$World, a Multi-view and Multimodal...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.14005",
      "pdf_url": "https://arxiv.org/pdf/2607.14005.pdf"
    },
    {
      "id": "2607.13927",
      "anchor": "paper-2607-13927",
      "title": "Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data",
      "authors": [
        "Thang-Anh-Quan Nguyen",
        "Moussab Bennehar",
        "Luis Guillermo Roldao Jimenez",
        "Nathan Piasco",
        "Dzmitry Tsishkou",
        "Laurent Caraffa",
        "Jean-Philippe Tarel",
        "Roland Brémond"
      ],
      "abstract": "Reliable perception under diverse weather conditions remains a major challenge for autonomous driving systems. A common strategy to improve robustness is either to synthesize adverse weather conditions for training perception models or to apply weather-removal techniques to recover clean inputs. However, existing approaches typically rely on synthetic data augmentation or physics-based, task-specific models that require paired training data and often struggle to generate realistic weather effects or generalize robustly to out-of-domain scenarios. Toward this problem, we present Cyclone, a unified framework for weather editing based on latent diffusion, equipped with cycle-consistent constraints and knowledge from image-text models. Cyclone enables the generation of multiple weather conditions across diverse scenes while eliminating the need for paired data. Experimental results show that our approach produces more realistic, structure-preserving outputs than existing baselines and leads to consistent improvements across several downstream driving perception tasks. Furthermore, we demonstrate that Cyclone can be distilled to a video diffusion model for temporally consistent weather editing.",
      "short_abstract": "Reliable perception under diverse weather conditions remains a major challenge for autonomous driving systems. A common strategy to improve robustness is either to synthesize adverse weather conditions for training perception models or to apply...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13927",
      "pdf_url": "https://arxiv.org/pdf/2607.13927.pdf"
    },
    {
      "id": "2607.13926",
      "anchor": "paper-2607-13926",
      "title": "S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving",
      "authors": [
        "Jianguo Yu",
        "Rukang Wang",
        "Duanfeng Chu",
        "Chen Wang",
        "Renju Feng",
        "Liping Lu"
      ],
      "abstract": "Vision-Language Models (VLMs) have demonstrated remarkable potential for high-level reasoning in autonomous driving, yet they fundamentally struggle to generate precise, low-level control actions. This limitation is rooted in a semantic-physical gap caused by the inherent mismatch between discrete language tokens and continuous trajectory planning. While Vision-Language-Action (VLA) architectures attempt to bridge this gap by unifying perception and control into a single policy, this entanglement creates a new bottleneck. Standard VLAs experience a severe spatial representation collapse, which irreversibly degrades the fine-grained spatial and geometric priors essential for safe, boundary-aware navigation. To address this limitation, we propose the S-squared-VLA, which explicitly decouples the semantic and spatial streams in Vision-Language-Action models. The semantic stream leverages hierarchical bridging to extract multi-scale VLM features for robust intent reasoning. In parallel, an independent spatial stream bypasses the autoregressive language bottleneck, directly preserving uncompressed spatial features from the visual encoder. By integrating auxiliary perception supervision, this stream explicitly equips the model with rich spatial and geometric priors. Finally, a dual-stream planning adapter fuses high-level semantic intent with precise spatial constraints via cascaded attention mechanisms. Evaluations on the NAVSIM closed-loop benchmark show that S-squared-VLA achieves a Predictive Driver Model Score (PDMS) of 87.1, establishing a new state-of-the-art for VLA models under a purely supervised fine-tuning (SFT) setting. By mitigating the spatial representation collapse of traditional VLMs, our framework significantly outperforms baselines, achieving the highest No Collision (NC) rate of 98.4 among all evaluated methods.",
      "short_abstract": "Vision-Language Models (VLMs) have demonstrated remarkable potential for high-level reasoning in autonomous driving, yet they fundamentally struggle to generate precise, low-level control actions. This limitation is rooted in a semantic-physical gap caused by...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 14,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13926",
      "pdf_url": "https://arxiv.org/pdf/2607.13926.pdf"
    },
    {
      "id": "2607.13704",
      "anchor": "paper-2607-13704",
      "title": "nuTruck: Benchmarking Autonomous Driving Planning for Distributed Electric-drive Trucks",
      "authors": [
        "Jinyu Miao",
        "Pu Zhang",
        "Yifei He",
        "Chengyao Zhang",
        "Kun Jiang",
        "Ke Wang",
        "Mengmeng Yang",
        "Diange Yang"
      ],
      "abstract": "The dominance of traditional rule-based methods in autonomous driving has gradually been replaced by learning-based approaches. While learning-based planners have achieved considerable success in passenger vehicles, their performance on heavy-duty trucks, particularly modern distributed electric-drive trucks (DETs), remains largely unexplored. To facilitate research and application of learning-based planners in DETs, this letter presents the first high-fidelity benchmark, called nuTruck, designed to support large-scale neural network training and closed-loop evaluation. Given the complex dynamics and high rollover susceptibility of DETs, we first incorporate a highly accurate nonlinear truck dynamical model into the simulation, which enables independent driving and steering of all wheels and captures dynamic load transfer caused by acceleration, deceleration, and cornering, thereby allowing quantitative assessment of rollover risk in closed-loop simulation. Second, we adapt several rule-based and learning-based planners as baselines for DETs and evaluate their performance in closed-loop simulation. Finally, using real-world driving scenarios from the nuPlan dataset, we conduct extensive closed-loop evaluations, analyzing not only conventional collision-free planning performance, but also the dynamical safety of the planned trajectories. The proposed nuTruck benchmark is expected to serve as a new standard for fair and realistic evaluation of autonomous driving planners on DETs.",
      "short_abstract": "The dominance of traditional rule-based methods in autonomous driving has gradually been replaced by learning-based approaches. While learning-based planners have achieved considerable success in passenger vehicles, their performance on heavy-duty trucks,...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13704",
      "pdf_url": "https://arxiv.org/pdf/2607.13704.pdf"
    },
    {
      "id": "2607.13561",
      "anchor": "paper-2607-13561",
      "title": "Design and Implementation of a Microservice-Architecture Master Control System for AIMS",
      "authors": [
        "Li-Yue Tong",
        "Jia-Ben Lin",
        "Jun-Feng Hou",
        "Yuan-Yong Deng",
        "Dong-Guang Wang",
        "Guang-Qian Liu",
        "Song-Bo Xu",
        "Shang-Jie Ren",
        "Lian-Wei Zhao",
        "Zhi-Wei Feng",
        "Wei Duan",
        "Ming-Fu Shao",
        "Hui Wang",
        "Chen Yang"
      ],
      "abstract": "The mid-infrared solar magnetic field telescope AIMS (An Infrared System for the Accurate Measurement of Solar Magnetic Field) is the first ground-based telescope designed to directly measure solar magnetic fields via Zeeman splitting in the 8-14 um band, overcoming the century-long bottleneck of model-dependent indirect measurements. Its remote high-altitude site, heterogeneous multi-institute components, and complex observation modes comprising Fourier Transform Infrared (FTIR) spectropolarimetry and broadband imaging demand a highly autonomous Master Control System (MCS). We present the design and implementation of the AIMS MCS, featuring three key contributions: (1) an L0-L5 telescope automation classification inspired by the SAE J3016 autonomous driving standard, providing well-defined boundaries and a progressive evolution roadmap; (2) a three-layer system framework device control, autonomy support, and central decision-implemented with a microservice software architecture that achieves loose coupling, high cohesion, and continuous integration of heterogeneous components; and (3) a suite of key enabling tech-nologies including automatic pointing/tracking, autofocus via lucky-frame selection combined with power spectral ratio analysis, and environment-adaptive observation integrating auto-exposure, cloud detection, and power/thermal monitoring. The MCS has been validated across three telescopes at progressive automation levels: AIMS itself, the WenQuan Solar Magnetic Field Telescope, and the Solar Full-disk Multi-layer Magnetograph (SFMM). Collectively, these deployments demonstrate the feasibility and stability of the proposed architecture for progressive telescope automation.",
      "short_abstract": "The mid-infrared solar magnetic field telescope AIMS (An Infrared System for the Accurate Measurement of Solar Magnetic Field) is the first ground-based telescope designed to directly measure solar magnetic fields via Zeeman splitting in the 8-14 um band,...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13561",
      "pdf_url": "https://arxiv.org/pdf/2607.13561.pdf"
    },
    {
      "id": "2607.13494",
      "anchor": "paper-2607-13494",
      "title": "A VAE-Driven Multi-Task Satellite-Aided Semantic Communication Framework for 6G-Enabled Connected Autonomous Vehicles",
      "authors": [
        "S. M. Abtahiul Alam",
        "Niloy Das",
        "Apurba Adhikary",
        "Yu Qiao",
        "Zhu Han",
        "Choong Seon Hong"
      ],
      "abstract": "The development of smart transportation systems and the introduction of 6G wireless communication technologies have significantly changed vehicle network topologies. Future connected autonomous vehicle (CAV) networks require bandwidth-efficient, reliable, and low-latency communication for safety-critical applications such as traffic sign recognition and decision-making. Conventional communication systems transmit raw data regardless of task relevance, which is inefficient in resource-constrained satellite channels where uplink bandwidth is scarce and propagation losses are large. Semantic communication addresses this limitation by transmitting task-relevant information instead of full signal representations. It extracts and conveys essential semantic features and leverages deep learning to optimize task performance at the receiver. Therefore, we present a Variational Autoencoder (VAE)-based multi-task semantic communication framework for satellite-assisted autonomous driving. Unlike deterministic autoencoder-based methods, the proposed model uses probabilistic latent representations for more robust and efficient encoding. The learned features are transmitted over noisy wireless channels to perform traffic sign reconstruction and classification. The framework is trained end-to-end to jointly optimize both tasks. Results show that the proposed approach achieves significant bandwidth reduction of up to 87.23\\% to 98.17\\% while maintaining stable performance across varying signal-to-noise ratio conditions.",
      "short_abstract": "The development of smart transportation systems and the introduction of 6G wireless communication technologies have significantly changed vehicle network topologies. Future connected autonomous vehicle (CAV) networks require bandwidth-efficient, reliable, and...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 20,
        "evidence": [
          "title: connected vehicles",
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13494",
      "pdf_url": "https://arxiv.org/pdf/2607.13494.pdf"
    },
    {
      "id": "2607.13481",
      "anchor": "paper-2607-13481",
      "title": "GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors",
      "authors": [
        "Changqing Zhou",
        "Yueru Luo",
        "Yulan Guo",
        "Bing Wang",
        "Jie Qin",
        "Changhao Chen"
      ],
      "abstract": "Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of objects, structures, and free space. However, recovering such a complete volumetric representation from visual observations remains challenging, particularly in occluded and unobserved regions. Visual geometry priors offer strong and generalizable geometric cues for addressing this challenge, but their outputs are inherently surface-centric, whereas occupancy prediction requires reasoning about volumetric interiors and free space. To bridge this gap, we introduce GPOcc, which transforms visual geometry priors into occupancy-aware sparse Gaussian representations for efficient and expressive volumetric scene modeling. Building on GPOcc, GPOcc++ models multi-view observations and temporal sequences within a unified framework, allowing spatial and temporal evidence to be handled through the same representation. We further extend GPOcc++ from indoor scenes to outdoor occupancy prediction. Extensive experiments on both indoor and outdoor benchmarks demonstrate consistently strong performance across both multi-view and temporal settings, together with favorable efficiency and generalization. Code will be released at https://github.com/JuIvyy/GPOcc.",
      "short_abstract": "Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of objects, structures, and free space. However, recovering such a complete volumetric representation from...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 3,
        "evidence": [
          "title: occupancy",
          "abstract: occupancy",
          "abstract: scene understanding"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13481",
      "pdf_url": "https://arxiv.org/pdf/2607.13481.pdf"
    },
    {
      "id": "2607.14436",
      "anchor": "paper-2607-14436",
      "title": "Global drivers and barriers to the public acceptance of autonomous vehicles: Evidence from 17 countries",
      "authors": [
        "Antonios Saravanos"
      ],
      "abstract": "This study investigated the public acceptance of Society of Automotive Engineers Level 3 conditionally automated cars, which can self-drive under certain specified conditions but require the human driver to remain ready to resume control when requested. Previous Unified Theory of Acceptance and Use of Technology 2 (UTAUT2)-based research has focused mainly on European samples, and so it is still unclear whether the same factors shape acceptance across broader world regions. This knowledge gap was addressed using the L3Pilot Global User Acceptance Survey. From an original dataset of 18,631 respondents, the final analytic sample comprised 18,603 respondents from 17 countries across Africa, Asia, Europe, North America, and South America. The data were analyzed using a UTAUT2-based structural equation model to examine how performance expectancy, effort expectancy, social influence, facilitating conditions, and hedonic motivation shape the intention to use Level 3 cars. The model showed strong explanatory power. Across the analytic sample, the intention to use Level 3 cars was driven mainly by performance expectancy, social influence, and hedonic motivation. Effort expectancy and facilitating conditions also contributed, but they played smaller direct roles. Age, gender, and previous experience with advanced driver assistance systems were statistically significant, but comparatively weak predictors. Overall, the findings suggest that the acceptance of Level 3 automated cars depends less on demographic characteristics or ease-of-use concerns and more on whether people see the technology as useful, socially supported, and enjoyable to use.",
      "short_abstract": "This study investigated the public acceptance of Society of Automotive Engineers Level 3 conditionally automated cars, which can self-drive under certain specified conditions but require the human driver to remain ready to resume control when requested....",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: human driver",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.14436",
      "pdf_url": "https://arxiv.org/pdf/2607.14436.pdf"
    },
    {
      "id": "2607.13806",
      "anchor": "paper-2607-13806",
      "title": "A Deployed Hybrid Vehicle-in-the-Loop Platform for Validating Cooperative Perception",
      "authors": [
        "Anastasia Bolovinou",
        "Giorgos Hadjipavlis",
        "Markos Antonopoulos",
        "Panagiotis Tachtalis",
        "Konstantinos Petousakis",
        "Konstantinos Lazaridis",
        "Alexandros Siskos",
        "Bill Roungas",
        "Angelos Amditis"
      ],
      "abstract": "European safety regulation now permits a large share of automated-driving homologation evidence to be produced virtually, provided a validated physical-virtual facility generates it. We present a deployed hybrid Vehicle-in-the-Loop (ViL) platform that couples a real instrumented vehicle with a CARLA-based digital twin (DT) through a V2X message pipeline, and we report its first integrated operation on a public-road-representative test track. A real vehicle streams ETSI-compliant CAM/CPM messages into the DT, where a GPU-accelerated Cooperative Perception (CP) module fuses them into a probabilistic occupancy grid during scenario runtime. We demonstrate the platform on a multi-vehicle double T-intersection scenario, characterise the CP workload across nominal, rain and night conditions and five localization-noise levels, and discuss the platform's current architectural limits and the engineering targets they define. The results show that CP substantially widens field-of-view (FoV) coverage and improves occupied-cell recall, and that beyond a moderate localization-noise threshold, positioning uncertainty, and not weather, becomes the dominant error source. We outline the platform's trajectory toward a Mediterranean operational design domain (ODD) testing service.",
      "short_abstract": "European safety regulation now permits a large share of automated-driving homologation evidence to be produced virtually, provided a validated physical-virtual facility generates it. We present a deployed hybrid Vehicle-in-the-Loop (ViL) platform that couples...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 18,
        "margin": 4,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13806",
      "pdf_url": "https://arxiv.org/pdf/2607.13806.pdf"
    },
    {
      "id": "2607.20554",
      "anchor": "paper-2607-20554",
      "title": "AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optimize Graph Neural Networks",
      "authors": [
        "Giambattista Amati",
        "Federica Mangiatordi",
        "Simone Angelini",
        "Emiliano Pallotti",
        "Pierpaolo Salvo"
      ],
      "abstract": "Reliable and low-latency NR-V2X communications are essential for smart mobility in dense urban environments. However, limited Road-Side Unit (RSU) density, frequent non-line-of-sight conditions, and highly dynamic vehicular topologies often prevent many Connected and Automated Vehicles (CAVs) from maintaining stable single-hop connectivity. Although multi-hop relay-assisted communication can extend infrastructure coverage, selecting relay links in real time under practical flow, capacity, and connectivity constraints remains challenging. Mixed-Integer Linear Programming (MILP) yields optimal multi-hop relay decisions, but its computational complexity scales sharply with network density, limiting real-time applicability. To address this, we propose a Learning-to-Optimise (L2O) framework based on Graph Neural Networks (GNNs) for real-time NR-V2X relay selection. Vehicular communication states are modeled as attributed graphs, where CAVs and RSUs are nodes and candidate radio links are enriched with propagation-aware features. An offline MILP oracle provides optimal supervision, while an edge-aware Graph Isomorphism Network (GINE) approximates oracle decisions with near-constant inference latency. Experiments on large-scale urban datasets generated by an integrated SUMO--GEMV2 simulation pipeline show that the proposed approach achieves connectivity comparable to that of the MILP oracle while reducing execution time by orders of magnitude. The framework enables cost-effective enhancement of urban V2X connectivity by leveraging existing vehicular assets and supporting scalable, real-time NR-V2X operation in smart city environments.",
      "short_abstract": "Reliable and low-latency NR-V2X communications are essential for smart mobility in dense urban environments. However, limited Road-Side Unit (RSU) density, frequent non-line-of-sight conditions, and highly dynamic vehicular topologies often prevent many...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 41,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.20554",
      "pdf_url": "https://arxiv.org/pdf/2607.20554.pdf"
    },
    {
      "id": "2607.13692",
      "anchor": "paper-2607-13692",
      "title": "Proactive URLLC Adaptation for Connected Vehicles Through ML-Based Channel Prediction",
      "authors": [
        "Andrea Giovannini",
        "Lorenzo Mario Amorosa",
        "Vittorio Todisco",
        "Claudia Campolo",
        "Antonella Molinaro",
        "Su Hongjia",
        "Alessandro Bazzi"
      ],
      "abstract": "Connected and automated vehicles (CAVs) are expected to increasingly rely on 5G and future 6G ultra-reliable and low-latency communication (URLLC) services to support safety-critical and time-sensitive applications. Since wireless link conditions can vary rapidly in urban vehicular environments, proactively adapting service parameters based on future channel conditions is essential to maintain service continuity and reliability. In this paper, we investigate the use of machine learning (ML) techniques for channel quality prediction in vehicular URLLC scenarios. Specifically, we evaluate deep neural network (DNN) and long short-term memory (LSTM) models to forecast future channel conditions and enable proactive service adaptation with minimized performance degradation. The analysis is conducted using realistic simulations combining the SUMO traffic simulator and the Sionna-RT ray-tracing framework in a real urban environment reconstructed from OpenStreetMap data. Results show that ML-based prediction significantly outperforms approaches relying solely on past channel measurements and achieves performance close to the ideal case in which future channel conditions are perfectly known in advance. These findings demonstrate the potential of ML-driven prediction techniques to enhance the reliability and robustness of URLLC services for connected vehicular systems.",
      "short_abstract": "Connected and automated vehicles (CAVs) are expected to increasingly rely on 5G and future 6G ultra-reliable and low-latency communication (URLLC) services to support safety-critical and time-sensitive applications. Since wireless link conditions can vary...",
      "published": "2026-07-15",
      "updated": "2026-07-15",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 12,
        "evidence": [
          "title: connected vehicles",
          "abstract: connected vehicles",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 35,
      "arxiv_url": "https://arxiv.org/abs/2607.13692",
      "pdf_url": "https://arxiv.org/pdf/2607.13692.pdf"
    },
    {
      "id": "2607.13410",
      "anchor": "paper-2607-13410",
      "title": "Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation",
      "authors": [
        "Zhidong Wang",
        "Jingsong Liang",
        "Zirui Li",
        "Zhan Chen",
        "Han Yu",
        "Chen Lv"
      ],
      "abstract": "World model (WM)-based reinforcement learning enables sample-efficient end-to-end autonomous driving learning by imagining long-horizon trajectories in latent space. However, most driving WMs operate on bird's-eye-view (BEV) representations that are inherently egocentric: the transition between consecutive frames entangles the ego vehicle's own motion with scene dynamics. As a result, the WM devotes significant capacity to recovering ego-motion from warped observations, at the cost of scene modeling fidelity and imagination accuracy. This work proposes DynaDreamer, a dynamics-augmented Dreamer-style reinforcement learning method to address this problem by augmenting the WM with an explicit ego-dynamics prior. A physics-informed ego-dynamics encoder-decoder extracts the ego-state history into a compact and identifiable context, which modulates a causal Transformer WM to condition both its prior and posterior latents. During imagination, the ego-dynamics predictor propagates this context forward to keep the ego-dynamics prior synchronized with the rollout. An information-theoretic analysis shows that conditioning on this context reduces both the predictive entropy of the observation transition and the prior--posterior Kullback--Leibler divergence, confining the WM's modeling burden to the scene dynamics beyond ego-motion. An additional benefit is zero-shot cross-chassis adaptation: the ego-dynamics context depends on identifiable chassis parameters, so that a vehicle with previously unseen dynamic characteristics can adapt the WM to the new chassis without retraining. Experiments demonstrate that DynaDreamer improves task success rates over the strongest baseline by 28% and 61% in urban and highway driving scenarios, respectively, with the advantage rising to 73% when extrapolating to unseen chassis.",
      "short_abstract": "World model (WM)-based reinforcement learning enables sample-efficient end-to-end autonomous driving learning by imagining long-horizon trajectories in latent space. However, most driving WMs operate on bird's-eye-view (BEV) representations that are...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 7,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.13410",
      "pdf_url": "https://arxiv.org/pdf/2607.13410.pdf"
    },
    {
      "id": "2607.13028",
      "anchor": "paper-2607-13028",
      "title": "TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale",
      "authors": [
        "Zhouchonghao Wu",
        "Akshay Rangesh",
        "Weixin Li",
        "Wei-Jer Chang",
        "Zachary Lee",
        "Saeed Bonab",
        "Tim Wang",
        "Wei Zhan"
      ],
      "abstract": "Training robust autonomous driving agents requires a simulator fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack that meets these goals. A configurable C engine runs simulation on the CPU and policy inference on the GPU over a zero-copy path, sustaining 1.3M agent-steps per second on a single server-grade GPU, far faster than existing object-level simulators, while keeping fidelity lighter single-agent systems omit: heterogeneous agents, multiple dynamics models, and full traffic-rule enforcement. TerraZero uses logged data only as a source of real-world map geometry, populating each map with randomized rule-based road users and signal controllers and randomizing agent dynamics, rewards, and sizes per episode, so one map yields an effectively unbounded set of scenarios. Every reported policy trains from scratch by reinforcement learning alone, with zero human demonstrations, no imitation, no logged trajectories, and no fallback planner at inference, on a compute-efficient self-play recipe scaled across GPUs. The policies generalize zero-shot across cities and datasets, including emergent left-hand-traffic driving without explicit supervision. As an ego policy, a single checkpoint is, to our knowledge, the first fully learned policy to top both val14 and the interactive long-tail InterPlan suite. On Waymo Open Sim Agents realism the same recipe outperforms other demonstration-free methods and is competitive with the strongest reference-anchored self-play method. One stack serves both roles: state-of-the-art demonstration-free driving policies across dynamics for cars and trucks, and sim agents that jointly control vehicles, pedestrians, and cyclists.",
      "short_abstract": "Training robust autonomous driving agents requires a simulator fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.13028",
      "pdf_url": "https://arxiv.org/pdf/2607.13028.pdf"
    },
    {
      "id": "2607.12858",
      "anchor": "paper-2607-12858",
      "title": "LARAD: Layout-Aware Road Anomaly Detection via Spatial-Logic Reasoning",
      "authors": [
        "Shiyi Mu",
        "Xujie Chen",
        "Shugong Xu"
      ],
      "abstract": "Accurate open-world obstacle detection is critical for autonomous driving. Current anomaly segmentation methods suffer from a fundamental blind spot: they over-rely on texture novelty to identify out-of-distribution (OoD) objects while ignoring contextual spatial logic. Furthermore, mitigating the resulting false positives often requires cascading massive vision models, introducing unacceptable inference latency. To address these issues, we propose Layout-Aware Road Anomaly Detection (LARAD), shifting the paradigm from appearance matching to spatial-logic reasoning. First, we introduce the Spatial-Logic Violation Synthesis (SLVS) pipeline, which generates training samples that are texture-consistent yet spatially invalid, forcing the model to learn contextual violations. Second, we augment a standard closed-set segmentation network with a lightweight, OoD-guided attention branch. Extensive experiments demonstrate that LARAD significantly enhances robustness against logical anomalies and establishes a new state-of-the-art, all while retaining the high efficiency of a single-model architecture.",
      "short_abstract": "Accurate open-world obstacle detection is critical for autonomous driving. Current anomaly segmentation methods suffer from a fundamental blind spot: they over-rely on texture novelty to identify out-of-distribution (OoD) objects while ignoring contextual...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: communications",
          "Systems, Deployment & Connectivity: abstract: systems constraints",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.12858",
      "pdf_url": "https://arxiv.org/pdf/2607.12858.pdf"
    },
    {
      "id": "2607.12598",
      "anchor": "paper-2607-12598",
      "title": "Hierarchical Fault Localization for Autonomous Driving Systems with Hypothesis Validation and Intent Analysis",
      "authors": [
        "Rui Zheng",
        "Changwen Li",
        "Yi Ji",
        "Rongjie Yan"
      ],
      "abstract": "Comprehensive testing is essential for the safety and reliability of Autonomous Driving Systems (ADS). Existing techniques can detect system-level failures or attribute them to coarse-grained modules, but they often fall short of localizing the root cause in source code. As a result, debugging remains labor-intensive, requiring developers to connect behavioral violations with complex implementation logic. To address this gap, we present HINT, a two-phase framework for hierarchical ADS fault localization based on hypothesis validation and intent analysis. In Phase I, HINT transforms failure-triggering execution recordings into multi-modal abstractions and uses causal reasoning to identify the responsible module. In Phase II, it reconstructs design-side intent and implementation-side behavior, then localizes suspicious code through reliability-aware consistency checking, without costly re-simulation. We evaluate HINT on Apollo across diverse failure modes and modules. The results show that HINT achieves the strongest overall performance across module-level diagnosis and code-level localization metrics, with 77.8% end-to-end Class@5 accuracy on real-world bugs.",
      "short_abstract": "Comprehensive testing is essential for the safety and reliability of Autonomous Driving Systems (ADS). Existing techniques can detect system-level failures or attribute them to coarse-grained modules, but they often fall short of localizing the root cause in...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 20,
        "margin": 2,
        "evidence": [
          "title: localization",
          "abstract: localization"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.12598",
      "pdf_url": "https://arxiv.org/pdf/2607.12598.pdf"
    },
    {
      "id": "2607.12593",
      "anchor": "paper-2607-12593",
      "title": "Improving Autonomous Nano-drones Performance via Automated End-to-End Optimization and Deployment of DNNs",
      "authors": [
        "Vlad Niculescu",
        "Lorenzo Lamberti",
        "Francesco Conti",
        "Luca Benini",
        "Daniele Palossi"
      ],
      "abstract": "The evolution of energy-efficient ultra-low-power (ULP) parallel processors and the diffusion of convolutional neural networks (CNNs) are fueling the advent of autonomous driving nano-sized unmanned aerial vehicles (UAVs). These sub-10 cm robotic platforms are envisioned as next-generation ubiquitous smart-sensors and unobtrusive robotic-helpers. However, the limited computational/memory resources available aboard nano-UAVs introduce the challenge of minimizing and optimizing vision-based CNNs -- which to date require error-prone, labor-intensive iterative development flows. This work explores methodologies and software tools to streamline and automate all the deployment of vision-based CNN navigation on a ULP multicore system-on-chip acting as a mission computer on a Crazyflie 2.1 nano-UAV. We focus on the deployment of PULP-Dronet, a state-of-the-art CNN for autonomous navigation of nano-UAVs, from the initial training to the final closed-loop evaluation. Compared to the original hand-crafted CNN, our results show a 2x reduction of memory footprint and a speedup of 1.6x in inference time while guaranteeing the same prediction accuracy and significantly improving the behavior in the field, achieving: i) obstacle avoidance with a peak braking-speed of 1.65 m/s and improving the speed/braking-space ratio of the baseline, ii) free flight in a familiar environment up to 1.96 m/s (0.5 m/s for the baseline), and iii) lane following on a path featuring a 90 deg turn -- all while using for computation less than 1.6% of the drone's power budget. To foster new applications and future research, we open-source all the software design in a ready-to-run project compatible with the Crazyflie 2.1",
      "short_abstract": "The evolution of energy-efficient ultra-low-power (ULP) parallel processors and the diffusion of convolutional neural networks (CNNs) are fueling the advent of autonomous driving nano-sized unmanned aerial vehicles (UAVs). These sub-10 cm robotic platforms...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 5,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.12593",
      "pdf_url": "https://arxiv.org/pdf/2607.12593.pdf"
    },
    {
      "id": "2607.12419",
      "anchor": "paper-2607-12419",
      "title": "DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection",
      "authors": [
        "Haifa Zhang",
        "Yijing Wang",
        "Peixi Peng",
        "Zhiqiang Zuo"
      ],
      "abstract": "In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current multi-modal frameworks heavily rely on massive visual backbones pretrained on 2D semantic tasks. This reliance introduces substantial parameter redundancy and a structural misalignment, as 2D priors are ill-equipped to handle the extreme sparsity of LiDAR projections required for Bird's-Eye-View geometry. To address this, we present DeGuNet, an ultra-compact and plug-and-play image backbone explicitly designed for depth-guided representation learning. By incorporating sparsity-aware feature extraction mechanisms, DeGuNet effectively aligns multi-view images with unstructured LiDAR depth while strictly preventing invalid-region contamination. Extensive experiments on the nuScenes dataset demonstrate DeGuNet's broad plug-and-play applicability and superior efficiency. When integrated into established baselines, it fundamentally eliminates architectural redundancy, reducing GPU memory consumption by up to 66.5% and achieving a 1.16x inference speedup. Concurrently, DeGuNet delivers up to a 6.20 absolute mAP gain, establishing a new paradigm for parameter-efficient multi-modal 3D perception.",
      "short_abstract": "In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current multi-modal frameworks heavily rely on massive visual backbones pretrained on 2D semantic tasks. This...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 27,
        "evidence": [
          "abstract: perception",
          "abstract: object detection",
          "title: LiDAR",
          "abstract: LiDAR",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.12419",
      "pdf_url": "https://arxiv.org/pdf/2607.12419.pdf"
    },
    {
      "id": "2607.13274",
      "anchor": "paper-2607-13274",
      "title": "Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners",
      "authors": [
        "Haseeb Shah",
        "Lingwei Zhu",
        "Adam White",
        "Martha White"
      ],
      "abstract": "Reinforcement learning is increasingly being considered for controlling real-world systems, from fusion plasma and autonomous vehicles to drug discovery and drinking water treatment, where reliability is essential and tuning budgets are limited. Actor-critic algorithms share a set of design decisions, such as how the policy is updated, how it represents the distribution over actions, how its gradient is estimated, and how often it is updated relative to the value estimator. Using a control task derived from a real water treatment plant, we analyze over 33,000 experiments to determine how these components affect variability across runs and sensitivity to hyperparameters. Common defaults, such as Gaussian action distributions with pathwise gradient estimators, are among the least reliable configurations, whereas bounded distributions with adaptive update schedules remain robust across a wide range of settings. These findings offer empirical guidance to practitioners across scientific and engineering domains for understanding and making component-level decisions when adapting actor-critic methods to new real-world control settings.",
      "short_abstract": "Reinforcement learning is increasingly being considered for controlling real-world systems, from fusion plasma and autonomous vehicles to drug discovery and drinking water treatment, where reliability is essential and tuning budgets are limited. Actor-critic...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.13274",
      "pdf_url": "https://arxiv.org/pdf/2607.13274.pdf"
    },
    {
      "id": "2607.13254",
      "anchor": "paper-2607-13254",
      "title": "Attitude Estimation Using Inertial and Barometric Measurements",
      "authors": [
        "Melone Nyoba Tchonkeu",
        "Soulaimane Berkane",
        "Tarek Hamel"
      ],
      "abstract": "Accurate and robust attitude estimation is a key challenge for autonomous vehicles, particularly in GNSS-denied conditions and during highly accelerated flight. In such conditions, Inertial Measurement Units (IMUs) alone are insufficient for reliable tilt estimation due to the ambiguity between gravitational and inertial accelerations. Although auxiliary velocity sensors such as GNSS, Pitot tubes, Doppler radar, or Visual Inertial Odometry are commonly used, they may be unavailable, intermittent, or costly. This paper introduces a barometer-aided attitude estimation architecture that exploits barometric altitude measurements to provide complementary information on the vehicle's vertical motion, thereby enhancing attitude estimation within nonlinear observers on SO(3). The contributions are twofold. First, we design a deterministic Riccati observer cascaded with a complementary filter, ensuring almost-global asymptotic stability (AGAS) under a uniform observability (UO) condition while preserving the geometric structure of the attitude dynamics. Second, we propose a nonlinear observer evolving on SO(3)xR2, which integrates IMU measurements as inputs and barometer and magnetometer measurements as outputs within a unified framework, guaranteeing local exponential stability (LES) under relaxed uniform observability conditions. The proposed approaches are validated using both simulated and real flight data. The results demonstrate that barometer-aided estimation provides a lightweight, reliable, and effective complementary sensing modality for attitude estimation in minimal-sensing configurations, offering a practical alternative when conventional velocity measurements are unavailable or degraded.",
      "short_abstract": "Accurate and robust attitude estimation is a key challenge for autonomous vehicles, particularly in GNSS-denied conditions and during highly accelerated flight. In such conditions, Inertial Measurement Units (IMUs) alone are insufficient for reliable tilt...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: radar",
          "Mapping & Localization: abstract: GNSS or GPS"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.13254",
      "pdf_url": "https://arxiv.org/pdf/2607.13254.pdf"
    },
    {
      "id": "2607.12833",
      "anchor": "paper-2607-12833",
      "title": "ANGLE: Angular Neural Generative Learning via Engression",
      "authors": [
        "Rajdeep Pathak",
        "Archi Roy",
        "Tanujit Chakraborty"
      ],
      "abstract": "Circular data, representing angles or directions, are frequently encountered in computer vision, biology, geology, and meteorology. Traditional regression targets the conditional mean, which is often geometrically misleading for circular responses under multimodal, skewed, or asymmetric data structures. To address these limitations, a lightweight deep generative framework, namely ANGLE, is introduced for non-parametric distributional regression on the circle. The full conditional distribution of an angular response, given Euclidean and circular covariates, is learned through a generative map optimized via a generalized circular energy score (GCES) loss. Desirable theoretical properties, including the strict propriety of the loss and the rotational equivariance of the estimators, are established. Furthermore, both pre- and post-additive noise models are accommodated. A unified toolbox is provided for advancing previously underexplored challenges in circular statistics: extrapolation, sufficient dimension reduction, and conditional distribution equality testing. The framework's efficacy is demonstrated through extensive simulations and real-world applications. Specifically, the proposal is utilized for object pose estimation from imagery and wind direction prediction, which are integral to surveillance, autonomous vehicles, and energy systems, respectively. Superior predictive performance and robust uncertainty quantification of the proposed method in these tasks are revealed.",
      "short_abstract": "Circular data, representing angles or directions, are frequently encountered in computer vision, biology, geology, and meteorology. Traditional regression targets the conditional mean, which is often geometrically misleading for circular responses under...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 3,
        "evidence": [
          "abstract: validation or testing",
          "abstract: robustness",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.12833",
      "pdf_url": "https://arxiv.org/pdf/2607.12833.pdf"
    },
    {
      "id": "2607.12452",
      "anchor": "paper-2607-12452",
      "title": "Automatic Testing of Interacting Autonomous Vehicles",
      "authors": [
        "Fabio Cavaleri",
        "Alessio Gambi",
        "Paolo Arcaini",
        "Dejan Ničković",
        "Mauro Pezzè"
      ],
      "abstract": "Autonomous vehicles (AVs) must be thoroughly tested to meet high safety standards and avoid endangering both AV passengers and road users. Scenario-based testing implements driving scenarios in virtual simulation environments as a cost-effective alternative to field testing. Common scenario-based testing approaches set the environment and the surrounding traffic and test a single AV. Recent studies show that the approaches that test single AVs miss critical behaviors that emerge from interactions among multiple AVs. Effective approaches to test scenarios that emerge from n-way interactions must address the combinatorial explosion that the presence of multiple AVs further exacerbates. In this paper, we propose EVITA, an approach that leverages multi-objective optimization to generate scenarios that trigger multiple and diverse AVs interactions, while minimizing the complexity of the generated scenarios, to effectively test multiple interacting AVs and reveal safety-critical scenarios that current approaches overlook. The experimental results that we discuss in this paper confirm that EVITA triggers a higher variety of AVs interactions than state-of-the-art approaches, thus improving the likelihood to reveal safety-critical behaviors.",
      "short_abstract": "Autonomous vehicles (AVs) must be thoroughly tested to meet high safety standards and avoid endangering both AV passengers and road users. Scenario-based testing implements driving scenarios in virtual simulation environments as a cost-effective alternative...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 16,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.12452",
      "pdf_url": "https://arxiv.org/pdf/2607.12452.pdf"
    },
    {
      "id": "2607.12352",
      "anchor": "paper-2607-12352",
      "title": "Filtering-out poor-quality images for data preparation",
      "authors": [
        "Roopdeep Kaur",
        "Gour Karmakar",
        "Muhammad Imran"
      ],
      "abstract": "Filtering noise is a fundamental part of data preparation that enhances image quality for applications such as object segmentation, detection, and recognition. Various noise reduction techniques are proposed in the literature, including the use of median, Gaussian, and bilateral filters. Convolutional neural networks (CNNs) have gained popularity in image denoising owing to their ability to extract complex patterns and features from data. CNNs are highly adaptable, making them effective tools for various image-denoising tasks. One drawback of CNN-based techniques is that they require an appropriate training dataset and all images to be resized. Another notable drawback of all these filtering techniques is that they work for certain types of environmental and camera noises. To bridge this research gap, in this paper, for the first time, instead of denoising, we propose an approach that filters out poor-quality images for various environmental and camera impacts. In our approach, quality is assessed using an image quality assessment metric and an optimum threshold is used to filter out poor-quality images. We also ensure that a sufficient number of images remain to develop the deep learning (DL) model. The results produced using real and simulated traffic and object recognition data demonstrate the performance supremacy of the proposed approach compared with the state-of-the-art approaches. The average recognition accuracy for our proposed approach is 93.8% for the traffic sign recognition dataset and 84.9% for the object recognition dataset. This indicates our model's potential for real-life applications such as autonomous vehicles.",
      "short_abstract": "Filtering noise is a fundamental part of data preparation that enhances image quality for applications such as object segmentation, detection, and recognition. Various noise reduction techniques are proposed in the literature, including the use of median,...",
      "published": "2026-07-14",
      "updated": "2026-07-14",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera",
          "Perception & Sensor Fusion: abstract: traffic sign recognition",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 36,
      "arxiv_url": "https://arxiv.org/abs/2607.12352",
      "pdf_url": "https://arxiv.org/pdf/2607.12352.pdf"
    },
    {
      "id": "2607.12293",
      "anchor": "paper-2607-12293",
      "title": "Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction",
      "authors": [
        "Md Mahfuzur Rahman",
        "Pengzhan Zhou",
        "A F M Abdun Noor",
        "Md Imam Ahasan",
        "Kah Ong Michael Goh",
        "S. M. Hasan Mahmud",
        "Md Mustafizur Rahman",
        "Kaixin Gao"
      ],
      "abstract": "Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or dense multimodal fusion strategies that inadequately capture complementary visual and kinematic information while introducing redundant inter-modal interactions. We propose ADAPT (Adaptive Domain-Aware Pedestrian Crossing Transformer), a multimodal framework that jointly models local and global visual context together with temporal motion dynamics for accurate pedestrian crossing intention prediction. ADAPT processes four spatially aligned visual modalities, including RGB images, local depth maps, global semantic maps, and global depth maps, together with ego-vehicle speed, pedestrian bounding boxes, and skeleton pose information through five specialized modules: a weight-shared Swin Transformer V2 backbone for visual feature extraction, a Cross-Modality Guided Attention module for hierarchical visual fusion, a Mamba-based Motion Feature Encoding module for efficient temporal modeling, a Sparse Cross-Modal Attention module that selectively preserves the most informative inter-modal interactions, and a Vision Transformer-based Temporal Feature Fusion module for sequence-level prediction. Extensive experiments on the JAAD and PIE benchmark datasets demonstrate that ADAPT consistently outperforms existing state-of-the-art methods while maintaining low computational complexity. On JAAD, the proposed method achieves an AUC of 0.73 on JAADbeh and 0.85 on JAADall, while on PIE it achieves an accuracy of 0.92 and an AUC of 0.90. Furthermore, ADAPT performs inference in only 17.23 ms per sample, offering an effective balance between predictive accuracy and real-time deployment efficiency for intelligent transportation and autonomous driving applications.",
      "short_abstract": "Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or dense multimodal fusion strategies that inadequately capture complementary visual and kinematic information...",
      "published": "2026-07-13",
      "updated": "2026-07-13",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 20,
        "evidence": [
          "title: behavior prediction",
          "abstract: behavior prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 37,
      "arxiv_url": "https://arxiv.org/abs/2607.12293",
      "pdf_url": "https://arxiv.org/pdf/2607.12293.pdf"
    },
    {
      "id": "2607.11560",
      "anchor": "paper-2607-11560",
      "title": "Technical Report on the CVPR 2026@AdvML Workshop Challenge",
      "authors": [
        "Tianyuan Zhang",
        "Zonglei Jing",
        "Jiangfan Liu",
        "Ligong Zhang",
        "Ke Ma",
        "Chengzhi Sun",
        "Xiaohai Xu",
        "Zhirui Zhang",
        "Qianqian Xu",
        "Qingming Huang",
        "Hanyu Fang",
        "Junhua Liu",
        "Zheng Wang",
        "Xiaoliang Liu",
        "Yuanbo Li",
        "Shuai Gui",
        "Bin Wang",
        "Menghe Zheng",
        "Jing Nie",
        "Hanyang Meng",
        "Zeyang Zhang",
        "Xiang Zhang",
        "Yongxuan Zhu",
        "Rui Ding",
        "Hainan Li"
      ],
      "abstract": "Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images and suffix-only textual perturbations that induce model responses to deviate from reference answers while preserving image fidelity and limiting textual cost. The competition comprises two phases, with Phase II adding a hidden black-box model to assess transferability. We describe the task design, submission rules, evaluation protocol, and leaderboard results, and then examine five leading submissions for which technical reports were available. Across these reports, several recurring patterns emerge: image-side attacks are favored by the suffix penalty; scene-level, multi-view optimization is more effective than treating views in isolation; QA types and graph structure provide useful priors for allocating attack budget; feature-space objectives can improve black-box transfer; and typographic content embedded in camera images exposes a persistent vulnerability in driving VLAs. These findings provide a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.",
      "short_abstract": "Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built...",
      "published": "2026-07-13",
      "updated": "2026-07-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 8,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 37,
      "arxiv_url": "https://arxiv.org/abs/2607.11560",
      "pdf_url": "https://arxiv.org/pdf/2607.11560.pdf"
    },
    {
      "id": "2607.12113",
      "anchor": "paper-2607-12113",
      "title": "Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap",
      "authors": [
        "Rafael Ferreira da Silva",
        "Milad Abolhasani",
        "Peter Beaucage",
        "Laura Biven",
        "Michael Bussmann",
        "Kyle Chard",
        "Ryan Coffee",
        "Stephen DeWitt",
        "Sagar Dolas",
        "Carrie Eckert",
        "David Elbert",
        "Ian Foster",
        "Tirthankar Ghosal",
        "Anna Giannakou",
        "Tom Gibbs",
        "Leslie Hamilton",
        "Glenn Lockwood",
        "Theresa Mayer",
        "Ben Mintz",
        "Raffi Nazikian",
        "Sal Nimer",
        "Amanda Randles",
        "Woong Shin",
        "Sreenivas Rangan Sukumar",
        "Frédéric Suter"
      ],
      "abstract": "One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have produced experimentally validated hypotheses, self-driving laboratories have grown more interoperable and orchestrated, reasoning-trained and domain foundation models have raised the capability ceiling, and the Genesis Mission has placed autonomous experimentation at the center of U.S. federal science strategy, with industry emerging as a primary actor. Progress has met a sobering counter-current, including a corrected flagship discovery result, benchmarks showing that agents which rival experts on closed-ended questions still complete only a fraction of open-ended research, and fabricated citations surfacing at leading venues. We read this as the defining tension of the field. Producing a candidate discovery is no longer the hard part, but verifying it is, and this asymmetry now limits autonomous science more than raw model capability. We update the roadmap around seven dimensions, revisiting the original five and elevating two former cross-cutting concerns, trust, verification, and reproducibility, and safety, security, and governance, to first-class status. We assess the original milestones (M1 through M14) as achieved, partially achieved, reframed, or open, add four new milestones (M15 through M18), and scope the path forward to a two-year horizon. The first year concentrates on interfaces, protocol adoption, and the scaffolding of verification, and the second targets federation, zero-trust coordination, and governance. Throughout, we position the grassroots network as the interoperability fabric that lets national programs, international initiatives, and commercial platforms connect rather than re-silo.",
      "short_abstract": "One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have produced...",
      "published": "2026-07-13",
      "updated": "2026-07-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: verification"
        ]
      },
      "recency": "fresh",
      "age_days": 37,
      "arxiv_url": "https://arxiv.org/abs/2607.12113",
      "pdf_url": "https://arxiv.org/pdf/2607.12113.pdf"
    },
    {
      "id": "2607.11232",
      "anchor": "paper-2607-11232",
      "title": "Modular Autonomous Transit: From Vehicle Modularity to Deployment-Ready Transit Operations",
      "authors": [
        "Dongyang Xia",
        "Shadi Sharif Azadeh"
      ],
      "abstract": "Audience: Tech companies such as NExT leadership; public transport operators; autonomous transit developers; charging partners; public authorities, and industrial partners considering pilots. Content: The strongest aspect of Modular Autonomous Vehicles (MAVs) is not only that vehicles can physically couple. A MAV is composed of Modular Autonomous Units (MAUs), and modularity creates operating advantages across line-level operations, intermodal services, network-level operations, stochastic real-time control, and electrified vehicle operations and charging infrastructure. Research basis and disclaimer: The brief summarizes scientific findings of four studies conducted by Dr. Xia and Dr. Sharif Azadeh on modular transit operations, including a charging infrastructure study using technical data associated with NExT MAUs.",
      "short_abstract": "Audience: Tech companies such as NExT leadership; public transport operators; autonomous transit developers; charging partners; public authorities, and industrial partners considering pilots. Content: The strongest aspect of Modular Autonomous Vehicles (MAVs)...",
      "published": "2026-07-13",
      "updated": "2026-07-13",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 12,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 37,
      "arxiv_url": "https://arxiv.org/abs/2607.11232",
      "pdf_url": "https://arxiv.org/pdf/2607.11232.pdf"
    },
    {
      "id": "2607.11128",
      "anchor": "paper-2607-11128",
      "title": "Comparison-Based Ordinal Learning for Proactive Driving Risk Assessment",
      "authors": [
        "Zhuoren Li",
        "Yi Zhong",
        "Weiqi Zhang",
        "Xinrui Zhang",
        "Lu Xiong",
        "Chongfeng Wei",
        "Bo Leng"
      ],
      "abstract": "Real-time driving risk assessment provides an essential basis for proactive safety by identifying and quantifying the danger of ongoing road interactions before adverse outcomes occur. However, due to the scarcity of collision data and frame-level risk labels, existing driving risk assessment methods often rely on surrogate objectives, which may imperfectly align with true collision risk and not faithfully reflect the relative danger of driving interaction. This paper proposes a comparison-based ordinal risk learning framework that learns collision-relevant risk scores from pairwise supervision in driving data, directly modeling relative risk ordering without requiring numerical frame-level risk labels. We derive pairwise comparisons from three sources of event-structured driving data for such ordinal risk learning: temporal progression within safety-critical sequences, event-level contrast between dangerous and normal interactions, and physics-based counterfactual perturbations. On this basis, instantiations with three risk-scoring function parameterizations are implemented, including directly learning risk scores from comparison data, and aligning existing single or multiple surrogate-based risk models. The proposed framework is evaluated on the 100-Car and SHRP2 naturalistic driving datasets using a proactive collision warning task. Results show that the proposed framework improves high-recall risk discrimination, warning precision, and warning lead time over representative surrogate-based baselines across both in-distribution and out-of-distribution evaluations. These results suggest that the proposed framework can contribute to proactive safety research by providing more reliable risk assessment for automated driving systems and safety-critical driving interactions.",
      "short_abstract": "Real-time driving risk assessment provides an essential basis for proactive safety by identifying and quantifying the danger of ongoing road interactions before adverse outcomes occur. However, due to the scarcity of collision data and frame-level risk...",
      "published": "2026-07-13",
      "updated": "2026-07-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 17,
        "evidence": [
          "abstract: safety",
          "title: risk assessment",
          "abstract: risk assessment"
        ]
      },
      "recency": "fresh",
      "age_days": 37,
      "arxiv_url": "https://arxiv.org/abs/2607.11128",
      "pdf_url": "https://arxiv.org/pdf/2607.11128.pdf"
    },
    {
      "id": "2607.11403",
      "anchor": "paper-2607-11403",
      "title": "Decentralized Model Predictive Control of Connected and Automated Vehicles with Coupled Safety Constraints",
      "authors": [
        "Philip Schultheis",
        "Kimia Chavoshi",
        "John Lygeros"
      ],
      "abstract": "Connected and Automated Vehicles (CAVs) operating on lane-free highways offer substantial gains in traffic efficiency. However, their inherent nonlinear dynamics and the presence of coupled, nonconvex safety constraints present critical challenges to control design. Centralized Model Predictive Control (MPC) ensures safety, but suffers from scalability and communication limitations. To address these challenges, this paper investigates decentralized MPC (DMPC) for CAV coordination, focusing on iterative, non-cooperative algorithms, including Jacobi-type and Gauss-Seidel-type. A novel decoupling method is developed to transform nonconvex safety constraints into convex, locally enforceable constraints, inspired by buffered Voronoi cells. The simulation results show that the proposed DMPC algorithms achieve safe and efficient vehicle trajectories while substantially improving scalability, highlighting their potential for future lane-free CAV traffic systems. Ultimately, the results indicate that the most suitable decentralized control strategy depends on the desired trade-off between safety, performance, and computational efficiency.",
      "short_abstract": "Connected and Automated Vehicles (CAVs) operating on lane-free highways offer substantial gains in traffic efficiency. However, their inherent nonlinear dynamics and the presence of coupled, nonconvex safety constraints present critical challenges to control...",
      "published": "2026-07-13",
      "updated": "2026-07-13",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 7,
        "evidence": [
          "title: model predictive control",
          "abstract: model predictive control",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 37,
      "arxiv_url": "https://arxiv.org/abs/2607.11403",
      "pdf_url": "https://arxiv.org/pdf/2607.11403.pdf"
    },
    {
      "id": "2607.11964",
      "anchor": "paper-2607-11964",
      "title": "LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving",
      "authors": [
        "Yongzhi Liu",
        "Yang Xiao",
        "Zhong Cao",
        "Zeng Kang",
        "Sunan Zhang",
        "Zhaozhi Dong",
        "Guojun Yu",
        "Weichao Zhuang"
      ],
      "abstract": "Autonomous driving requires long-horizon closedloop decision making in dynamic traffic environments. Latent world models offer an effective framework for this problem by enabling imagination-based decision making in compact latent spaces. However, multi-source observations contain controlirrelevant redundancy, whereas reliable driving decisions rely on risk-relevant relations, future dynamics, and continuous action adjustments. This mismatch makes observation reconstruction and absolute action modeling suboptimal for learning decisionrelevant latent dynamics. We propose LIDAR-AD, a decoderfree Latent-Interaction Dreamer with Action-Residual Chains for autonomous driving. LIDAR-AD replaces observation reconstruction with redundancy-reduced latent alignment, encouraging compact representations of risk-relevant relations in multi-source driving inputs. It further models vehicle control as residual action updates and uses residual-action sequence contrastive learning to align multi-step residual-driven rollouts with future latent states. A deterministic analysis shows that the latent-tanh residual parameterization preserves interior action reachability while representing smooth long-horizon control as compact local updates. Together, these designs improve risk-aware state abstraction, continuous-control modeling, and long-horizon dynamics prediction. Extensive experiments across diverse simulated driving scenarios demonstrate that LIDAR-AD consistently outperforms world-model baselines, achieving the highest reward and the best success rate among learning-based methods. Evaluations on nuPlan-derived log-reconstructed scenarios further demonstrate the transferability of LIDAR-AD under real-world traffic layouts.",
      "short_abstract": "Autonomous driving requires long-horizon closedloop decision making in dynamic traffic environments. Latent world models offer an effective framework for this problem by enabling imagination-based decision making in compact latent spaces. However,...",
      "published": "2026-07-12",
      "updated": "2026-07-12",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 38,
      "arxiv_url": "https://arxiv.org/abs/2607.11964",
      "pdf_url": "https://arxiv.org/pdf/2607.11964.pdf"
    },
    {
      "id": "2607.10975",
      "anchor": "paper-2607-10975",
      "title": "Real-Time Rulebook-Aware Nonlinear MPC for Autonomous Driving with Priority-Biased Tiered Slacks",
      "authors": [
        "Hadi Hajieghrary",
        "Benedikt Walter",
        "Chaitanya Shinde",
        "Paul Schmitt",
        "Miguel Hurtado"
      ],
      "abstract": "Autonomous-vehicle motion planners must resolve conflicts among safety, regulation, comfort, and efficiency in real time while exposing those decisions for audit. We present W-SQP, a weighted tiered-slack nonlinear model predictive controller (NMPC) that compiles nine driving-rule families into a four-tier shared-slack nonlinear program solved online with CasADi and IPOPT; the name denotes the weighted quadratic slack penalty, not a sequential-quadratic-programming solver. Strongly separated tier penalties bias residual violations toward lower-priority rules while leaving actuation bounds hard. The controller replans from its executed state at $10$\\,Hz and records per-rule residuals on every cycle. A $90$\\,ms solver-time limit returns an anytime iterate that is projected through the vehicle dynamics before execution; median and maximum observed wall-clock solve times were $28$ and $104$\\,ms. We evaluate W-SQP in closed loop on 150 Waymo Open Motion Dataset scenarios in Waymax against reactive and proposal-and-select baselines, and introduce a log-independent protocol that separates safety and regulatory compliance from resemblance to the recorded human trajectory. Under this protocol, W-SQP shows no systematic group-level deficit relative to expert replay on the log-independent safety and regulatory rules, with several localized regressions in the hardest, highest-divergence scenarios. The results characterize W-SQP as an auditable, priority-biased, anytime-capable NMPC prototype rather than a hard-real-time or formally safe controller.",
      "short_abstract": "Autonomous-vehicle motion planners must resolve conflicts among safety, regulation, comfort, and efficiency in real time while exposing those decisions for audit. We present W-SQP, a weighted tiered-slack nonlinear model predictive controller (NMPC) that...",
      "published": "2026-07-12",
      "updated": "2026-07-12",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 11,
        "evidence": [
          "title: model predictive control",
          "abstract: vehicle dynamics",
          "abstract: controller"
        ]
      },
      "recency": "fresh",
      "age_days": 38,
      "arxiv_url": "https://arxiv.org/abs/2607.10975",
      "pdf_url": "https://arxiv.org/pdf/2607.10975.pdf"
    },
    {
      "id": "2607.10942",
      "anchor": "paper-2607-10942",
      "title": "Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles",
      "authors": [
        "Ashiyana Abdul Majeed",
        "Mahmoud Meribout",
        "Neethu Joseph",
        "Abel Kidane Haile",
        "Mohammad Abdullah Al Faruque"
      ],
      "abstract": "Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints. However, heterogeneous edge-GPU deployment remains limited by underutilized hardware engines and accelerator-incompatible operators, causing fragmented execution and lower throughput per watt. This paper presents Heterogeneous Frame Dispatch Scheduling (H-FraDS), a hardware-aware frame scheduling methodology for transformer inference on a recent NVIDIA edge GPU. H-FraDS routes frames across the GPU and dual deep learning accelerator (DLA) cores using fixed dispatch ratios to improve utilization under latency and power constraints. To enable scheduling, incompatible transformer components are adapted for DLA execution by reshaping tensors, approximating error function (ERF) with tanh, and replacing layer normalization with bounded tanh. The adapted model maintains a 92% F1 score, with only a 2% reduction from the original. Optical flow accelerator (OFA) is further used for inference-side optical-flow estimation. To the best of the authors' knowledge, prior work has not addressed these combined issues. Using Swin Transformer for autonomous-driving perception, H-FraDS Balanced Dispatch (1:2) achieves 125.93 FPS, a 2.36x speedup over standalone adapted-DLA execution, 4.0 FPS/W, and approximately 24 ms DLA latency, satisfying 30 FPS real-time operation; the GPU-DLA-OFA case achieves a 2.02x DLA throughput speedup.",
      "short_abstract": "Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints. However, heterogeneous edge-GPU deployment remains limited by underutilized...",
      "published": "2026-07-12",
      "updated": "2026-07-12",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 14,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: hardware",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 38,
      "arxiv_url": "https://arxiv.org/abs/2607.10942",
      "pdf_url": "https://arxiv.org/pdf/2607.10942.pdf"
    },
    {
      "id": "2607.10565",
      "anchor": "paper-2607-10565",
      "title": "BucketKD: A Safety-Aware Bucket-Based Knowledge Distillation Framework for End-to-End Motion Planning",
      "authors": [
        "Md Nahidul Islam",
        "Mohd Hasan Ali",
        "Dipankar Dasgupta",
        "Myounggyu Won"
      ],
      "abstract": "End-to-end motion planning has emerged as a promising paradigm in autonomous driving, directly mapping raw sensor data to control commands via deep neural networks. Despite its advantages, its large model size hinders deployment in resource-constrained platforms. In this paper, we present BucketKD, a bucket-based knowledge distillation framework that yields compact and safety-aware end-to-end planners. Compared to the state-of-the-art approach, which relies on simplified planning state representations, BucketKD discretizes critical environmental variables into adaptive buckets that capture richer scene semantics while preserving efficiency. In addition, we design a safety-aware waypoint attention mechanism that evaluates each waypoint's risk level by accounting for both obstacle proximity and relative motion through a time-to-collision (TTC) formulation widely used in transportation research. This enables the student model to better retain safety-critical behaviors during distillation. Extensive experiments in CARLA using the Bench2Drive dataset show that BucketKD significantly outperforms the state-of-the-art in both planning accuracy and safety while maintaining strong compression ratios.",
      "short_abstract": "End-to-end motion planning has emerged as a promising paradigm in autonomous driving, directly mapping raw sensor data to control commands via deep neural networks. Despite its advantages, its large model size hinders deployment in resource-constrained...",
      "published": "2026-07-12",
      "updated": "2026-07-12",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 16,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 38,
      "arxiv_url": "https://arxiv.org/abs/2607.10565",
      "pdf_url": "https://arxiv.org/pdf/2607.10565.pdf"
    },
    {
      "id": "2607.10712",
      "anchor": "paper-2607-10712",
      "title": "Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud",
      "authors": [
        "Bálint Gyevnár",
        "Atoosa Kasirzadeh",
        "Nihar B. Shah"
      ],
      "abstract": "Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required the resources of a company: deep pockets, ghostwritten articles, and corrupt academics. Today, Artificial Intelligence (AI) is increasingly automating scientific research, so we ask: Can a remote adversary weaponize the honest use of AI in science to compromise scientific integrity? We envision and empirically evaluate a new attack, indirect data poisoning, in which an adversary corrupts an open dataset and uploads the poisoned variant to a public repository. Autonomous research agents may independently retrieve and process this data, turning honest scientists into the unpaid and unwitting distributors of fraud at scale. Across five socially-salient topics, from hiring discrimination to the safety of autonomous vehicles, three widely used frontier AI systems (Claude Code with Claude Opus 4.7, Codex with GPT-5.5, Gemini CLI with Gemini 3.1 Pro), and 450 ethically contained experimental runs, we find that poisoning succeeds in 49.56% of runs, while the rate of poisoning detection is only 6.0%. The attack requires no topic-specific trigger-words, agent access, indirect prompt injection, or fabricated papers, only the open data ecosystem and misleading metadata. To mitigate the attacks, we propose and evaluate two measures: a scientist persona and a data provenance audit with five checks (referencing papers, social markers, statistical anomalies, related datasets, poisoning caution). We find that the persona still leaves 16.67% of runs with a poisoned conclusion, but provenance auditing reduces attack success rate to zero. Our results suggest that indirect data poisoning may enable scientific fraud at unprecedented scale, but these attacks can be mitigated with suitable auditing by agents during data retrieval.",
      "short_abstract": "Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required the resources of a company: deep pockets, ghostwritten articles, and corrupt academics. Today, Artificial Intelligence...",
      "published": "2026-07-12",
      "updated": "2026-07-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 38,
      "arxiv_url": "https://arxiv.org/abs/2607.10712",
      "pdf_url": "https://arxiv.org/pdf/2607.10712.pdf"
    },
    {
      "id": "2607.10630",
      "anchor": "paper-2607-10630",
      "title": "World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning",
      "authors": [
        "Tong Nie",
        "Yuewen Mei",
        "Junlin He",
        "Yihong Tang",
        "Jian Sun",
        "Wei Ma"
      ],
      "abstract": "Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods often rely on external scenario generators, heuristic perturbations, or simulator-heavy rollouts, which makes them difficult to integrate with modern autoregressive planners. Here, we cast adversarially robust planner learning as a constrained min-max game and propose Adversarial World Modeling (AWM), a theoretically grounded multi-agent self-play fine-tuning framework. Since solving the exact game is intractable, AWM introduces a principled decoupled solver. In the inner minimization, the planner's predictive world model is converted into a role-conditioned adversary that learns sparse, scene-adaptive attack coalitions via counterfactual credit assignment. In the outer maximization, the ego planner optimizes a regret-aware robust best response against the frozen AWM, utilizing tail-risk weighting and reference-anchored trust regions to improve hard-case recovery while preserving nominal driving behavior. Experiments on the nuPlan and InterPlan benchmarks demonstrate that our method generates transferable adversarial interactions and yields a robust planner that achieves competitive closed-loop performance in both nominal and highly interactive long-tail scenarios. Theoretical analysis justifies the decoupled solver and the main optimization components.",
      "short_abstract": "Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods often...",
      "published": "2026-07-12",
      "updated": "2026-07-12",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 16,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 38,
      "arxiv_url": "https://arxiv.org/abs/2607.10630",
      "pdf_url": "https://arxiv.org/pdf/2607.10630.pdf"
    },
    {
      "id": "2607.10945",
      "anchor": "paper-2607-10945",
      "title": "When Context Dominates: Multimodal Signatures of Takeover Readiness Under Varying Hazard and Cognitive Load Conditions",
      "authors": [
        "Shiva Azimi",
        "Yasaman Hakiminejad",
        "Luis Gomero",
        "Elizabeth Pantesco",
        "Irene P. Kan",
        "Meltem Izzetoglu",
        "Arash Tavakoli"
      ],
      "abstract": "Semi-automated driving systems promise to reduce crashes by assisting with perception and control, yet they simultaneously introduce additional human factors challenges by requiring drivers to monitor automation and rapidly resume control when failures occur. Prolonged passive monitoring can degrade vigilance, delay reactions, and increase takeover risk, but the extent to which distraction, hazard context, and drivers' underlying cognitive and physiological states jointly shape takeover performance remains insufficiently understood. This study investigates these interacting factors using a controlled, within-subjects driving simulator experiment that crosses two hazard types (dynamic pedestrian and static crash events) with three levels of secondary task engagement (no task, conversation, and working memory load). Driver responses were assessed using a multimodal sensing framework that integrates vehicle-dynamics measures, subjective workload ratings, autonomic physiology (electrodermal activity and heart rate variability), and prefrontal cortical activation measured with functional near-infrared spectroscopy. Results show that hazard context is the primary determinant of takeover behavior, with pedestrian events producing longer and more variable maneuvers and crash events yielding faster and more stable responses. Secondary tasks exerted smaller effects on objective vehicle control, while internal-state measures showed more variable task-related patterns. These findings highlight the importance of jointly considering environmental context and human state when evaluating takeover readiness and designing driver monitoring systems. This study lays the groundwork for adaptive, context-aware strategies that support safer human-automation collaboration in semi-automated vehicles.",
      "short_abstract": "Semi-automated driving systems promise to reduce crashes by assisting with perception and control, yet they simultaneously introduce additional human factors challenges by requiring drivers to monitor automation and rapidly resume control when failures occur....",
      "published": "2026-07-12",
      "updated": "2026-07-12",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 25,
        "evidence": [
          "abstract: human factors",
          "abstract: driver monitoring",
          "title: takeover",
          "abstract: takeover"
        ]
      },
      "recency": "fresh",
      "age_days": 38,
      "arxiv_url": "https://arxiv.org/abs/2607.10945",
      "pdf_url": "https://arxiv.org/pdf/2607.10945.pdf"
    },
    {
      "id": "2607.10438",
      "anchor": "paper-2607-10438",
      "title": "Large Language Model Enhanced Differentiable Trajectory Planning for IoT-Enabled Autonomous Driving",
      "authors": [
        "Shihao Zhang",
        "Jing Yang",
        "Ziyu Song",
        "Zheng Lin",
        "Sunil Prajapat",
        "Zhaochen Xia",
        "Hemant Ghayvat",
        "Haitao Ding",
        "Lip Yee Por",
        "Ashok Kumar Das"
      ],
      "abstract": "Autonomous driving planning is a key component of IoT-enabled intelligent transportation systems, requiring vehicles to generate safe, efficient, and executable trajectories in complex urban environments from multi-source contextual information. While imitation learning (IL) has shown promise on large-scale datasets, IL-based planners still suffer from limited coverage of complex long-tail interactions, weak consistency with downstream constrained refinement, and insufficient use of high level scene semantics under real time constraints. To address these issues, this paper proposes a large language model (LLM) enhanced differentiable trajectory planning framework for IoT-enabled autonomous driving. Specifically, we introduce a surrounding agent centric data augmentation strategy to reorganize sur rounding agent trajectories as additional planning supervision, thereby improving the training distribution without collecting additional raw data. We further design a complexity-aware asyn chronous LLM-based semantic enhancement module to extract scene-related high-level semantic features with controlled online overhead. In addition, a differentiable optimization module is incorporated to refine generated trajectories with explicit residual penalties while backpropagating optimization gradients to the upstream planner. Experiments show that the proposed method achieves the best overall scores of 83.63 and 78.29 on the nuPlan closed-loop nonreactive and reactive Hard20 benchmarks, respectively, and CARLA-ROS tests further verify its online deployment and real time closed-loop execution capability.",
      "short_abstract": "Autonomous driving planning is a key component of IoT-enabled intelligent transportation systems, requiring vehicles to generate safe, efficient, and executable trajectories in complex urban environments from multi-source contextual information. While...",
      "published": "2026-07-11",
      "updated": "2026-07-11",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 29,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 39,
      "arxiv_url": "https://arxiv.org/abs/2607.10438",
      "pdf_url": "https://arxiv.org/pdf/2607.10438.pdf"
    },
    {
      "id": "2607.10336",
      "anchor": "paper-2607-10336",
      "title": "PrismAD: Decoupled Planning via Semantic Mixture-of-Planners for End-to-End Autonomous Driving",
      "authors": [
        "Kang Ding",
        "Zhigui Lin",
        "Hongsong Wang",
        "Jie Gui",
        "Qi Liu",
        "Zhe Wang",
        "Luqi Tang",
        "Lei He"
      ],
      "abstract": "This letter presents PrismAD, a decoupled end-to-end autonomous driving framework based on a Semantic Mixture-of-Planners. Existing planners usually aggregate heterogeneous scene tokens into a coupled representation space, forcing a single planning branch to jointly model agent interaction, road geometry, and driving intention. Such coupling may weaken factor-specific reasoning and obscure the contribution of different planning cues. To address this limitation, PrismAD partitions scene tokens into interaction, geometry, and intent groups, and assigns them to independent planning experts with the same architecture but separate parameters. Each expert learns a specialized motion-planning representation, while a semantics-aware router adaptively aggregates expert predictions with separate routing weights for motion prediction and ego planning. Sparse top-$K$ activation with noisy gating is further introduced to improve routing robustness and reduce unnecessary expert computation. Extensive experiments on the nuScenes open-loop dataset and NeuroNCAP closed-loop benchmark demonstrate that PrismAD exhibits competitive performance. Our code will be released soon.",
      "short_abstract": "This letter presents PrismAD, a decoupled end-to-end autonomous driving framework based on a Semantic Mixture-of-Planners. Existing planners usually aggregate heterogeneous scene tokens into a coupled representation space, forcing a single planning branch to...",
      "published": "2026-07-11",
      "updated": "2026-07-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 39,
      "arxiv_url": "https://arxiv.org/abs/2607.10336",
      "pdf_url": "https://arxiv.org/pdf/2607.10336.pdf"
    },
    {
      "id": "2607.10138",
      "anchor": "paper-2607-10138",
      "title": "Channel Knowledge Empowered Finite-Blocklength Rate-Splitting Transmission for High-Mobility Autonomous Driving",
      "authors": [
        "Yi Wang",
        "Yingyang Chen",
        "Feng Bai",
        "Li Wang",
        "Gang Feng"
      ],
      "abstract": "To meet the extended ultra-low latency and high reliability (xURLLC) requirements for autonomous driving systems, multiple access schemes must operate reliably in high-mobility and complex propagation environments. Recently, rate-splitting multiple access (RSMA) has emerged as a promising multi-user transmission framework, showing robustness in dynamic situations where imperfect and outdated channel state information (CSI) is prevalent.Moreover, the advanced sensing, localization, and on-board computation capabilities of autonomous driving vehicles facilitate the construction of a channel knowledge map (CKM), which is a key enabler for environment-aware communications in future 6G networks.In this work, we propose a CKM empowered finite-blocklength (FBL) RSMA for downlink autonomous driving system. The location-dependent large-scale channel information provided by CKM is exploited in RSMA to develop a refined rate-splitting design. The min-rate performance of FBL rate splitting is analyzed explicitly to ensure user fairness. We derive a new and tight closed-form bound for the private-stream ergodic rate. Combined with the closed-form common-stream expression, an efficient optimization design of rate-splitting ratios has been formulated. Numerical results show that the CKM empowered FBL RSMA outperforms space-division multiple access (SDMA) and non-orthogonal multiple access (NOMA), particularly in high-mobility scenarios. Its performance is improved by a data-based CKM, which provides more accurate large-scale channel information than model-based approaches and enables more precise common-stream allocation. The results also reveal that RSMA is sensitive to errors in large-scale channel knowledge, emphasizing the importance of accurate CKM information for optimal rate-splitting.",
      "short_abstract": "To meet the extended ultra-low latency and high reliability (xURLLC) requirements for autonomous driving systems, multiple access schemes must operate reliably in high-mobility and complex propagation environments. Recently, rate-splitting multiple access...",
      "published": "2026-07-11",
      "updated": "2026-07-11",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: localization",
          "Systems, Deployment & Connectivity: abstract: communications",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 39,
      "arxiv_url": "https://arxiv.org/abs/2607.10138",
      "pdf_url": "https://arxiv.org/pdf/2607.10138.pdf"
    },
    {
      "id": "2607.10394",
      "anchor": "paper-2607-10394",
      "title": "CSI-Assisted Edge SLAM Testbed Platform for 5G Connected Unmanned Autonomous Vehicles",
      "authors": [
        "Boris Radovanovic",
        "Sasa Talosi",
        "Srdjan Sobot",
        "Dejan Vukobratovic"
      ],
      "abstract": "The evolution from 5G towards 6G reinforces interest in connected robotics, where mobile robots offload compute-intensive tasks to edge servers over ultra-reliable low-latency communication (URLLC) links. Simultaneous localization and mapping (SLAM), a fundamental yet demanding robotics function, is increasingly considered for edge deployment within mobile edge computing (MEC) frameworks. In parallel, integrated sensing and communications (ISAC) enables the use of radio channel information, such as channel state information (CSI), as an additional sensing modality in radio-based SLAM. In this paper, we design and implement a CSI-assisted Edge SLAM testbed integrating a custom unmanned ground vehicle (UGV), a ROS2-based SLAM framework, and a 5G Open Radio Access Network (O-RAN) system. The proposed architecture provides an end-to-end, cross-layer view of ROS2 sensor data streaming over 5G, explicitly enabling CSI exposure and integration into the SLAM pipeline. We analyze ROS2 DDS communication, RTPS packetization, and 5G user-plane transport, and discuss mechanisms for CSI extraction and delivery via O-RAN components. The platform enables realistic experimentation with communication-aware SLAM and reveals key challenges related to latency, data streaming, synchronization, and cross-system integration, providing insights for future 6G-enabled robotic platforms.",
      "short_abstract": "The evolution from 5G towards 6G reinforces interest in connected robotics, where mobile robots offload compute-intensive tasks to edge servers over ultra-reliable low-latency communication (URLLC) links. Simultaneous localization and mapping (SLAM), a...",
      "published": "2026-07-11",
      "updated": "2026-07-11",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 22,
        "evidence": [
          "abstract: localization",
          "title: SLAM",
          "abstract: SLAM",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 39,
      "arxiv_url": "https://arxiv.org/abs/2607.10394",
      "pdf_url": "https://arxiv.org/pdf/2607.10394.pdf"
    },
    {
      "id": "2607.10071",
      "anchor": "paper-2607-10071",
      "title": "FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness",
      "authors": [
        "Shunsuke Yokokawa",
        "Hironori Kasahara"
      ],
      "abstract": "Bird's-eye-view (BEV) perception is a core component of camera-based 3D understanding in autonomous driving, where view transformation (VT) maps multi-camera image features into a unified BEV representation. Sampling-based view transformation (Sampling-VT) is attractive because it supports dense and continuous BEV aggregation for high-resolution and long-range perception. Its deployment bottleneck, however, is systems-level: standard tensorized implementations of Sampling-VT -- which we refer to as Tensorized Sampling-VT -- explicitly materialize large height-dependent intermediate tensors, causing memory and latency costs that scale poorly with vertical resolution and the number of cameras. We revisit Tensorized Sampling-VT from an operator-execution perspective and show that it follows a gather-reduction pattern: each BEV query independently accumulates contributions across cameras and height bins, enabling thread-local accumulation with on-the-fly recomputation that eliminates the need to materialize height- and camera-dependent intermediates. Based on this insight, we propose FlashBEV, a fully fused and IO-aware execution strategy mathematically equivalent to Tensorized Sampling-VT (same operator output) while substantially reducing global memory traffic and kernel-launch overhead. Experiments show that FlashBEV achieves more than an order of magnitude lower peak GPU memory and significant inference-latency speedups, with memory effectively independent of the number of height bins, reducing the operator's peak memory to O(BCXY) (output only). This unlocks higher BEV range/resolution and vertical discretization within fixed deployment budgets on memory-constrained devices. Our contribution is an execution redesign -- same math, different execution -- that removes a key scalability barrier for deployment-ready Sampling-VT. Code available at https://github.com/yokosyun/FlashBEV",
      "short_abstract": "Bird's-eye-view (BEV) perception is a core component of camera-based 3D understanding in autonomous driving, where view transformation (VT) maps multi-camera image features into a unified BEV representation. Sampling-based view transformation (Sampling-VT) is...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 8,
        "evidence": [
          "abstract: perception",
          "abstract: camera",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.10071",
      "pdf_url": "https://arxiv.org/pdf/2607.10071.pdf"
    },
    {
      "id": "2607.10037",
      "anchor": "paper-2607-10037",
      "title": "Plug-and-Play Reweighting for Resilient Collaborative Decision-Making in Connected Autonomous Driving",
      "authors": [
        "Jiewen Liu",
        "Rui Liu",
        "Matthew Lee",
        "Ming C. Lin",
        "Xiaorui Liu",
        "Peng Gao"
      ],
      "abstract": "Collaborative decision-making is a fundamental capability in multi-robot systems, such as connected autonomous vehicles. However, perceptual noise and adversarial attacks in collaborators can severely affect decision reliability. Overall, existing methods typically rely on retraining with attack-specific defenses or on restrictive perturbation assumptions to improve resilience, which limits their practicality. In this paper, we propose a novel Resilient Collaborative Decision-Making (RCDM) framework that consists of an attention-based encoder for extracting individual robot perceptual embeddings and an attention-based decoder for fusing collaborator perceptions and making decisions. To improve resilience to corrupted observations, we design a novel plug-and-play reweighting module that down-weights the influence of corrupted inputs by analyzing the consistency of neighborhood points relative to the local structure and assigning smaller weights to points that deviate strongly from the local median. This module can be seamlessly integrated into attention-based collaborative decision-making without requiring additional training. We evaluate our method in high-fidelity simulations, considering perceptual noise and five types of attacks across diverse accident-prone scenarios. Experimental results demonstrate that our approach consistently outperforms existing methods by up to 26% and achieves state-of-the-art resilient performance.",
      "short_abstract": "Collaborative decision-making is a fundamental capability in multi-robot systems, such as connected autonomous vehicles. However, perceptual noise and adversarial attacks in collaborators can severely affect decision reliability. Overall, existing methods...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: title: decision-making",
          "Planning & Decision-Making: abstract: decision-making",
          "Systems, Deployment & Connectivity: title: cooperative systems",
          "Systems, Deployment & Connectivity: abstract: cooperative systems",
          "Systems, Deployment & Connectivity: abstract: connected vehicles"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.10037",
      "pdf_url": "https://arxiv.org/pdf/2607.10037.pdf"
    },
    {
      "id": "2607.09655",
      "anchor": "paper-2607-09655",
      "title": "OpenLongTail: Generative Scaling of Long-Tail Driving Data",
      "authors": [
        "Lulin Liu",
        "Nuo Chen",
        "Yan Wang",
        "Bangya Liu",
        "Wenyan Cong",
        "Hezhen Hu",
        "Boris Ivanovic",
        "Hao Wang",
        "Ziyao Zeng",
        "Xinyu Gong",
        "Yang Zhou",
        "Zixiang Xiong",
        "Dilin Wang",
        "Zhangyang Wang",
        "Weisong Shi",
        "Ruohan Zhang",
        "Marco Pavone",
        "Zhiwen Fan"
      ],
      "abstract": "Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.",
      "short_abstract": "Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Perception & Sensor Fusion: abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.09655",
      "pdf_url": "https://arxiv.org/pdf/2607.09655.pdf"
    },
    {
      "id": "2607.09629",
      "anchor": "paper-2607-09629",
      "title": "4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception",
      "authors": [
        "Xiaokai Bai",
        "Lianqing Zheng",
        "Runwei Guan",
        "Songkai Wang",
        "Siyuan Cao",
        "Hui-liang Shen"
      ],
      "abstract": "Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary for comprehensive scene understanding. Existing radar-camera methods mainly optimize detection, while dual-task systems usually decode boxes and occupancy with limited interaction. To address this gap and advance radar-based multi-task learning, we propose \\method, a 4D radar-camera framework for 360$^\\circ$ full-scene perception, which models semantic occupancy as a persistent scene state rather than a terminal output. \\method{} follows a cross-modal state reasoning paradigm, where the occupancy state is modeled and propagated through stages for coarse-to-fine feature aggregation. Specifically, State-guided BEV Enhancement (SBE) strengthens intra-frame BEV representation, while Doppler-guided Temporal Fusion (DTF) preserves state evidence over longer temporal horizons. Beyond the model, we further extend ManTruckScenes with satellite-map-based generated occupancy labels and pair it with OmniHD-Scenes in a unified cross-dataset detection-and-occupancy protocol. The resulting experiments cover accuracy, robustness, ablation, and efficiency under one radar-camera multi-task evaluation framework. Code and labels will be released upon acceptance.",
      "short_abstract": "Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 45,
        "margin": 39,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "title: occupancy",
          "abstract: occupancy",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.09629",
      "pdf_url": "https://arxiv.org/pdf/2607.09629.pdf"
    },
    {
      "id": "2607.09428",
      "anchor": "paper-2607-09428",
      "title": "Multimodal Scenario Similarity Search for Autonomous Driving",
      "authors": [
        "Tamás Matuszka",
        "András Tamásy",
        "Balázs Szolár"
      ],
      "abstract": "Large-scale autonomous-driving datasets contain vast numbers of recorded scenarios, creating a need for efficient retrieval methods that can identify situations similar to a given query. Existing approaches typically rely on either visual representations or motion-based descriptions, making it difficult to understand their relative strengths and limitations for scenario retrieval. In this work, we present a multimodal framework for autonomous-driving scenario retrieval that combines visual and trajectory-based representations within a unified retrieval pipeline. We investigate two trajectory-based approaches: Exo-Trajectory, an explicit matching method based on surrounding-agent motion, and ScenarioFormer, a transformer-based representation learned from object trajectories using contrastive learning. We compare these approaches against strong vision-based baselines and analyze their behavior across a diverse set of driving scenarios. Experimental results show that trajectory representations provide strong retrieval performance for motion-centric events such as cut-ins, turning maneuvers, and traffic queueing, while visual embeddings excel when appearance cues are informative. Most importantly, combining visual and trajectory information consistently improves retrieval quality, yielding the best overall performance. These findings demonstrate that appearance and motion capture are complementary notions of scenario similarity and motivate multimodal retrieval systems for autonomous-driving data mining, dataset curation, and scenario-based validation.",
      "short_abstract": "Large-scale autonomous-driving datasets contain vast numbers of recorded scenarios, creating a need for efficient retrieval methods that can identify situations similar to a given query. Existing approaches typically rely on either visual representations or...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.09428",
      "pdf_url": "https://arxiv.org/pdf/2607.09428.pdf"
    },
    {
      "id": "2607.09138",
      "anchor": "paper-2607-09138",
      "title": "BeyondSight: Object Permanence for End-to-End Autonomous Driving",
      "authors": [
        "Sandro Papais",
        "Letian Wang",
        "Mudit Jain",
        "Behnaz Rezaei",
        "Steven L. Waslander"
      ],
      "abstract": "Autonomous driving operates in partially observable environments where actors may become fully occluded by other vehicles or infrastructure. Most end-to-end driving systems implicitly couple actor existence to instantaneous observations, causing actor hypotheses to degrade or disappear during prolonged occlusion and removing potentially critical agents from downstream prediction and planning. We introduce BeyondSight, a permanence-aware end-to-end driving framework that decouples actor existence from observability by maintaining persistent actor hypotheses over time. BeyondSight propagates actor queries temporally and updates them with observation-conditioned evidence, enabling joint perception, prediction, and planning to reason about actors even when they are temporarily unobservable. To enable principled training and evaluation of persistence-aware models, we further introduce nuScenes-Permanence, an extension of nuScenes that provides supervision and observability-conditioned evaluation for unobservable actors. Experiments show that BeyondSight substantially improves reasoning under occlusion, increasing detection performance for unobservable actors from 0 to 0.249 mAP while reducing planning error from 0.61 to 0.54 L2avg. These results highlight object permanence as an important modeling principle for robust end-to-end autonomous driving.",
      "short_abstract": "Autonomous driving operates in partially observable environments where actors may become fully occluded by other vehicles or infrastructure. Most end-to-end driving systems implicitly couple actor existence to instantaneous observations, causing actor...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 33,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.09138",
      "pdf_url": "https://arxiv.org/pdf/2607.09138.pdf"
    },
    {
      "id": "2607.09903",
      "anchor": "paper-2607-09903",
      "title": "The Precursor Genome: A Pairwise Reaction Dataset for Solid-State Synthesis",
      "authors": [
        "Lauren N. Walters",
        "Matthew J. McDermott",
        "Bernardus Rendy",
        "Yuxing Fei",
        "Kristin A. Persson",
        "Gerbrand Ceder"
      ],
      "abstract": "Solid-state reactions remain the dominant route to inorganic materials, yet no large, machine-readable dataset reports their experimental protocols and outcomes with consistent provenance; this gap obstructs first-principles, data-driven, and machine-learning approaches to synthesis science. Here, we present the Precursor Genome, a dataset of 1,035 pairwise solid-state reactions generated autonomously by the A-Lab self-driving laboratory, spanning 46 precursors and 39 elements. Every reaction is reported together with its full experimental metadata, including measured thermal profiles, precursor and recovered masses, and instrument configuration. Every product mixture is identified from raw X-ray diffraction (1,351 scans) through automated Rietveld refinement with the Dara framework, yielding 1,950 refinement cases that are independently validated by human experts on a three-tier quality scale. Raw pattern files, serialized refinement objects, and reviewer annotations are distributed through a Pydantic-validated JSON ledger, preserving full traceability from each precursor pair to its final phase assignment. The Precursor Genome establishes a FAIR, reusable benchmark for training and evaluating predictive models of solid-state reactivity.",
      "short_abstract": "Solid-state reactions remain the dominant route to inorganic materials, yet no large, machine-readable dataset reports their experimental protocols and outcomes with consistent provenance; this gap obstructs first-principles, data-driven, and machine-learning...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.09903",
      "pdf_url": "https://arxiv.org/pdf/2607.09903.pdf"
    },
    {
      "id": "2607.10061",
      "anchor": "paper-2607-10061",
      "title": "Federated Cybersecurity Testbed as a Service (FCTaaS): A framework to federate cybersecurity testbeds",
      "authors": [
        "Josh Dean",
        "Yu-Zheng Lin",
        "John Paul Martin Encinas",
        "Ibrahim Almazyad",
        "Qinxuan Shi",
        "Zhanglong Yang",
        "Shalaka Satam",
        "Tingjun Lei",
        "Jielun Zhang",
        "Sicong Shao",
        "Salim Hariri",
        "Pratik Satam"
      ],
      "abstract": "Rapid technological change is reshaping society through emerging domains such as autonomous vehicles and smart manufacturing, creating new research challenges in system design, operation, security, and training. Researchers often rely on testbeds to reproduce experimental scenarios, collect and analyze data, observe system behavior, and evaluate proposed solutions. However, the fast pace of innovation makes it difficult and costly for individual testbeds to remain representative of state-of-the-art systems, as doing so requires frequent upgrades and new capabilities. Moreover, access to specialized testbeds is often limited to a small group of researchers, leaving valuable infrastructure underutilized during its operational lifetime. This paper presents FCTaaS, a Federated Cybersecurity Testbed as a Service framework that enables heterogeneous cybersecurity testbeds to participate in a single experiment across geographical boundaries. By connecting independently managed testbeds through a Virtual Private Network (VPN), FCTaaS supports remote testbed discovery, experiment design, and coordinated experimentation. We evaluate FCTaaS across three case studies involving denial-of-service scenarios on smart infrastructure and intrusion detection and prevention workflows using a Suricata-based IDS/IPS testbed. The results show that FCTaaS enables effective cross-testbed experimentation while preserving visibility into attack traffic, IDS alerts, and detection-system resource stress. Even under resource-intensive attack scenarios, FCTaaS achieves limiting network utilization of 49%, introduces only 1% overhead, and supports latency ranging from 5.63 ms between local nodes to 147 ms between geographically dispersed nodes.",
      "short_abstract": "Rapid technological change is reshaping society through emerging domains such as autonomous vehicles and smart manufacturing, creating new research challenges in system design, operation, security, and training. Researchers often rely on testbeds to reproduce...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 20,
        "evidence": [
          "title: security",
          "abstract: security",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.10061",
      "pdf_url": "https://arxiv.org/pdf/2607.10061.pdf"
    },
    {
      "id": "2607.09103",
      "anchor": "paper-2607-09103",
      "title": "Equivariant Filter for High Performance Image Tracking using an Event Camera",
      "authors": [
        "Angus Apps",
        "Yixiao Ge",
        "Timothy L. Molloy",
        "Robert Mahony"
      ],
      "abstract": "Image tracking is the problem of estimating the transformation that relates a moving image of a scene to an original reference image. The problem is important in control of autonomous vehicles or robots, where the image encodes information about the motion of the camera or environment, as well as in pure computer vision applications. In this paper, we present an equivariant filter design for high performance tracking of planar image transformations using an event camera. The design exploits the Asynchronous Event Blob (AEB) tracker (Wang et al., 2024) to extract feature-position measurements from the raw event stream, and an equivariant filter to compute an affine image translation and rotation using the special Euclidean group symmetry. The equivariant filter incorporates an equivalent-measurement update step that de-correlates the (highly temporally correlated) feature-position measurements provided by the AEB tracker. We evaluate the design experimentally using two datasets involving general and fast rotational motion. We benchmark results against direct optimisation (estimating the relative transformation from the raw blob tracks), and a covariance intersection approach for overcoming data correlation. Our design provides smooth image tracking for features moving up to 7000 pixels per second on the image plane.",
      "short_abstract": "Image tracking is the problem of estimating the transformation that relates a moving image of a scene to an original reference image. The problem is important in control of autonomous vehicles or robots, where the image encodes information about the motion of...",
      "published": "2026-07-10",
      "updated": "2026-07-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 6,
        "evidence": [
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 40,
      "arxiv_url": "https://arxiv.org/abs/2607.09103",
      "pdf_url": "https://arxiv.org/pdf/2607.09103.pdf"
    },
    {
      "id": "2607.09045",
      "anchor": "paper-2607-09045",
      "title": "Can the Cloud Drive? Infrastructure Feasibility of Offloading Autonomous Driving Across 5G and 6G",
      "authors": [
        "Pouya Parsa",
        "Kawon Han",
        "Seongjin Choi"
      ],
      "abstract": "Frontier autonomous-driving models -- especially vision-language-action (VLA) models, whose forward pass approaches $\\sim$60~TFLOPs -- are outgrowing economical onboard deployment, since peak hardware sits idle most of the day. Cloud inference can instead share GPUs across active vehicles, but the vehicle must upload through a capacity-limited uplink, reach a GPU without queueing, and return a decision within the closed-loop budget. This paper asks: can the cloud drive? We answer with an analytical framework coupling communication limits, a roofline GPU service model, stochastic latency, and utilization-aware cost across three model classes, three offloading strategies, and three communication generations, applied to New York City. Separating a reactive 100~ms budget from a 300~ms deliberative tier (presuming an onboard reactive fallback), we find three \\emph{nested} binding regimes. Communication binds first in dense cells: 5G fails early, 5G-Advanced is the practical threshold for feature-level offloading, and 6G adds headroom. Compute binds next under the reactive budget: near-term VLA is latency-infeasible regardless of bandwidth, because autoregressive FP16 decode is memory-bandwidth-bound (~114 ms on 2025 hardware). Its floor clears 100 ms around 2027; 6G then admits feature-level VLA by ~2028, 5G-Advanced only at light loading and not the dense corridor, and the deliberative tier from 2026. Cost binds last: once admissible, utilization-pooled cloud GPUs undercut onboard hardware for VLA, whose baseline (up to \\$8,500 per vehicle-year) is expensive and idle; feature-level offloading (S2) is where the VLA cost crossover concentrates. Latency decides which model is admissible in which year; cost decides whether it is economical.",
      "short_abstract": "Frontier autonomous-driving models -- especially vision-language-action (VLA) models, whose forward pass approaches $\\sim$60~TFLOPs -- are outgrowing economical onboard deployment, since peak hardware sits idle most of the day. Cloud inference can instead...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 4,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.09045",
      "pdf_url": "https://arxiv.org/pdf/2607.09045.pdf"
    },
    {
      "id": "2607.08745",
      "anchor": "paper-2607-08745",
      "title": "AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding",
      "authors": [
        "Siddharth Damodharan",
        "Radhika Gupta",
        "Ali Alshami",
        "Ryan Rabinowitz",
        "Jugal Kalita"
      ],
      "abstract": "Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving incidents and near-incidents. The benchmark covers diverse safety-relevant categories, including weather and lighting conditions, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability-related reasoning. By requiring models to answer grounded questions about both contextual scene properties and event-level incident details, AUTOPILOT-VQA moves beyond object recognition toward temporally grounded, safety-aware reasoning. The dataset is released as part of the AUTOPILOT CVPR 2026 competition and provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios. Our benchmark support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.",
      "short_abstract": "Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However,...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Benchmark",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.08745",
      "pdf_url": "https://arxiv.org/pdf/2607.08745.pdf"
    },
    {
      "id": "2607.08391",
      "anchor": "paper-2607-08391",
      "title": "On Exploring Input Resolution Scaling For Anytime LiDAR Object Detection",
      "authors": [
        "Ahmet Soyyigit",
        "Shuochao Yao",
        "Heechul Yun"
      ],
      "abstract": "Making tradeoffs between execution latency and result utility (i.e., anytime computing) for adapting to dynamic operational requirements has been shown to enhance the performance of cyber-physical systems. In this work, we focus on enabling anytime computing for deep neural networks (DNNs) that process LiDAR point clouds for 3D object detection. We propose a novel method that enables multi-resolution inference for models that process point clouds as pillars or voxels, allowing the input to be dynamically scaled and processed at the resolution needed to meet timing requirements. Importantly, our memory-efficient approach requires the deployment of only a single DNN model, avoiding the need to deploy multiple models, each trained for a different input resolution. We also introduce a deadline-aware scheduler that selects the highest possible resolution for any given input by accurately predicting the execution time for all possible resolutions at runtime, which is challenging due to the irregularity of LiDAR point clouds. Experimental results on the nuScenes autonomous driving dataset demonstrate that our method significantly outperforms existing anytime computing approaches for LiDAR object detection. Finally, we deploy our approach in a simulated autonomous driving system, where it consistently enables collision-free navigation while avoiding unnecessary stalls caused by environmental complexity.",
      "short_abstract": "Making tradeoffs between execution latency and result utility (i.e., anytime computing) for adapting to dynamic operational requirements has been shown to enhance the performance of cyber-physical systems. In this work, we focus on enabling anytime computing...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 24,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.08391",
      "pdf_url": "https://arxiv.org/pdf/2607.08391.pdf"
    },
    {
      "id": "2607.08375",
      "anchor": "paper-2607-08375",
      "title": "WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving",
      "authors": [
        "Xuerun Yan",
        "Zhexi Lian",
        "Nuoheng Zhang",
        "Shiyu Fang",
        "Haoran Wang",
        "Chen Lv",
        "Jia Hu",
        "Binyang Song"
      ],
      "abstract": "Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world cognition or suffer from fragmented world foresight, inherently confining these models to reactive driving. To address this limitation, we propose WCog-VLA, a novel dual-level World-Cognitive VLA framework that successfully bridges semantic world forecasting with generative world evolution to achieve proactive autonomous driving. At the semantic level, WCog-VLA unifies world cognition and reasoning by incorporating 3D spatial perception and injecting agent tokens to capture the world dynamics, while concurrently enabling Game-theoretic Chain-of-Thought (Game-CoT) reasoning. At the generative level, we introduce the Aligned Decoupled Diffusion Transformer (ADDT) as a powerful generative world model that synthesizes physically-plausible joint multi-agent trajectories. Through scene representation alignment, ADDT reduces the number of denoising steps required and thus significantly accelerates inference. To facilitate strategic reasoning, we further construct a large-scale dataset featuring 85k Game-CoT annotations. Extensive experiments on the NAVSIM benchmark demonstrate that WCog-VLA achieves a State-Of-The-Art (SOTA) PDMS score of 92.9.",
      "short_abstract": "Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world cognition or suffer from fragmented world foresight, inherently confining these models to reactive driving. To address...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "VLA",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 60,
        "margin": 53,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.08375",
      "pdf_url": "https://arxiv.org/pdf/2607.08375.pdf"
    },
    {
      "id": "2607.08137",
      "anchor": "paper-2607-08137",
      "title": "Securing Autonomous Vehicle Systems via Twin-Aware Federated Reinforcement Learning",
      "authors": [
        "Zifan Zhang",
        "Minghong Fang",
        "Dianwei Chen",
        "Zhuqing Liu",
        "Prashant Khanduri",
        "Xianfeng Yang",
        "Anupam Das",
        "Yuchen Liu"
      ],
      "abstract": "Federated reinforcement learning (FRL) is crucial for enabling collaborative learning across multiple agents without sharing raw data, thereby enhancing privacy and scalability in the decision-making process within dynamic vehicular environments. However, poisoning attacks pose a significant threat to the security and reliability of FRL-based systems, particularly in safety-critical autonomous driving, where this vulnerability remains largely unexplored. These attacks can compromise the global control model by subtly injecting malicious system parameters, leading to potential hazards. To counter these challenges, we present \\alg (\\underline{Sec}ure \\underline{A}ggregation with \\underline{p}oisoning-\\underline{p}revention and historical reinforcement) as a defensive framework aimed at enhancing the robustness of FRL systems designed for safety-critical driving scenarios. \\alg strategically integrates digital twins for rehearsal-based learning and leverages historical aggregated model parameters along with a selected central gradient to ensure that only benign data is aggregated, effectively mitigating the influence of malicious agents. Theoretical guarantees are provided for the convergence performance of \\alg in the presence of poisoning attacks. We also validate the effectiveness of \\alg using developed digital twins that model realistic highway environments to evaluate the control of autonomous vehicles under adversarial conditions.",
      "short_abstract": "Federated reinforcement learning (FRL) is crucial for enabling collaborative learning across multiple agents without sharing raw data, thereby enhancing privacy and scalability in the decision-making process within dynamic vehicular environments. However,...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 11,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.08137",
      "pdf_url": "https://arxiv.org/pdf/2607.08137.pdf"
    },
    {
      "id": "2607.09795",
      "anchor": "paper-2607-09795",
      "title": "Large Multimodal Model-Based Environment-Aware Mobility Management",
      "authors": [
        "Seokhyun Jeong",
        "Sangmok Shin",
        "Seungnyun Kim",
        "Jiao Wu",
        "Byonghyo Shim"
      ],
      "abstract": "Recently, large language models (LLMs) have been successfully adopted in various fields, including wireless communications, robotics, and autonomous vehicles, owing to their outstanding adaptability and reasoning abilities. Despite their huge potential, the application of LLMs for mobility management is relatively scarce since it requires not only analyzing wireless measurements but also predicting dynamic user trajectories and making real-time handover decisions across densely deployed small base stations (SBSs). In this paper, we propose an environment-aware mobility management scheme based on large multimodal models (LMMs), which extend capabilities of LLMs to process multimodal sensing data. By leveraging LMMs, the proposed scheme extracts contextual information on the surrounding environments from RGB-D images to capture user equipment (UE) mobility patterns and identify signal reflections and blockages caused by static reflectors and dynamic obstacles. Using the extracted environmental information, the proposed scheme learns the intrinsic mapping from UE and SBS positions to channel capacity, referred to as channel capacity map (CCM), from which future channel capacities along UE trajectories are predicted. Based on the predicted channel capacities, we determine proactive handover decisions maximizing the cumulative channel capacities. Simulation results demonstrate that the proposed scheme achieves substantial channel capacity improvements over conventional deep learning (DL)-based approaches.",
      "short_abstract": "Recently, large language models (LLMs) have been successfully adopted in various fields, including wireless communications, robotics, and autonomous vehicles, owing to their outstanding adaptability and reasoning abilities. Despite their huge potential, the...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: mapping",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.09795",
      "pdf_url": "https://arxiv.org/pdf/2607.09795.pdf"
    },
    {
      "id": "2607.09070",
      "anchor": "paper-2607-09070",
      "title": "Latency-Aware Digital Twin-Assisted Cooperative Perception for Autonomous Vehicles",
      "authors": [
        "Boniface Uwizeyimana",
        "Manobendu Sarker",
        "Abraham O. Fapojuwo"
      ],
      "abstract": "This paper introduces a digital-twin (DT)-assisted cooperative perception framework designed to improve perception accuracy under end-to-end (E2E) latency constraints and to balance perception accuracy and E2E latency under communication resource constraints in autonomous vehicles. We formulate an optimization problem that maximizes perception accuracy subject to latency and communication limitations, and solve it using a newly proposed coarse-to-fine search (CTFS) algorithm. Simulation results show that the proposed CTFS algorithm achieves 96.6% perception accuracy, close to exhaustive search, under latency constraints while reducing computational complexity by approximately 85.78%. The DT-assisted framework further achieves a 50% reduction in the non-DT communication cost through estimated, time-synchronized state updates.",
      "short_abstract": "This paper introduces a digital-twin (DT)-assisted cooperative perception framework designed to improve perception accuracy under end-to-end (E2E) latency constraints and to balance perception accuracy and E2E latency under communication resource constraints...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 10,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.09070",
      "pdf_url": "https://arxiv.org/pdf/2607.09070.pdf"
    },
    {
      "id": "2607.08610",
      "anchor": "paper-2607-08610",
      "title": "Sharing economy in the era of full automation: Evidence from autonomous vehicle on-demand mobility services",
      "authors": [
        "Xiaoyan Wang",
        "Kenan Zhang",
        "Yaochen Ma"
      ],
      "abstract": "The digital age has facilitated the sharing of underutilized assets. This paper focuses on privately owned autonomous vehicles (AVs), a unique class of robots that can move independently and provide transportation services. When not in personal use, private AV owners can lease their vehicles to a platform that operates an on-demand mobility service (MoD). We refer to this service as AV crowdsourcing, and develop a time-expanded network flow model that captures temporal and spatial heterogeneity in AV usage of both owners and passengers while preserving analytical tractability. We analyze the conditions under which AV crowdsourcing reduces MoD operating costs and identify their key factors, namely, the complementarity of the mobility pattern between AV owners and MoD passengers, the slack time reserved by vehicle owners, and the vehicle repositioning distance. A case study of Chicago further reveals substantial spatiotemporal heterogeneity in optimal prices and service quality. The results demonstrate how centralized dispatching can simultaneously fulfill the high demand in downtown areas while maintaining relatively high service quality in peripheral regions. Our findings provide insights into how supply heterogeneity and market conditions jointly shape the performance of AV crowdsourcing systems that leverage the underutilized private robotic assets.",
      "short_abstract": "The digital age has facilitated the sharing of underutilized assets. This paper focuses on privately owned autonomous vehicles (AVs), a unique class of robots that can move independently and provide transportation services. When not in personal use, private...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.08610",
      "pdf_url": "https://arxiv.org/pdf/2607.08610.pdf"
    },
    {
      "id": "2607.08402",
      "anchor": "paper-2607-08402",
      "title": "Swapping Faces, Saving Features: A Dual-Purpose Pipeline for Pedestrian Privacy in ITS",
      "authors": [
        "Roba H. Farouk",
        "Catherine M. Elias"
      ],
      "abstract": "Large-scale and diverse datasets are needed to train AI models to take real-time decisions for autonomous vehicles (AVs), an intelligent transportation system (ITS) application. Pedestrian intention and trajectory prediction are critical models used in AVs, requiring datasets involving diverse pedestrian images. Unrestricted access to these datasets imposes serious security risks, like identity theft and pedestrian tracking. The challenge is to apply privacy preservation procedures while maintaining the image attributes needed to train the models. Existing privacy methods may preserve the pedestrian's privacy, but degrade the image usability, which hinders the models' effectiveness. This work's focus is to implement a five-stage pipeline to protect pedestrians' privacy through face swapping while keeping the essential facial attributes intact. It should be tailored to satisfy the privacy needs of the Egy-DRiVeS dataset. Moreover, Roop and Ghost-v2 face-swapping models are evaluated. Provenly, Roop outperforms Ghost-v2 in various aspects, as will be discussed. Consequently, Roop is the face-swapping model to be used in the pipeline to strike the balance between pedestrian privacy via identity concealment and data usability via facial attribute preservation.",
      "short_abstract": "Large-scale and diverse datasets are needed to train AI models to take real-time decisions for autonomous vehicles (AVs), an intelligent transportation system (ITS) application. Pedestrian intention and trajectory prediction are critical models used in AVs,...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 2,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.08402",
      "pdf_url": "https://arxiv.org/pdf/2607.08402.pdf"
    },
    {
      "id": "2607.08316",
      "anchor": "paper-2607-08316",
      "title": "INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis",
      "authors": [
        "Logine M. Zaki",
        "Catherine M. Elias"
      ],
      "abstract": "Vehicle intention prediction is a pivotal aspect in the agility and safety of autonomous vehicles in all driving scenarios; if genuine enhancement of autonomous vehicles are required, we need to make them adopt human interpretation of driver's intention especially in cases that require a lot of human interaction as well as complex driving behaviors like the ones at intersections, roundabouts and emergency cases such as sudden stops where vehicle intention prediction helps in taking the correct evasive action within a real time period where every second of action makes an impact and can prevent a catastrophe from taking place. In the worst case, it helps minimize the damage and make safety a priority. Intention prediction can also be used to enhance trajectory prediction (intention conditioned trajectory prediction). In this study, The INTENT framework is proposed using LSTM model to predict the vehicle's intention at intersections 2 seconds ahead of the event occurrence to predict whether the cars in intersections are going straight, turning left, or turning right. Various model experiments and ablation study are thoroughly tested on InD dataset achieving 99.71% accuracy.",
      "short_abstract": "Vehicle intention prediction is a pivotal aspect in the agility and safety of autonomous vehicles in all driving scenarios; if genuine enhancement of autonomous vehicles are required, we need to make them adopt human interpretation of driver's intention...",
      "published": "2026-07-09",
      "updated": "2026-07-09",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 25,
        "evidence": [
          "abstract: trajectory prediction",
          "title: behavior prediction",
          "abstract: behavior prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 41,
      "arxiv_url": "https://arxiv.org/abs/2607.08316",
      "pdf_url": "https://arxiv.org/pdf/2607.08316.pdf"
    },
    {
      "id": "2607.09787",
      "anchor": "paper-2607-09787",
      "title": "Adversarially Guided Diffusion for LiDAR Range Image Synthesis",
      "authors": [
        "Stavros Bouras",
        "Antonios Makris",
        "Alexandros Gkillas",
        "Aris S. Lalos",
        "Konstantinos Tserpes"
      ],
      "abstract": "LiDAR semantic segmentation is a key perception task in autonomous driving, where false predictions can affect downstream planning and safety-critical decision-making. Although adversarial attacks, and specifically adversarial examples, have been widely studied for image classification and 3D point cloud segmentation, unrestricted adversarial examples remain largely unexplored in the space of 2D range images, which are projections of 3D point clouds. The proposed method is, to the best of our knowledge, the first diffusion-based unrestricted adversarial attack against 2D range-image segmentation, using adversarial guidance from a segmentation loss. By applying guidance directly during sampling, the method produces unrestricted adversarial examples that remain close to the learned LiDAR data manifold while inducing structured segmentation errors. Experiments on the SemanticKITTI dataset using RangeNet++ and CENet segmentation networks demonstrate that the attack provides adjustable degradation across guidance strengths and transfers across segmentation architectures. Compared with norm-bounded FGSM and SegPGD baselines, the proposed attack offers a distinct effectiveness-realism trade-off, achieving controllable white-box and transfer degradation while maintaining competitive distributional and visual realism.",
      "short_abstract": "LiDAR semantic segmentation is a key perception task in autonomous driving, where false predictions can affect downstream planning and safety-critical decision-making. Although adversarial attacks, and specifically adversarial examples, have been widely...",
      "published": "2026-07-08",
      "updated": "2026-07-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 14,
        "evidence": [
          "abstract: perception",
          "abstract: scene segmentation",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 42,
      "arxiv_url": "https://arxiv.org/abs/2607.09787",
      "pdf_url": "https://arxiv.org/pdf/2607.09787.pdf"
    },
    {
      "id": "2607.08072",
      "anchor": "paper-2607-08072",
      "title": "Post-Training in End-to-End Autonomous Driving",
      "authors": [
        "Ruining Yang",
        "Muxing Wang",
        "Yixiao Chen",
        "Tongfei Guo",
        "Yi Xu",
        "Can Cui",
        "Zichong Yang",
        "Yitian Zhang",
        "Ziran Wang",
        "Yun Fu",
        "Lili Su"
      ],
      "abstract": "End-to-end models that map multimodal inputs directly to future trajectories/maneuvers have emerged as an increasingly prominent research paradigm in autonomous driving. This class of models includes both Vision-Language-Action models and trajectory-generative planners. Unlike classic machine learning applications, autonomous vehicles operate in safety-critical and interaction-intensive environments where traditional open-loop imitation of expert demonstrations is not sufficient to ensure reliability. In particular, small execution errors can accumulate over time, while recovery behaviors are scarce in training data. In addition, long-horizon objectives such as safety and driving comfort are not captured by pointwise labels either. These limitations have motivated a shift toward post-training techniques, which further refine driving policies beyond pure imitation. This survey presents a unified view of post-training for autonomous driving by defining its scope and organizing the existing literature into four major families based on the form of supervision they use. For each family, we discuss its capabilities, limitations, and open challenges. We aim to facilitate a systematic understanding of this emerging area and stimulate future research on reliable and efficient post-training for autonomous driving.A collection of related papers is available at https://github.com/RYNing/Awesome-Post-Training-In-Autonomous-Driving-Papers.",
      "short_abstract": "End-to-end models that map multimodal inputs directly to future trajectories/maneuvers have emerged as an increasingly prominent research paradigm in autonomous driving. This class of models includes both Vision-Language-Action models and...",
      "published": "2026-07-08",
      "updated": "2026-07-08",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Survey"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 32,
        "evidence": [
          "title: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 42,
      "arxiv_url": "https://arxiv.org/abs/2607.08072",
      "pdf_url": "https://arxiv.org/pdf/2607.08072.pdf"
    },
    {
      "id": "2607.07844",
      "anchor": "paper-2607-07844",
      "title": "Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning",
      "authors": [
        "Alessandro Canevaro",
        "Hang Yu",
        "Julian Schmidt",
        "Peizheng Li",
        "Silvan Lindner",
        "Wilhelm Stork",
        "Georg Martius",
        "Julian Jordan"
      ],
      "abstract": "While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) performance, their generalization to novel urban topologies and recovery mechanisms following execution perturbations remain under-explored. To address this, we present Shift & Drift, a novel dual-track benchmark designed to rigorously stress-test motion planners across two critical axes of distribution shift: (1) The Semantic Shift Track leverages a novel conversion pipeline that transforms the aerial, DeepScenario Open 3D dataset into the nuPlan simulation framework. This enables zero-shot evaluation of planners trained on North American and Singaporean data against 1,182 scenarios spanning four German cities and the US city of San Francisco featuring dense pedestrian-cyclist interactions. (2) The State-Distribution Drift Track injects stochastic perturbations into the ego vehicle's dynamics to quantify robustness against compounding execution errors. Based on this, we systematically evaluate the failure modes of diverse planning paradigms under semantic and state-distribution shifts. While imitation learning methods achieve high scores in ID benchmarks, they exhibit significant failures under semantic shift, particularly in pedestrian-dense environments, and suffer from persistent drift when subjected to temporally correlated actuation noise. In contrast, the evaluated reinforcement-learning-based planner demonstrates more graceful degradation, maintaining higher safety and progress metrics across both tracks. Our findings reveal an empirical trade-off between imitation fidelity and closed-loop resilience, providing the community with a rigorous benchmark to evaluate progress toward reliable deployment.",
      "short_abstract": "While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) performance, their generalization to novel urban topologies and recovery mechanisms following execution perturbations remain...",
      "published": "2026-07-08",
      "updated": "2026-07-08",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Benchmark",
        "Simulation",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 13,
        "evidence": [
          "title: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 42,
      "arxiv_url": "https://arxiv.org/abs/2607.07844",
      "pdf_url": "https://arxiv.org/pdf/2607.07844.pdf"
    },
    {
      "id": "2607.07601",
      "anchor": "paper-2607-07601",
      "title": "CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis",
      "authors": [
        "Kaicong Huang",
        "Meng Ma",
        "Ruimin Ke"
      ],
      "abstract": "Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem spanning visual representation, scene reasoning, and vehicle trajectory generation and control. Prior knowledge- and model-based approaches typically focus on scene or trajectory components in isolation, while diffusion-based methods attempt end-to-end generation but still struggle to ensure spatiotemporal consistency and physical realism. To unify these aspects within a single framework, we propose CARLA-GS, a modular corner-case synthesis pipeline that decouples visual representation, semantic reasoning, and physics-based execution while maintaining tight cross-module coupling. Starting from real driving data, we reconstruct an editable gaussian scene with additional geometry-consistent constraints. A multi-agent LLM then performs scene-level reasoning to identify risky interactions and generate intent-level waypoint trajectories, while the low-level motion control is delegated to CARLA, where a PID controller ensures kinematic and dynamic feasibility. The simulated vehicle states are finally re-projected into the gaussian scene for ego-centric rendering. This design enables high-level semantic reasoning, low-level physically executable motion, and photorealistic corner-case generation within a unified pipeline. Experiments on the Waymo Open Dataset show, both quantitatively and qualitatively, that our framework enables controllable corner-case generation and produces photorealistic, spatiotemporally consistent videos aligned with semantic intent and physically feasible motion.",
      "short_abstract": "Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem...",
      "published": "2026-07-08",
      "updated": "2026-07-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: controller",
          "Control & Vehicle Dynamics: abstract: control",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 42,
      "arxiv_url": "https://arxiv.org/abs/2607.07601",
      "pdf_url": "https://arxiv.org/pdf/2607.07601.pdf"
    },
    {
      "id": "2607.07196",
      "anchor": "paper-2607-07196",
      "title": "Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators",
      "authors": [
        "Christian Oefinger",
        "Finn Rasmus Schäfer",
        "Korbinian Moller",
        "Mattia Piccinini",
        "Johannes Betz"
      ],
      "abstract": "Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagined world, and returning a success or safety verdict. Yet a verdict is only as trustworthy as the WM that produced it, and the WM itself needs to be certified. In video-generation WMs, fidelity metrics such as Fréchet Video Distance (FVD) reward visual realism, but ignore whether the world responds correctly to the policy's actions, including those unseen in training. Classical simulation-based validation assumes a trusted simulator evaluating an untrusted policy, whereas generative WMs are themselves unverified learned artifacts. Hence, we argue that any WM used as a test oracle must first be accredited before its verdicts can serve as evidence. Building on credibility practices from safety-critical simulation, including Verification, Validation & Accreditation (VV&A), Safety of the Intended Functionality (SOTIF), and scenario-based testing standards, we define an admissibility ladder (L0-L4) that a WM must climb before its closed-loop verdicts are accepted as assurance evidence. Our framework is embodiment-agnostic, and is instantiated in autonomous driving (AD), where assurance methods for traditional simulation are most mature. Applied to two driving WMs, the lower rungs reveal a reversal: the model that ranks higher on visual generation quality (L0) ranks lower on action-following (L1-L2), so visual fidelity does not predict the action-robustness a closed-loop verdict depends on.",
      "short_abstract": "Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagined world, and returning a success or safety verdict. Yet a verdict is only as trustworthy as the WM that produced it,...",
      "published": "2026-07-08",
      "updated": "2026-07-08",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 16,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "abstract: validation or testing",
          "abstract: SOTIF",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 42,
      "arxiv_url": "https://arxiv.org/abs/2607.07196",
      "pdf_url": "https://arxiv.org/pdf/2607.07196.pdf"
    },
    {
      "id": "2607.07103",
      "anchor": "paper-2607-07103",
      "title": "A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving",
      "authors": [
        "Heye Huang",
        "Jingguang Li",
        "Zhiyuan Zhou",
        "Paul Liang",
        "Mingyu Wu",
        "Kitae Jang",
        "Jianqiang Wang"
      ],
      "abstract": "Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving data, and existing trajectory and language-augmented datasets seldom provide high-risk event labels, semantic annotations, and verifiable safety signals. Here we present K-Risk, a knowledge-augmented dataset that combines structured driving trajectories with large language model generated semantic annotations for safety-critical driving scenarios. K-Risk integrates 20 human-driven and autonomous-vehicle trajectory datasets from Europe, China, and the United States, covering highways, urban freeways, intersections, and roundabouts. Using a unified risk-centric extraction pipeline, K-Risk curates 31,398 high-risk events, together with a 1,036-event extreme subset of near-collision cases. Each event is released as a synchronized trajectory, metadata, and language triplet containing structured scenario descriptions, abnormal-behavior notifications, and, for a representative subset, causal risk analyses and action recommendations validated through a closed-loop simulator with iterative reflection. By combining multi-dimensional risk annotations, interpretable language supervision, and verifiable decisions, K-Risk bridges structured traffic trajectories, semantic reasoning, and decision supervision, providing a standardized foundation for developing and evaluating next-generation risk-aware autonomous driving agents.",
      "short_abstract": "Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving data, and existing...",
      "published": "2026-07-08",
      "updated": "2026-07-08",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset",
        "Simulation",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "abstract: safety",
          "abstract: risk assessment"
        ]
      },
      "recency": "fresh",
      "age_days": 42,
      "arxiv_url": "https://arxiv.org/abs/2607.07103",
      "pdf_url": "https://arxiv.org/pdf/2607.07103.pdf"
    },
    {
      "id": "2607.07114",
      "anchor": "paper-2607-07114",
      "title": "How the Fusion of Onboard Sensors and V2X Data can Improve (or not) the Cooperative Perception of Connected Automated Vehicles",
      "authors": [
        "Amir Mohammadisarab",
        "Miguel Sepulcre",
        "Luca Lusvarghi",
        "Javier Gozalvez"
      ],
      "abstract": "Automated vehicles rely on onboard sensors to perceive their surroundings and navigate autonomously. However, sensor performance may degrade under adverse weather conditions or when line-of-sight is obstructed. Cooperative perception (or collective perception) is expected to mitigate these limitations by enabling Connected and Automated Vehicles (CAVs) to share sensor data and collaboratively enhance situational awareness. Several studies have analyzed the potential of cooperative perception, yet the fusion of V2X data with information from onboard sensors has received limited focus. V2X data may contain errors that affect the quality of the fused data, and hence the effectiveness of cooperative perception. This study analyzes the impact of sensing measurement errors, V2X packet losses, and GNSS inaccuracies on the effectiveness of cooperative perception. The results highlight the potential of cooperative perception to enhance perception levels and range compared to using onboard sensors alone. However, they also identify key challenges related to the generation of ghost vehicles during the fusion process, which must be addressed to prevent V2X data from introducing additional errors when fused with onboard sensor data.",
      "short_abstract": "Automated vehicles rely on onboard sensors to perceive their surroundings and navigate autonomously. However, sensor performance may degrade under adverse weather conditions or when line-of-sight is obstructed. Cooperative perception (or collective...",
      "published": "2026-07-08",
      "updated": "2026-07-08",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 52,
        "margin": 40,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "title: connected vehicles",
          "abstract: connected vehicles"
        ]
      },
      "recency": "fresh",
      "age_days": 42,
      "arxiv_url": "https://arxiv.org/abs/2607.07114",
      "pdf_url": "https://arxiv.org/pdf/2607.07114.pdf"
    },
    {
      "id": "2607.12214",
      "anchor": "paper-2607-12214",
      "title": "Robust 4D Driving Scene Reconstruction from Imperfect Visual Priors",
      "authors": [
        "Xiaoyun Dong",
        "Qian Xu",
        "Yun Wang",
        "Yang Lu",
        "Jen-Ming Wu",
        "Jianping Wang"
      ],
      "abstract": "Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recent Gaussian Scene Graph (GSG) methods achieve impressive visual quality, they heavily rely on precise priors, such as accurate camera poses and LiDAR depth, or manual annotations. When initialized with noisy priors estimated from in-the-wild videos, existing GSG methods suffer from optimization ambiguity (e.g., entangling camera and agent poses) and topological failures (e.g., missing objects), causing severe rendering artifacts. To enable robust in-the-wild reconstruction, we introduce Adaptive Gaussian Graph (AGG), a self-correcting 4D framework. Our Semantically-Guided Tick-Tock Strategy leverages 2D foundation features to explicitly decouple static background and camera pose updates from dynamic agent learning. Concurrently, our Adaptive Topology Evolution module actively rectifies graph structures by spawning missing agents, reassigning misclassified Gaussians, and pruning false positives. To rigorously evaluate this in-the-wild setting, we introduce Wild-30, a challenging benchmark of internet and generative videos. Extensive experiments on KITTI and Wild-30 validate that AGG consistently outperforms state-of-the-art approaches in visual fidelity and robustness under noisy priors.",
      "short_abstract": "Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recent Gaussian Scene Graph (GSG) methods achieve impressive visual quality, they heavily rely on precise...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 5,
        "evidence": [
          "title: robustness",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.12214",
      "pdf_url": "https://arxiv.org/pdf/2607.12214.pdf"
    },
    {
      "id": "2607.09772",
      "anchor": "paper-2607-09772",
      "title": "A Risk-Field Enhanced Closed-Loop Digital Twin Framework for Autonomous Driving Safety Validation",
      "authors": [
        "Yongzhi Liu"
      ],
      "abstract": "Autonomous driving systems require reliable safety validation before real-world deployment. However, large-scale road testing is costly, difffcult to reproduce, and inefffcient for exposing rare safety-critical scenarios. Conventional simulation improves repeatability, but an offfine simulator alone cannot continuously connect physical trafffc states, virtual reconstruction, algorithm evaluation, and scenario evolution. This paper proposes a risk-ffeld enhanced closed-loop digital twin framework for autonomous driving safety validation. The framework integrates physical data acquisition, data synchronization, virtual twin reconstruction, risk-aware scenario generation, autonomous driving algorithm evaluation, and safety analysis. A driving risk ffeld is introduced as a uniffed intermediate representation to describe obstacle, lane-departure, road-boundary, time-to-collision, and comfort-related risks around the ego vehicle. The risk ffeld ranks high-risk scenarios in the digital twin scenario library and provides dense safety guidance for reinforcement learning-based driving policies. A simulation-style evaluation protocol is designed to compare conventional reinforcement learning baselines, risk-penalty baselines, and the proposed risk-ffeld guided method. The study indicates that embedding explicit risk structure into digital twins can make autonomous driving validation more targeted, interpretable, and reusable, while its practical effectiveness remains bounded by model ffdelity, risk calibration, and sim-to-real transfer.",
      "short_abstract": "Autonomous driving systems require reliable safety validation before real-world deployment. However, large-scale road testing is costly, difffcult to reproduce, and inefffcient for exposing rare safety-critical scenarios. Conventional simulation improves...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 29,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing",
          "abstract: risk assessment"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.09772",
      "pdf_url": "https://arxiv.org/pdf/2607.09772.pdf"
    },
    {
      "id": "2607.09764",
      "anchor": "paper-2607-09764",
      "title": "OmniSCS: Omni Safety-Critical Scenario Synthesis for Autonomous Driving via a Fully Editable Driving World",
      "authors": [
        "Xiaoyun Dong",
        "Qian Xu",
        "Yang Lu",
        "Yang Lou",
        "Yung-Hui Li",
        "Jianping Wang"
      ],
      "abstract": "The synthesis of safety-critical scenarios (SCS) and their evaluation through closed-loop simulations are crucial for developing robust autonomous driving systems. A key aspect of this process involves editing agent states in both appearance and trajectory levels within existing scenes. However, current methods struggle to preserve data fidelity after scene editing and fail to efficiently generate high-quality SCS through such modifications. To overcome these limitations, we propose OmniSCS, an innovative system that generates photorealistic SCS with high physical fidelity while enabling closed-loop testing in synthetic environments. OmniSCS comprises two key modules: 1) A Fully Editable Driving World Construction module that maintains high-fidelity agent appearance and background during scene editing via dual-strategy agent reconstruction and depth-refinement background reconstruction methods. 2) A SCS Synthesis module that facilitates object insertion and agent trajectory editing to synthesize diverse SCS while preserving data fidelity. Experiments on nuScenes, Waymo, and KITTI datasets show that OmniSCS outperforms state-of-the-art methods in edited scene fidelity. We further validate its ability to enhance autonomous driving algorithms and support real-time (13Hz) closed-loop testing. Overall, OmniSCS provides a safer, more effective, and cost-efficient solution for SCS optimization and testing in autonomous driving.",
      "short_abstract": "The synthesis of safety-critical scenarios (SCS) and their evaluation through closed-loop simulations are crucial for developing robust autonomous driving systems. A key aspect of this process involves editing agent states in both appearance and trajectory...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 18,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.09764",
      "pdf_url": "https://arxiv.org/pdf/2607.09764.pdf"
    },
    {
      "id": "2607.06957",
      "anchor": "paper-2607-06957",
      "title": "Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation",
      "authors": [
        "Seulbin Hwang",
        "Kiyoung Om",
        "Daejung Kim",
        "Jinhan Lee"
      ],
      "abstract": "Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks predominantly reward realism, and recent methods have optimized accordingly, leaving diversity underexplored. We introduce \\textbf{Flow-ERD}, a multi-agent simulator that pursues realism and diversity jointly. Its backbone, \\textbf{Agent-Type Aware Flow Matching} (AFM), couples flow matching's multi-modal expressiveness with type-specific kinematic execution. It preserves fine-grained diversity while keeping motions consistent with each agent type. A second stage, \\textbf{Entropy-Regularized Distillation} (ERD), fine-tunes the closed-loop rollout distribution with an entropy-regularized reverse-KL objective. This mitigates covariate shift while explicitly preventing collapse onto high-density modes. We evaluate Flow-ERD with a log-free diversity metric alongside standard realism scores. Flow-ERD ranks first on the WOSAC test benchmark and dominates the realism--diversity Pareto front among reproducible baselines. Our project page is available \\href{https://seulbinhwang.github.io/flow-erd-project-page/}{here}.",
      "short_abstract": "Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks predominantly reward realism, and recent methods have optimized accordingly, leaving diversity underexplored. We introduce \\textbf{Flow-ERD}, a...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06957",
      "pdf_url": "https://arxiv.org/pdf/2607.06957.pdf"
    },
    {
      "id": "2607.06516",
      "anchor": "paper-2607-06516",
      "title": "Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation",
      "authors": [
        "Songbur Wong",
        "Xiaosong Jia",
        "Junqi You",
        "Bo Zhang",
        "Pei Xu",
        "Renqiu Xia",
        "Yuping Qiu",
        "Shaofeng Zhang",
        "Zelin Zhao",
        "Xuechao Yan",
        "Yuchen Zhou",
        "Yurui Chen",
        "Wen Guo",
        "Hang Xu",
        "Junchi Yan"
      ],
      "abstract": "Evaluating end-to-end autonomous driving (E2E-AD) remains challenging, as existing driving simulation methods often trade off closed-loop interactivity (e.g., CARLA) and real-world visual fidelity (e.g., nuScenes). We present \\textbf{\\emph{Point as Skeleton}}, a generative sensor simulation framework for state-updated autoregressive driving video generation, in which an autoregressive generator synthesizes visual observations from step-wise updated ego states, actor states, scene maps, and point-cloud skeleton conditions. To support closed-loop rollout, we introduce Reset-and-Roll, which adapts rolling diffusion inference to simulation by preventing future-conditioned latent states from being committed across simulation steps. To stabilize error accumulation during step-wise autoregressive rollout, we introduce point-cloud skeletons that decouple foreground and background assets and project them into camera-view painted-point and template-depth conditions, providing appearance and geometric cues. We further implement a nuPlan-based renderer-level closed-loop generative interface for evaluating generation under ego deviations from the original log. Experiments on nuScenes and nuPlan show that \\textit{Point as Skeleton} improves autoregressive generation quality during closed-loop rollout, demonstrating its potential for visually faithful closed-loop driving simulation. The code is available at https://github.com/krauwu/point-as-skeleton.",
      "short_abstract": "Evaluating end-to-end autonomous driving (E2E-AD) remains challenging, as existing driving simulation methods often trade off closed-loop interactivity (e.g., CARLA) and real-world visual fidelity (e.g., nuScenes). We present \\textbf{\\emph{Point as...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 11,
        "margin": 2,
        "evidence": [
          "title: point cloud",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06516",
      "pdf_url": "https://arxiv.org/pdf/2607.06516.pdf"
    },
    {
      "id": "2607.06328",
      "anchor": "paper-2607-06328",
      "title": "Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models",
      "authors": [
        "Franz Motzkus",
        "Sebastian Bernhard"
      ],
      "abstract": "The increasing adoption of end-to-end learning for autonomous driving introduces increased model complexity and opacity, raising the risk of learning undesired or erroneous behavior. In this work, we integrate unsupervised dictionary learning as a post hoc interpretability module within state-of-the-art driving models to decompose driving behavior into semantically meaningful concepts while demonstrating their causal influence on the model's driving decisions. We propose a stepwise framework for extracting and interpreting meaningful concepts from the end-to-end model and connecting them to the multifaceted model outputs, thereby revealing the underlying decision-making logic for the prediction of future trajectories. Furthermore, targeted interventions at the concept level allow us to manipulate and correct driving decisions, resulting in measurable improvements in overall driving performance. We thus demonstrate how interpretability can effectively be used to reduce model opacity, uncover erroneous behavior, and enable targeted mitigation, ultimately boosting model performance.",
      "short_abstract": "The increasing adoption of end-to-end learning for autonomous driving introduces increased model complexity and opacity, raising the risk of learning undesired or erroneous behavior. In this work, we integrate unsupervised dictionary learning as a post hoc...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Planning & Decision-Making: abstract: decision-making",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06328",
      "pdf_url": "https://arxiv.org/pdf/2607.06328.pdf"
    },
    {
      "id": "2607.06319",
      "anchor": "paper-2607-06319",
      "title": "Synthetic-to-Real Translation for Class-Agnostic Motion Prediction",
      "authors": [
        "Yizheng Wu",
        "Hongwei Fan",
        "Kewei Wang",
        "Ruibo Li",
        "Xingyi Li",
        "Xiao Song",
        "Zhe Wang",
        "Chenjing Ding",
        "Dongliang Wang",
        "Zhiguo Cao",
        "Guosheng Lin"
      ],
      "abstract": "Motion understanding is critical for ensuring safety and robustness in autonomous driving systems, driving increasing interest in motion prediction. A key challenge in this domain is the high cost associated with acquiring real-world motion labels. It is therefore ideal if we could transfer motion knowledge from synthetic data to real data. In this context, we explore the potential of synthetic-to-real translation for motion prediction (SRMP). However, the most used naive motion regression methods are notably sensitive to the synthetic-to-real domain shift, resulting in unreliable knowledge translation. To address this, we propose a novel approach integrating a motion knowledge translation framework with two key components: (1) objectness-aware motion prediction, which explicitly models the joint distribution of motion patterns and objectness priors to improve domain-invariant feature learning, and (2) objectness-aided motion enhancement, a motion label refinement mechanism that leverages learned objectness priors to filter motion noise. Furthermore, we present a physically-based pipeline for generating Motion4D, the first synthetic 4D LiDAR dataset tailored for SRMP research, addressing the lack of synthetic motion datasets. Experimental results demonstrate that our approach effectively bridges the domain gaps and yields superior performance on real scenes.",
      "short_abstract": "Motion understanding is critical for ensuring safety and robustness in autonomous driving systems, driving increasing interest in motion prediction. A key challenge in this domain is the high cost associated with acquiring real-world motion labels. It is...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 22,
        "evidence": [
          "title: motion forecasting",
          "abstract: motion forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06319",
      "pdf_url": "https://arxiv.org/pdf/2607.06319.pdf"
    },
    {
      "id": "2607.06078",
      "anchor": "paper-2607-06078",
      "title": "Collaborative Multi-Agent Testing for Emergent Failure Discovery in Autonomous Driving Systems",
      "authors": [
        "Ruizhen Gu",
        "Konstantinos Koufos",
        "Donghwan Shin",
        "Vahid Garousi",
        "Mehrdad Dianati"
      ],
      "abstract": "Autonomous Driving Systems (ADS) can fail because of faults within individual modules as well as from interactions across perception, planning, and control. Yet existing ADS testing research often treats key testing functions, such as perturbation generation, behavioural assessment, and test case selection and exploration, as loosely coupled steps rather than coordinated roles for discovering such failures. We present CREAD, a collaborative multi-agent testing framework for testing ADS that organises perturbation generation, behavioural validation, and search coordination through a shared blackboard and an orchestrator. In the current work-in-progress instantiation, the framework focuses on perception-oriented perturbation generation, while remaining extensible to other ADS modules, including planning and control. It currently comprises a Perception Fuzzer Agent, a Metamorphic Validator Agent, and an Orchestrator Agent. Respectively, they generate perturbations, assess behavioural consistency across related scenario pairs, and coordinate further exploration. Experiments in HighwayEnv simulator show that the collaborative configuration improves failure discovery in the highway environment and remains competitive in the roundabout setting. Across the two environments, it yields about 2.1x as many failures per 100 scenarios as the single-agent baseline on average, while gains over a non-collaborative two-agent baseline vary across environments. These results suggest that collaborative multi-agent testing is a promising research direction for emergent ADS behaviour discovery.",
      "short_abstract": "Autonomous Driving Systems (ADS) can fail because of faults within individual modules as well as from interactions across perception, planning, and control. Yet existing ADS testing research often treats key testing functions, such as perturbation generation,...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 8,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06078",
      "pdf_url": "https://arxiv.org/pdf/2607.06078.pdf"
    },
    {
      "id": "2607.06484",
      "anchor": "paper-2607-06484",
      "title": "Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles",
      "authors": [
        "Marwan Lazrag",
        "Badis Hammi",
        "Lorena Gonzalez-Manzano",
        "Joaquin Garcia-Alfaro"
      ],
      "abstract": "Poisoning attacks against public datasets lead to major concerns, such as (i) misclassification of perceived objects when the poisoned data is used for training and (ii) embedding of backdoors that may eventually be triggered later on, when specific conditions in the system apply over the learned models. Its impact over data augmentation models is unclear. While data augmentation reduces the likelihood of poisoning attack success, some valid questions remain. Is data augmentation affecting the impact of poisoning attacks? can it increase the number of poisoned samples or injected backdoors? We explore in this paper some of these questions. We assess the effects of augmenting poisoned 3D point cloud datasets and validate that poisoning is able to evade the sanitizing nature of augmentation techniques when using the concrete case of Generative Adversarial Network (GAN) techniques to exemplify the case of data augmentation processing. We also validate that poisoning propagates over the augmented datasets and perturbs the decision made by general-purpose classifiers, in the end. All the experimental material (including tools, datasets, and classifiers) is publicly available, to facilitate reproducibility and to foster further research in the topic.",
      "short_abstract": "Poisoning attacks against public datasets lead to major concerns, such as (i) misclassification of perceived objects when the poisoned data is used for training and (ii) embedding of backdoors that may eventually be triggered later on, when specific...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 2,
        "evidence": [
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06484",
      "pdf_url": "https://arxiv.org/pdf/2607.06484.pdf"
    },
    {
      "id": "2607.06039",
      "anchor": "paper-2607-06039",
      "title": "Automating Quality Assessment with NLP of LLM-Generated Defeaters",
      "authors": [
        "Tihomir Rohlinger",
        "Daniel Ratiu",
        "Stefan Wagner"
      ],
      "abstract": "High-integrity systems, such as autonomous vehicle fleets and large-scale energy infrastructures, rely on structured assurance cases to justify safety claims. To remain valid under evolving operational conditions, such cases must be examined against potential challenges, known as defeaters. While large language models (LLMs) can support the scalable generation of candidate defeaters, assessing their quality remains largely manual and subjective process. This paper presents an automated approach for supporting the assessment of LLM-generated defeaters using natural language processing techniques. The method combines structural features from assurance case graphs with semantic embeddings and meta-classifiers trained on expert-assessed defeater annotations. We evaluate the approach through two case studies in the automotive and energy domains. The results show substantial human reviewer dissensus, with Cohen's kappa values below 0.442, highlighting the difficulty of consistent manual assessment. Against this background, the proposed classifiers achieve an average F1-score of 0.84 in validation and show improved alignment with individual expert ratings. The findings suggest that automated assessment can help reduce subjective variance and provide scalable decision support for assurance case review, while leaving final judgment to domain experts.",
      "short_abstract": "High-integrity systems, such as autonomous vehicle fleets and large-scale energy infrastructures, rely on structured assurance cases to justify safety claims. To remain valid under evolving operational conditions, such cases must be examined against potential...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06039",
      "pdf_url": "https://arxiv.org/pdf/2607.06039.pdf"
    },
    {
      "id": "2607.06771",
      "anchor": "paper-2607-06771",
      "title": "Integrated Automated Car Following and Lane-changing control based on a Parametrized Deep Q-network with Hybrid Action Space",
      "authors": [
        "Hao Zhang",
        "Zihao Li",
        "Yang Zhou"
      ],
      "abstract": "Lane-change, a triggering of traffic disturbances to the upstream vehicles, is detrimental to traffic safety and efficiency. Coupled with car-following behavior, the joint maneuvers depict the general picture of how traffic disturbances generate and propagate through vehicle streams, especially under traffic congestion. This study proposes an integrated control framework for lane-changing and car-following for connected and automated vehicles (CAVs), where those two tasks are largely treated as independent driving tasks by prevailing methods. Utilizing the Parametrized Deep Q-Network (P-DQN) with a hybrid action space, the framework adeptly models multiple objectives in CAV control. The P-DQN's high-level control is employed for discrete lane-change decisions, while its low-level control manages continuous acceleration actions, i.e., lateral and longitudinal acceleration. These actions are interdependently determined, seamlessly integrating car-following and lane-changing control. By training to maximize cumulative rewards, the proposed control strategy ensures driving safety as well as the efficiency of car-following, lane-changing, and lane-keeping. Through numerical experiments, it is indicated that the P-DQN outperforms separated control methods, e.g., the combination of the Minimizing Overall Braking Decelerations Induced by Lane Changes (MOBIL) model and the Intelligent Driver Model (IDM), in terms of safety and comfort.",
      "short_abstract": "Lane-change, a triggering of traffic disturbances to the upstream vehicles, is detrimental to traffic safety and efficiency. Coupled with car-following behavior, the joint maneuvers depict the general picture of how traffic disturbances generate and propagate...",
      "published": "2026-07-07",
      "updated": "2026-07-07",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 2,
        "evidence": [
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 43,
      "arxiv_url": "https://arxiv.org/abs/2607.06771",
      "pdf_url": "https://arxiv.org/pdf/2607.06771.pdf"
    },
    {
      "id": "2607.05801",
      "anchor": "paper-2607-05801",
      "title": "TRIG: Trajectory-Rig Decoupled Metric Geometry Learning",
      "authors": [
        "Lizhou Liao",
        "Wentao Xu",
        "Handong Wang",
        "Lirong Yang",
        "Shuai Yang",
        "Weiwei Liu",
        "Chang Huang"
      ],
      "abstract": "Vision-centric autonomous driving requires accurate metric geometry and ego-motion estimation from synchronized multi-camera observations. Recent visual geometry models show strong performance in pose estimation, depth prediction, and 3D reconstruction, but are not tailored to rigid multi-camera driving systems. They often encode camera poses as entangled representations, in which time-varying ego-motion and static camera-rig geometry are jointly modeled, limiting the utilization of vehicle-side geometric priors. We propose Trajectory-Rig Decoupled Metric Geometry Learning (TRIG), a geometry perception framework for autonomous driving. TRIG factorizes camera poses into ego-trajectory and camera-rig components, enabling separate modeling of ego-motion and static multi-camera topology. We introduce decoupled pose encoding and supervision, which separately constrain trajectory evolution and rig geometry for metric-consistent learning. Moreover, sparse Temporal--Spatial attention separates cross-camera interaction from temporal aggregation, reducing global attention cost while preserving geometric reasoning. Experiments on five autonomous driving benchmarks show that TRIG achieves state-of-the-art performance in pose estimation, metric depth prediction, and 3D reconstruction.",
      "short_abstract": "Vision-centric autonomous driving requires accurate metric geometry and ego-motion estimation from synchronized multi-camera observations. Recent visual geometry models show strong performance in pose estimation, depth prediction, and 3D reconstruction, but...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: camera",
          "Mapping & Localization: abstract: pose estimation"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.05801",
      "pdf_url": "https://arxiv.org/pdf/2607.05801.pdf"
    },
    {
      "id": "2607.05783",
      "anchor": "paper-2607-05783",
      "title": "Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective",
      "authors": [
        "Tianyuan Zhang",
        "Xianglong Liu",
        "Aishan Liu",
        "Lu Wang",
        "Yitong Zhang",
        "Peng Yue",
        "Mingchuan Zhang",
        "Siyuan Liang",
        "Dacheng Tao"
      ],
      "abstract": "Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments. They can disturb visual perception, leading to misinterpretation of the scene and posing serious safety risks to autonomous driving (AD) systems. However, existing researches largely overlook these phenomena, leaving a critical gap. To address this issue, we study AD robustness through the lane perception perspective, a fundamental task supporting core functions like cruise control and lane centering. We focus on two representative models: conventional lane detection (LD) and vision-language model-based systems (ADVLMs). In this work, we introduce the first benchmark, LanEvil++, for evaluating the robustness of lane perception under environmental illusions. LanEvil++ encompasses 14 types of illusions and leverages the CARLA simulator to generate 94 high-fidelity, fully controllable 3D scenes, yielding a dataset of 90,292 annotated images, 1,596 video clips, and 41,855 visual question answering pairs. Extensive evaluations demonstrate that environmental illusions substantially degrade the performance of state-of-the-art LD methods. On average, LD models experience a 5.27% drop in Accuracy and a 10.49% decline in F1-score, while ADVLMs show a 2.03% reduction in GPT-score and a 0.75% drop in Language-score. Among all illusions, shadows emerge as the most disruptive factor, reducing accuracy by up to 7.20%. Furthermore, closed-loop simulations reveal that these illusions can lead to incorrect driving decisions. Complementary real-world case studies highlight safety-critical failures in actual traffic scenes. To enhance robustness, we propose the Multimodal Illusion Defense Approach (MIDA). MIDA achieves substantial gains under challenging conditions, boosting robustness by 4.23% on LD models and 3.82% on ADVLMs.",
      "short_abstract": "Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments. They can disturb visual perception, leading to misinterpretation of the scene and posing serious safety...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 15,
        "margin": 1,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: lane detection"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.05783",
      "pdf_url": "https://arxiv.org/pdf/2607.05783.pdf"
    },
    {
      "id": "2607.05180",
      "anchor": "paper-2607-05180",
      "title": "VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving",
      "authors": [
        "Tianjia Yang",
        "Ke Li",
        "Ruwen Qin",
        "Xianbiao Hu"
      ],
      "abstract": "Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the vehicle can perceive and what it can physically do. Human drivers cope by anticipation, reasoning about the scene and re-budgeting speed, following distance, and steering before grip or sight is lost, whereas current autonomous driving systems at best react after the fact. This paper proposes VLM-CASE, a framework that gives an autonomous vehicle this anticipatory capacity while keeping its motion bounded by a formal safety model at all times. A vision-language model (VLM), fine-tuned with low-rank adaptation (LoRA), reasons about the scene from the front-camera image and reports the road surface and visibility conditions. This output parametrizes a context-adaptive safety envelope (CASE), derived from physical limits and the guarantees of responsibility-sensitive safety, that couples braking and steering through a shared friction budget. A model predictive controller then drives freely within the envelope, while the VLM runs asynchronously so it never blocks the real-time control loop. We validate the framework in closed-loop CARLA simulation on tasks that demand both lateral and longitudinal control, across a range of weather, road-surface, and lighting conditions. The resulting controller, VLM-CASE-MPC, completes all trials, outperforming a conventional MPC baseline and a state-of-the-art VLM-integrated controller. Ablations confirm that the gains come from context adaptation, with the friction and visibility adaptations proving complementary. Furthermore, the framework is controller-agnostic and pairs with almost any low-level controller, offering a promising direction for safe autonomous driving. The dataset and supplementary materials for VLM-CASE are available at https://github.com/ytj254/VLM-CASE.",
      "short_abstract": "Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the vehicle can perceive and what it can physically do. Human drivers cope by anticipation, reasoning about the...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: model predictive control",
          "Control & Vehicle Dynamics: abstract: vehicle control",
          "Control & Vehicle Dynamics: abstract: controller",
          "Control & Vehicle Dynamics: abstract: control",
          "Control & Vehicle Dynamics: abstract: steering or braking",
          "Safety, Security & Verification: title: safety",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.05180",
      "pdf_url": "https://arxiv.org/pdf/2607.05180.pdf"
    },
    {
      "id": "2607.05133",
      "anchor": "paper-2607-05133",
      "title": "UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation",
      "authors": [
        "Mengmeng Liu",
        "Diankun Zhang",
        "Jiuming Liu",
        "Jianfeng Cui",
        "Hongwei Xie",
        "Guang Chen",
        "Hangjun Ye",
        "Francesco Nex",
        "Hao Cheng",
        "Michael Ying Yang"
      ],
      "abstract": "World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as dense supervision for scene dynamics and temporal causality. However, it remains unclear which architecture better transfers video-modeling benefits to trajectory generation. Existing cascaded or dual-DiT designs separate video imagination from action prediction, weakening the transfer of video-learned world dynamics to the trajectory branch: the action model may still overfit dataset-specific driving priors, while the video model only indirectly regularizes planning. We propose UNIVERSE, a unified video-action model built upon a single mask-modulated Diffusion Transformer. By co-training future video latents and ego-trajectory tokens within shared generative parameters, UNIVERSE allows dense video supervision to directly shape trajectory denoising, leading to stronger cross-domain action generalization. To ensure causal validity and efficient deployment, we introduce a Modality-Decoupling Visibility Mask, which shares historical context across modalities while blocking mutual attention between future video and trajectory tokens. This prevents future-target leakage and enables trajectory-only inference by removing future-video denoising at test time, achieving a $4.3\\times$ speedup over joint video-action rollout while maintaining comparable planning accuracy. The same model also supports video-only and joint video-action rollouts. Experiments show that UNIVERSE achieves 91.0 PDMS on NAVSIM (vs. 89.6 for the Two-DiT variant), and demonstrates strong zero-shot transfer to nuScenes and Bench2Drive without fine-tuning, while ablations confirm the importance of single-DiT unification, video co-training, and mask-based modality decoupling.",
      "short_abstract": "World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as dense supervision for scene dynamics and temporal causality. However, it remains unclear which architecture...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: abstract: planning",
          "Planning & Decision-Making: abstract: trajectory generation",
          "Prediction & World Models: abstract: world model",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.05133",
      "pdf_url": "https://arxiv.org/pdf/2607.05133.pdf"
    },
    {
      "id": "2607.04974",
      "anchor": "paper-2607-04974",
      "title": "A Comprehensive Study of Implementation Bugs in Multi-modal Agents",
      "authors": [
        "Suwan Li",
        "Lei Bu",
        "Shangqing Liu",
        "Yile Wang",
        "Guangdong Bai",
        "Fuman Xie",
        "Kai Chen",
        "Chang Yue"
      ],
      "abstract": "Multi-Modal Agents (M-agents), empowered by Large Language Models (LLMs), excel in various complex, open-world scenarios such as autonomous driving and robotics. However, their unique requirements to interact with dynamic and diverse multi-modal environments introduce novel implementation challenges beyond those faced by traditional agents. Outdated perception, untrustworthy planning and inapplicable execution could cause traffic accident and financial loss. Despite growing study on agent issues, there has not been a systematic study focusing on M-agent-specific implementation bugs. To address this gap, we conducted the first systematic study of implementation bugs in M-agents. We collected 34 representative M-agents from diverse sources and, through meticulous filtering,identified 158 M-agent-specific bugs from 1,268 issue reports. Using a top-down strategy, we developed a comprehensive taxonomy that classifies bugs by global symptoms, functionality component-level symptoms, and root causes. We then implemented MATester, an automatic proof-of-concept bug identifier by analyzing runtime inter-component outputs. When applied to 12 extra M-agents, MATester successfully covered 61.4% of known open issues and discovered 31 additional bugs, demonstrating the practical usefulness of our study. Our work provides a comprehensive reference and guideline for classification, prevention and fix of M-agent bugs.",
      "short_abstract": "Multi-Modal Agents (M-agents), empowered by Large Language Models (LLMs), excel in various complex, open-world scenarios such as autonomous driving and robotics. However, their unique requirements to interact with dynamic and diverse multi-modal environments...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04974",
      "pdf_url": "https://arxiv.org/pdf/2607.04974.pdf"
    },
    {
      "id": "2607.04953",
      "anchor": "paper-2607-04953",
      "title": "Real-World Perturbation Testing of Autonomous Driving Systems",
      "authors": [
        "Stefano Carlo Lambertenghi",
        "Matthias Weil",
        "Andrea Stocco"
      ],
      "abstract": "Autonomous Driving Systems (ADS) must operate reliably under diverse conditions, yet representative data for rare or adverse scenarios is difficult to obtain. Perturbation-based testing is widely used to assess robustness, but most studies focus on offline datasets or simulation, leaving open questions about how such results translate to real-world driving. We present a large-scale study of 72 camera and LiDAR perturbations, evaluated across three testing modalities: offline model-level analysis, hardware-in-the-loop execution, and closed-loop system-level testing on a full-scale autonomous vehicle. The study covers both an end-to-end vision-based driving model and a modular LiDAR-based perception and planning stack. Our results reveal a clear gap between testing levels. For camera-based systems, perturbations with limited offline impact can still induce unstable control and failures in real-world driving. For LiDAR-based systems, degradation is more consistent at the perception level but weakly predictive of system-level failures. Across both modalities, model-level metrics alone are insufficient to identify the most harmful perturbations. We further show that real-time feasibility is a key constraint in real-world testing, and that robustness observations obtained from recorded data do not consistently transfer to closed-loop behavior on a physical vehicle, highlighting the importance of complementary real-world, system-level evaluation.",
      "short_abstract": "Autonomous Driving Systems (ADS) must operate reliably under diverse conditions, yet representative data for rare or adverse scenarios is difficult to obtain. Perturbation-based testing is widely used to assess robustness, but most studies focus on offline...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04953",
      "pdf_url": "https://arxiv.org/pdf/2607.04953.pdf"
    },
    {
      "id": "2607.04812",
      "anchor": "paper-2607-04812",
      "title": "TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving",
      "authors": [
        "Miguel Antunes-García",
        "Santiago Montiel-Marín",
        "Fabio Sánchez-García",
        "Rodrigo Gutiérrez-Moreno",
        "Rafael Barea",
        "Luis M. Bergasa"
      ],
      "abstract": "Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles. This absence of explicit semantic awareness imposes limitations on the capacity of the model to solve ambiguities in complex scenarios, particularly those where object-specific behavior is essential for accurate forecasting (e.g. overtaking, intersections). In this paper, we introduce Text-Guided Representation for Instance Prediction (TGRIP), a novel framework that bridges this gap by injecting rich semantic priors into the instance prediction loop. The proposed teacher-student pipeline employs Vision-Language Foundation Models to generate dense, semantic-enhanced BEV maps from multi-camera images. These maps serve as auxiliary supervision during training, guiding the network to learn spatio-temporal representations that are not only geometrically consistent but also semantically discriminative. To the best of our knowledge, this represents the first attempt to unify semantic guidance with the temporal task of future instance prediction. The experimental results demonstrate that TGRIP surpasses existing state-of-the-art models in nuScenes, validating the hypothesis that semantic enrichment is a fundamental element for robust, end-to-end motion prediction. Code is available on https://github.com/miguelag99/TGRIP.",
      "short_abstract": "Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 7,
        "evidence": [
          "abstract: motion forecasting",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04812",
      "pdf_url": "https://arxiv.org/pdf/2607.04812.pdf"
    },
    {
      "id": "2607.04803",
      "anchor": "paper-2607-04803",
      "title": "E-CoDrive: A Co-Simulation Framework for Testing Energy-Critical Driving Scenarios",
      "authors": [
        "Manfredi Napolitano",
        "Alessandra Somma",
        "Alessio Gambi",
        "Andrea Stocco",
        "Nicola Mazzocca"
      ],
      "abstract": "Autonomous driving research has largely focused on safety while giving limited attention to non-functional aspects such as energy consumption and sustainability. As Autonomous Electric Vehicles (AEVs) become increasingly common in urban traffic, understanding how complex traffic dynamics influence their energy consumption is paramount to test whether AEVs can complete trips before battery depletion. To support energy-aware scenario-based testing of AEVs, we present E-CoDrive, a framework for reproducible closed-loop driving co-simulations that integrates an energy consumption model, a micro-traffic simulator, and a high-fidelity driving simulator to test AEV software stacks in urban scenarios. This tool paper describes the architecture of E-CoDrive and demonstrates its applicability by testing an Autoware-based AEV stack. Our evaluation shows that varying traffic conditions produce substantial differences in vehicle energy consumption. The artifact is publicly available at https://doi.org/10.6084/m9.figshare.32244783, and a screencast showing the tool is available at https://youtu.be/yX9fWHqCvgc.",
      "short_abstract": "Autonomous driving research has largely focused on safety while giving limited attention to non-functional aspects such as energy consumption and sustainability. As Autonomous Electric Vehicles (AEVs) become increasingly common in urban traffic, understanding...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04803",
      "pdf_url": "https://arxiv.org/pdf/2607.04803.pdf"
    },
    {
      "id": "2607.04770",
      "anchor": "paper-2607-04770",
      "title": "Cam2Sim: Neural Scenario Reconstruction for Closed-Loop Autonomous Driving Simulation",
      "authors": [
        "Davide Jannussi",
        "Stefano Carlo Lambertenghi",
        "Constantin Carste",
        "Andrea Stocco"
      ],
      "abstract": "Simulation-based testing enables safe and repeatable evaluation of autonomous driving systems, but its effectiveness is limited by the gap between synthetic simulator outputs and real-world camera observations. To address this problem, we present Cam2Sim, a tool that transforms real-world driving recordings into playable CARLA simulation scenarios. Starting from camera images and poses, Cam2Sim reconstructs road geometry, ego trajectories, parked vehicles, and simulation assets, and augments the reconstructed environment with Gaussian Splatting to render camera observations that resemble the original recording. The framework supports ROS-based data extraction, parked-vehicle detection, OpenStreetMap-based map generation, CARLA scenario construction, Gaussian Splatting training, trajectory replay, and closed-loop execution with a system under test. We validate Cam2Sim on a real-world urban-driving scenario with a camera-based end-to-end driving model, comparing reconstruction quality, image-generation quality, and closed-loop behavior against both a simulation-only baseline and the real-world target. Results show that Gaussian-Splatting-based rendering reduces the visual gap with respect to standard simulator rendering and improves behavioral similarity to the real-world reference runs. The artifact is publicly available at https: //github.com/ast-fortiss-tum/cam2sim, and a screencast showing the tool is available at https://youtu.be/KmZ74l1__lI",
      "short_abstract": "Simulation-based testing enables safe and repeatable evaluation of autonomous driving systems, but its effectiveness is limited by the gap between synthetic simulator outputs and real-world camera observations. To address this problem, we present Cam2Sim, a...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 6,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04770",
      "pdf_url": "https://arxiv.org/pdf/2607.04770.pdf"
    },
    {
      "id": "2607.04732",
      "anchor": "paper-2607-04732",
      "title": "SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction",
      "authors": [
        "Pin Tang",
        "Zhongdao Wang",
        "Guoqing Wang",
        "Xiangxuan Ren",
        "Chao Ma"
      ],
      "abstract": "Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty space, while BEV and TPV projections compromise fine-grained 3D structure. Fully sparse representations offer an attractive alternative, but existing methods, including SparseOcc, entangle scene completion with semantic prediction by indiscriminately propagating high-dimensional features into empty regions and applying voxel-wise classification. This creates excessive activations, computational overhead, and geometric ambiguity. We present SparseOcc++, a geometry-aware sparse framework that explicitly decouples scene completion from semantic segmentation. SparseOcc++ reformulates completion as signed-distance regression on sparse anchor voxels through a scene completion field (SCF). To model complex outdoor geometry robustly, it combines orthogonal decomposition with discretized distance learning. A geometry-guided propagation module then converts the SCF into a complete volumetric scene and restricts semantic segmentation to geometrically verified regions. Experiments establish new state of the art: SparseOcc++ improves IoU by 2.3 points and is 3.9x faster than SparseOcc on nuScenes, while achieving a 5.9x speedup over OccFormer on SemanticKITTI.",
      "short_abstract": "Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty space, while BEV and TPV projections compromise fine-grained 3D structure. Fully sparse representations offer...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 6,
        "evidence": [
          "abstract: scene segmentation",
          "title: occupancy",
          "abstract: occupancy",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04732",
      "pdf_url": "https://arxiv.org/pdf/2607.04732.pdf"
    },
    {
      "id": "2607.04689",
      "anchor": "paper-2607-04689",
      "title": "A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving",
      "authors": [
        "Argho Dey",
        "Yunfei Yin",
        "Swachha Ray",
        "Md Minhazul Islam",
        "Zheng Yuan",
        "Sijing Xiong",
        "Hongyu Liu",
        "Zhiqiu Huang"
      ],
      "abstract": "Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded. In real driving conditions, camera observations are frequently corrupted by occlusion, motion blur, illumination change, and sensor noise, and when such degraded observations are aggregated indiscriminately over time, trajectory planning becomes unstable and collision risk rises for both the ego vehicle and surrounding road users. Recent Bird's-Eye-View (BEV) approaches unify perception and planning through a shared spatial representation, but most fuse temporal information across frames without assessing the reliability of the underlying observations. We present a Reliable Context-Aware and Temporal Planning framework for Autonomous Driving (RCT-AD) that explicitly models feature quality and temporal consistency to support safer, more consistent planning. A Reliable Context Awareness module scores per-frame reliability and selectively retains trustworthy features through a quality-gated First-In-Last-Out (FILO) memory mechanism, reconstructing degraded observations from reliable historical context so that corrupted inputs do not destabilize the scene representation. A Temporal Trajectory Planner captures long-term dependencies and multi-agent interactions to produce smoother, safety-aware trajectories, while a joint detection-and-segmentation head injects semantic and motion cues into the shared BEV space to strengthen scene understanding. Experiments on the nuScenes autonomous driving benchmark show that RCT-AD improves perception accuracy, motion prediction, and planning robustness over recent end-to-end baselines, achieving 61.5 nuScenes Detection Score, 52.9 mean Average Precision, and 52.3 mean Intersection over Union, while maintaining competitive computational efficiency suitable for real-time deployment.",
      "short_abstract": "Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded. In real driving conditions, camera observations are frequently corrupted by occlusion, motion blur,...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 7,
        "evidence": [
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04689",
      "pdf_url": "https://arxiv.org/pdf/2607.04689.pdf"
    },
    {
      "id": "2607.04681",
      "anchor": "paper-2607-04681",
      "title": "Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning",
      "authors": [
        "Matthew Foutter",
        "Matteo Cercola",
        "Lena Wild",
        "Yunshan Wang",
        "Michelle Li",
        "Daniele Gammelli",
        "Marco Pavone"
      ],
      "abstract": "Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models. However, whether this verbalized Chain-of-Thought truthfully reflects the policy's underlying decision process remains poorly understood. We distinguish between functional reasoning, in which reasoning improves task performance, and faithful reasoning, in which reasoning truly reflects the policy's internal decision process. We argue that SoTA alignment strategies offer a necessary but insufficient notion of faithfulness, admitting reasoning whose intermediate steps can mask the causal links in action prediction through confounding factors (e.g., reasoning that is ungrounded in the environment and internally disconnected or inconsistent), restricting policy generalization. We study this gap through a human evaluation of a SoTA reasoning model for autonomous driving, revealing an inconsistent coupling between reasoning quality and downstream trajectory improvement. We then operationalize a behavioral surrogate for embodied faithfulness through a learned critic, Pinocchio, scoring observation grounding and stepwise coherence, and use this critic as a dense reward signal in post-training an embodied policy with reinforcement learning. Across withheld driving benchmarks, our post-trained planner improves faithfulness by 4% and 18% over SoTA alignment and trajectory error post-training baselines, respectively, while maintaining competitive downstream task performance. Finally, on a synthetic out-of-distribution test set, post-training for faithfulness improves policy responsiveness to rare counterfactual scenarios by 1.6x that of a SoTA policy, suggesting that faithful reasoning traces contribute to more robust, generalizable, and interpretable embodied intelligence. Project page: https://mjf-su.github.io/pinocchio/",
      "short_abstract": "Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models. However, whether this verbalized Chain-of-Thought truthfully reflects the policy's...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 20,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04681",
      "pdf_url": "https://arxiv.org/pdf/2607.04681.pdf"
    },
    {
      "id": "2607.04661",
      "anchor": "paper-2607-04661",
      "title": "Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving",
      "authors": [
        "Guoqing Wang",
        "Pin Tang",
        "Xiangxuan Ren",
        "Liping Hou",
        "Chao Ma"
      ],
      "abstract": "Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent state-of-the-art frameworks achieve promising results by incorporating voxel-based Gaussians, but incur substantial computational redundancy due to a uniform volumetric processing strategy. To bridge the gap between the efficiency of pixel-based Gaussian methods and the structural completeness of voxel-based Gaussian approaches, we propose FocusGS, a simple yet effective framework that shifts the paradigm from global densification to targeted structural completion. Our central insight is that structural completion should be decoupled from deterministic regions, with computation concentrated exclusively on areas exhibiting geometric ambiguity. Specifically, FocusGS addresses the localization challenge by deriving a 3D Geometric Ambiguity Manifold to accurately isolate localized areas prone to occlusion and high geometric uncertainty. To overcome the subsequent manifold completion challenge, we design a lightweight targeted structure completion module that selectively instantiates and optimizes continuous Gaussian queries strictly within this unstructured, sparse topological subspace. Extensive experiments demonstrate that FocusGS achieves a superior efficiency-quality trade-off, advancing state-of-the-art performance on driving-centric benchmarks while naturally reducing the total number of Gaussians by ~74% and decreasing rendering time by ~34%.",
      "short_abstract": "Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent state-of-the-art frameworks achieve promising results by incorporating voxel-based Gaussians, but incur substantial...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Mapping & Localization: abstract: localization",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04661",
      "pdf_url": "https://arxiv.org/pdf/2607.04661.pdf"
    },
    {
      "id": "2607.05649",
      "anchor": "paper-2607-05649",
      "title": "REVIVE: A Multi-Modal Framework for Vandalism Detection and Recovery in Autonomous Vehicles",
      "authors": [
        "Abdullah Tariq Choudhry",
        "Tapadhir Das"
      ],
      "abstract": "Autonomous vehicles (AVs) face increasing threats from vandalism-induced occlusion attacks (VOAs) that compromise camera-based perception. While detection frameworks can identify vandalized images, restoring camera-stream utility after physical occlusion remains underexplored. This paper presents present the Recovery and Enhancement of Vandalized Images for Vision Excellence (REVIVE) framework, a vandalism recovery pipeline integrating: (1) binary VOA detection, (2) multi-class VOA pattern identification, (3) EfficientNet-based U-Net segmentation, and (4) type-aware recovery using Bootstrapping Language-Image Pre-training (BLIP)-guided Stable Diffusion inpainting, direct pixel replacement, or adaptive median filtering. Stable Diffusion shows variable reconstruction performance (per-pattern SSIM 0.667-0.867, PSNR 15.4-26.7dB) across VOA patterns, while aligned direct pixel replacement achieves near-identical reconstruction under the aligned-reference condition. On 500 tracked clean/vandalized image pairs, unrecovered VOAs reduce YOLOv8l object-detection recall to 0.588, while direct pixel replacement restores recall to 0.967 and F1-score to 0.970 under that aligned-reference condition. LaMa, Telea, and Navier-Stokes baselines improve image similarity but provide more limited downstream detection recovery, and Stable Diffusion is treated as an asynchronous recovery branch subject to a quality gate rather than a blocking real-time perception step. We evaluate a reference-available quality gate that filters recovered candidates before downstream use: without it, type-aware routing degrades per-image recall to 0.304, whereas with it, recall returns to 0.608, at or above the unrecovered baseline, ensuring the forwarded stream is never worse than the unrecovered frame. REVIVE therefore, provides a structured recovery framework from VOAs in AVs.",
      "short_abstract": "Autonomous vehicles (AVs) face increasing threats from vandalism-induced occlusion attacks (VOAs) that compromise camera-based perception. While detection frameworks can identify vandalized images, restoring camera-stream utility after physical occlusion...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: camera",
          "Safety, Security & Verification: abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.05649",
      "pdf_url": "https://arxiv.org/pdf/2607.05649.pdf"
    },
    {
      "id": "2607.04670",
      "anchor": "paper-2607-04670",
      "title": "Who Responds When the Driver Is Gone? A Framework for Holistic Passenger Intent Understanding",
      "authors": [
        "Xuewen Luo",
        "Ding Fan",
        "Ruiqi Chen",
        "Ye Cao",
        "Xiujin Liu",
        "Bo Yu",
        "Fengze Yang",
        "Chenxi Liu"
      ],
      "abstract": "As autonomous vehicles advance toward driverless mobility, understanding and responding to passenger needs and intentions becomes increasingly important in the absence of a human driver. We propose Intent2Drive, a unified framework for holistic passenger intent understanding and passenger-aligned planning. Unlike existing methods that rely on explicit commands, Intent2Drive models passenger intent as a latent cognitive state inferred from language, personal attributes, emotions, behaviors, and situational context. To support this task, we construct the Holistic Passenger Intent Dataset (HPID) with structured annotations of explicit and implicit passenger-intent cues. A Theory-of-Mind-inspired Passenger Intent Reasoner (PIR) infers a Latent Passenger State (LPS) and converts it into a planner-compatible Passenger Intent Objective (PIO). We validate the downstream utility of PIO by conditioning an existing hierarchical planning pipeline at the route and trajectory levels. Experiments demonstrate that the proposed method understands and responds to passenger needs, enabling passenger-aligned driving while maintaining competitive closed-loop planning performance.",
      "short_abstract": "As autonomous vehicles advance toward driverless mobility, understanding and responding to passenger needs and intentions becomes increasingly important in the absence of a human driver. We propose Intent2Drive, a unified framework for holistic passenger...",
      "published": "2026-07-06",
      "updated": "2026-07-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: abstract: human driver",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 44,
      "arxiv_url": "https://arxiv.org/abs/2607.04670",
      "pdf_url": "https://arxiv.org/pdf/2607.04670.pdf"
    },
    {
      "id": "2607.04637",
      "anchor": "paper-2607-04637",
      "title": "PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving",
      "authors": [
        "Pin Tang",
        "Guoqing Wang",
        "Xiangxuan Ren",
        "Zhongdao Wang",
        "Guodongfang Zhao",
        "Bailan",
        "Chao Ma"
      ],
      "abstract": "Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous driving scenarios. Existing VLAs typically predict and optimize 3D trajectories from 2D images. While intuitive, this 2D-to-3D prediction is inherently entangled with camera parameters, leading to limited data scalability across heterogeneous driving datasets. Moreover, directly optimizing in 3D space induces severe convergence to trivial solutions, where VLAs rely on ego-status rather than visual scene understanding. To address these issues, we propose PixelPilot, a novel VLA featuring a decoupled planning and lifting paradigm. In the planning phase, PixelPilot reformulates scene understanding and trajectory prediction as sensor-agnostic 2D-to-2D tasks in the image plane, thereby facilitating scalable training across diverse datasets. The planned 2D trajectories are then deterministically lifted to 3D only during inference, ensuring the full exploitation of visual cues and generalization across different vehicles. To realize this paradigm, we propose a knowledge-instilled policy learning strategy that applies dense, intermediate rewards via Group Relative Policy Optimization (GRPO) to enforce a rigorous causal chain from visual perception to spatial planning. Extensive experiments demonstrate that PixelPilot achieves state-of-the-art performance in both open-loop and closed-loop settings, validating its superior scalability and visual reasoning capabilities.",
      "short_abstract": "Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous driving scenarios. Existing VLAs typically predict and optimize 3D trajectories from...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 51,
        "margin": 43,
        "evidence": [
          "title: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "title: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04637",
      "pdf_url": "https://arxiv.org/pdf/2607.04637.pdf"
    },
    {
      "id": "2607.04610",
      "anchor": "paper-2607-04610",
      "title": "RoboVista: Evaluating Vision Language Models for Diverse Robot Applications",
      "authors": [
        "Shuangyu Xie",
        "Kaiyuan Chen",
        "Ziyang Chen",
        "Simeon Adebola",
        "Yixuan Huang",
        "Zehan Ma",
        "Tianshuang Qiu",
        "Wentao Yuan",
        "Dhruv Shah",
        "Pannag R. Sanketi",
        "Ken Goldberg"
      ],
      "abstract": "Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-Language Models (VLMs) offer a promising foundation for general-purpose and interpretable robotic reasoning. Aligning VLMs with diverse robot applications requires a modular understanding of the individual decision components that underlie robotic behavior. Capturing such structure is challenging for conventional robot benchmarks that are primarily based on teleoperated, end-to-end datasets. We propose Robot Question Answering (RQA), a modular evaluation framework and RoboVista, a benchmark curated from real robotic systems, research papers, and expert annotations. RoboVista contains 474 Visual Question Answering (VQA) instances with human annotated reasoning and covers 39 unique task types in agricultural, industrial, domestic, surgical robotics, autonomous driving, and open robot datasets. Experiments on RoboVista show that state-of-the-art VLMs exhibit substantial gaps. Physical robot experiments suggest strong correlation between RoboVista performance and real-world task execution.",
      "short_abstract": "Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-Language Models (VLMs) offer a promising foundation for general-purpose and...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: teleoperation",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04610",
      "pdf_url": "https://arxiv.org/pdf/2607.04610.pdf"
    },
    {
      "id": "2607.04541",
      "anchor": "paper-2607-04541",
      "title": "CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining",
      "authors": [
        "Jingyu Song",
        "Yi Liu",
        "Katherine A. Skinner"
      ],
      "abstract": "Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-specific supervision, limiting reusable representation learning. We present CRISP, a spatiotemporal CR backbone pretrained through forecasting-based representation learning. Given historical multi-view images and radar sweeps, CRISP learns a unified bird's-eye-view (BEV) representation by predicting future LiDAR point clouds. LiDAR is used only as privileged supervision during pretraining; the deployed model requires only camera and radar. To make forecasting-based pretraining effective for CR fusion, CRISP introduces an enhanced radar encoder, radar-enhanced temporal self-attention, and multimodal feature rendering with modality innovation gating. These components inject radar range and Doppler cues into BEV temporal propagation and allow BEV tokens to selectively incorporate camera and radar evidence. Experiments on nuScenes show that CRISP improves long-horizon point cloud forecasting and transfers effectively to downstream tasks, including 3D detection, tracking, online mapping, motion forecasting, future occupancy prediction, and planning, suggesting that predictive CR pretraining is a promising path toward scalable driving representations under practical sensor configurations. The project website is https://umfieldrobotics.github.io/CRISP.",
      "short_abstract": "Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-specific supervision, limiting reusable representation learning. We present CRISP, a spatiotemporal CR backbone...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 11,
        "evidence": [
          "abstract: occupancy",
          "abstract: point cloud",
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04541",
      "pdf_url": "https://arxiv.org/pdf/2607.04541.pdf"
    },
    {
      "id": "2607.04331",
      "anchor": "paper-2607-04331",
      "title": "Agent-driven Long-tail Simulation for Autonomous Driving",
      "authors": [
        "Junru Gu",
        "Lijin Yang",
        "Jianing Huang",
        "Shu Liu",
        "Zhongzhan Huang",
        "Hang Zhao"
      ],
      "abstract": "Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on log replay or rule-based agents, limiting behavioral diversity and long-tail coverage. We propose an agent-driven simulation framework in which surrounding road participants are controlled by instruction-following large language models through a structured action interface, enabling intentional and reactive behaviors while preserving physical plausibility. Furthermore, we introduce SemanticPlan, a benchmark of closed-loop planning in long-tail and semantically rich scenarios that augment real nuPlan scenes with multiple interactive agents following diverse language instructions. Evaluation results show that state-of-the-art planners still struggle to consistently achieve safe and effective task completion, suggesting that these long-tail scenarios remain challenging.",
      "short_abstract": "Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on log replay or rule-based agents, limiting behavioral diversity and long-tail coverage. We propose an...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04331",
      "pdf_url": "https://arxiv.org/pdf/2607.04331.pdf"
    },
    {
      "id": "2607.04330",
      "anchor": "paper-2607-04330",
      "title": "Framework and Multi-modal Dataset for Roadwork Zone Detection and Geo-localization",
      "authors": [
        "Zhiran Yan",
        "Yutong Xin",
        "S Shyam Shenoi",
        "Rui Song",
        "Gordon Elger"
      ],
      "abstract": "Autonomous vehicles often rely on high-definition (HD) maps for navigation; however, these maps are not frequently updated and often lack semi-static information, such as temporary roadwork zones, which can significantly alter the road network. This limitation underscores the urgent need for an accurate global position of roadwork zones. However, the absence of publicly available datasets for evaluating roadwork zone detection and geo-localization models has hindered the development of reliable autonomous driving systems. To address this challenge, we propose the Roadwork Zone Detection and Geo-localization (RZDG) dataset, which includes both simulated and real-world data, providing multimodal sensor inputs along with comprehensive annotations. The dataset supports multiple perception tasks, including image semantic segmentation, 3D object detection, and object geo-localization. In addition, we introduce a tracker-based roadwork zone detection and geo-localization (RZDG) pipeline, an extension of AB3DMOT, for accurate object geo-localization in roadwork zones. We benchmark our approach on the RZDG dataset, demonstrating its effectiveness in detecting roadwork zones and transforming object positions from the local coordinate system to the global coordinate system. A prediction is considered a true positive (TP) if its estimated position falls within one meter of the ground truth. Our experimental results show that our approach achieves high accuracy on both real and simulated data. Specifically, we report: Precision: 0.565 (real) / 0.615 (simulated) Recall: 0.898 (real) / 0.809 (simulated) F1-score: 0.597 (real) / 0.665 (simulated).",
      "short_abstract": "Autonomous vehicles often rely on high-definition (HD) maps for navigation; however, these maps are not frequently updated and often lack semi-static information, such as temporary roadwork zones, which can significantly alter the road network. This...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 9,
        "evidence": [
          "title: localization",
          "abstract: localization"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04330",
      "pdf_url": "https://arxiv.org/pdf/2607.04330.pdf"
    },
    {
      "id": "2607.04304",
      "anchor": "paper-2607-04304",
      "title": "Road-Aware Anomaly Segmentation with Query-Guided Polygons and CLIP in Autonomous Driving",
      "authors": [
        "Zhiran Yan",
        "Gordon Elger"
      ],
      "abstract": "Traditional semantic segmentation models operate under a closed-set assumption and struggle to recognize unknown or unexpected objects-an essential capability for autonomous driving. As a result, such models often misclassify or overlook out-of-distribution (OOD) road anomalies, posing safety risks in open-world environments. We present a lightweight, postprocessing, road-aware anomaly segmentation framework that requires no retraining, no OOD data, and no auxiliary supervision. Our approach builds on a mask transformer-based segmentation network by exploiting query-level mask confidence and deriving a polygonal road prior to detect gap regions that may correspond to anomalies. To further suppress false positives, we introduce a CLIP-based zero-shot semantic filtering module using in-distribution prompts, with optional generalized OOD prompts. By jointly leveraging spatial priors and semantic verification, our framework produces robust and interpretable anomaly predictions. Evaluation on three public benchmarks-Fishyscapes, SMIYC, and RoadAnomaly-shows consistently strong performance. In particular, our method outperforms the training-free baseline Maskomaly on most metrics and achieves the highest AP on Fishyscapes LostAndFound. These results demonstrate the practicality and deployability of our approach for real-world autonomous driving systems.",
      "short_abstract": "Traditional semantic segmentation models operate under a closed-set assumption and struggle to recognize unknown or unexpected objects-an essential capability for autonomous driving. As a result, such models often misclassify or overlook out-of-distribution...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04304",
      "pdf_url": "https://arxiv.org/pdf/2607.04304.pdf"
    },
    {
      "id": "2607.04179",
      "anchor": "paper-2607-04179",
      "title": "CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving",
      "authors": [
        "Zhaohong Liu",
        "Hao Ye",
        "Xianlin Zhang",
        "Mengshi Qi"
      ],
      "abstract": "End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers from reasoning hallucinations and conservative biases. While traditional tool-augmented frameworks and Chain-of-Thought (CoT) approaches mitigate these issues, they incur exorbitant token consumption and unacceptable latency, rendering real-time deployment impractical. To resolve this reliability-efficiency trade-off, we propose CritiqueDriveVLM, a novel unified three-stage framework internalizing reasoning directly into the VLM. First, we introduce Critique-Driven Multi-Turn Reinforcement Learning (RL) guided by a multi-dimensional verifier. By providing granular scalar feedback and a multi-turn penalty, we force the policy to internalize logical deduction, cultivating a robust System-2 Teacher that achieves high accuracy without fragile external tools. Subsequently, we propose Latent Thought Distillation to overcome the latency bottleneck. By aligning the Student's latent representations with the Teacher's fully converged reasoning states, we compress deep logical capabilities into a fast, CoT-free System-1 Student. Extensive experiments on the widely-used DriveLMM-01 benchmark demonstrate remarkable improvements. Compared to the base model, our tool-free Teacher significantly boosts Multiple Choice Quality (MCQ) from 55.54% to a state-of-the-art 76.54%. Crucially, our distilled Student preserves competitive reasoning depth while drastically minimizing generation length to an average of merely 28 tokens. This slashes inference latency by 88% (from 3482 ms to 416 ms), paving a highly robust pathway for low-latency autonomous driving.Our source code is available at https://github.com/MICLAB-BUPT/CritiqueDriveVLM.",
      "short_abstract": "End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers from reasoning hallucinations and conservative biases. While traditional tool-augmented frameworks and...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: systems constraints",
          "Human Factors & Policy: abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04179",
      "pdf_url": "https://arxiv.org/pdf/2607.04179.pdf"
    },
    {
      "id": "2607.04129",
      "anchor": "paper-2607-04129",
      "title": "Toward the Right Analytical Model and System Software for Autonomous Driving Systems: Open Problems and Research Directions",
      "authors": [
        "Atsushi Yano",
        "Takuya Azumi"
      ],
      "abstract": "Autonomous driving (AD) systems continuously transform multi-rate and asynchronous sensor streams into vehicle actuation through graphs of callbacks, nodes, and middleware components. In such systems, temporal correctness cannot be characterized by the execution time or deadline of an individual task alone: localization and perception chains run in parallel, fuse data with different timestamps, converge at planning, and propagate through control to actuation. Moreover, the demand for high processing capability places AD systems on high-performance processors with multicore parallelism and GPU acceleration, where execution times vary strongly with the input scene, hardware state, and co-running work. Rare deadline misses at runtime therefore cannot be ruled out, and safety is preserved through fail-safe mechanisms such as the minimal-risk maneuver (MRM). This raises a two-sided question: what analytical models are needed to reason about timing in AD systems, and what system software is needed to realize, observe, and enforce those models on real platforms? On the analytical side, real-time research has evolved from periodic/sporadic tasks, directed acyclic graphs (DAGs), pipelines, mixed-criticality systems, and timer-/event-driven models toward end-to-end latency along cause-effect chains, data freshness, timing disparity, probabilistic timing, highest-criticality fail-safe operation, and early deadline-miss detection. On the system-software side, AD stacks and middleware, such as Autoware and ROS~2, expose both the opportunities and limitations of implementing analyzable timing behavior through executors, communication layers, tracing tools, and evaluation frameworks. This paper surveys these two lines of work and identifies the remaining gaps along five dimensions: units of timing constraints, timing metrics, resource models, execution-time variability, and safety integration...",
      "short_abstract": "Autonomous driving (AD) systems continuously transform multi-rate and asynchronous sensor streams into vehicle actuation through graphs of callbacks, nodes, and middleware components. In such systems, temporal correctness cannot be characterized by the...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 9,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware",
          "abstract: communications",
          "abstract: software architecture",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04129",
      "pdf_url": "https://arxiv.org/pdf/2607.04129.pdf"
    },
    {
      "id": "2607.04451",
      "anchor": "paper-2607-04451",
      "title": "CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation",
      "authors": [
        "Ke Li",
        "Kaidi Liang",
        "Yuxin Ding",
        "Debojyoti Biswas",
        "Xianbiao Hu",
        "Ruwen Qin"
      ],
      "abstract": "Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating controllable safety-critical scenarios remains challenging. Existing approaches use soft guidance that provides only probabilistic preferences and cannot guarantee the satisfaction of geometric and severity constraints associated with specific collision types. We introduce Collision-Constrained Flow Matching (CCFM), a novel framework that guarantees precise collision control through hard physical constraints. CCFM consists of three key components: (i) a heuristic collision selector that optimally identifies an adversarial agent and collision type via composite scoring; (ii) structured hard constraints that explicitly define four collision types (rear-end, side, cut-in, head-on) through contact point, heading, and severity requirements; and (iii) a collision-constrained flow matching sampler that enforces the constraints via Gauss-Newton manifold projection. CCFM achieves collision rate up to 46.4% on nuScenes and 83.1% on nuPlan, significantly outperforming baselines while preserving realistic driving behavior. By enabling controllable collision characteristics in safety-critical scenario generation, CCFM provides a reliable foundation for AV safety evaluation and sim-to-real crash data generation. The code and implementation details are available at https://github.com/KELISBU/CCFM.",
      "short_abstract": "Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating controllable safety-critical scenarios remains challenging. Existing approaches use soft guidance that...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 17,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04451",
      "pdf_url": "https://arxiv.org/pdf/2607.04451.pdf"
    },
    {
      "id": "2607.04259",
      "anchor": "paper-2607-04259",
      "title": "Integrated Graph Search and Model Predictive Control for Smooth and Efficient Path Planning in Autonomous Vehicles",
      "authors": [
        "Duc-Tien Bui",
        "Ngoc Thinh Nguyen",
        "Hung Duy Nguyen",
        "Dong Bi",
        "Tomislav Mihalj",
        "Arno Eichberger"
      ],
      "abstract": "Path planning is a fundamental component of autonomous vehicles, where achieving safe, comfortable, and dynamically feasible paths while ensuring computational efficiency remains a significant challenge. This paper presents a sequential path planning framework in which a rough path obtained from graph search is explicitly exploited to guide a Model Predictive Control (MPC)-based path refinement. A rough path is first obtained via Dijkstra search on a discretized grid and is then used to construct a spatially varying convex lateral safety corridor that explicitly captures obstacle avoidance constraints, transforming discrete obstacle avoidance decisions into continuous feasibility constraints for optimization. Within this corridor, an MPC problem is formulated to refine the path, enabling efficient optimization while maintaining path smoothness by penalizing the third-order spatial derivative of the lateral offset over a prediction horizon. The proposed algorithm is evaluated in multiple overtaking scenarios on both straight and curved roads, including cases with single and multiple target vehicles, using high-fidelity environment simulations (i.e., CarMaker). Compared with the previous study, which used polynomial fitting and a quadratic programming method, the proposed approach consistently achieves lower lateral acceleration, curvature, and jerk while reducing computational cost by 28.08% on straight roads and 29.52% on curved roads. These results demonstrate that exploiting graph-search structure within an MPC formulation provides an effective balance between path smoothness and computational efficiency for autonomous vehicles in structured driving environments.",
      "short_abstract": "Path planning is a fundamental component of autonomous vehicles, where achieving safe, comfortable, and dynamically feasible paths while ensuring computational efficiency remains a significant challenge. This paper presents a sequential path planning...",
      "published": "2026-07-05",
      "updated": "2026-07-05",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 32,
        "margin": 2,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 45,
      "arxiv_url": "https://arxiv.org/abs/2607.04259",
      "pdf_url": "https://arxiv.org/pdf/2607.04259.pdf"
    },
    {
      "id": "2607.04098",
      "anchor": "paper-2607-04098",
      "title": "Sparse4D-Radar: An Efficient and Robust Framework for Surround-View 3D Object Detection via 4D Radar-Camera Fusion",
      "authors": [
        "Fuyuan Ai",
        "Yuchen Tan",
        "Jiehui Chen",
        "Zhiwei Xu",
        "Chunyi Song"
      ],
      "abstract": "In recent years, 4D imaging radar has gained wide attention in autonomous driving for its robustness against harsh weather and ability to output target velocity. Nevertheless, mainstream 4D radar-camera fusion methods only support front-view perception, lacking mature solutions for surround-view sensing. Directly expanding these pipelines to full 360° coverage introduces excessive computation cost and limits real-world deployment. To tackle these limitations, this work proposes Sparse4D-Radar, an efficient robust surround-view multi-modal fusion framework. We first design a Deformable Fusion module to embed radar-camera features into sparse queries, constructing the lightweight base version Sparse4D-Radar-Base. Two dedicated modules are further introduced to boost localization accuracy and modality stability: Velocity-Consistency Sampling (VCS) refines features via radar velocity cues for motion awareness, and Adaptive Modality Gating (AMG) dynamically adjusts cross-modal fusion weights according to feature confidence. Combining all components, we build Sparse4D-Radar-Acc for high-precision detection demands. Comprehensive experiments on OmniHD-Scenes verify that our approach achieves state-of-the-art surround-view 3D detection performance. Compared with prior arts, our method obtains over 7% mAP and 10% ODS improvements under complex driving scenes while running at nearly 10 FPS, striking a favorable trade-off among detection accuracy, environmental robustness and inference efficiency. Our open-source code is available at https://github.com/Aiuan/Sparse4D-Radar.",
      "short_abstract": "In recent years, 4D imaging radar has gained wide attention in autonomous driving for its robustness against harsh weather and ability to output target velocity. Nevertheless, mainstream 4D radar-camera fusion methods only support front-view perception,...",
      "published": "2026-07-04",
      "updated": "2026-07-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 27,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 46,
      "arxiv_url": "https://arxiv.org/abs/2607.04098",
      "pdf_url": "https://arxiv.org/pdf/2607.04098.pdf"
    },
    {
      "id": "2607.03755",
      "anchor": "paper-2607-03755",
      "title": "EvoEye: Self-Evolving Runtime Monitoring for Autonomous Driving Systems",
      "authors": [
        "Mingfei Cheng",
        "Lionel Briand",
        "Xiaofei Xie"
      ],
      "abstract": "Runtime monitoring is essential for detecting impending hazards in autonomous driving systems (ADSs). However, existing ADS runtime monitors have fixed detection capabilities: rule-based monitors cover only manually specified hazards, while learning-based monitors depend heavily on their initial training data and may retain substantial prediction errors. We therefore propose EvoEye, which identifies the current monitor's errors, generates informative executions accordingly, and updates the monitor through self-evolution. To enable effective self-evolution, EvoEye combines a capable runtime monitor with targeted scenario acquisition. FusionMonitor learns cross-module temporal interactions for collision prediction, while BlindSpotEvolver converts current prediction errors into search guidance and uses density-aware mutation to acquire informative executions for subsequent monitor updates. We evaluate EvoEye on Baidu Apollo with CARLA in representative highway and urban scenarios. FusionMonitor improves frame-level Recall by up to 37.8 percentage points at a false positive rate of 0.05, with 2.49 ms latency and 2.8-4.2 seconds of median warning time. Under the same budget, BlindSpotEvolver outperforms uniform and violation-oriented sampling by up to 13.2 F1 points on previously missed unsafe contexts.",
      "short_abstract": "Runtime monitoring is essential for detecting impending hazards in autonomous driving systems (ADSs). However, existing ADS runtime monitors have fixed detection capabilities: rule-based monitors cover only manually specified hazards, while learning-based...",
      "published": "2026-07-04",
      "updated": "2026-07-04",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Prediction & World Models: abstract: prediction",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 46,
      "arxiv_url": "https://arxiv.org/abs/2607.03755",
      "pdf_url": "https://arxiv.org/pdf/2607.03755.pdf"
    },
    {
      "id": "2608.14603",
      "anchor": "paper-2608-14603",
      "title": "Extend the Safety Horizon for Intelligent Transportation Systems through Semantic-Aware Cooperative Perception",
      "authors": [
        "Chun-Yeow Yeoh",
        "Chee Keong Tan",
        "Joanne Mun-Yee Lim",
        "Heng-Siong Lim"
      ],
      "abstract": "Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending sensing coverage beyond occlusions and mitigating blind spots. While critical for autonomous driving and safety, practical deployments often rely on bandwidth-efficient late fusion. Recently, intermediate fusion has emerged as a promising approach for an optimal bandwidth-accuracy trade-off. However, in dense urban environments, cumulative bandwidth demands can overwhelm network capacity, potentially compromising safety-critical Cooperative Intelligent Transport Systems (C-ITS) functions. To alleviate these problems, this paper proposes Hierarchical Multi-Scale Semantic-Aware Cooperative Perception (HMS-SCP), a robust noise-resilient and bandwidth-efficient framework for task-oriented semantic communication in cooperative perception. HMS-SCP employs a spatial importance predictor to identify task-relevant grid elements at each scale, which are then directly mapped into complex-valued symbols for Joint Source-Channel Coding (JSCC). Unlike prior methods that rely on high-dimensional symbol projections for robustness, HMS-SCP exploits structural semantic redundancy across multiple scales to enhance resilience against channel noise, while maintaining an ultra-low symbol rate. This design significantly reduces bandwidth consumption and mitigates network congestion in high-density vehicular environments. Extensive evaluations on the simulated OPV2V and real-world DAIR-V2X datasets demonstrate that HMS-SCP effectively prevents performance collapse under severe Rayleigh fading and extreme compression ratio, maintaining high-confidence far-field detection with a real-time latency of below 16~ms, well within the safety-critical thresholds for dynamic V2X environments.",
      "short_abstract": "Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending sensing coverage beyond occlusions and mitigating blind spots. While critical for autonomous driving and safety,...",
      "published": "2026-07-03",
      "updated": "2026-07-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 7,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 47,
      "arxiv_url": "https://arxiv.org/abs/2608.14603",
      "pdf_url": "https://arxiv.org/pdf/2608.14603.pdf"
    },
    {
      "id": "2607.09741",
      "anchor": "paper-2607-09741",
      "title": "SWIFT: A Small-World Interaction Framework for Flow-Aware Trajectory Prediction in Autonomous Driving",
      "authors": [
        "Chengyue Wang",
        "Bin Rao",
        "Haicheng Liao",
        "Bonan Wang",
        "Chengzhong Xu",
        "Zhenning Li"
      ],
      "abstract": "Accurate trajectory prediction in autonomous driving hinges on modeling dynamic and context-dependent interactions among traffic agents. However, most existing approaches are purely data-driven and lack structural priors, which limits their generalization under distribution shifts. In this work, interaction modeling is revisited through the structure and dynamics of traffic networks, and SWIFT (Small-World Interaction Framework for Trajectory prediction) is proposed as a unified framework that integrates small-world networks with traffic flow theory. SWIFT introduces structural inductive biases via a Small-World Interaction Network that captures both local and global dependencies, and a Flow Regime Encoder that adapts the interaction structure to scene-level traffic states. Interaction reasoning is further enhanced through a multi-relational graph module that explicitly encodes direct and higher-order agent relationships. Extensive experiments on three real-world datasets, nuScenes, MoCAD, and NGSIM, show that SWIFT consistently outperforms strong baselines in prediction accuracy across diverse traffic regimes. Beyond accuracy, SWIFT exhibits improved generalization to unseen locations and regimes, robustness under noisy observations, and strong performance with limited training data, supporting the effectiveness of its structure-aware design.",
      "short_abstract": "Accurate trajectory prediction in autonomous driving hinges on modeling dynamic and context-dependent interactions among traffic agents. However, most existing approaches are purely data-driven and lack structural priors, which limits their generalization...",
      "published": "2026-07-03",
      "updated": "2026-07-03",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 26,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 47,
      "arxiv_url": "https://arxiv.org/abs/2607.09741",
      "pdf_url": "https://arxiv.org/pdf/2607.09741.pdf"
    },
    {
      "id": "2607.03182",
      "anchor": "paper-2607-03182",
      "title": "AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning",
      "authors": [
        "Qi Liu",
        "Yabei Li",
        "Hongsong Wang",
        "Heng Zhang",
        "Lei He"
      ],
      "abstract": "Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable continuous trajectories. Vision-Language-Action models have been introduced into driving planning to improve long-tail generalization, commonsense reasoning, high-level semantic understanding, and explainability. However, existing VLA planners mainly follow planning-head-based trajectory prediction or full-trajectory autoregressive generation. The former only weakly constrains continuous trajectory generation with VLA reasoning, while the latter relies on long sequences of low-information-density coordinate tokens, making semantic-action alignment difficult and leading to discretization errors and inefficient inference. To address these limitations, we propose AnchorVLA, a hierarchical decision-anchored VLA planning framework that uses trajectory-pattern anchors as an explicit interface between high-level VLA reasoning and continuous trajectory execution. Specifically, Decision-as-Anchor Representation represents behavior-level driving decisions with anchor tokens, each encoding an entire local motion pattern rather than a single coordinate point. Decision-Anchored Residual Flow then generates fine-grained continuous trajectories in the selected anchor-defined residual space, capturing multi-modal execution refinements after high-level decision making. By reasoning over compact and semantically meaningful anchors instead of autoregressively generating waypoint sequences, AnchorVLA preserves LLM-based decision making while improving inference efficiency, semantic-action alignment, and continuous generation flexibility. Experiments on the Bench2Drive closed-loop benchmark show that AnchorVLA achieves a state-of-the-art Success Rate of 77.28 and a competitive Driving Score of 89.92.",
      "short_abstract": "Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable continuous trajectories. Vision-Language-Action models have been introduced into driving planning to improve...",
      "published": "2026-07-03",
      "updated": "2026-07-03",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 24,
        "margin": 3,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 47,
      "arxiv_url": "https://arxiv.org/abs/2607.03182",
      "pdf_url": "https://arxiv.org/pdf/2607.03182.pdf"
    },
    {
      "id": "2607.03554",
      "anchor": "paper-2607-03554",
      "title": "Interception-Driven Inverse Reachability for Engagement Zone Construction",
      "authors": [
        "Grant Stagg",
        "Cameron K. Peterson",
        "Alexander Von Moll",
        "Isaac Weintraub"
      ],
      "abstract": "In contested environments, autonomous vehicles may need to plan around adversarial pursuers whose launch locations are unknown. This paper presents an interception-driven inverse-reachability framework for inferring a feasible pursuer launch region directly from observed interception events for a single pursuer. Each interception induces a geometric constraint on the unknown launch location, and intersecting these constraints yields a bounded set guaranteed to contain the true origin under maximum-capability assumptions. Mapping this inferred set through the pursuer reachable region produces deterministic engagement zones with an explicit worst-case safety interpretation. A probabilistic extension models uncertainty in the pursuer launch location and yields graded engagement-risk fields for risk-aware planning. To accelerate localization, we introduce an information-driven planner for sacrificial agents that selects trajectories to maximize expected contraction of the feasible launch region. Monte Carlo simulations show that the proposed framework rapidly reduces launch-location uncertainty and enables substantially shorter safe trajectories after only a small number of sacrificial deployments.",
      "short_abstract": "In contested environments, autonomous vehicles may need to plan around adversarial pursuers whose launch locations are unknown. This paper presents an interception-driven inverse-reachability framework for inferring a feasible pursuer launch region directly...",
      "published": "2026-07-03",
      "updated": "2026-07-03",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: risk assessment",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 47,
      "arxiv_url": "https://arxiv.org/abs/2607.03554",
      "pdf_url": "https://arxiv.org/pdf/2607.03554.pdf"
    },
    {
      "id": "2607.03247",
      "anchor": "paper-2607-03247",
      "title": "Learning to Suppress SPAD-based LiDAR Flare",
      "authors": [
        "Xuanya Zhu",
        "Linghao Shen"
      ],
      "abstract": "Single-Photon Avalanche Diode (SPAD)-based Light Detection and Ranging (LiDAR) is emerging for autonomous vehicles due to its high sensitivity and precise depth sensing capabilities. However, flare caused by excessive photon returns or pile-up effects can lead to incorrect depth estimation and exaggerated boundaries in point clouds, resulting in severe distortions of geometric measurements, making flare suppression essential for safety-critical applications. Existing flare mitigation methods primarily operate at the hardware or signal-processing levels. While effective under specific configurations, they are largely rule-based and configuration-dependent, lacking learnable representations that generalize across diverse sensing scenarios. In this work, we reformulate flare suppression as a semantic segmentation problem, enabling data-driven learning of geometric and photometric cues directly from SPAD measurements. We first benchmark representative segmentation models on the newly introduced SPAD flare dataset and observe that they struggle to exploit the intrinsic multi-echo characteristics of SPAD signals. Motivated by this observation, we propose Physically-Informed segmentation for LiDAR Flare (PILF), a learning-based approach that treats the first and second echoes, together with ambient illumination, as distinct modalities, aggregating cross-echo information while jointly encoding geometric and photometric features. Experiments across multiple real-world scenes demonstrate that PILF significantly outperforms compared segmentation models, achieving up to 79.32% mIoU, and providing an effective solution for SPAD-based LiDAR flare suppression.",
      "short_abstract": "Single-Photon Avalanche Diode (SPAD)-based Light Detection and Ranging (LiDAR) is emerging for autonomous vehicles due to its high sensitivity and precise depth sensing capabilities. However, flare caused by excessive photon returns or pile-up effects can...",
      "published": "2026-07-03",
      "updated": "2026-07-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 18,
        "evidence": [
          "abstract: scene segmentation",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: depth estimation"
        ]
      },
      "recency": "fresh",
      "age_days": 47,
      "arxiv_url": "https://arxiv.org/abs/2607.03247",
      "pdf_url": "https://arxiv.org/pdf/2607.03247.pdf"
    },
    {
      "id": "2607.03155",
      "anchor": "paper-2607-03155",
      "title": "Hope for the Best, Prepare for the Worst: Occlusion-Aware Contingency Planning for Autonomous Vehicles",
      "authors": [
        "Truls Nyberg",
        "Anna Gautier",
        "Jana Tumova"
      ],
      "abstract": "The deployment of autonomous vehicles in urban environments introduces significant safety challenges, particularly in scenarios with occlusions, where critical traffic participants may be hidden from view. Recent accidents involving driverless vehicles highlight the importance of motion planners that explicitly addresses the risks posed by occlusions. In this work, we propose a formal, occlusion-aware trajectory planning framework that guarantees collision avoidance even when there are possible hidden traffic participants. Building on our previous methods that apply reachability analysis to sequentially determine the possible states of hidden traffic participants, we integrate a tree-based motion planner capable of reasoning over future observations and the absence thereof. This approach reduces conservativeness while maintaining safety guarantees. We demonstrate the effectiveness of our framework in a challenging simulated occluded scenario, showing that it pro-actively and efficiently guarantees collision-avoidance.",
      "short_abstract": "The deployment of autonomous vehicles in urban environments introduces significant safety challenges, particularly in scenarios with occlusions, where critical traffic participants may be hidden from view. Recent accidents involving driverless vehicles...",
      "published": "2026-07-03",
      "updated": "2026-07-03",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 13,
        "evidence": [
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 47,
      "arxiv_url": "https://arxiv.org/abs/2607.03155",
      "pdf_url": "https://arxiv.org/pdf/2607.03155.pdf"
    },
    {
      "id": "2607.02841",
      "anchor": "paper-2607-02841",
      "title": "CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving",
      "authors": [
        "Yunxiao Shi",
        "Hong Cai",
        "Mohammad Ghavamzadeh",
        "Fatih Porikli"
      ],
      "abstract": "End-to-end autonomous driving (E2E-AD) aims to directly map raw sensor information to driving actions. Recently, with the rapid advancement of multi-modal large language models (MLLMs), researchers have proposed the paradigm of Vision-Language-Action (VLA) models for E2E-AD, where it seeks to integrate visual perception, language understanding and action prediction within a single policy. However, existing VLA-based policies largely adopts imitation learning, where it only learns to drive by optimizing distance-based metrics w.r.t. logged expert trajectories. Such distribution shift between open-loop training and closed-loop inference leads to suboptimal performance in closed-loop planning. To close this gap, we present CLEAR, a system that enables closed-loop training using Reinforcement Learning (RL) at scale for E2E-AD. We propose to learn a novel residual waypoint policy around the waypoint prior from pretrained VLA policies, effectively harnessing the knowledge within. On another front, one of the key challenges to scale up RL for vision-based policies is the number of parallel simulation environments since RL is data hungry. To that end, we design a heterogeneous pipeline that places the simulator and the VLA learner on distinct compute groups, which allows us to dramatically increase the number of simulation environments running in parallel while avoiding resource contention and maintaining training stability. We show that with a simple reward, CLEAR significantly outperforms previous methods and sets new state-of-the-art performance on the challenging benchmarks of CARLA longest6 v2 and Bench2Drive.",
      "short_abstract": "End-to-end autonomous driving (E2E-AD) aims to directly map raw sensor information to driving actions. Recently, with the rapid advancement of multi-modal large language models (MLLMs), researchers have proposed the paradigm of Vision-Language-Action (VLA)...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation",
        "VLA",
        "Reinforcement Learning",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 38,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.02841",
      "pdf_url": "https://arxiv.org/pdf/2607.02841.pdf"
    },
    {
      "id": "2607.02494",
      "anchor": "paper-2607-02494",
      "title": "Towards Robustness against Typographic Attack with Training-free Concept Localization",
      "authors": [
        "Bohan Liu",
        "Wenqian Ye",
        "Guangzhi Xiong",
        "Zhenghao He",
        "Sanchit Sinha",
        "Aidong Zhang"
      ],
      "abstract": "Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, commonly described as a Typographic Attack (TA), exposes a vulnerability that poses a significant risk to safety-critical applications such as autonomous driving. To achieve interpretable and effective robustness against TA, we propose a novel, training-free mechanistic interpretability method. Our method provides sampling-based interpretations of hidden state representations and quantitatively attributes semantic versus lexical focus to individual attention heads. Through probabilistic analysis and circuit mining, we isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information, thereby identifying the mechanistic source of TA. We further show that simple interventions applied directly to the identified circuits, without any additional training, can substantially improve robustness against Typographic Attacks in object classification. These interventions, such as selective adjustment of attention weights, also outperform both supervised and training-free defense methods. Our experiments demonstrate that applying the proposed intervention to the vision encoders of several state-of-the-art LVLMs yields substantial gains in Visual Question Answering accuracy under Typographic Attack interference on RIO-Bench. These results confirm both the efficacy and the generalizability of our mechanistic approach. Code is released at https://github.com/Liu-524/SamplingTAR.",
      "short_abstract": "Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode:...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 15,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "title: robustness",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.02494",
      "pdf_url": "https://arxiv.org/pdf/2607.02494.pdf"
    },
    {
      "id": "2607.02074",
      "anchor": "paper-2607-02074",
      "title": "Comprehensive Robustness Analysis of LiDAR-based 3D Object Detection in Autonomous Driving",
      "authors": [
        "Adwait Chandorkar",
        "Kai Krink",
        "Yerdana Maulenbay",
        "Hasan Tercan",
        "Tobias Meisen"
      ],
      "abstract": "Recent advancements in LiDAR-only 3D object detection have demonstrated improved detection accuracy over benchmark datasets. However, the adversarial robustness of these models remains untested. Very few adversarial robustness studies exist for LiDAR-only 3D object detection and unfortunately, even they are limited to legacy models. Moreover, there is a systemic gap in the existing evaluation frameworks that rely simply on mAP ignoring other structural and predictive factors. To fill this gap, we propose a holistic framework that evaluates adversarial robustness using two structural factors (point cloud density and point cloud localization) and three predictive factors (misclassification, localization error, distance from ego). Using this framework, we perform an empirical study and critical analysis on recent and legacy state-of-the-art models using adversarial attacks specifically designed for LiDAR-based models. Our key finding is that high-capacity, voxel-based detectors are more susceptible to structured coordinate perturbations than pillar-based detectors. Additionally, non-anchor-based detectors demonstrate poor adversarial robustness, which necessitates rethinking model training techniques. Overall, our results demonstrate that recent models are as vulnerable to adversarial attacks as their predecessors. Therefore, we argue that there is a need to improve the evaluation benchmarks for 3D object detection that not only reward architectural modifications for improving detection accuracy, but also evaluate whether the design choices improve adversarial robustness.",
      "short_abstract": "Recent advancements in LiDAR-only 3D object detection have demonstrated improved detection accuracy over benchmark datasets. However, the adversarial robustness of these models remains untested. Very few adversarial robustness studies exist for LiDAR-only 3D...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 19,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.02074",
      "pdf_url": "https://arxiv.org/pdf/2607.02074.pdf"
    },
    {
      "id": "2607.01983",
      "anchor": "paper-2607-01983",
      "title": "Open-Weather Robust 3D Detection via Dual-Critic Diffusion Alignment",
      "authors": [
        "Shuyao Li",
        "Chuanxing Geng",
        "Heyang Sun",
        "Qiang Zhou",
        "Jingjing Gu"
      ],
      "abstract": "Robust 3D object detection under adverse weather remains a critical hurdle for autonomous driving. Despite progress with LiDAR-4D radar fusion, most methods are constrained by a closed-world assumption, implicitly requiring training and test weather to align in both type and severity. This premise fails in practice: the open-ended nature of weather, and even variations within a single type like rain, cause dramatically different LiDAR degradation patterns, leading to significant performance drops in unseen conditions. To address this, we present Dual-Critic Guided Diffusion Alignment (DCDA), a weather-agnostic framework that learns to recover degraded LiDAR features toward a clean manifold. Rather than modeling specific weather types, DCDA employs a 4D radar-conditioned diffusion process to progressively refine features, guided by two complementary critics. (i) A detection-guided critic, anchored by a pre-trained clean-weather model, ensures that the refined features retain object-level discriminability and localization accuracy. (ii) A weather adversarial critic enforces holistic distributional consistency with clean-weather representations. By aligning features through semantic and distributional constraints rather than explicit weather modeling, DCDA generalizes effectively to unseen weather types and severities without requiring paired data or weather labels. We further introduce a structured open-weather benchmark with held-out type-severity combinations and extensive experiments verify DCDA's advantages.",
      "short_abstract": "Robust 3D object detection under adverse weather remains a critical hurdle for autonomous driving. Despite progress with LiDAR-4D radar fusion, most methods are constrained by a closed-world assumption, implicitly requiring training and test weather to align...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 2,
        "evidence": [
          "abstract: attack or threat",
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.01983",
      "pdf_url": "https://arxiv.org/pdf/2607.01983.pdf"
    },
    {
      "id": "2607.01827",
      "anchor": "paper-2607-01827",
      "title": "C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation",
      "authors": [
        "Jinlong Wang",
        "Xun Huang",
        "Qiming Xia",
        "Shijia Zhao",
        "Chenglu Wen"
      ],
      "abstract": "LiDAR-based 3D object detection is essential for autonomous driving systems. However, traditional Ego-only Perception (Eo-Perception) suffers from limited perspective and occlusions in a complex outdoor environment, leading to performance bottlenecks. Recently, research on multi-agent Collaborative Perception (Co-Perception) has demonstrated excellent performance, but high communication costs and accumulated pose error hinder its application. To address this, we explore a novel C2E (Co-Perception to Eo-Perception) paradigm through the Multi-to-Single (M2S) agent contrastive knowledge distillation framework. Our M2S framework first designs Multi-Level Feature Enhancement module to provide more stable features, and introduces Auxiliary Point Cloud Reconstruction and Multi-Teacher Contrastive Distillation mechanisms to mitigate domain gaps in point cloud and feature distributions within the C2E paradigm. Benefiting from this, our M2S can retain the excellent performance of collaborative perception while effectively avoiding the drawbacks, such as communication delays and positioning errors. Extensive experiments on the V2XSet, V2V4Real and DAIR-V2X datasets show the effectiveness and generalizability of our M2S framework when combined with the state-of-the-art CoSDH model and other excellent 3D detectors. Our M2S framework can deliver up to a 8.64% improvement in 3D mAP performance without introducing any communication costs.",
      "short_abstract": "LiDAR-based 3D object detection is essential for autonomous driving systems. However, traditional Ego-only Perception (Eo-Perception) suffers from limited perspective and occlusions in a complex outdoor environment, leading to performance bottlenecks....",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 14,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.01827",
      "pdf_url": "https://arxiv.org/pdf/2607.01827.pdf"
    },
    {
      "id": "2607.01772",
      "anchor": "paper-2607-01772",
      "title": "LLM-Empowered Multimodal Fusion Framework for Autonomous Driving: Semantic Enhancement and Channel-Adaptive Design",
      "authors": [
        "Wen Wang",
        "Yaping Sun",
        "Yejun He",
        "Hao Chen",
        "Zhiyong Chen",
        "Xiaodong Xu",
        "Nan Ma",
        "Shuguang Cui"
      ],
      "abstract": "Vision-radar fusion is central to robust autonomous driving, combining dense visual semantics with precise range and velocity measurements from radar. However, real-world fusion quality is fundamentally challenged by dynamically varying input quality, stemming from occlusion, adverse weather, and channel noise. To address this, we re-frame the problem from static data fusion to channel-aware semantic reasoning and propose a Large Language Model-centric Semantic-layer Channel-aware Integrated Perception (LM-SCIP) framework. It places a Large Language Model (LLM) as a central reasoning core to fuse a local visual stream with a quality-varying external radar stream used to cover perception-blind spots. Concretely, LM-SCIP couples a hierarchical radar-vision encoder with a Channel-Adaptive Semantic Module (CASM) that maps link indicators into a \"Channel Prompt\" to dynamically gate external radar features. A parameter-efficient, LoRA-tuned LLM, in conjunction with a heterogeneous Mixture-of-Experts (H-MoE), then arbitrates between local visual cues and the channel-conditioned radar context. Finally, a decoupled multi-task decoder outputs localization, trajectory forecasting, and image reconstruction. Experiments on nuScenes and VIRAT validate our approach. On nuScenes, under a controlled toggle of radar input, LM-SCIP reduces localization RMSE by 40.0% versus a vision-only baseline. On VIRAT, the model attains a 0.214m localization RMSE and 0.179m minFDE (k=1). These results reveal that the proposed LM-SCIP enables a robust vision-dominant fallback at low SNR and synergistic fusion at high SNR.",
      "short_abstract": "Vision-radar fusion is central to robust autonomous driving, combining dense visual semantics with precise range and velocity measurements from radar. However, real-world fusion quality is fundamentally challenged by dynamically varying input quality,...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: perception",
          "abstract: radar"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.01772",
      "pdf_url": "https://arxiv.org/pdf/2607.01772.pdf"
    },
    {
      "id": "2607.09740",
      "anchor": "paper-2607-09740",
      "title": "A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles",
      "authors": [
        "Joshua Kofi Asamoah",
        "Blessing Agyei Kyem",
        "Eugene Denteh",
        "Armstrong Aboah"
      ],
      "abstract": "Safe motion planning in advanced driver-assistance systems and autonomous vehicles requires an accurate understanding of how the surrounding traffic scene is likely to evolve. However, many existing lane-change prediction methods remain centered on a single target vehicle, while multi-agent forecasting approaches often describe scene evolution only through future positions and provide limited explicit information about the maneuver associated with each vehicle. This study proposes a dynamic scene graph attention framework that predicts the lane-change intention and future trajectory of every relevant vehicle within a local traffic scene. The scene is represented as a time-varying interaction graph in which vehicles are modeled as nodes and their spatial and kinematic relationships are encoded through explicit edge features. Temporal graph-attention message passing captures evolving inter-vehicle dependencies and pre-maneuver cues, while an intention-guided decoder links each predicted maneuver to its corresponding future motion. A scene-level consistency objective further encourages compatible multi-vehicle futures. Experiments on the NGSIM I-80, NGSIM US-101, and highD datasets demonstrate consistent improvements over competing baselines. DSiGAT achieves intention prediction accuracies of 90.12% and 90.97% on NGSIM I-80 and US-101, respectively, and reduces trajectory RMSE by up to 52.94% relative to the strongest baseline. It also produces lower inter-agent collision rates and joint displacement errors, indicating more coherent scene-level predictions. Ablation, sensitivity, robustness, and qualitative analyses further validate the contribution of the proposed components and the effectiveness of the scene-focused formulation.",
      "short_abstract": "Safe motion planning in advanced driver-assistance systems and autonomous vehicles requires an accurate understanding of how the surrounding traffic scene is likely to evolve. However, many existing lane-change prediction methods remain centered on a single...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 22,
        "evidence": [
          "title: trajectory prediction",
          "abstract: behavior prediction",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.09740",
      "pdf_url": "https://arxiv.org/pdf/2607.09740.pdf"
    },
    {
      "id": "2607.02924",
      "anchor": "paper-2607-02924",
      "title": "Longitudinal-Motion-Aware Lateral Control for Autonomous Vehicles: A Robust Nonlinear Control Framework",
      "authors": [
        "Sixu Li",
        "Nitesh Kumar",
        "Reyshwanth Ganeshan",
        "Sivakumar Rathinam",
        "Swaroop Darbha",
        "Yang Zhou"
      ],
      "abstract": "As autonomous vehicles (AVs) operate in increasingly dynamic traffic conditions, lateral control must be performed while longitudinal speed and acceleration vary. Yet many existing lateral controllers rely on constant-speed or operating-point-based assumptions, which can degrade performance during transient longitudinal maneuvers. Moreover, most methods assume precisely known vehicle parameters, despite real-world parametric uncertainties. To address these limitations, this paper presents a longitudinal-motion-aware robust nonlinear lateral control framework for AVs. It first derives a tracking error model that depends on varying longitudinal speed and acceleration. Using this model, feedback linearization is employed to obtain a linear input-output relation for lateral error tracking while embedding longitudinal motion into the control law. The resulting internal dynamics are then analyzed to ensure overall system stability. To address parameter uncertainty, two robust control designs with distinct implementation trade-offs are proposed: (i) a Lyapunov redesign (LR) approach inspired by sliding mode control, and (ii) an incremental nonlinear dynamic inversion (INDI) method. Both are rigorously analyzed and proven to ensure ultimate boundedness, with key robustness-tuning parameters explicitly identified. Simulations demonstrate enhanced tracking accuracy, consistent performance across varying speeds and accelerations, and robustness to model uncertainties, while also examining the effects of the robustness-related parameters. Real-vehicle tests further confirm real-time implementation and practical path-tracking performance on actual hardware.",
      "short_abstract": "As autonomous vehicles (AVs) operate in increasingly dynamic traffic conditions, lateral control must be performed while longitudinal speed and acceleration vary. Yet many existing lateral controllers rely on constant-speed or operating-point-based...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 19,
        "evidence": [
          "title: vehicle control",
          "abstract: vehicle control",
          "abstract: controller",
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.02924",
      "pdf_url": "https://arxiv.org/pdf/2607.02924.pdf"
    },
    {
      "id": "2607.02478",
      "anchor": "paper-2607-02478",
      "title": "Docking of Autonomous Vehicles with a Stationary Docking Station in 3D Space",
      "authors": [
        "Ram Milan Kumar Verma",
        "Shashi Ranjan Kumar",
        "Hemendra Arya"
      ],
      "abstract": "In this letter, we present a strategy for autonomous docking of autonomous vehicles in three-dimensional space. Docking is a safety-critical task and requires expert piloting skills. Vehicles with autonomous docking capabilities are highly desirable in various applications, such as marine vehicle docking, aerial vehicle docking, spacecraft docking, and landing. To dock autonomously with the docking station, the vehicle must align itself to a specific desired orientation relative to the docking station and also reduce speed as it approaches. The vehicle achieves near-zero speed to dock successfully and safely without colliding with the docking station. Inspired by the philosophies from the guidance literature, we present a finite-time sliding mode-based strategy to achieve the same. The range and line-of-sight kinematics relations describing the motion of the vehicle with respect to the stationary docking station are used to steer the vehicle to achieve the desired orientation for docking. This docking strategy is validated in MATLAB\\textsuperscript{\\textregistered} simulations for various initial locations and orientations of both the vehicle and the docking station.",
      "short_abstract": "In this letter, we present a strategy for autonomous docking of autonomous vehicles in three-dimensional space. Docking is a safety-critical task and requires expert piloting skills. Vehicles with autonomous docking capabilities are highly desirable in...",
      "published": "2026-07-02",
      "updated": "2026-07-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 48,
      "arxiv_url": "https://arxiv.org/abs/2607.02478",
      "pdf_url": "https://arxiv.org/pdf/2607.02478.pdf"
    },
    {
      "id": "2607.01658",
      "anchor": "paper-2607-01658",
      "title": "Teaching Vision-Language-Action Models What to See and Where to Look",
      "authors": [
        "Yuguang Yang",
        "Canyu Chen",
        "Zhewen Tan",
        "Yizhi Wang",
        "Zichao Feng",
        "Chunyang Liu",
        "Kehua Sheng",
        "Juan Zhang",
        "Linlin Yang",
        "Baochang Zhang",
        "Yan Wang",
        "Bo Zhang",
        "Xianbin Cao"
      ],
      "abstract": "Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought reasoning data, which emphasizes linguistic reasoning rather than action-grounded planning. As a result, the learned representations capture semantic knowledge but lack spatial dependencies crucial for reliable trajectory prediction. We propose DriveTeach-VLA, a framework that explicitly teaches VLAs what to see and where to look. Driving-aware Vision Distillation (DVD) injects driving-specific perceptual priors into the vision encoder, while 2D Trajectory-Guided Prompts (2D-TGP) provide spatial conditioning aligned with feasible driving trajectories. Together, they form a vision-guided learning pipeline: what to see (DVD pretraining) - where to look (TGP-guided SFT) - how to act (TGP-guided GRPO). DriveTeach-VLA achieves the state-of-the-art performance on NAVSIM and nuScenes. Our code is available at: https://github.com/ShivaTeam/DriveTeach-VLA.",
      "short_abstract": "Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought reasoning data, which emphasizes...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 26,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.01658",
      "pdf_url": "https://arxiv.org/pdf/2607.01658.pdf"
    },
    {
      "id": "2607.01382",
      "anchor": "paper-2607-01382",
      "title": "CommonRoad-Game: A Human-in-the-Loop Simulation Framework for Autonomous Driving",
      "authors": [
        "Yunfei Bi",
        "Youran Wang"
      ],
      "abstract": "Motion planning algorithms should be evaluated in human-in-the-loop environments to ensure they produce safe and efficient behaviors during interactions. However, existing simulation platforms often rely on recorded datasets, lack dedicated interfaces for real-time human interaction, or remain weakly integrated with an autonomous driving ecosystem. Moreover, many human-in-the-loop simulators are computationally intensive by design, making them less suitable for rapid prototyping and flexible experimentation in early-stage autonomous driving research. To address these limitations, we present CommonRoad-Game, a lightweight human-in-the-loop simulation framework tightly integrated with the CommonRoad platform, focusing on the systematic testing of motion planners with human participation and the analysis of human driving behaviors in interactive scenarios. We introduce a multi-threaded architecture with a robust synchronization mechanism that aligns simulation time with wall-clock time, enabling deterministic and temporally consistent interaction between autonomous and human-driven vehicles. In addition, the framework provides a scenario generation module that records driving logs, allowing diverse and reproducible test cases to be constructed from human-in-the-loop experiments. Experimental results demonstrate that CommonRoad-Game achieves stable temporal synchronization, supports scalable multi-agent simulation, and seamlessly integrates CommonRoad-compatible motion planners to generate interactive driving scenarios. The source code is publicly available at https://github.com/Yunfei-Bi8/CommonRoad-Game.",
      "short_abstract": "Motion planning algorithms should be evaluated in human-in-the-loop environments to ensure they produce safe and efficient behaviors during interactions. However, existing simulation platforms often rely on recorded datasets, lack dedicated interfaces for...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 12,
        "evidence": [
          "title: human-machine interaction",
          "abstract: human-machine interaction"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.01382",
      "pdf_url": "https://arxiv.org/pdf/2607.01382.pdf"
    },
    {
      "id": "2607.01370",
      "anchor": "paper-2607-01370",
      "title": "MapDreamer: Aerial Imagery Conditioned Latent Diffusion for Lane-Level Map Generation",
      "authors": [
        "Julian Brandes",
        "Philipp Crocoll",
        "Wolfram Burgard"
      ],
      "abstract": "High definition map generation is essential for autonomous driving, yet remains a labor-intensive process at scale. We present MapDreamer, a generative diffusion model that synthesizes lane-level vector maps with explicit topology directly from a single aerial image. MapDreamer learns a compact latent representation of lane centerlines and their topological relations using a variational autoencoder and predicts graphs with a transformer-based latent diffusion model. To align generated maps with the observed scene, we condition each denoising step on dense aerial features injected through cross-attention. To handle the varying number of lanes across scenes, we propose a lane cardinality module paired with background ghost lane latents, a learned buffer that prevents slot collapse during diffusion. Furthermore, we introduce a sliding-window global graph aggregation strategy that stitches local tiles into city-scale maps while preserving connectivity through encoded lane boundaries. Experiments on UrbanLaneGraph derived from Argoverse 2 show improved geometric and topological fidelity over non-generative baselines.",
      "short_abstract": "High definition map generation is essential for autonomous driving, yet remains a labor-intensive process at scale. We present MapDreamer, a generative diffusion model that synthesizes lane-level vector maps with explicit topology directly from a single...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Mapping & Localization: abstract: HD map",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.01370",
      "pdf_url": "https://arxiv.org/pdf/2607.01370.pdf"
    },
    {
      "id": "2607.01133",
      "anchor": "paper-2607-01133",
      "title": "Towards Metric-Agnostic Trajectory Forecasting",
      "authors": [
        "Markus Knoche",
        "Daan de Geus",
        "Bastian Leibe"
      ],
      "abstract": "Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavior and plan safe maneuvers. We observe that current state-of-the-art forecasting models on Argoverse 2 and the Waymo Open Motion Dataset tailor their training objectives to the different benchmark metrics. Because these metrics encourage conflicting behavior, we propose a paradigm change for trajectory forecasting: training models with metric-agnostic probabilistic objectives and treating metric optimization as a downstream task applied to the predictive distribution. Concretely, we introduce Trajectory Distribution Evaluation (TraDiE) policies, metric-specific policies that map a predictive distribution to the set of $K$ trajectories and confidences required by trajectory forecasting metrics. We evaluate this framework by introducing DONUT-NLL, which adapts the training objective of the state-of-the-art trajectory forecasting model DONUT to directly optimize the predictive distribution. Using our policies, DONUT-NLL achieves state-of-the-art results on all metrics of the Waymo motion prediction benchmark.",
      "short_abstract": "Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavior and plan safe maneuvers. We observe that current state-of-the-art forecasting models on Argoverse 2 and...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 19,
        "evidence": [
          "abstract: motion forecasting",
          "title: forecasting",
          "abstract: forecasting",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.01133",
      "pdf_url": "https://arxiv.org/pdf/2607.01133.pdf"
    },
    {
      "id": "2607.00839",
      "anchor": "paper-2607-00839",
      "title": "Rethinking Multi-Label Image Classification With Deep Learning: Taxonomy, Challenge, and Outlook",
      "authors": [
        "Xuelin Zhu",
        "Xiu-Shen Wei",
        "Jiawei Ge",
        "Shuai Xu",
        "Bing Wang"
      ],
      "abstract": "Multi-label image classification (MLIC), a fundamental task in computer vision, focuses on identifying multiple objects or concepts within an image, underpinning numerous read-world applications, such as autonomous driving, disease diagnosis, recommendation system, and mobile service robot. Over the past decade, deep learning paradigms based on convolutional neural networks, recurrent neural networks, and Transformers have significantly advanced this field, owing to their powerful capability in visual representation and relationship modeling. These advances have markedly improved the robustness, scalability, and generalization ability of MLIC models across diverse datasets and application domains. In this survey, we provide a comprehensive review of the deep learning-based literature on MLIC. Concretely, we first revisit the background, including problem definition, datasets, backbones and evaluation metrics. Next, we develop a plausible taxonomy for the deep learning-based MLIC approaches, organizing them into six groups: region-oriented methods, label-oriented methods, architecture-oriented methods, representation-oriented methods, learning-oriented methods, and data-oriented methods. Finally, we provide an insightful exposition of the underlying learning game in MLIC and its implications for other vision domains, and we empirically summarize the key challenges and research directions in MLIC while outlining promising avenues for future development. We believe this survey offers the research community a holistic and systematic perspective on MLIC, thereby facilitating subsequent exploration and innovation in this field and beyond.",
      "short_abstract": "Multi-label image classification (MLIC), a fundamental task in computer vision, focuses on identifying multiple objects or concepts within an image, underpinning numerous read-world applications, such as autonomous driving, disease diagnosis, recommendation...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Survey"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.00839",
      "pdf_url": "https://arxiv.org/pdf/2607.00839.pdf"
    },
    {
      "id": "2607.00710",
      "anchor": "paper-2607-00710",
      "title": "Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark",
      "authors": [
        "Richard Schwarzkopf",
        "Jonas Merkert",
        "Frank Bieder",
        "Annika Bätz",
        "Alexander Blumberg",
        "Carlos Fernandez",
        "Felix Hauser",
        "Fabian Immel",
        "Christian Kinzig",
        "Hendrik Königshof",
        "Fabian Konstantinidis",
        "Martin Lauer",
        "Willi Poh",
        "Nils Rack",
        "Kevin Rösch",
        "Yinzhe Shen",
        "Marlon Steiner",
        "Gleb Stepanov",
        "Dominik Strutz",
        "Ömer Şahin Taş",
        "Julian Truetsch",
        "Kaiwen Wang",
        "Royden Wagner",
        "Jan-Hendrik Pauls",
        "Christoph Stiller"
      ],
      "abstract": "Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a data problem or an evaluation problem, and proceeds by selecting the minimal data operator(s) that closes the resulting gap, recording new data only when no cheaper operator(s) suffices. We analyze the evolution of major autonomous driving (AD) datasets through this lens and distill a strategic framework spanning gap identification, operator choice, sensor suite design, and annotation strategy. We ground the framework in a running case study of our KITScenes dataset family. The datasets are available at: https://kitscenes.com/",
      "short_abstract": "Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.00710",
      "pdf_url": "https://arxiv.org/pdf/2607.00710.pdf"
    },
    {
      "id": "2607.00545",
      "anchor": "paper-2607-00545",
      "title": "ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation",
      "authors": [
        "Yu-Hsiang Chen",
        "Wei-Jer Chang",
        "Yi-Ting Chen",
        "Masayoshi Tomizuka"
      ],
      "abstract": "Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative models with extensive annotated data. We introduce a lightweight control adaptation framework that enables multi-modal controllability (sketch, latent behavior codes, and text) for pretrained state-of-the-art diffusion and autoregressive traffic models. By modulating intermediate features through identity-initialized FiLM layers, our method efficiently adds new control modalities while preserving the base model's generative prior. Evaluated on Waymo Open Sim Agents Challenge, our approach demonstrates strong controllability with less than 1% of the paired control data. Through context-aware condition transfer, our framework enables counterfactual scenario generation and long-tail synthesis while maintaining stable closed-loop driving realism and safety. Our framework unlocks new possibilities for controllable traffic simulation, enabling targeted scenario generation through lightweight adaptation of pretrained generative models. Project page: https://ecosim-web.github.io/",
      "short_abstract": "Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative models with extensive annotated data. We introduce a lightweight control adaptation framework that enables...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.00545",
      "pdf_url": "https://arxiv.org/pdf/2607.00545.pdf"
    },
    {
      "id": "2607.01524",
      "anchor": "paper-2607-01524",
      "title": "Flexible and Reliable Network Design for Emerging Transportation Services: Multi-stage Stochastic Programming Approach",
      "authors": [
        "Riki Kawase",
        "Koki Satsukawa",
        "Toru Seo"
      ],
      "abstract": "This paper proposes a general framework for flexible and reliable network design problems (FR-NDPs). The framework enables planners to change infrastructure investments in response to realized uncertainties, while ensuring desired levels of reliability. Motivated by emerging transportation services such as shared autonomous vehicle (SAV) systems, where historical data are scarce and technological developments uncertain, FR-NDPs integrate strategic investment decisions with operational control. We formulate the FR-NDPs as risk-averse multi-stage stochastic problems to be solvable by stochastic dual dynamic programming (SDDP) and establish sufficient conditions under which strategic and operational subproblems converge to the global optimum. We illustrate applications to SAV capacity expansion and integrated SAV-BRT (Bus Rapid Transit) route design, and numerical experiments on a Midtown Manhattan network highlight three key findings: (i) flexibility and reliability act complementarily to hedge against severe scenarios while mitigating the loss of expected performance; (ii) flexibility in investment planning allows dynamic risk hedging, with risk-averse planners reducing early-stage investments to preserve adaptability; and (iii) differences in operational flexibility between SAV and BRT systems are reflected in strategic decisions, with risk-averse planners tending to refrain investment in transport modes with lower operational flexibility.",
      "short_abstract": "This paper proposes a general framework for flexible and reliable network design problems (FR-NDPs). The framework enables planners to change infrastructure investments in response to realized uncertainties, while ensuring desired levels of reliability....",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 5,
        "evidence": [
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.01524",
      "pdf_url": "https://arxiv.org/pdf/2607.01524.pdf"
    },
    {
      "id": "2607.00874",
      "anchor": "paper-2607-00874",
      "title": "Beyond Line of Sight: Hybrid Validation of V2X Collective Perception in Complex Scenarios",
      "authors": [
        "Markos Antonopoulos",
        "Anastasia Bolovinou",
        "Bill Roungas",
        "Elena Daskalaki",
        "Angelos Amditis"
      ],
      "abstract": "This paper introduces a probabilistic framework and hybrid validation methodology for V2X-enabled Collective Perception (CP) in complex traffic scenarios. The proposed Bayesian fusion algorithm extends the perceptual horizon of connected and autonomous vehicles by integrating heterogeneous sensor observations from multiple agents into a shared probabilistic occupancy grid. Each cell of this grid encapsulates both occupancy likelihood and uncertainty, enabling explainable and trustworthy situational awareness beyond the ego vehicle's field of view. To bridge the gap between simulation and real-world evaluation, a hybrid testing framework is developed, combining CARLA-based virtual environments with vehicle-in-the-loop experimentation. Experimental results in a roundabout scenario demonstrate a 260 percent increase in field-of-view coverage and a rise in occupied-cell recall from 0.82 (ego-only) to 0.94 (six-agent CP) under nominal localization conditions. Overall, the proposed approach provides a reproducible and interpretable foundation for validating CP systems, supporting the safe and certifiable deployment of cooperative autonomous vehicles.",
      "short_abstract": "This paper introduces a probabilistic framework and hybrid validation methodology for V2X-enabled Collective Perception (CP) in complex traffic scenarios. The proposed Bayesian fusion algorithm extends the perceptual horizon of connected and autonomous...",
      "published": "2026-07-01",
      "updated": "2026-07-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Cooperative / V2X",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 34,
        "margin": 20,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 49,
      "arxiv_url": "https://arxiv.org/abs/2607.00874",
      "pdf_url": "https://arxiv.org/pdf/2607.00874.pdf"
    },
    {
      "id": "2607.00399",
      "anchor": "paper-2607-00399",
      "title": "DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving",
      "authors": [
        "Chong He",
        "Yuechen Luo",
        "Fang Li",
        "Shaoqing Xu",
        "Fuxi Wen"
      ],
      "abstract": "End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing marginal returns. Existing planners typically adopt a one-shot generation paradigm, lacking secondary validation and active correction mechanisms to detect and revise suboptimal or unsafe trajectories during inference. To address this issue, we propose DriveVer, a lightweight, plug-and-play Test-Time Verifier that leverages the test-time scaling paradigm to enable autonomous driving systems to validate and refine trajectories without costly and heavy training. We construct a dedicated trajectory dataset based on the NAVSIM benchmark through condition-driven clustering and balanced sampling according to ego-vehicle states and navigation commands. Employing a dual-head architecture, DriveVer efficiently fuses candidate trajectories with multi-view visual representations and ego-vehicle kinematic features to simultaneously predict a safety confidence score and an absolute geometric refinement vector. Extensive experiments on the NAVSIM benchmark show that DriveVer significantly improves the performance of base planning models. Notably, as an extremely compact model with only 34M parameters, DriveVer introduces minimal computational overhead, achieving competitive results while maintaining real-time inference efficiency.",
      "short_abstract": "End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing marginal returns. Existing planners typically adopt a one-shot generation paradigm, lacking secondary...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 2,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2607.00399",
      "pdf_url": "https://arxiv.org/pdf/2607.00399.pdf"
    },
    {
      "id": "2607.00283",
      "anchor": "paper-2607-00283",
      "title": "What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models",
      "authors": [
        "Amirhosein Chahe",
        "Tyler Naes",
        "Jovin D'sa",
        "Faizan M. Tariq",
        "Sangjae Bae",
        "Lifeng Zhou",
        "David Isele"
      ],
      "abstract": "Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform conservatism, yielding needlessly defensive driving, or they infer hidden spaces without estimating the impact on the planner. This work bridges the critical gap between perception and planning by enabling Vision-Language Models (VLMs) to identify and reason about the specific hidden agents that are most critical to the ego-vehicle's trajectory. We introduce a novel framework that uses Planning KL-divergence (PKL), an information-theoretic metric, to systematically identify and rank occluded agents based on their impact on the ego vehicle's plan. Using this planning-aware ranking, we employ an expert VLM (GPT-5) to generate rich, structured annotations that capture the visual evidence and reasoning required for this task. We apply this framework to the nuScenes dataset to create a new benchmark focused on high-impact scenarios. We conduct comprehensive experiments on a wide range of general-purpose and domain-adapted VLMs, demonstrating that fine-tuning on our PKL-guided data yields dramatic performance improvements across all models. Notably, our results show that smaller, fine-tuned models significantly outperform their much larger zero-shot counterparts, and that our PKL-guided data selection strategy improves performance by approximately 30\\% over random sampling. Our work presents the first systematic approach for training VLMs to focus on planning-critical occlusions, enabling more semantically grounded and efficient risk assessment in autonomous driving.",
      "short_abstract": "Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform conservatism, yielding needlessly defensive driving, or they infer hidden spaces...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2607.00283",
      "pdf_url": "https://arxiv.org/pdf/2607.00283.pdf"
    },
    {
      "id": "2606.31918",
      "anchor": "paper-2606-31918",
      "title": "DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation",
      "authors": [
        "Junzhe Jiang",
        "Zipei Ma",
        "Zijie Pan",
        "Li Zhang"
      ],
      "abstract": "A pivotal step in autonomous driving simulation involves inserting foreground vehicles with predefined trajectories into simulated scenes. This process enhances scene diversity and facilitates the creation of various corner cases for testing and improving autonomous driving models. However, existing methods often rely on pre-reconstructed 3D assets, which frequently lead to lighting inconsistencies between the inserted foreground and the background. Moreover, the reliance on limited, manually-curated 3D assets hinders large-scale deployment. To address these challenges, we propose DriveWeaver, a novel framework for controllable vehicle insertion in autonomous driving simulation. Specifically, for a masked target insertion area, DriveWeaver performs video inpainting conditioned on vehicle point clouds to generate high-quality, temporally consistent vehicles. This video-inpainting-based approach ensures seamless blending between the foreground and background, while the readily available point cloud conditions enable superior generalization. To support long-term generation, we further design a global-to-local hierarchical inpainting strategy, ensuring the consistent identity and appearance of the inserted vehicles. Meanwhile, we extract explicit 3D Gaussian representations of the inserted vehicles through an urban reconstruction pipeline to enable real-time rendering for autonomous driving simulation. Extensive experiments across diverse datasets demonstrate that our method outperforms existing baselines in visual realism and geometric consistency, providing a robust tool for scalable autonomous driving scene augmentation.",
      "short_abstract": "A pivotal step in autonomous driving simulation involves inserting foreground vehicles with predefined trajectories into simulated scenes. This process enhances scene diversity and facilitates the creation of various corner cases for testing and improving...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing",
          "Safety, Security & Verification: abstract: robustness",
          "Perception & Sensor Fusion: abstract: point cloud"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31918",
      "pdf_url": "https://arxiv.org/pdf/2606.31918.pdf"
    },
    {
      "id": "2606.31875",
      "anchor": "paper-2606-31875",
      "title": "SENSE-VAD: Sentient and Semantic Video Anomaly Detection for Autonomous Driving",
      "authors": [
        "Nghia T. Nguyen",
        "Lokman Bekit",
        "Yasin Yilmaz"
      ],
      "abstract": "Autonomous vehicles (AVs) must navigate not only motion-based hazards but also socially complex situations whose danger is constituted by inter-agent relationships rather than movement statistics alone. A child running away from a guardian, a person being carried by another, or a pursuer chasing a pedestrian across a sidewalk are all anomalous in social context, yet none produces an obvious motion signal that current anomaly detectors are equipped to flag. We introduce SENSE-VAD, the first synthetic video anomaly detection benchmark for autonomous driving explicitly designed around socially complex anomalies. Using the CARLA simulator and Unreal Engine (UE), we generate distinct anomaly scenarios across multiple categories: individual behaviors, group behaviors, person--object interactions, cyclist interactions, vehicle & agent, each annotated with per-frame binary labels. A key design principle is the separation of social anomaly from motion-based or appearance-based anomaly: many scenarios involve motion of objects that appears unremarkable in isolation but is anomalous in relational context. We additionally provide real-world normal and anomalous videos as a sim-to-real transfer probe. We evaluate state-of-the-art video anomaly detection baselines and demonstrate that socially complex anomalies constitute a distinct and currently unsolved challenge. Our dataset, annotations, and generation code are publicly available.",
      "short_abstract": "Autonomous vehicles (AVs) must navigate not only motion-based hazards but also socially complex situations whose danger is constituted by inter-agent relationships rather than movement statistics alone. A child running away from a guardian, a person being...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31875",
      "pdf_url": "https://arxiv.org/pdf/2606.31875.pdf"
    },
    {
      "id": "2606.31834",
      "anchor": "paper-2606-31834",
      "title": "Real-Time Source-Free Object Detection",
      "authors": [
        "Sairam VCR",
        "Varun Gopal",
        "Poornima Jain",
        "Vineeth N Balasubramanian",
        "Muhammad Haris Khan"
      ],
      "abstract": "Real-world detectors for autonomous driving, surveillance, and robotics must handle domain-shifts under strict latency and memory constraints, yet existing source-free object detection (SFOD) methods rely on heavyweight architectures that prioritize accuracy alone. We show this trade-off is unnecessary: building on YOLOv10, an NMS-free dual-head detector, we achieve state-of-the-art adaptation accuracy while being faster and more compact. We observe that directly applying vanilla mean-teacher self-training to dual-head detectors leads to suboptimal adaptation performance due to two key factors. First, simple pseudo-label generation strategies, such as using a single head or directly combining high-confidence predictions from both heads, yield suboptimal supervision under domain-shift. We propose DHF (Dual-Head Pseudo-Label Fusion) which selectively admits one-to-one (O2O) and one-to-many (O2M) head predictions, preserving precision and recovering missed objects. Second, we observe domain-shift collapses multi-scale feature discriminability. We propose the use of our MARD (Multi-scale Adaptive Representation Diversification) loss which mitigates this by enforcing detection-aware variance and covariance constraints on multi-scale feature maps. Both modules are training-time only, leaving inference unchanged. Across domain-shift benchmarks, our method, RT-SFOD yields 1.4 to 3.5\\% mAP gains, 1.3$\\times$ higher throughput, with $\\sim$2$\\times$ fewer parameters than prior state-of-the-art SFOD methods, thus advancing the Pareto frontier of the speed-accuracy-model size trade-off. We report main results with YOLOv10, and demonstrate generalizability with additional YOLO- and DETR-based dual-head detectors. Code is available here: https://github.com/Sairam13001/RT-SFOD/",
      "short_abstract": "Real-world detectors for autonomous driving, surveillance, and robotics must handle domain-shifts under strict latency and memory constraints, yet existing source-free object detection (SFOD) methods rely on heavyweight architectures that prioritize accuracy...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 5,
        "evidence": [
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31834",
      "pdf_url": "https://arxiv.org/pdf/2606.31834.pdf"
    },
    {
      "id": "2606.31830",
      "anchor": "paper-2606-31830",
      "title": "PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving",
      "authors": [
        "Kyuhwan Yeon",
        "Benjamin Ramtoula",
        "Daniele De Martini"
      ],
      "abstract": "Most end-to-end autonomous driving methods rely solely on instantaneous sensor observations, limiting them to reactive behavior without the anticipatory foresight human drivers employ through prior experience. We introduce geospatial visual priors, street-level visual context anchored to the intended driving route, providing visual-spatial foresight independent of real-time sensors. We propose a memory augmentation module featuring a dual-memory architecture and an adaptive memory gate, which can be easily integrated into existing end-to-end approaches. This design pairs a contextual memory for retrieved priors with a persistent fallback memory, and dynamically regulates the influence of memories based on current state compatibility. Evaluated on the NAVSIM-v2 benchmark, our approach consistently improves performance across diverse end-to-end baselines. Furthermore, because these priors are independent of onboard sensors, our method inherently improves robustness against sensor corruption, while the dual-memory design ensures safe fallback when the retrieved priors themselves become unreliable. Our project page is available at https://ori-mrg.github.io/PriorEye.",
      "short_abstract": "Most end-to-end autonomous driving methods rely solely on instantaneous sensor observations, limiting them to reactive behavior without the anticipatory foresight human drivers employ through prior experience. We introduce geospatial visual priors,...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 33,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31830",
      "pdf_url": "https://arxiv.org/pdf/2606.31830.pdf"
    },
    {
      "id": "2606.31716",
      "anchor": "paper-2606-31716",
      "title": "Gaussian Belief Propagation for Tracking With Unresolved Measurements",
      "authors": [
        "Augustin A. Saucan",
        "Florian Meyer",
        "Peter Willett"
      ],
      "abstract": "Unresolved measurements occur in many inference problems where two or more hidden processes may, at times, jointly generate a single measurement. For instance, such phenomena are encountered in multiobject tracking owing to the limited resolution capabilities of practical sensors; or in camera-aided autonomous driving due to shadowing or occlusions. Substantial performance degradation, such as track losses, are incurred when unresolved measurements are not accounted for. In this paper, we address multiobject tracking under a generalized unresolved measurement model, where any subset of objects may generate a single unresolved measurement according to a probabilistic model. Our innovation lies both in modeling and algorithm-design directions. First, we develop a probability distribution for object partitions based on a model of pairwise coupling of objects and subsequently a probability distribution for object-to-measurement association variables. This generic model incorporates sensor resolution capabilities, sensor detection, and sensor noise characteristics for object groups. Second, a generic Loopy Belief Propagation (LBP) method as well as a specialized Gaussian-LBP (GLBP) algorithm are proposed that perform object state inference under the aforementioned model. In contrast to direct marginalization methods, which involve a computational complexity of $O(m^n)$, for $m$ measurements and $n$ objects, the proposed GLBP algorithm achieves a computational complexity on the order of $O(m n 2^{n})$. Numerical results demonstrate the effectiveness of our proposed GLBP, with estimation performance that closely matches that of exact marginalization for only a fraction of the computational resources.",
      "short_abstract": "Unresolved measurements occur in many inference problems where two or more hidden processes may, at times, jointly generate a single measurement. For instance, such phenomena are encountered in multiobject tracking owing to the limited resolution capabilities...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31716",
      "pdf_url": "https://arxiv.org/pdf/2606.31716.pdf"
    },
    {
      "id": "2606.31688",
      "anchor": "paper-2606-31688",
      "title": "Semantic Occupancy Prediction with Dual Range-Voxel Representation",
      "authors": [
        "Sitao Chen",
        "Zhuangwei Zhuang",
        "Hui Luo",
        "Lizhao Liu",
        "Qingyao Wu",
        "Mingkui Tan"
      ],
      "abstract": "LiDAR-based 3D semantic occupancy prediction, which aims to provide accurate and comprehensive scene representation, is crucial for autonomous driving systems. As point clouds suffer from sparsity and incompleteness, leading to insufficient semantic learning and difficult occupancy perception, existing methods often stack multi-sweep point clouds to obtain dense spatial information. However, such a naive strategy also results in efficiency (e.g., additional computational burden) and robustness (e.g., pose transformation noise) concerns, which hinder their practical applications. In this work, we propose a Dual Range-Voxel Representation (DRVR) that leverages the range-view context and voxel-view geometry of single-sweep point clouds for 3D semantic occupancy prediction, eliminating the concerns associated with the multi-sweeps. Specifically, we use the range-view encoder to extract the compact context of the scene. To fully exploit the spatial information, we design a geometry-aware voxel-view encoder that extracts multi-scale voxel-view features separately and combines them for better geometric occupancy prediction. Moreover, we propose a range-voxel fusion module to cooperate range- and voxel-view features via voxel-to-range and range-to-voxel fusions. Extensive experiments on nuScenes-Occupancy, SemanticKITTI and SemanticPOSS show the superiority of our method. Especially on nuScenes-Occupancy, our single-sweep DRVR achieves 5.4% improvement in mIoU and 2.1x acceleration compared to the multi-sweep method.",
      "short_abstract": "LiDAR-based 3D semantic occupancy prediction, which aims to provide accurate and comprehensive scene representation, is crucial for autonomous driving systems. As point clouds suffer from sparsity and incompleteness, leading to insufficient semantic learning...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "title: occupancy",
          "abstract: occupancy",
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31688",
      "pdf_url": "https://arxiv.org/pdf/2606.31688.pdf"
    },
    {
      "id": "2606.31589",
      "anchor": "paper-2606-31589",
      "title": "From Failure to Alignment: A Requirements Engineering Framework for Machine Learning Systems",
      "authors": [
        "Amel Bennaceur",
        "Gopi Krishnan Rajbahadur",
        "Prince Mercy",
        "Bashar Nuseibeh",
        "Faeq Alrimawi"
      ],
      "abstract": "Organisations designing, developing, and deploying machine learning systems (MLS) need to be able to check that these systems are trustworthy, and communicate this clearly to their stakeholders, be they different categories of users, engineers, or wider society. By focusing on stakeholders, Requirements Engineering is well positioned to drive the design and engineering of MLS that align with the needs of their stakeholders. Yet, we still need a systematic process for modelling and reasoning about requirements for MLS that is driven both by stakeholders' needs and constraints for MLS development. This paper proposes a framework entitled REAL (Requirements Engineering for mAchines that Learn - and Fail) to help develop MLS that align with stakeholders' needs by adopting a requirements engineering approach. This model-based framework is based on three principles. First, weaving together requirements for data, models, and the system as a whole. Second, using failure to drive the exploration of alternative requirements. Third, iterative and traceable refinement of MLS requirements. We demonstrate the proposed framework using an example from autonomous driving and show that REAL supports the development of MLS that better align with stakeholders' requirements. A replication package is available online.",
      "short_abstract": "Organisations designing, developing, and deploying machine learning systems (MLS) need to be able to check that these systems are trustworthy, and communicate this clearly to their stakeholders, be they different categories of users, engineers, or wider...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31589",
      "pdf_url": "https://arxiv.org/pdf/2606.31589.pdf"
    },
    {
      "id": "2606.31512",
      "anchor": "paper-2606-31512",
      "title": "Support of Teleoperated Driving with 5G Networks",
      "authors": [
        "M. Carmen Lucas-Estañ",
        "Baldomero Coll-Perales",
        "Mohammad Irfan Khan",
        "Sergei S. Avedisov",
        "Onur Altintas",
        "Javier Gozalvez",
        "Miguel Sepulcre"
      ],
      "abstract": "Teleoperated driving (ToD) can support autonomous driving under complex or unexpected traffic scenarios that an autonomous vehicle may not understand or be able to handle. In ToD, autonomous vehicles transmit video feeds and perception data to the remote control center. The operator uses this data to understand the driving environment and remotely control the vehicle that can take over the control once the scenario is resolved. ToD requires reliable and low latency communications between the vehicle and the ToD control center. This study analyzes the feasibility to support ToD with 5G networks. The study demonstrates that the feasibility strongly depends on the bandwidth and the Time Division Duplexing (TDD) frame structure that conditions how the bandwidth is distributed between uplink and downlink transmissions. The study also shows that scaling the number of 5G-supported ToD vehicles requires the vehicles to reduce the video bitrates. The study also shows that traditional centralized 5G network deployments may be challenged by some of the most stringent ToD latency requirements due to the latency introduced by the Internet connection to the ToD control center.",
      "short_abstract": "Teleoperated driving (ToD) can support autonomous driving under complex or unexpected traffic scenarios that an autonomous vehicle may not understand or be able to handle. In ToD, autonomous vehicles transmit video feeds and perception data to the remote...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 27,
        "evidence": [
          "title: teleoperation",
          "abstract: teleoperation",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31512",
      "pdf_url": "https://arxiv.org/pdf/2606.31512.pdf"
    },
    {
      "id": "2606.31488",
      "anchor": "paper-2606-31488",
      "title": "DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation",
      "authors": [
        "Chi Huang",
        "Wenhao Zhang",
        "Hang Yin",
        "YuAn Wang",
        "Hao Li",
        "Bosheng Wang",
        "Xun Sun",
        "Liang Wang"
      ],
      "abstract": "Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, while projected LiDAR provides metric anchors that are sparse, noisy, and misaligned with image structures. Existing sparse-prompted methods incorporate LiDAR by regenerating depth from scratch, overriding the foundation model's coherent geometry and producing structural artifacts on visually continuous surfaces. Our key insight is that foundation models already capture geometrically coherent relative depth; no additional surface structure learning is required-only a per-pixel scale factor mapping relative geometry to metric coordinates. Based on this, we propose DrivingDepth, which treats sparse LiDAR as geometric prompts that locally calibrate a frozen foundation prior through residual pixel-wise scale correction, preserving dense visual geometry by construction. On nuScenes with 4-frame surround-view input, DrivingDepth achieves an AbsRel of 11.19 and an EdgeCR of 5.741, outperforming MapAnything (11.99/1.914) by simultaneously delivering SOTA metric accuracy and geometric consistency.",
      "short_abstract": "Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, while projected LiDAR provides metric anchors that are sparse, noisy, and...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 11,
        "evidence": [
          "abstract: LiDAR",
          "title: depth estimation",
          "abstract: depth estimation"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31488",
      "pdf_url": "https://arxiv.org/pdf/2606.31488.pdf"
    },
    {
      "id": "2606.31372",
      "anchor": "paper-2606-31372",
      "title": "Failure-Based Testing for Deep Reinforcement Learning Agents",
      "authors": [
        "Weibin Lin",
        "Jiangtao Meng",
        "Zheng Zheng"
      ],
      "abstract": "Deep Reinforcement Learning (DRL) agents have been widely adopted across diverse domains to address challenging decision-making problems, such as autonomous driving and robotic control. Given that many of these applications are safety- and security-critical, rigorous testing of DRL agents is indispensable. Existing testing methods are typically guided by reward signals to detect failures. However, for well-trained agents, whose performance approaches optimal levels in standard operating conditions, reward signals remain generally high, making current methods ineffective at uncovering critical failures. To address these challenges, we propose a novel failure-based method that leverages task-induced failure insights to enhance failure detection capability while reducing the number of tests required. Since DRL agents are inherently designed with human-defined tasks, they provide valuable cues about task difficulty. Intuitively, a DRL agent is more likely to fail when confronted with a more difficult task; therefore, PRT prioritizes these tasks. Building on this foundation, we propose Prior Random Testing, a black-box failure-based testing method that enables targeted prioritization while preserving the diversity of generated test cases. Guided by task-induced failure insights, PRT prioritizes failure-prone regions of the input domain, thereby facilitating efficient failure detection. PRT is evaluated on four widely used benchmarks and compared with different state-of-the-art methods including fuzzing, search-based and generative-based methods. PRT ranks among the top performers in terms of both the cost of finding the first failure and the diversity of test cases. Notably, compared to random testing, PRT achieves better diversity and reduces the testing cost by over 50%.",
      "short_abstract": "Deep Reinforcement Learning (DRL) agents have been widely adopted across diverse domains to address challenging decision-making problems, such as autonomous driving and robotic control. Given that many of these applications are safety- and security-critical,...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 25,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "title: validation or testing",
          "abstract: validation or testing",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31372",
      "pdf_url": "https://arxiv.org/pdf/2606.31372.pdf"
    },
    {
      "id": "2606.31226",
      "anchor": "paper-2606-31226",
      "title": "ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving",
      "authors": [
        "Xuchang Zhong",
        "He Zheng",
        "Chenxu Zhao",
        "Tianxiong Lv",
        "Hangqi Fan",
        "Bohua Wang",
        "Yushan Liu",
        "Li Gao",
        "Zhihao Liao",
        "Leigang Luo",
        "Congyang Zhao",
        "Yang Cai"
      ],
      "abstract": "World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing imagine-then-act paradigm, which allows errors from the more challenging visual generation stage to cascade into action planning. We introduce ForgeDrive, a unified autoregressive diffusion framework with visual-action cross-conditioning that closes this gap through act-then-imagine paradigm. ForgeDrive factorizes the future as a sequence of per-timestep frame-action pairs, intertwining each action with its corresponding visual observation. During training, we decouple the diffusion timesteps of the two modalities and introduce a UniDiffuser-style noise scheduler to get the ability to infer either modality from its counterpart and deepen understanding of relationships between images and actions. At inference, we propose a novel act-then-imagine inference paradigm, and find that at each step, action generation is a capability internalized during training, requiring no clean future frame as a prerequisite at inference time; instead, the generated action can improve the accuracy of future frame generation, which in turn enhances the quality of the next action. Additionally, we augment each step with future ego-status prediction, further sharpening planning ability. Extensive experiments on NAVSIM demonstrate that ForgeDrive not only unifies driving simulation, planning, and visual odometry into a single model, but also outperforms existing strong planners without any post-training strategy.",
      "short_abstract": "World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing imagine-then-act paradigm, which allows errors from the more challenging visual generation stage to cascade...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 0,
        "evidence": [
          "Mapping & Localization: abstract: visual odometry",
          "Prediction & World Models: abstract: future-frame prediction",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31226",
      "pdf_url": "https://arxiv.org/pdf/2606.31226.pdf"
    },
    {
      "id": "2606.31209",
      "anchor": "paper-2606-31209",
      "title": "Long-term Traffic Simulation via Structured Autoregressive Modeling",
      "authors": [
        "Lingyu Xiao",
        "Zexin Feng",
        "Xintao Yan"
      ],
      "abstract": "Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling sustained multi-agent interactions, which is further exacerbated by dynamic token cardinality as agents continuously enter and exit the scene. In this work, we propose that the solution lies in the synergy between the architectural inductive biases and statistical priors of large-scale sequence models, e.g., Large Language Models (LLMs). Our probing experiments reveal that the transferability of attention mechanisms and the distributional consistency between motion tokens and natural language enable small-scale, heavily frozen LLMs to rapidly adapt to traffic modeling. Building on this insight, we introduce RosettaSim, a unified framework that projects scene topology, agent states, and spawning intents into a structured autoregressive stream with variable length, achieving both strong short-term accuracy and stable long-horizon simulation fidelity. Furthermore, evaluating extended rollouts presents yet another hurdle, as one-to-one agent correspondence inevitably fades over time. To address this, we introduce Retrieval-based Traffic Evaluation (RTE), which retrieves semantically similar real-world scenarios as context-aware reference anchors. Experiments on the Waymo Open Sim Agent Challenge (WOSAC) demonstrate that RosettaSim achieves state-of-the-art performance in both short- and long-term simulation. Furthermore, RTE exhibits a stronger correlation with standard metrics ($r=0.83$) than existing approaches ($r=0.74$), indicating improved alignment with long-horizon simulation fidelity.",
      "short_abstract": "Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling sustained multi-agent interactions, which is further exacerbated by dynamic token cardinality as agents continuously enter...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Prediction & World Models: abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31209",
      "pdf_url": "https://arxiv.org/pdf/2606.31209.pdf"
    },
    {
      "id": "2606.31177",
      "anchor": "paper-2606-31177",
      "title": "GaussianMap: Learning Gaussian Representation for Multi-Sensor Online HD Map Construction",
      "authors": [
        "Hongyu Lyu",
        "Julie Stephany Berrio Perez",
        "Mao Shan",
        "Stewart Worrall"
      ],
      "abstract": "Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction of HD maps offers a scalable approach to generate local vectorized maps from onboard sensor observations. Existing methods commonly adopt bird's-eye-view (BEV) features as the intermediate scene representation, encoding the surrounding space with fixed-resolution dense grids. However, map elements are spatially sparse yet require fine-grained geometric localization, making uniformly allocated BEV representations redundant and less effective for vectorized map prediction. In this work, we propose GaussianMap, an online HD map construction framework that learns an adaptive Gaussian representation of the surrounding scene. This representation consists of a set of Gaussian primitives on the BEV plane, each encoding a flexible local region with geometric properties and a feature vector, allowing the model to allocate representational capacity to map-relevant regions. To generate such a representation from sensor observations, we introduce a feed-forward Gaussian encoder that progressively refines these primitives through Gaussian interaction modeling and multi-sensor feature aggregation. The refined Gaussian representation is then splatted into a BEV feature map and decoded into vectorized map predictions. Extensive experiments on nuScenes and Argoverse 2 datasets demonstrate that GaussianMap achieves state-of-the-art performance in both camera-only and camera-LiDAR fusion settings. Our code will be made publicly available.",
      "short_abstract": "Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction of HD maps offers a scalable approach to generate local vectorized maps from onboard sensor observations....",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 34,
        "evidence": [
          "abstract: localization",
          "title: HD map",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31177",
      "pdf_url": "https://arxiv.org/pdf/2606.31177.pdf"
    },
    {
      "id": "2606.31172",
      "anchor": "paper-2606-31172",
      "title": "HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection",
      "authors": [
        "Jiyong Boo",
        "Byeongin Joung",
        "Hyemin Yang",
        "Kuk-Jin Yoon"
      ],
      "abstract": "Monocular 3D lane detection plays a critical role in autonomous driving, yet recovering reliable 3D geometry from a single image remains challenging due to inherent depth ambiguity. Prior methods project image features into Bird's-Eye-View (BEV) space under a flat-ground assumption, causing geometric distortion on real-world roads. Recent methods instead predict explicit height maps to capture non-planar surfaces, but still rely on sparse anchor-based regression and exploit the recovered geometry merely for spatial transformation rather than semantic understanding. To overcome these limitations, we propose HSDF-Lane, which implicitly models the road surface as a Height-aligned Signed Distance Field (HSDF) over a densely sampled 3D feature volume. Through differentiable rendering, the HSDF jointly produces an accurate height map and surface-aligned features. We further introduce Lane-aware Semantic Positional Encoding (LSPE), which injects a lane-existence prior derived from the surface-aligned features into the transformer queries, coupling geometric structure with semantic guidance. Extensive experiments on the OpenLane benchmark show that HSDF-Lane achieves state-of-the-art performance in both 3D lane detection and height map estimation.",
      "short_abstract": "Monocular 3D lane detection plays a critical role in autonomous driving, yet recovering reliable 3D geometry from a single image remains challenging due to inherent depth ambiguity. Prior methods project image features into Bird's-Eye-View (BEV) space under a...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 14,
        "evidence": [
          "title: lane detection",
          "abstract: lane detection",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31172",
      "pdf_url": "https://arxiv.org/pdf/2606.31172.pdf"
    },
    {
      "id": "2606.31160",
      "anchor": "paper-2606-31160",
      "title": "Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving",
      "authors": [
        "Anh Dung Dinh",
        "Simon Khan",
        "Flora Salim"
      ],
      "abstract": "Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step \\emph{before} producing a trajectory. The reasoning is autoregressive and dominates inference latency, while the trajectory head is parallel and cheap. Latency is an operational constraint in autonomous driving, so accelerating the reasoning step is the central problem we address. We observe that CoC reasoning has two qualitatively different needs: most tokens continue routine setup that follows naturally from the ego-trajectory history, and a small fraction encode commitments that require fresh visual evidence about an unexpected situation. We split this reasoning into two specialized paths: a \\emph{routine reasoner} that handles the predictable continuation by attending to trajectory history, and a \\emph{deliberative reasoner} (the unmodified VLA target) that handles novel cases by attending to current visual evidence, using the speculative decoding framework as the architectural template for how the two paths cooperate. Unlike standard speculative decoding, our routine reasoner is not a smaller replica of the target; the two reasoners are deliberately specialized to read different parts of the prompt. We propose two techniques to realize this. First, we introduce \\textbf{FlatRoPE}, a 1D rotary positional embedding in the draft that breaks the rotational symmetry of the target's 3D M-RoPE, redirecting attention away from visual tokens and onto trajectory-history tokens. Second, we introduce \\textbf{Action-aware RL (AARL)}, a post-training stage that uses an action-quality reward together with a static-reference KL anchor. Together, our two-reasoner system reduces the reasoning-step running time by approximately $4\\times$ relative to the original Alpamayo planner.",
      "short_abstract": "Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step \\emph{before} producing a trajectory. The reasoning is autoregressive and dominates inference latency, while the trajectory head is parallel and...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 22,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31160",
      "pdf_url": "https://arxiv.org/pdf/2606.31160.pdf"
    },
    {
      "id": "2606.31131",
      "anchor": "paper-2606-31131",
      "title": "Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records",
      "authors": [
        "Anjali Parashar",
        "Chuchu Fan"
      ],
      "abstract": "To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day simulation based testing methods focus largely on mathematical models for efficient search of optimal scenarios, assuming a fixed scenario representation. On the other hand, real-world testing involves substantial manual effort to design scenario templates for testing. These templates represent distinct failure scenarios consisting of pre-deployment vehicle movements, map types, etc. Historical failure records for ADS are a reliable source of real-world failure conditions, which can be used for scenario generation. In this work, we propose a scenario generation pipeline using categorical and contextual information available from historical records in natural language format. Our approach consists of modular LLM based synthetic scenario generation, compatible with the testing constraints of a given system. We successfully apply our method to generate a diverse set of scenarios for testing autonomous navigation on Metadrive simulator using the NHTSA ADS crash records. Our approach results in accurate and diverse scenario generation with a combination of 4 road types, 3 non ego vehicle movement types, including on road anomalies in the form of working zones. Generated scenarios align with the provided testing conditions, and reveals interesting failures of the system within a limited testing budget of 20 scenarios. Code is available at https://github.com/anjaliParashar/crash2scenario.",
      "short_abstract": "To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day simulation based testing methods focus largely on mathematical models for efficient search of optimal scenarios, assuming...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 17,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31131",
      "pdf_url": "https://arxiv.org/pdf/2606.31131.pdf"
    },
    {
      "id": "2606.31109",
      "anchor": "paper-2606-31109",
      "title": "InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving",
      "authors": [
        "Xiaoyu Ye",
        "Leheng Li",
        "Xinyu Ji",
        "Yingjie Cai",
        "Hongda He",
        "Xu Yan",
        "Guanyi Zhao",
        "Ying-Cong Chen",
        "Bingbing Liu",
        "Shuguang Cui",
        "Zhen Li"
      ],
      "abstract": "Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and controllable synthesis of dynamic urban scenes from a single frame. In practice, our approach first reconstructs a 3D occupancy representation from the input multi-view frame. This representation serves as a foundation for autoregressive scene extension along arbitrary trajectories. Subsequently, a video diffusion model translates the coarse occupancy grid into realistic, spatiotemporally consistent video sequences. Moreover, we propose a hierarchical sketch-and-refine paradigm, in which the generated videos are re-projected as image-conditioned feedback to enhance the 3D occupancy representation, establishing cross-modal alignment and mutual enhancement between the visual and spatial domains. Extensive evaluations on the Waymo Open Dataset and nuScenes demonstrate that InfiniVerse achieves state-of-the-art performance, with a FID of 6.4 and FVD of 67.97, significantly outperforming existing benchmarks in both duration and stability.",
      "short_abstract": "Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "title: occupancy",
          "abstract: occupancy"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31109",
      "pdf_url": "https://arxiv.org/pdf/2606.31109.pdf"
    },
    {
      "id": "2606.31106",
      "anchor": "paper-2606-31106",
      "title": "What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning",
      "authors": [
        "Hyeonchang Jeon",
        "Kyungbeom Kim",
        "Eugene Vinitsky",
        "Kyung-Joong Kim"
      ],
      "abstract": "Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Summary scores from closed-loop simulators do not give significant insight into the policy, making it difficult to determine whether they truly predict the motion of surrounding vehicles, how the ego vehicle generates future plans, or whether they merely rely on brittle heuristics that happen to succeed in nominal scenarios. To better understand the limits and weaknesses of driving policies, we focus on probing for forms of prediction, i.e., where surrounding vehicles will move next, and planning, i.e., understanding how to generate safe trajectories. We focus on these two capabilities because they reflect behaviors expected of effective driving policies, and use their presence or absence to assess policy quality across data-driven behavior cloning and simulation-driven reinforcement learning policies. To evaluate the presence of these capabilities, we investigate them as a function of scale, asking whether the closed-loop gains from larger datasets and longer simulation training reflect stronger prediction and planning or merely better behavioral heuristics. We use linear probing and targeted perturbations in both imitation learning and reinforcement learning models to track when these internal signals emerge, plateau, or fail. Despite good closed-loop performance, policies often fail to form timely surrounding-vehicle predictions during near-collision events, revealing a limitation in the predictive signals available for ego planning. Finally, causal intervention shows that correcting mistaken predictions improves ego planning toward safer trajectories.",
      "short_abstract": "Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Summary scores from closed-loop simulators...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31106",
      "pdf_url": "https://arxiv.org/pdf/2606.31106.pdf"
    },
    {
      "id": "2606.31273",
      "anchor": "paper-2606-31273",
      "title": "The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims",
      "authors": [
        "Hongmin Li"
      ],
      "abstract": "AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether their scientific claims are calibrated to the evidence that supports them. This Perspective-style paper develops a conceptual and methodological framework for evidence-licensed claims in AI-assisted research. Motivated by representative routes including specialized scientific foundation models, LLM research assistants, multi-agent co-scientists, AI Scientist pipelines, mathematical discovery agents, and self-driving laboratories, it represents AI-assisted research as five operators: hypothesis generation, model-mediated consequence derivation, external validation, belief update, and claim calibration. The central claim is that calibration is not merely cautious wording but a mechanism for managing scientific assertion rights: evidence licenses some forms of speech and withholds others. The paper distinguishes linguistic, consequence-based, interventional, and evidence-licensed semantics; defines the claim-evidence gap and epistemic debt; and treats minimal structural reconstruction across heterogeneous outputs as an upward form of claim calibration. AISim-Cal is included as an illustrative synthetic dynamics exercise, not as an empirical forecast or benchmark. The resulting principles are: no claim without license, validation does not determine claim level, and automation amplifies the need for calibration. Reliable AI-assisted research is therefore evaluated as a loop that generates hypotheses, derives testable consequences, accepts independent adjudication, updates beliefs, and outputs only evidence-licensed claims.",
      "short_abstract": "AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether their scientific claims are calibrated to the evidence that supports them. This...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31273",
      "pdf_url": "https://arxiv.org/pdf/2606.31273.pdf"
    },
    {
      "id": "2606.31219",
      "anchor": "paper-2606-31219",
      "title": "CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization",
      "authors": [
        "Bo Wu",
        "Ruoshen Mo",
        "Justin Yue",
        "Yanyu Zhang",
        "Janice Nguyen",
        "Guoyuan Wu",
        "Amit Roy-Chowdhury",
        "Matthew J. Barth",
        "Hang Qiu"
      ],
      "abstract": "Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dynamics, heterogeneous sensing modalities, and scalability beyond a single cooperative partner. In this paper, we introduce CooperScene, a high-fidelity cooperative autonomy dataset with real-world C-V2X communication characterization. The dataset is organized into diverse scenes, including intersections, highway ramps, and parking lots. These scenes involve three connected and autonomous vehicles (CAVs) and one infrastructure roadside unit (RSU), all equipped with multi-modal sensors and commercial off-the-shelf C-V2X communication radios. All scenes are annotated with globally consistent 3D labels at 10 Hz, totaling 344K objects across 59K frames, underpinned by tight sensor- and agent-synchronization, centimeter-level localization and spatial alignment, precise cross-modality calibration, and 3GPP-standard-compliant C-V2X communication. CooperScene establishes a rigorous benchmark for evaluating multi-agent scaling and actual performance in real-world deployable settings. Project website for data and benchmark: https://cisl.ucr.edu/CooperScene",
      "short_abstract": "Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Dataset",
        "Benchmark",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 56,
        "margin": 51,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31219",
      "pdf_url": "https://arxiv.org/pdf/2606.31219.pdf"
    },
    {
      "id": "2606.31483",
      "anchor": "paper-2606-31483",
      "title": "A Large-Language-Model Supported Personalized Driving Framework for Lane Change in Highway Scenarios",
      "authors": [
        "Dong Bi",
        "Yongqi Zhao",
        "Paul Kovacevic",
        "Tomislav Mihalj",
        "Ji Zhou",
        "Jiayuan Gong",
        "Arno Eichberger"
      ],
      "abstract": "Personalized driving can improve the user acceptance of automated driving systems. However, existing methods still provide limited support for translating natural-language driving preferences, especially when such preferences are expressed implicitly, into executable and distinguishable driving behaviors. This paper proposes a large language model (LLM)-supported personalized driving framework for highway lane-change scenarios. The framework maps natural-language driving commands to executable planning parameters in the open-source Apollo automated driving stack according to three driving styles: aggressive, normal, and conservative. To establish this mapping, candidate planning parameters are evaluated based on the resulting lane-change behaviors, and style-specific parameter sets are constructed through clustering and style-intensity ranking. For command interpretation, a retrieval dataset is constructed to support retrieval-augmented generation (RAG), enabling LLM-based interpretation of implicit user commands. Experimental results show that the derived parameter sets generate distinguishable personalized lane-change behaviors, while RAG consistently improves preference interpretation, particularly for implicit commands. These results indicate the potential of integrating LLM-based natural-language interaction with Apollo to support personalized lane-change behavior generation. The source code and the relevant datasets are available at: https://github.com/ftgTUGraz/LLM-Personalized-Driving.",
      "short_abstract": "Personalized driving can improve the user acceptance of automated driving systems. However, existing methods still provide limited support for translating natural-language driving preferences, especially when such preferences are expressed implicitly, into...",
      "published": "2026-06-30",
      "updated": "2026-06-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: mapping",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 50,
      "arxiv_url": "https://arxiv.org/abs/2606.31483",
      "pdf_url": "https://arxiv.org/pdf/2606.31483.pdf"
    },
    {
      "id": "2607.09725",
      "anchor": "paper-2607-09725",
      "title": "Learning High-Level Decision Making with an Interaction-Aware Attention-Based Network in Autonomous Driving",
      "authors": [
        "Marcelo Contreras",
        "Willi Poh",
        "Christoph Stiller",
        "Ehsan Hashemi"
      ],
      "abstract": "Reliable learning-based high-level decision making for lane changes and speed control in automated driving must accommodate dynamically sized inputs due to varying scene traffic flow. DeepSet and its variants represent the state of the art among shared-encoder approaches; however, they neglect explicit traffic interaction modeling, limiting performance in negotiation-intensive scenarios such as intersections. Attention-based methods capture interactions among static and dynamic agents, but incur quadratic memory and computational complexity and provide limited control over representation granularity. Inspired by Perceiver IO, an attention-based architecture, DecisionPerceiver, is proposed to project dynamic agent features into a fixed-size latent space, where feature granularity is regulated by the number of latent queries, improving scalability for larger networks. A finer discretization of the action set is further proposed to increase the performance gain due to interaction awareness. Extensive evaluations across three driving scenarios that require different levels of interaction awareness demonstrate consistent performance gains and generalization across various navigation objectives. In addition, the proposed architecture is assessed in scenarios with an increasing number of vehicles to demonstrate scalability.",
      "short_abstract": "Reliable learning-based high-level decision making for lane changes and speed control in automated driving must accommodate dynamically sized inputs due to varying scene traffic flow. DeepSet and its variants represent the state of the art among...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 10,
        "evidence": [
          "title: decision-making",
          "abstract: decision-making",
          "abstract: navigation"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2607.09725",
      "pdf_url": "https://arxiv.org/pdf/2607.09725.pdf"
    },
    {
      "id": "2606.31096",
      "anchor": "paper-2606-31096",
      "title": "Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving",
      "authors": [
        "Geonho Bang",
        "Geunju Baek",
        "Dongyoung Lee",
        "Wonjun Jeong",
        "Jun Won Choi"
      ],
      "abstract": "Long-range 3D object detection is critical for safe autonomous driving at highway speeds, yet existing radar-camera fusion methods remain limited at extended ranges. BEV-based methods capture scene-level context but incur rapidly growing computation and often lose fine-grained object detail, while query-based methods are efficient but provide limited scene-level context. Temporal fusion further requires both multi-frame accumulation for sparse distant observations and object-level motion modeling for fast-moving objects. We propose Horizon3D, a sparse radar-camera fusion framework for long-range 3D object detection that combines Gaussian primitives with sparse BEV features. Horizon3D initializes Gaussian primitives at radar- and camera-estimated object keypoints using Keypoint-Guided Gaussian Initialization, refines them through Object-Centric Sparse Fusion, and splats them onto the BEV plane to fuse object-level detail with sparse radar BEV context. It further introduces Dual-Path Temporal Fusion, which aggregates temporal cues through a BEV path for scene-level accumulation and a Gaussian path for object-level motion propagation. Experiments on TruckScenes show that Horizon3D achieves state-of-the-art radar-camera 3D detection performance. On the validation set, it outperforms the previous best method by +3.0 NDS and +1.6 mAP while maintaining competitive inference speed.",
      "short_abstract": "Long-range 3D object detection is critical for safe autonomous driving at highway speeds, yet existing radar-camera fusion methods remain limited at extended ranges. BEV-based methods capture scene-level context but incur rapidly growing computation and often...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 32,
        "evidence": [
          "title: perception",
          "abstract: object detection",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.31096",
      "pdf_url": "https://arxiv.org/pdf/2606.31096.pdf"
    },
    {
      "id": "2606.30896",
      "anchor": "paper-2606-30896",
      "title": "Knowledge-Driven Dimension Estimation from a Single Image -3D Asset Generation Technology for Digital Twin Construction",
      "authors": [
        "Hidenori Sakaniwa",
        "Akihito Akai",
        "Akihiko Hyodo"
      ],
      "abstract": "In the verification of in-vehicle cameras, simulation technology using virtual spaces has advanced, enabling pre-evaluation of false detections and missed detections in various scenarios. However, discrepancies in the scale of the object being verified between the virtual and real environments can lead to a decrease in camera recognition performance. For traffic signs installed at high altitudes, distance measurement using LiDAR or stereo cameras is difficult, requiring size estimation from monocular images. This paper proposes a method for estimating the scale of an object by decomposing it into multiple structural elements and integrating external knowledge regarding design rules, geometric relationships, and conventional dimensions. Specifically, this method detects each component from a monocular image and estimates the size of each component by considering its structural relationships and dimensional consistency with surrounding elements. Furthermore, it generates a 3D asset of the object by reconstructing the estimated components. This method makes it possible to place 3D assets with a scale approximating the real environment within a digital twin space and is expected to contribute to improving the verification accuracy of in-vehicle cameras for autonomous driving in virtual environments.",
      "short_abstract": "In the verification of in-vehicle cameras, simulation technology using virtual spaces has advanced, enabling pre-evaluation of false detections and missed detections in various scenarios. However, discrepancies in the scale of the object being verified...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Perception & Sensor Fusion: abstract: camera",
          "Safety, Security & Verification: abstract: verification"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30896",
      "pdf_url": "https://arxiv.org/pdf/2606.30896.pdf"
    },
    {
      "id": "2606.30838",
      "anchor": "paper-2606-30838",
      "title": "A Practical Implementation of Day-3 Cooperative Intersection with Automated Connected Mini-Cars",
      "authors": [
        "Lorenzo Farina",
        "Vittorio Todisco",
        "Federico Gavioli",
        "Salvatore Iandolo",
        "Francesco Moretti",
        "Giuseppe Perrone",
        "Matteo Piccoli",
        "Francesco Raviglione",
        "Marco Rapelli",
        "Antonio Solida",
        "Claudio Casetti",
        "Paolo Burgio",
        "Carlo Augusto Grazia",
        "Alessandro Bazzi"
      ],
      "abstract": "Cooperative driving enabled by connected and automated vehicles is expected to improve traffic efficiency and safety, particularly at intersections where traditional control mechanisms such as traffic lights introduce delays and unnecessary stops. Although cooperative intersection management algorithms have been widely studied, experimental demonstrations remain limited. This paper presents a real-time demonstration of cooperative intersection management using connected autonomous mini-cars. The testbed consists of multiple 1:10 scale vehicles equipped with autonomous driving capabilities and wireless communication modules that interact with a centralized controller responsible for scheduling their crossing of the intersection. Vehicles approaching the intersection exchange messages with the controller to set the appropriate mobility profile to traverse the intersection without stopping. The demonstration integrates autonomous driving, wireless communication, and cooperative control in a single experimental platform, providing a practical environment for validating cooperative intersection management concepts for future intelligent transportation systems.",
      "short_abstract": "Cooperative driving enabled by connected and automated vehicles is expected to improve traffic efficiency and safety, particularly at intersections where traditional control mechanisms such as traffic lights introduce delays and unnecessary stops. Although...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 18,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30838",
      "pdf_url": "https://arxiv.org/pdf/2606.30838.pdf"
    },
    {
      "id": "2606.30807",
      "anchor": "paper-2606-30807",
      "title": "Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations",
      "authors": [
        "Halima Bouzidi",
        "Mboutidem Ekemini Mkpong",
        "Haoyu Liu",
        "Mohammad Abdullah Al Faruque"
      ],
      "abstract": "Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based retrieval becoming the dominant trajectory-decoding paradigms. Despite their architectural diversity, current generative AD planners share a common inference pattern: a fixed set of candidate trajectories (anchors, vocabulary entries, or proposal queries) is scored by one or more learned heads conditioned on the Bird's-Eye-View (BEV) features, and the highest-scored candidate is returned as the final trajectory. Under this design, the scoring head is the only barrier between perception and the motion command, and its decision margins between competing candidates are often small. We introduce \\textsc{Derail}, an adversarial framework that exploits this scoring-head attack surface. Evaluated on various generative planners, \\textsc{Derail} flips the trajectory selection from a safe to an unsafe candidate, with score drops of $39$--$80\\%$ and collision rates of up to $50\\%$, consistently outperforming generic loss-maximization and feature-divergence attacks. Our analysis suggests that safety-violating objectives govern attack effectiveness against generative AD planners, and that the scoring-head inference pattern itself is a recurring attack surface worth explicit defensive consideration.",
      "short_abstract": "Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based retrieval becoming the dominant trajectory-decoding paradigms. Despite their architectural diversity, current...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 32,
        "margin": 2,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30807",
      "pdf_url": "https://arxiv.org/pdf/2606.30807.pdf"
    },
    {
      "id": "2606.30537",
      "anchor": "paper-2606-30537",
      "title": "Learning from Mistakes: Rollout-Retrieval Lifelong Policy Learning for Autonomous Driving",
      "authors": [
        "Cheng Gong",
        "Haoyang Wang",
        "Chao Lu",
        "Zirui Li",
        "Jianwei Gong"
      ],
      "abstract": "Autonomous driving policies should be able to improve continually as deployment exposes them to increasingly diverse and long-tail traffic situations. However, most learning-based policies are trained or fine-tuned on expert demonstrations and then rely largely on generalization to handle challenging closed-loop scenarios, lacking an explicit mechanism to correct and retain the mistakes exposed in these scenarios. This paper studies autonomous driving policy improvement from a lifelong learning perspective: Can a pretrained policy improve continually by accumulating corrective knowledge derived from its own mistakes, while retaining previously acquired driving competence? To answer this question, we propose Rollout-Retrieval Lifelong Policy Learning (R$^2$LPL), a policy learning framework that retrieves corrective targets from recoverable policy-induced mistakes and retains the resulting knowledge through lifelong policy learning. R^2LPL addresses a key bottleneck in continual policy improvement: closed-loop mistakes reveal where the policy is weak, but do not directly specify what the policy should learn. By filtering recoverable mistake-related states and retrieving feasible corrective targets, R$^2$LPL turns sparse failure evidence into compact supervised knowledge for stable and sample-efficient policy improvement. We evaluate R$^2$LPL on large-scale closed-loop nuPlan benchmarks. With only a few rollout and continual-learning cycles, R$^2$LPL elevates a learning-based planner with moderate initial performance to state-of-the-art performance across the evaluated benchmarks, especially on the challenging and long-tail Test14-hard split. These results demonstrate the effectiveness of R$^2$LPL in converting recoverable closed-loop mistakes into corrective knowledge for sustained policy improvement.",
      "short_abstract": "Autonomous driving policies should be able to improve continually as deployment exposes them to increasingly diverse and long-tail traffic situations. However, most learning-based policies are trained or fine-tuned on expert demonstrations and then rely...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: law or policy",
          "abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30537",
      "pdf_url": "https://arxiv.org/pdf/2606.30537.pdf"
    },
    {
      "id": "2606.30421",
      "anchor": "paper-2606-30421",
      "title": "OWMDrive: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World Model",
      "authors": [
        "Junjie Cheng",
        "Ruiqi Song",
        "Ye Wu",
        "Nanxing Zeng",
        "Ximiao Li",
        "Yunfeng Ai"
      ],
      "abstract": "Autonomous driving systems are steadily moving toward end-to-end paradigms to mitigate the limited adaptability of rule-based pipelines in complex traffic environments. However, most existing learning-based methods still make decisions from static representations of the current scene, without explicit future rollouts or modeling of the temporal causal dynamics in traffic interactions. This limitation often results in unstable or overly conservative planning under high-uncertainty conditions, such as occlusions and unexpected events. To overcome these challenges, we introduce OWMDrive, a generative end-to-end driving framework built upon an Occupancy World Model for multi-step 3D occupancy forecasting, which serves as a conditional prior to guide diffusion-based planning. Conditioned on both current observations and predicted future states, the planner iteratively refines trajectory candidates to generate a reinforced driving trajectory. By explicitly modeling scene evolution over future horizons, OWMDrive captures key spatiotemporal causal dependencies, which leads to more foresighted and robust trajectory generation. Extensive experiments demonstrate that OWMDrive significantly improves planning reliability and safety, especially in challenging and partially observable driving scenarios.",
      "short_abstract": "Autonomous driving systems are steadily moving toward end-to-end paradigms to mitigate the limited adaptability of rule-based pipelines in complex traffic environments. However, most existing learning-based methods still make decisions from static...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 17,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30421",
      "pdf_url": "https://arxiv.org/pdf/2606.30421.pdf"
    },
    {
      "id": "2606.29879",
      "anchor": "paper-2606-29879",
      "title": "LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving",
      "authors": [
        "Chen Yang",
        "Yuhao Wei",
        "Ze Xu",
        "Ziheng Zou",
        "Shuang Liang",
        "Delin Ouyang",
        "Lingfeng Qi",
        "Jie Li",
        "Guofa Li"
      ],
      "abstract": "Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain insufficient for geometrically accurate, future-aware, and multi-view-grounded planning. To address these limitations, we develop the Layer-Wise World-Model-Guided Driving framework (LWDrive). LWDrive is a VLM planning framework that refines coarse trajectories through layer-wise world-model guidance. Instead of treating the VLM output as the final trajectory, LWDrive uses it as an intent-aware coarse plan, expands a diverse candidate space around it, and progressively refines the candidates through a Foresight Cascade Planner (FCP). Specifically, we introduce future-frame generation supervision to encourage the VLM to learn forward-looking scene representations, thereby injecting planning-relevant predictive dynamics into its internal hidden states. Built upon these world-model-supervised representations, FCP exploits VLM features across multiple layers and integrates historical temporal states, Action-Query representations, and current-frame multi-view Bird's-Eye-View (BEV) features to refine candidate trajectories in a coarse-to-fine manner. This design enables progressive correction of spatial positions and motion trends while grounding trajectory refinement with multi-view scene cues and preserving the high-level driving intention produced by the large model. Finally, a score head evaluates the refined candidates and selects the best trajectory as the final planning output. Experiments show that LWDrive achieves a score of 92.0 on the NAVSIM benchmark and 89.6 on NAVSIM-v2. Code and models will be made publicly available.",
      "short_abstract": "Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 3,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.29879",
      "pdf_url": "https://arxiv.org/pdf/2606.29879.pdf"
    },
    {
      "id": "2606.30940",
      "anchor": "paper-2606-30940",
      "title": "Motion Planning in Compressed Representation Spaces",
      "authors": [
        "Lukas Lao Beyer",
        "Sertac Karaman"
      ],
      "abstract": "Deep learning methods have vastly expanded the capabilities of motion planning in robotics applications, as learning priors from large-scale data has been shown to be essential in capturing the highly complex behavior required for solving tasks such as manipulation or navigation for autonomous vehicles. At the same time, model-based planning algorithms based on search or optimization remain an essential tool due to their flexibility, efficiency, and the ability to incorporate domain knowledge via expert-designed algorithms and objective functions. We propose a new generative framework to unify these two paradigms. First, we learn an autoencoder with a high compression ratio and a latent space of hierarchically ordered, discrete-valued tokens. Leveraging both the dimensionality reduction and the hierarchical coarse-to-fine structure learned by this autoencoder, we then perform motion planning by directly searching in the latent space of tokens. This search can optimize arbitrary objective functions specified at test time, providing a large degree of flexibility while maintaining efficiency and producing realistic solutions by relying on the generative capabilities of the highly compressed autoencoder. We evaluate our method on nuPlan and the Waymo Open Motion Dataset, showing how latent space search can be used for a variety of guided behavior generation tasks, achieving strong performance for closed-loop motion planning and multi-agent guided scenario synthesis without requiring any task-specific training.",
      "short_abstract": "Deep learning methods have vastly expanded the capabilities of motion planning in robotics applications, as learning priors from large-scale data has been shown to be essential in capturing the highly complex behavior required for solving tasks such as...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 34,
        "margin": 34,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: navigation"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30940",
      "pdf_url": "https://arxiv.org/pdf/2606.30940.pdf"
    },
    {
      "id": "2606.30554",
      "anchor": "paper-2606-30554",
      "title": "SubEdge: A Subscriber-Centric Edge Computing Subsystem in 6G Networks for AI",
      "authors": [
        "Abdirazak Ali Asir Rage",
        "Riccardo Pozza",
        "Rahim Tafazolli"
      ],
      "abstract": "Beyond traditional connectivity, 6G is envisioned to transform mobile networks into a distributed fabric that provides native integrated communication, computing, and intelligence services. AI-native terminals (e.g., robots, autonomous vehicles, and smart glasses) require real-time inference from individualised, manufacturer-specific models that cannot be executed on-board nor shared across subscribers, making per-subscriber edge compute the necessary complement to per-subscriber connectivity. Existing Network for AI (Net4AI) architectures provision compute for application providers through shared deployments and do not address per-subscriber provisioning. This paper proposes SubEdge, a Net4AI subsystem that provisions integrated communication and compute resources on a per-subscriber basis, ensuring the coupled migration of both dimensions to maintain service continuity during mobility. SubEdge contributes the computing context--a per-subscriber data structure binding a Subscription Permanent Identifier (SUPI) to its inference container, edge node, and service entitlement--and a mobility-event-driven mechanism that simultaneously migrates the subscriber's compute instance and its traffic-routing policy when the serving cell changes. SubEdge operates as an Application Function over existing Network Exposure Function (NEF) APIs with zero 3GPP core modifications. Experimental evaluation on a real-world testbed shows that SubEdge's mobility-driven joint communication-and-compute migration reduces 95th-percentile latency from 22.9 ms to 12.2 ms with zero packet loss across six mobility events, sustains 99.92% frame delivery for an end-to-end 30 fps inference workload, and completes 1,560 migration operations across batches of up to 50 simultaneously migrating subscribers with 100% success.",
      "short_abstract": "Beyond traditional connectivity, 6G is envisioned to transform mobile networks into a distributed fabric that provides native integrated communication, computing, and intelligence services. AI-native terminals (e.g., robots, autonomous vehicles, and smart...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 21,
        "evidence": [
          "abstract: deployment or real-time",
          "title: edge or cloud",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30554",
      "pdf_url": "https://arxiv.org/pdf/2606.30554.pdf"
    },
    {
      "id": "2606.30153",
      "anchor": "paper-2606-30153",
      "title": "The Spectrum Strikes Back: Infrared POV Attacks on Traffic Sign Classification",
      "authors": [
        "Michael Kühr",
        "Mevlüt Yildirim",
        "Maximilian Luedecke",
        "Mohammad Hamad",
        "Sebastian Steinhorst"
      ],
      "abstract": "Traffic sign classification is a crucial task for autonomous vehicles, and numerous attacks against it have been identified. A majority of physical adversarial attacks involve attaching patches to traffic signs or projecting perturbations on them. While they demonstrate high effectiveness, they are perceptible to humans. At the same time, light-based attacks outside the human visible spectrum are known but have limitations in their dynamic adaptability. We propose a persistence-of-vision-based attack that operates in the near-infrared light spectrum. With the possibility of showing dynamic, remotely triggered content, this allows a stealthy physical adversarial attack against traffic sign classification. By identifying the optimal position through digital simulation, we conduct extensive real-world evaluations using two different traffic signs, 12 machine learning models from different families, multiple distances up to 20 meters, and varying illumination conditions. Our evaluation shows high attack success rates across our test scenarios. We propose near-infrared cutoff filters and a software-based detection mechanism as defenses, and tackle limitations of the near-infrared persistence of vision display by prototyping a human-visible RGB version of it.",
      "short_abstract": "Traffic sign classification is a crucial task for autonomous vehicles, and numerous attacks against it have been identified. A majority of physical adversarial attacks involve attaching patches to traffic signs or projecting perturbations on them. While they...",
      "published": "2026-06-29",
      "updated": "2026-06-29",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 16,
        "evidence": [
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 51,
      "arxiv_url": "https://arxiv.org/abs/2606.30153",
      "pdf_url": "https://arxiv.org/pdf/2606.30153.pdf"
    },
    {
      "id": "2606.29684",
      "anchor": "paper-2606-29684",
      "title": "Evolutionary Hyperparameter Optimization to Find Lightweight CNN Models for Autonomous Steering",
      "authors": [
        "Devson Butani",
        "Ryan Kaddis",
        "Chan-Jin Chung"
      ],
      "abstract": "This research investigates the optimization of Convolutional and Dense Neural Networks (CNNs and DNNs) for autonomous steering using the (N+M) Evolution Strategy (ES) with the 1/5th success rule. The primary objective is to develop a lightweight CNN based model capable of real-time steering angle prediction, mimicking human driving behavior on predefined paths. The ES algorithm automates hyperparameter tuning, dynamically adjusting parameters such as filter sizes and layer configurations. Data collection encompasses driving scenarios recorded via the LTU ACTor autonomous driving platform, including variations in path direction and driving style. The very small dataset consists of timestamped images labeled with steering angles and pre-processed to focus on relevant visual information. Initial experiments involve training a baseline CNN model, which is then refined using ES to significantly reduce the size of the model while maintaining competitive predictive accuracy. The results highlight the viability of lightweight neural network architectures for real-time autonomous systems, striking a balance between computational efficiency and performance. This study not only advances research initiatives on the use of evolutionary algorithms for autonomous driving applications but also lays the foundation for the deployment of cost-effective and scalable solutions in self-driving technology.",
      "short_abstract": "This research investigates the optimization of Convolutional and Dense Neural Networks (CNNs and DNNs) for autonomous steering using the (N+M) Evolution Strategy (ES) with the 1/5th success rule. The primary objective is to develop a lightweight CNN based...",
      "published": "2026-06-28",
      "updated": "2026-06-28",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 3,
        "evidence": [
          "title: steering or braking",
          "abstract: steering or braking"
        ]
      },
      "recency": "fresh",
      "age_days": 52,
      "arxiv_url": "https://arxiv.org/abs/2606.29684",
      "pdf_url": "https://arxiv.org/pdf/2606.29684.pdf"
    },
    {
      "id": "2606.29374",
      "anchor": "paper-2606-29374",
      "title": "L2D2-GS: Learning to Densify for Feedforward Dynamic Gaussian Scene Reconstruction",
      "authors": [
        "Zetian Song",
        "Chenming Wu",
        "Junnan Liu",
        "Chitian Sun",
        "Liangliang He",
        "Hangjun Ye",
        "Jiaqi Zhang",
        "Siwei Ma",
        "Wen Gao"
      ],
      "abstract": "High-fidelity reconstruction of dynamic urban environments is a cornerstone of autonomous driving simulation and large-scale world modeling. While 3D Gaussian Splatting (3DGS) has established a new standard for real-time rendering, its reliance on expensive per-scene optimization limits scalability. Conversely, recent feedforward methods that infer Gaussian parameters offer faster speed but face fundamental bottlenecks: they are memory-prohibitive at high resolutions and struggle to fuse dense multi-view observations consistently. This paper presents L2D2-GS, a unified framework that reformulates generalizable reconstruction not as a one-shot regression, but as a robust iterative process of optimization and densification. To resolve the ambiguity of supervision in primitive generation, we propose a self-supervised densification policy that derives explicit reward signals from global reconstruction gains to guide local densification. Furthermore, we mitigate irreversible early-stage artifacts through a geometric regularization mechanism, utilizing reparameterization to constrain the optimization manifold and prevent convergence to poor local optima. Extensive experiments on the PandaSet and Waymo datasets demonstrate that our method achieves state-of-the-art reconstruction fidelity and strong zero-shot generalization, while using fewer primitives than competing baselines.",
      "short_abstract": "High-fidelity reconstruction of dynamic urban environments is a cornerstone of autonomous driving simulation and large-scale world modeling. While 3D Gaussian Splatting (3DGS) has established a new standard for real-time rendering, its reliance on expensive...",
      "published": "2026-06-28",
      "updated": "2026-06-28",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 52,
      "arxiv_url": "https://arxiv.org/abs/2606.29374",
      "pdf_url": "https://arxiv.org/pdf/2606.29374.pdf"
    },
    {
      "id": "2606.29286",
      "anchor": "paper-2606-29286",
      "title": "ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving",
      "authors": [
        "Dingyi Yao",
        "Xinqi Zhang",
        "Lihui Peng",
        "Jianming Hu",
        "Danya Yao",
        "Yi Zhang"
      ],
      "abstract": "Synthetic data mitigates the data scarcity problem in autonomous driving perception. However, the synthetic-to-real gap leads to performance degradation, hindering real-world model generalization. Although current methods leverage diffusion models for photorealistic style transfer to bridge this gap, they critically ignore a practical asymmetry: while synthetic data possesses perfect pixel-level annotations, real-world style reference images generally lack corresponding labels. Consequently, existing methods relying on symmetric semantic guidance suffer from either prohibitive annotation costs or severe semantic misalignment. To address this dilemma, we formally propose a novel task: Asymmetric Style Transfer for Autonomous Driving (ASTAD), which requires semantically consistent transfer using only labeled synthetic content and unlabeled real-world references. We further introduce the ASTModel, a training-free two-stage framework designed to bridge this domain gap under asymmetric constraints. ASTModel first extracts a coarse semantic prior from the unlabeled target, followed by dynamic prior refinement and class-consistent style injection during the denoising process. Extensive experiments demonstrate that ASTModel significantly outperforms existing methods in downstream perception utility and structural fidelity, while offering a 3.2$\\times$ inference speedup. This work aligns synthetic-to-real adaptation with practical constraints, holding the potential to accelerate the scalable deployment of robust autonomous driving systems. Code: https://github.com/Dingyi-Yao/ASTAD.",
      "short_abstract": "Synthetic data mitigates the data scarcity problem in autonomous driving perception. However, the synthetic-to-real gap leads to performance degradation, hindering real-world model generalization. Although current methods leverage diffusion models for...",
      "published": "2026-06-28",
      "updated": "2026-06-28",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Prediction & World Models: abstract: world model",
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "fresh",
      "age_days": 52,
      "arxiv_url": "https://arxiv.org/abs/2606.29286",
      "pdf_url": "https://arxiv.org/pdf/2606.29286.pdf"
    },
    {
      "id": "2606.29677",
      "anchor": "paper-2606-29677",
      "title": "Lateral String Stability for Vehicle Platoons",
      "authors": [
        "Sixu Li",
        "Swaroop Darbha",
        "Yang Zhou"
      ],
      "abstract": "Connected and automated vehicle (CAV) platooning promises gains in energy efficiency and traffic throughput and, most critically, in safety. These safety benefits hinge on string stability, which determines how disturbances propagate along a platoon. While longitudinal string stability is well studied, lateral string stability, which governs the propagation of path-tracking errors that can lead to unsafe deviations from the intended path, remains underexplored. Its importance is increasing as autonomous vehicles rely more heavily on onboard sensing and map-free navigation, where sensor occlusion and dense formations amplify safety risks. This paper presents a new framework for lateral string stability that directly addresses safety-critical path-relative tracking errors and enables consistent comparison across vehicles following the same road geometry. Central to this framework is an arc-length (Eulerian) viewpoint, a departure from traditional analyses, that clarifies how tracking errors at a given point on the path propagate from one vehicle to the next. A formal definition of lateral string stability is introduced along with two control strategies: an onboard-sensing-only controller and a novel learn-from-predecessor approach utilizing vehicle-to-vehicle (V2V) communication. We show that onboard sensing alone cannot guarantee attenuation of path-tracking errors, imposing a fundamental safety limitation, whereas V2V communication enables true error attenuation.",
      "short_abstract": "Connected and automated vehicle (CAV) platooning promises gains in energy efficiency and traffic throughput and, most critically, in safety. These safety benefits hinge on string stability, which determines how disturbances propagate along a platoon. While...",
      "published": "2026-06-28",
      "updated": "2026-06-28",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "abstract: V2X",
          "abstract: connected vehicles",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 52,
      "arxiv_url": "https://arxiv.org/abs/2606.29677",
      "pdf_url": "https://arxiv.org/pdf/2606.29677.pdf"
    },
    {
      "id": "2606.29566",
      "anchor": "paper-2606-29566",
      "title": "Analyzing Uncertainty in the Spatial Representation of the Kinematic Bicycle Model",
      "authors": [
        "Shafayat Abrar",
        "M. Zaeem Baig",
        "Shahir Ul Islam Anzal",
        "Abdul Basit Memon"
      ],
      "abstract": "Locating a vehicle and determining its orientation in an uncertain environment is a critical challenge in autonomous vehicle navigation and path planning. To address these challenges, a vehicle estimates its pose while depending on sensor data that offer noisy measurements. These uncertainties in pose quantities are expressed mathematically as a covariance matrix. The real-time computation of the covariance matrix is critical because of the non-linearity involved in the kinematic model. The challenge is thus to evaluate the evolution of the covariance matrix of a vehicle's discretized stochastic kinematics. The purpose of this study is to obtain a near-accurate evolution of the covariance matrix of the rear-wheel bicycle kinematic model under uncertainties in wheel displacement and steering angle. We used Taylor's series to linearize the nonlinear trigonometric functions and provided closed-form expectations of random variables with the required accuracy. Our analytical findings are in good agreement with those obtained from Monte-Carlo simulations. Our contribution is probably the first detailed closed-form presentation of the covariance matrix constituents of the vehicle under evaluation, which were previously reported either incorrectly or incompletely. These findings aid in identifying the potential and constraints of the discretized kinematic model as well as its stochastic analysis. The techniques presented here are useful for the simultaneous localization and odometry self-calibration of certain mobile robots and autonomous vehicles.",
      "short_abstract": "Locating a vehicle and determining its orientation in an uncertain environment is a critical challenge in autonomous vehicle navigation and path planning. To address these challenges, a vehicle estimates its pose while depending on sensor data that offer...",
      "published": "2026-06-28",
      "updated": "2026-06-28",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 10,
        "margin": 2,
        "evidence": [
          "abstract: motion planning",
          "abstract: planning",
          "abstract: navigation"
        ]
      },
      "recency": "fresh",
      "age_days": 52,
      "arxiv_url": "https://arxiv.org/abs/2606.29566",
      "pdf_url": "https://arxiv.org/pdf/2606.29566.pdf"
    },
    {
      "id": "2606.29390",
      "anchor": "paper-2606-29390",
      "title": "Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems",
      "authors": [
        "Heidy Khlaaf"
      ],
      "abstract": "Novel safety, socio-economic, and ethical harms arising from the deployment of AI-based systems have led to a breadth of work seeking to map, measure, and mitigate against newly found risks. These works have heavily leveraged techniques and terminology from the fields of System Safety Engineering and Cybersecurity, yet they have fallen short in accounting for the limitations and nuances that reduce the efficacy and correct application of adopted methodologies. Furthermore, misuse of terminology entailing compliance with established safety and security properties can mislead stakeholders with regard to the claims an AI system satisfies and provide a false sense of safety. In this paper, we seek to align overlapping, AI-adjacent communities on a consistent and comprehensive assurance terminology crucial for the safe deployment of AI-based systems. We outline why previous attempts to adapt risk assessment techniques and terminology from the safety and security fields have been insufficient. We then propose a novel end-to-end AI risk framework that integrates the concept of an Operational Design Domains (ODD), initially introduced for ADS (Automated Driving Systems) [1], for more general AI-based systems. The purpose of an ODD is to provide a description of the specific operating conditions for which an AI-system is designed to properly behave, thus outlining the safety envelope for which system hazards and harms can be determined against. We believe that by defining a more concrete operational envelope, developers and auditors can better assess potential risks and required safety mitigations for AI-based systems.",
      "short_abstract": "Novel safety, socio-economic, and ethical harms arising from the deployment of AI-based systems have led to a breadth of work seeking to map, measure, and mitigate against newly found risks. These works have heavily leveraged techniques and terminology from...",
      "published": "2026-06-28",
      "updated": "2026-06-28",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 9,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: risk assessment"
        ]
      },
      "recency": "fresh",
      "age_days": 52,
      "arxiv_url": "https://arxiv.org/abs/2606.29390",
      "pdf_url": "https://arxiv.org/pdf/2606.29390.pdf"
    },
    {
      "id": "2606.30694",
      "anchor": "paper-2606-30694",
      "title": "DSIP: A Dynamic Coordination Planner for Signal-Free Intersections using Diffusion-Model-Based Multi-Agent Motion Planning",
      "authors": [
        "Qian Hu",
        "Haoyang Peng",
        "Songan Zhang",
        "Ming Yang",
        "Hongtei Eric Tseng"
      ],
      "abstract": "Traffic signal control at urban intersections inherently introduces stop-and-go behavior, resulting in increased delays and reduced traffic efficiency, especially under high traffic demand. With the emergence of connected and automated vehicles (CAVs), trajectory-level coordination has emerged as a high-potential strategy to augment or transcend conventional phase-based management. This paper proposes DSIP (Diffusion-model-based Signal-free Intersection Planner), a multi-agent motion planning framework driven by a generative diffusion process. DSIP shifts the intersection management paradigm from discrete temporal phasing to continuous multi-vehicle trajectory optimization. This work evaluates the theoretical upper-bound performance of this coordination strategy under idealized communication and execution conditions to isolate the core benefits of the diffusion-driven approach. Using the SUMO platform, we evaluate DSIP across diverse four-leg intersection configurations. Experimental results demonstrate that DSIP significantly reduces average delay and maintains higher average speed compared to both fixed-time signal control and state-of-the-art reinforcement-learning-based controllers, particularly in medium- to high-density traffic. These findings suggest that diffusion-based trajectory planning provides a scalable and robust foundation for future autonomous intersection management. By unlocking latent intersection capacity through software-defined coordination, this approach offers a cost-effective pathway to improve urban traffic flow efficiency without requiring physical infrastructure expansion.",
      "short_abstract": "Traffic signal control at urban intersections inherently introduces stop-and-go behavior, resulting in increased delays and reduced traffic efficiency, especially under high traffic demand. With the emergence of connected and automated vehicles (CAVs),...",
      "published": "2026-06-28",
      "updated": "2026-06-28",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 26,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 52,
      "arxiv_url": "https://arxiv.org/abs/2606.30694",
      "pdf_url": "https://arxiv.org/pdf/2606.30694.pdf"
    },
    {
      "id": "2606.29115",
      "anchor": "paper-2606-29115",
      "title": "When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems",
      "authors": [
        "Yash Tandon",
        "Giovanni Tapia Lopez",
        "Marcus Blennemann",
        "Mohan Trivedi",
        "Ross Greer"
      ],
      "abstract": "Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoidance and minimal risk condition (MRC) behaviors such as slowing or stopping when uncertainty arises. Although effective in reducing immediate crash risk, real-world deployments indicate that stopping alone does not guarantee safe integration into human-governed roadway systems. Incidents reported by municipalities and public records show that AV fallback behaviors can obstruct traffic, interfere with emergency response operations, and create accessibility challenges for passengers and pedestrians. This paper presents an analysis of publicly documented incidents involving AV stopping behavior and human-AV interaction failures. We categorize these incidents according to limitations in perception, planning, and control within current AV architectures. Using this taxonomy, we identify key gaps in existing safety paradigms, particularly the lack of mechanisms for interpreting human authority, responding to multimodal instructions, and adapting to dynamic, socially regulated traffic conditions. We then review emerging research directions that support human-interactive perception, language-grounded and accessibility-aware planning, and assisted control through remote guidance and teleoperation. The analysis highlights the need to augment current AV safety frameworks with capabilities that enable cooperative interaction with human agents and infrastructure. These findings suggest that reliable urban deployment of AVs requires moving beyond passive fallback strategies toward human-interactive autonomy.",
      "short_abstract": "Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoidance and minimal risk condition (MRC) behaviors such as slowing or stopping when uncertainty arises....",
      "published": "2026-06-27",
      "updated": "2026-06-27",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 3,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: teleoperation",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 53,
      "arxiv_url": "https://arxiv.org/abs/2606.29115",
      "pdf_url": "https://arxiv.org/pdf/2606.29115.pdf"
    },
    {
      "id": "2606.29097",
      "anchor": "paper-2606-29097",
      "title": "TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation",
      "authors": [
        "Zhi Tu",
        "Liangkun Niu",
        "Tianyi Zhang"
      ],
      "abstract": "Recent research has investigated the use of large language models (LLMs) to generate traffic scenarios for autonomous driving. However, pretrained LLMs often fail to align with real-world traffic distributions. In this work, we present TrafficAlign, an automated framework that synthesizes traffic scenarios based on real-world driving videos, performs data validation, and aligns LLMs with the synthesized scenarios. The evaluation shows that traffic scenarios generated by TrafficAlign are highly effective, revealing up to 10.8% more collisions on average across three autonomous driving models than state-of-the-art methods. Furthermore, fine-tuning these driving models with TrafficAlign-generated scenarios significantly reduced collision rates by 36.1% compared with the original models. A qualitative study using traffic datasets from six geographically diverse regions shows that TrafficAlign-generated scenarios exhibit strong alignment with corresponding traffic distributions in these regions.",
      "short_abstract": "Recent research has investigated the use of large language models (LLMs) to generate traffic scenarios for autonomous driving. However, pretrained LLMs often fail to align with real-world traffic distributions. In this work, we present TrafficAlign, an...",
      "published": "2026-06-27",
      "updated": "2026-06-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 53,
      "arxiv_url": "https://arxiv.org/abs/2606.29097",
      "pdf_url": "https://arxiv.org/pdf/2606.29097.pdf"
    },
    {
      "id": "2606.29020",
      "anchor": "paper-2606-29020",
      "title": "Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis",
      "authors": [
        "Chenghao Qian",
        "Nedko Savov",
        "Lingdong Kong",
        "Yeying Jin",
        "Rui Song",
        "Wenjing Li",
        "Zhun Zhong",
        "Jiaqi Ma",
        "Gustav Markkula",
        "Luc Van Gool"
      ],
      "abstract": "Weather synthesis aims to add weather effects to input videos while preserving scene identity, structure, and motion. The key limitation of existing methods is the lack of diversity in weather appearance and effective control over weather dynamics (e.g., temporal evolution and particle motion). Most approaches rely on text prompts, which are inherently underspecified and often fail to produce detailed weather characteristics. Additionally, general-purpose video editors optimized for clean and aesthetic outputs tend to suppress heavy weather phenomena, making dense particle effects difficult to generate. To address these, we propose a Semantic-Aware, Physics-Informed, and Geometry-Grounded framework that steers an off-the-shelf video editor to synthesize diverse global appearances and detailed particle dynamics. We factorize the synthesis into three conditional signals, so that each provides a distinct and stable source of guidance: semantics specifies what the weather should look like, dynamics governs how it evolves over time, and geometry determines where it should appear in the scene. Specifically, we introduce (1) semantic-aware appearance anchoring to establish the target appearance from scene semantics and user input; (2) physics-informed dynamic simulation to generate particle effects by simulating a Gaussian-represented particle field under gravity, wind, and turbulence; and (3) geometry-grounded video synthesis to align the simulated particles with target scene geometry and synthesize the final video. Experiments demonstrate that our method produces diverse, physically and visually realistic weather effects. Furthermore, we show that our synthesized data significantly improves the robustness of autonomous driving semantic segmentation under adverse weather conditions. Project page: https://jumponthemoon.github.io/w-crafter/.",
      "short_abstract": "Weather synthesis aims to add weather effects to input videos while preserving scene identity, structure, and motion. The key limitation of existing methods is the lack of diversity in weather appearance and effective control over weather dynamics (e.g.,...",
      "published": "2026-06-27",
      "updated": "2026-06-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Perception & Sensor Fusion: abstract: scene segmentation",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 53,
      "arxiv_url": "https://arxiv.org/abs/2606.29020",
      "pdf_url": "https://arxiv.org/pdf/2606.29020.pdf"
    },
    {
      "id": "2606.28716",
      "anchor": "paper-2606-28716",
      "title": "TrajRS: Towards Certified Robustness in Pedestrian Trajectory Prediction",
      "authors": [
        "Liang Zhang",
        "Gaojie Jin",
        "Yao Shi",
        "Quanzhi Li",
        "Cheng-Chao Huang",
        "David N. Jansen",
        "Lijun Zhang"
      ],
      "abstract": "The robustness of trajectory prediction models is crucial for developing safe autonomous driving systems. Adversarial attacks on trajectory prediction can significantly impair the accuracy of predicted trajectories, leading to hazardous driving behaviors. While heuristic defense strategies have been implemented to enhance the robustness of trajectory prediction models, these measures often fail against more sophisticated, targeted adversarial attacks. Hence, there is a pressing need to establish verifiable safety assurances for trajectory prediction models. In this paper, we extend the traditional Randomized Smoothing framework to \"TrajRS\", which provides a certified robust radius for smoothed trajectory predictors. We clarify and expand the formal definitions of robustness in trajectory prediction and tailor the practical TrajRS scheme specifically to \"robustness for the optimal prediction\" and \"robustness for all possible predictions\". An extensive set of experiments demonstrates that TrajRS effectively achieves robustness certification for all smoothed pedestrian trajectory predictors in this work.",
      "short_abstract": "The robustness of trajectory prediction models is crucial for developing safe autonomous driving systems. Adversarial attacks on trajectory prediction can significantly impair the accuracy of predicted trajectories, leading to hazardous driving behaviors....",
      "published": "2026-06-26",
      "updated": "2026-06-26",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 12,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 54,
      "arxiv_url": "https://arxiv.org/abs/2606.28716",
      "pdf_url": "https://arxiv.org/pdf/2606.28716.pdf"
    },
    {
      "id": "2606.27900",
      "anchor": "paper-2606-27900",
      "title": "Long-Term Prediction of Local and Global Human Motion with Occlusion Recovery",
      "authors": [
        "Qiaoyue Yang",
        "Sven Heutger",
        "Christopher Niemann",
        "Magnus Jung",
        "Ayoub Al-Hamadi",
        "Sven Wachsmuth"
      ],
      "abstract": "Human motion describes the three-dimensional full-body movement of a person. Anticipating such motion holds significant relevance across a wide range of application domains such as human-robot interaction, autonomous driving, animation, and healthcare. In recent research, spatial and temporal dependencies are modeled by bidirectional attention mechanisms. These typically anticipate human motion in an autoregressive manner which could cause an accumulation of errors over time. As a consequence, they solely focus on local pose forecasting. To address these limitations, we propose a non-autoregressive transformer based on spatio-temporal attention, and train it not only for local pose anticipation, but also for global motion prediction in space. Furthermore, to enhance its applicability in real-world scenarios, our model is also trained to recover missing joints due to occlusions, and is capable of processing varying lengths of history observations. Our code is publicly available at https://github.com/Q-Y-Yang/Prediction-of-Local-and-Global-Human-Motion.",
      "short_abstract": "Human motion describes the three-dimensional full-body movement of a person. Anticipating such motion holds significant relevance across a wide range of application domains such as human-robot interaction, autonomous driving, animation, and healthcare. In...",
      "published": "2026-06-26",
      "updated": "2026-06-26",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 16,
        "evidence": [
          "abstract: motion forecasting",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 54,
      "arxiv_url": "https://arxiv.org/abs/2606.27900",
      "pdf_url": "https://arxiv.org/pdf/2606.27900.pdf"
    },
    {
      "id": "2606.27750",
      "anchor": "paper-2606-27750",
      "title": "Lightweight Multi-Vehicle Collaborative Perception Acceleration with Fusion Position Adjustment",
      "authors": [
        "Wenzhao Zhang",
        "Shujun Han",
        "Haixiao Gao",
        "Mengying Sun",
        "Bizhu Wang",
        "Xiaodong Xu"
      ],
      "abstract": "Multi-vehicle collaborative perception (MvCP) is considered as a key technology to facilitate automated driving (AD), where real-time MvCP under limited resources is significant for reliable AD. In this paper, we formulate a lightweight acceleration scheme for intermediate-fusion (IF) MvCP, which can adapt to both situations of limited computation and communication resources. We provide a relaxed definition conditional additivity and analyze the conditional additivity for various DNN linear layers. On this basis, we focus on the IF-MvCP based on additive feature fusion, and derive the MvCP precision consistency of the forward and backward feature fusion position (FP) adjustments among linear layers. Through experiments, we further validate the precision consistency of the FP adjustment method. Moreover, we propose an FP adjustment among linear layers (FALL) scheme for MvCP acceleration without precision loss theoretically. Simulation results show that the proposed FALL can reduce MvCP latency by up to 74.8% under limited communication resources and by up to 30.3% under limited computation resources.",
      "short_abstract": "Multi-vehicle collaborative perception (MvCP) is considered as a key technology to facilitate automated driving (AD), where real-time MvCP under limited resources is significant for reliable AD. In this paper, we formulate a lightweight acceleration scheme...",
      "published": "2026-06-26",
      "updated": "2026-06-26",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 7,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 54,
      "arxiv_url": "https://arxiv.org/abs/2606.27750",
      "pdf_url": "https://arxiv.org/pdf/2606.27750.pdf"
    },
    {
      "id": "2606.27660",
      "anchor": "paper-2606-27660",
      "title": "MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving",
      "authors": [
        "Nan Yang",
        "Zhanwen Liu",
        "Linfeng Zhang",
        "Shangyu Xie",
        "Yang Wang",
        "Wenzhuo Zhou",
        "Xiangmo Zhao"
      ],
      "abstract": "Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly in standard multi-view settings. Existing token pruning methods employ fixed pruning rate allocation and static importance metrics, ignoring dynamic inter-view importance differences and the evolving information importance during inference. Our analysis reveals that multi-view VLMs inherently encode task-related view priors in deeper layers and exhibit dynamic information requirements. Motivated by these findings, we propose MVPruner, a two-stage adaptive token pruning method that aligns pruning behavior with the model's dynamic information requirements. The first stage allocates pruning budgets based on the information diversity of each view, and retains tokens with consistent contribution across stages, ensuring semantic representational capacity. The second stage allocates budgets and selects tokens guided by instruction text to guarantee task alignment. Experimental results on four benchmarks demonstrate the superior performance of our method. For example, DriveMM equipped with MVPruner achieves 87.3% reduction in FLOPs, 4.97* speedup in prefilling phase while retaining 98.5% accuracy on DriveLM benchmark.",
      "short_abstract": "Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly in standard multi-view settings. Existing token pruning methods employ fixed...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.27660",
      "pdf_url": "https://arxiv.org/pdf/2606.27660.pdf"
    },
    {
      "id": "2606.27644",
      "anchor": "paper-2606-27644",
      "title": "CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations",
      "authors": [
        "Kyumin Hwang",
        "Wonhyeok Choi",
        "Jaeyeul Kim",
        "Jihun Park",
        "Daehee Park",
        "Sunghoon Im"
      ],
      "abstract": "This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for autonomous driving. Occupancy world models -- forecasting the future driving environment and planning the driving trajectory -- effectively bridge perception and planning, but current approaches often heavily rely on external modalities or large language models, failing to fully exploit the inherent structural potential of occupancy representations themselves. To enhance representational capacity for complex 3D scenes, we integrate a cascaded Vector Quantized (VQ) mechanism into an autoregressive framework. Following a coarse-to-fine principle, CascadeOcc progressively refines fine-grained details from global structures through a multi-scale architecture. Additionally, we incorporate a TimeMixer to capture multi-scale temporal dependencies, establishing a dual-hierarchy mechanism in both space and time. Experimental results on 4D occupancy forecasting and motion planning benchmarks demonstrate that CascadeOcc achieves superior performance among vision-centric approaches, validating that optimizing inherent representations is a powerful alternative to relying on external foundation models.",
      "short_abstract": "This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for autonomous driving. Occupancy world models -- forecasting the future driving environment and planning the...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 8,
        "evidence": [
          "abstract: forecasting",
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.27644",
      "pdf_url": "https://arxiv.org/pdf/2606.27644.pdf"
    },
    {
      "id": "2606.27554",
      "anchor": "paper-2606-27554",
      "title": "Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics",
      "authors": [
        "Tim Alexander Bader",
        "Tim Dieter Eberhardt",
        "Maximilian Dillitzer",
        "Wilhelm Stork"
      ],
      "abstract": "Camera-based perception systems for autonomous driving are typically developed and evaluated using fixed sensor rigs, while real-world vehicle fleets exhibit substantial variation in camera placement, orientation, field of view, and camera count. This mismatch introduces a cross-rig domain gap in which only the geometric observation process changes. To study this effect under controlled conditions, we introduce Plentiful CARLA Camera Rigs, a benchmark that renders identical driving scenes under 14 systematically designed camera rigs. This setup enables direct analysis of cross-rig generalization without confounding changes in scene content or appearance. Using the benchmark, we analyze cross-rig transfer behavior of representative multi-view perception architectures and observe substantial performance shifts induced by geometric rig variation. To facilitate structured analysis, we further introduce two calibration-based descriptors derived from rig metadata: Rig Variance, capturing internal rig diversity, and Rig Contrastive Distance, measuring geometric discrepancy between rigs. Our experiments show that geometric rig differences strongly correlate with relative cross-rig performance shifts and that Rig Contrastive Distance provides a reliable proxy for ranking transfer difficulty between sensor rigs.",
      "short_abstract": "Camera-based perception systems for autonomous driving are typically developed and evaluated using fixed sensor rigs, while real-world vehicle fleets exhibit substantial variation in camera placement, orientation, field of view, and camera count. This...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark",
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 14,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.27554",
      "pdf_url": "https://arxiv.org/pdf/2606.27554.pdf"
    },
    {
      "id": "2606.27504",
      "anchor": "paper-2606-27504",
      "title": "ReWorld: Learning Better Representations for World Action Models",
      "authors": [
        "Tianze Xia",
        "Lijun Zhou",
        "Kaixin Xiong",
        "Jingfeng Yao",
        "Yu Zhu",
        "Zhenxin Zhu",
        "Bing Wang",
        "Guang Chen",
        "Hangjun Ye",
        "Wenyu Liu",
        "Haiyang Sun",
        "Xinggang Wang"
      ],
      "abstract": "World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, existing approaches focus largely on model architecture design, and how a WAM can efficiently learn better world representations for planning remains underexplored. To address this gap, we propose ReWorld, the first representation learning framework specifically designed for autonomous-driving world action models. In WAMs, standard training supervises only the output ends of the generation and planning modules, leaving the intermediate representations that carry world knowledge to be shaped only indirectly, as byproducts of fitting these outputs. The core idea of ReWorld is to treat intermediate representations as direct targets of optimization, shaping them along three complementary dimensions. On the Video DiT responsible for generation, we impose future-predictive supervision on its intermediate representations. On the Action DiT responsible for planning, we first align its intermediate representations cross-modally with the video world representation, then further shape them to be discriminative around safety-critical boundaries via hard-negative supervision. In addition, we systematically analyze the effectiveness of existing representation learning methods in video generation world models, and discuss why their performance is limited on this task. Experiments on nuScenes and NAVSIM show that ReWorld improves fine-tuned video generation by 23.9% in FVD (81.3 to 61.9), raises closed-loop PDMS from 89.1 to 90.4 without any post-training such as RL or post-processing, and accelerates from-scratch convergence by approximately 2x.",
      "short_abstract": "World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, existing approaches focus largely on model architecture design, and how a WAM can efficiently learn better...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.27504",
      "pdf_url": "https://arxiv.org/pdf/2606.27504.pdf"
    },
    {
      "id": "2606.26661",
      "anchor": "paper-2606-26661",
      "title": "LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction",
      "authors": [
        "Sangjin Han",
        "Hoseong Jung",
        "Jeongtae Her",
        "Changhyun Choi",
        "H. Jin Kim"
      ],
      "abstract": "Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. While existing predictors excel at minimizing standard displacement errors, they often overlook the adherence to lane topology of multimodal predictions, particularly for lower-probability modes. Consequently, predicted trajectories may violate physical and logical constraints, making the prediction set unreliable for safety-critical planning. In this paper, we propose LAMP (Lane-Aligned Motion Primitives), a topology-aware forecasting framework that anchors multimodal prediction to structured motion primitives aligned with lane topology. Specifically, we use a VQ-VAE to learn shape-aware motion primitives as discrete intention queries, capturing spatiotemporal patterns beyond endpoint-based intentions. We further introduce a feasibility-aware intention selector trained with a lane-topology prior for filtering unreachable intention queries, guiding the decoder to prioritize topology-consistent intentions while preserving behavioral diversity. Extensive experiments on the Argoverse 2 dataset demonstrate that LAMP achieves prediction accuracy comparable to state-of-the-art baselines while outperforming them in feasibility and diversity metrics.",
      "short_abstract": "Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. While existing predictors excel at minimizing standard displacement errors, they often overlook the adherence to lane...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 24,
        "evidence": [
          "title: trajectory prediction",
          "abstract: motion forecasting",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.26661",
      "pdf_url": "https://arxiv.org/pdf/2606.26661.pdf"
    },
    {
      "id": "2606.27656",
      "anchor": "paper-2606-27656",
      "title": "Characterizing Driver Interactions with Autonomous Vehicles via Response Maps",
      "authors": [
        "Dave Broaddus",
        "Rachel DiPirro",
        "Chishang",
        "Yang",
        "Dan Calderone",
        "Wendy Ju",
        "Meeko Oishi"
      ],
      "abstract": "Understanding human responses to autonomous vehicle (AV) behaviors is essential for socially aware interaction, which is crucial for socially compatible navigation in shared traffic environments. We characterize human driving responses in interactions with AVs as feedback laws over the coupled state space of the human driven vehicle and the AV. We model the human driver's actions using a response map, a concept based in game theory, and employ a linear representation to capture driver behaviors as a function of AV behaviors, based on empirical data from a driving simulator study. Our results show that 1) human driver acceleration behavior can be captured using response maps, and 2) human driver responses differ significantly with respect to AV behaviors of yielding, non-yielding, and responsive to the human driver.",
      "short_abstract": "Understanding human responses to autonomous vehicle (AV) behaviors is essential for socially aware interaction, which is crucial for socially compatible navigation in shared traffic environments. We characterize human driving responses in interactions with...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: human driver",
          "Control & Vehicle Dynamics: abstract: steering or braking"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.27656",
      "pdf_url": "https://arxiv.org/pdf/2606.27656.pdf"
    },
    {
      "id": "2606.27123",
      "anchor": "paper-2606-27123",
      "title": "Proposal-Conditioned Latent Diffusion for Closed-Loop Traffic Scenario Generation",
      "authors": [
        "Shubham Vaijanath Phoolari",
        "Aleyna Kara",
        "Christoph Lauer",
        "Steven Peters"
      ],
      "abstract": "Closed-loop traffic simulation remains challenging because it must generate interactive multi-agent behaviors that are scene-consistent and controllable throughout rollout. Prior diffusion-based approaches achieve strong realism, but their computational cost can hinder deployment in time-constrained replanning loops for autonomous vehicle planning and simulation. We present a diffusion-based scenario generation framework conditioned on instance-centric scene context and multimodal proposal priors, with optional test-time guidance for shaping safety-critical behaviors. A compact action-latent representation and proposal-based initialization improve sampling efficiency and reduce per-step runtime without retraining. Experiments on the Waymo Open Motion Dataset demonstrate a favorable balance among realism, safety, and controllability across diverse interactive scenarios, while showing that test-time guidance enables systematic trade-offs among competing objectives.",
      "short_abstract": "Closed-loop traffic simulation remains challenging because it must generate interactive multi-agent behaviors that are scene-consistent and controllable throughout rollout. Prior diffusion-based approaches achieve strong realism, but their computational cost...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.27123",
      "pdf_url": "https://arxiv.org/pdf/2606.27123.pdf"
    },
    {
      "id": "2606.26858",
      "anchor": "paper-2606-26858",
      "title": "PlanRL: A Trajectory Planning Architecture for Reinforcement Learning-based Driving Experts",
      "authors": [
        "Joonhee Lim",
        "Yongjae Lee",
        "Jangho Shin",
        "Dongsuk Kum"
      ],
      "abstract": "Reinforcement learning (RL) has become a prominent framework for developing driving experts in autonomous vehicles. However, most existing RL-based experts are designed to output direct control commands (e.g., throttle, steering), which suffer from a lack of interpretability, high spatial complexity in learning road geometries, and poor compatibility with modern end-to-end planning architectures. To address these limitations, we propose a novel trajectory planning architecture for RL driving experts that integrates an RL policy with a polynomial-based trajectory planner. By employing a Frenet-frame coordinate system, our method simplifies complex road geometries into a curvilinear framework, offering a structured coordinate prior that facilitates policy learning. Furthermore, we incorporate a kinematic feasibility check into the planning stage to ensure that generated trajectories remain within the vehicle's physical limits, effectively mitigating cumulative tracking errors typically found in planning-based systems. We evaluate our approach on key CARLA benchmarks, where it significantly outperforms existing state-of-the-art control-based RL experts. On the CARLA Offline Leaderboard v1 and NoCrash benchmarks, our method improves the driving score by 5% and 11%, respectively, and increases the success rate by 8% and 19%.",
      "short_abstract": "Reinforcement learning (RL) has become a prominent framework for developing driving experts in autonomous vehicles. However, most existing RL-based experts are designed to output direct control commands (e.g., throttle, steering), which suffer from a lack of...",
      "published": "2026-06-25",
      "updated": "2026-06-25",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 28,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 55,
      "arxiv_url": "https://arxiv.org/abs/2606.26858",
      "pdf_url": "https://arxiv.org/pdf/2606.26858.pdf"
    },
    {
      "id": "2606.26574",
      "anchor": "paper-2606-26574",
      "title": "Revisiting Action Factorization for Complex Action Spaces",
      "authors": [
        "Timothy Flavin",
        "Sandip Sen"
      ],
      "abstract": "Many real-world control problems involve hybrid discrete-continuous action spaces. For example, steering and signaling in autonomous driving, and aiming and firing in robotics or video-games. Despite real-world hybrid factorization and reinforcement learning framework support for complex action spaces (e.g., Gymnasium, PettingZoo, TorchRL, SeedRL, Mujoco, etc), the default environments within those frameworks often implement uniform action space configurations (LunarLander, Walker2D, Cheetah, SMAC, SUMO, Ant, Atari). Landmark hybrid-action benchmarks (RoboCup 2D HFO, SC2LE, Platform, CARLA, etc) are mostly heavyweight or archival implementations originating from papers which test one or a small number of competing factorization methods on one kind of control. This article provides a cross-sectional study of factorization methods [independent networks, shared encoder, VDN, QPLEX, Joint, Auto-Regressive] on each of three families of algorithms [PPO, SAC, DQN] across three action spaces [discretized, hybrid, continuous] over four lightweight environments [Platform, hybrid-LunarLander, Hybrid-Shoot, CoopPush]. Accounting for some invalid pairings such as joint-continuous, we are left with 220 configurations to analyze each method. We provide two new C++ parallel gymnasium and petting-zoo compliant environments [CoopPush, Hybrid-Shoot] to isolate particular challenges such as state-dependent inter-action dependence. Finally, we introduce VDN-PPO and PPO-MIX which use a branching critic to assign credit to multi-headed PPO. These variants out-perform all other tested PPO factorizations. Our results suggest that branching dueling architectures balance compute and performance most effectively, with Auto-Regressive actions reaching the highest performance overall and native continuous SAC outperforming discrete and hybrid algorithms, albiet both at increased computational cost.",
      "short_abstract": "Many real-world control problems involve hybrid discrete-continuous action spaces. For example, steering and signaling in autonomous driving, and aiming and firing in robotics or video-games. Despite real-world hybrid factorization and reinforcement learning...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control",
          "Control & Vehicle Dynamics: abstract: steering or braking",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.26574",
      "pdf_url": "https://arxiv.org/pdf/2606.26574.pdf"
    },
    {
      "id": "2606.26456",
      "anchor": "paper-2606-26456",
      "title": "Towards Safety-Aware Mutation Testing for Autonomous Driving Systems",
      "authors": [
        "Donghwan Shin"
      ],
      "abstract": "Simulation-based testing is essential for ensuring the safety of Autonomous Driving Systems (ADS), yet the community lacks a systematic criterion for determining when we can safely stop additional test scenario generation. Existing coverage metrics typically focus on individual component reliability or treat the ADS as a black box, failing to capture certain component interactions that cause most ADS accidents. While traditional mutation testing provides a falsifiable measure of test adequacy, directly porting code- and deep learning model-level mutations to the corresponding modules of ADS is insufficient. In this vision paper, we propose a paradigm shift toward Safety-Aware Mutation Testing (SAMT). Unlike traditional mutation testing, which creates mutants (i.e., faulty versions of the software under test) by injecting artificial faults into individual components, SAMT systematically injects temporally bounded faults into the messages exchanged between ADS modules to simulate realistic interaction failures. To ensure these mutants represent genuine hazards, we propose deriving mutant generation rules directly from top-down safety engineering frameworks, such as System-Theoretic Process Analysis (STPA). By embedding systems thinking into the mutation testing pipeline, SAMT provides a rigorous mechanism for evaluating test adequacy, enabling automated scenario generation, and guiding ADS repair. We also outline critical open challenges.",
      "short_abstract": "Simulation-based testing is essential for ensuring the safety of Autonomous Driving Systems (ADS), yet the community lacks a systematic criterion for determining when we can safely stop additional test scenario generation. Existing coverage metrics typically...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 30,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.26456",
      "pdf_url": "https://arxiv.org/pdf/2606.26456.pdf"
    },
    {
      "id": "2606.26424",
      "anchor": "paper-2606-26424",
      "title": "Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs",
      "authors": [
        "Qiyuan Wu",
        "Katie Z Luo",
        "Bharath Hariharan",
        "Wei-Lun Chao",
        "Mark Campbell"
      ],
      "abstract": "Trajectory forecasting for autonomous driving has advanced rapidly, yet representative models often produce uninformative posteriors over forecast modes, causing problems for mode pruning. We trace this to a modeling-training mismatch: forecasters are typically modeled as conditional Gaussian mixture models (GMMs) but trained with a winner-take-all (WTA) loss that assigns each sample to its nearest mode. We argue that this K-means-like hard assignment (one-hot), while preventing mode collapse, is the source of uninformative mode probabilities: it over-segments the trajectory space, ignores relatedness among nearby modes, and yields assignment instability under small perturbations. Guided by this lens, we introduce two post-hoc treatments: (1) test-time posterior-weighted merging that aggregates nearby candidate trajectories; and (2) a one-step expectation-maximization (EM) update that replaces hard labels with soft responsibilities, sharing probability mass across neighboring modes. Across several WTA-trained architectures, these lightweight steps produce more informative, faithfully ranked mode posteriors and strengthen final forecasts on popular displacement metrics -- without retraining. Our analysis unifies recent design choices through a GMM-vs-K-means perspective and offers principled, practical corrections that better align training objectives with inference.",
      "short_abstract": "Trajectory forecasting for autonomous driving has advanced rapidly, yet representative models often produce uninformative posteriors over forecast modes, causing problems for mode pruning. We trace this to a modeling-training mismatch: forecasters are...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "title: forecasting",
          "abstract: forecasting"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.26424",
      "pdf_url": "https://arxiv.org/pdf/2606.26424.pdf"
    },
    {
      "id": "2606.26017",
      "anchor": "paper-2606-26017",
      "title": "G2DP: Diffusion Planning with Spatio-Temporal Grid Guidance",
      "authors": [
        "Hang Yu",
        "Ye Jin",
        "Alessandro Canevaro",
        "Julian Schmidt",
        "Julian Jordan",
        "Peizheng Li",
        "Marc Kaufeld",
        "Silvan Lindner",
        "Johannes Betz",
        "Wilhelm Stork"
      ],
      "abstract": "In autonomous driving, diffusion-based planners have emerged as a promising paradigm for robust motion planning in dense and interactive traffic, as they can effectively model diverse driving behaviors. However, their inherent stochasticity often requires explicit guidance during denoising to ensure safety and route adherence for robust closed-loop execution. Existing guidance typically relies on sparse, entity-centric geometric queries or post-hoc refinement, yielding limited situational awareness and fragile performance in interactive scenes. To address this issue, we propose G2DP (Grid-Guided Diffusion Planning), a diffusion-based planner that directly enforces dense environmental constraints through inference-time guidance. Specifically, G2DP constructs a differentiable spatio-temporal cost volume by fusing probabilistic future occupancy distributions with a route-progress map. By formulating this volume as a continuous safety energy functional, it injects dense gradients directly into the denoising loop, actively steering trajectory generation toward collision-free and progress-optimal regions. Extensive closed-loop evaluations show that G2DP achieves state-of-the-art performance on nuPlan, outperforming the strongest imitation-learning baseline by +7.2 points in reactive score. It further maintains top scores in zero-shot transfers to interPlan and DeepScenario benchmarks, with collision avoidance improving by +10.15 over the unguided approach on interPlan. These results demonstrate that spatio-temporal cost grids serve as an effective representation for robust guidance in diffusion-based planning. Code is available at https://github.com/HangYuu/G2DP.",
      "short_abstract": "In autonomous driving, diffusion-based planners have emerged as a promising paradigm for robust motion planning in dense and interactive traffic, as they can effectively model diverse driving behaviors. However, their inherent stochasticity often requires...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 14,
        "evidence": [
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: trajectory generation"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.26017",
      "pdf_url": "https://arxiv.org/pdf/2606.26017.pdf"
    },
    {
      "id": "2606.25953",
      "anchor": "paper-2606-25953",
      "title": "DSP-SLAM++: A Unified Framework for Multi-Class, High-Fidelity Object SLAM in the Wild",
      "authors": [
        "Ahmad Kourani",
        "Ghina Daoud",
        "Daniel Asmar",
        "Imad Elhajj"
      ],
      "abstract": "Existing object-aware SLAM systems force a trade-off between real-time performance, multi-class support, and the generation of high-fidelity, semantically coherent object models. To address this trade-off, we present DSP-SLAM++, which extends the DSP-SLAM framework with an asynchronous mapping pipeline for real-time performance and dedicated sensor fusion adaptations for a monocular fisheye-LiDAR suite. Experiments demonstrate that our system generates fine-grained, geometrically-complete shapes for multiple object classes while eliminating severe mapping thread bottlenecks by reducing maximum object processing latency by up to 70\\% compared to the state-of-the-art baseline, enabling robust, real-time performance on a challenging 25 Hz multi-class datasets. This work makes high-fidelity, multi-class object SLAM more practical for real-world applications like autonomous driving and robotic manipulation by enabling its use on platforms with common fisheye-LiDAR sensor setups. The open-source code is available at: [github.com/AUBVRL/DSP-SLAMpp].",
      "short_abstract": "Existing object-aware SLAM systems force a trade-off between real-time performance, multi-class support, and the generation of high-fidelity, semantically coherent object models. To address this trade-off, we present DSP-SLAM++, which extends the DSP-SLAM...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 21,
        "evidence": [
          "title: SLAM",
          "abstract: SLAM",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.25953",
      "pdf_url": "https://arxiv.org/pdf/2606.25953.pdf"
    },
    {
      "id": "2606.25736",
      "anchor": "paper-2606-25736",
      "title": "UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving",
      "authors": [
        "Bo Zhao",
        "Xinting Zhao",
        "Naifan Li",
        "Erkang Cheng",
        "Haibin Ling"
      ],
      "abstract": "Diffusion models have shown strong potential for multi-modal planning in end-to-end autonomous driving. However, most existing methods confine diffusion to the planning module, conditioning on fixed outputs from separate discriminative perception networks. This decoupled design propagates perception errors to the planner, increasing optimization difficulty and reducing robustness. To overcome these limitations, we propose UniTeD, a Unified Temporal Diffusion framework that jointly models perception and planning through iterative denoising in a shared generative space. By enabling bidirectional information exchange, the framework facilitates mutual refinement between tasks and improves robustness via noise-conditioned multi-task training. We further extend this unified diffusion paradigm to a streaming setting by incorporating temporal context. A Temporal Transition Module (TTM) is introduced to resolve the noise-level mismatch between historical and current frames. In addition, we propose an Anchor Refresh Strategy (ARS) to alleviate the training-inference distribution shift commonly observed in sparse diffusion-based end-to-end driving frameworks. Without bells and whistles, UniTeD achieves state-of-the-art performance across multiple benchmarks, surpassing both recent discriminative end-to-end methods and diffusion-based planning approaches.",
      "short_abstract": "Diffusion models have shown strong potential for multi-modal planning in end-to-end autonomous driving. However, most existing methods confine diffusion to the planning module, conditioning on fixed outputs from separate discriminative perception networks....",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: title: perception",
          "Perception & Sensor Fusion: abstract: perception",
          "Planning & Decision-Making: title: planning",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.25736",
      "pdf_url": "https://arxiv.org/pdf/2606.25736.pdf"
    },
    {
      "id": "2606.25652",
      "anchor": "paper-2606-25652",
      "title": "Auto-Labelling-Based Domain Transfer for 3D Object Detection on a Bicycle-Mounted LiDAR Platform",
      "authors": [
        "Mario Finkbeiner",
        "Max A. Buettner",
        "Kanak Mazumder",
        "Fabian B. Flohr"
      ],
      "abstract": "Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requirement for autonomous driving (AD). Alongside advances in vehicle-based perception, research increasingly equips bicycles with sensors to study traffic from a perspective native to VRUs. Such platforms still rely on LiDAR detectors originally trained on vehicle data, yet annotated 3D data from a cyclist's perspective is scarce. How well these detectors generalise to this setting has not been evaluated. We present a 3D object detection benchmark of 1,027 annotated LiDAR keyframes (over 18,000 3D bounding boxes) from the FUSE-Bike platform in urban Munich. We evaluate four nuScenes-pre-trained detectors against 1,854 human-verified ground-truth (GT) boxes both in their original form and after finetuning on training labels produced by a VRU-dedicated auto-labelling pipeline that requires no manual annotation. The zero-shot domain gap is concentrated on the VRU classes. Finetuning recovers most of it, improving mean average precision (mAP) by up to 23.4 points with the largest gains on pedestrians and cyclists, and the adapted detectors even surpass the quality of the auto-labels they were trained on. The benchmark provides a reproducible baseline for VRU-centric 3D detection and shows that auto-labels are a viable substitute for manual annotation when adapting vehicle-trained detectors to a cyclist platform.",
      "short_abstract": "Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requirement for autonomous driving (AD). Alongside advances in vehicle-based perception, research increasingly...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 27,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.25652",
      "pdf_url": "https://arxiv.org/pdf/2606.25652.pdf"
    },
    {
      "id": "2606.25626",
      "anchor": "paper-2606-25626",
      "title": "Reasonable Motion: A General ASP Foundation for Environment Constrained Movement Trajectory Computation",
      "authors": [
        "Julius Monsen",
        "Jakob Suchan",
        "Mehul Bhatt",
        "Lars Karlsson"
      ],
      "abstract": "We present a general answer set programming based hybrid quantitative-qualitative method for computing constrained branching trajectory modes for moving objects in real-world settings. The method performs constrained traversal of an environment graph, enumerating geometrically admissible motion behaviours as stable models, each constituting a distinct trajectory mode characterised by both domain-dependent and independent factors such as derived event sequence, map topology, and domain norms. The hybrid trajectory computation method is generally applicable across motion characteristics typically encountered in diverse dynamic domains with moving objects, e.g., autonomous driving. We demonstrate applicability and highlight how computed trajectories are traceable to their underlying stable model, thereby affording verifiable interpretability that purely learned approaches cannot provide. We also perform an empirical evaluation with Argoverse 2, a large-scale real-world autonomous driving benchmark representative of the class of dynamic domains within the scope of the proposed method.",
      "short_abstract": "We present a general answer set programming based hybrid quantitative-qualitative method for computing constrained branching trajectory modes for moving objects in real-world settings. The method performs constrained traversal of an environment graph,...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.25626",
      "pdf_url": "https://arxiv.org/pdf/2606.25626.pdf"
    },
    {
      "id": "2606.25509",
      "anchor": "paper-2606-25509",
      "title": "ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving",
      "authors": [
        "Sining Ang",
        "Yuan Chen",
        "Liu Haiyan",
        "Xuanyao Mao",
        "Jason Bao",
        "Xuliang",
        "Bingchuan Sun",
        "Yan Wang"
      ],
      "abstract": "Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast-slow planners often rely on hand-designed triggering rules that either over-call the slow system or call it at the wrong times. We formulate slow-system invocation as a resource-aware sequential decision problem and propose the Adaptive Slow-System Control Gate (ASSCG), which makes frame-level Query/Cache/Drop decisions to refresh, reuse, or suppress slow guidance. ASSCG uses an RWKV backbone for efficient long-horizon gating and is trained with supervised fine-tuning followed by GRPO-style compute-aware reinforcement fine-tuning. We apply ASSCG to two different fast-slow architectures: (i) AsyncDriver on nuPlan Hard20 closed-loop evaluation, where ASSCG improves score to 67.28 (+2.28) while reducing average end-to-end inference latency by 60%; and (ii) a RecogDrive-based dual system that we build by replacing its original VLM-2B module with a lightweight ViT-based fast planner and adding an LLM slow planner, evaluated on NAVSIM, where ASSCG achieves 91.4 PDMS (+0.6) and increases average speed by 25%. The project page, including video visualizations and additional results, is available at https://williamxuanyu.github.io/asscg/.",
      "short_abstract": "Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast-slow planners often rely on hand-designed triggering rules that either over-call the slow system or call it at the wrong times. We formulate...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 9,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.25509",
      "pdf_url": "https://arxiv.org/pdf/2606.25509.pdf"
    },
    {
      "id": "2606.25711",
      "anchor": "paper-2606-25711",
      "title": "Distributed SDN-Based Communication Architecture for the Pods4Rail System",
      "authors": [
        "Dingyang Liu",
        "Dereje Mechal Molla",
        "Léo Mendiboure",
        "Marion Berbineau"
      ],
      "abstract": "Future multimodal transportation systems require reliable, low-latency communication infrastructures to coordinate autonomous vehicles and moving infrastructure across rail and road networks. Traditional centralized control architectures struggle to meet these requirements in highly dynamic environments due to increased latency, limited scalability, and poor adaptability to changing network conditions. To address this, we propose a distributed communication architecture integrating Software-Defined Networking (SDN) and Multi-Access Edge Computing (MEC), that can create a flexible, programmable and lowlatency network. Results show controller communication and flow setup latency between edge SDN controllers and Pods are lower than reported in literature. The framework uses hierarchical control with regional and edge controllers to support low-latency interface management and edge autonomy. Operational workflows and control logic are defined as representative scenarios. The architecture combines regional policy coordination with edge-level autonomy, enabling local failover and adaptive interface management without central dependency.",
      "short_abstract": "Future multimodal transportation systems require reliable, low-latency communication infrastructures to coordinate autonomous vehicles and moving infrastructure across rail and road networks. Traditional centralized control architectures struggle to meet...",
      "published": "2026-06-24",
      "updated": "2026-06-24",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 9,
        "evidence": [
          "abstract: edge or cloud",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 56,
      "arxiv_url": "https://arxiv.org/abs/2606.25711",
      "pdf_url": "https://arxiv.org/pdf/2606.25711.pdf"
    },
    {
      "id": "2606.25296",
      "anchor": "paper-2606-25296",
      "title": "SafeGen: LLM-Driven Assertion Generation and Fault Criticality Evaluation for Functional Safety",
      "authors": [
        "Xuanyi Tan",
        "Arjun Chaudhuri",
        "Rubin Parekhji",
        "Krishnendu Chakrabarty"
      ],
      "abstract": "With advances in autonomous driving and electric vehicle technologies, functional safety has become a critical requirement in automotive chip design. Traditional simulation-based fault analysis is often overly conservative at the module level and fails to accurately reflect fault criticality. This paper presents SafeGen, an LLM-driven, formal-verification-assisted framework for functional-safety-oriented fault criticality assessment. SafeGen leverages large language models (LLMs) and a document-level Hyper Knowledge Graph (HyperKG) that incorporates Failure Modes, Effects, and Diagnostic Analysis (FMEDA) guidelines to extract verifiable specifications from design and safety documents and evaluate their relevance to overall system safety. The HyperKG is further enriched with register-transfer-level (RTL) information to guide the generation of Functional Safety Assertions (FSAs) that are both semantically grounded and design-aware. Each assertion is linked to its corresponding specification, enabling traceable reasoning throughout the assessment process. A gate-to-RTL fault-mapping mechanism supporting both stuck-at and bridging faults, combined with formal property verification (FPV), enables semantic-level fault criticality grading based on specification-linked assertion violations. A digital-physical co-simulation platform for a field-oriented control (FOC) system is developed to validate SafeGen. Experimental results demonstrate that SafeGen generates higher-quality assertions than existing LLM-based assertion generation frameworks while providing greater semantic interpretability in fault criticality assessment compared with traditional simulation-based approaches.",
      "short_abstract": "With advances in autonomous driving and electric vehicle technologies, functional safety has become a critical requirement in automotive chip design. Traditional simulation-based fault analysis is often overly conservative at the module level and fails to...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 19,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: verification",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.25296",
      "pdf_url": "https://arxiv.org/pdf/2606.25296.pdf"
    },
    {
      "id": "2606.25239",
      "anchor": "paper-2606-25239",
      "title": "Tensor-Based Batch Fuzzing with Adaptive Perturbation Scaling for Deep Neural Networks",
      "authors": [
        "Guanqin Zhang",
        "Yulei Sui"
      ],
      "abstract": "Deep neural networks are increasingly deployed in safety-critical domains such as autonomous driving and medical diagnosis, yet their opaque, high-dimensional parameter spaces make it difficult to systematically assess model reliability on unseen inputs. Existing coverage-guided sequential fuzzing frameworks for DNNs inherit a one-input-per-iteration design from traditional software fuzzing and apply uniform perturbation budgets across all input dimensions, limiting both testing throughput (i.e., inputs processed per unit time) and the precision of input-space exploration. We present a new specification-aware batch fuzzing framework with adaptive perturbation scaling that addresses both limitations. Rather than relying on a fixed global perturbation radius epsilon, our approach derives mutation step sizes from specification-defined feasible ranges (the gap between lower and upper bounds) using a shared scale factor. This scaling can be applied either as a global scalar (isotropic) or as per-dimension step sizes (anisotropic), keeping perturbations consistent with the underlying constraint structure. As a result, the fuzzer can explore input spaces with heterogeneous feature scales more effectively across all specifications in the batch. We embed input constraints and output property checks directly into the network as non-trainable layers, yielding a wrapped model that processes B specification instances in a single batched iteration, substantially improving fuzzing efficiency and counterexample exploration. We evaluate our framework extensively on three benchmarks, covering six networks and over 400 specifications across TrafficSigns, Cifar100, and TinyImageNet. Our tensor-based fuzzing achieves up to 40X higher throughput and 4X more violations than the sequential baseline under the same time budget, demonstrating significantly improved effectiveness in specification-guided fuzzing.",
      "short_abstract": "Deep neural networks are increasingly deployed in safety-critical domains such as autonomous driving and medical diagnosis, yet their opaque, high-dimensional parameter spaces make it difficult to systematically assess model reliability on unseen inputs....",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 1,
        "evidence": [
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.25239",
      "pdf_url": "https://arxiv.org/pdf/2606.25239.pdf"
    },
    {
      "id": "2606.25134",
      "anchor": "paper-2606-25134",
      "title": "Causality-Based Parametric Control Barrier Function for Safe Multi-Vehicle Interaction",
      "authors": [
        "Yiwei Lyu",
        "Caleb Chang",
        "John M. Dolan"
      ],
      "abstract": "Safe control has been widely studied in various safety-critical applications, for instance, autonomous driving. In order to ensure the autonomous vehicle does not collide with other vehicles, it is essential to obtain an accurate expectation of surrounding vehicles' behavior and react adaptively. Instead of assuming fully cooperative and homogeneous vehicles using the same safety-critical controllers, recent works have been exploring different data-driven approaches to model the neighboring vehicles' underlying controllers with observed data. However, existing works either suffer from 1) the inter-vehicle influence during the multi-vehicle interaction, which makes it hard to determine the causality of surrounding vehicles' behavior in controller modeling, or 2) being dominated by the worst-case analysis, which may lead to overly conservative behavior. In this paper, we extend the prior work on Parametric-Control Barrier Function (Parametric-CBF) to multi-robot interactions with embedded causality inference to explicitly reason over the inter-vehicle influence. Given the learned Causality-based Parametric-CBF, we present an adaptive safety-critical controller that allows the ego vehicle to safely react to surrounding vehicles with the learned expectation. We demonstrate that by leveraging the motion flexibility among multi-vehicle systems, task efficiency can be greatly improved in various interaction-intensive scenarios.",
      "short_abstract": "Safe control has been widely studied in various safety-critical applications, for instance, autonomous driving. In order to ensure the autonomous vehicle does not collide with other vehicles, it is essential to obtain an accurate expectation of surrounding...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.25134",
      "pdf_url": "https://arxiv.org/pdf/2606.25134.pdf"
    },
    {
      "id": "2606.25127",
      "anchor": "paper-2606-25127",
      "title": "Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See",
      "authors": [
        "Mohamed Benabdelouahad",
        "Ahmed Djalal Hacini",
        "Nadir Farhi",
        "Aissa Boulmerka"
      ],
      "abstract": "We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Perceiver-based agents that share identical architectures and training data but differ only in their reward configurations$\\unicode{x2014}$ranging from basic violation penalties to continuous proximity penalties$\\unicode{x2014}$we analyze cross-attention allocation across 50 real-world scenarios from the Waymo Open Motion Dataset. A central methodological finding is that naïve pooling of timesteps across episodes substantially underestimates the attention$\\unicode{x2013}$risk relationship; within-episode correlation with Fisher z-transform aggregation is the appropriate statistic and reveals a robustly positive link between collision risk and agent-directed attention. Building on this validated methodology, we demonstrate two reward-conditioned effects: agents trained with navigation rewards allocate up to $2.0\\times$ more attention to GPS-path tokens than those trained with additional proximity penalties$\\unicode{x2014}$and $4.7\\times$ more than agents with no navigation incentive$\\unicode{x2014}$revealing that reward content directly determines which scene elements the encoder prioritizes, and continuous time-to-collision penalties create a $\\textit{learned vigilance prior}$$\\unicode{x2014}$elevated resting agent surveillance maintained throughout collision-free phases. In several scenarios, the complete-reward and minimal-reward models exhibit opposite attention$\\unicode{x2013}$risk correlation directions, demonstrating that reward design can qualitatively reverse attentional strategy rather than merely modulating its magnitude. These results suggest that attention analysis is a practical diagnostic for verifying that a reward function produces the intended representational behaviour in safety-critical RL systems.",
      "short_abstract": "We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Perceiver-based agents that share identical architectures and training data but differ only in their reward...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Mapping & Localization: abstract: GNSS or GPS"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.25127",
      "pdf_url": "https://arxiv.org/pdf/2606.25127.pdf"
    },
    {
      "id": "2606.24962",
      "anchor": "paper-2606-24962",
      "title": "Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models",
      "authors": [
        "Thibaut Kulak"
      ],
      "abstract": "Recent progress in large-scale sequence modeling has shown that a single model can learn useful representations across highly diverse data distributions. Inspired by these advances, we investigate whether a unified transformer policy can be trained across large collections of heterogeneous reinforcement learning environments. We introduce LDM-v0, a Large Decision Model trained offline on trajectories collected from thousands of environments spanning multiple domains and modalities. LDM-v0 is a multi-task, multi-modal transformer policy conditioned on histories of observations, actions, rewards, and termination signals, and trained through supervised next-action prediction over offline trajectories. We describe the environment infrastructure, automated data generation pipeline, model architecture, and training methodology used to build LDM-v0, and evaluate its performance across diverse environments. We show that a single pretrained model matches the performance of independently trained task-specific reference policies on approximately 1,000 environments including robotics, autonomous driving, inventory management, cybersecurity, trading, and video games. These results demonstrate the feasibility of large-scale offline pretraining across heterogeneous reinforcement learning environments using a single transformer policy.",
      "short_abstract": "Recent progress in large-scale sequence modeling has shown that a single model can learn useful representations across highly diverse data distributions. Inspired by these advances, we investigate whether a unified transformer policy can be trained across...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: security",
          "Human Factors & Policy: abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24962",
      "pdf_url": "https://arxiv.org/pdf/2606.24962.pdf"
    },
    {
      "id": "2606.24805",
      "anchor": "paper-2606-24805",
      "title": "DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection",
      "authors": [
        "Shiyi Mu",
        "Zichong Gu",
        "Zhiqi Ai",
        "Yilin Gao",
        "Shugong Xu"
      ],
      "abstract": "Stereo-based 3D obstacle perception for autonomous driving is currently constrained by an imbalanced triplet: deployment cost, detection accuracy, and open-set adaptability. While existing methods struggle to balance these three competing objectives, there is an urgent demand for high-precision, real-time algorithms capable of detecting arbitrary obstacles in the wild. In this paper, we present DDStereo, a novel Dual-Decoder Stereo Transformer that achieves a synergistic integration of 3D object detection and Out-of-Distribution (OoD) road anomaly detection. Leveraging the geometric priors of stereo disparity, our approach effectively couples 3D attribute regression with open-set foreground detection within a streamlined dual-branch decoder architecture. Conventional methods rely on complex feature-level fusion; DDStereo maintains execution efficiency by employing a decoupled decoding strategy and shared object-level queries to ensure cross-modal target alignment. Extensive evaluations of public benchmarks demonstrate that DDStereo not only achieves state-of-the-art accuracy under open-set and closed-set protocols. Our method delivers real-time performance comparable to monocular 3D detection baselines, providing a cost-effective solution for the perception of obstacles of the normal and OoD category. Code and models are available at https://github.com/shiyi-mu/DDStereo.",
      "short_abstract": "Stereo-based 3D obstacle perception for autonomous driving is currently constrained by an imbalanced triplet: deployment cost, detection accuracy, and open-set adaptability. While existing methods struggle to balance these three competing objectives, there is...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 4,
        "evidence": [
          "abstract: perception",
          "abstract: object detection"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24805",
      "pdf_url": "https://arxiv.org/pdf/2606.24805.pdf"
    },
    {
      "id": "2606.24796",
      "anchor": "paper-2606-24796",
      "title": "Pocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAM",
      "authors": [
        "Leshu Li",
        "Jie Peng",
        "Yang Zhao"
      ],
      "abstract": "3D Gaussian Splatting (3DGS) has garnered significant attention in Simultaneous Localization and Mapping (SLAM) due to its advances in capturing fine-grained geometry features and synthesizing novel views. For SLAM in large-scale scenes, such as autonomous driving, 3DGS-SLAM faces a critical limitation: memory consumption increases continuously over time as Gaussian points accumulate, leading to poor memory efficiency and limiting its applicability. In this work, we propose a rendering-area-aware pruning strategy that selectively removes Gaussians based on their contribution to the effective rendering area, rather than solely relying on Gaussian-level heuristics such as opacity or gradient magnitude. This perspective directly targets the sources of memory redundancy, effectively reducing the peak memory footprint of 3DGS-SLAM during runtime. Evaluations on the EuRoC and KITTI datasets demonstrate that our method consistently outperforms existing pruning approaches in large-scale outdoor scenes, achieving over 60% memory reduction and more than 2 times FPS improvement while preserving localization and mapping accuracy. These results highlight rendering-area-aware pruning as a promising direction for scaling 3DGS-SLAM to real-world autonomous driving scenarios. Our code is publicly available at https://github.com/UMN-ZhaoLab/Pocket-SLAM.git.",
      "short_abstract": "3D Gaussian Splatting (3DGS) has garnered significant attention in Simultaneous Localization and Mapping (SLAM) due to its advances in capturing fine-grained geometry features and synthesizing novel views. For SLAM in large-scale scenes, such as autonomous...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 33,
        "evidence": [
          "abstract: localization",
          "title: SLAM",
          "abstract: SLAM",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24796",
      "pdf_url": "https://arxiv.org/pdf/2606.24796.pdf"
    },
    {
      "id": "2606.24759",
      "anchor": "paper-2606-24759",
      "title": "UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving",
      "authors": [
        "Xiaowei Gao",
        "Pengxiang Li",
        "Yitai Cheng",
        "Ruihan Xu",
        "James Haworth",
        "Stephen Law",
        "Yun Ye"
      ],
      "abstract": "Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations. To address this gap, we propose UniDrive, a unified visual-language and grounding framework for interpretable risk understanding in autonomous driving. UniDrive combines a temporal reasoning branch that models scene dynamics from multi-frame visual input with a high-resolution perception branch that preserves fine-grained spatial details from the latest frame. The two branches are integrated through a gated cross-attention fusion module, enabling dynamic context to be aligned with precise spatial evidence. Based on the fused representation, UniDrive jointly generates natural-language risk descriptions and grounded bounding-box outputs for risk objects. Experiments on the DRAMA-Reasoning benchmark show that UniDrive outperforms representative image-based and video-based baselines in both captioning and risk-object grounding. In particular, UniDrive achieves the best overall performance on the validation split and demonstrates clear advantages in small-object localization, zero-shot generalization to NuScenes and BDD100K, and human-rated interpretability and trustworthiness. These results suggest that explicitly combining temporal semantics and high-resolution perception provides a stronger foundation for interpretable and safety-oriented autonomous driving systems. The code is available at https://github.com/pixeli99/unidrive-dev.",
      "short_abstract": "Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24759",
      "pdf_url": "https://arxiv.org/pdf/2606.24759.pdf"
    },
    {
      "id": "2606.24353",
      "anchor": "paper-2606-24353",
      "title": "Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints",
      "authors": [
        "Hojun Choi",
        "Seulbin Hwang",
        "Daejung Kim",
        "Kisung Kim",
        "Hyunjung Shim",
        "Jinhan Lee"
      ],
      "abstract": "Bird's-eye view (BEV) perception fuses multi-camera images into a unified top-down representation for autonomous driving. Despite recent progress, state-of-the-art methods remain confined to closed-set scenarios, making them vulnerable to unpredictable real-world environments. In this work, we introduce open-vocabulary BEV segmentation (OVBS), which leverages vision-language models (VLMs) to recognize categories beyond the training set while maintaining precise BEV perception and real-time efficiency. A key challenge in OVBS lies in the 3D geometric inconsistency inherent in the ill-posed lifting of 2D VLM semantics into BEV. To address this, we propose OVBEVSeg, a geometry-aware OVBS framework that enhances efficient Gaussian splatting (GS)-based unprojection by leveraging robust 3D geometric constraints across three progressive stages: (1) 2D-to-BEV pseudo-labeling via reliable 3D projection for OV generalization; (2) joint 2D-BEV per-scene optimization with BEV structural constraints for 3D geometric consistency; and (3) 3D geometric distillation for online efficiency. On the nuScenes dataset, OVBEVSeg achieves state-of-the-art performance, outperforming closed-set methods by 15.3 mIoU on unseen categories. Remarkably, even with no novel-class ground-truth labels, it remains competitive with self- and semi-supervised baselines trained with up to 40% of ground-truth annotations. Furthermore, it achieves 2.5x faster inference with only 0.22x the memory consumption of projection-based methods. Project page: https://hchoi256.github.io/projects/ovbevseg/.",
      "short_abstract": "Bird's-eye view (BEV) perception fuses multi-camera images into a unified top-down representation for autonomous driving. Despite recent progress, state-of-the-art methods remain confined to closed-set scenarios, making them vulnerable to unpredictable...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 10,
        "evidence": [
          "abstract: perception",
          "abstract: camera",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24353",
      "pdf_url": "https://arxiv.org/pdf/2606.24353.pdf"
    },
    {
      "id": "2606.24301",
      "anchor": "paper-2606-24301",
      "title": "MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving",
      "authors": [
        "Hongli Xiao",
        "Youjian Zhang",
        "Yucai Bai",
        "Chaoyue Wang",
        "Yaohui Jin",
        "Xiaoguang Ren",
        "Wenjing Yang",
        "Long Lan"
      ],
      "abstract": "Recovering realistic 3D vehicle models from autonomous driving scenes is crucial for synthesizing training data and building simulation environment. However, most existing vehicle generation methods fail to fully exploit multimodal sensors i.e. multi-view images and LiDAR point clouds) and rely on neural rendering based reconstruction, leading to low-quality mesh. Recently, native 3D generative models have made significant progress, yet they are not built for arbitrary multi-view inputs and often struggle with in-the-wild driving images. In this work, we present MM-TRELLIS, a multi-modal version of TRELLIS for in-the-wild 3D vehicle generation that integrates LiDAR and image sensors from autonomous driving datasets into native 3D generative models. Specifically, multi-view images are cycled as conditioning inputs, while LiDAR point clouds provide test-time guidance to ensure geometric accuracy and cross-view consistency. During denoising, we first align the guidance point cloud with the model priors, then enforce consistency between the generated geometry and the guidance point cloud. Finally, we introduce a voxel filtering strategy based on the opacity of 3D Gaussian Splatting to suppress floaters and produce clean meshes. Comprehensive experiments on Waymo dataset demonstrate our method outperforms existing methods in high-fidelity 3D vehicle generation. Code is available at https://github.com/HongliXiao/MM-TRELLIS.",
      "short_abstract": "Recovering realistic 3D vehicle models from autonomous driving scenes is crucial for synthesizing training data and building simulation environment. However, most existing vehicle generation methods fail to fully exploit multimodal sensors i.e. multi-view...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 6,
        "evidence": [
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24301",
      "pdf_url": "https://arxiv.org/pdf/2606.24301.pdf"
    },
    {
      "id": "2606.24257",
      "anchor": "paper-2606-24257",
      "title": "3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view Synthesis",
      "authors": [
        "Hongli Xiao",
        "Youjian Zhang",
        "Yaohui Jin",
        "Xiaoguang Ren",
        "Wenjing Yang",
        "Long Lan"
      ],
      "abstract": "High-quality 3D vehicle assets are essential for autonomous driving simulation. Although multi-view diffusion-based paradigms enable controllable single-image reconstruction, they typically produce limited viewpoints and exhibit cross-view geometric inconsistencies, thereby reducing reconstruction fidelity in real-world scenarios. In this work, we introduce 3DCarGen, a scalable single-view 3D car generation framework designed for real-world images by synthesizing an arbitrary number of 3D-consistent multi-view images. Specifically, given a single image as input, we first synthesize a set of images from fixed viewpoints. These images are then fed into a feed-forward reconstruction model, resulting in a coarse 3D representation based on 3D Gaussian Splatting. Conditioned on this explicit 3D prior, our multi-view diffusion model generates 3D-consistent images from arbitrary camera viewpoints. We further extend a fast mesh reconstruction algorithm by incorporating color-normal joint optimization to recover detailed and coherent 3D vehicle models from the synthesized dense views. Extensive experiments on synthetic and real-world datasets demonstrate that our approach achieves robust geometric consistency and reconstruction fidelity compared to existing methods. Project page: https://honglixiao.github.io/3dcargen.github.io/.",
      "short_abstract": "High-quality 3D vehicle assets are essential for autonomous driving simulation. Although multi-view diffusion-based paradigms enable controllable single-image reconstruction, they typically produce limited viewpoints and exhibit cross-view geometric...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24257",
      "pdf_url": "https://arxiv.org/pdf/2606.24257.pdf"
    },
    {
      "id": "2606.24546",
      "anchor": "paper-2606-24546",
      "title": "Explaining Failures of Cyber-Physical Systems with Actual Causality",
      "authors": [
        "Khen Elimelech",
        "Tom Yaacov",
        "David A. Kelly",
        "Hana Chockler",
        "Moshe Y. Vardi"
      ],
      "abstract": "Modern autonomous Cyber-Physical Systems (CPSs), such as self-driving cars, face increasingly complex demands, and yet are expected to act reliably. The black-box nature often characterizing such systems, especially those relying on neural components, makes it impossible to fully verify the system behavior prior to deployment. Unfortunately, unexpected failures-when the system does not comply with its specification-are inevitable and may have catastrophic implications. To improve trust in the system and facilitate future mitigation after a failure occurs, it is important to try to derive an explanation for the unexpected system behavior. This paper introduces the novel concept of leveraging the framework of actual causality for CPS failure explanation. Up until now, this framework was only used to derive explanations in the context of simple systems, such as image classifiers. This paper addresses the theoretical gaps and provides the guidance needed to allow for correct explanation derivation in the CPS domain. Beyond the theoretical contribution, the paper presents two novel, practical, system-agnostic explanation derivation algorithms, allowing to prioritize either explanation optimality or derivation efficiency. The approach is demonstrated and evaluated in the context of a neural-network-controlled autonomous car, designed to avoid collisions.",
      "short_abstract": "Modern autonomous Cyber-Physical Systems (CPSs), such as self-driving cars, face increasingly complex demands, and yet are expected to act reliably. The black-box nature often characterizing such systems, especially those relying on neural components, makes...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 3,
        "evidence": [
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24546",
      "pdf_url": "https://arxiv.org/pdf/2606.24546.pdf"
    },
    {
      "id": "2606.24784",
      "anchor": "paper-2606-24784",
      "title": "AerialFusionMapNet: Online HD Map Construction with Aerial-Onboard BEV Fusion",
      "authors": [
        "Daniel Lengerer",
        "Mathias Pechinger",
        "Klaus Bogenberger",
        "Carsten Markgraf"
      ],
      "abstract": "High-resolution aerial imagery has recently emerged as a complementary modality for automated driving perception and has shown potential to improve birds-eye-view (BEV) scene understanding when fused with onboard sensors. Prior work demonstrated performance gains for online high-definition (HD) map construction through aerial-onboard fusion; however, conventional end-to-end fusion does not fully exploit the structural information contained in aerial representations. In this work, we introduce AerialFusionMapNet, a fusion-based mapping framework with a structured two-stage training strategy that explicitly enhances the contribution of aerial features within a unified pipeline. The proposed training scheme enables more effective integration of structural aerial priors. On the nuScenes geographic split, AerialFusionMapNet achieves up to 54.7 mAP, improving over prior aerial-onboard fusion baselines from 48.8 mAP by +5.9 absolute and +12.1% relative. The results suggest that structured training design, rather than increased architectural complexity, plays a more decisive role in unlocking the full potential of aerial imagery for online HD map construction. Code and trained models are available at https://github.com/DriverlessMobility/AerialFusionMapNet.",
      "short_abstract": "High-resolution aerial imagery has recently emerged as a complementary modality for automated driving perception and has shown potential to improve birds-eye-view (BEV) scene understanding when fused with onboard sensors. Prior work demonstrated performance...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 22,
        "evidence": [
          "title: HD map",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24784",
      "pdf_url": "https://arxiv.org/pdf/2606.24784.pdf"
    },
    {
      "id": "2606.24201",
      "anchor": "paper-2606-24201",
      "title": "FORESEE: A Cooperative Lane Change Model for Connected and Automated Driving",
      "authors": [
        "Rafael Molina-Masegosa",
        "Sergei S. Avedisov",
        "Miguel Sepulcre",
        "Javier Gozalvez",
        "Yashar Z. Farid",
        "Onur Altintas"
      ],
      "abstract": "This paper presents FORESEE, a novel cooperative lane change model for connected and automated driving. FORESEE leverages Vehicle-to-Everything (V2X) data to anticipate traffic conditions and effectively organize lane changes. Specifically, it uses V2X data to organize vehicles into lanes based on their desired speeds, which helps to homogenize traffic flow and reduce disturbances caused by speed differences among vehicles within the same lane. The study demonstrates that implementing cooperative lane changes with FORESEE enhances average vehicle speed and energy efficiency compared to non-cooperative lane changes, which typically rely on short-term and local information about the ego vehicle and its immediate neighbors. This is achieved through fewer but more effective lane changes. Additionally, vehicles can maintain speeds closer to their desired speeds, resulting in fewer fluctuations in speed and acceleration and enhanced driving comfort. Moreover, cooperative lane changes can better manage road traffic disturbances, such as obstacles, by anticipating traffic conditions and organizing lane changes ahead. FORESEE serves as a valuable framework for the future design and testing of V2X-based maneuver coordinations as their effectiveness depends on how vehicles change lanes and their ability to plan and organize maneuvers in consideration of the upcoming traffic conditions.",
      "short_abstract": "This paper presents FORESEE, a novel cooperative lane change model for connected and automated driving. FORESEE leverages Vehicle-to-Everything (V2X) data to anticipate traffic conditions and effectively organize lane changes. Specifically, it uses V2X data...",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 15,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24201",
      "pdf_url": "https://arxiv.org/pdf/2606.24201.pdf"
    },
    {
      "id": "2606.24203",
      "anchor": "paper-2606-24203",
      "title": "Importance of Intent-Sharing for V2X-based Maneuver Coordination",
      "authors": [
        "Rafael Molina-Masegosa",
        "Sergei S. Avedisov",
        "Miguel Sepulcre",
        "Javier Gozalvez",
        "Yashar Z. Farid",
        "Onur Altintas"
      ],
      "abstract": "This paper examines the critical role of intent-sharing in enabling effective maneuver coordination for connected and automated vehicles (CAVs). Successful maneuver coordinations require vehicles to accurately know other vehicles' driving intentions. Intent-sharing can be achieved by the remote vehicles directly communicating their plans with the ego vehicle, as opposed to the ego vehicle predicting the trajectory on the remote vehicles' behalf. In this paper, we investigate the potential of intent-sharing on maneuver coordination effectiveness by quantifying the percentage of successful coordinations. We analyze the potential of intent-sharing by comparing its effectiveness for coordinated lane changes in a highway scenario with the effectiveness of a trajectory prediction method based on current kinematic data. Our analysis demonstrates in two scenarios substantial improvements in maneuver coordination when CAVs have direct access to the nearby vehicles' driving intentions through intent sharing. These findings highlight the importance of including intent-sharing in the maneuver coordination protocol.",
      "short_abstract": "This paper examines the critical role of intent-sharing in enabling effective maneuver coordination for connected and automated vehicles (CAVs). Successful maneuver coordinations require vehicles to accurately know other vehicles' driving intentions....",
      "published": "2026-06-23",
      "updated": "2026-06-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 15,
        "evidence": [
          "title: V2X",
          "abstract: connected vehicles"
        ]
      },
      "recency": "fresh",
      "age_days": 57,
      "arxiv_url": "https://arxiv.org/abs/2606.24203",
      "pdf_url": "https://arxiv.org/pdf/2606.24203.pdf"
    },
    {
      "id": "2606.28384",
      "anchor": "paper-2606-28384",
      "title": "A Query-Driven Communication-Efficient Digital Twins Design for Autonomous Driving",
      "authors": [
        "Nuocheng Yang",
        "Longyu Zhou",
        "Sihua Wang",
        "Changchuan Yin",
        "Tony Q. S. Quek"
      ],
      "abstract": "Digital twins (DTs) have become a potential technology to perform risk-free simulation of physical entities for deterministic and high-reliability services in diverse scenarios such as autonomous driving and low-altitude economy. In the autonomous driving scenario, traditional DT methods that rely solely on vehicle's real-time state synchronization, however, might lead to unacceptable computing and communication consumption for construction of high-fidelity DT with redundant data. To address this issue, we first propose a query-driven DT architecture to enable the DT to actively request the desired environment data from vehicles based on its simulation result. Then, we formulate an optimization problem whose goal is to minimize autonomous driving position error while accounting for DT fidelity and communication constraints. We also design a cross-time-step progressive query mechanism to further improve communication efficiency. The simulation results show that our proposed method achieves a 24% reduction in planning position error compared to traditional methods, while reducing communication overhead by 40%.",
      "short_abstract": "Digital twins (DTs) have become a potential technology to perform risk-free simulation of physical entities for deterministic and high-reliability services in diverse scenarios such as autonomous driving and low-altitude economy. In the autonomous driving...",
      "published": "2026-06-22",
      "updated": "2026-06-22",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 8,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 58,
      "arxiv_url": "https://arxiv.org/abs/2606.28384",
      "pdf_url": "https://arxiv.org/pdf/2606.28384.pdf"
    },
    {
      "id": "2606.24096",
      "anchor": "paper-2606-24096",
      "title": "Beyond Bayer: Task-Optimal Sensor Co-Design for Robust Autonomous-Driving Segmentation",
      "authors": [
        "Reeshad Khan",
        "John Gauch"
      ],
      "abstract": "Robust perception underpins autonomous driving, and most recent progress comes from scaling the model-larger backbones, foundation models, and cooperative multi-agent fusion. We pursue a complementary, upstream question: what should the camera itself measure? Using a differentiable RAW-to-task pipeline, we decompose which sensor degrees of freedom benefit dense prediction. Learning the spectral colour-filter-array (CFA) weights is the dominant lever, improving mIoU by +0.017 (KITTI-360) and +0.023 (ACDC) over a fixed camera. In contrast, point-spread-function (optics) co-design is net-negative (-0.020 mIoU on KITTI-360) - a consequence of the data-processing inequality, which also bounds the task information that any downstream model, however large or cooperative, can recover. Noise co-optimisation is marginal, and counter to intuition enlarging the CFA tile beyond 2x2 consistently hurts, as the filters are confined to the rank three sRGB input. Because the intervention is at the sensor, the gains are model-agnostic; we validate robustness on ACDC's fog, night, rain, and snow, and conclude with a simple recipe: learn the 2x2 CFA weights and keep an identity PSF.",
      "short_abstract": "Robust perception underpins autonomous driving, and most recent progress comes from scaling the model-larger backbones, foundation models, and cooperative multi-agent fusion. We pursue a complementary, upstream question: what should the camera itself measure?...",
      "published": "2026-06-22",
      "updated": "2026-06-22",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 3,
        "evidence": [
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 58,
      "arxiv_url": "https://arxiv.org/abs/2606.24096",
      "pdf_url": "https://arxiv.org/pdf/2606.24096.pdf"
    },
    {
      "id": "2606.23046",
      "anchor": "paper-2606-23046",
      "title": "UECP: Uncertainty-Enhanced Collaborative Perception",
      "authors": [
        "Kang Yang",
        "Tianci Bu",
        "Peng Wang",
        "Deying Li",
        "Wen Jie",
        "Yongcai Wang"
      ],
      "abstract": "Collaborative perception serves as a pivotal solution to enhance the perception capability of individual agents in autonomous driving, where a core challenge lies in seeking reliable evidence to quantify and weight the contribution of each participating agent. Existing methods typically rely on a confidence map, which is co-trained with the detection head, but it is inherently correlated with the detection results and thus fails to provide unbiased physical evidence. Furthermore, how to deeply integrate evidence into the cooperative fusion process remains an open question. To address these issues, this paper first proposes an uncertainty map, a physically grounded and unambiguous metric for evaluating perception quality. This map is directly supervised by real-time sensor signals, i.e., LiDAR point density, ensuring decoupling from detection noise and thereby providing physical scenario-aware evidence for weighting agent contribution. Based on this map, we develop the Uncertainty-Enhanced Collaborative Perception (UECP) framework, centered on the Uncertainty-Aware Pyramid Fusion (UAPF) module. UAPF uses a coarse-to-fine strategy, with two key components: Uncertainty-Weighted Downsampling (UWD) for high-fidelity feature preservation, and Uncertainty-Guided Residual Fusion (UGRF) to reinforce ego features, suppressing noise and ensuring robust fusion. Extensive experiments on real-world datasets show UECP outperforms state-of-the-art methods in effectiveness and robustness by embedding the uncertainty map into fusion. Code will be publicly available.",
      "short_abstract": "Collaborative perception serves as a pivotal solution to enhance the perception capability of individual agents in autonomous driving, where a core challenge lies in seeking reliable evidence to quantify and weight the contribution of each participating...",
      "published": "2026-06-22",
      "updated": "2026-06-22",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 15,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: title: perception",
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Systems, Deployment & Connectivity: title: cooperative systems",
          "Systems, Deployment & Connectivity: abstract: cooperative systems",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 58,
      "arxiv_url": "https://arxiv.org/abs/2606.23046",
      "pdf_url": "https://arxiv.org/pdf/2606.23046.pdf"
    },
    {
      "id": "2606.22971",
      "anchor": "paper-2606-22971",
      "title": "Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI",
      "authors": [
        "Xianda Guo",
        "Bohao Zhang",
        "Chenwei Huang",
        "Shiyuan Chen",
        "Ruilin Wang",
        "Yiqun Duan",
        "Cong Yang",
        "Qin Zou",
        "Wei Sui"
      ],
      "abstract": "Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases -- forward-facing cameras, far-field geometry, and static road priors -- limiting their applicability to embodied humanoid perception. We present Humanoid-OmniOcc, a large-scale panoramic stereo-based occupancy dataset tailored for humanoid robots. The dataset encompasses 15 diverse simulated indoor scenes and 5 real-world environments, yielding over 155K samples with broad scene and style diversity. Importantly, the dataset is designed around a Real2Sim2Real closed-loop paradigm: real sensor specifications drive physically accurate simulation, simulation produces large-scale annotated training data, and models trained in simulation are directly evaluated on real-world captures -- enabling iterative refinement of the sim-to-real pipeline. We further propose \\textbf{H}umanoid \\textbf{S}urround \\textbf{S}tereo-guided \\textbf{Occ}upancy model (Humanoid-OmniOcc) that exploits robust depth priors for accurate 2D-to-3D lifting. Extensive experiments show that Humanoid-OmniOcc consistently outperforms monocular baselines and generalizes well to both unseen simulated test scenes and real-world environments, validating the effectiveness of the Real2Sim2Real design. Code and data will be available upon acceptance at https://d-robotics-ai-lab.github.io/humanoid-omniocc.",
      "short_abstract": "Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases --...",
      "published": "2026-06-22",
      "updated": "2026-06-22",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 11,
        "evidence": [
          "abstract: perception",
          "title: occupancy",
          "abstract: occupancy",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 58,
      "arxiv_url": "https://arxiv.org/abs/2606.22971",
      "pdf_url": "https://arxiv.org/pdf/2606.22971.pdf"
    },
    {
      "id": "2606.22913",
      "anchor": "paper-2606-22913",
      "title": "Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving",
      "authors": [
        "Zisheng Chen",
        "Yuping Qiu",
        "Jianhua Han",
        "Tao Tang",
        "Xiuwei Chen",
        "Likui Zhang",
        "Ying-Cong Chen",
        "Hang Xu",
        "Xiaodan Liang"
      ],
      "abstract": "Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most existing approaches directly generate the final trajectory without explicitly examining its future consequences, which limits their reliability in complex and dynamic environments. To address this limitation, we propose IRR-Drive (Intend, Reflect, Refine), an adaptive multimodal reflection framework for autonomous driving. Specifically, to tightly couple high-level reasoning with physical constraints, IRR-Drive first generates a preliminary textual intention and anticipates potential interactions by predicting future semantic bird's-eye view (BEV) representations. This dual-modality (Text + BEV) reflection space explicitly models anticipated scene evolution, enabling the model to rigorously self-correct and refine its initial intent before generating the final trajectory. Furthermore, to balance planning performance and computational efficiency, we construct reflection-oriented training data and design an adaptive reflection reward, enabling the model to adaptively select its reasoning mode according to scene complexity. Instead of using reasoning primarily as an auxiliary interpretation, IRR-Drive directly integrates an adaptive reflection mechanism into the planning framework, enabling grounded, decision-aware trajectory correction that is driven by scene complexity. Our method achieves state-of-the-art performance on the NAVSIM benchmark in both PDMS and EPDMS. Extensive experiments demonstrate the effectiveness of our multimodal reflection framework and validate the efficacy of the proposed adaptive reflection strategy.",
      "short_abstract": "Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most existing approaches directly generate the final trajectory without explicitly...",
      "published": "2026-06-22",
      "updated": "2026-06-22",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 12,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 58,
      "arxiv_url": "https://arxiv.org/abs/2606.22913",
      "pdf_url": "https://arxiv.org/pdf/2606.22913.pdf"
    },
    {
      "id": "2606.22881",
      "anchor": "paper-2606-22881",
      "title": "A Vendor-Agnostic LiDAR Data Conversion System with Multi-Signal Detection and Multi-Format Output",
      "authors": [
        "Param Patel",
        "Jay Dave",
        "Pratyush Chakraborty"
      ],
      "abstract": "LiDAR (Light Detection and Ranging) sensors capture the surrounding environment as dense 3D point clouds by measuring the time-of-flight of emitted laser pulses, making them foundational across autonomous vehicles, robotics, and large-scale mapping. PCAP (Packet Capture) files from these sensors are the starting point of most 3D perception pipelines, yet internal packet structures, UDP (User Datagram Protocol) port conventions and encoding schemes differ enough across manufacturers that no single tool reads them all. Ouster, Velodyne, Hesai, and Livox each require their own SDK (Software Development Kit), their own environment setup, and their own conversion workflow. Supporting all four means maintaining four disconnected pipelines with no shared infrastructure. The pipeline described here takes a raw PCAP as input and handles vendor identification automatically, scoring six independent file characteristics through a weighted multi-signal approach to determine the source sensor. C++ SDKs handle Ouster and Velodyne, while Hesai and Livox rely on Python-based dpkt parsing where no open source SDK exists. From there, a single command writes output to any of five industry-standard formats. We tested on real outdoor captures. Ouster peaks at 2.08M points per second, Velodyne at 1.47M, both running through native C++ packet decoding. Hesai and Livox land at 110K and 150K respectively, where Python-layer parsing introduces overhead that compounds under sustained load. The 8-10x gap held consistently across runs. Tested on a consumer-grade i3 with 8GB RAM, no vendor configuration required",
      "short_abstract": "LiDAR (Light Detection and Ranging) sensors capture the surrounding environment as dense 3D point clouds by measuring the time-of-flight of emitted laser pulses, making them foundational across autonomous vehicles, robotics, and large-scale mapping. PCAP...",
      "published": "2026-06-22",
      "updated": "2026-06-22",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 14,
        "evidence": [
          "abstract: perception",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 58,
      "arxiv_url": "https://arxiv.org/abs/2606.22881",
      "pdf_url": "https://arxiv.org/pdf/2606.22881.pdf"
    },
    {
      "id": "2606.22813",
      "anchor": "paper-2606-22813",
      "title": "Active Inference as the Test-Time Scaling Law for Physical AI Agents",
      "authors": [
        "Omar Hashash",
        "Christo Kurisummoottil Thomas",
        "Walid Saad",
        "Merouane Debbah",
        "Karl Friston",
        "Adeel Razi"
      ],
      "abstract": "In this paper, a novel test-time scaling law for physical artificial intelligence (AI) agents is introduced. This scaling law enables physical AI agents to reason with their world models to generalize in unforeseen scenarios at test time. The derived scaling law is grounded in the first principle of active inference, which equips agents with the general objective to survive in the real world, under which their specific task objectives are subsumed. Active inference achieves this by providing the reasoning to resolve prediction errors that arise when the agent encounters unforeseen situations outside its training distribution, enabling generalization in non-stationary environments. The proposed scaling law captures this by dynamically updating the agent's policy with this reasoning at test time. This policy update is modeled as a soft Bayesian inference process in which beliefs about the policy are updated using the reasoning that reduces expected prediction errors under allowable policies as a likelihood. The resulting posterior policy admits a biological interpretation, recovering the scaling mechanism that engages the brain's basal ganglia and prefrontal cortex at test time. To solve this analytically intractable problem, a variational inference solution minimizing free energy bounds is developed. This solution extends to enable learning beyond training by reinforcing new instances, resolved at test time, in both the policy and world model. Unlike existing scaling laws constrained by model size and training data, the derived solution scales with the continuous real-world experience of a physical AI agent. Simulation results on an autonomous driving task demonstrate that the proposed solution outperforms model-free Q-learning and model-based Bayesian reinforcement learning, achieving robust generalization to unforeseen scenarios while improving inference efficiency by over 36%.",
      "short_abstract": "In this paper, a novel test-time scaling law for physical artificial intelligence (AI) agents is introduced. This scaling law enables physical AI agents to reason with their world models to generalize in unforeseen scenarios at test time. The derived scaling...",
      "published": "2026-06-21",
      "updated": "2026-06-21",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 2,
        "evidence": [
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 59,
      "arxiv_url": "https://arxiv.org/abs/2606.22813",
      "pdf_url": "https://arxiv.org/pdf/2606.22813.pdf"
    },
    {
      "id": "2607.00027",
      "anchor": "paper-2607-00027",
      "title": "Urban Deceleration Behavior Modes Under Scene Context: An Early-Kinematic Classifier from Argoverse 2 Multi-Agent Trajectories",
      "authors": [
        "Eni Solomon Laughter"
      ],
      "abstract": "Urban deceleration is one of the most empirically studied yet least taxonomically organized behaviors in car-following research. Recent perception-equipped autonomous-vehicle datasets enable trajectory-anchored mode discovery. We extract 1,219 sustained deceleration events from 234 urban driving logs of the Argoverse 2 Sensor dataset, encode each event in a 19-dimensional kinematic feature vector, discover behavioral modes via K-means clustering with bootstrap stability analysis, and quantify modulation by eleven scene-context variables. A HistGradientBoosting classifier predicts mode membership from the first 1.0 s of each event. Four stable modes emerge with a bootstrap Adjusted Rand Index of 0.897 across 50 resamples: anticipatory soft (62.8%), reactive closing (30.6%), brake-like jerk (4.8%), and an outlier category (1.8%). Only pair age shows a medium effect (epsilon^2 = 0.085); scene geometry and vulnerable-road-user proximity show negligible effects. The early-event classifier achieves macro-F1 = 0.758 at 1.0 s, with scene context contributing +0.059 F1 over kinematics alone. Modes are regime-invariant in medium-speed driving (ARI = 0.817) but regime-dependent at low speed (ARI = 0.166). A small set of stable kinematic modes structures urban deceleration; early-window jerk dominates predictive signal; and pair age is the primary contextual modulator.",
      "short_abstract": "Urban deceleration is one of the most empirically studied yet least taxonomically organized behaviors in car-following research. Recent perception-equipped autonomous-vehicle datasets enable trajectory-anchored mode discovery. We extract 1,219 sustained...",
      "published": "2026-06-21",
      "updated": "2026-06-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "fresh",
      "age_days": 59,
      "arxiv_url": "https://arxiv.org/abs/2607.00027",
      "pdf_url": "https://arxiv.org/pdf/2607.00027.pdf"
    },
    {
      "id": "2606.22617",
      "anchor": "paper-2606-22617",
      "title": "OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs",
      "authors": [
        "Hao Vo",
        "Phu Loc Nguyen",
        "Khoa Vo",
        "Sieu Tran",
        "Duc Minh Nguyen",
        "Ngo Xuan Cuong",
        "Nghi D. Q. Bui",
        "Anh Nguyen",
        "Duy Minh Ho Nguyen",
        "Ngan Le"
      ],
      "abstract": "Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Autonomous Vehicles (AV) remains an open challenge. Existing geometry-aware MLLMs typically rely on auxiliary 3D models at inference time, introducing pipeline complexity and the risk of cascading failures. In this paper, we present OmniSpace, a simple yet effective plug-and-play paradigm for geometry-aware spatial reasoning from purely 2D observations. Motivated by our finding that current MLLMs are bottlenecked by weak cross-view correspondence and depth estimation, OmniSpace introduces a Camera Pose Injector, a Multi-view Epipolar Attention module, and a 3D Geometric Distillation objective that jointly address these two limitations by transferring geometric knowledge into the model. Extensive experiments show that OmniSpace surpasses existing methods on planning benchmarks (nuScenes, Bench2Drive), risk detection (nuInstruct), language (Omnidrive), and generalization (DriveBench).",
      "short_abstract": "Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Autonomous Vehicles (AV) remains an open challenge. Existing geometry-aware MLLMs...",
      "published": "2026-06-21",
      "updated": "2026-06-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera",
          "Perception & Sensor Fusion: abstract: depth estimation",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 59,
      "arxiv_url": "https://arxiv.org/abs/2606.22617",
      "pdf_url": "https://arxiv.org/pdf/2606.22617.pdf"
    },
    {
      "id": "2606.21993",
      "anchor": "paper-2606-21993",
      "title": "From Driving Videos to Simulatable Scenarios",
      "authors": [
        "Alexandre Levy",
        "Ernest Valveny Llobet",
        "Antonio Manuel López"
      ],
      "abstract": "Autonomous vehicles (AVs) face driving scenarios ranging from routine traffic to rare events. To assess safety it is crucial to reproduce these scenarios in a controllable, repeatable, and scalable manner, with simulation playing a key role. This paper introduces D-V2S, a novel framework that automatically generates simulatable driving scenarios from driving videos. D-V2S operates in two stages: a Driving Record Analyzer (DRA) uses a vision language model (VLM) with our designed prompt to produce natural-language descriptions from input videos, capturing road layouts and dynamic traffic interactions; subsequently, a Scenario Generator (SG) uses a large language model (LLM) and our conditioning context to translate these descriptions into executable scenarios. Using simulations, we show that D-V2S generates scenarios where 90% of the relevant semantic elements of the videos are present. We also provide qualitative results demonstrating D-V2S's capability to transform real-world driving videos into simulatable scenarios. Moreover, we provide both semantic and human driven ablative analyses of D-V2S's modules. In particular, we show how the VLM choice matters for DRA, and how our SG achieves a 75% preference rate over other state-of-the-art methods.",
      "short_abstract": "Autonomous vehicles (AVs) face driving scenarios ranging from routine traffic to rare events. To assess safety it is crucial to reproduce these scenarios in a controllable, repeatable, and scalable manner, with simulation playing a key role. This paper...",
      "published": "2026-06-20",
      "updated": "2026-06-20",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 60,
      "arxiv_url": "https://arxiv.org/abs/2606.21993",
      "pdf_url": "https://arxiv.org/pdf/2606.21993.pdf"
    },
    {
      "id": "2606.22055",
      "anchor": "paper-2606-22055",
      "title": "Multi-Target Maneuver Coordinations: Unlocking Coordination Opportunities in Connected Automated Driving",
      "authors": [
        "Rafael Molina-Masegosa",
        "Sergei S. Avedisov",
        "Miguel Sepulcre",
        "Takayuki Shimizu",
        "Javier Gozalvez",
        "Onur Altintas"
      ],
      "abstract": "Maneuver coordination is a key enabler of connected and automated driving, allowing vehicles to negotiate and execute maneuvers that would otherwise be difficult, inefficient or unsafe. Existing approaches and use cases typically assume coordination with a single predefined target vehicle, which limits the number of coordination opportunities. This paper introduces a maneuver coordination approach based on multi-target selection, which allows a vehicle to identify and select among multiple potential coordination vehicles for a given maneuver. Multi-target maneuver coordination does not require modifications to the maneuver execution logic or to the underlying coordination protocol. Instead, it extends the decision-making process preceding coordination, enabling vehicles to exploit a broader set of feasible cooperative interactions. Results show that multi-target maneuver coordination significantly increases triggered and successfully executed coordinations while maintaining a low computational cost, as the proposed approach achieves these gains without requiring the analysis of a large number of potential target vehicles. These improvements preserve coordination success rates while enabling earlier maneuver initiation.",
      "short_abstract": "Maneuver coordination is a key enabler of connected and automated driving, allowing vehicles to negotiate and execute maneuvers that would otherwise be difficult, inefficient or unsafe. Existing approaches and use cases typically assume coordination with a...",
      "published": "2026-06-20",
      "updated": "2026-06-20",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Planning & Decision-Making: abstract: decision-making",
          "Systems, Deployment & Connectivity: abstract: cooperative systems"
        ]
      },
      "recency": "fresh",
      "age_days": 60,
      "arxiv_url": "https://arxiv.org/abs/2606.22055",
      "pdf_url": "https://arxiv.org/pdf/2606.22055.pdf"
    },
    {
      "id": "2606.22052",
      "anchor": "paper-2606-22052",
      "title": "When Cooperation Should End: Maneuver Coordination Cancellation for Connected Automated Driving",
      "authors": [
        "Rafael Molina-Masegosa",
        "Sergei S. Avedisov",
        "Miguel Sepulcre",
        "Javier Gozalvez",
        "Onur Altintas"
      ],
      "abstract": "Maneuver coordination is essential for cooperative connected automated driving, enabling vehicles to negotiate maneuvers and interactions through V2X communication. While prior work has largely focused on how to initiate and execute coordinations, considerably less attention has been given to how ongoing coordinations should be terminated when they become unsuitable. This paper introduces the first complete design and implementation of maneuver coordination cancellation, including a state machine, message set, and decision-making logic. Our evaluation shows that cancellation significantly reduces the time vehicles spend in coordinations that cannot succeed, allowing them to become available for new maneuvers sooner. This increases the number of triggered coordinations and improves the number of successful maneuver coordinations. Overall, the study demonstrates that maneuver coordination cancellation improves cooperative driving, and establishes a foundation for further refinements that can enhance the efficiency and robustness of connected automated driving.",
      "short_abstract": "Maneuver coordination is essential for cooperative connected automated driving, enabling vehicles to negotiate maneuvers and interactions through V2X communication. While prior work has largely focused on how to initiate and execute coordinations,...",
      "published": "2026-06-20",
      "updated": "2026-06-20",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 60,
      "arxiv_url": "https://arxiv.org/abs/2606.22052",
      "pdf_url": "https://arxiv.org/pdf/2606.22052.pdf"
    },
    {
      "id": "2606.21623",
      "anchor": "paper-2606-21623",
      "title": "A DVDrive Approach for doScenes Instructed Driving Challenge",
      "authors": [
        "Zijian Fu",
        "Xiangyang Chu",
        "Mengshi Qi",
        "Huadong Ma",
        "Guanghao Zhang",
        "Wei Li"
      ],
      "abstract": "Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and historical motion, but also from a natural-language maneuver instruction. This paper presents our submission to the doScenes Instructed Driving Challenge, built upon OmniDrive, a vision-language-action driving agent with 3D perception, reasoning, and planning capabilities. We adapt OmniDrive to the doScenes setting by training it on instruction-annotated nuScenes scenes and generating a 6-second ego trajectory represented by 12 future waypoints. To improve multi-view visual grounding, we further introduce a DVPE-style divided-view perception module into the OmniDrive perception head. Instead of attending globally to all camera features, the proposed module groups query features and image tokens into divided local view spaces and performs visibility-aware cross-attention within each view. This design reduces irrelevant cross-view interference and helps the model better align language instructions with local driving-relevant visual evidence. The code is publicly available at: https://github.com/feel12348/doscenes-omnidrive.",
      "short_abstract": "Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and historical motion, but also from a natural-language maneuver instruction. This...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21623",
      "pdf_url": "https://arxiv.org/pdf/2606.21623.pdf"
    },
    {
      "id": "2606.21587",
      "anchor": "paper-2606-21587",
      "title": "FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving",
      "authors": [
        "Bonan Wang",
        "Letian Tao",
        "Bin Shuai",
        "Jiaxin Gao",
        "Wenxin Zhao",
        "Wei Xiong",
        "Kehua Sheng",
        "Bo Zhang",
        "Yang Guan",
        "Shengbo Eben Li"
      ],
      "abstract": "Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effect, where the premature termination of a single environment necessitates a synchronized batch re-initialization, leading to suboptimal sample utilization and prohibitive re-initialization latency. To address this, we propose FAST, a synchronous parallel framework tailored for closed-loop simulation. Specifically, FAST employs Dynamic Parallel Sampling Alignment (DPSA) to maintain vectorization synchronization by extending terminated episodes via virtual continuation, thereby decoupling the sampling loop from individual terminations. By dynamically triggering global truncation based on the termination rate of parallel clips, FAST effectively eliminates the bottleneck of premature resets without sacrificing data diversity. Furthermore, to strictly preserve theoretical consistency, we incorporate a Scaled Mask-Padding Optimization (SMPO) that leverages validity masking and adaptive loss normalization to nullify the bias from auxiliary padding data. Empirical evaluations demonstrate that FAST achieves at least a 1.78 times wall-clock speedup over the single-clip baseline while preserving statistical unbiasedness.",
      "short_abstract": "Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effect, where the premature termination of...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21587",
      "pdf_url": "https://arxiv.org/pdf/2606.21587.pdf"
    },
    {
      "id": "2606.21509",
      "anchor": "paper-2606-21509",
      "title": "A Stitch in Time Saves Nine: Preserving Policy Compatibility Under Perception Updates in End-to-End Autonomous Driving",
      "authors": [
        "Yueyuan Li",
        "Yifei Xiao",
        "Mingyang Jiang",
        "Xiang Zuo",
        "Songan Zhang",
        "Ming Yang"
      ],
      "abstract": "End-to-end autonomous driving systems tightly couple perception and decision-making through latent representations. Consequently, updates to perception models can alter these representations and degrade the performance of downstream policies that remain fixed. Existing solutions typically rely on policy retraining or architectural decoupling, both of which incur substantial computation and validation costs. In this paper, we formulate the model stitching problem for end-to-end autonomous driving and test the hypothesis that policy compatibility can be preserved through lightweight latent-space alignment. We study low-complexity model stitching methods, including linear and convolutional stitchers, for restoring compatibility between updated perception modules and frozen downstream policy modules. Experiments demonstrate that stitching effectively preserves downstream driving behavior under diverse perception updates, including changes in random initialization, sensor configuration, and training domain. In the most challenging cross-domain setting from nuScenes to CARLA, convolutional stitching retains over 91\\% of the no-shift driving score while reducing adaptation time from \\SI{22.18}{h} to \\SI{0.91}{h}. These results suggest that model stitching provides an effective and computationally efficient alternative to retraining or fine-tuning for maintaining end-to-end autonomous driving systems. The model will be open-sourced upon paper acceptance at https://github.com/SCP-CN-001/model-stitching to support further research and development in autonomous driving.",
      "short_abstract": "End-to-end autonomous driving systems tightly couple perception and decision-making through latent representations. Consequently, updates to perception models can alter these representations and degrade the performance of downstream policies that remain...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 20,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21509",
      "pdf_url": "https://arxiv.org/pdf/2606.21509.pdf"
    },
    {
      "id": "2606.21486",
      "anchor": "paper-2606-21486",
      "title": "Reference-Free, Long-Horizon Trajectory Optimization for Aggressive Autonomous Driving in Milliseconds",
      "authors": [
        "Prayag Sharma",
        "Jonathan Y. M. Goh",
        "Franck Djeumou"
      ],
      "abstract": "Autonomous vehicles must generate long-horizon and dynamically feasible trajectories in real time-even when operating at the limits of vehicle handling-to ensure safe operation in adverse conditions. However, existing work rarely quantifies the computational demands of generating such trajectories without prior references, warm starts and often defaults to low-fidelity models, compromising accuracy and control authority. We investigate the modeling and solver design choices that enable real-time solution of long-horizon, reference-free optimal control problems (OCPs) using full vehicle dynamics. To this end, we analyze vehicle stiffness properties to justify the OCP's integration scheme and show that lower-order A-stable methods consistently outperform alternatives, with solve time differences reaching two orders of magnitude. We show that robust nonlinear solver performance hinges on understanding barrier parameter update strategies and safeguarding techniques for Hessian indefiniteness, inherent in some interior point methods. Lastly, we propose a computationally efficient method for generating initial guesses using dynamic equilibrium, unlocking real-time performance and reducing initial infeasibility by up to four orders of magnitude. Extensive benchmarking and high-fidelity BeamNG simulation demonstrate compute times as low as 55 ms over a 260 m horizon, including high-speed obstacle avoidance scenarios where drifting emerges as a necessary component of feasible trajectory generation.",
      "short_abstract": "Autonomous vehicles must generate long-horizon and dynamically feasible trajectories in real time-even when operating at the limits of vehicle handling-to ensure safe operation in adverse conditions. However, existing work rarely quantifies the computational...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 4,
        "evidence": [
          "abstract: vehicle dynamics",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21486",
      "pdf_url": "https://arxiv.org/pdf/2606.21486.pdf"
    },
    {
      "id": "2606.21270",
      "anchor": "paper-2606-21270",
      "title": "Non-line-of-sight imaging with arbitrary relay surface geometries via 3D Gaussian Transient Rendering",
      "authors": [
        "Yi Wang",
        "Ziyu Zhan",
        "Yuran Wang",
        "Hao Wang",
        "Qiang Liu",
        "Zuoqiang Shi",
        "Lingyun Qiu",
        "Xing Fu"
      ],
      "abstract": "Imaging objects hidden outside the direct line of sight expands the effective field of view and is critical for applications such as autonomous driving and robotic perception. Despite impressive progress in time-of-flight (ToF)-based non-line-of-sight (NLOS) imaging, real-world deployment remains challenging because practical measurements are often collected over spatially limited, arbitrarily shaped relay regions-conditions that violate the planar-wall and dense-sampling assumptions made by most existing methods. To address these limitations, we propose a LOS-guided NLOS imaging pipeline that imposes no geometric assumptions on the relay surface and naturally supports both confocal and non-confocal configurations. Our method represents the hidden scene using 3D Gaussian primitives and couples them with an efficient, differentiable transient rendering model, enabling end-to-end optimization directly from measured transients. We validate our approach on real-world measurements from both a public dataset and a custom-built capture system. Across settings, our method achieves state-of-the-art reconstruction fidelity under spatially limited, sparsely sampled conditions, and significantly outperforms existing methods on complex, arbitrary relay surface geometries.",
      "short_abstract": "Imaging objects hidden outside the direct line of sight expands the effective field of view and is critical for applications such as autonomous driving and robotic perception. Despite impressive progress in time-of-flight (ToF)-based non-line-of-sight (NLOS)...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end",
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21270",
      "pdf_url": "https://arxiv.org/pdf/2606.21270.pdf"
    },
    {
      "id": "2606.21172",
      "anchor": "paper-2606-21172",
      "title": "BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving",
      "authors": [
        "Zhe Shuai",
        "Xiaopeng Xie",
        "Yikun Zeng"
      ],
      "abstract": "Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representations for downstream action prediction. In perception-to-action pipelines, these representations can directly influence ego-vehicle waypoint planning, making the learned future dynamics a critical security-sensitive component. Despite their promise, the training-time security risks of autonomous-driving video world models remain largely unexplored. We present BadDreamer, a transferable spatio-temporal backdoor attack that targets the perception side of this pipeline. Unlike conventional backdoors that manipulate image labels, prompt outputs, or action supervision, BadDreamer poisons the learned transition dynamics of a video world model. It constructs trigger-erasure sequences in which an oncoming yellow delivery rider is visible in the observed context frames but erased from the future frames. After fine-tuning on a small fraction of such sequences, the compromised world model learns a hidden conditional association: when the physical trigger appears, it hallucinates a future where the rider disappears and the road appears clear. We further show that this corrupted future-aware representation can transfer to the downstream action module without directly modifying ego-trajectory labels, inducing unsafe non-evasive waypoint predictions. Our experiments instantiate this attack on a representative open-source perception-to-action pipeline, revealing a representation-level safety risk in autonomous-driving video world models and highlighting the need for backdoor-aware validation beyond clean generation quality.",
      "short_abstract": "Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representations for downstream action prediction. In perception-to-action pipelines, these representations can directly...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 10,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21172",
      "pdf_url": "https://arxiv.org/pdf/2606.21172.pdf"
    },
    {
      "id": "2606.21344",
      "anchor": "paper-2606-21344",
      "title": "Mind the Noise: Sensitivity of Transformer-based Interaction-Aware Trajectory Prediction Models to Noisy Data",
      "authors": [
        "Shahab Salehi",
        "Luca Lusvarghi",
        "Miguel Sepulcre",
        "Javier Gozalvez"
      ],
      "abstract": "Trajectory prediction allows autonomous vehicles to anticipate the future behavior of surrounding objects (or agents) and, accordingly, maximize the safety and efficiency of their driving. State-of-the-art Transformed-based interaction-aware trajectory prediction models, which rely on attention mechanisms to capture multi-agent interactions and maximize prediction accuracy, are commonly trained and evaluated on long-range high-quality datasets. These datasets are typically obtained by aggregating data from multiple vehicles or drones and removing any object detection or tracking noise offline. Yet, information about a surrounding object's state (its position, speed, heading) is far from being noiseless in real-world deployments. Object state estimation is affected by perception uncertainties and localization errors that can be particularly large for objects received via Vehicle-to-Everything (V2X) communications. In this paper, we analyze the impact of noisy object state information on the trajectory prediction accuracy of a state-of-the-art Transformer-based interaction-aware trajectory prediction model. Our study demonstrates that trajectory prediction accuracy can rapidly deteriorate as the noise intensity increases. Numerical results show that the prediction accuracy can reduce by a 1.3x factor under small noise levels and by as much as a 3.9x factor under the highest (yet realistic) noise conditions. These findings reveal the strong sensitivity of trajectory prediction models to noisy data, underscoring the need for more realistic training and evaluation datasets as well as noise mitigation strategies.",
      "short_abstract": "Trajectory prediction allows autonomous vehicles to anticipate the future behavior of surrounding objects (or agents) and, accordingly, maximize the safety and efficiency of their driving. State-of-the-art Transformed-based interaction-aware trajectory...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 21,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21344",
      "pdf_url": "https://arxiv.org/pdf/2606.21344.pdf"
    },
    {
      "id": "2606.21223",
      "anchor": "paper-2606-21223",
      "title": "Ultra-Fusion: A Resilient Tightly-Coupled Multi-Sensor Fusion SLAM Framework under Sensor Degradation and Spatiotemporal Perturbation for Intelligent Transportation Systems",
      "authors": [
        "Yihong Tian",
        "Junjie Zhang",
        "Liuyang Li",
        "Deteng Zhang",
        "Yunfei Zuo",
        "Jie Yin"
      ],
      "abstract": "Reliable localization is essential for intelligent transportation systems (ITS), including autonomous vehicles, quadruped last-mile carriers, and infrastructure-inspection unmanned aerial vehicles (UAVs). Although tightly-coupled multi-sensor fusion improves accuracy in favorable conditions, deployed systems remain vulnerable to sensor degradation -- poor illumination, LiDAR degeneracy, wheel slippage, and GNSS outage -- and to spatiotemporal calibration errors. These failures are common in urban canyons, tunnels, and high-speed corridors, where localization drift can degrade route tracking, tunnel passage continuity, and local map alignment. This paper presents Ultra-Fusion, a tightly-coupled multi-sensor localization framework based on a unified sliding-window estimator. Asynchronous measurements are timestamp-ordered and converted into optional factors within one optimization window, supporting WIO, VIO, LIO, and LVIO with optional wheel and GNSS augmentation. Observability-aware initialization selects the bootstrap mode, factor-wise reliability scheduling gates degraded measurements, and online LiDAR--IMU spatiotemporal calibration refines temporal offsets and rotational extrinsics during operation. We extend the M3DGR benchmark with simulation trajectories and evaluate more than 60 open-source SLAM systems on M3DGR, M2DGR-Plus, KAIST, GrandTour, and MARS-LVIG. The results show competitive accuracy across wheeled, legged, and aerial platforms under long-duration and high-speed operation, degradation, and calibration perturbation, improving localization availability for road-level autonomy, campus and warehouse mobility, and low-altitude aerial inspection. To benefit the industrial and academic community, we will release source code and datasets upon paper acceptance.",
      "short_abstract": "Reliable localization is essential for intelligent transportation systems (ITS), including autonomous vehicles, quadruped last-mile carriers, and infrastructure-inspection unmanned aerial vehicles (UAVs). Although tightly-coupled multi-sensor fusion improves...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 12,
        "evidence": [
          "abstract: localization",
          "title: SLAM",
          "abstract: SLAM",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21223",
      "pdf_url": "https://arxiv.org/pdf/2606.21223.pdf"
    },
    {
      "id": "2606.21154",
      "anchor": "paper-2606-21154",
      "title": "Virginia Tech Transportation Safety Index (VTTSI)",
      "authors": [
        "Jason Cusati",
        "Cheng-Shun Chuang"
      ],
      "abstract": "The Virginia Tech Transportation Safety Index (VTTSI) is a real-time, cloud-native framework for quantifying intersection safety using multimodal connected-vehicle telemetry and multi-year VDOT crash history. Traditional crash-based methods rely on lagged, aggregated data and cannot reflect rapidly changing operational conditions. VTTSI addresses this gap through a hybrid modeling approach that fuses Empirical Bayes (EB) crash stabilization, uplift factors derived from speed and conflict behavior, and a CRITIC-weighted multi-criteria decision-making (MCDM) module combining SAW, EDAS, and CODAS. The system produces interpretable, exposure-adjusted safety scores on a 0--100 scale every 15 minutes. A cloud-deployed architecture built on FastAPI, PostgreSQL, PostGIS, and Streamlit supports interactive visualization of traffic volumes, VRU exposure, speed variance, and real-time incident activity. Validation across intersections demonstrates coherent diurnal patterns, consistency among MCDM methods, and sensitivity to observable operational turbulence. Sensitivity analysis further shows that the RT--SI is robust to parameter perturbations, with deviations typically remaining below one point on the 0--100 scale. By integrating long-term crash risk with short-term behavioral dynamics, VTTSI provides a transparent, adaptive, and proactive safety-monitoring framework suitable for transportation agencies, traffic management centers, fleet operators, and autonomous vehicle systems.%~\\cite{persaud2007, montella2020systemic, schultz2025_surrogate, Amraji2025CombinedSafetyIndex}.",
      "short_abstract": "The Virginia Tech Transportation Safety Index (VTTSI) is a real-time, cloud-native framework for quantifying intersection safety using multimodal connected-vehicle telemetry and multi-year VDOT crash history. Traditional crash-based methods rely on lagged,...",
      "published": "2026-06-19",
      "updated": "2026-06-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 17,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 61,
      "arxiv_url": "https://arxiv.org/abs/2606.21154",
      "pdf_url": "https://arxiv.org/pdf/2606.21154.pdf"
    },
    {
      "id": "2608.14586",
      "anchor": "paper-2608-14586",
      "title": "Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures",
      "authors": [
        "Haibo HU",
        "Lianming Huang",
        "Qiao Li",
        "Nan Guan",
        "Chun Jason Xue"
      ],
      "abstract": "Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a full autonomous driving stack, this problem is even more pronounced: legacy vehicle platforms were provisioned for modular pipelines, yet after several planning-related functions are absorbed into a unified VLA model, part of the original CPU budget becomes underutilized, while the visual encoder and the main reasoning path still concentrate most computation and memory demand on the GPU. As a result, directly deploying VLA together with the rest of the onboard system can be hard under realistic GPU memory constraints. To address this issue, we present a hybrid CPU--GPU inference framework with flexible resource scheduling for autonomous driving. Our design partitions the VLA backbone at the block-layer granularity, executes the visual encoder and LLM prefix on the GPU, and offloads the LLM suffix to the CPU through a cross-frame asynchronous pipeline, thereby exposing a schedulable boundary for redistributing compute and memory pressure across heterogeneous processors. We evaluate the proposed framework on two representative driving VLA models, Orion and MindDrive. On Bench2Drive, our method reduces average latency from 521ms to 408.0ms for Orion and from 443ms to 306.2ms for MindDrive, corresponding to 21.7% and 30.9% reduction, respectively. For Orion, the estimated peak GPU memory is further reduced from 45GB to 29GB. In real-vehicle deployment under coexistence with Autoware.Universe, native Orion cannot run because the onboard GPU memory budget is insufficient, whereas the hybrid version runs successfully together with the full vehicle stack.",
      "short_abstract": "Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 15,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2608.14586",
      "pdf_url": "https://arxiv.org/pdf/2608.14586.pdf"
    },
    {
      "id": "2606.20980",
      "anchor": "paper-2606-20980",
      "title": "Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City",
      "authors": [
        "Adrian Cespedes",
        "Marcelo Chincha",
        "Dunant Cusipuma",
        "Victor Flores-Benites",
        "David Ortega",
        "Arturo Deza"
      ],
      "abstract": "As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new settings, in particular out-of-distribution (OOD) edge-case scenarios in new geographies? In this paper, we study this open question by providing a full factorial analysis with human drivers of Lima, human drivers from New York City, and VLMs and showing them dashcam footage collected from Lima and New York City -- prompting them with a variety of questions under a Visual Question Answering (VQA) paradigm. In particular, we pick these two cities as they are highly challenging driving locations where no Self-Driving Car company currently operates in, and ask questions that span 4 categories: Factual, Ratings, Counterfactual and Reasoning. We find that Humans and VLMs diverge in their responses -- though this is modulated by the type of questions asked, and that Humans answer similarly independent of where they are from (Lima/NYC). To our surprise, we did not find a strong difference in terms of answers (Humans or VLMs) that was modulated by geography, likely due to their high out-of-distribution nature. Our dataset is available at: https://huggingface.co/datasets/Artificio/robusto-2",
      "short_abstract": "As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new settings, in particular out-of-distribution (OOD) edge-case...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Human Factors & Policy: abstract: human driver"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20980",
      "pdf_url": "https://arxiv.org/pdf/2606.20980.pdf"
    },
    {
      "id": "2606.20772",
      "anchor": "paper-2606-20772",
      "title": "Mind the Privileged-to-Camera Gap: Actor-Centric Sidecar Supervision for Camera-First Open-Loop Waypoint Prediction",
      "authors": [
        "Feeza Khan Khanzada",
        "Jaerock Kwon"
      ],
      "abstract": "Camera-first autonomous-driving models predict future ego waypoints from images, ego-state features, and route commands, but waypoint supervision alone does not explicitly supervise actor-level representations of nearby road users. We study this as supervised representation learning for open-loop waypoint prediction. The deployable model uses multi-view RGB, ego state, and route command at inference. During training, simulator-derived sidecar labels supervise actor grounding, privileged hindsight actor relevance relative to the logged ego trajectory, and selected-actor short-horizon motion; these labels are never inference inputs. We evaluate route-disjoint splits with matched architecture, optimizer, validation criterion, checkpoint selection, and three seeds. A plain waypoint-only RGB baseline obtains 1.815$\\pm$0.02 m final displacement error (FDE), and the matched no-teacher non-sidecar RGB control obtains 1.716$\\pm$0.02 m. Road-user sidecar supervision (RU-sidecar) reduces FDE to 1.223$\\pm$0.01 m, a 32.6% reduction over the plain baseline and 28.7% over the matched no-teacher non-sidecar RGB control. It improves over the plain baseline on 1445/1494 routes and over the matched no-teacher non-sidecar RGB control on 1417/1494 routes. Actor-conditioned slices show gains in all nonempty subsets, including 29.1% reduction for samples with at least four valid sidecar actors and 30.0% when a vulnerable road user is present. Optional simulator-state teacher alignment reaches 1.186$\\pm$0.15 m FDE, but higher seed variability makes it secondary. Non-deployable simulator-state diagnostics remain stronger, indicating a privileged-to-camera gap. The evidence is limited to open-loop simulation diagnostics.",
      "short_abstract": "Camera-first autonomous-driving models predict future ego waypoints from images, ego-state features, and route commands, but waypoint supervision alone does not explicitly supervise actor-level representations of nearby road users. We study this as supervised...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: title: camera",
          "Perception & Sensor Fusion: abstract: camera",
          "Prediction & World Models: title: prediction",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20772",
      "pdf_url": "https://arxiv.org/pdf/2606.20772.pdf"
    },
    {
      "id": "2606.20764",
      "anchor": "paper-2606-20764",
      "title": "One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception",
      "authors": [
        "Keqin Zeng",
        "Shuting Su",
        "Shihao Lin",
        "Ziyue Li",
        "Rui Zhao"
      ],
      "abstract": "Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, real-world spatiotemporal data exhibits severe heterogeneity, often manifesting as extreme long-tail distributions for safety-critical scenarios. This data scarcity induces dataset shift that degrades detection performance and pose safety risks. While synthetic data generation offers a potential solution, existing generative approaches, such as diffusion models and Generative Adversarial Networks (GANs), often lack explicit spatial grounding and structural constraints, resulting in spatial and physical inconsistencies in generated scenes. To address these challenges, we introduce WMGen-v1, an agentic text-based world model framework for long-tail spatial data generation. WMGen-v1 employs a Large Vision-Language Model (LVLM) to construct a structured scene representation from a single reference image, while a Large Language Model (LLM) performs guidance-based scene expansion under physical plausibility and commonsense constraints. Subsequently, conditioned on the structured semantic representations produced by this reasoning process, a diffusion model generates diverse and physically grounded long-tail training data. Experiments on internal industrial datasets, ROADWork, and LaRS benchmarks demonstrate that WMGen-v1 outperforms baseline approaches. Notably, detectors trained solely on WMGen-v1 synthetic data approach real-only performance on aggregate dataset-level metrics, highlighting its potential to alleviate long-tail data scarcity for downstream spatial perception.",
      "short_abstract": "Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, real-world spatiotemporal data exhibits severe heterogeneity, often manifesting as extreme long-tail...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Synthetic Data",
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20764",
      "pdf_url": "https://arxiv.org/pdf/2606.20764.pdf"
    },
    {
      "id": "2606.20394",
      "anchor": "paper-2606-20394",
      "title": "Agentic AutoResearch forSpace Autonomy: An Auditable, LLM-Driven Research Agent for Aerospace Control Problems",
      "authors": [
        "Amit Jain",
        "Richard Linares"
      ],
      "abstract": "Spacecraft guidance, navigation, and control functions are increasingly realized as learned policies distilled from expert solvers. Developing such a policy is itself a research process: an investigator selects an architecture and hyperparameters, runs experiments, and must determine whether an apparent improvement is genuine or merely seed noise. This paper presents AutoResearch, a framework in which a large language model autonomously drives that loop for aerospace control problems, coupled with a credibility layer, built into the loop, that certifies each reported result against the problem's own measured seed noise. The language model serves only as the offline research agent that develops the control policy; the trained policy it produces is then deployed onboard the spacecraft, while the model itself never operates the vehicle. At each iteration the agent reads a plain-language problem description and the run history, proposes a single edit to the training script, executes it, and logs the outcome. No reported result is credited until it passes the same three checks: measured per-problem seed noise, reseeded verification of the best configuration, and leave-one-out pruning of the agent's edits. The same loop is applied, unchanged, to two aerospace control problems: a Clohessy-Wiltshire relative rendezvous and a safety-constrained collision-avoidance docking past a keep-out zone, each calibrated against a known optimal control benchmark. In both, the audited policy clears the measured seed noise by many standard deviations; an undirected search over the same parameters does not. On the docking problem the gap becomes categorical: undirected search yields no feasible policy, while the learned policy stays outside the keep-out zone on every seed.",
      "short_abstract": "Spacecraft guidance, navigation, and control functions are increasingly realized as learned policies distilled from expert solvers. Developing such a policy is itself a research process: an investigator selects an architecture and hyperparameters, runs...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: verification"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20394",
      "pdf_url": "https://arxiv.org/pdf/2606.20394.pdf"
    },
    {
      "id": "2606.20376",
      "anchor": "paper-2606-20376",
      "title": "CRAX: Fast Safe Reinforcement Learning Benchmarking",
      "authors": [
        "Tristan Tomilin",
        "Mourad Boustani",
        "Mickey Beurskens",
        "Thiago D. Simão"
      ],
      "abstract": "Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progress in RL, existing safety benchmarks with high-fidelity 3D physics remain computationally slow, limiting large-scale experimentation and rapid prototyping. To address this gap, we propose CRAX (Constrained RL Accelerated with JAX). Built on top of the MuJoCo XLA (MJX) physics engine with realistic 3D dynamics, CRAX leverages vectorized operations and hardware acceleration, yielding up to ~100x speedups over comparable CPU-based safety benchmarks. The benchmark features six environment suites and three agent-specific tasks, each spanning three difficulty levels. Evaluating six popular safe RL methods shows that no single approach dominates across all tasks, and reveals the trade-offs between performance and safety. We find that curriculum learning across difficulty levels and safety transfer can improve performance over direct training in harder settings.",
      "short_abstract": "Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progress in RL, existing safety benchmarks with high-fidelity 3D physics remain...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Systems, Deployment & Connectivity: abstract: hardware"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20376",
      "pdf_url": "https://arxiv.org/pdf/2606.20376.pdf"
    },
    {
      "id": "2606.20336",
      "anchor": "paper-2606-20336",
      "title": "Autonomous Driving with Priority-Ordered STL Specifications Under Multimodal Uncertainty",
      "authors": [
        "Taha Bouzid",
        "Shuhao Qi",
        "Mircea Lazar",
        "Sofie Haesaert"
      ],
      "abstract": "Autonomous vehicles must plan trajectories that satisfy multiple requirements, such as safety, traffic-rule compliance, and passenger comfort. However, in safety-critical scenarios, it is not always possible to satisfy all requirements simultaneously, necessitating their prioritization based on importance. At the same time, the uncertainty in the predicted trajectories of surrounding road users, such as other vehicles and pedestrians, must be explicitly accounted for. In this work, we propose an uncertainty-aware trajectory planning framework that incorporates a predefined priority ordering over Signal Temporal Logic (STL) specifications and preserves the induced lexicographic ordering under multimodal uncertainty. We implement this formulation with Model Predictive Path Integral (MPPI) control and demonstrate the effectiveness of our method on simulation scenarios, showing that our framework efficiently handles conflicting objectives under realistic multimodal uncertainty.",
      "short_abstract": "Autonomous vehicles must plan trajectories that satisfy multiple requirements, such as safety, traffic-rule compliance, and passenger comfort. However, in safety-critical scenarios, it is not always possible to satisfy all requirements simultaneously,...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "title: uncertainty",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20336",
      "pdf_url": "https://arxiv.org/pdf/2606.20336.pdf"
    },
    {
      "id": "2606.20274",
      "anchor": "paper-2606-20274",
      "title": "Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving",
      "authors": [
        "Shihao Ji",
        "HongXi Li",
        "Zihui Song",
        "Mingyu Li"
      ],
      "abstract": "Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories. Existing paradigms face a distinct dichotomy between representational efficiency and generalization capacity. Dense models (e.g., occupancy networks), while geometrically robust, incur critical computational bottlenecks and struggle with high-level semantic reasoning. Conversely, sparse, query-based planners are efficient but reliant on closed-set definitions, rendering them vulnerable to out-of-distribution (OOD) events. Although recent Vision-Language-Action (VLA) models offer open-vocabulary reasoning, their autoregressive, discrete token generation fundamentally conflicts with the continuous, high-frequency control requirements of vehicle dynamics. To address this, we propose Lagrange, an open-vocabulary, computationally sparse driving framework based on Masked Latent Fields (MLF). Rather than relying on dense volumetric reconstructions or closed-set query mechanisms, Lagrange exploits Vision-Language Models (VLMs) to encode class-agnostic object proposals into continuous semantic visual tokens. We introduce an intent-driven masked cross-attention module that temporally filters irrelevant entities, decoding the attended tokens into an implicit continuous energy field defined over spatial coordinates. By framing decision-making as a Lagrangian action minimization problem spanning this energy field, we enforce strict compliance with vehicle kinematics while executing collision avoidance. Extensive offline evaluations on both standard (nuScenes) and long-tail (CODA) benchmarks demonstrate that Lagrange establishes a promising framework for robust, interpretable, and kinematically feasible open-world autonomy.",
      "short_abstract": "Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories. Existing paradigms face a distinct dichotomy between...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 35,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20274",
      "pdf_url": "https://arxiv.org/pdf/2606.20274.pdf"
    },
    {
      "id": "2606.20189",
      "anchor": "paper-2606-20189",
      "title": "HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training",
      "authors": [
        "Maciej Wozniak",
        "Jesper Ericsson",
        "Hariprasath Govindarajan",
        "Truls Nyberg",
        "Thomas Gustafsson",
        "Patric Jensfelt",
        "Olov Andersson"
      ],
      "abstract": "Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD). However, current approaches typically treat VFMs as black-box teachers, relying exclusively on frame-wise feature similarity. Consequently, they do not fully exploit the teacher's layer-wise semantic structure and global context, as well as the rich spatiotemporal information inherent in LiDAR sequences. We propose HilDA, a self-supervised pretraining framework for LiDAR backbones that better captures the semantic what and geometric where needed for driving tasks. HilDA combines hierarchical distillation comprising multi-layer distillation for progressive semantic alignment and global context distillation for scene-level semantics, with a temporal occupancy diffusion objective promoting spatiotemporal consistency. Models pre-trained with HilDA achieve state-of-the-art results on cross-modal distillation benchmarks and outperform models trained via prior distillation approaches on 3D object detection, scene flow, and semantic occupancy prediction. Code available at: https://maxiuw.github.io/hilda.",
      "short_abstract": "Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD)....",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 18,
        "evidence": [
          "abstract: object detection",
          "abstract: occupancy",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20189",
      "pdf_url": "https://arxiv.org/pdf/2606.20189.pdf"
    },
    {
      "id": "2606.20110",
      "anchor": "paper-2606-20110",
      "title": "FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model",
      "authors": [
        "Yuhwan Jeong",
        "Hyeonseong Kim",
        "Daehyun We",
        "Seonkyu Song",
        "Jinnyeong Yang",
        "Hyun-Kurl Jang",
        "Youngho Yoon",
        "Kuk-Jin Yoon"
      ],
      "abstract": "Synthetic data for autonomous driving is surging, powered by diffusion models that promise scalable scene generation. Yet key obstacles remain, as enforcing multi-view and temporal consistency often relies on backbone fine-tuning or added layers, which erodes pre-trained knowledge and weakens text alignment. Models also stay close to the training distribution, struggling under adverse weather and unseen configurations, and fidelity favors frequent over rare classes. We address these gaps with FrozenDrive, a controllable generative framework that preserves a pretrained diffusion models knowledge while achieving strong consistency. FrozenDrive conditions on rich driving-stack signals and text prompts, and introduces knowledge-preserving spatio-temporal attention to impose cross-view alignment and temporal coherence in a single pass within a parameter-free frozen diffusion backbone. An additional object-focused constraint improves per-object fidelity for rare categories. Without any weather- or scene-specific fine-tuning, our model synthesizes globally coherent multi-view driving scenes from text, particularly under adverse and rare conditions, and surpasses prior baselines. On nuScenes, FrozenDrive augmented data significantly improves AD models performance, especially at night and in rain, demonstrating stronger robustness when trained with our scenario-targeted data.",
      "short_abstract": "Synthetic data for autonomous driving is surging, powered by diffusion models that promise scalable scene generation. Yet key obstacles remain, as enforcing multi-view and temporal consistency often relies on backbone fine-tuning or added layers, which erodes...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20110",
      "pdf_url": "https://arxiv.org/pdf/2606.20110.pdf"
    },
    {
      "id": "2606.19836",
      "anchor": "paper-2606-19836",
      "title": "World Engine: Towards the Era of Post-Training for Autonomous Driving",
      "authors": [
        "Tianyu Li",
        "Li Chen",
        "Caojun Wang",
        "Haochen Liu",
        "Kashyap Chitta",
        "Zhenjie Yang",
        "Yuhang Lu",
        "Naisheng Ye",
        "Yihang Qiu",
        "Yufei Wang",
        "Luoxi Zou",
        "Jiaxin Peng",
        "Jin Pan",
        "Zhaoyu Su",
        "Andrei Bursuc",
        "Shengbo Eben Li",
        "Andreas Geiger",
        "Peng Su",
        "Hongyang Li"
      ],
      "abstract": "Autonomous vehicles must operate safely in the real world, where errors can have severe consequences. Although modern end-to-end driving policies excel in routine scenarios, their reliability is limited by the scarcity of safety-critical ``long-tail'' events in real driving datasets. These rare interactions define the practical safety boundary of the learned policy, yet they are difficult to collect at scale in the real world. Here we show that this fundamental limitation can be addressed by post-training pre-trained driving models on synthesized high-stakes interactions. We introduce World Engine, a generative framework that reconstructs high-fidelity interactive environments from real-world logs and systematically extrapolates them into realistic safety-critical variations. This paradigm enables reinforcement-based post-training to align policies with safety constraints, circumventing the physical risks inherent in real-world exploration. On a public benchmark built on nuPlan, World Engine substantially reduces failures in rare safety-critical scenarios and yields significantly larger gains than scaling pre-training data alone. Furthermore, when deployed on a production-scale autonomous driving system, the resulting policy reduces simulated collisions and demonstrates measurable improvements in on-road testing, showing that post-training on synthesized, safety-critical interactions offers a scalable and effective pathway to safer autonomous driving. The full codebase suite, including training, is released to the public.",
      "short_abstract": "Autonomous vehicles must operate safely in the real world, where errors can have severe consequences. Although modern end-to-end driving policies excel in routine scenarios, their reliability is limited by the scarcity of safety-critical ``long-tail'' events...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 0,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end driving",
          "End-to-End & VLA: abstract: end-to-end",
          "Safety, Security & Verification: abstract: safety",
          "Safety, Security & Verification: abstract: validation or testing",
          "Safety, Security & Verification: abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.19836",
      "pdf_url": "https://arxiv.org/pdf/2606.19836.pdf"
    },
    {
      "id": "2606.20120",
      "anchor": "paper-2606-20120",
      "title": "Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform",
      "authors": [
        "Hyeonna Choi",
        "Jung Yup Kim",
        "Hyuneui Lim",
        "Seunggyu Jeon"
      ],
      "abstract": "Biological experiment protocols are written in natural language, whereas automation systems rely on predefined control commands, creating a semantic gap that limits autonomous execution. Microplate-based automatic experiments are particularly challenging due to the need to simultaneously control well mapping, sample-reagent combinations, replicate placement, and parallel dispensing. This study proposes an agent-based protocol translation framework that converts natural-language microplate-based protocols into executable control commands for a robotic laboratory platform. A Parser Agent formalizes the natural-language protocol into a structured representation, and a rule-based mapping engine deterministically incorporates the operational constraints of the robotic laboratory platform to generate device-level control commands. A heterogeneous LLM Validation Agent verifies completeness, parameter accuracy, and execution order, and triggers a self-correction loop with structured feedback when errors are detected. A sweep involving 7 Parsers and 3 Validators on randomly selected ELISA protocols evaluates how model scale and Validator type affect translation accuracy and pass rates under cross-model verification. The accuracy-latency trade-off is further verified by comparing the rule-based mapping of the proposed framework with LLM end-to-end direct mapping. Finally, Bradford assay-based protein quantification using a microplate was demonstrated on a robotic laboratory platform, validating end-to-end autonomous execution from natural-language protocols to real-world experiments. The proposed framework provides a flexible approach to narrowing the semantic gap between natural-language protocols and microplate-based self-driving laboratories.",
      "short_abstract": "Biological experiment protocols are written in natural language, whereas automation systems rely on predefined control commands, creating a semantic gap that limits autonomous execution. Microplate-based automatic experiments are particularly challenging due...",
      "published": "2026-06-18",
      "updated": "2026-06-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "abstract: verification",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 62,
      "arxiv_url": "https://arxiv.org/abs/2606.20120",
      "pdf_url": "https://arxiv.org/pdf/2606.20120.pdf"
    },
    {
      "id": "2606.19641",
      "anchor": "paper-2606-19641",
      "title": "Scaling Self-Play for End-to-End Driving",
      "authors": [
        "Luke Rowe",
        "Roger Girgis",
        "Rodrigue de Schaetzen",
        "Daphne Cornelisse",
        "Alaap Grandhi",
        "Felix Heide",
        "Eugene Vinitsky",
        "Christopher Pal",
        "Liam Paull"
      ],
      "abstract": "End-to-end autonomous driving models are typically trained on offline human-demonstration datasets that provide limited state coverage and often no closed-loop feedback, making them prone to compounding errors when deployed in closed-loop and brittle to long-tail agent interactions. To overcome these limitations, we propose an alternative strategy for training end-to-end driving models: large-scale self-play directly from pixels in simulation. While prior self-play approaches have shown promising transfer to real-world driving, they typically assume vectorized Bird's-Eye-View (BEV) observations that are incompatible with end-to-end policies operating directly on sensor observations. To this end, we introduce Gigapixel, a high-throughput batched driving simulator with perspective rendering, enabling scalable self-play directly from pixel observations. Rather than targeting compute-costly photorealistic sensor simulation, Gigapixel renders a simplified bounding-box world that preserves essential scene structure while achieving throughput at 50k agent steps per second. Since direct pixel-space self-play RL is prohibitively sample-inefficient at end-to-end model scale, we propose self-play DAgger training: we train pixel-based policies in self-play via on-policy distillation from a privileged RL teacher. To bridge the sim-to-real gap, we subsequently transfer the self-play trained policies to real-world sensor data through lightweight perception adaptation. Policies trained in Gigapixel and adapted to real-world sensor data achieve competitive performance on the HUGSIM and NAVSIM-v2 benchmarks without human trajectory supervision. Moreover, scaling self-play training yields proportional gains in policy performance, establishing self-play as a practical and scalable strategy for training end-to-end models.",
      "short_abstract": "End-to-end autonomous driving models are typically trained on offline human-demonstration datasets that provide limited state coverage and often no closed-loop feedback, making them prone to compounding errors when deployed in closed-loop and brittle to...",
      "published": "2026-06-17",
      "updated": "2026-06-17",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 31,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 63,
      "arxiv_url": "https://arxiv.org/abs/2606.19641",
      "pdf_url": "https://arxiv.org/pdf/2606.19641.pdf"
    },
    {
      "id": "2606.19122",
      "anchor": "paper-2606-19122",
      "title": "Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning",
      "authors": [
        "Yukai Ma",
        "Joe Lin",
        "Liu Liu",
        "Honglin He",
        "Lulu Ricketts",
        "Brad Squicciarini",
        "Yong Liu",
        "Bolei Zhou"
      ],
      "abstract": "Sidewalks in the real world are crowded, cluttered, and less structured than roads, making 3D occupancy prediction a key ingredient for the safe navigation of mobile robots such as delivery bots and electric wheelchairs. Existing occupancy learning pipelines are largely designed for on-road autonomous driving and often train on large-scale paired LiDAR-RGB datasets with dense 3D supervision and multiple camera inputs, which are costly to collect and do not adequately capture sidewalk-specific characteristics. We propose WalkOCC, a hybrid Ray-marching monocular 3D occupancy perception framework for robots operating on sidewalks. WalkOCC explicitly couples geometric grounding from LiDAR-RGB paired data with scalable learning from large-scale unpaired monocular images. It bootstraps pseudo occupancy supervision from paired sequences and jointly learns image-level representations on additional 2D-only data. It yields stable optimization and improved generalization without requiring costly 3D occupancy annotations. Extensive experiments demonstrate consistent gains in prediction accuracy, fine-grained segmentation of subtle urban structures such as curbs and gutters, and robustness to environmental and cross-embodiment shifts compared with self-supervised image-based baselines. To facilitate evaluation and benchmarking, we also introduce Sidewalk3D, a large-scale sidewalk perception dataset with LiDAR-camera paired sequences collected across multiple locations and time periods, along with 3D semantic occupancy annotations for evaluation. Code and data will be made available.",
      "short_abstract": "Sidewalks in the real world are crowded, cluttered, and less structured than roads, making 3D occupancy prediction a key ingredient for the safe navigation of mobile robots such as delivery bots and electric wheelchairs. Existing occupancy learning pipelines...",
      "published": "2026-06-17",
      "updated": "2026-06-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 23,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "title: occupancy",
          "abstract: occupancy",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 63,
      "arxiv_url": "https://arxiv.org/abs/2606.19122",
      "pdf_url": "https://arxiv.org/pdf/2606.19122.pdf"
    },
    {
      "id": "2606.19632",
      "anchor": "paper-2606-19632",
      "title": "Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation",
      "authors": [
        "Ahmad Farooq",
        "Kamran Iqbal"
      ],
      "abstract": "Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets. We present the first end-to-end framework for safety verification of learned multi-agent communication policies through policy abstraction: neural policies are distilled into interpretable decision trees, then formally verified, with empirical validation confirming that verified safety properties transfer to original networks. Our four-stage pipeline consists of domain-specific feature extraction from agent observations, decision tree distillation achieving 97.9% +/- 1.2% fidelity to neural policies, automated translation to PRISM probabilistic model checker specifications with complete feature-to-state-variable correspondence, and compositional verification of Probabilistic Computation Tree Logic (PCTL) properties via pairwise decomposition with union-bound aggregation and empirical neighbor modeling. Evaluating Vector-Quantized Variational Information Bottleneck (VQ-VIB) policies for multi-drone coordination with 5-7 agents, we verify 18 temporal logic properties across safety, liveness, and cooperation, achieving 88.9% property satisfaction with all five safety thresholds satisfied (0.3% collision probability vs. 1% threshold). Monte Carlo validation of original neural policies confirms that verified safety properties transfer with <=0.6 percentage-point deviation (95% CI). Discrete VQ-VIB messages provide +11.6 to +13.6 percentage-point fidelity advantages over continuous methods, enabling 3-4x faster verification. Our framework provides empirically validated safety verification for distilled policy abstractions, serving as a practical bridge between deep MARL and formal safety workflows for multi-robot deployment.",
      "short_abstract": "Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous...",
      "published": "2026-06-17",
      "updated": "2026-06-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 16,
        "evidence": [
          "abstract: safety",
          "title: verification",
          "abstract: verification",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 63,
      "arxiv_url": "https://arxiv.org/abs/2606.19632",
      "pdf_url": "https://arxiv.org/pdf/2606.19632.pdf"
    },
    {
      "id": "2606.19267",
      "anchor": "paper-2606-19267",
      "title": "A Mixed-Reality Testbed for Autonomous Vehicles",
      "authors": [
        "H. M. Sabbir Ahmad",
        "Ehsan Sabouni",
        "Emrullah Celik",
        "Zean Wan",
        "Damola Ajeyemi",
        "Christos G. Cassandras",
        "Wenchao Li"
      ],
      "abstract": "We propose a mixed-reality, hardware-in-the-loop (HIL) testbed for autonomous vehicles that seamlessly integrates a physical testbed of mobile robots with a high-fidelity simulation environment. The virtual simulation enables the creation of diverse, safety-critical driving scenarios to validate state-of-the-art perception, planning, and control algorithms, while augmenting simulations with physical robots equipped with multimodal sensors in photorealistic virtual environments further facilitating rigorous validation. Our testbed also features vehicular connectivity using wireless communication and can accommodate a large number of agents through the combination of physical robots and virtual simulated agents, supporting research on multi-agent systems including Connected and Autonomous Vehicles (CAVs). Finally, we present a safety-guaranteed framework combining perception, planning and a novel online learning-based controller using Control Barrier Functions (CBFs) for CAVs. Experiments using the proposed framework are used to validate and demonstrate the key functionalities and the overall utility of the testbed to bridge the gap between simulation and real-world hardware deployment.",
      "short_abstract": "We propose a mixed-reality, hardware-in-the-loop (HIL) testbed for autonomous vehicles that seamlessly integrates a physical testbed of mobile robots with a high-fidelity simulation environment. The virtual simulation enables the creation of diverse,...",
      "published": "2026-06-17",
      "updated": "2026-06-17",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 5,
        "evidence": [
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: hardware",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 63,
      "arxiv_url": "https://arxiv.org/abs/2606.19267",
      "pdf_url": "https://arxiv.org/pdf/2606.19267.pdf"
    },
    {
      "id": "2606.18537",
      "anchor": "paper-2606-18537",
      "title": "Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents",
      "authors": [
        "Caleb Chang",
        "Davin Win Kyi",
        "Natasha Jaques",
        "Karen Leung"
      ],
      "abstract": "Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act in an environment. However, observations drawn from a heterogeneous population introduce conflicting behavioral signals, making it difficult to determine which behaviors are worth imitating. We address this challenge with General Reward Inference and Disentanglement (GRID), a social learning method that extracts universally useful behaviors from a heterogeneous population of demonstrators pursuing different goals. GRID decomposes per-agent reward functions into a general reward, capturing behaviors shared across all agents, and specific rewards, capturing individual preferences and objectives. Training exclusively on the general reward provides a new paradigm of generalist pretraining. It yields a generalist agent that internalizes universal environmental competencies, such as safety and basic task proficiency, without the mode-averaging bias that afflicts standard learning from demonstration techniques. This generalist serves as a superior prior for fine-tuning to downstream tasks, including preferences unseen during training. Experiments across a synthetic basis function decomposition, multi-agent Craftax, and a continuous autonomous driving simulator (Highway-Env) confirm that GRID successfully disentangles reward structure in a semantically meaningful way, outperforms standard learning from demonstration baselines, and enables more efficient and stable specialization.",
      "short_abstract": "Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act in an environment. However, observations drawn from a heterogeneous population introduce conflicting behavioral signals, making it difficult to...",
      "published": "2026-06-16",
      "updated": "2026-06-16",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 64,
      "arxiv_url": "https://arxiv.org/abs/2606.18537",
      "pdf_url": "https://arxiv.org/pdf/2606.18537.pdf"
    },
    {
      "id": "2606.18242",
      "anchor": "paper-2606-18242",
      "title": "EventDrive: Event Cameras for Vision-Language Driving Intelligence",
      "authors": [
        "Dongyue Lu",
        "Rong Li",
        "Ao Liang",
        "Lingdong Kong",
        "Wei Yin",
        "Lai Xing Ng",
        "Benoit R. Cottereau",
        "Camille Simon Chane",
        "Wei Tsang Ooi"
      ],
      "abstract": "Event cameras sense the world through asynchronous brightness changes with microsecond latency and high dynamic range, offering motion fidelity far beyond frame-based sensors and capturing temporal structure that conventional exposures often miss. These properties make events a powerful complement to RGB in autonomous driving, especially under blur, glare, and rapid motion, where frame-based perception can become unreliable. However, existing event-aware vision-language models remain limited to generic perception and do not reveal how event sensing contributes to reasoning and decision-making across the full driving loop. We present EventDrive, a large-scale benchmark and model suite that unifies event streams, RGB frames, and language supervision across four core dimensions: Perception, Understanding, Prediction, and Planning, covering captions, structured QA, grounding, motion-state recognition, trajectory forecasting, and planning tasks. Building on this foundation, EventDrive-VLM introduces a multi-horizon event pyramid and a temporal-horizon mixture-of-experts module to adaptively encode and fuse asynchronous and frame-based information for downstream reasoning. Comprehensive evaluation across diverse tasks shows that event streams provide substantial gains in temporal precision, motion awareness, and robustness, bringing event sensing into the center of driving intelligence.",
      "short_abstract": "Event cameras sense the world through asynchronous brightness changes with microsecond latency and high dynamic range, offering motion fidelity far beyond frame-based sensors and capturing temporal structure that conventional exposures often miss. These...",
      "published": "2026-06-16",
      "updated": "2026-06-16",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 4,
        "evidence": [
          "abstract: perception",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 64,
      "arxiv_url": "https://arxiv.org/abs/2606.18242",
      "pdf_url": "https://arxiv.org/pdf/2606.18242.pdf"
    },
    {
      "id": "2606.18112",
      "anchor": "paper-2606-18112",
      "title": "Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System",
      "authors": [
        "Jiazhao Zhang",
        "Gengze Zhou",
        "Hale Yin",
        "Yiyang Huang",
        "Zixing Lei",
        "Qihang Peng",
        "Haoqi Yuan",
        "Jie Zhang",
        "Xudong Guo",
        "Xiaoyue Chen",
        "An Yang",
        "Fei Huang",
        "Zhibo Yang",
        "Junyang Lin",
        "Dayiheng Liu",
        "Jingren Zhou",
        "Zhuoyuan Yu",
        "Jingyang Fan",
        "Zhixuan Liang",
        "Pei Lin",
        "Ye Wang",
        "Haoyang Li",
        "Anzhe Chen",
        "Kun Yan",
        "Xiao Xu"
      ],
      "abstract": "Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, because instruction following, object search, target tracking, and autonomous driving share the same perception-planning backbone yet demand fundamentally different strategies for consuming the visual stream. We present Qwen-RobotNav, a scalable navigation model built on Qwen-RobotNav that addresses it through a parameterised interface with two complementary dimensions: multiple task modes that select the navigation behaviour, and controllable observation parameters (e.g., token budget, per-camera weights) that govern how visual history is encoded. With training-time randomization over all parameters, Qwen-RobotNav is robust to any inference-time configuration requiring zero architectural modification to the Qwen-RobotNav backbone. We train Qwen-RobotNav on 15.6M samples; co-training with vision-language data prevents the collapse into reactive action-sequence mappers observed in trajectory-only training. The parameterised interface also makes Qwen-RobotNav a natural building block for agentic systems: for long-horizon scenarios, an upper-level planner decomposes goals into sub-tasks and dynamically switches Qwen-RobotNav's task mode and context strategy mid-episode, composing complex behaviours from repeated calls to the same model. Extensive experiments show that Qwen-RobotNav sets new state-of-the-art results across major navigation benchmarks. The model exhibits favourable scaling from 2B to 8B parameters, with joint multi-task training developing a shared spatial-planning substrate that transfers across task families, and demonstrates strong zero-shot generalisation to real-world robots across diverse environments.",
      "short_abstract": "Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, because instruction following, object search, target tracking, and autonomous driving share the same perception-planning...",
      "published": "2026-06-16",
      "updated": "2026-06-16",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 6,
        "evidence": [
          "abstract: planning",
          "title: navigation",
          "abstract: navigation"
        ]
      },
      "recency": "fresh",
      "age_days": 64,
      "arxiv_url": "https://arxiv.org/abs/2606.18112",
      "pdf_url": "https://arxiv.org/pdf/2606.18112.pdf"
    },
    {
      "id": "2606.17897",
      "anchor": "paper-2606-17897",
      "title": "Learn to Quantify Social Interaction with Constraints for Pedestrian Walking",
      "authors": [
        "Xiaodan Shi"
      ],
      "abstract": "Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous driving cars and social robots) to avoid collision and make high-quality planning. Although the current research take into account social interactions for prediction, they don't reveal the exact kinds of social interactions happened among people and how the social interactions affect the decision-making process of pedestrians, which further limits its robustness. Social interactions in pedestrian walking are intuitively massive and hard to label and quantify. In this paper, we explore creatively to quantify and interpret how pedestrians interact with others by proposing Learn to Cluster. Our clustering social interactions is probabilistic latent variable generative, learning directly from sequential trajectory observations, scalable to arbitrary number of pedestrians. Learn to cluster is label-free and can be naturally integrated into the training process of the prediction model. The latent variables will then serve as 'labels' to categorize social interactions. Extensive experiments over several trajectory prediction benchmarks demonstrate that our method is able to learn the patterns of social interactions and effectively integrate the patterns to pedestrian trajectory prediction.",
      "short_abstract": "Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous driving cars and social robots) to avoid collision and make high-quality planning. Although the current research take into account social interactions for...",
      "published": "2026-06-16",
      "updated": "2026-06-16",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 3,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: forecasting",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 64,
      "arxiv_url": "https://arxiv.org/abs/2606.17897",
      "pdf_url": "https://arxiv.org/pdf/2606.17897.pdf"
    },
    {
      "id": "2606.17536",
      "anchor": "paper-2606-17536",
      "title": "OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation",
      "authors": [
        "Zijie Meng",
        "Yufei Liu",
        "Chengqian Ma",
        "Zhiyu Li",
        "Jiyuan Liu",
        "Wenhua Nie",
        "Bingcai Wei",
        "Shuqin Chen",
        "Weichen Xu",
        "Jiquan Yuan",
        "Miao Zhang"
      ],
      "abstract": "Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry. We trace both to a single root cause: the absence of a shared symbolic interlingua aligning language, geometry, and pixels at the latent-token level. We present DRIVE-CHOREO, an LLM-choreographed multi-agent world model that recasts controllable multi-view video generation as latent choreography. Three Qwen2.5-VL agents - a Director parsing user intent into a structured WorldScript, a Cartographer grounding it into spatially-anchored layout tokens, and an Auditor feeding cross-view critiques back as auxiliary supervision - jointly author a single position-aware token sequence. This sequence is co-compressed with the multi-view video via a view-time permutation that enforces inter-camera geometry within the convolutional receptive field of a 3-D VAE. On nuScenes, DRIVE-CHOREO sets new state-of-the-art multi-view consistency and BEV mAP (21.6) with competitive FVD (45.7); a detector trained purely on our synthetic data gains +2.4 NDS on the real validation split, validating downstream utility.",
      "short_abstract": "Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion,...",
      "published": "2026-06-16",
      "updated": "2026-06-16",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Synthetic Data",
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 64,
      "arxiv_url": "https://arxiv.org/abs/2606.17536",
      "pdf_url": "https://arxiv.org/pdf/2606.17536.pdf"
    },
    {
      "id": "2606.20725",
      "anchor": "paper-2606-20725",
      "title": "D2HDMap: Non-visible Driveline Map Prior for Online Vectorized HD Map Prediction",
      "authors": [
        "Seojun Shon",
        "Chikao Tsuchiya",
        "Dhaval Bhanderi",
        "David Ilstrup",
        "Hsinmin Cheng",
        "Christopher Ostafew"
      ],
      "abstract": "Accurate, up-to-date representations of road structures are critical for the safe operation of autonomous vehicles. Existing systems rely either on costly, maintenance-heavy high-definition (HD) maps which compromise safety when outdated, or purely sensor-based online mapping which struggles with long-range reliability and occlusion. Systems incorporating map prior information into online mapping seek to overcome drawbacks of both approaches by combining them in some way. We propose 'Driveline To HD Map' (D2HDMap), an online mapping system that injects a lightweight, non-visible driveline prior to guide the estimation of visible road structures such as lane dividers, road boundaries and crosswalks. This prior incurs less effort to create and update compared to full HD map priors used in other approaches. We also show that training with such a prior can improve generalization at inference time when no prior is available. Ablation studies conducted on the nuScenes and Argoverse 2 dataset demonstrate that models trained using a driveline prior largely retain performance even when priors are not available. On a geographically disjoint split, D2HDMap achieves 44.8 mAP, surpassing recent state-of-the-art. Additionally, noise-aware training substantially increases robustness to realistic localization error.",
      "short_abstract": "Accurate, up-to-date representations of road structures are critical for the safe operation of autonomous vehicles. Existing systems rely either on costly, maintenance-heavy high-definition (HD) maps which compromise safety when outdated, or purely...",
      "published": "2026-06-16",
      "updated": "2026-06-16",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 23,
        "evidence": [
          "abstract: localization",
          "title: HD map",
          "abstract: HD map",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 64,
      "arxiv_url": "https://arxiv.org/abs/2606.20725",
      "pdf_url": "https://arxiv.org/pdf/2606.20725.pdf"
    },
    {
      "id": "2606.17562",
      "anchor": "paper-2606-17562",
      "title": "Anywhere, Any-Stymie: Remote Activation of Trojan Malware on LiDAR with Modulated Signals",
      "authors": [
        "R. Spencer Hallyburton",
        "Miroslav Pajic"
      ],
      "abstract": "LiDAR sensors are widely deployed in autonomous systems for 3D perception and safety-critical decision-making. We identify a previously unexplored attack surface in which dormant malware embedded in the LiDAR sensing pipeline remains inactive during normal operation and can be externally triggered after deployment, without requiring access to sensor hardware or networking at attack time. To operationalize this threat, we design malware capable of low-level point-cloud manipulation and embed it into LiDAR firmware. This malware was developed in a closed research test environment with vendor technical support, rather than by exploiting an inherent production supply-chain vulnerability. To selectively trigger attack activation, we design and implement an optical trigger that remotely activates the malware by delivering a modulated signal into the sensing environment. Once triggered, the malware performs real-time point cloud manipulation, and we demonstrate false object injection and real object suppression on static and mobile victim platforms. Our evaluation first establishes attack feasibility, including static operation at 300~ft and recorded drive-by runs reaching 35~mph. We then illustrate quantitatively that injected person-like artifacts can remain semantically detectable by a state-of-the-art 3D object detector. Finally, we demonstrate multiple modes of safety-critical impact on a deployed tactical autonomous vehicle. Together, these results highlight the need for stronger integrity guarantees throughout the LiDAR sensor development and deployment pipeline.",
      "short_abstract": "LiDAR sensors are widely deployed in autonomous systems for 3D perception and safety-critical decision-making. We identify a previously unexplored attack surface in which dormant malware embedded in the LiDAR sensing pipeline remains inactive during normal...",
      "published": "2026-06-16",
      "updated": "2026-06-16",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 10,
        "evidence": [
          "abstract: perception",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 64,
      "arxiv_url": "https://arxiv.org/abs/2606.17562",
      "pdf_url": "https://arxiv.org/pdf/2606.17562.pdf"
    },
    {
      "id": "2606.17386",
      "anchor": "paper-2606-17386",
      "title": "TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations",
      "authors": [
        "Zikang Xiong",
        "Weixin Li",
        "Zhouchonghao Wu",
        "Akshay Rangesh",
        "Saarth Bonde",
        "Grantland Hall",
        "Chen Tang",
        "Yihan Hu",
        "Wei Zhan"
      ],
      "abstract": "End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deployments. Its standard training recipe, however, is expensive across all stages: collecting and labeling millions of driving frames is costly, and closed-loop RL on images is bottlenecked by the per-step cost of photorealistic rendering plus a forward pass through a large vision backbone. Self-play in vectorized simulators changes the economics: millions of rollout steps per second, and a state distribution naturally rich in collisions, near-misses, and recoveries that no driving log contains. Our approach exploits this asymmetry by decoupling learning to drive from learning to see. We pretrain a single policy by self-play, then align its latent space with a pretrained vision backbone, through the action KL divergence and a batch-relational low-rank structural loss. The action target comes from the self-play policy, so alignment never supervises against a logged trajectory: a paired dataset of (image, scene-state) frames suffices, with no need for the curated expert demonstrations that imitation pretraining is built on. On photorealistic 3D Gaussian splatting closed-loop scenarios, the resulting end-to-end policy matches or exceeds prior end-to-end methods.",
      "short_abstract": "End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deployments. Its standard training recipe, however, is expensive across all stages: collecting and labeling millions of driving frames is costly, and...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 32,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.17386",
      "pdf_url": "https://arxiv.org/pdf/2606.17386.pdf"
    },
    {
      "id": "2606.17362",
      "anchor": "paper-2606-17362",
      "title": "DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models",
      "authors": [
        "Xinglong Sun",
        "Kevin Xie",
        "Jenny Schmalfuss",
        "Despoina Paschalidou",
        "Xiuming Zhang",
        "Sanja Fidler",
        "Kashyap Chitta",
        "Jose M. Alvarez"
      ],
      "abstract": "Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-dependent. Commonly used rule-based driving metrics like EPDMS are interpretable but lack context-awareness, while recent VLMbased evaluations are context-aware but limited by ambiguous VLM outputs and weak physical grounding. To evaluate driving in a manner that is both interpretable and context-aware, we introduce DriveJudge. DriveJudge is a driving evaluation agent that combines rule-grounded evaluation with Vision-Language Model (VLM) reasoning and selectively invokes physically-grounded deterministic rule functions after interpreting the environmental context. To train and evaluate DriveJudge, we curate a large-scale dataset of 33,577 challenging driving samples with human annotations on whether the driving behavior is reasonable in the given scenario. With this dataset, we address the underexplored problem of driving metric evaluation, and introduce two human-aligned benchmark tasks: Driving Quality Classification and Trajectory Preference Selection. DriveJudge outperforms EPDMS for driving quality classification by 21.23 AUC, and the recent VLM-based DriveCritic for trajectory preference selection by 6.5%, setting a new standard for interpretable and precise driving evaluation.",
      "short_abstract": "Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-dependent. Commonly used rule-based driving metrics like EPDMS are interpretable...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.17362",
      "pdf_url": "https://arxiv.org/pdf/2606.17362.pdf"
    },
    {
      "id": "2606.17049",
      "anchor": "paper-2606-17049",
      "title": "BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering",
      "authors": [
        "Yi-Ruei Liu",
        "Jie-Ying Lee",
        "Zheng-Hui Huang",
        "Yu-Lun Liu",
        "Chih-Hao Lin"
      ],
      "abstract": "Inverse rendering of urban scenes from captured videos enables numerous applications, including content creation and autonomous driving simulation. Physically-based rendering methods follow and control lighting physics, but suffer from reconstruction and rendering artifacts. While generative models produce realistic videos, they offer limited consistency and controllability. We present BRDFusion, a unified framework that combines two complementary models for inverse and forward rendering. Specifically, BRDFusion recovers explicit, consistent scene properties with physical modeling and alleviates optimization ambiguity with generative priors. During forward rendering, the physical model provides controllable rendering from the scene configuration, and the generative model denoises and fixes artifacts. Therefore, our method produces high-quality videos while allowing precise control, outperforming baselines in real and synthetic scenes. Moreover, BRDFusion supports novel-view relighting, night simulation, and dynamic object insertion/editing. Project page: https://shigon255.github.io/brdfusion-page/",
      "short_abstract": "Inverse rendering of urban scenes from captured videos enables numerous applications, including content creation and autonomous driving simulation. Physically-based rendering methods follow and control lighting physics, but suffer from reconstruction and...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.17049",
      "pdf_url": "https://arxiv.org/pdf/2606.17049.pdf"
    },
    {
      "id": "2606.17030",
      "anchor": "paper-2606-17030",
      "title": "Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation",
      "authors": [
        "Jie Zhang",
        "Xiaoyue Chen",
        "Anzhe Chen",
        "Dayiheng Liu",
        "Deqing Li",
        "Gengze Zhou",
        "Hale Yin",
        "Haoqi Yuan",
        "Haoyang Li",
        "Jiahao Li",
        "Jiazhao Zhang",
        "Jingren Zhou",
        "Kaiyuan Gao",
        "Kun Yan",
        "Lihan Jiang",
        "Ningyuan Tang",
        "Pei Lin",
        "Qihang Peng",
        "Shengming Yin",
        "Tianhe Wu",
        "Tianyi Yan",
        "Xiao Xu",
        "Yan Shu",
        "Yanran Zhang",
        "Ye Wang"
      ],
      "abstract": "We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual trajectories from current observations across robotic manipulation, autonomous driving, indoor navigation, and human-to-robot transfer. This unified formulation provides three promising application directions: synthetic data generation for policy training augmentation, scalable virtual environments for policy evaluation, and language-guided planning signals for downstream robot control. This is achieved through a three-part design: a) Double-Stream MMDiT with MLLM Action Encoding, where a 60-layer double-stream diffusion transformer couples frozen Qwen2.5-VL semantics with video-VAE latents through layer-wise joint attention; b) Embodied World Knowledge (EWK), an 8.6M video-text corpus (200M+ frames) with action-language mapping over 20+ embodiments and 500+ action categories; and c) General+Expert Progressive Curriculum, a two-stage training strategy that first learns general visual priors and then injects embodied specialization under a shared language interface. Extensive results show strong competitiveness: ranks 1st overall on EWMBench and DreamGen Bench, outperforms all open-source models on WorldModelBench and PBench. Additional zero-shot analyses on RoboTwin-IF benchmark further support robust generalization and multi-view consistency.",
      "short_abstract": "We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual trajectories from current observations across robotic...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Planning & Decision-Making: abstract: planning",
          "Planning & Decision-Making: abstract: navigation",
          "Human Factors & Policy: abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.17030",
      "pdf_url": "https://arxiv.org/pdf/2606.17030.pdf"
    },
    {
      "id": "2606.16996",
      "anchor": "paper-2606-16996",
      "title": "ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation",
      "authors": [
        "Tran Dinh Tien",
        "Zhiqiang Shen"
      ],
      "abstract": "Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes. We introduce ActiveSAM, a training-free, zero-shot inference framework that turns SAM 3 into an active-vocabulary segmenter. ActiveSAM first canonicalizes and expands class prompts, then estimates an image-conditioned active set from a low-resolution presence preview. Only the retained classes are decoded at full resolution, using bucketed prompt multiplexing with the frozen SAM 3 decoder. The preview stage uses only class-presence evidence and skips unnecessary segmentation-head computation, while the final stage applies margin-aware background calibration to suppress low-confidence pixels. ActiveSAM requires no target-dataset training, no weight updates, and no oracle class-presence labels. Across eight OVSS benchmarks, ActiveSAM improves the speed-accuracy tradeoff of training-free open-vocabulary semantic segmentation, outperforming the current state-of-the-art SegEarth-OV3 by approximately +1.4 mIoU on average while running up to 5.5x faster on large-vocabulary datasets. ActiveSAM also demonstrates the strongest robustness under image corruption that simulates real-world distribution shift, making it well-suited for deployment in noisy-input domains such as autonomous driving and embodied AI. Code is available at https://github.com/VILA-Lab/ActiveSAM.",
      "short_abstract": "Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: scene segmentation",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16996",
      "pdf_url": "https://arxiv.org/pdf/2606.16996.pdf"
    },
    {
      "id": "2606.16960",
      "anchor": "paper-2606-16960",
      "title": "SurroundNEXO: Ego-Centric Metric Bridging for Spatially Consistent Geometry in Autonomous Driving",
      "authors": [
        "Shuai Yuan",
        "Runxi Tang",
        "Yuzhou Ji",
        "Fudong Ge",
        "Hanshi Wang",
        "Yifei Wang",
        "Xianming Zeng",
        "Jianyun Xu",
        "Xingliang Liu",
        "Yanfeng Wang",
        "Zhipeng Zhang"
      ],
      "abstract": "Modern autonomous driving depends on accurate metric 3D understanding for perception, reconstruction, and planning, which in turn requires reliable multi-camera depth prediction. However, the outward-facing nature of vehicle-mounted surround-view camera rigs inherently limits visual overlap across views, challenging the correspondence-based assumptions that underpin conventional multi-view geometry. To bridge this gap, we present SurroundNEXO, named after the Spanish word nexo for a geometric link, a low-overlap multi-camera metric depth framework that grounds cross-view reasoning in ego-centric geometry rather than dense visual correspondences. Instead of directly enforcing early global fusion, SurroundNEXO first assigns image tokens globally comparable ego-frame viewing directions through Ego-Ray Positional Encoding, then uses sparse LiDAR measurements as metric anchors to propagate absolute scale cues, and finally expands feature interaction progressively from view-local modeling to decomposed spatio-temporal reasoning and global integration. This design enables metric-scale depth prediction with improved spatial consistency across weakly overlapping cameras. Across low-overlap autonomous driving benchmarks, including NuScenes, Waymo and DDAD, SurroundNEXO reduces single-view error by 33.2%, improves cross-view consistency by 10.5%, and enhances metric reconstruction quality by 25.6% compared with SOTA methods. It further remains robust under extremely sparse depth prompts and exhibits strong zero-shot generalization to unseen camera layouts.",
      "short_abstract": "Modern autonomous driving depends on accurate metric 3D understanding for perception, reconstruction, and planning, which in turn requires reliable multi-camera depth prediction. However, the outward-facing nature of vehicle-mounted surround-view camera rigs...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 5,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16960",
      "pdf_url": "https://arxiv.org/pdf/2606.16960.pdf"
    },
    {
      "id": "2606.16480",
      "anchor": "paper-2606-16480",
      "title": "HOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy Optimization",
      "authors": [
        "Youngjae Min",
        "Jovin D'sa",
        "Faizan M. Tariq",
        "David Isele",
        "Navid Azizan",
        "Sangjae Bae"
      ],
      "abstract": "Robots deployed in the real world must plan motions across diverse scenarios without per-scenario retuning. End-to-end reinforcement learning (RL) can generalize across scenarios but often becomes brittle under distribution shift, reward misspecification, and stochastic interactions. Model predictive path integral (MPPI) control enables strong real-time refinement without gradients, but its performance depends on a well-shaped sampling prior, while manually designing the priors does not scale to multi-scenario deployment. We present HOLO-MPPI (High-level Offline, Low-level Online MPPI), a multi-scenario motion planning framework that combines high-level policy learning with low-level stochastic optimal control. Offline, we learn a high-level policy that proposes scenario-robust plans in an abstract action space, with a learned world model for online rollout. Online, the policy serves as a data-driven prior generator that parameterizes MPPI's sampling distribution conditioned on the current observation and goal. MPPI then optimizes low-level control sequences around this prior in real time to adapt to local disturbances. We instantiate HOLO-MPPI in autonomous driving by designing an effective high-level action space and tailored model architectures. Our evaluation across diverse driving scenarios shows that HOLO-MPPI improves upon MPPI and end-to-end RL baselines while maintaining real-time control.",
      "short_abstract": "Robots deployed in the real world must plan motions across diverse scenarios without per-scenario retuning. End-to-end reinforcement learning (RL) can generalize across scenarios but often becomes brittle under distribution shift, reward misspecification, and...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "World Model",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 16,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16480",
      "pdf_url": "https://arxiv.org/pdf/2606.16480.pdf"
    },
    {
      "id": "2606.16354",
      "anchor": "paper-2606-16354",
      "title": "GraphBEV++: Multi-Modal Feature Alignment for Autonomous Driving",
      "authors": [
        "Ziying Song",
        "Caiyan Jia",
        "Lin Liu",
        "Shaoqing Xu",
        "Lei Yang",
        "Yadan Luo"
      ],
      "abstract": "Feature misalignment in BEV perception is a critical yet often overlooked challenge in autonomous driving, especially under calibration uncertainties between LiDAR and camera sensors. To address this issue, we propose a robust multi-modal fusion framework, GraphBEV++, which systematically mitigates projection-induced misalignment. The framework consists of two key modules: LocalAlign-v2 and GlobalAlign-v2. LocalAlign-v2 introduces neighborhood-aware depth features via graph matching to correct local misalignment. It supports both LSS-based and query-based BEV representations, making it compatible with BEVFusion and BEVFormer architectures for consistent cross-paradigm alignment. GlobalAlign-v2 encompasses two variants: Deformable and Diffusion. The Deformable variant addresses global misalignment in LSS-based multi-modal BEV by explicitly learning cross-modal feature offsets. In contrast, the Diffusion variant targets implicit misalignment in query-based BEV by injecting noise to simulate misalignment and employing a denoising process to recover aligned features. Experimental results show that GraphBEV++ achieves state-of-the-art performance under misalignment noise on nuScenes and Waymo subset, improves long-range detection on Argoverse2, and generalizes effectively to the 3D occupancy prediction task, consistently improving occupancy estimation accuracy and robustness under both clean and noisy settings. Furthermore, GraphBEV++ effectively alleviates misalignment issues in end-to-end autonomous driving. Compared with five baselines (UniAD, VAD, FusionAD, MomAD, and WoTE), it demonstrates superior performance in both open-loop (nuScenes) and closed-loop (Bench2Drive and NAVSIM) evaluations across perception, prediction, and planning tasks.",
      "short_abstract": "Feature misalignment in BEV perception is a critical yet often overlooked challenge in autonomous driving, especially under calibration uncertainties between LiDAR and camera sensors. To address this issue, we propose a robust multi-modal fusion framework,...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 3,
        "evidence": [
          "abstract: perception",
          "abstract: occupancy",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16354",
      "pdf_url": "https://arxiv.org/pdf/2606.16354.pdf"
    },
    {
      "id": "2606.16313",
      "anchor": "paper-2606-16313",
      "title": "Is Your Trajectory Displacement Safe in Long-tail?",
      "authors": [
        "Qiao Sun",
        "Weicheng Zheng",
        "Yixin Huang",
        "Hang Zhao"
      ],
      "abstract": "Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligned, safety-aware, verifiable, and explainable at the same time: closed-loop metrics often saturate among strong planners, while unstructured human ratings can be noisy without a carefully designed protocol. We formulate planning evaluation as additional-threat detection: given a planner trajectory and an expert reference, does the planner's displacement introduce new unsafe driving behavior? We propose FluidTest, an evaluation pipeline with three components: a pairwise WebUI protocol for reliable human annotation; a taxonomy of 32 semantic threats with evidence-grounded decision graphs; and a three-agent verification system with reflection for precision and auditability. Experiments on the WOD-E2E dataset show that FluidTest produces consistent labels among trained annotators and identifies additional threats in 65% of Poutine trajectories and 51% of RAP trajectories. These results show that state-of-the-art planners can still exhibit substantial safety-relevant failures despite high Rater Feedback Scores (RFS) and low Average Displacement Error (ADE). Additional details, guidance, and code are available at https://fluidtest.web.app.",
      "short_abstract": "Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligned, safety-aware, verifiable, and explainable at the same time: closed-loop...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16313",
      "pdf_url": "https://arxiv.org/pdf/2606.16313.pdf"
    },
    {
      "id": "2606.16278",
      "anchor": "paper-2606-16278",
      "title": "RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos",
      "authors": [
        "Zhenhua Wu",
        "Yun Pang",
        "Mingkun Chang",
        "Yuwei Ning",
        "Liangzhi Wang",
        "Yi Xiao",
        "Guanbin Li"
      ],
      "abstract": "Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D Gaussian Splatting (3DGS) simulation offers a scalable alternative through real-scene reconstruction and controllable editing. However, edited 3DGS-rendered videos often exhibit a significant Sim-to-Real gap, manifested as rendering artifacts, degraded foreground assets, illumination mismatch, and temporal flickering. Addressing these coupled defects requires jointly restoring local appearance, harmonizing edited content, and maintaining temporal consistency, whereas existing methods typically address only a subset of these requirements. To fill this gap, we propose RealityBridge, a video restoration and harmonization framework that converts edited 3DGS renderings into realistic driving footage while preserving simulator-defined structure, edits, and dynamics. RealityBridge conditions a video foundation model on complementary modality signals, with a lightweight GateNet adaptively controlling their injection across backbone blocks. We further develop a task-oriented curation pipeline to construct training data, and design a four-stage supervised training strategy followed by reward-guided post-training. Extensive experiments demonstrate that RealityBridge outperforms existing methods in restoration and harmonization while preserving strong temporal consistency.",
      "short_abstract": "Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D Gaussian Splatting (3DGS) simulation offers a scalable alternative through real-scene reconstruction and controllable...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16278",
      "pdf_url": "https://arxiv.org/pdf/2606.16278.pdf"
    },
    {
      "id": "2606.16274",
      "anchor": "paper-2606-16274",
      "title": "GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving",
      "authors": [
        "Ziying Song",
        "Caiyan Jia",
        "Lin Liu",
        "Lei Yang",
        "Shengkai Zhang",
        "Feiyang Jia",
        "Fengda Zhao",
        "Peiliang Wu",
        "Shaoqing Xu",
        "Chen Lv",
        "Yadan Luo"
      ],
      "abstract": "End-to-end autonomous driving has made significant progress by unifying perception, prediction, and planning within a single learning framework, achieving strong performance in short-horizon decision making. However, most existing E2E-AD methods remain confined to short-horizon planning and lack the ability to model long-term temporal dependencies, which severely limits their generalization and security in complex and highly interactive driving scenarios. In this work, we propose GraphWorld, an E2E-AD framework that explicitly enhances long-horizon planning through latent world modeling. We introduce an Ego-Centric Interaction Graph, which adaptively models critical neighboring agents based on spatial proximity, and propagates relational context to planning queries via cross-node cross-attention. We present a World-State-Conditioned Planning that learns ego-centric latent world representations by modeling interactions between an ego vehicle and surrounding agents. This latent world state captures key interaction dynamics and safety-relevant semantics, and serves as a conditioning signal to guide long-horizon, safety-aware trajectory planning. Extensive experiments on Bench2Drive, NAVSIMv1/2, and nuScenes demonstrate that GraphWorld significantly reduces collision rates and improves long-horizon planning performance, validating its effectiveness in complex driving environments.",
      "short_abstract": "End-to-end autonomous driving has made significant progress by unifying perception, prediction, and planning within a single learning framework, achieving strong performance in short-horizon decision making. However, most existing E2E-AD methods remain...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 15,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16274",
      "pdf_url": "https://arxiv.org/pdf/2606.16274.pdf"
    },
    {
      "id": "2606.17482",
      "anchor": "paper-2606-17482",
      "title": "SPHINX: First Explain, Then Explore",
      "authors": [
        "Nguyen Do",
        "Tue M. Cao",
        "Tien Van Do",
        "András Hajdu",
        "Tamás Bérczes",
        "My T. Thai"
      ],
      "abstract": "Generating adversarial driving scenarios is critical for evaluating and improving autonomous vehicle decision-making systems in simulation. Recent approaches rely primarily on the prior knowledge of Large Language Models and Vision-Language Models to generate driving scenarios procedurally. We argue that adversarial scenes should be generated based on the failure diagnosis (e.g., indecisiveness, multi-frame inconsistency) of the driving policy to specifically address the policy's weaknesses instead of relying on prior assumptions. In this paper, we propose SPHINX, a closed-loop framework for adversarial scenario synthesis guided by a simple principle: first explain, then explore. Beyond blindly exploring the scenario space, SPHINX leverages explainable artificial intelligence methods to analyze the policy, identifying key visual concepts and their influence on policy outputs, and the uncertainty of the decisions. Given the interpretable evidence extracted from the policy's own decision process, we use a vision language model to rationalize and criticize failure modes of the current policy. These critics are then used to generate targeted adversarial scenarios for policy retraining and improvement. We demonstrate that SPHINX can highlight an interpretable account of policy failures while other adversarial scene generation cannot. Across the evaluated benchmarks and test suites, SPHINX can be applied to diverse state-of-the-art autonomous vehicle architectures and yields consistent robustness improvements over existing scenario-generation methods.",
      "short_abstract": "Generating adversarial driving scenarios is critical for evaluating and improving autonomous vehicle decision-making systems in simulation. Recent approaches rely primarily on the prior knowledge of Large Language Models and Vision-Language Models to generate...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Synthetic Data",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 3,
        "evidence": [
          "abstract: attack or threat",
          "abstract: robustness",
          "abstract: uncertainty",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.17482",
      "pdf_url": "https://arxiv.org/pdf/2606.17482.pdf"
    },
    {
      "id": "2606.17451",
      "anchor": "paper-2606-17451",
      "title": "Credibility-Weighted Pricing of Autonomous Vehicle Liability Under Operational Design Domain Shift",
      "authors": [
        "Doyeon Jang"
      ],
      "abstract": "Automated Driving System deployments create a foundational ratemaking challenge: sparse experience, shifting operational design domains, and non-stationary risk across software releases. We propose a hierarchical Bayesian credibility framework pooling across cities, software versions, and territories via a learned ODD-similarity kernel, nesting Buhlmann-Straub as a limiting case. Demonstrated on 648 verified-engaged Waymo crashes across four U.S. metros from the NHTSA Standing General Order database against 116 million matched miles, city-aggregate credibility weights are moderate (0.12-0.46), partial pooling decisively outperforms no pooling, and a power analysis shows the learned kernel's advantage becomes detectable at approximately twelve deployed cities.",
      "short_abstract": "Automated Driving System deployments create a foundational ratemaking challenge: sparse experience, shifting operational design domains, and non-stationary risk across software releases. We propose a hierarchical Bayesian credibility framework pooling across...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.17451",
      "pdf_url": "https://arxiv.org/pdf/2606.17451.pdf"
    },
    {
      "id": "2606.16735",
      "anchor": "paper-2606-16735",
      "title": "Pride and Prejudice: Toward an Information-Theoretic Framework for Mutually Communicative Driver Behavior Modeling",
      "authors": [
        "Tingjun Li",
        "Nan Xu",
        "Shuo Feng",
        "Hassan Askari",
        "Bruno Henrique Groenner Barbosa",
        "Konghui Guo"
      ],
      "abstract": "Mixed autonomy driving becomes unsafe and inefficient when autonomous vehicles (AVs) and human-driven vehicles (HVs) misread each other's intentions. We study this problem as implicit mutual communication in lane changes. The proposed framework models how the ego vehicle both expresses its intent and probes the other driver's preference under epistemic uncertainty. It combines a level-k Bayesian persuasion game with virtual features for proactive signaling, information-theoretic rewards for mutual communication, and adaptive weights of communication affordances. We further introduce the Pride-Inquiry (P-I) and Pride-Prejudice (P-P) planes to analyze communication intensity and tendency. The model is calibrated with a Communication-Based Multi-Agent Inverse Reinforcement Learning algorithm (C-MIRL) on the naturalistic NGSIM dataset. Compared with the non-communicative baseline, the proposed model reduces the prediction error of mandatory lane changes by up to 20% while maintaining strong generalization. Driver-In-the-Loop questionnaire scores are positively correlated with the calibrated communication variables, supporting the subjective validity of the model. The learned rewards further show that inquiry and listening affordances contribute more than pride and expression alone, and that inquiry preference varies more strongly across drivers. These results support explicit modeling of mutual communication and epistemic uncertainty in interactive driving.",
      "short_abstract": "Mixed autonomy driving becomes unsafe and inefficient when autonomous vehicles (AVs) and human-driven vehicles (HVs) misread each other's intentions. We study this problem as implicit mutual communication in lane changes. The proposed framework models how the...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 13,
        "evidence": [
          "title: driver monitoring"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16735",
      "pdf_url": "https://arxiv.org/pdf/2606.16735.pdf"
    },
    {
      "id": "2606.16589",
      "anchor": "paper-2606-16589",
      "title": "DRIFT: Risk-Constrained Diffusion with Imitation Priors for Mixed-Autonomy Traffic Generation",
      "authors": [
        "Yaoshen Yu",
        "Minghui Liwang",
        "Wenbo Zhu",
        "Xinlei Yi",
        "Zhang Liu",
        "Yiguang Hong",
        "Xianbin Wang",
        "Seyyedali Hosseinalipour"
      ],
      "abstract": "Future intelligent transportation systems are envisioned to evolve toward a long-term mixed-autonomy paradigm, where human-driven vehicles (HVs) and autonomous vehicles (AVs) coexist within highly coupled traffic ecosystems. Such coexistence introduces pronounced heterogeneity, amplified uncertainty, and increasingly intricate interaction dynamics. In this context, it remains fundamentally challenging to simultaneously capture the heterogeneous behavioral distribution shifts arising from dynamic AV penetration, generate diverse yet executable trajectories under strong inter-vehicle coupling, and conduct reliable closed-loop safety and stability diagnostics for rare but high-impact events. To this end, we present Diffusion with Risk constraints, Imitation priors, and long-tail Feedback for mixed-autonomy Traffic generation (DRIFT), a mixed-autonomy traffic generation framework that unifies heterogeneity-aware conditional encoding, conditional diffusion-based executable trajectory generation, and progressive adversarial alignment enhanced by risk-aware long-tail feedback, thereby enabling traffic behaviors to be iteratively generated, filtered, selected, and validated within a closed-loop execution pipeline. In addition, a unified evaluation protocol is developed to jointly characterize safety, efficiency, and closed-loop stability across representative traffic scenarios and AV penetration regimes. Experimental results demonstrate that DRIFT achieves a strong safety-efficiency trade-off in closed-loop mixed-autonomy benchmarks, while further revealing the critical influence of candidate executability, online selection, and long-tail feedback on executable traffic evolution.",
      "short_abstract": "Future intelligent transportation systems are envisioned to evolve toward a long-term mixed-autonomy paradigm, where human-driven vehicles (HVs) and autonomous vehicles (AVs) coexist within highly coupled traffic ecosystems. Such coexistence introduces...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 11,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: risk assessment",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16589",
      "pdf_url": "https://arxiv.org/pdf/2606.16589.pdf"
    },
    {
      "id": "2606.16374",
      "anchor": "paper-2606-16374",
      "title": "Transient-Safe Platooning via Dynamic Headway",
      "authors": [
        "Gal Barkai",
        "J Ér Émie Kreiss",
        "Vineeth S Varma",
        "Irinel-Constantin Morarescu"
      ],
      "abstract": "Managing autonomous vehicle platoons requires a delicate balance between string stability and rigorous safety. This challenge is intensified by aggressive transients, such as highway merging. Although Constant Time Headway (CTH) spacing is the industry standard for Cooperative Adaptive Cruise Control, it lacks formal safety guarantees during significant velocity deviations. This letter proposes a computationally efficient control framework that considers a linear time-invariant model for the dynamics of each vehicle, while ensuring formal transient safety and stability. By introducing a spacing policy that naturally converges to CTH at steady state, we establish platoon safety as an inductive property. We derive a non-linear and saturated control law for the lead follower and provide sufficient initial conditions to guarantee velocity non-negativity and safety throughout the platoon for any CTH-based followers' control law. Numerical examples indicate the proposed methodology may be applicable even under non-nominal setups.",
      "short_abstract": "Managing autonomous vehicle platoons requires a delicate balance between string stability and rigorous safety. This challenge is intensified by aggressive transients, such as highway merging. Although Constant Time Headway (CTH) spacing is the industry...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16374",
      "pdf_url": "https://arxiv.org/pdf/2606.16374.pdf"
    },
    {
      "id": "2606.16570",
      "anchor": "paper-2606-16570",
      "title": "Automated Digital Twin Construction for Highway Scenarios Using LiDAR Point Clouds and OpenStreetMap",
      "authors": [
        "Yongqi Zhao",
        "Dong Bi",
        "Paul Kovacevic",
        "Tomislav Mihalj",
        "Martin Schabauer",
        "Johannes Betz",
        "Arno Eichberger"
      ],
      "abstract": "Accurate road environment modeling is fundamental to the simulation and validation of automated driving systems. However, constructing road maps in standardized formats such as ASAM OpenDRIVE from real-world sensor data remains a time-consuming and costly process. Mobile mapping LiDAR captures accurate lane-level geometry but is confined to the driven corridor, while OpenStreetMap (OSM) provides broad road network topology but lacks geometric precision at the lane level. To address this, an automated workflow is proposed to fuse LiDAR point clouds with OSM data to generate georeferenced ASAM OpenDRIVE maps of highway environments, requiring minimal manual intervention. The pipeline reconstructs mainline roads from LiDAR-derived measurements and infers ramp geometry and topology from the OSM road graph, enabling complete highway interchange modeling without full sensor coverage. Experiments demonstrate a mean lateral RMSE of 0.740 m, and the generated maps are directly usable in mainstream simulation platforms including IPG CarMaker and Esmini. These results validate the effectiveness of combining measurement-derived geometry with map-derived topology for automated OpenDRIVE digital twin generation. The project code is available at https://github.com/ftgTUGraz/opendrive-digital-twin-generator",
      "short_abstract": "Accurate road environment modeling is fundamental to the simulation and validation of automated driving systems. However, constructing road maps in standardized formats such as ASAM OpenDRIVE from real-world sensor data remains a time-consuming and costly...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 20,
        "evidence": [
          "title: point cloud",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16570",
      "pdf_url": "https://arxiv.org/pdf/2606.16570.pdf"
    },
    {
      "id": "2606.16558",
      "anchor": "paper-2606-16558",
      "title": "ROSA-RL: Uncertainty-Aware Roundabout Optimized Speed Advisory with Reinforcement Learning",
      "authors": [
        "Anna-Lena Schlamp",
        "Jeremias Gerner",
        "Klaus Bogenberger",
        "Werner Huber",
        "Stefanie Schmidtner"
      ],
      "abstract": "Roundabouts challenge automated driving in mixed traffic, as heterogeneous and non-deterministic human behavior, unknown driving intentions, and high interaction complexity create uncertainty about whether the conflict zone will be blocked or available at the moment of entry. We present ROSA-RL -- uncertainty-aware Roundabout Optimized Speed Advisory with Reinforcement Learning. It enables safe and efficient roundabout entry for automated and human-driven vehicles in mixed traffic through probabilistic conflict forecasting. A Transformer-based model predicts conflict zone occupancy over a five-second horizon, capturing multi-agent interactions to anticipate upcoming conflicts and available gaps. The prediction outputs encode uncertainty in future motion and intent, and augment the state of a classical RL framework, enabling uncertainty-aware speed coordination. Evaluated in simulations grounded in real-world data, ROSA-RL can effectively handle uncertainty and outperform a comparable model-based baseline, closing the gap to an ideal setting assuming fully known occupancy while improving traffic efficiency and safety. The source code of this work is available under: github.com/urbanAIthi/ROSA-RL.",
      "short_abstract": "Roundabouts challenge automated driving in mixed traffic, as heterogeneous and non-deterministic human behavior, unknown driving intentions, and high interaction complexity create uncertainty about whether the conflict zone will be blocked or available at the...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "title: uncertainty",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16558",
      "pdf_url": "https://arxiv.org/pdf/2606.16558.pdf"
    },
    {
      "id": "2606.16453",
      "anchor": "paper-2606-16453",
      "title": "Impact of ADAS and V2X Penetration Rates on Cooperative Active Safety",
      "authors": [
        "M. C. Lucas-Estañ",
        "B. Coll-Perales",
        "M. I. Khan",
        "J. Gozalvez",
        "O. Altintas"
      ],
      "abstract": "A major driver of connected and automated driving is cooperative active safety. The effectiveness of cooperative safety applications depends on the ability of vehicles to detect traffic safety risks in advance. Such risks can be identified either through ADAS (Advanced Driving Assistance Systems) or via V2X (Vehicle to Everything) communications. More vehicles are gradually being deployed with ADAS, but ADAS sensors can be limited by their sensing range and field of view. On the other hand, V2X can experience communication ranges beyond the ADAS sensing range, but its impact is highly dependent on the V2X penetration rate. This paper analyzes the impact of ADAS and V2X penetration rates on the effectiveness of cooperative active safety applications considering an emergency braking maneuver use case in a highway scenario. Results show that while ADAS and V2X each enhance traffic safety, their combined deployment further amplifies these gains, with the effect becoming more pronounced as V2X is deployed more rapidly.",
      "short_abstract": "A major driver of connected and automated driving is cooperative active safety. The effectiveness of cooperative safety applications depends on the ability of vehicles to detect traffic safety risks in advance. Such risks can be identified either through ADAS...",
      "published": "2026-06-15",
      "updated": "2026-06-15",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 25,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 65,
      "arxiv_url": "https://arxiv.org/abs/2606.16453",
      "pdf_url": "https://arxiv.org/pdf/2606.16453.pdf"
    },
    {
      "id": "2606.20687",
      "anchor": "paper-2606-20687",
      "title": "ARGUSTRACK: A Multi-View Annotation System for Multi-Object Tracking",
      "authors": [
        "Hao Vo",
        "Duc Nguyen",
        "Ngan Le"
      ],
      "abstract": "Multi-Camera Multi-Target (MCMT) tracking has emerged as a critical capability for applications ranging from autonomous driving to animal behavior monitoring. While recent advances have yielded sophisticated tracking algorithms, the availability of annotated multi-view data remains a significant bottleneck. Existing annotation tools predominantly support single-camera workflows or rely on LiDAR sensors, making cross-view labeling tedious and impractical for camera-only setups. We present ARGUS-TRACK, a multi-camera annotation system that addresses these limitations by enabling annotators to work directly on a bird's-eye-view (BEV) plane. Given calibrated camera parameters, a single ground-plane annotation is automatically projected into 2D bounding boxes across all relevant views, inherently ensuring identity consistency without manual cross-view alignment. To further accelerate the labeling process, ARGUSTRACK incorporates two complementary mechanisms: a Temporal Aware module that propagates annotations from preceding frames to initialize new ones, requiring only minor positional adjustments; and a Multi-camera Semi-annotation module that leverages off-the-shelf 2D detectors combined with foot-point estimation to automatically generate candidate BEV positions for annotator verification. We evaluate ARGUSTRACK through a pilot study on multi-camera broiler tracking and demonstrate that it substantially reduces annotation time compared to conventional single-camera labeling workflows.",
      "short_abstract": "Multi-Camera Multi-Target (MCMT) tracking has emerged as a critical capability for applications ranging from autonomous driving to animal behavior monitoring. While recent advances have yielded sophisticated tracking algorithms, the availability of annotated...",
      "published": "2026-06-14",
      "updated": "2026-06-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 2,
        "evidence": [
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 66,
      "arxiv_url": "https://arxiv.org/abs/2606.20687",
      "pdf_url": "https://arxiv.org/pdf/2606.20687.pdf"
    },
    {
      "id": "2606.16048",
      "anchor": "paper-2606-16048",
      "title": "CloudDiffusion: Diffusion-Based Scene Completion in the Point Cloud Domain",
      "authors": [
        "Chidera Agbasiere",
        "Mikhail Sannikov",
        "Faith Ogunwoye",
        "Erik Shaikhiev",
        "Alex Kozinov",
        "Ilya Mikhalchuk",
        "Iana Zhura",
        "Dzmitry Tsetserukou"
      ],
      "abstract": "Reconstructing dense 3D scenes from sparse LiDAR point clouds (LiDAR scene completion) is a fundamental challenge in autonomous driving, where diffusion models offer a promising solution. However, existing approaches rely on object-level autoencoders that collapse into unstable global representations at outdoor scale, and suffer from ground truth data corrupted by odometry drift that systematically degrades supervision quality. Furthermore, multi-step diffusion inference incurs prohibitive latency for real-time deployment. We present CloudDiffusion, addressing these issues with three independent components. First, a multi-token Gaussian VAE with cross-attention pooling provides stable scene-scale LiDAR compression as a standalone reconstruction module, avoiding the global-pooling and codebook-collapse failure modes of prior point-cloud autoencoders. Second, an anchor-based ICP ground truth refinement pipeline eliminates drift-induced noise from training supervision, reducing our single-step x0 diffusion teacher's squared Chamfer distance by approximately 16x on SemanticKITTI seq. 08 (0.396 to 0.024 m^2) with no model change (partly aided by the denser, more compact refined references). Third, the same teacher completes scenes in a single x0 step, operating directly in coordinate space, not in the VAE latent. It runs in near real time at 209ms/frame, 65-138x lower inference latency than iterative diffusion baselines. Our results indicate that data quality dominates model design in this regime, and suggest that multi-token latent spaces could serve as a stable first stage for future latent diffusion-based scene completion.",
      "short_abstract": "Reconstructing dense 3D scenes from sparse LiDAR point clouds (LiDAR scene completion) is a fundamental challenge in autonomous driving, where diffusion models offer a promising solution. However, existing approaches rely on object-level autoencoders that...",
      "published": "2026-06-14",
      "updated": "2026-06-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 10,
        "evidence": [
          "title: point cloud",
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 66,
      "arxiv_url": "https://arxiv.org/abs/2606.16048",
      "pdf_url": "https://arxiv.org/pdf/2606.16048.pdf"
    },
    {
      "id": "2606.15930",
      "anchor": "paper-2606-15930",
      "title": "ControlMap: Controllable High-Definition Map Generation for Traffic Scenario Simulation",
      "authors": [
        "Marwan Farag",
        "Steffen Wäldele",
        "Yu Yao"
      ],
      "abstract": "Simulation is central to validating autonomous driving systems, yet current pipelines are limited by insufficient scenario diversity due to costly High Definition (HD) map creation. Scaling HD maps requires expensive data collection and manual processing. Moreover, existing generative models lack the fine-grained control necessary to target specific road topologies during generation. This paper presents a data-driven pipeline for controllable HD map generation using latent diffusion and ControlNet for spatial conditioning. To our knowledge, we are the first to inject spatial guidance signals into a diffusion model for HD map synthesis. Furthermore, our model supports adjustable conditioning strength through classifier-free guidance and city-level style transfer via city label conditioning. To complement existing metrics, we introduce two novel metrics to evaluate adherence to the control signal and similarity to ground-truth maps. Experiments demonstrate that our model generates realistic HD maps that faithfully follow input road topologies while accurately preserving city-specific details.",
      "short_abstract": "Simulation is central to validating autonomous driving systems, yet current pipelines are limited by insufficient scenario diversity due to costly High Definition (HD) map creation. Scaling HD maps requires expensive data collection and manual processing....",
      "published": "2026-06-14",
      "updated": "2026-06-14",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 18,
        "evidence": [
          "title: HD map",
          "abstract: HD map"
        ]
      },
      "recency": "fresh",
      "age_days": 66,
      "arxiv_url": "https://arxiv.org/abs/2606.15930",
      "pdf_url": "https://arxiv.org/pdf/2606.15930.pdf"
    },
    {
      "id": "2606.15869",
      "anchor": "paper-2606-15869",
      "title": "Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation",
      "authors": [
        "Jingyu Li",
        "Zhe Liu",
        "Dongnan Hu",
        "Junjie Wu",
        "Zipei Ma",
        "Wenxiao Wu",
        "Chao Han",
        "Zhihui Hao",
        "Zhikang Liu",
        "Kun Zhan",
        "Jiankang Deng",
        "Xiatian Zhu",
        "Li Zhang"
      ],
      "abstract": "World action models~(WAMs) have shown great promise for autonomous driving and urban navigation. Built upon Vision-Language-Action models or video generation models, existing approaches suffer key limitations: (1) High inference latency due to future observation prediction at test time, and (2) tightly coupled video and action modeling leading to representational mismatch and degraded generalization. To address both issues, we propose Metis, an end-to-end WAM framework that decouples video generation and action prediction. Specifically, Metis employs a Mixture-of-Transformers architecture with dedicated experts for video generation and action prediction, preserving the intrinsic distributional properties of each task. To enhance efficiency, we introduce an asymmetric attention mask that enables joint training of both experts while allowing the action model to bypass explicit video generation during inference. This design ensures training-inference consistency and significantly reduces computational costs without compromising planning performance. Extensive experiments demonstrate state-of-the-art performance on the NAVSIM navhard and navtest benchmarks and the CityWalker navigation benchmark, validating both the generalizability and efficiency across diverse tasks. Real-robot deployments further confirm the practical feasibility of our approach.",
      "short_abstract": "World action models~(WAMs) have shown great promise for autonomous driving and urban navigation. Built upon Vision-Language-Action models or video generation models, existing approaches suffer key limitations: (1) High inference latency due to future...",
      "published": "2026-06-14",
      "updated": "2026-06-14",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "VLA",
        "World Model",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 11,
        "margin": 2,
        "evidence": [
          "abstract: planning",
          "title: navigation",
          "abstract: navigation"
        ]
      },
      "recency": "fresh",
      "age_days": 66,
      "arxiv_url": "https://arxiv.org/abs/2606.15869",
      "pdf_url": "https://arxiv.org/pdf/2606.15869.pdf"
    },
    {
      "id": "2606.15756",
      "anchor": "paper-2606-15756",
      "title": "From Correlation to Causation in Lane Change Prediction for Automated Driving: A Causal Explanation Framework",
      "authors": [
        "Mohamed Manzour",
        "Aditya Kumar",
        "Augusto Luis Ballardini",
        "Miguel Ángel Sotelo"
      ],
      "abstract": "Lane-change prediction is a central task in intelligent vehicles, where early maneuver anticipation can support safer decision-making. However, many existing approaches mainly learn statistical associations between observed driving variables and future maneuvers, while overlooking the causal dependencies among the input variables themselves. This limits interpretability, especially when physically related variables such as longitudinal gap, relative longitudinal velocity, and Time-To-Collision (TTC) are treated as independent flat inputs. This article presents a causal-inference-based framework for lane-change prediction and explanation. The proposed approach combines linguistic feature construction, expert-constrained causal discovery, deep structural causal modeling with Deep End-to-end Causal Inference (DECI), intervention-based effect analysis, refutation testing, and recursive causal-chain explanation. The objective is not only to predict the future maneuver, but also to identify candidate variables that directly contribute to the prediction, the upstream factors influencing them, and the causal chains through which these effects propagate. The framework achieves average F1-scores above 95% during the first three seconds before the lane-marking crossing event. Beyond prediction accuracy, the framework uses intervention-based effect analysis to distinguish influential from weakly influential variables under the learned causal structure. It further distinguishes candidate direct contributors from mediated effects and generates contrastive causal-chain explanations that clarify why the predicted maneuver is favored and why the alternative maneuvers are less supported. The main contribution is therefore a mechanism-aware lane-change prediction pipeline that moves beyond correlation-based classification toward more interpretable causal reasoning for maneuver prediction.",
      "short_abstract": "Lane-change prediction is a central task in intelligent vehicles, where early maneuver anticipation can support safer decision-making. However, many existing approaches mainly learn statistical associations between observed driving variables and future...",
      "published": "2026-06-14",
      "updated": "2026-06-14",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 66,
      "arxiv_url": "https://arxiv.org/abs/2606.15756",
      "pdf_url": "https://arxiv.org/pdf/2606.15756.pdf"
    },
    {
      "id": "2606.15341",
      "anchor": "paper-2606-15341",
      "title": "CausalDrive: Real-time Causal World Models for Autonomous Driving",
      "authors": [
        "Tianyi Yan",
        "Huan Zheng",
        "Dubing Chen",
        "Meizhi Qu",
        "Yingying Shen",
        "Lijun Zhou",
        "Mingfei Tu",
        "Bing Wang",
        "Guang Chen",
        "Hangjun Ye",
        "Haiyang Sun",
        "Cheng-zhong Xu",
        "Jianbing Shen"
      ],
      "abstract": "World models have emerged as a promising paradigm for scaling autonomous driving (AD) data, yet existing video generative models fall short as interactive simulators. Layout-conditioned renderers rely on \"oracle\" future trajectories of all background agents, rendering them strictly non-reactive. Conversely, pure action-conditioned predictors lack semantic control over complex interactions and suffer from prohibitive diffusion latencies, hindering closed-loop policy learning. To bridge this gap, we present CausalDrive, a controllable, real-time foundation driving world renderer. CausalDrive operates solely on the initial front-view frame, the ego-vehicle's trajectory, and a macroscopic text prompt. By excluding future NPC layouts, we compel the model to intrinsically predict causal interactions, enabling text-driven control over Driving Sociology, allowing users to dynamically orchestrate diverse counterfactual reactions to identical ego-actions. To overcome the efficiency bottleneck and address the covariate shift in autoregressive generation, we propose a novel Context-Forced DMD architecture. This combines continuous flow-matching with a self-correcting distillation objective, achieving interactive speeds of 12 FPS. This breakthrough transforms the passive video generator into a playable neural simulator. We demonstrate its versatility across three downstream applications: (1) generative closed-loop evaluation with significantly mitigated collision artifacts, (2) large-scale Reinforcement Learning (RL) post-training driven by a Video2Reward module, and (3) real-time human-in-the-loop simulation. Extensive experiments validate that policies trained within CausalDrive's reactive scenarios exhibit superior interaction capabilities in the real world.",
      "short_abstract": "World models have emerged as a promising paradigm for scaling autonomous driving (AD) data, yet existing video generative models fall short as interactive simulators. Layout-conditioned renderers rely on \"oracle\" future trajectories of all background agents,...",
      "published": "2026-06-13",
      "updated": "2026-06-13",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "World Model",
        "Reinforcement Learning",
        "Hardware / Real-Time",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 2,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 67,
      "arxiv_url": "https://arxiv.org/abs/2606.15341",
      "pdf_url": "https://arxiv.org/pdf/2606.15341.pdf"
    },
    {
      "id": "2606.15139",
      "anchor": "paper-2606-15139",
      "title": "Self-Driving Negotiator: An interactive, verifiable benchmark for social negotiation and theory of mind under hidden intent",
      "authors": [
        "Ashutosh Kumar"
      ],
      "abstract": "Autonomous driving is full of tiny social negotiations: a driver presses forward, another yields, a pedestrian fakes toward the curb, or a lane vehicle chooses whether to open a merge gap. Such interactions require inferring hidden intent from behavior under partial observability and then acting safely and efficiently. Existing autonomous-driving language benchmarks mostly focus on perception, visual question answering, or open-loop planning, while existing language-agent negotiation benchmarks typically make the negotiation explicit in text. Self-Driving Negotiator bridges the gap between the two: a text-only, multi-turn, procedurally generated environment for measuring implicit social coordination in driving. Agents generate specific driving actions. Reward and diagnostics are computed from the privileged simulator state, not from the explanation of the model. This report covers task design, reward and anti-gaming invariants, validated scenarios, non-LLM baselines, and a six-model inference leaderboard. Current models are far removed from the scripted expert. The best average success rate across three scenarios is 0.68; contested merge is statistically flat across models; and difficulty tiers separate cue-following from true wait-for-commitment behavior.",
      "short_abstract": "Autonomous driving is full of tiny social negotiations: a driver presses forward, another yields, a pedestrian fakes toward the curb, or a lane vehicle chooses whether to open a merge gap. Such interactions require inferring hidden intent from behavior under...",
      "published": "2026-06-13",
      "updated": "2026-06-13",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 67,
      "arxiv_url": "https://arxiv.org/abs/2606.15139",
      "pdf_url": "https://arxiv.org/pdf/2606.15139.pdf"
    },
    {
      "id": "2606.15094",
      "anchor": "paper-2606-15094",
      "title": "Adaptive Deep Koopman Operator for Vehicle Dynamics Modeling: A Physics-Informed and Tire-Force-Driven Approach",
      "authors": [
        "Wenjie Wang",
        "Hao Chen",
        "Ran Shu",
        "Solyeon Kwon",
        "Kyoungseok Han",
        "Hongyu Shu"
      ],
      "abstract": "Accurate and adaptive modeling of vehicle dynamics is paramount for the safety of autonomous driving systems, particularly under extreme maneuvers and time-varying parameters. While Deep Koopman operator theory offers a promising global linearization framework, its online application faces a theoretical bottleneck: the high-dimensional lifted state space inherently induces a rank-deficient problem, rendering traditional recursive least squares based updates numerically unstable. To address this, we propose a novel tire-force-driven modeling framework with guaranteed online stability. First, an offline Deep Koopman model is constructed by embedding 7DOF dynamic equilibrium constraints into the learning objective, ensuring the structural fidelity and physical interpretability of the lifted manifold. Second, we theoretically reformulate the operator update in the rank-deficient space as a minimum-norm solution problem. A Physics-Informed Variable Step-Size Normalized Least Mean Squares (PI-VSS-NLMS) algorithm is proposed, which leverages the projection property of NLMS to act as a stable pseudo-inverse solver while incorporating an anchoring mechanism to suppress parameter drift. Extensive simulations on CarSim and Hardware-in-the-Loop validation on dSPACE MicroAutobox III confirm the superiority of the proposed algorithm. It achieves robust prediction accuracy under unseen excitations while guaranteeing real-time feasibility with an average execution time of 0.421 ms, thus bridging the gap between theoretical models and practical deployment.",
      "short_abstract": "Accurate and adaptive modeling of vehicle dynamics is paramount for the safety of autonomous driving systems, particularly under extreme maneuvers and time-varying parameters. While Deep Koopman operator theory offers a promising global linearization...",
      "published": "2026-06-13",
      "updated": "2026-06-13",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 11,
        "evidence": [
          "title: vehicle dynamics",
          "abstract: vehicle dynamics"
        ]
      },
      "recency": "fresh",
      "age_days": 67,
      "arxiv_url": "https://arxiv.org/abs/2606.15094",
      "pdf_url": "https://arxiv.org/pdf/2606.15094.pdf"
    },
    {
      "id": "2606.17082",
      "anchor": "paper-2606-17082",
      "title": "ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking",
      "authors": [
        "Hauteng Wu",
        "Xu Li",
        "Dong Kong",
        "Zihang Wang",
        "Xieyuanli Chen",
        "Benwu Wang",
        "Wenkai Zhu"
      ],
      "abstract": "End-to-end autonomous parking has emerged as a critical task within the realm of autonomous driving. However, existing methods suffer from black-box characteristics, lacking high-level semantic understanding and interpretability, which impedes the realization of seamless long-distance autonomous parking from the road to the target spot. To address these limitations, we propose ParkingTransformer, a novel framework that leverages multi-view perception and the scene understanding capability of Large Language Models (LLMs). By combining trajectory queries with LLMs implicit state features, our method interacts directly with historical information and raw sensor data to output planning trajectories, eliminating the need for dense Bird's-View (BEV) representations. To compensate for the inadequate spatial reasoning ability of LLMs, we introduce 3D positional encoding to explicitly inject spatial geometric awareness. Furthermore, a fixed-window streaming mechanism is designed for historical information processing, significantly improving long-term temporal processing efficiency and inference speed. Additionally, a coarse-to-fine decoding strategy is employed to progressively enhance trajectory precision. Extensive closed-loop experiments are conducted on the CARLA simulator and real-world vehicle platforms. The results demonstrate that our method achieves a driving score of 61.32 in CARLA simulator and an average success rate of 88.70% in real-world experiments, validating the feasibility and effectiveness of the proposed algorithms.",
      "short_abstract": "End-to-end autonomous parking has emerged as a critical task within the realm of autonomous driving. However, existing methods suffer from black-box characteristics, lacking high-level semantic understanding and interpretability, which impedes the realization...",
      "published": "2026-06-12",
      "updated": "2026-06-12",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 15,
        "evidence": [
          "title: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 68,
      "arxiv_url": "https://arxiv.org/abs/2606.17082",
      "pdf_url": "https://arxiv.org/pdf/2606.17082.pdf"
    },
    {
      "id": "2606.14992",
      "anchor": "paper-2606-14992",
      "title": "KATANA: A Fast, Low-Power Mapping of Kalman Filters onto Edge NPUs for Real-Time Tracking",
      "authors": [
        "Bodhisatwa Kundu",
        "Anish Rooj",
        "Sumit Saha",
        "Abhradeep Sarkar",
        "Arghadip Das",
        "Arnab Raha",
        "Mrinal K. Naskar"
      ],
      "abstract": "State estimation is the closed-loop core of every real-time tracking system, from radar surveillance and counter-UAV defense to autonomous driving and robotics. These deployments run on edge platforms, where defense systems mount on vehicles and drones, and civilian pipelines live on cars and handheld devices. Here, every additional watt of compute erodes mission duration or operational range. Two hard constraints follow: each new measurement must be fused before the next control cycle, and the total compute must fit within a strict battery and thermal power envelope. The Linear and Extended Kalman Filters (LKF, EKF) are dominant estimators on these systems, but today they execute almost exclusively on CPUs, which serialize multi-object tracking (MOT) updates, or on custom FPGA/ASIC accelerators that lengthen design cycles. Contemporary AI-PC SoCs, like the Intel Core Ultra Series 1 and 2, integrate a low-power, data-parallel Neural Processing Unit (NPU). We therefore ask whether the Kalman filter can be mapped onto this existing matrix engine to meet real-time and low-power budgets simultaneously, avoiding a dedicated accelerator and keeping the CPU and GPU free for primary workloads. We present KATANA, an NPU-aware optimization framework delivering the first end-to-end mapping of the LKF and EKF onto a commercial NPU, alongside a cross-platform characterization on shipping AI-PC silicon. KATANA applies three algebraic graph rewrites: subtract-to-add reformulation via a precomputed negative-projection matrix H_neg, static-shape tensor fusion, and block-diagonal batched parallelization, ensuring 100% of operations execute on the DPU matrix engine. On the Series 2, the optimized batched EKF reaches 223.35 FPS at 13.43 W active power, and the LKF reaches 408.73 FPS at 14.05 W, delivering up to a 97.9% reduction in dynamic energy versus the CPU implementation.",
      "short_abstract": "State estimation is the closed-loop core of every real-time tracking system, from radar surveillance and counter-UAV defense to autonomous driving and robotics. These deployments run on edge platforms, where defense systems mount on vehicles and drones, and...",
      "published": "2026-06-12",
      "updated": "2026-06-12",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 1,
        "evidence": [
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 68,
      "arxiv_url": "https://arxiv.org/abs/2606.14992",
      "pdf_url": "https://arxiv.org/pdf/2606.14992.pdf"
    },
    {
      "id": "2606.14956",
      "anchor": "paper-2606-14956",
      "title": "A Comparative Study of Graph Neural Network Layer Selection for Interaction Modelling in Driving Trajectory Prediction",
      "authors": [
        "George Daoud",
        "Mohamed El-Darieby"
      ],
      "abstract": "Autonomous driving systems rely on precise trajectory prediction to plan safe and efficient movement. Graph Neural Networks (GNNs) have become a promising approach for modelling spatiotemporal interactions among road agents. However, designing GNN architectures for trajectory prediction remains non-standardized, with little guidance on which graph layers effectively capture spatial interactions and temporal dynamics. This paper offers a detailed comparative study of 19 graph layer types, focusing on their spatial and temporal processing capabilities to discover the most effective architectures for trajectory prediction. Within the explored hyperparameter setting, we highlight five standout layer combinations, with ARMA, Chebyshev, and topology-aware layers consistently performing better than others. Beyond performance metrics, our findings yield practical design principles: sum-based aggregation is more effective than mean-based methods, multi-head attention mechanisms enable richer interactions, and assigning different weights to different hop distances significantly improves prediction accuracy. These findings offer useful guidance for designing more interpretable and effective trajectory prediction models.",
      "short_abstract": "Autonomous driving systems rely on precise trajectory prediction to plan safe and efficient movement. Graph Neural Networks (GNNs) have become a promising approach for modelling spatiotemporal interactions among road agents. However, designing GNN...",
      "published": "2026-06-12",
      "updated": "2026-06-12",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 20,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 68,
      "arxiv_url": "https://arxiv.org/abs/2606.14956",
      "pdf_url": "https://arxiv.org/pdf/2606.14956.pdf"
    },
    {
      "id": "2606.14438",
      "anchor": "paper-2606-14438",
      "title": "Physics-Grounded Causal Auditing of End-to-End Driving Planners",
      "authors": [
        "Zikun Guo",
        "Minglan Chen",
        "Jinyou Zhai",
        "Rongjin Zou"
      ],
      "abstract": "End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that merely co-occur with expert actions (a roadside object, a building facade) with driving decisions, rather than the variables that causally determine them. Such causal confusion silently compromises reliability in long-tail scenarios, and it is difficult to detect, because prevailing open-loop metrics (L2 displacement and collision rate) are dominated by ego status and do not indicate whether a planner depends on spurious cues. Existing remedies based on causal-intervention training require retraining large models and cannot audit a planner that is already deployed. We present CADET, a training-free framework that audits, benchmarks, and repairs spurious reliance in pretrained E2E planners without any parameter update.",
      "short_abstract": "End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that merely co-occur with expert actions (a roadside object, a building facade) with driving decisions, rather than the...",
      "published": "2026-06-12",
      "updated": "2026-06-12",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 26,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 68,
      "arxiv_url": "https://arxiv.org/abs/2606.14438",
      "pdf_url": "https://arxiv.org/pdf/2606.14438.pdf"
    },
    {
      "id": "2606.14658",
      "anchor": "paper-2606-14658",
      "title": "Giving AI a Headache: Acoustic Adversarial Attacks to Computer Vision Applications",
      "authors": [
        "Nicole Villavicencio-Garduño",
        "Maksim Ekin Eren",
        "Milo Prisbrey",
        "Ben Migliori",
        "Michael Teti"
      ],
      "abstract": "Artificial Intelligence (AI) is increasingly used to automate a variety of real-world computer vision (CV) applications, such as autonomous vehicle control, facial recognition, and security cameras. Recent research has shown that acoustic vibration can induce real physical motion in cameras, interfering with their internal stabilization mechanisms. Because the motion falls outside the conditions the stabilization system was designed to handle, the system introduces artifacts into the frame, causing AI-based CV models to misclassify, miss targets, or hallucinate objects. Previous work used ultrasonic frequencies (>20 kHz) to perform short-range attacks, which limits them to short distances due to the attenuation exhibited by high frequencies. In this work, we investigate acoustic attacks using lower frequencies in the audible range (<20 kHz), and we further expand our analysis to include how various image and object features are affected by the attacks. Specifically, we performed physical experiments to demonstrate the viability of our attacks on an off-the-shelf object detection model (YOLO11) by resonating a commercially available camera with various frequencies. Based on our results, we provide insights into several factors that make an AI CV system more vulnerable to these attacks, which could help inform the development of future mitigation strategies.",
      "short_abstract": "Artificial Intelligence (AI) is increasingly used to automate a variety of real-world computer vision (CV) applications, such as autonomous vehicle control, facial recognition, and security cameras. Recent research has shown that acoustic vibration can induce...",
      "published": "2026-06-12",
      "updated": "2026-06-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 15,
        "evidence": [
          "abstract: security",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 68,
      "arxiv_url": "https://arxiv.org/abs/2606.14658",
      "pdf_url": "https://arxiv.org/pdf/2606.14658.pdf"
    },
    {
      "id": "2606.14609",
      "anchor": "paper-2606-14609",
      "title": "Safe Reinforcement Learning of Autonomous Highway Driving: A Unified Framework for Safety and Efficiency",
      "authors": [
        "Chufei Yan",
        "Zhihao Cui",
        "Yiyan Lv",
        "Taojie Chen",
        "Ning Bian",
        "Yulei Wang"
      ],
      "abstract": "Deep reinforcement learning (DRL) offers a compelling route to decision-making for advanced autonomous vehicles (AVs), yet its trial-and-error nature makes it difficult to guarantee safety during training and to achieve both safety and efficiency at deployment. We propose a unified safe reinforcement learning (SRL) framework that integrates safe distance (SD), reward machines (RM), and mixture-of-experts (MoE), termed MoE-RM-SRL. For deployment, SD and RM jointly shape a rule-aware reward that encodes highway traffic regulations and stage-wise objectives, enabling safe and reliable behavior without sacrificing efficiency. For training, we introduce a sparsely gated MoE layer comprising up to 11 deep Q-networks (DQNs); an SD-based gating rule activates a minimal set of experts for lane-keeping and lane-changing, mitigating the instability, discontinuities, and impulsive transients commonly induced by switching between heterogeneous controllers (e.g., MPC/rule-based modules and learned policies). We implement the proposed architecture in CARLA and integrate it with a 6-DoF driver-in-the-loop virtual-reality (DiL-VR) platform. Experiments in stochastic two-lane traffic show that MoE-RM-SRL substantially improves safety and efficiency over state-of-the-art baselines, and the framework naturally extends to multi-lane driving as well as on-ramp merging and exiting scenarios.",
      "short_abstract": "Deep reinforcement learning (DRL) offers a compelling route to decision-making for advanced autonomous vehicles (AVs), yet its trial-and-error nature makes it difficult to guarantee safety during training and to achieve both safety and efficiency at...",
      "published": "2026-06-12",
      "updated": "2026-06-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 68,
      "arxiv_url": "https://arxiv.org/abs/2606.14609",
      "pdf_url": "https://arxiv.org/pdf/2606.14609.pdf"
    },
    {
      "id": "2606.14216",
      "anchor": "paper-2606-14216",
      "title": "Short-Horizon Position Accuracy of Single-Track Models: Implications for Motion Planning of Autonomous Vehicles",
      "authors": [
        "Aron J. Aertssen",
        "Lars A. T. H. van Alen",
        "Igo J. M. Besselink",
        "Rudolf G. M. Huisman",
        "René M. J. G. van de Molengraft"
      ],
      "abstract": "Accurate and computationally efficient vehicle models are essential for motion planning of autonomous vehicles, where positional accuracy directly affects trajectory feasibility and safety. However, the positional accuracy has not been systematically evaluated against real measurements. Therefore, this paper compares the short-horizon positional accuracy of three single-track vehicle models against vehicle measurements across various driving maneuvers. Model parameters are identified through dedicated experiments with the instrumented test vehicle. Rather than identifying a single best model, this work aims to provide insight into the trade-offs between model complexity, parameterization quality, and positional accuracy for informed model selection in Model Predictive Control applications.",
      "short_abstract": "Accurate and computationally efficient vehicle models are essential for motion planning of autonomous vehicles, where positional accuracy directly affects trajectory feasibility and safety. However, the positional accuracy has not been systematically...",
      "published": "2026-06-12",
      "updated": "2026-06-12",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 25,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 68,
      "arxiv_url": "https://arxiv.org/abs/2606.14216",
      "pdf_url": "https://arxiv.org/pdf/2606.14216.pdf"
    },
    {
      "id": "2606.17080",
      "anchor": "paper-2606-17080",
      "title": "HRDX: A Large-Scale Vector HD-Map Dataset",
      "authors": [
        "Sahith Reddy Chada",
        "Isht Dwivedi",
        "Nirav Savaliya"
      ],
      "abstract": "Reliable autonomous driving requires vectorized HD maps that are geometrically accurate, semantically rich, and scalable to long-horizon driving. However, existing public HD map datasets are limited in scale, provide sparse semantic attributes, and lack modalities such as aerial imagery that could enable new research directions. We present HRDX, a large-scale dataset for vector HD-map construction, spanning about 40 hours (1,400 km) of minimally overlapping drives, which is several times larger than prior public HD map datasets. Data is captured using six synchronized surround cameras, a 128-beam LiDAR, and centimeter-level RTK GNSS/IMU, and is further complemented by precisely aligned aerial orthoimagery. Annotations cover 10 vector map classes, complemented with over 20 semantic and topological attributes. To evaluate this richer ontology, we introduce the Composite Score (CS) to jointly assess geometric fidelity and attribute correctness. Benchmark experiments show that HRDX's scale improves online vector-map construction, and that aligned aerial imagery provides a useful structural prior: using aerial imagery at training and/or inference improves geometric map quality, while aerial-augmented teachers can transfer part of this benefit to camera-only students without increasing inference-time sensor requirements. HRDX is intended to support reproducible research on large-scale HD-map learning, multimodal BEV fusion, and training-time privileged information. HRDX dataset and benchmarks are available at https://github.com/honda-research-institute/HRDX",
      "short_abstract": "Reliable autonomous driving requires vectorized HD maps that are geometrically accurate, semantically rich, and scalable to long-horizon driving. However, existing public HD map datasets are limited in scale, provide sparse semantic attributes, and lack...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 4,
        "evidence": [
          "abstract: HD map",
          "abstract: mapping",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.17080",
      "pdf_url": "https://arxiv.org/pdf/2606.17080.pdf"
    },
    {
      "id": "2606.14067",
      "anchor": "paper-2606-14067",
      "title": "Vision-Based Efficient Joint Trajectory and Channel Tracking in Near-Field XL-MIMO Systems",
      "authors": [
        "Mengyuan Li",
        "Yu Han",
        "Hao Xu",
        "Yongxu Zhu",
        "Shi Jin",
        "Chao-Kai Wen"
      ],
      "abstract": "Accurate joint tracking of mobile users, surrounding scatterers, and dynamic channels is a critical task for sixth-generation (6G) wireless systems, essential for both ensuring high-quality communications and empowering advanced selsing applications such as autonomous driving and immersive extended reality. While extremely large-scale multiple-input multiple-output (XL-MIMO) inherently offers strong support for this task through its high spatial resolution and spectral efficiency, its massive scale of antenna arrays, coupled with near-field propagation characteristics, makes joint trajectory and channel tracking time-consuming and hardware-intensive. To address these challenges, we rethink the problem from a vision-based signal perspective. Specifically, we design a subarray-based partially connected hybrid beamforming (PC-HBF) architecture with a tailored time-multiplexed (TM) mechanism. This effectively compensates for the aperture loss caused by limited radio frequency (RF) chains, generating high-fidelity Cartesian-domain signal images that inherently capture near-field spatial features. Based on this visual representation, we propose an improved CenterNet to perform accurate one-shot path localization, circumventing the path-iterative search required by conventional compressed-sensing-based methods. Building upon this to further improve the accuracy and exploit temporal correlation, a local small-scale orthogonal matching pursuit (OMP) refiner and a lightweight cascaded OMP tracker are developed. Finally, a Hungarian-based trajectory association module is incorporated to maintain track continuity and provide trajectory-level information for environment monitoring. Simulation results show that the proposed framework consistently outperforms representative baselines in position and channel tracking accuracy, especially under low-SNR and limited-hardware conditions.",
      "short_abstract": "Accurate joint tracking of mobile users, surrounding scatterers, and dynamic channels is a critical task for sixth-generation (6G) wireless systems, essential for both ensuring high-quality communications and empowering advanced selsing applications such as...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Mapping & Localization: abstract: localization",
          "Systems, Deployment & Connectivity: abstract: hardware"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.14067",
      "pdf_url": "https://arxiv.org/pdf/2606.14067.pdf"
    },
    {
      "id": "2606.14058",
      "anchor": "paper-2606-14058",
      "title": "ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving",
      "authors": [
        "Zhiyuan Zhang",
        "Yanlun Peng",
        "Jianing Zhang",
        "Xianda Guo",
        "Zehan Huang",
        "Haoran Liu",
        "Qifeng Li",
        "Shaofeng Zhang",
        "Xiaosong Jia",
        "Junchi Yan"
      ],
      "abstract": "Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can respond feasibly to autonomous vehicle (AV) behaviors that differ from the log. However, existing behavior simulation benchmarks do not directly measure reactive capability. They often let the simulator jointly control the AV and surrounding agents and evaluate realism through log similarity or open-loop prediction metrics. In this work, we introduce ReactSim-Bench for evaluating the reactive capability of behavior world model simulation in autonomous driving. We decouple the control of agents and the AV, using AV behaviors that differ from the log and require agents to respond as independent AV inputs. To obtain these AV behaviors, we construct a pipeline that uses an AV planner model to generate candidate behaviors and filters the data using rules and manual verification. Collision metrics, map-based metrics, and kinematic feasibility metrics are used to evaluate the safety and rule compliance of reactive responses. We construct 2,636 test scenarios with three categories and conduct a systematic evaluation of state-of-the-art models across multiple architectures, including Transformer-based, diffusion-based, and next-token-prediction-based models. We further analyze how replan frequency affects performance and provide insights for future studies.",
      "short_abstract": "Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can respond feasibly to autonomous vehicle (AV) behaviors that differ from the log....",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 9,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.14058",
      "pdf_url": "https://arxiv.org/pdf/2606.14058.pdf"
    },
    {
      "id": "2606.14032",
      "anchor": "paper-2606-14032",
      "title": "From Attacks to Curricula: Learnability-Guided Adversarial Training for Safe Autonomous Driving",
      "authors": [
        "Yuewen Mei",
        "Tong Nie",
        "Jie Sun",
        "Haotian Shi",
        "Wei Ma",
        "Jian Sun"
      ],
      "abstract": "Closed-loop adversarial training improves autonomous driving safety by exposing policies to rare safety-critical scenarios. Standard pipelines first generate adversarial scenarios and then sample them for policy optimization. However, most existing frameworks remain attack-oriented: collision-driven generators often synthesize unsolvable extreme situations, which can degrade learning, while heuristic samplers ignore the evolving capability of the driving policy, causing sample inefficiency and delayed convergence. We propose AlignADV, a learnability-guided closed-loop adversarial training framework that converts adversarial scenarios into resolvable and capability-aligned curricula. First, we reformulate adversarial scenario generation as a preference alignment problem and employ direct preference optimization to guide the generator toward critical yet resolvable scenarios. Second, we introduce behavioral fingerprints to capture the intrinsic characteristics of the evolving policy and construct a multi-modal capability prediction model that estimates policy performance without expensive closed-loop simulations. By combining resolvability-aligned scenarios with capability predictions, AlignADV develops a dynamic curriculum sampling mechanism that prioritizes scenarios targeting the current policy's vulnerabilities. Experiments on the Waymo Open Motion Dataset demonstrate that AlignADV improves convergence efficiency and final performance, reducing training steps by up to 40.6 percent compared with baseline methods while lowering collision rate and improving route completion under both normal and adversarial traffic conditions. These results highlight a shift from attack-oriented scenario generation to learnability-guided policy improvement, offering a principled direction for safer and more efficient autonomous driving training. Project page: https://meiyuewen.github.io/AlignADV/.",
      "short_abstract": "Closed-loop adversarial training improves autonomous driving safety by exposing policies to rare safety-critical scenarios. Standard pipelines first generate adversarial scenarios and then sample them for policy optimization. However, most existing frameworks...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 16,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.14032",
      "pdf_url": "https://arxiv.org/pdf/2606.14032.pdf"
    },
    {
      "id": "2606.14010",
      "anchor": "paper-2606-14010",
      "title": "RT-VLA: Real-Time Vision-Language-Action Models via Knowledge Distillation",
      "authors": [
        "Xiangyu Huang",
        "Zhenlin Hua",
        "Han Zhou",
        "Shounak Sural",
        "Ragunathan Rajkumar"
      ],
      "abstract": "Vision-Language-Action (VLA) models have shown strong potential for end-to-end autonomous driving by jointly modeling visual perception, language reasoning, explainability and action prediction. However, their large vision-language backbones and reasoning modules introduce substantial inference latency and thereby prevent their deployment in the unforgiving reality of the road networks. We propose RT-VLA, a lightweight, distilled VLA model that transfers the driving and reasoning capabilities of the state-of-the-art SimLingo model into a compact student through multi-level supervised distillation. RT-VLA preserves language-based reasoning and supports post-hoc explanation through offline language analysis of safety-critical driving moments without adding latency to real-time control. Compared to the SimLingo teacher, RT-VLA maintains competitive closed-loop driving and language reasoning performance while reducing inference time by 44.8X in vision-only mode and 7.9X in vision+language mode. These results suggest that supervised distillation is a practical approach for building real-time, explainable VLA-style autonomous driving models.",
      "short_abstract": "Vision-Language-Action (VLA) models have shown strong potential for end-to-end autonomous driving by jointly modeling visual perception, language reasoning, explainability and action prediction. However, their large vision-language backbones and reasoning...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Hardware / Real-Time",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 17,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.14010",
      "pdf_url": "https://arxiv.org/pdf/2606.14010.pdf"
    },
    {
      "id": "2606.13951",
      "anchor": "paper-2606-13951",
      "title": "Accuracy of Joint Time-Based and Carrier-Phase Positioning in 5G Networks under Correlated Measurement Errors",
      "authors": [
        "Nahidul Islam",
        "Mohammad Razzaghpour",
        "Marwan Hammouda",
        "Carsten Bockelmann",
        "Armin Dekorsy"
      ],
      "abstract": "High-accuracy positioning is critical for emerging applications such as autonomous driving, industrial automation, augmented reality, and smart cities. 3GPP Release 18 introduced carrier-phase (CP) positioning for 5G that offers superior accuracy compared to conventional time-based methods such as time of arrival (ToA). However, CP-based positioning requires resolving the integer phase ambiguity, which refers to the unknown number of full-wavelength cycles completed during signal propagation. Joint processing of ToA and CP can mitigate this integer ambiguity by narrowing down the search space of possible integers, particularly for short wavelengths. This paper investigates the performance of a positioning method that integrates ToA and CP measurements. As a main contribution, the analysis explicitly accounts for the error correlation between ToA and CP measurements. Furthermore, the study analyzes the impact of key 5G system parameters on positioning accuracy using this correlation-aware joint method in both factory and urban environments, where many 5G positioning applications are expected to emerge. The results highlight that exploiting this correlation can further improve positioning performance by approximately 7 percent. Moreover, the findings of this study provide insight into how 5G system parameters can be tuned to achieve centimeter-level accuracy under favorable conditions.",
      "short_abstract": "High-accuracy positioning is critical for emerging applications such as autonomous driving, industrial automation, augmented reality, and smart cities. 3GPP Release 18 introduced carrier-phase (CP) positioning for 5G that offers superior accuracy compared to...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 6,
        "evidence": [
          "title: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.13951",
      "pdf_url": "https://arxiv.org/pdf/2606.13951.pdf"
    },
    {
      "id": "2606.13840",
      "anchor": "paper-2606-13840",
      "title": "Multi-Agent Embodied Autonomous Driving (MAEAD): From V2X Information Exchange to Shared World Models",
      "authors": [
        "Senkang Hu",
        "Zhengru Fang",
        "Yihang Tao",
        "Zihan Fang",
        "Yiqin Deng",
        "Yuguang Fang"
      ],
      "abstract": "Autonomous driving is shifting from isolated vehicle intelligence toward multi-agent embodied systems that share perception, infer intent, and coordinate action under uncertainty. This survey examines this transition through the lens of Shared World Models (SWMs): predictive cross-agent representations maintained across vehicles, infrastructure, and other traffic participants. We review approximately 400 publications covering vehicle-to-everything (V2X) communication, collaborative perception, inter-agent cognition, cooperative planning, end-to-end cooperative driving, and simulation and data engines for closed-loop validation. The organizing question is how exchanged observations become aligned state, intent-aware interaction, and coordinated downstream action. Across the surveyed literature, evaluation remains concentrated in simulation, curated benchmarks, and offline protocols. Foundation-model-based coordination also lacks verifiable real-time safety guarantees in open traffic. These gaps motivate key research priorities for multi-agent embodied autonomous driving (MAEAD): verifiable shared-state maintenance, robust intent and plan alignment, and safe coordinated action under communication and computing constraints in real-world deployment. We maintain an open-source project to continuously track the latest developments at https://github.com/dl-m9/Multi-Agent-Embodied-Autonomous-Driving.",
      "short_abstract": "Autonomous driving is shifting from isolated vehicle intelligence toward multi-agent embodied systems that share perception, infer intent, and coordinate action under uncertainty. This survey examines this transition through the lens of Shared World Models...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "World Model",
        "Survey"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 16,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.13840",
      "pdf_url": "https://arxiv.org/pdf/2606.13840.pdf"
    },
    {
      "id": "2606.13460",
      "anchor": "paper-2606-13460",
      "title": "VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models",
      "authors": [
        "Ruiqi Xian",
        "Yuehan Xian",
        "Jing Liang",
        "Xuewei Qi",
        "Dinesh Manocha"
      ],
      "abstract": "Semantic 3D occupancy provides a voxelized world state for autonomous driving and robot decision making, but object and rare-class errors can affect free-space interpretation, collision checking, and temporal state propagation. We show that a common VLM strategy, aligning 3D voxel or object features with crop-caption embeddings, improves text-space similarity without reliably improving closed-set occupancy mIoU. Motivated by this mismatch, we propose VISA, a training-time semantic auditing approach for existing occupancy world models. VISA queries an offline VLM on a representative crop of each physical object instance, obtains a structured audit with class hypotheses, plausible confusions, reliability, attributes, and evidence, and propagates it along the object track. The audit is grounded to matched 3D object voxels and distilled into semantic logits through reliability-weighted taxonomy, attribute-factor, and scene-level audit graph losses, while inference remains unchanged and requires no VLM. On nuScenes, averaged across three runs, VISA improves OccWorld from 19.06 to 20.05 mIoU and GaussianWorld from 21.36 to 21.91 mIoU; on GaussianWorld, object mIoU improves from 18.18 to 19.16 and rare-class mIoU from 15.60 to 16.79. These results suggest that VLMs are better suited to closed-set occupancy as reliability-aware semantic auditors than as generic caption-embedding targets.",
      "short_abstract": "Semantic 3D occupancy provides a voxelized world state for autonomous driving and robot decision making, but object and rare-class errors can affect free-space interpretation, collision checking, and temporal state propagation. We show that a common VLM...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.13460",
      "pdf_url": "https://arxiv.org/pdf/2606.13460.pdf"
    },
    {
      "id": "2606.13899",
      "anchor": "paper-2606-13899",
      "title": "ARTSN: Exact and Adaptive Self-triggered Traffic Scheduling for ARTS Networks",
      "authors": [
        "Ruide Cao",
        "Shuangping Zhan",
        "Jiashuo Lin",
        "Yan Liu",
        "Chenxi Ling",
        "Yi Wang",
        "Guoming Tang"
      ],
      "abstract": "Autonomous real-time systems (ARTS), such as self-driving vehicles and robotic assembly lines, are increasingly deployed to improve efficiency, accuracy, and responsiveness with reduced human intervention. In ARTS networks, self-triggered (ST) traffic-initiated by internal decision-making rather than fixed schedules or external events-is becoming prevalent and plays a critical role in enabling timely autonomous actions. However, existing network schedulers do not adequately support ST traffic due to two inherent challenges: volatility, where bounded processing jitter leads to uncertain arrival times, and absence, where reserved network resources remain underutilized when ST traffic does not materialize. To address these challenges, we propose ARTSN, an ST-tailored scheduling paradigm built upon time-sensitive networking (TSN). ARTSN introduces two key techniques: (1) an exact offline scheduling method that leverages the inferable arrival information of ST traffic for precise time-slot reservation, and (2) an adaptive online slot-release mechanism that dynamically reclaims unused reservations when ST traffic is absent. Extensive experiments on both a TSN simulator and a real-world testbed show that ARTSN significantly improves schedulability, scalability, and efficiency over state-of-the-art methods while maintaining reliable transmission guarantees.",
      "short_abstract": "Autonomous real-time systems (ARTS), such as self-driving vehicles and robotic assembly lines, are increasingly deployed to improve efficiency, accuracy, and responsiveness with reduced human intervention. In ARTS networks, self-triggered (ST)...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 15,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.13899",
      "pdf_url": "https://arxiv.org/pdf/2606.13899.pdf"
    },
    {
      "id": "2606.13891",
      "anchor": "paper-2606-13891",
      "title": "TetraRL: A Self-Adaptive Runtime for On-Device Deep Reinforcement Learning Systems",
      "authors": [
        "Zexin Li",
        "Soheil Shirvani",
        "Cong Liu"
      ],
      "abstract": "Autonomous robotic systems, including autonomous vehicles, drones, and mobile robots, increasingly rely on on-device Deep Reinforcement Learning (DRL) to adapt to dynamic environments. Unlike cloud-based solutions, embedded DRL must perform training and inference directly on resource-constrained hardware while maintaining timely decision-making. This creates a fundamental challenge: balancing four tightly coupled objectives, real-time performance, task reward, memory utilization, and energy consumption. Optimizing these objectives independently often leads to suboptimal behavior, while conventional multi-objective methods may violate resource constraints and compromise reliability. This paper presents TetraRL, a self-adaptive runtime framework for tetra-objective on-device DRL. TetraRL formulates embedded DRL as a unified optimization problem over real-time, reward, RAM, and reserve (energy) objectives, and employs a preference-conditioned reinforcement learning controller to dynamically navigate the resulting trade-off space. The framework integrates a unified resource-management abstraction, hardware-aware DVFS control, and a runtime Override Layer for robust constraint enforcement. We implement TetraRL on NVIDIA Jetson AGX Orin and Orin Nano platforms and evaluate it across diverse DRL environments. Results show that TetraRL effectively balances all four objectives, achieves competitive trade-offs under varying runtime preferences, and incurs negligible overhead. Moreover, a single trained policy can support runtime-switchable optimization goals, providing a practical foundation for resource-aware and self-adaptive on-device DRL.",
      "short_abstract": "Autonomous robotic systems, including autonomous vehicles, drones, and mobile robots, increasingly rely on on-device Deep Reinforcement Learning (DRL) to adapt to dynamic environments. Unlike cloud-based solutions, embedded DRL must perform training and...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Reinforcement Learning",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.13891",
      "pdf_url": "https://arxiv.org/pdf/2606.13891.pdf"
    },
    {
      "id": "2606.13282",
      "anchor": "paper-2606-13282",
      "title": "ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space",
      "authors": [
        "Pratyush Chaudhari"
      ],
      "abstract": "As AI systems are deployed in high-stakes ethical contexts such as healthcare triage, autonomous vehicle control, and employment screening, formal methods for evaluating their robustness against adversarial manipulation of ethical reasoning remain underdeveloped. This paper introduces the Ethical Robustness Testing System (ERTS), a closed-pipeline framework that: (1) encodes ethical dilemmas into a 22-dimensional Ethical Consequence Space (ECS) grounded in established ethical theory; (2) applies 17 semantic perturbation functions subject to 6 validity constraint classes including a novel semantic coherence constraint; (3) measures decision deviation via a 4-component Ethical Instability Index (EII); and (4) produces domain-adaptive pre-deployment robustness assessment verdicts. We evaluate 4 structured baseline models and 2 production LLMs (Gemini 2.0 Flash and Llama 3.2) across 50 ethical scenarios spanning 8 deployment domains, generating 1,500 adversarial test cases. Results demonstrate that only 33% of models achieve assessment clearance, with the local Llama-3.2 model proving particularly vulnerable to fairness corruption and information degradation attacks (ERS = 0.737). To the best of our knowledge, no existing framework combines a bounded ethical consequence space, semantic coherence constraints, and domain-adaptive assessment in a single adversarial testing pipeline.",
      "short_abstract": "As AI systems are deployed in high-stakes ethical contexts such as healthcare triage, autonomous vehicle control, and employment screening, formal methods for evaluating their robustness against adversarial manipulation of ethical reasoning remain...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 20,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.13282",
      "pdf_url": "https://arxiv.org/pdf/2606.13282.pdf"
    },
    {
      "id": "2606.12987",
      "anchor": "paper-2606-12987",
      "title": "Diffusion Transformer World-Action Model for AV Scene Prediction",
      "authors": [
        "Ruslan Sharifullin",
        "Benjamin Jiang",
        "Kai Xi Chew"
      ],
      "abstract": "Action-conditioned world models let an autonomous vehicle predict future camera scenes from its own planned controls, enabling planning and simulation without real-world rollouts, but at compact, trainable scale the futures are ambiguous and the field's standard distortion metrics actively mislead: they reward a blurry regression mean over a realistic prediction. We confront this with a compact latent world model that, given the present front-camera latent and a sequence of ego-actions, predicts future scene latents a frozen decoder renders to $256 \\times 256$ frames up to 8 seconds ahead, evaluated on 150 held-out nuScenes scenes. We first benchmark where to predict: across six frozen encoders spanning four representation families, V-JEPA2 with temporal context reduces steering RMSE by 40% over the best single-frame encoder. We then train a latent Diffusion Transformer (DiT) and, through a controlled diagnosis, identify the four ingredients it needs: spatial tokens, the $x_0$ objective, residual anchoring, and sampling matched to target uncertainty. In a Stable-Diffusion-VAE encode-predict-decode pipeline we expose the central tension: distortion metrics (cosine similarity, SSIM) favor the blurry mean, masking that the diffusion model is far closer to the real frame distribution. Inception-based FID and KID reveal a clean perception-distortion frontier: diffusion attains KID 0.078 versus 0.375 for regression ($4.8\\times$ better), and a deployable train-derived calibration makes this practical without test-time ground truth. The model is genuinely action-controllable (steering drives scene displacement, Spearman $ρ= 0.81$, vs $-0.18$ for regression). We trace limited single-pass motion to a shared-present anchor and engineer a compact 1.7M-parameter \"jump\" model that recovers full ground-truth motion magnitude ($1.02\\times$ GT), where single-pass models capture less than half.",
      "short_abstract": "Action-conditioned world models let an autonomous vehicle predict future camera scenes from its own planned controls, enabling planning and simulation without real-world rollouts, but at compact, trainable scale the futures are ambiguous and the field's...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 16,
        "evidence": [
          "abstract: world model",
          "title: scene prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.12987",
      "pdf_url": "https://arxiv.org/pdf/2606.12987.pdf"
    },
    {
      "id": "2606.13292",
      "anchor": "paper-2606-13292",
      "title": "Feasibility Assessment of Remote Driving via Latency Analysis of ITS-G5 and Cellular Networks in the MASA Living Lab",
      "authors": [
        "Gaetano Orazio Cauchi",
        "Antonio Solida",
        "Salvatore Iandolo",
        "Marco Savarese",
        "Martin Klapez",
        "Enrico Rossini",
        "Marcello Pietri",
        "Marco Picone",
        "Marco Mamei",
        "Maurizio Casoni",
        "Carlo Augusto Grazia"
      ],
      "abstract": "Remote driving has gained increasing attention as a key enabler for connected and automated vehicles. Yet its practical deployment hinges on wireless networks' ability to guarantee low, predictable latency. In this paper, we present an extensive latency analysis of ITS-G5 and cellular (5G) technologies within the Modena Automotive Smart Area (MASA), a real-world, city-scale testbed equipped with a distributed intelligent transportation infrastructure. By conducting controlled experiments under varying network loads and traffic conditions, we measure network and end-to-end latency components relevant to remote driving, in which the uplink consists of a continuous video stream transmitted from the vehicle to the remote operator, and the downlink conveys control commands back to the car. Measurements conducted under diverse conditions reveal how latency and variability differ across the two technologies and how infrastructure coverage impacts video-stream transmission performance. Based on the observed latency distributions and reliability metrics, we assess the practical feasibility and safety margins of remote driving in mixed network environments. The results provide actionable insights for future teleoperation deployments and motivate hybrid communication strategies that combine the strengths of ITS-G5 and cellular networks.",
      "short_abstract": "Remote driving has gained increasing attention as a key enabler for connected and automated vehicles. Yet its practical deployment hinges on wireless networks' ability to guarantee low, predictable latency. In this paper, we present an extensive latency...",
      "published": "2026-06-11",
      "updated": "2026-06-11",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 24,
        "evidence": [
          "abstract: connected vehicles",
          "abstract: teleoperation",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 69,
      "arxiv_url": "https://arxiv.org/abs/2606.13292",
      "pdf_url": "https://arxiv.org/pdf/2606.13292.pdf"
    },
    {
      "id": "2606.12826",
      "anchor": "paper-2606-12826",
      "title": "DIMOS: Disentangling Instance-level Moving Object Segmentation",
      "authors": [
        "Hongxiang Huang",
        "Hongwei Ren",
        "Xiaopeng Lin",
        "Yulong Huang",
        "Zeke Xie",
        "Bojun Cheng"
      ],
      "abstract": "Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal resolution and dynamic range, which makes them highly sensitive to motion information. By fusing event and image features, motion cues from events can complement spatial details from images, enhancing the performance of MIS. However, current multimodal MIS methods still struggle to segment small moving instances, as event cameras often yield sparse features under limited resolution. Moreover, event features entangle appearance attributes with motion cues, which further restricts effective cross-modal fusion. To address these challenges, we first propose a dual-disentangling feature extraction framework that separates and extracts appearance and motion information within both image and event modalities, thereby improving feature density. Subsequently, a multi-granularity cross-modal alignment is introduced to align distributionally and semantically consistent features across modalities, enabling more effective fusion with rich spatial and temporal details. The experiment results demonstrate that our method achieves state-of-the-art performance in multimodal MIS, especially for small instances under challenging conditions such as fast motion and low-light settings.",
      "short_abstract": "Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal resolution and...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 6,
        "evidence": [
          "abstract: scene segmentation",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12826",
      "pdf_url": "https://arxiv.org/pdf/2606.12826.pdf"
    },
    {
      "id": "2606.12783",
      "anchor": "paper-2606-12783",
      "title": "A Tutorial on World Models and Physical AI",
      "authors": [
        "Il-Seok Oh"
      ],
      "abstract": "World modeling is emerging as a central principle for building intelligent systems capable of prediction, reasoning, and decision making. A central distinction can be drawn between explicit world models, which learn structured dynamics for rollout-based reasoning and planning, and implicit world models, which encode predictive structure within scalable learned representations. These complementary paradigms provide a foundation for physical AI in domains such as robotics and autonomous driving, enabling intelligence beyond reactive control under real-world constraints. Recent foundation models further suggest a pathway toward unified systems integrating perception, prediction, and action. Despite rapid progress, major challenges remain in hierarchical reasoning, long-horizon planning, and autonomous goal formation, which are critical for advancing toward artificial general intelligence. This tutorial presents a coherent framework in which diverse world modeling approaches are unified through shared predictive structure and differentiated by how such structure is represented and exploited.",
      "short_abstract": "World modeling is emerging as a central principle for building intelligent systems capable of prediction, reasoning, and decision making. A central distinction can be drawn between explicit world models, which learn structured dynamics for rollout-based...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model",
        "Survey"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 11,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12783",
      "pdf_url": "https://arxiv.org/pdf/2606.12783.pdf"
    },
    {
      "id": "2606.12774",
      "anchor": "paper-2606-12774",
      "title": "Agentic MPC for Semantic Control System Resynthesis",
      "authors": [
        "Yuya Miyaoka",
        "Masaki Inoue"
      ],
      "abstract": "While MPC effectively handles structured, diverse, and low-level specifications, it lacks the capability to dynamically incorporate high-level contextual information such as social norms, user intent, or natural language instructions. To address this limitation, this manuscript introduces an agentic MPC framework that enables context-aware, semantically adaptive control synthesis by integrating with large language model-based agents. The agent interprets heterogeneous inputs, including natural language messages, environmental observations, and external knowledge, to resynthesize the control specifications. The effectiveness of the framework is demonstrated in an autonomous driving scenario, where the system aligns with personal preferences or responds to social situations such as emergency vehicle yielding.",
      "short_abstract": "While MPC effectively handles structured, diverse, and low-level specifications, it lacks the capability to dynamically incorporate high-level contextual information such as social norms, user intent, or natural language instructions. To address this...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 28,
        "evidence": [
          "title: model predictive control",
          "abstract: model predictive control",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12774",
      "pdf_url": "https://arxiv.org/pdf/2606.12774.pdf"
    },
    {
      "id": "2606.12706",
      "anchor": "paper-2606-12706",
      "title": "VLADriveBench: Evaluating CoT-Action Relationship in VLA for Autonomous Driving",
      "authors": [
        "Thach Nguyen",
        "Danhua Guo",
        "Tom Lampo",
        "Fei Wu",
        "Burhan Yaman"
      ],
      "abstract": "Vision-language-action (VLA) models generate chain-of-thought (CoT) reasoning alongside driving trajectories, but existing benchmarks evaluate only trajectory quality and do not assess whether the CoT is relevant, consistent, or causally connected to the driving action. We introduce VLADriveBench, a framework that combines observational metrics (mentioning, hallucination, contradiction, action alignment) with a CoT intervention protocol to provide complementary views of the CoT-action relationship. Applying VLADriveBench to three models across two architectures, we find that the two analyses can diverge sharply: ORION scores highest on observational alignment yet its CoT is epiphenomenal, while Alpamayo v1.5 scores lower yet its CoT is strongly causal, with visual salience gating the extent of CoT influence.",
      "short_abstract": "Vision-language-action (VLA) models generate chain-of-thought (CoT) reasoning alongside driving trajectories, but existing benchmarks evaluate only trajectory quality and do not assess whether the CoT is relevant, consistent, or causally connected to the...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 24,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12706",
      "pdf_url": "https://arxiv.org/pdf/2606.12706.pdf"
    },
    {
      "id": "2606.12628",
      "anchor": "paper-2606-12628",
      "title": "Context-Aware Feature-Fusion for Co-occurring Object Detection in Autonomous Driving",
      "authors": [
        "Binay Kumar Singh",
        "Niels Da Vitoria Lobo"
      ],
      "abstract": "Object detection in autonomous driving requires precise localization and an inherent understanding of the relational context between co-occurring objects. In extremely complex heterogeneous environments rare classes, small-scale objects, and frequently appearing objects are difficult for standard object detection frameworks to handle. In this paper, we propose a novel framework called Context-Centric Feature Fusion (CCFF), which utilizes two attention-based modules, Local Context Fusion Module (LCFM) uses the RoI-to-RoI self-attention mechanism to resolve spatial interactions, mainly considering small and partially obscured objects, while Global Context Attention Module (GCAM) converts the co-occurrence of objects priors by pooling top-K RoI features into a global context attention token, avoiding the computational overhead of pixel-level global pooling. This fusion of local and object-centric global features yields contextualized embeddings that enhance classification results and co-occurring objects detection. Our method is evaluated on two datasets, Cityscapes and BDD100K which demonstrate significant improvement on relational consistency, achieving a Category-level Consistency Strategy (CCS) of 0.973 and 0.969, respectively. Furthermore, our approach produces substantial gains in small object detection (AP_S: 14.1%) and successfully recovers rare classes such as \"Train\" that are typically lost in large distributions. Our efficiency report shows that the framework processes images in real time with a 0.2 FPS overhead. The code is available at https://github.com/BinayKSingh/CCFF.",
      "short_abstract": "Object detection in autonomous driving requires precise localization and an inherent understanding of the relational context between co-occurring objects. In extremely complex heterogeneous environments rare classes, small-scale objects, and frequently...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 11,
        "evidence": [
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12628",
      "pdf_url": "https://arxiv.org/pdf/2606.12628.pdf"
    },
    {
      "id": "2606.12603",
      "anchor": "paper-2606-12603",
      "title": "From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation",
      "authors": [
        "Honglin He",
        "Zhizheng Liu",
        "Yukai Ma",
        "Bolei Zhou"
      ],
      "abstract": "Autonomous long-horizon sidewalk navigation is essential for micro-mobility applications such as robotic food delivery and assistive electronic wheelchairs. Unlike autonomous driving on the road, long-horizon sidewalk navigation requires precise maneuvering through unpredictable sidewalk terrains and pedestrians, with a lightweight perception stack as minimal as a single monocular RGB camera. While imitation learning (IL) from demonstrations offers a practical solution, the resulting autopilot policy often suffers from compounding errors, a lack of social compliance on sidewalks, and deficiencies in counterfactual reasoning to handle complex situations. To address these challenges, we introduce FlowPilot, a mapless navigation policy that achieves robust and efficient long-horizon navigation performance using only a monocular RGB camera. We first propose to use anchored flow matching as an action representation for policy pre-training on large-scale robot fleet data and to capture the diverse, complex, multimodal distribution of sidewalk navigation behaviors. To bridge the gap between imitation and alignment, we further design a human-in-the-loop preference learning scheme to tune the policy on a small amount of human intervention data. It strengthens the model's counterfactual reasoning and social compliance on sidewalks. We evaluate FlowPilot through extensive simulation and real-world experiments in diverse sidewalk environments. FlowPilot achieves 42% success rate and 66% route completion in simulation, while FlowPilot-HP further improves real-world robustness and social compliance, reducing IR by 40.0% and NIR by 52.1% relative to the base model.",
      "short_abstract": "Autonomous long-horizon sidewalk navigation is essential for micro-mobility applications such as robotic food delivery and assistive electronic wheelchairs. Unlike autonomous driving on the road, long-horizon sidewalk navigation requires precise maneuvering...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Imitation Learning",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: human-machine interaction",
          "abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12603",
      "pdf_url": "https://arxiv.org/pdf/2606.12603.pdf"
    },
    {
      "id": "2606.12396",
      "anchor": "paper-2606-12396",
      "title": "VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving",
      "authors": [
        "Jin Yao",
        "Dhruva Dixith Kurra",
        "Tom Lampo",
        "Zezhou Cheng",
        "Danhua Guo",
        "Burhan Yaman"
      ],
      "abstract": "Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision-language-action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open-loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50\\,m average) and 3-second collision rate (0.18\\%). On closed-loop Bench2Drive, VLGA attains the state-of-the-art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.",
      "short_abstract": "Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 2,
        "evidence": [
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12396",
      "pdf_url": "https://arxiv.org/pdf/2606.12396.pdf"
    },
    {
      "id": "2606.12236",
      "anchor": "paper-2606-12236",
      "title": "DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems",
      "authors": [
        "Zhongyu Xia",
        "Wenhao Chen",
        "Yongtao Wang",
        "Ming-Hsuan Yang"
      ],
      "abstract": "Many autonomous driving systems are increasingly incorporating foundation models to improve generalization and handle long-tail scenarios. However, this trend introduces two key challenges: (i) the manual and labor-intensive process of designing and integrating new models, and (ii) the lack of intelligent, dynamic scheduling mechanisms to meet strict real-time constraints. While Large Language Model (LLM)-based agents offer a promising avenue for automation, existing frameworks are ill-suited for autonomous driving. Specifically, they fail to distinguish between the fundamentally different requirements of system design and real-time scheduling, treat modules as opaque black boxes, and are not designed for continuous operation. To address these limitations, we propose DrivingAgent, a novel agent framework tailored to the dual challenges of autonomous driving system design and scheduling. In the design phase, DrivingAgent automates module development by interpreting system architecture, generating code, and validating modules via super-network training. In the scheduling phase, it employs a lightweight LLM trained with reinforcement learning to dynamically orchestrate system modules in real time, supported by a structured memory that integrates long-term storage with timestamped short-term context. Experimental results demonstrate that DrivingAgent achieves a superior speed--accuracy trade-off on both the nuScenes and Bench2Drive benchmarks.",
      "short_abstract": "Many autonomous driving systems are increasingly incorporating foundation models to improve generalization and handle long-tail scenarios. However, this trend introduces two key challenges: (i) the manual and labor-intensive process of designing and...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 17,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: software architecture",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12236",
      "pdf_url": "https://arxiv.org/pdf/2606.12236.pdf"
    },
    {
      "id": "2606.12207",
      "anchor": "paper-2606-12207",
      "title": "Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends",
      "authors": [
        "Jinshan Lai",
        "Jianwei Hu",
        "Baoyang Jiang",
        "Fengchun Zhang",
        "Leyuan Wang",
        "Haotian Li",
        "Yida Wang",
        "Tingxuan Huang",
        "Xi Ren",
        "Qiang Ma"
      ],
      "abstract": "Embodied intelligence now spans navigation, household assistance, manipulation, autonomous driving, aerial agents, and multimodal large-model control. This expansion has made benchmark construction a central bottleneck for reliable evaluation. Unlike static datasets, embodied benchmarks combine task specifications, environments, robot data, demonstrations, annotations, metrics, evaluation scripts, and release policies into a single evaluation system. This survey reviews the literature through a five-stage construction pipeline: requirement and task construction, data acquisition, data cleaning and annotation, benchmark suite generation and metric definition, and evaluation execution with diagnostic feedback. For each stage, the survey analyzes the transition from manual curation to traditional automation, foundation-model assistance, and agentic closed-loop workflows. It also compares qualitative construction costs across human labor, data and asset acquisition, compute and simulation, validation and debugging, governance and maintenance, and rework risk. The main conclusion is that automation does not simply reduce benchmark cost. Instead, it often shifts cost toward validation, auditability, version control, and long-term governance. Progress in embodied evaluation will therefore depend not only on larger benchmark suites, but also on construction pipelines that are diagnosable, auditable, and responsibly refreshable.",
      "short_abstract": "Embodied intelligence now spans navigation, household assistance, manipulation, autonomous driving, aerial agents, and multimodal large-model control. This expansion has made benchmark construction a central bottleneck for reliable evaluation. Unlike static...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark",
        "Simulation",
        "Survey"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12207",
      "pdf_url": "https://arxiv.org/pdf/2606.12207.pdf"
    },
    {
      "id": "2606.12066",
      "anchor": "paper-2606-12066",
      "title": "Performance Analysis of YOLOv11 and YOLOv8 for Mixed Traffic Object Detection under Adverse Weather Conditions in Developing Countries",
      "authors": [
        "Quoc Thuan Nguyen",
        "Ha Anh Vu",
        "Ngo Dang Thanh Ngan",
        "Minh Phuc Hoang Ngoc"
      ],
      "abstract": "In modern vehicular systems, robust performance under harsh conditions has become a critical problem of autonomous driving. Our study delivers a comprehensive evaluation of the newest iteration of the YOLO series, which is YOLOv11 Nano architecture benchmarked against the widely adopted YOLOv8 Nano as a baseline on a custom fused dataset that combines the Indian Driving Dataset (IDD) [1] and Berkeley Deep Drive Dataset (BDD100K) [2]. We have analyzed the trade-offs among detection accuracy, inference speed, and computational efficiency in high-entropy scenarios involving dense mixed traffic, rain, and low-light conditions. Specifically, YOLOv11n achieves a mean Average Precision (mAP@50) of 46.6%, with a notable 3.2% improvement in Precision over the baseline, effectively reducing false positives in cluttered scenes. Furthermore, the proposed model exhibits enhanced energy efficiency, requiring 22% fewer FLOPs (6.3G vs. 8.1G) while maintaining real-time inference speed of 70.9 FPS on a Tesla T4 GPU, offering an optimal trade-off for safety-critical edge deployment.",
      "short_abstract": "In modern vehicular systems, robust performance under harsh conditions has become a critical problem of autonomous driving. Our study delivers a comprehensive evaluation of the newest iteration of the YOLO series, which is YOLOv11 Nano architecture...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "title: object detection"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12066",
      "pdf_url": "https://arxiv.org/pdf/2606.12066.pdf"
    },
    {
      "id": "2606.11989",
      "anchor": "paper-2606-11989",
      "title": "From Nominal Intensity to Equivalent Rainfall: A Path-Based Credibility Evaluation Framework for Simulated Rainfall in Autonomous-Driving Perception Tests",
      "authors": [
        "Tian Xia",
        "Xin Zhao",
        "Shaolingfeng Ye",
        "Junyi Chen"
      ],
      "abstract": "Credible simulated-rainfall conditions are essential for identifying perception-system boundaries and supporting SOTIF-oriented risk assessment in automated driving. However, closed-field tests are often described only by nominal rainfall intensity or single-point measurements, making it difficult to align simulated rain fields with real rainfall and map test results to real-world scenarios. This paper proposes a path-based credibility evaluation method for simulated rainfall in autonomous-driving perception tests. Using the drop size and velocity joint distribution of real rainfall as the reference, each candidate path is represented by path-equivalent rainfall intensity, an uncertainty band, and a path-averaged Realism of Raindrop Distribution (RRD) score. Lidar target point-cloud count and mean reflectivity are further used for perception-consistency correction, quantifying the proxy capability of each simulated-rainfall path for real-rainfall perception effects. Experiments are conducted using about 10,000 real-rainfall raindrop-spectrum samples, 728 RainSense perception samples, and 45 spatial sampling points in a 2.4 m x 7.2 m simulated-rainfall area. Results show that spatial non-uniformity remains under the same nominal condition, confirming the need for path-based evaluation. The method identifies Path IV and Path VI as preferable candidates, with results of 11.54 +/- 0.31 mm/h, RRD = 0.43, and 8.28 +/- 0.34 mm/h, RRD = 0.46, respectively. These paths show more balanced performance in rainfall-intensity stability, raindrop-spectrum realism, and perception consistency. The proposed method supports path selection, condition description, and credible interpretation of autonomous-driving perception tests under rainfall.",
      "short_abstract": "Credible simulated-rainfall conditions are essential for identifying perception-system boundaries and supporting SOTIF-oriented risk assessment in automated driving. However, closed-field tests are often described only by nominal rainfall intensity or...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 3,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.11989",
      "pdf_url": "https://arxiv.org/pdf/2606.11989.pdf"
    },
    {
      "id": "2606.11889",
      "anchor": "paper-2606-11889",
      "title": "Task-Aligned Stability Analysis of Vision-Language Models for Autonomous Driving Hazard Detection",
      "authors": [
        "Everett Richards"
      ],
      "abstract": "Vision-language models (VLMs) are increasingly used for scene understanding in autonomous driving, but robustness analysis often relies on task-agnostic embedding stability alone. We study whether corruption-induced embedding drift predicts changes in a task-aligned hazard score derived from CLIP image-text similarities. Using controlled corruptions on BDD100K road scenes, we compare embedding drift against margin drift, defined as the change in hazard score under perturbation. The relationship is highly corruption-dependent: some families exhibit strong coupling between representation drift and decision drift, while others induce hazardous decision instability despite relatively modest embedding change. Furthermore, corruption families differ in failure direction: most suppress hazard detections via false negatives, while occlusion instead triggers false alarms, suggesting that benchmark design should account for asymmetric failure modes, not just overall instability rates. These results suggest that robustness benchmarks should include task-aligned stability measures in addition to embedding-level perturbation statistics.",
      "short_abstract": "Vision-language models (VLMs) are increasingly used for scene understanding in autonomous driving, but robustness analysis often relies on task-agnostic embedding stability alone. We study whether corruption-induced embedding drift predicts changes in a...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness",
          "Safety, Security & Verification: abstract: failure",
          "Perception & Sensor Fusion: abstract: scene understanding"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.11889",
      "pdf_url": "https://arxiv.org/pdf/2606.11889.pdf"
    },
    {
      "id": "2606.11874",
      "anchor": "paper-2606-11874",
      "title": "AutoMine Solution for AV2 2026 Scenario Mining Challenge",
      "authors": [
        "Songliang Cao",
        "Jiele Zhao",
        "Yuru Wang",
        "Hao Li",
        "Daqi Liu",
        "Zehan Zhang",
        "Fangzhen Li",
        "Yu Wang",
        "Yue Zhang",
        "Bing Wang",
        "Guang Chen",
        "Hao Lu",
        "Hangjun Ye"
      ],
      "abstract": "With the development of autonomous driving systems, mining high-value, safety-critical, and planning-relevant scenarios from large-scale driving logs has become essential for data-driven evaluation. In this paper, we propose AutoMine, a robust self-refining scenario mining method based on LLMs and VLMs. AutoMine uses semantics-preserving prompt augmentation to reduce LLM prompt sensitivity, combines robust trajectory atomic functions with VLM-based functions to handle perception noise and open-world visual cues, and refines generated code through execution feedback from real logs. In the Argoverse 2 Scenario Mining Competition at CVPR 2026, AutoMine achieves a HOTA-Temporal score of 36.38 and a Timestamp BA score of 77.21.",
      "short_abstract": "With the development of autonomous driving systems, mining high-value, safety-critical, and planning-relevant scenarios from large-scale driving logs has become essential for data-driven evaluation. In this paper, we propose AutoMine, a robust self-refining...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.11874",
      "pdf_url": "https://arxiv.org/pdf/2606.11874.pdf"
    },
    {
      "id": "2606.11839",
      "anchor": "paper-2606-11839",
      "title": "Systematic Cybersecurity Risk Analysis of European Rail Traffic Management System",
      "authors": [
        "Kacper Darowski",
        "Sebastian N. Peters",
        "Lukas Lautenschlager"
      ],
      "abstract": "European Rail Traffic Management System (ERTMS) is a widely adopted standard unifying train management in the EU. While the standard allows for use cases like fully autonomous driving, cybersecurity has been an afterthought. Risk analysis enables the systematic assessment and prioritization of threats and mitigations. To date, it remains unclear which threats are most significant in ERTMS. This study systematically models components of ERTMS and analyzes their security in light of threats identified in the underlying technologies. The results suggest a concerning state of ERTMS, despite its critical role in railway safety. The use of legacy standards like EuroBalises and GSM-Railway (GSM-R) introduces vulnerabilities that persist across minimal ERTMS implementations, deployments incorporating various optional safety measures, and prospective future evolutions of the system, e.g., adopting Future Railway Mobile Communication System (FRMCS). Fully transitioning to European Train Control System (ETCS) level 2 was identified as the most significant measure for advancing ERTMS cybersecurity. The results indicate that a shift of ERTMS toward security is required to ensure availability and safe operation. While the chosen methodology proved its feasibility and shows remaining weaknesses of ERTMS, future work is needed to develop railway-centric adaptations to improve the quantification and evaluation of the computed risks.",
      "short_abstract": "European Rail Traffic Management System (ERTMS) is a widely adopted standard unifying train management in the EU. While the standard allows for use cases like fully autonomous driving, cybersecurity has been an afterthought. Risk analysis enables the...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 26,
        "evidence": [
          "abstract: safety",
          "title: security",
          "abstract: security",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.11839",
      "pdf_url": "https://arxiv.org/pdf/2606.11839.pdf"
    },
    {
      "id": "2606.12614",
      "anchor": "paper-2606-12614",
      "title": "DARRMS -- An Efficient Algorithm for Dynamic Attention Radius in Resource-Constrained Multi-Agent Systems",
      "authors": [
        "Benjamin Alcorn",
        "Eman Hammad"
      ],
      "abstract": "Multi-agent systems are integral tools for various domains such as robotics, cybersecurity, and autonomous vehicle planning. These types of systems often have constraints on the computational resources, leading to a need for efficient lightweight algorithms. Traditional decision making frameworks often assume ideal conditions, such as full observability and unlimited computational capacity, which do not align with real-world challenges. In this paper, we introduce a new algorithm that allows for reduced demand on computational resources without a large cost of other performance metrics. Agents will limit their observability to some attention radius, which intentionally allows them to ignore parts of the environment that might be unnecessary for action planning. By optimizing both the attention radius and decision-making, our approach enhances coordination and scalability in uncertain environments. Through both theoretical analysis and empirical validation, we demonstrate the effectiveness of adaptive observation in improving system performance and maintaining robust decision-making strategies in resource-constrained systems.",
      "short_abstract": "Multi-agent systems are integral tools for various domains such as robotics, cybersecurity, and autonomous vehicle planning. These types of systems often have constraints on the computational resources, leading to a need for efficient lightweight algorithms....",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 3,
        "evidence": [
          "abstract: security",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.12614",
      "pdf_url": "https://arxiv.org/pdf/2606.12614.pdf"
    },
    {
      "id": "2606.11971",
      "anchor": "paper-2606-11971",
      "title": "Cooperative Switched Formation Control of Autonomous Vehicles: An Event-triggered Approach to Input Saturation and Time-delay Challenges",
      "authors": [
        "Ziming Wang",
        "Guanxuan Jiang",
        "Yihuai Zhang",
        "Karl H. Johansson",
        "Apostolos I. Rikos"
      ],
      "abstract": "This paper presents a collaborative adaptive formation control framework for autonomous vehicles (AVs), that explicitly handles system uncertainties, input saturation, and communication delays. To overcome the inherent physical torque limits of steering and braking actuators, an input saturation compensation mechanism is introduced to render nonlinearities tractable and improve control reliability. Additionally, a delay-compensating auxiliary system is designed to mitigate the effects of communication delays and reduce tracking errors. Our framework incorporates a dynamic-threshold event-triggered control (ETC) strategy to optimize resource usage. Additionally, uncertainty observers and symmetric barrier Lyapunov functions are developed to ensure robust and safe formation maneuvers. Finally, the effectiveness of the proposed approach is validated through numerical simulations of vehicle formations, complemented by a 3D visualization video demonstrating the dynamic fleet reconfiguration process.",
      "short_abstract": "This paper presents a collaborative adaptive formation control framework for autonomous vehicles (AVs), that explicitly handles system uncertainties, input saturation, and communication delays. To overcome the inherent physical torque limits of steering and...",
      "published": "2026-06-10",
      "updated": "2026-06-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 4,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 70,
      "arxiv_url": "https://arxiv.org/abs/2606.11971",
      "pdf_url": "https://arxiv.org/pdf/2606.11971.pdf"
    },
    {
      "id": "2606.11569",
      "anchor": "paper-2606-11569",
      "title": "ConsistencyPlanner: Real-time Planning with Fast-Sampling Consistency Models",
      "authors": [
        "Qichao Zhang",
        "Xing Fang",
        "Jiaqi Fang",
        "Zhenwen Cai",
        "Jie Ling",
        "Qiankun Yu",
        "Dongbin Zhao"
      ],
      "abstract": "Closed-loop planning in complex, real-world driving scenarios presents a critical challenge for autonomous driving systems. While traditional rule-based methods are interpretable, their predefined heuristics lack the adaptability for dynamic traffic environments. Learning-based approaches have shown considerable promise. Conversely, learning-based approaches, despite their promise, struggle to balance the modeling diverse and multimodal driving behaviors and real-time planning, often leading to indecisive or unsafe actions. To address this limitation, we propose Consistency Planner, a real-time planning framework with fast-sampling consistency models. Our approach is built upon two key technical contributions. Efficient Multimodal Sampling: We employ fast-sampling consistency models to generate a diverse set of plausible future trajectories. This enables efficient, real-time exploration of multimodal actions, overcoming the computational bottlenecks of previous iterative generative methods. Heterogeneous Feature Fusion: We introduce an attention-enhanced decoder that dynamically integrates heterogeneous input features (including scene feature and action token) into a cohesive representation for robust planning. Extensive evaluation in the Waymax simulator demonstrates superior performance in safety metrics compared to existing methods, with particularly strong results in challenging dynamic scenarios.",
      "short_abstract": "Closed-loop planning in complex, real-world driving scenarios presents a critical challenge for autonomous driving systems. While traditional rule-based methods are interpretable, their predefined heuristics lack the adaptability for dynamic traffic...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "Hardware / Real-Time",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: title: planning",
          "Planning & Decision-Making: abstract: planning",
          "Systems, Deployment & Connectivity: title: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.11569",
      "pdf_url": "https://arxiv.org/pdf/2606.11569.pdf"
    },
    {
      "id": "2606.11120",
      "anchor": "paper-2606-11120",
      "title": "Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football",
      "authors": [
        "Andrew Kang",
        "Priya Narasimhan"
      ],
      "abstract": "We recast pass evaluation in football (soccer) as a Monte Carlo Tree Search (MCTS)-like evaluation problem whose components mostly exist in the literature under different names: a value model (possession value), a world model (multi-agent trajectories with ball interactions), and a policy over counterfactual actions (sampling pass variants with noise). Building on the first public high-fidelity tracking dataset with 3D ball trajectories from the Bundesliga, we introduce Monte Carlo Pass Search (MCPS), which infers kick parameters for each observed pass, samples execution variants and option variants, rolls each candidate forward with a ball-conditioned world model until the next ball interaction, and scores outcomes with a learned value model to obtain a distribution over gained value. This distribution enables distribution-aware attribution with two complementary execution-surplus scores used for analysis and ranking: mean-based and percentile-based scores. To make the world model sample-efficient under limited public data, we adapt a discrete-token, autoregressive trajectory generator from autonomous driving (SMART) and show it yields strong best-of-20 forecasting accuracy compared to baselines, while supporting fully hypothetical rollouts for downstream evaluation. We have released model checkpoints and code.",
      "short_abstract": "We recast pass evaluation in football (soccer) as a Monte Carlo Tree Search (MCTS)-like evaluation problem whose components mostly exist in the literature under different names: a value model (possession value), a world model (multi-agent trajectories with...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Dataset",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 2,
        "evidence": [
          "title: trajectory generation"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.11120",
      "pdf_url": "https://arxiv.org/pdf/2606.11120.pdf"
    },
    {
      "id": "2606.11019",
      "anchor": "paper-2606-11019",
      "title": "Diffusion Forcing Planner: History-Annealed Planning with Time-Dependent Guidance for Autonomous Driving",
      "authors": [
        "Zehan Zhang",
        "Neng Zhang",
        "Yaoyi Li",
        "Jia Cai",
        "Zhiling Wang"
      ],
      "abstract": "Learning-based motion planners, despite recent progress, often suffer from temporal inconsistency. Small perturbations across frames can accumulate into unstable trajectories, degrading comfort and safety in closed-loop driving. Several methods attempt to inject history as a static conditioning signal to stabilize outputs, only to induce the planner to copy historical patterns instead of adapting to environment contexts. To address this limitation, we propose Diffusion Forcing Planner (DFP), a diffusion-based planning framework driven by history-guided control. Specifically, DFP decomposes the full trajectory into history, current and future segments, and assign independent noise levels to each segment. The model jointly denoises the historical and the future segments, enforcing a heterogeneous joint diffusion process. At inference, classifier-free guidance (CFG) is applied to steer future sampling using annealed history in a controllable manner. Closed-loop evaluation and comprehensive ablations on nuPlan show that DFP achieves competitive performance while producing continuous, stable, and controllable motion plans in complex driving scenarios.",
      "short_abstract": "Learning-based motion planners, despite recent progress, often suffer from temporal inconsistency. Small perturbations across frames can accumulate into unstable trajectories, degrading comfort and safety in closed-loop driving. Several methods attempt to...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.11019",
      "pdf_url": "https://arxiv.org/pdf/2606.11019.pdf"
    },
    {
      "id": "2606.10974",
      "anchor": "paper-2606-10974",
      "title": "Language-Driven Cost Optimization for Autonomous Driving",
      "authors": [
        "Diego Martinez-Baselga",
        "Khaled Mustafa",
        "Javier Alonso-Mora"
      ],
      "abstract": "The driving behavior of autonomous vehicles is typically governed by the cost function of their motion planner, which encodes objectives such as speed tracking, smoothness, lane keeping, and collision avoidance. However, tuning the parameters that shape this cost function is a challenging task that requires technical expertise, limiting the vehicle's ability to adapt to evolving traffic scenarios or end-user preferences. This work presents a language-driven framework for adaptive cost design in autonomous driving. A Large Language Model (LLM) interprets structured scenario descriptions and natural language user queries to generate the parameters applied to a risk-aware Model Predictive Path Integral (MPPI) controller. The system incorporates a human-in-the-loop validation stage in which the proposed behavioral changes are described in non-technical language and confirmed prior to deployment. Users may additionally provide feedback either before or after deployment, enabling iterative refinement of the vehicle's motion behavior. The framework is evaluated across multiple queries in realistic driving scenarios to assess its effectiveness. Simulation results demonstrate that the method successfully induces behavioral changes that align with the intended requirements in an intuitive manner, thereby bridging the gap between intelligent vehicle control systems and end users.",
      "short_abstract": "The driving behavior of autonomous vehicles is typically governed by the cost function of their motion planner, which encodes objectives such as speed tracking, smoothness, lane keeping, and collision avoidance. However, tuning the parameters that shape this...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 2,
        "evidence": [
          "abstract: vehicle control",
          "abstract: controller",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.10974",
      "pdf_url": "https://arxiv.org/pdf/2606.10974.pdf"
    },
    {
      "id": "2606.10856",
      "anchor": "paper-2606-10856",
      "title": "An Exposure-Time-Aligned Primary-Path Architecture for Autonomous-Driving ECUs",
      "authors": [
        "Toru Saito",
        "Yuki Hagura",
        "Tatsuya Konishi",
        "Satoru Mizusawa",
        "Takumi Yajima"
      ],
      "abstract": "While end-to-end (E2E) autonomous driving has become the dominant research direction, production vehicles continue to rely on modular multi-NN pipelines for a non-trivial transitional period. The subject of this paper is the design of an architecture that, during this phase, supports a modular pipeline and an E2E path side by side and embeds a path for staged migration. Transplanted to a production SoC, egalitarian late fusion is compute-inefficient and offers no natural unit for staged E2E substitution. As an alternative, we propose three design principles: (i) Primary-Path, which selects a primary perception-to-planner chain and prioritizes its enclosure within a single SoC over the non-critical paths, (ii) Exposure-Time-Aligned, which propagates the primary sensor's exposure time $τ_{\\mathrm{exp}}$ as a tag along the chain and event-drives the fusion node on matched $τ_{\\mathrm{exp}}$ rather than a fixed cycle, and (iii) Co-Path Coexistence, which, building on (i) and (ii), lets an E2E output path co-run with the modular pipeline within the same $τ_{\\mathrm{exp}}$ cycle. On a Dual-SoC production AD-ECU, the implementation achieves a camera-shutter to planner-output latency of 296 ms on average within a 350 ms design budget, with the modular pipeline and the E2E path coexisting within the shared $τ_{\\mathrm{exp}}$ cycle.",
      "short_abstract": "While end-to-end (E2E) autonomous driving has become the dominant research direction, production vehicles continue to rely on modular multi-NN pipelines for a non-trivial transitional period. The subject of this paper is the design of an architecture that,...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: camera",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.10856",
      "pdf_url": "https://arxiv.org/pdf/2606.10856.pdf"
    },
    {
      "id": "2606.10851",
      "anchor": "paper-2606-10851",
      "title": "Modular2Simple: A Tool for Modular Scenario Creation Based on the OpenSCENARIO Format",
      "authors": [
        "Nikolai Khriapov",
        "Mohamed Taha Drif",
        "Renjue Li",
        "Cas Widdershoven"
      ],
      "abstract": "The rapid advancement of autonomous driving systems (ADS) has introduced significant challenges, particularly in the creation of realistic and complex scenarios for testing and validation. This paper introduces Modular2Simple, a tool designed to address these challenges by simplifying and enhancing the process of creating complex ADS scenarios. Modular2Simple seamlessly integrates with the CARLA simulator and is applicable to any software that supports the OpenSCENARIO format. By leveraging existing simple scenarios in the OpenSCENARIO format, the tool enables developers to create easily customizable modular scenarios through the combination of multiple simple or modular scenarios, significantly simplifying the scenario creation process while maintaining flexibility in scenario design. This approach not only facilitates the development of complex scenarios, reducing both development time and effort, but also promotes scenario reuse and customization, which leads to a significant reduction in code complexity and enhanced efficiency in scenario design and testing compared to traditional scenario development methods.",
      "short_abstract": "The rapid advancement of autonomous driving systems (ADS) has introduced significant challenges, particularly in the creation of realistic and complex scenarios for testing and validation. This paper introduces Modular2Simple, a tool designed to address these...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.10851",
      "pdf_url": "https://arxiv.org/pdf/2606.10851.pdf"
    },
    {
      "id": "2606.10688",
      "anchor": "paper-2606-10688",
      "title": "Self-Supervised Relevance Modelling in Autonomous Driving via Counterfactual Analysis",
      "authors": [
        "Luca Lusvarghi",
        "Javier Gozalvez",
        "Pablo Urbano Hidalgo"
      ],
      "abstract": "Autonomous driving relies on computationally intensive perception pipelines to continuously detect and track objects in the surrounding environment. While some objects are key to plan safe and effective maneuvers, others may not be relevant and have no impact on the autonomous vehicle's driving decisions. Focusing on relevant objects allows a more efficient usage of available computational resources, reduces processing latencies, and limits the downstream propagation of perception noise. In this work, we propose a novel self-supervised approach based on counterfactual analysis to develop a relevance model - an AI-based tool that quantifies the relevance of objects for an autonomous vehicle. To demonstrate the potential of the proposed approach, we train a relevance model on a synthetic causal dataset generated in a selected urban scenario. Results show that the relevance model is able to accurately estimate the objects' relevance with millisecond-level latency, enabling real-time relevance estimation also in high-density scenarios. We also show that the relevance model can be used to build relevance heatmaps that offer valuable insights into the autonomous vehicle's driving policy and can be used to proactively inform perception and planning tasks. We openly release both the relevance model and the causal dataset.",
      "short_abstract": "Autonomous driving relies on computationally intensive perception pipelines to continuously detect and track objects in the surrounding environment. While some objects are key to plan safe and effective maneuvers, others may not be relevant and have no impact...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 2,
        "evidence": [
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.10688",
      "pdf_url": "https://arxiv.org/pdf/2606.10688.pdf"
    },
    {
      "id": "2606.10656",
      "anchor": "paper-2606-10656",
      "title": "Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving",
      "authors": [
        "Qi Song",
        "Yifei He",
        "Chi Zhang",
        "Zheng Fu",
        "Xuhe Zhao",
        "Mengmeng Yang",
        "Kun Jiang",
        "Rui Huang",
        "Diange Yang"
      ],
      "abstract": "Forecasting the future evolution of dynamic scenes is crucial in autonomous driving. However, existing feed-forward paradigms are primarily designed for interpolation. When extended to future extrapolation, they suffer from ghosting artifacts under large displacements and are constrained by simplified motion assumptions or strict future priors. To overcome these challenges, we propose Envision4D, a fully self-supervised feed-forward framework for pose-free future extrapolation. Specifically, we introduce a Future Pose Prediction module that infers future camera parameters via an iterative denoising process. Furthermore, to capture non-linear dynamics, we propose In-layer Temporal Attention and employ Conditioned Motion Lifting, which transforms the highly uncertain extrapolation process into robust relational mappings. Finally, a Progressive Training Strategy is utilized to stabilize unsupervised motion learning against error accumulation. Extensive experiments demonstrate that Envision4D achieves state-of-the-art performance, significantly outperforming existing methods in future view synthesis.",
      "short_abstract": "Forecasting the future evolution of dynamic scenes is crucial in autonomous driving. However, existing feed-forward paradigms are primarily designed for interpolation. When extended to future extrapolation, they suffer from ghosting artifacts under large...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Prediction & World Models: abstract: forecasting",
          "Prediction & World Models: abstract: prediction",
          "Perception & Sensor Fusion: abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.10656",
      "pdf_url": "https://arxiv.org/pdf/2606.10656.pdf"
    },
    {
      "id": "2606.14776",
      "anchor": "paper-2606-14776",
      "title": "Deep Learning-Based Lunar Crater Terrain Relative Navigation",
      "authors": [
        "Batu Candan",
        "Simone Servadio"
      ],
      "abstract": "Accurate position estimation is crucial for the successful implementation of future lunar landings using autonomous vehicles, especially in dangerous environments with sparse terrain features. In this paper, we propose a terrain relative navigation (TRN) algorithm combining our deep-learning crater detector, which was designed specifically for the NASA Crater Detection Challenge problem, and an Extended Kalman Filter (EKF). Our detector analyzes crater features from the monocular images acquired from orbit, and their matches with craters from a global database are identified via a Hungarian assignment approach followed by the consensus-based outliers removal method. The estimated measurements are then used to refine an EKF, where spacecraft pose estimation in the Lunar-Centered Lunar-Fixed (LCLF) frame of reference, augmented with altitude aiding information, constrains radial drift. The simulation results indicate that even if the spacecraft is off from its actual location up to 5 km, TRN could recover from this situation, achieving navigation error reduction to a few hundred meters. It should be noted that in order to maintain crater feature correspondences, it is important to match the image resolution and the scales within the scene to the detector training set distribution.",
      "short_abstract": "Accurate position estimation is crucial for the successful implementation of future lunar landings using autonomous vehicles, especially in dangerous environments with sparse terrain features. In this paper, we propose a terrain relative navigation (TRN)...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: navigation",
          "abstract: navigation"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.14776",
      "pdf_url": "https://arxiv.org/pdf/2606.14776.pdf"
    },
    {
      "id": "2606.10997",
      "anchor": "paper-2606-10997",
      "title": "A Companion App for an Autonomous Family Vehicle: Identification of Values for an Autonomous Mobility System",
      "authors": [
        "Leon Johann Brettin",
        "Tobias Schräder",
        "Kerstin Kuhlmann",
        "Vanessa Schmidt",
        "Markus Maurer"
      ],
      "abstract": "In this paper, we present a companion app for an autonomous vehicle aimed at user groups who would normally require an accompanying person to drive them. Two aspects of a companion app are presented in this paper: First, the possibility for a trusted person to track the ride of the person in need of support and second, to put the settings of the vehicle for persons in need of support in the hands of a trusted person. In addition, this article describes the requirements and addressed values and discusses the safety-relevant aspects of such a companion app. We also discuss and identify the values that influence passengers and trusted persons using the companion app. Overall, a companion app can provide new perspectives and opportunities for people in need of support, allowing them to take advantage of the features offered by autonomous vehicles. It enables trusted individuals to configure the vehicle according to the passengers needs. Also such an app can be a mechanism to involve trusted persons in the options given by the vehicle and give them the possibility to adapt the vehicle to the needs of the person in need of support.",
      "short_abstract": "In this paper, we present a companion app for an autonomous vehicle aimed at user groups who would normally require an accompanying person to drive them. Two aspects of a companion app are presented in this paper: First, the possibility for a trusted person...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.10997",
      "pdf_url": "https://arxiv.org/pdf/2606.10997.pdf"
    },
    {
      "id": "2606.10774",
      "anchor": "paper-2606-10774",
      "title": "Asynchronous Decentralized Federated Learning over Lossy Wireless Links via Reception- and Age-Aware Aggregation",
      "authors": [
        "Chanuka A. S. Hewa Kaluannakkage",
        "Rajkumar Buyya"
      ],
      "abstract": "Decentralized Federated Learning(DFL) enables collaborative model training across wireless edge nodes, including IoT deployments, autonomous vehicles, UAV swarms, and satellite constellations. Operating over lossy wireless links under constraints, these systems cannot rely on retransmissions, so model parameters must be accepted as partial chunks, leading to two key failure modes, which are selection bias, where poor-quality links are systematically under-represented in gossip aggregation, and update staleness, where asynchronous nodes contribute outdated models. We prove that classical gossip aggregation introduces irreducible selection bias proportional to the link-loss rate. We propose DFL-AA (Decentralized Federated Learning with Adaptive AoI-weighted Aggregation), which corrects selection bias using Inverse Probability Weighting (IPW) with online channel estimation and mitigates staleness via Age-of-Information (AoI) decay without requiring a global clock. We prove that DFL-AA removes link-quality distortion in expectation and consistently outperforms state-of-the-art baselines across varying loss rates and heterogeneous channel conditions on fixed directed topologies.",
      "short_abstract": "Decentralized Federated Learning(DFL) enables collaborative model training across wireless edge nodes, including IoT deployments, autonomous vehicles, UAV swarms, and satellite constellations. Operating over lossy wireless links under constraints, these...",
      "published": "2026-06-09",
      "updated": "2026-06-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: cooperative systems",
          "Safety, Security & Verification: abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 71,
      "arxiv_url": "https://arxiv.org/abs/2606.10774",
      "pdf_url": "https://arxiv.org/pdf/2606.10774.pdf"
    },
    {
      "id": "2606.10303",
      "anchor": "paper-2606-10303",
      "title": "Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-based Accelerators",
      "authors": [
        "Chenguang Zhang",
        "Yuanpeng Zhang",
        "Chenhao Xue",
        "Yihan Yin",
        "Chen Zhang",
        "Guangyu Sun"
      ],
      "abstract": "Level-4+ autonomous driving systems (ADS) must run dozens of heterogeneous deep neural networks (DNNs) as end-to-end (E2E) pipelines under a strict latency constraint (<=100 ms), even as execution time varies by up to 3.3x. Cost rules out dedicating isolated hardware to each function in mass-produced ADS, so these DNNs must be densely colocated on a single chip, which introduces shared-resource contention. Tile-based accelerators expose two scheduling opportunities that conventional ADS schedulers do not exploit. First, they provide a tunable degree of parallelism (DoP): assigning more tiles raises DoP and can shorten DNN execution time. Second, they provide hardware-native isolation: tiles can be physically partitioned among co-located DNNs. But using this flexibility is expensive: changing a task's DoP triggers a stop-migrate-restart reallocation of its weights and intermediate features. At ADS task rates of 10-240 Hz, these stalls accumulate along E2E chains and threaten deadlines. Reservation-based schedulers fix DoP and leave this flexibility unused; work-conserving schedulers exploit it but assume reallocation is cheap and treat deadlines as independent. We present ADS-Tile that combines configurable isolation and elastic reservation into a spatio-temporal isolation-sharing space that bounds where and when reallocation occurs; a probabilistic latency model and a DAG-aware runtime scheduler then use this space to decide task colocation and DoP under shared E2E deadlines. On an industry- and academia- derived ADS benchmark, ADS-Tile uses up to 32% fewer tiles than the work-conserving baseline in deadline-critical settings and cuts reallocation-induced wasted processing capacity from 17%-44% to below 1.2%. Controlled spatio-temporal sharing improves resource efficiency and latency predictability for tile-based ADS.",
      "short_abstract": "Level-4+ autonomous driving systems (ADS) must run dozens of heterogeneous deep neural networks (DNNs) as end-to-end (E2E) pipelines under a strict latency constraint (<=100 ms), even as execution time varies by up to 3.3x. Cost rules out dedicating isolated...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 8,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.10303",
      "pdf_url": "https://arxiv.org/pdf/2606.10303.pdf"
    },
    {
      "id": "2606.10044",
      "anchor": "paper-2606-10044",
      "title": "Business World Model",
      "authors": [
        "Cecil Pang",
        "Hiroki Sayama"
      ],
      "abstract": "World model has emerged as a powerful paradigm in artificial intelligence, enabling agents to represent their environments, predict future states, and evaluate possible actions before acting. However, existing world model approaches have largely been developed for domains such as computer vision, robotics, gaming, and autonomous driving, where the world is primarily visual or physical and governed by relatively stable dynamics. These formulations are not directly applicable to business practice, where the relevant environment is semantic, organizational, and market-driven rather than physical. Business outcomes depend on context-sensitive factors such as customer behavior, pricing, competition, regulation, resources, and operational constraints. This paper introduces the concept and architecture of a Business World Model (BWM), which is a world model specialized for business and organizational environments. A BWM encodes business states, dynamics, and feasible actions space to support autonomous business planning and decision-making. We propose a business-semantics-centric formulation in which states, dynamics, and actions are linked to key business entities, their attributes, and their relationships. Within this framework, intelligent agents can simulate alternative action sequences, estimate their effects on future business outcomes, and evaluate trade-offs under uncertainty. The proposed architecture integrates semantic data representations, probabilistic machine learning models, deterministic business rules, and explicit action spaces into a coherent internal simulator. This work establishes a conceptual foundation for autonomous business systems capable of moving from instruction-based execution toward goal-driven planning, optimization, and execution.",
      "short_abstract": "World model has emerged as a powerful paradigm in artificial intelligence, enabling agents to represent their environments, predict future states, and evaluate possible actions before acting. However, existing world model approaches have largely been...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 9,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.10044",
      "pdf_url": "https://arxiv.org/pdf/2606.10044.pdf"
    },
    {
      "id": "2606.09958",
      "anchor": "paper-2606-09958",
      "title": "Uncertainty-Aware Motion Planning for Autonomous Driving in Mixed Traffic Environment",
      "authors": [
        "Ming Cheng",
        "Hao Chen",
        "Ziyi Yang",
        "Ziluowen Luo",
        "Senzhang Wang"
      ],
      "abstract": "In mixed-traffic environments where autonomous and human-driven vehicles may co-exist, motion planning for autonomous vehicles requires anticipating the future behaviors of surrounding human drivers. Existing reinforcement learning-based methods generally directly incorporate the predicted human intents into the observation to enable a proactive planning. However, human intent is inherently uncertain due to the behavioral diversity, perception noise, and partial observability. Treating predicted intends as deterministic states can result in unsafe decisions for autonomous vehicles. To address this problem, we propose Uncertainty-Aware Motion Planning (UAMP), which incorporates uncertainty in human intent prediction for AV decision-making. Specifically, UAMP first introduces a proximity-aware uncertainty estimator to quantify the interaction-conditioned intent uncertainty and constructs an uncertainty-guided joint intent distribution over surrounding human-driven vehicles. Within this uncertainty set, UAMP further introduces Uncertainty-Calibrated Value Learning (UCVL) to correct value function learning biases arising from directly incorporating uncertain human intent predictions into the observation. Extensive experiments in various mixed-traffic scenarios show that UAMP significantly improves safety and driving comfort, while maintaining traffic efficiency compared with existing approaches. The code is released at https://anonymous.4open.science/r/UAMP-5638.",
      "short_abstract": "In mixed-traffic environments where autonomous and human-driven vehicles may co-exist, motion planning for autonomous vehicles requires anticipating the future behaviors of surrounding human drivers. Existing reinforcement learning-based methods generally...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 24,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09958",
      "pdf_url": "https://arxiv.org/pdf/2606.09958.pdf"
    },
    {
      "id": "2606.09746",
      "anchor": "paper-2606-09746",
      "title": "Hybrid Robustness Verification for Spatio-Temporal Neural Networks",
      "authors": [
        "Sherwin Varghese",
        "Matthew Wicker",
        "Alessio Lomuscio"
      ],
      "abstract": "With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential. Existing verification methods either rely on overly conservative approximations or incur prohibitive computational costs. For example, the use of lp-norm perturbations in video settings encodes the belief that the adversary can inject noise in every video frame. In practice, adversarial perturbations exhibit structured spatial and temporal correlations, constrained to lower-dimensional, semantically meaningful subspaces. In this work, we study robustness verification of 3D CNNs processing video and volumetric inputs, targeting applications in action recognition (UCF-101), autonomous driving (Udacity), and medical imaging (MedMNIST) exploiting realistic assumptions on adversarial strength by modelling them as spatio-temporal constraints - where the attacker can modify either a subset of frames or patches within a set of consecutive frames. We demonstrate that modelling realistic constraints enables tighter approximations. We introduce Spatio-Temporal Bound Propagation (STBP), a verification framework that computes an exact closed-form characterization of the first convolutional layer and propagates certified bounds through subsequent layers using scalable approximations. Computing the exact closed form provides the tightest bounds for the first convolutional layer. Thus, we utilise approximation methods in the remainder of the network. To spur further progress in this field, we propose ST-Bench, a verification benchmark for autonomous driving and activity recognition, to systematically evaluate verifiable robustness. Compared to existing verification-based approaches, STBP provides stronger robustness guarantees with significantly improved scalability, achieving 1.7x higher certified robust accuracy under identical perturbation budgets.",
      "short_abstract": "With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential. Existing verification methods either rely on overly conservative approximations or incur prohibitive computational costs....",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 28,
        "evidence": [
          "abstract: safety",
          "title: verification",
          "abstract: verification",
          "abstract: attack or threat",
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09746",
      "pdf_url": "https://arxiv.org/pdf/2606.09746.pdf"
    },
    {
      "id": "2606.09644",
      "anchor": "paper-2606-09644",
      "title": "Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving",
      "authors": [
        "Yimu Wang",
        "Yee Man Choi",
        "Barry Zhang",
        "Mozhgan Nasr Azadani",
        "Sean Sedwards",
        "Krzysztof Czarnecki"
      ],
      "abstract": "Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a model relied on the correct visual evidence. This gap is particularly important in multi-view driving scenes used for autonomous driving, where a model can produce a plausible answer while grounding it in the wrong camera view. We introduce a multi-view visual question answering benchmark for evaluating evidence-source identification: given six synchronized NuScenes views and a question, the model must identify the supporting camera view and answer the question. The benchmark contains 122 conflict-centric question-answer pairs from 73 scenes, spanning causality, counterfactual reasoning, and intent prediction. View labels are proposed by an automatic conflict-mining pipeline and manually verified by annotators. We evaluate three settings: camera-view selection, oracle QA given the golden view, and joint prediction in which the model selects a view and answers in one pass. Answers are evaluated in both multiple-choice and free-form formats, using exact match for structured predictions and an LLM judge for free-form responses. By explicitly separating visual-source identification from answer correctness, the benchmark exposes grounding failures that answer-only evaluation misses.",
      "short_abstract": "Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a model relied on the correct visual evidence. This gap is particularly important in multi-view driving scenes...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 4,
        "evidence": [
          "abstract: behavior prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09644",
      "pdf_url": "https://arxiv.org/pdf/2606.09644.pdf"
    },
    {
      "id": "2606.09569",
      "anchor": "paper-2606-09569",
      "title": "Efficient Minimal Solvers for Relative Pose Estimation in Autonomous Driving Applications",
      "authors": [
        "Tao Li",
        "Liang Liu",
        "Jianli Han",
        "Weimin Lv"
      ],
      "abstract": "With the advancement of visual sensing systems, computer vision is playing an increasingly important role in autonomous driving and robot navigation. Relative pose estimation in multi-camera systems is essential for accurate vehicle localization and environment perception, demanding high real-time performance and robustness. Existing methods, however, often involve high computational costs and rely heavily on abundant feature matches, limiting their applicability in time-sensitive driving scenarios. To address these limitations, this paper introduces a unified framework for efficient relative pose estimation, built upon a novel translation parameterization and first-order rotation approximation. Within this framework, we propose three efficient minimal solvers specifically designed for autonomous vehicles. The first solver integrates the vertical direction prior from Inertial Measurement Units (IMUs), the second utilizes the rotation axis direction prior during steering maneuvers, and the third is designed for planar motion - a realistic assumption for ground vehicles operating on structured roads. By reducing both the minimal number of point correspondences and the algebraic complexity, our methods enable faster hypothesis generation within RANSAC-based pipelines, improving suitability for real-time systems. Extensive experiments on synthetic datasets and the KITTI autonomous driving benchmark demonstrate that the proposed solvers achieve a favorable balance between speed and accuracy compared to existing state-of-the-art algorithms.",
      "short_abstract": "With the advancement of visual sensing systems, computer vision is playing an increasingly important role in autonomous driving and robot navigation. Relative pose estimation in multi-camera systems is essential for accurate vehicle localization and...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 16,
        "evidence": [
          "abstract: localization",
          "title: pose estimation",
          "abstract: pose estimation"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09569",
      "pdf_url": "https://arxiv.org/pdf/2606.09569.pdf"
    },
    {
      "id": "2606.09362",
      "anchor": "paper-2606-09362",
      "title": "Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study",
      "authors": [
        "Eduardo Borges",
        "Manuel Abreu",
        "Luís Garrote",
        "Urbano J. Nunes"
      ],
      "abstract": "Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues. However, purely visual representations may be sensitive to viewpoint, occlusion, illumination, and sensor-domain variations, limiting their interpretability and robustness in complex driving scenes. We propose a baseline study of a zero-shot pipeline using Vision-Language Models (VLMs) to generate textual descriptions of detected traffic participants and evaluate whether these descriptions can support identity matching across observations. Instead of relying only on low-level visual similarity, the proposed formulation represents each object through structured semantic attributes, including category, color, shape, pose, visible parts, spatial context, and distinctive visual cues. This study provides an initial benchmark for language-based re-identification in autonomous-driving scenarios, discussing and evaluating the strengths and limitations of current VLMs for this task. Results demonstrate that zero-shot semantic descriptions can support effective object re-identification, achieving retrieval performance comparable to a supervised CNN baseline while offering greater interpretability through explicit identity cues. However, the experiments also reveal important challenges, including attribute inconsistency across viewpoints and limited fine-grained discrimination between visually similar instances.",
      "short_abstract": "Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09362",
      "pdf_url": "https://arxiv.org/pdf/2606.09362.pdf"
    },
    {
      "id": "2606.09350",
      "anchor": "paper-2606-09350",
      "title": "Taming Perception Jitter: Uncertainty-Aware LiDAR Object Detection for Reliable Motion Classification",
      "authors": [
        "Cornelius Schröder",
        "Žygimantas Marcinkus",
        "Markus Lienkamp"
      ],
      "abstract": "Reliable motion classification is critical for autonomous driving, as false dynamic predictions of static objects can cascade into unnecessary planner interventions. Unstable bounding box predictions can lead to spurious velocity estimates in tracking and falsely predicted trajectories. We present a deployment-friendly mitigation strategy that augments a 3D object detector with aleatoric uncertainty estimates and applies a two-sample z-test over short observation windows to separate true motion from jitter. Integrated into Autoware with minimal changes, the approach reuses existing data association for minimal compute overhead. Empirical results show parity with velocity thresholding on nuScenes, but substantially fewer false dynamic predictions and unnecessary stops in real-world test drives, explained by the presence of an intermediate jitter band in the recorded data that speed-only rules misclassify. This demonstrates that uncertainty-aware detection and lightweight statistical testing can deliver practical performance gains for autonomous driving in noisier real-world settings.",
      "short_abstract": "Reliable motion classification is critical for autonomous driving, as false dynamic predictions of static objects can cascade into unnecessary planner interventions. Unstable bounding box predictions can lead to spurious velocity estimates in tracking and...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 19,
        "evidence": [
          "title: perception",
          "title: object detection",
          "title: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09350",
      "pdf_url": "https://arxiv.org/pdf/2606.09350.pdf"
    },
    {
      "id": "2606.09273",
      "anchor": "paper-2606-09273",
      "title": "EditSSC: Toward Editable Semantic Occupancy Scenes with Unconditional Diffusion Models",
      "authors": [
        "Fatima Balde",
        "Raoul de Charette",
        "Alexandre Boulch"
      ],
      "abstract": "3D semantic scene generation is crucial for autonomous driving applications, yet most methods rely on complex 3D-specific architectures such as triplane encoders and adapted diffusion networks, limiting both their simplicity and their editing capabilities. We propose EditSSC, an editing-ready method for 3D semantic scene generation using 2D Bird's Eye View (BEV) representations and off-the-shelf latent diffusion network. Our approach reshapes 3D semantic occupancy grids into multi-channel BEV images and leverages the quantized autoencoder and UNet from Stable Diffusion with minimal modifications. We perform diffusion on the latents after quantization, which enables training-free editing capabilities. By exploiting class-to-code correspondences in the codebook, our method supports sketch-guided generation, inpainting, and outpainting without any retraining. On SemanticKITTI, EditSSC outperforms existing 3D-specific baselines on unconditional generation, demonstrating that well-established 2D architectures can be effectively repurposed for 3D scene generation and editing.",
      "short_abstract": "3D semantic scene generation is crucial for autonomous driving applications, yet most methods rely on complex 3D-specific architectures such as triplane encoders and adapted diffusion networks, limiting both their simplicity and their editing capabilities. We...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 8,
        "evidence": [
          "title: occupancy",
          "abstract: occupancy",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09273",
      "pdf_url": "https://arxiv.org/pdf/2606.09273.pdf"
    },
    {
      "id": "2606.09236",
      "anchor": "paper-2606-09236",
      "title": "Self-Paced Curriculum Reinforcement Learning for Autonomous Superbike Racing in Simulation",
      "authors": [
        "Luca Ghisi",
        "Jacopo Essenziale",
        "Carlo D'Eramo",
        "Matteo Luperto"
      ],
      "abstract": "Autonomous Racing has seen remarkable progress through deep Reinforcement Learning (RL), primarily for four-wheeled vehicles. However, motorbikes introduce substantially greater complexity due to the need to manage balance and lean angle, in addition to more reactive steering and throttle control, and a smaller weight. In this work, we present a framework for training an autonomous agent to race a superbike in VRider SBK, a physics-accurate Unity-based motorbike simulator. Our approach integrates Soft Actor-Critic (SAC) with Self-Paced curriculum Deep reinforcement Learning (SPDL), which dynamically generates progressively more challenging tasks based on the agent's performance, without requiring manual curriculum design. The agent's state space comprises proprioceptive features extended with lean-angle history, along with global track features via course points. The reward signal is shaped to encourage progress along the track while penalizing instability-inducing behaviors specific to two-wheeled dynamics. Preliminary experimental results demonstrate that SPDL outperforms SAC alone in training efficiency, lap time, and driving stability across multiple tracks and motorbike models, establishing a first baseline for RL-based autonomous motorbike racing.",
      "short_abstract": "Autonomous Racing has seen remarkable progress through deep Reinforcement Learning (RL), primarily for four-wheeled vehicles. However, motorbikes introduce substantially greater complexity due to the need to manage balance and lean angle, in addition to more...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control",
          "Control & Vehicle Dynamics: abstract: steering or braking"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09236",
      "pdf_url": "https://arxiv.org/pdf/2606.09236.pdf"
    },
    {
      "id": "2606.09109",
      "anchor": "paper-2606-09109",
      "title": "Driving Video Retrieval for Complex Queries with Structured Grounding",
      "authors": [
        "Manyi Yao",
        "Sparsh Garg",
        "Christian Shelton",
        "Amit Roy-Chowdhury",
        "Abhishek Aich"
      ],
      "abstract": "Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be explicitly described in text or captured by lexical overlap. Rule-based retrieval can encode such events more directly, but it is brittle: generated or hand-written rules often fail when their assumptions do not match real driving data. We propose STRIVE-D, a data-calibrated retrieval framework for driving videos. It uses weakly labeled in-domain videos to estimate when a query rule is reliable, adapt rules that mismatch observed data, and fuse calibrated rule scores with vision-language and keyword-based retrieval signals. Across three driving benchmarks, including newly released human-annotated event data on DrivingDojo, STRIVE-D delivers up to 84% relative improvement in top-1 accuracy over state-of-the-art methods.",
      "short_abstract": "Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09109",
      "pdf_url": "https://arxiv.org/pdf/2606.09109.pdf"
    },
    {
      "id": "2606.09477",
      "anchor": "paper-2606-09477",
      "title": "Efficient Minimal Solvers for Visual-Inertial Relative Pose Estimation in Multi-Camera Systems",
      "authors": [
        "Tao Li",
        "Zhenbao Yu",
        "Banglei Guan",
        "Jianli Han",
        "Weimin Lv"
      ],
      "abstract": "Estimating the relative poses of multi-camera systems is a fundamental problem in computer vision, with critical applications in autonomous vehicles, mobile devices, and unmanned aerial vehicles (UAVs). However, existing solutions often suffer from high computational complexity or rely on an excessive number of point correspondences, limiting their real-world applicability. To address these limitations, we propose two efficient minimal solvers for estimating the relative poses of multi-camera systems using a novel parameterization. The first solver leverages the vertical direction prior provided by Inertial Measurement Units (IMUs), while the second utilizes the rotation axis direction prior from IMUs. Our methods require only four point correspondences and reduce the problem of multi-camera relative pose estimation to solving a univariate 6th-degree polynomial, a significant improvement over existing approaches, which typically involve 8th-degree polynomials. This reduction in computational complexity and correspondence requirements makes our solvers particularly effective when integrated into RANSAC frameworks, demonstrating strong potential for visual odometry applications. Through rigorous evaluations on synthetic data and the KITTI benchmark, our methods achieved superior computational efficiency and competitive accuracy compared to state-of-the-art algorithms.",
      "short_abstract": "Estimating the relative poses of multi-camera systems is a fundamental problem in computer vision, with critical applications in autonomous vehicles, mobile devices, and unmanned aerial vehicles (UAVs). However, existing solutions often suffer from high...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 14,
        "evidence": [
          "abstract: visual odometry",
          "title: pose estimation",
          "abstract: pose estimation"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09477",
      "pdf_url": "https://arxiv.org/pdf/2606.09477.pdf"
    },
    {
      "id": "2606.09292",
      "anchor": "paper-2606-09292",
      "title": "Dual Quaternion-Based Unscented Kalman Filter with Visual Inertial Odometry for Navigation in GPS-Denied Environments",
      "authors": [
        "Mohamed Khalifa",
        "Hashim A. Hashim"
      ],
      "abstract": "Reliable navigation in GPS-denied environments remains a fundamental challenge in robotics, aerospace, and autonomous vehicle applications. This paper presents a Dual Quaternion-Based Unscented Kalman Filter (DQUKF) equipped with a Visual Inertial Odometry (VIO) algorithm for accurate state estimation enabling navigation in GPS denied locations. The proposed framework formulates the DQUKF in an error state manner, where the nominal pose is represented by a unit dual quaternion and the local pose error is represented by a 6-dimensional twistor parameterization used for sigma point generation, covariance propagation, and measurement correction. In parallel, the VIO algorithm tracks features across image frames, synchronizes measurements between the IMU and camera, and provides visual constraints that complement inertial propagation. Simulation results on the EuRoC MAV dataset show that the proposed DQUKF converges under high initialization uncertainty and achieves a position RMSE of 0.2584~m in the difficult flight sequence, outperforming the benchmark filters.",
      "short_abstract": "Reliable navigation in GPS-denied environments remains a fundamental challenge in robotics, aerospace, and autonomous vehicle applications. This paper presents a Dual Quaternion-Based Unscented Kalman Filter (DQUKF) equipped with a Visual Inertial Odometry...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 0,
        "evidence": [
          "Mapping & Localization: title: GNSS or GPS",
          "Mapping & Localization: abstract: GNSS or GPS",
          "Planning & Decision-Making: title: navigation",
          "Planning & Decision-Making: abstract: navigation"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09292",
      "pdf_url": "https://arxiv.org/pdf/2606.09292.pdf"
    },
    {
      "id": "2606.09534",
      "anchor": "paper-2606-09534",
      "title": "A Continuification Approach to CAV Control in Mixed Traffic via Variable Speed Limits",
      "authors": [
        "Brian Block",
        "Cecilia Pasquale",
        "Silvia Siri",
        "Simona Sacone",
        "Stephanie Stockar"
      ],
      "abstract": "This paper presents a method for controlling traffic via the use of connected and automated vehicles (CAVs) acting as moving bottlenecks. Current methods for moving bottleneck control use a couple PDE-ODE model, based on the Lighthill-Whitham-Richard (LWR) model, to represent the influence of the CAV. Control of the CAV is normally achieved by designing the control on the ODE which models the speed of the moving bottleneck. The proposed method in this paper instead looks to reduce the computational burden of controlling multiple CAVs by designing the moving bottleneck controller first on the PDE. The original control designed on the PDE is a linear quadratic regulator (LQR) that determines the optimal variable speed limit (VSL) for the entire length of freeway in order to regulate density to a desired setpoint. Then, a continuification approach is utilized to determine the input speed for each CAV. Results show that multiple CAVs can be controlled via this method, with minimal computational burden, and that as the number of CAVs increases the solution approaches the global optimal solution determined by the LQR.",
      "short_abstract": "This paper presents a method for controlling traffic via the use of connected and automated vehicles (CAVs) acting as moving bottlenecks. Current methods for moving bottleneck control use a couple PDE-ODE model, based on the Lighthill-Whitham-Richard (LWR)...",
      "published": "2026-06-08",
      "updated": "2026-06-08",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 72,
      "arxiv_url": "https://arxiv.org/abs/2606.09534",
      "pdf_url": "https://arxiv.org/pdf/2606.09534.pdf"
    },
    {
      "id": "2606.08684",
      "anchor": "paper-2606-08684",
      "title": "BLUE: Toward Better Language Use in Efficient Vision-Language-Action Models for Autonomous Driving",
      "authors": [
        "George Ling",
        "Lijin Yang",
        "Hao Yang",
        "Zhongzhan Huang"
      ],
      "abstract": "We present BLUE, a minimal method for better language use in vision-language-action (VLA) models for autonomous driving (AD). Through extensive analysis, we reveal that language matters on only a small fraction of routes, but on those routes it can greatly improve or degrade performance. Generating language at every frame is therefore inefficient, since most computation is spent on frames that do not benefit from language. We further show that pretrained VLA hidden states potentially already encode whether language will benefit a given frame, even though scene complexity and kinematic features alone struggle to predict this. Based on this finding, BLUE trains a lightweight gate on frozen VLA hidden states to decide per frame whether to activate language generation or predict actions directly, without modifying the backbone or requiring additional human annotation. With just a 0.11M-parameter gate, BLUE sets a new state of the art on both benchmarks, achieving 76.2% success rate on Bench2Drive and 36 driving score on Longest6 v2, while delivering 2.54x inference speedup and 8.9% success rate improvement over the backbone. BLUE provides a practical path toward efficient language-augmented AD, showing that VLA models can retain the benefits of language at a fraction of the cost. Our code, data, logs and checkpoints are fully available on https://github.com/George-Ling3/BLUE.",
      "short_abstract": "We present BLUE, a minimal method for better language use in vision-language-action (VLA) models for autonomous driving (AD). Through extensive analysis, we reveal that language matters on only a small fraction of routes, but on those routes it can greatly...",
      "published": "2026-06-07",
      "updated": "2026-06-07",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 24,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 73,
      "arxiv_url": "https://arxiv.org/abs/2606.08684",
      "pdf_url": "https://arxiv.org/pdf/2606.08684.pdf"
    },
    {
      "id": "2606.08680",
      "anchor": "paper-2606-08680",
      "title": "Distortion-Aware PETR for BEV Object Detection with Mixed Pinhole-Fisheye Cameras",
      "authors": [
        "Xiangzhong Liu"
      ],
      "abstract": "Fisheye cameras are widely deployed in autonomous driving perception suites for their low cost and full-coverage field of view (FOV), yet their potential remains underleveraged in 3D object detection. Severe radial distortion challenges most BEV detectors by violating the fundamental assumption of uniform sampling. To bridge this gap, we propose Distortion-Aware PETR (DAPETR), a projection-free detector tailored for mixed pinhole-fisheye camera setups. DAPETR incorporates two key learned-adaptive modules: a unified distortion-aware positional embedding that harmonizes positional encodings for image representations with fisheye geometry, and a bidirectional feature-geometry co-modulation module that mutually adapts image features and 3D positional embeddings. In our experiments on a converted KITTI-360 benchmark, we systematically compare our learned adaptive approach against PETR in polar coordinates (PolarPETR). We find that while both methods improve over the baseline, our learned modules achieve superior performance. Crucially, we uncover a negative interaction when combining both strategies, revealing that learned adaptation and explicit geometric reparameterization can conflict. Our final DAPETR model significantly advances the research and benchmark for fisheye BEV detection, providing critical insights into effective distortion-aware 3D perception design other than image rectification.",
      "short_abstract": "Fisheye cameras are widely deployed in autonomous driving perception suites for their low cost and full-coverage field of view (FOV), yet their potential remains underleveraged in 3D object detection. Severe radial distortion challenges most BEV detectors by...",
      "published": "2026-06-07",
      "updated": "2026-06-07",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 35,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "title: camera",
          "abstract: camera",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 73,
      "arxiv_url": "https://arxiv.org/abs/2606.08680",
      "pdf_url": "https://arxiv.org/pdf/2606.08680.pdf"
    },
    {
      "id": "2606.08525",
      "anchor": "paper-2606-08525",
      "title": "DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving",
      "authors": [
        "Qimao Chen",
        "Fang Li",
        "Yuechen Luo",
        "Zehan Zhang",
        "Haiyang Sun",
        "Fangzhen Li",
        "Bing Wang",
        "Guang Chen",
        "Yang Ji",
        "Jiong Deng",
        "Hongwei Xie",
        "Hangjun Ye",
        "Long Chen",
        "Yi Zhang"
      ],
      "abstract": "Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such rewards typically relies on hand-crafted rule-based objectives or perception ground truth, which hinders generalization for data-scaling. While Vision-Language Models (VLMs) have demonstrated feasibility as reward models in other domains, their effectiveness in driving tasks remains underexplored. In this work, we bridge this gap by (1) introducing DriveReward, a reasoning trajectory evaluation dataset rigorously labeled via temporally-grounded visual guidance, and augmented with counterfactual driving behaviors., (2) alongside a specialized Vision-Language Reward Model. To address the scarcity of failure cases in conventional datasets, we propose a counterfactual data annotation scheme to construct cases encompassing diverse driving styles and erroneous behaviors. Evaluations on our proposed benchmark reveal that even leading open-source and proprietary VLMs fail to excel across all tasks, highlighting significant room for improvement in existing models. Building on these findings, we subsequently tailor a specialized 1B reward model that outperforms larger VLMs on task-specific reward alignment. Finally, we validate our reward model's effectiveness by integrating it into RL finetuning and multi-modal trajectory scoring across multiple baselines, achieving performance comparable to rule-based reward calculations in both open-loop and closed-loop evaluation.",
      "short_abstract": "Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such rewards typically relies on hand-crafted rule-based objectives or perception ground truth, which hinders...",
      "published": "2026-06-07",
      "updated": "2026-06-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Safety, Security & Verification: abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 73,
      "arxiv_url": "https://arxiv.org/abs/2606.08525",
      "pdf_url": "https://arxiv.org/pdf/2606.08525.pdf"
    },
    {
      "id": "2606.08470",
      "anchor": "paper-2606-08470",
      "title": "LUNA-AD: Lightweight Uncertainty-Aware Language Model with Lifelong Learning for Autonomous Driving",
      "authors": [
        "Ruoyu Yao",
        "Pei Liu",
        "Ruiguo Zhong",
        "Mingxing Peng",
        "Rui Yang",
        "Jun Ma"
      ],
      "abstract": "While large language models (LLMs) offer promising reasoning capabilities, their integration into safety-critical driving systems is hindered by limited reasoning diversity, high computational overhead, and static learning paradigms. To address these challenges, we propose LUNA-AD, a lightweight uncertainty-aware language model with lifelong learning for autonomous driving (AD). LUNA-AD features a tri-system architecture that reconciles complex multimodal behavioral reasoning, efficient deployment, and continual refinement. We design a multi-agent analytical system to generate uncertainty-aware decision-making demonstrations through diverse hypothesis exploration. A dual-head lightweight heuristic model is distilled to unify the inference of decision distributions and textual explanations while enabling efficient deployment. Furthermore, a reflection-driven lifelong learning mechanism operates on multimodal decision outputs and preserves strategic diversity, allowing for the refinement of candidate decisions and rationales via closed-loop feedback to enhance driving robustness. Extensive experiments on nuPlan benchmarks demonstrate that LUNA-AD achieves state-of-the-art success rates under both non-reactive and reactive modes, with drastically reduced inference latency compared to existing knowledge-driven AD frameworks.",
      "short_abstract": "While large language models (LLMs) offer promising reasoning capabilities, their integration into safety-critical driving systems is hindered by limited reasoning diversity, high computational overhead, and static learning paradigms. To address these...",
      "published": "2026-06-07",
      "updated": "2026-06-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: robustness",
          "title: uncertainty",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 73,
      "arxiv_url": "https://arxiv.org/abs/2606.08470",
      "pdf_url": "https://arxiv.org/pdf/2606.08470.pdf"
    },
    {
      "id": "2606.20648",
      "anchor": "paper-2606-20648",
      "title": "Platooning Connected, Autonomous, and Human-Driven Vehicles: A Deep Reinforcement Learning-based Approach",
      "authors": [
        "Zhen Qina",
        "Dong-Fan Xie",
        "Heng Ma",
        "Xiaomei Zhao",
        "Zhengbing He"
      ],
      "abstract": "Conventionally, existing vehicle platooning approaches are designed for connected vehicles, typically including connected autonomous vehicles and connected human-driven vehicles. Non-connected vehicles, such as non-connected autonomous or human-driven vehicles, are not incorporated. As a result, these platooning approaches may not properly reflect real-world mixed traffic conditions at the current stage. To address this limitation, this study proposes a hybrid platooning pattern that conditionally permits non-connected vehicles to join platoons, thereby enhancing platooning diversity and flexibility. However, it was found that the unregulated integration of non-connected vehicles can trigger rapid platoon expansion, significantly amplifying the risk of disturbance propagation in traffic flow. This, in turn, exacerbates the inherent conflict between traffic throughput and stability. To mitigate these challenges, this paper further develops a hybrid platooning control strategy based on deep reinforcement learning (DRL). This strategy integrates vehicle dynamics, platoon topology, and traffic flow states through a multi-level state representation network, enabling a dynamic trade-off between traffic capacity and stability. Numerical simulations demonstrate that the proposed strategy effectively suppresses velocity disturbance propagation by dynamically optimizing platoon structures, thereby significantly enhancing the stability and safety of mixed traffic while reducing fuel consumption and emissions.",
      "short_abstract": "Conventionally, existing vehicle platooning approaches are designed for connected vehicles, typically including connected autonomous vehicles and connected human-driven vehicles. Non-connected vehicles, such as non-connected autonomous or human-driven...",
      "published": "2026-06-07",
      "updated": "2026-06-07",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: vehicle dynamics",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 73,
      "arxiv_url": "https://arxiv.org/abs/2606.20648",
      "pdf_url": "https://arxiv.org/pdf/2606.20648.pdf"
    },
    {
      "id": "2606.08880",
      "anchor": "paper-2606-08880",
      "title": "Direct Data-driven Predictive Control: A Computationally Efficient Alternative to DeePC for Eco-driving in Mixed Traffic Flows",
      "authors": [
        "Dongjun Li",
        "Haoxuan Dong",
        "Liangcai Xu",
        "Ziyou Song"
      ],
      "abstract": "Improving energy efficiency in the transportation sector is critical for achieving sustainable mobility, with eco-driving emerging as a key strategy. However, implementing effective eco-driving for connected and automated vehicles (CAVs) in mixed traffic presents a significant control challenge due to the heterogeneous, uncertain behavior of human-driven vehicles (HDVs). Data-enabled Predictive Control (DeePC) offers a promising model-free approach but is often hindered by a high computational burden, limiting its real-time feasibility. This paper introduces a novel Direct Data-driven Predictive Control (D3PC) framework to address this limitation. By reformulating the data-driven prediction mechanism, the D3PC significantly reduces computational complexity, making its computation time nearly invariant to historical data size. This computational efficiency directly enables the formulation of a sophisticated eco-driving controller that can solve the complex energy optimization problem in real time, even within diverse and stochastic mixed-traffic environments. Comprehensive simulations demonstrate that the D3PC is orders of magnitude faster than existing DeePC-based methods while achieving superior energy efficiency. Specifically, it reduces total platoon energy consumption by up to 10.71% compared to rule-based cruise control baselines and 3.80% compared to the original DeePC, confirming its effectiveness for real-time, energy-efficient control.",
      "short_abstract": "Improving energy efficiency in the transportation sector is critical for achieving sustainable mobility, with eco-driving emerging as a key strategy. However, implementing effective eco-driving for connected and automated vehicles (CAVs) in mixed traffic...",
      "published": "2026-06-07",
      "updated": "2026-06-07",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 4,
        "evidence": [
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 73,
      "arxiv_url": "https://arxiv.org/abs/2606.08880",
      "pdf_url": "https://arxiv.org/pdf/2606.08880.pdf"
    },
    {
      "id": "2606.08197",
      "anchor": "paper-2606-08197",
      "title": "AlignFed: Alignment-Aware Asynchronous Federated Fine-Tuning for Large Language Models in Heterogeneous Edge Environments",
      "authors": [
        "Yan Wang",
        "Ziyi Gao",
        "Rui Wang"
      ],
      "abstract": "Large Language Models (LLMs) have significantly propelled the advancement of edge intelligence and have been widely deployed across various scenarios, including autonomous driving, industrial inspection, and personalized IoT services. However, the collaborative adaptation of LLMs on edge devices continues to face formidable challenges due to strict data privacy constraints, highly heterogeneous computing and communication resources, and the non-independent and identically distributed (non-IID) nature of local data. Federated Fine-Tuning (FFT) enables the collaborative optimization of distributed models without exposing raw data. Yet, traditional synchronous aggregation suffers from a severe straggler effect, resulting in high system latency and low resource utilization. Existing asynchronous federated learning methods are predominantly designed for small-to-medium-scale models and struggle to address the specific challenges inherent in LLM fine-tuning namely, model drift caused by stale updates, aggravated client drift stemming from data heterogeneity, and aggregation fairness imbalance resulting from the dominance of fast clients. To address these issues, this paper proposes AlignFed, an asynchronous federated fine-tuning framework for LLMs tailored to heterogeneous edge environments. AlignFed employs a lightweight multi-stage semantic alignment mechanism comprising three core modules: version-aware update grouping, cross-version semantic alignment based on a mini-batch calibration set, and fairness-aware aggregation that integrates both update freshness and client participation frequency. This framework effectively mitigates cross-version model drift and client drift while enhancing aggregation fairness, thereby achieving stable and efficient asynchronous federated optimization in scenarios characterized by high heterogeneity and significant update staleness.",
      "short_abstract": "Large Language Models (LLMs) have significantly propelled the advancement of edge intelligence and have been widely deployed across various scenarios, including autonomous driving, industrial inspection, and personalized IoT services. However, the...",
      "published": "2026-06-06",
      "updated": "2026-06-06",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 7,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 74,
      "arxiv_url": "https://arxiv.org/abs/2606.08197",
      "pdf_url": "https://arxiv.org/pdf/2606.08197.pdf"
    },
    {
      "id": "2606.08170",
      "anchor": "paper-2606-08170",
      "title": "Learning from Human Driving: A Human-in-the-Loop Online Behavior Cloning Framework for Autonomous Driving",
      "authors": [
        "Yuhong Shi",
        "Jianyi Liu",
        "Lihang Sun",
        "Li Li",
        "Xudong Dong"
      ],
      "abstract": "With the evolution of large foundation models (LFMs), data-driven autonomous driving has made significant strides. However, existing paradigms still face severe challenges in complex interaction and long-tail scenarios due to distribution shift and causal confusion. These limitations often result in a lack of human-level decision-making flexibility and safety in extreme conditions. To overcome this limitation, this paper proposes a Human-in-the-Loop Online Behavior Cloning frame work (HiL-OBC) for autonomous driving, which aims to deeply integrate the cross-modal perceptual capabilities of LFMs with the high-level driving intelligence of human experts. Specifically, HiL-OBC deployment is executed through three critical phases: policy initialization with human intervention, latent behavioral modeling with Bayesian policy adaptation, and online deploy ment and updates. Furthermore, we design a Multi-modal Online Behavior Cloning (MOBC) model, which optimizes the base driving policy online through a lightweight network architecture, a takeover trigger mechanism, and a multi-variant loss function, thereby enhancing the system's decision-making robustness in complex environments. We evaluated the HiL-OBC on the LangAuto-Human CARLA benchmark. Experimental results demonstrate that the driving policies optimized via the human-in-the-loop mechanism achieve substantial performance gains: the DS of StructNav, LFG, and LMDrive increased by 47.25%, 31.59%, and 32.12%, respectively, with a simultaneous of various experimental settings and key components highlights the advantages of human-in-the-loop learning in improving decision-making robustness and overall driving performance.",
      "short_abstract": "With the evolution of large foundation models (LFMs), data-driven autonomous driving has made significant strides. However, existing paradigms still face severe challenges in complex interaction and long-tail scenarios due to distribution shift and causal...",
      "published": "2026-06-06",
      "updated": "2026-06-06",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Imitation Learning",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 23,
        "evidence": [
          "title: human-machine interaction",
          "abstract: human-machine interaction",
          "abstract: takeover",
          "abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 74,
      "arxiv_url": "https://arxiv.org/abs/2606.08170",
      "pdf_url": "https://arxiv.org/pdf/2606.08170.pdf"
    },
    {
      "id": "2606.08173",
      "anchor": "paper-2606-08173",
      "title": "AI-Native Closed-Loop Security for 6G-Enabled Cyber-Physical Systems: From Edge Detection to Network-Wide Mitigation",
      "authors": [
        "Bilal Hussain",
        "Muhammad Bilal",
        "Tan Li",
        "Haris Pervaiz",
        "Xiao Tang",
        "Qinghe Du",
        "Fawad Ahmad",
        "Muhammad Azhar",
        "Jun Zhang"
      ],
      "abstract": "In sixth-generation (6G) networks, billions of cyber-physical systems (CPSs) - autonomous vehicles, smart grids, industrial robots, and remote-surgical equipment - will run over ultra-reliable low-latency slices, collapsing the gap between a remote breach and physical harm to milliseconds, a budget perimeter firewalls and centralised security operations centres cannot meet. This survey reframes 6G CPS security as a closed-loop, AI-native pipeline that senses at the multi-access edge computing (MEC) tier, using minute-scale call-detail records (CDRs) for baseline learning and sub-millisecond RAN/Open-RAN (O-RAN) telemetry for the latency-critical path. It decides locally with compressed deep models, mitigates network-wide via SDN, NFV, and O-RAN controllers, and retrains through federated learning (FL) and digital-twin (DT) replay. We formalise a per-slice, tail-bounded latency contract on the sense, detect, and mitigate stages, enforced at a slice-dependent tail percentile (p99 for safety-critical URLLC slices). Organising 128 peer-reviewed studies (2017-2026) under a PRISMA 2020 protocol, we (i) map the 6G/CPS threat surface to MITRE ATT&CK and a CDR-observable feature space; (ii) unify edge anomaly detection and DDoS classification across twelve datasets and statistical, graph, and transformer models; (iii) synthesise SDN/NFV/O-RAN primitives into one closed-loop reference architecture; (iv) treat FL, large language models (LLMs), DT, post-quantum cryptography (PQC), zero-trust architecture (ZTA), and explainable AI as cross-cutting enablers, not parallel pillars; and (v) consolidate open problems into five directions spanning data, latency, trust, standardisation, and evaluation.",
      "short_abstract": "In sixth-generation (6G) networks, billions of cyber-physical systems (CPSs) - autonomous vehicles, smart grids, industrial robots, and remote-surgical equipment - will run over ultra-reliable low-latency slices, collapsing the gap between a remote breach and...",
      "published": "2026-06-06",
      "updated": "2026-06-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Survey",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 14,
        "evidence": [
          "abstract: safety",
          "title: security",
          "abstract: security",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 74,
      "arxiv_url": "https://arxiv.org/abs/2606.08173",
      "pdf_url": "https://arxiv.org/pdf/2606.08173.pdf"
    },
    {
      "id": "2606.08136",
      "anchor": "paper-2606-08136",
      "title": "Learning Predictive Control with Deep Koopman Operators for Autonomous Vehicle Motion Planning",
      "authors": [
        "Xinglong Zhang",
        "Yongqian Xiao",
        "Haotian Cao",
        "Xing Zhou",
        "Xin Yin",
        "Xin Xu"
      ],
      "abstract": "Model Predictive Control (MPC) is widely used for autonomous-vehicle (AV) motion planning, but its real-time applicability is often limited by the need for accurate models and online solution of nonlinear, nonconvex optimization problems in dynamic road environments. Actor-critic reinforcement learning offers a promising alternative for online policy generation, yet its policy-learning process often lacks explicit control-theoretic structure. This article proposes a learning predictive control (LPC) framework with deep Koopman operators for efficient real-time motion planning under nonconvex constraints. To address nonlinear and uncertain vehicle dynamics, a deep-Koopman-based predictor is used to lift the system into an interpretable linear observable space in a data-driven manner. Unlike traditional MPC, which computes open-loop control sequences, the proposed LPC framework yields a closed-loop state-feedback policy within each prediction interval through receding-horizon actor-critic learning. To ensure safety under nonconvex environmental constraints, LPC constructs convex local surrogate representations of obstacles and defines corresponding potential-field functions. These functions and their gradients are directly embedded into the actor-critic structure, enabling efficient, safety-aware policy learning. Extensive simulations and real-world experiments on the HongQi-EHS3 platform demonstrate favorable performance in diverse obstacle-avoidance scenarios in terms of safety, computational efficiency, and driving comfort, compared with benchmark methods such as CBF-MPC and LMPCC.",
      "short_abstract": "Model Predictive Control (MPC) is widely used for autonomous-vehicle (AV) motion planning, but its real-time applicability is often limited by the need for accurate models and online solution of nonlinear, nonconvex optimization problems in dynamic road...",
      "published": "2026-06-06",
      "updated": "2026-06-06",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 14,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 74,
      "arxiv_url": "https://arxiv.org/abs/2606.08136",
      "pdf_url": "https://arxiv.org/pdf/2606.08136.pdf"
    },
    {
      "id": "2606.07697",
      "anchor": "paper-2606-07697",
      "title": "TianJi-Environ: An Autonomous AI Scientist for Atmospheric Environmental Research",
      "authors": [
        "Haoluo Zhao",
        "Hongchun Zhang",
        "Nan Li",
        "Jing-Jia Luo",
        "Kaikai Zhang",
        "Mengyang Yu",
        "Nan Chen",
        "Tao Song",
        "Fan Meng"
      ],
      "abstract": "As atmospheric environmental prediction continues to improve, interpretable validation of pollution mechanisms and feedback processes has become a main challenge in atmospheric chemistry. Yet mechanism validation based on complex numerical models still relies heavily on expert knowledge: mechanistic hypotheses must be operationalized into executable experiments, and model outputs must be organized into traceable evidence. We present TianJi-Environ, an auditable AI Scientist for atmospheric-chemistry mechanism validation. TianJi-Environ establishes the first WRF-Chem-based multi-agent framework that autonomously drives complex atmospheric-chemistry simulations, converting mechanistic hypotheses into executable configurations, testing experiments, and evidence criteria. Using ozone response and particulate-matter feedback as two representative examples, we demonstrate TianJi-Environ's capability for mechanism validation. In a summertime ozone case over the North China Plain, the system detects directionally consistent aerosol-radiation-interaction signals in shortwave radiation and boundary-layer height, but judges the evidence for ozone response to NOx control to be incomplete. In a wintertime PM2.5 case over the Guanzhong Basin, it localizes the unsupported link to insufficient propagation from black-carbon perturbation to particulate response and missing diagnostics of vertical absorptive heating. These results show that TianJi-Environ makes expert-driven mechanism validation explicit, structured, and auditable, offering a reproducible paradigm for multi-agent systems coupled with complex atmospheric-chemistry models.",
      "short_abstract": "As atmospheric environmental prediction continues to improve, interpretable validation of pollution mechanisms and feedback processes has become a main challenge in atmospheric chemistry. Yet mechanism validation based on complex numerical models still relies...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07697",
      "pdf_url": "https://arxiv.org/pdf/2606.07697.pdf"
    },
    {
      "id": "2606.07464",
      "anchor": "paper-2606-07464",
      "title": "Planning-aligned Token Compression for Long-Context Autonomous Driving",
      "authors": [
        "Zhixuan Liang",
        "Yuxiao Chen",
        "Yurong You",
        "Peter Karkus",
        "Wenhao Ding",
        "Boyi Li",
        "Alexander Popov",
        "Yan Wang",
        "Maximilian Igl",
        "Yiming Li",
        "Danfei Xu",
        "Nikolai Smolyanskiy",
        "Boris Ivanovic",
        "Ping Luo",
        "Marco Pavone"
      ],
      "abstract": "Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences that quickly exceed real-time computational budgets when encoding extended temporal context for complex interactions. While approaches like linear transformers and external memory try to make the context lightweight, token compression is most compatible with the architecture as it requires no backbone modifications. Yet existing compression adopts rule-based heuristics like temporal decay, decoupled from planning, risking loss of decision-critical information. We propose COMPACT-VA, a planning-aligned working memory framework built on conditional VQ-VAE, compressing extended context into bounded representations. Compression is conditioned on both historical trajectory and a learned planning intent that the posterior encoder distills from future trajectories during training, while the prior encoder learns to predict it from compressed observations. The compressed memory, concatenated with the predicted latent, feeds the policy for end-to-end optimization, planning with retained decision-critical information. We evaluate on high-signal dynamic scenarios where historical context is most critical for behavior correctness (e.g., stop, yield, or proceed), and accordingly design behavioral metrics. Under comparable token budgets, we achieve $>$6% improvement (68.3%) on success rates with consistent gains across metrics. Ablations validate planning-aligned coupling effectiveness. Closed-loop evaluation confirms that COMPACT-VA maintained general driving performance with 3.3* speedup and 2.7* memory reduction over uncompressed processing.",
      "short_abstract": "Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences that quickly exceed real-time computational budgets when encoding extended temporal context for complex interactions....",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07464",
      "pdf_url": "https://arxiv.org/pdf/2606.07464.pdf"
    },
    {
      "id": "2606.07186",
      "anchor": "paper-2606-07186",
      "title": "A Causal Probabilistic Framework for Perception-Informed Closed-Loop Simulation of Autonomous Driving",
      "authors": [
        "Zhennan Fei",
        "Rickard Johansson",
        "Mikael Andersson",
        "Matthias Eng",
        "Mattias Eriksson",
        "Kaveh Kianfar",
        "Sadegh Rahrovani",
        "Chris van der Ploeg",
        "Michael Borth",
        "Maren Buermann",
        "Michiel Braat",
        "Henk Goossens",
        "Zijian Han",
        "Majid Khorsand Vakilzadeh",
        "Gabriel Rodrigues de Campos"
      ],
      "abstract": "Software-in-the-loop (SIL) simulation is a cornerstone for the validation of modern automotive safety functions. However, many current frameworks utilize ideal sensing, which bypasses the functional insufficiencies of perception algorithms, leading to over-optimistic safety assessments. This paper proposes a perception-informed SIL testing methodology that bridges the gap between ground-truth simulation and real-world perception behavior. We present a framework for incorporating causal probabilistic models into standardized, scenario-based simulation toolchains, applicable to both Advanced Driver Assistance Systems (ADAS) and Autonomous Driving Systems (ADS). Our approach enables the systematic injection of realistic perception errors, such as loss of detection, sizing inaccuracies, and positioning offsets, derived from physical triggering conditions like fog, rain, and object-merging scenarios. By evaluating these ``faults'' within a standardized simulation environment, we demonstrate that perception-informed testing reveals latent operational risks that ideal SIL environments fail to capture, providing a scalable pathway for SOTIF (ISO 21448) validation.",
      "short_abstract": "Software-in-the-loop (SIL) simulation is a cornerstone for the validation of modern automotive safety functions. However, many current frameworks utilize ideal sensing, which bypasses the functional insufficiencies of perception algorithms, leading to...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 13,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: SOTIF"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07186",
      "pdf_url": "https://arxiv.org/pdf/2606.07186.pdf"
    },
    {
      "id": "2606.07170",
      "anchor": "paper-2606-07170",
      "title": "Test-Time Trajectory Optimization for Autonomous Driving",
      "authors": [
        "Yihong Xu",
        "Eloi Zablocki",
        "Yuan Yin",
        "Elias Ramzi",
        "Ellington Kirby",
        "Alexandre Boulch",
        "Matthieu Cord"
      ],
      "abstract": "End-to-end planners for autonomous driving typically generate a set of candidate trajectories, score each one, and return the highest-scoring candidate. However, the scorer is applied only after the proposals are generated and cannot influence the set of trajectories: a weak set of candidates limits planning performance regardless of the scorer's quality. We instead treat the scorer as a learned trajectory-level reward function and search for trajectories that maximize it. Our method, TOAD, runs the Cross-Entropy Method at test time, warm-started from the planner's proposals. It requires no retraining and is plug-and-play for existing planners. Across six base planners, TOAD improves results on NAVSIM-v1 (94.7 PDMS), NAVSIM-v2 (56.3 EPDMS), and the closed-loop HUGSIM benchmark. The code will be made publicly available via the project page: https://valeoai.github.io/TOAD/.",
      "short_abstract": "End-to-end planners for autonomous driving typically generate a set of candidate trajectories, score each one, and return the highest-scoring candidate. However, the scorer is applied only after the proposals are generated and cannot influence the set of...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07170",
      "pdf_url": "https://arxiv.org/pdf/2606.07170.pdf"
    },
    {
      "id": "2606.07067",
      "anchor": "paper-2606-07067",
      "title": "Extending Responsibility-Sensitive Safety for the Assessment of Offloaded Autonomous Driving Services",
      "authors": [
        "Robin Dehler",
        "Aryan Thakur",
        "Michael Buchholz"
      ],
      "abstract": "Safety is a fundamental requirement in the development of autonomous driving (AD) systems. While function offloading has demonstrated significant benefits in terms of computational efficiency and energy consumption, its application to safety-critical AD functionality introduces new challenges. In particular, offloaded service compositions incur increased and variable response times due to wireless vehicle-to-everything (V2X) communication, which directly affects the vehicle's reaction time and thus its safety guarantees. In this paper, we address this challenge by extending the definitions of Responsibility-Sensitive Safety (RSS) to explicitly account for different response times of local and offloaded AD service compositions. Based on this extension, we propose an integration into function offloading, using the RSS safety constraints for offloading decision-making and fallback mechanisms. Offloaded service compositions are only permitted if the current traffic situation remains safe under the corresponding end-to-end response time. If this condition is violated, the system performs a controlled fallback to local execution. Furthermore, we introduce an enhanced fallback strategy that includes a warm-standby phase for offloaded services, enabling faster and safer transitions from offloaded to local services. The proposed approach is integrated into our AD stack and evaluated in both simulation and the real world. Experimental results demonstrate that the proposed method improves safety compared to state-of-the-art function offloading and safety frameworks, while preserving the benefits of distributed computation when safety conditions allow.",
      "short_abstract": "Safety is a fundamental requirement in the development of autonomous driving (AD) systems. While function offloading has demonstrated significant benefits in terms of computational efficiency and energy consumption, its application to safety-critical AD...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07067",
      "pdf_url": "https://arxiv.org/pdf/2606.07067.pdf"
    },
    {
      "id": "2606.06996",
      "anchor": "paper-2606-06996",
      "title": "Mission-Level Runtime Assurance Framework for Autonomous Driving",
      "authors": [
        "Chieh Tsai",
        "Salim Hariri"
      ],
      "abstract": "This paper studies runtime safety for autonomous driving when high-level driving commands become faulty or unreliable. Unlike conventional runtime-safety approaches that mainly focus on immediate vehicle safety, the proposed framework evaluates both driving safety and whether the vehicle can still successfully complete its mission before a command is executed. The framework extends highway-env with mission-level fault scenarios such as skipping required checkpoints, entering restricted areas, and generating future routes that can no longer complete the mission successfully. A runtime monitoring system is introduced to detect and reject unsafe or mission-infeasible commands before execution. For comparison, an adapted Simplex-Drive runtime-safety baseline with learning-based driving control, safety fallback control, and runtime controller switching is implemented using the public Simplex-Drive framework. Experimental results show that platform-level runtime safety alone cannot detect mission-level planning faults, while the proposed framework successfully rejects mission-infeasible commands and improves mission success under randomized fault conditions.",
      "short_abstract": "This paper studies runtime safety for autonomous driving when high-level driving commands become faulty or unreliable. Unlike conventional runtime-safety approaches that mainly focus on immediate vehicle safety, the proposed framework evaluates both driving...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: controller",
          "Control & Vehicle Dynamics: abstract: control",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.06996",
      "pdf_url": "https://arxiv.org/pdf/2606.06996.pdf"
    },
    {
      "id": "2606.07366",
      "anchor": "paper-2606-07366",
      "title": "Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos",
      "authors": [
        "Anurag Ghosh",
        "Francesco Pittaluga",
        "Khiem Vuong",
        "Angela Chen",
        "Juan Alvarez-Padilla",
        "Manmohan Chandraker",
        "Srinivasa Narasimhan"
      ],
      "abstract": "Self-driving simulations typically rely on data collected in a small number of cities or on hand-authored synthetic scenarios. Dashcam videos cover a far broader range of locations and situations, including rare or long-tailed scenarios. They are considered less usable for simulation because it is difficult to recover accurate 4D scenes from monocular in-the-wild videos. Work zones are one such class of long-tailed situations that dashcams capture. We present Dash2Sim, a framework that turns in-the-wild monocular dashcam videos into metric, geo-referenced 4D driving logs compatible with existing simulators, and verifies eachone against an independently maintained map without annotations. We apply Dash2Sim to a large video corpus to create the ROADWork4D benchmark dataset, which spans 4,244 scenes with 2.7M 3D objects across 17 cities. On a verified subset ROADWork4D-CL (2,201 scenes), we study privileged closed-loop planners and find that work zone scenarios are difficult: while rule-based and hybrid planners generalize better than learning-based ones, all fall short, failing to make the lane changes that temporary work zone channels require. Beyond planning, dense depth recovered by Dash2Sim improves novel-view synthesis quality by up to 19% on perceptual metrics, suggesting its potential to provide rich conditioning for closed-loop sensor simulation from monocular videos.",
      "short_abstract": "Self-driving simulations typically rely on data collected in a small number of cities or on hand-authored synthetic scenarios. Dashcam videos cover a far broader range of locations and situations, including rare or long-tailed scenarios. They are considered...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07366",
      "pdf_url": "https://arxiv.org/pdf/2606.07366.pdf"
    },
    {
      "id": "2606.07855",
      "anchor": "paper-2606-07855",
      "title": "Path Planning Using Deep Deterministic Policy Gradient: A Reinforcement Learning Approach",
      "authors": [
        "Qiang Le",
        "Yaguang Yang",
        "Isaac E. Weintraub"
      ],
      "abstract": "Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge because the problem is nonlinear and nonconvex even in simplest scenarios. While traditional optimal control methods can be used to find ideal paths, the computational time is often too slow for real-time decision-making. To solve this challenge, we propose a method based on Deep Deterministic Policy Gradient (DDPG) and model the threat as possibly multiple circular 'no-go' zones. A mission is regarded as a failure if the vehicle enters this restricted zone at any time or does not reach a neighborhood of the destination. The DDPG agent is trained through trial and error in a simulated environment, learning a direct mapping from its current state (position and heading) to a series of feasible actions that guide the agent to safely reach its destination. The reword function has three parts: (a) an attractive field centered at the final destination, (b) some repulsive fields centered at the origins of circular obstacles, and (c) a penalty of control energy consumption (the magnitude of heading change) that indirectly in favor for straight path. The DDPG trains the agent using these incentives to find the largest possible set of starting points wherein a safe path to the destination is guaranteed. This provides critical information for mission planning, showing beforehand whether a task is achievable from a given starting point, assisting pre-mission planning activities. The approach is validated in simulation. A comparison between the DDPG method and a traditional optimal control (pseudo-spectral) method is carried out. The results show that the learning-based agent produces effective paths while being significantly faster, making it a better fit for real-time applications.",
      "short_abstract": "Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge because the problem is nonlinear and nonconvex even in simplest scenarios. While traditional optimal control methods can be used to find ideal paths, the...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 15,
        "evidence": [
          "title: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07855",
      "pdf_url": "https://arxiv.org/pdf/2606.07855.pdf"
    },
    {
      "id": "2606.07437",
      "anchor": "paper-2606-07437",
      "title": "Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability",
      "authors": [
        "Chaitanya Shinde",
        "Hadi Hajieghrary",
        "Paul Schmitt",
        "Adam Shoemaker",
        "Bodo Seifert",
        "Steve Kenner"
      ],
      "abstract": "The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded in a human-driven vehicle paradigm. In the context of autonomous vehicles (AVs), the absence of a human driver necessitates revisiting these principles. This paper decomposes the Controllability placeholder into two auditable evidence dimensions of ISO 26262 by introducing two measurable sub-concepts: Transferability and Predictability. Transferability extends Controllability to capture AV systems' ability to hand off control to dedicated fallback safety mechanisms, while Predictability captures how easily external agents can anticipate AV behavior. Predictability is formally defined from human-robot interaction-inspired principles, and a mathematical framework is provided to quantify it. A designed-versus-achievable gap is introduced to distinguish architectural fallback claims from scene-conditioned achievable fallback capability. The proposed metrics align with ISO 26262 and ISO/PAS 21448 (SOTIF), rendering fallback and interaction claims falsifiable and traceable across ODD slices. These dimensions complement rather than replace existing standards, and the enhancements preserve the structure of ISO 26262 while extending its applicability to driverless automated systems operating at SAE Levels 4 and 5.",
      "short_abstract": "The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded in a human-driven vehicle paradigm. In the context of autonomous vehicles (AVs), the absence of a human...",
      "published": "2026-06-05",
      "updated": "2026-06-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "abstract: SOTIF"
        ]
      },
      "recency": "fresh",
      "age_days": 75,
      "arxiv_url": "https://arxiv.org/abs/2606.07437",
      "pdf_url": "https://arxiv.org/pdf/2606.07437.pdf"
    },
    {
      "id": "2606.06423",
      "anchor": "paper-2606-06423",
      "title": "RiskFlow: Fast and Faithful Safety-Critical Traffic Scenario Generation",
      "authors": [
        "Qi Lan",
        "Yining Tang",
        "Yu Shen",
        "Yi Zhou",
        "Yuhao Wei",
        "Jie Li",
        "Guofa Li"
      ],
      "abstract": "Safety-critical traffic scenario generation is essential for evaluating autonomous driving systems under rare but high-risk interactions. Existing diffusion-based methods offer strong controllability in closed-loop generation, but their iterative denoising process is computationally expensive and may accumulate sampling and guidance errors over long rollouts, causing unrealistic motion artifacts such as jitter, abnormal acceleration, and off-road behavior. To address these issues, we propose RiskFlow, a closed-loop safety-critical multi-agent traffic generation framework that formulates future trajectory generation as transport in the action space. Instead of relying on iterative denoising, RiskFlow learns an average velocity field over a finite interval to transform Gaussian action sequences into future acceleration and yaw-rate commands with a single forward pass, using a JVP-based objective for efficient and stable training. At test time, RiskFlow applies output-space guidance to the generated actions, steering selected critical agents toward risky interactions while regularizing off-road behavior, and reconstructs physically feasible trajectories through vehicle dynamics. Experiments on nuScenes with tbsim closed-loop evaluation show that RiskFlow achieves a strong adversariality-realism trade-off across multi-agent and long-horizon settings. Compared with representative baselines, RiskFlow consistently improves realism while maintaining competitive safety-critical generation capability, and substantially reduces inference time for evaluation.",
      "short_abstract": "Safety-critical traffic scenario generation is essential for evaluating autonomous driving systems under rare but high-risk interactions. Existing diffusion-based methods offer strong controllability in closed-loop generation, but their iterative denoising...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 9,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.06423",
      "pdf_url": "https://arxiv.org/pdf/2606.06423.pdf"
    },
    {
      "id": "2606.06396",
      "anchor": "paper-2606-06396",
      "title": "Risk Assessment of Autonomous Driving: Integrating Technical Failures, Ethical Dilemmas, and Policy Frameworks",
      "authors": [
        "Boyi Chen",
        "Shengqin Chu",
        "Zicheng Wang",
        "Brian Baetz",
        "Zhen Gao"
      ],
      "abstract": "Autonomous driving technology has the potential to reduce the large number of road traffic accidents caused by human error each year, but it also brings new types of risks that need to be evaluated from the aspects of technology, ethics and regulations. Based on public crash data from the National Highway Traffic Safety Administration (NHTSA), disengagement reports from the California Department of Motor Vehicles (DMV), the MIT Moral Machines dataset, and a comparative regulatory analysis of five jurisdictions, we have found that the main types of technical failure modes are perception and classification errors. These account for a relatively large proportion of the reported accidents, and it can be concluded that there are different ethical frameworks for autonomous vehicle decision-making, and inconsistent regulations in different areas increase the uncertainty of widespread application. Generally speaking, the problems of technology, ethics and regulation are closely related and need to be solved together. Therefore, this paper recommends a more adaptive and cooperative governance approach that combines engineering standards, ethical discussion, and institutional supervision.",
      "short_abstract": "Autonomous driving technology has the potential to reduce the large number of road traffic accidents caused by human error each year, but it also brings new types of risks that need to be evaluated from the aspects of technology, ethics and regulations. Based...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 32,
        "margin": 6,
        "evidence": [
          "title: ethics",
          "abstract: ethics",
          "title: law or policy",
          "abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.06396",
      "pdf_url": "https://arxiv.org/pdf/2606.06396.pdf"
    },
    {
      "id": "2606.06366",
      "anchor": "paper-2606-06366",
      "title": "Waypoints Matter: A Systematic Study for Sampling-Based Trajectory Planning",
      "authors": [
        "Josep M. Barbera",
        "Antonio Artuñedo",
        "Jorge Villagra"
      ],
      "abstract": "Real-time autonomous driving commonly relies on sampling-based trajectory planners that link candidate trajectories to target waypoints along the road centerline. The placement of these waypoints directly impacts both the existence and quality of feasible trajectories. Yet, its effect on planner performance remains largely unexplored. In this paper, we treat waypoint placement as a first-class design variable. We hold the trajectory primitive and candidate budget fixed, and systematically sweep three placement strategies (uniform spacing, an augmented Ramer-Douglas-Peucker variant (RDP*), and a novel curvature-conditioned allocation) across 449 configurations and five CommonRoad maps of increasing geometric complexity. Our results show that the nominal inter-waypoint spacing $d_s$ is the primary performance driver, with large differences in planner reliability attributed to placement alone. Uniform sampling at a well-tuned spacing matches or surpasses both RDP* and the centered curvature variant. The curvature variant offers a small but consistent advantage on geometrically complex roads under reliability-first and balanced weightings, while RDP* never outperforms uniform sampling. These findings suggest that $d_s$ should be treated as the dominant tuning parameter, with geometry-aware strategies reserved for curvature-rich corridors where feasibility is the limiting factor.",
      "short_abstract": "Real-time autonomous driving commonly relies on sampling-based trajectory planners that link candidate trajectories to target waypoints along the road centerline. The placement of these waypoints directly impacts both the existence and quality of feasible...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 21,
        "evidence": [
          "title: motion planning",
          "title: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.06366",
      "pdf_url": "https://arxiv.org/pdf/2606.06366.pdf"
    },
    {
      "id": "2606.06312",
      "anchor": "paper-2606-06312",
      "title": "Meridian: Metric-Semantic Primitive Matching for Cross-View Geo-Localization Beyond Urban Environments",
      "authors": [
        "Mason Peterson",
        "Qingyuan Li",
        "Yixuan Jia",
        "Fernando Cladera",
        "Carlos Nieto-Granda",
        "Camillo Jose Taylor",
        "Jonathan P. How"
      ],
      "abstract": "Successful robot automation requires accurate global localization to support repeatability, task planning, goal specification, and safe operation. However, reliable localization in GNSS-denied environments remains an open problem. Overhead aerial imagery offers a promising solution, but existing approaches primarily target structured urban environments and have been rarely demonstrated in unstructured natural terrain. Limitations of the state-of-the-art include a reliance on models trained for specific environments, as well as difficulty handling repetitive geometries and featureless landscapes commonly found in natural outdoor areas. To overcome these challenges, we present Meridian, a method for matching high-level metric-semantic primitives across aerial images and ground robot RGB-D camera data that achieves accurate global localization and generalizes well across diverse environments, all without any training or algorithmic fine-tuning on area-specific data. We formulate novel consistency metrics to estimate a distribution over robot submap poses and to reject outlier hypotheses in a robust pose graph optimization step for accurate robot trajectory estimation. We demonstrate that our algorithm can localize a ground robot across a wide variety of environments, including an autonomous driving dataset, a park and campus area, and a wilderness camp, with an average optimized trajectory error of 2.4 m over 19 km of ground traversal.",
      "short_abstract": "Successful robot automation requires accurate global localization to support repeatability, task planning, goal specification, and safe operation. However, reliable localization in GNSS-denied environments remains an open problem. Overhead aerial imagery...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 19,
        "evidence": [
          "title: localization",
          "abstract: localization",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.06312",
      "pdf_url": "https://arxiv.org/pdf/2606.06312.pdf"
    },
    {
      "id": "2606.06255",
      "anchor": "paper-2606-06255",
      "title": "RadiusFPS: Efficient Farthest Point Sampling on CPUs and GPUs via Spherical Voxel Pruning",
      "authors": [
        "Ziyang Yu",
        "Xiang Li",
        "Qiong Chang",
        "Jun Miyazaki"
      ],
      "abstract": "Point clouds are a primary sensory representation for robotic perception, underpinning LiDAR-based autonomous driving, simultaneous localization and mapping (SLAM), and navigation. Within these pipelines, Farthest Point Sampling (FPS) is the most well-known downsampling operator, as its uniform coverage preserves the geometric structure on which downstream perception relies. However, the large time complexity of classical FPS scales poorly with the million-point-per-second rates of modern 3D sensors, making it a dominant latency bottleneck that conflicts with the real-time and limited onboard compute budgets of robotic systems. Therefore, we propose RadiusFPS, an FPS acceleration framework based on spherical voxel pruning that preserves the standard FPS update rule under the same initialization and tie-breaking policy. By indexing the point cloud with spherical voxels, RadiusFPS derives a conservative geometric bound that prunes redundant distance computations in each iteration, complemented by a coordinate-wise point-skip test that removes residual updates. We further introduce RadiusFPS-G, a warp-level GPU implementation that fuses voxel selection, pruning, and distance update into memory-coalesced kernels, eliminating costly global-memory round-trips. On indoor (S3DIS, ScanNet) and outdoor LiDAR (SemanticKITTI) benchmarks, RadiusFPS-G attains up to 2.5x speedup over GPU-based FPS and matches or exceeds QuickFPS among the evaluated methods while using roughly half its GPU memory, with comparable segmentation accuracy. When coupled with the learning-based FastPoint sampler, the resulting pipeline achieves the fastest End-to-End inference among all evaluated configurations. These properties make high-quality FPS-style sampling practical for latency- and memory-constrained robotic vision.",
      "short_abstract": "Point clouds are a primary sensory representation for robotic perception, underpinning LiDAR-based autonomous driving, simultaneous localization and mapping (SLAM), and navigation. Within these pipelines, Farthest Point Sampling (FPS) is the most well-known...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 6,
        "evidence": [
          "abstract: localization",
          "abstract: SLAM",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.06255",
      "pdf_url": "https://arxiv.org/pdf/2606.06255.pdf"
    },
    {
      "id": "2606.06219",
      "anchor": "paper-2606-06219",
      "title": "CLEAR: Cognition and Latent Evaluation for Adaptive Routing in End-to-End Autonomous Driving",
      "authors": [
        "Yining Xing",
        "Zehong Ke",
        "Zhiyuan Liu",
        "Yanbo Jiang",
        "Wenhao Yu",
        "Jianqiang Wang"
      ],
      "abstract": "End-to-end autonomous driving models often struggle to balance multi-modal maneuver generation with real-time inference constraints. While diffusion models successfully capture diverse driving behaviors, their iterative denoising process incurs unacceptable latency for safety-critical deployment. To address this, we propose CLEAR (Cognition and Latent Evaluation for Adaptive Routing), a framework that combines ultra-fast generative planning with deep semantic reasoning. CLEAR employs Drive-JEPA as the visual encoder and replaces the multi-step denoising chain with a single-step conditional drift in a VAE latent space, introducing a conditioning coefficient to balance diversity and expert precision. Meanwhile, we fully fine-tune Qwen~3.5~0.8B on driving QA pairs to extract scene-aware hidden states. These states guide both an Adaptive Scheduler, which selects the conditioning coefficient $α$ and sample count $N$ from a discrete set of predefined schemes, and a cross-attention scorer that selects the optimal trajectory from candidates. On the NAVSIM v1 benchmark, CLEAR achieves a state-of-the-art PDMS of 93.7. Our results demonstrate that high-fidelity, multi-modal planning can be executed efficiently without dense geometric annotations or iterative sampling.",
      "short_abstract": "End-to-end autonomous driving models often struggle to balance multi-modal maneuver generation with real-time inference constraints. While diffusion models successfully capture diverse driving behaviors, their iterative denoising process incurs unacceptable...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 31,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.06219",
      "pdf_url": "https://arxiv.org/pdf/2606.06219.pdf"
    },
    {
      "id": "2606.06014",
      "anchor": "paper-2606-06014",
      "title": "PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models",
      "authors": [
        "Xiaoyun Qiu",
        "Jingtao He",
        "Yijie Chen",
        "Yusong Huang",
        "Haotian Wang",
        "Yixuan Wang",
        "Xinhu Zheng"
      ],
      "abstract": "Latent world models (LWMs) have strengthened end-to-end autonomous driving by forecasting compact scene dynamics for downstream planning. However, existing LWM-based planners usually generate trajectories directly from entangled latent representations. This compact latent-to-planner pathway lacks explicit modeling of risk, drivability, and diverse style preferences, making driving-style dynamics difficult to supervise, inspect, or modulate before a final trajectory is selected. We propose PLAN-S (PLANning with latent Style dynamics), a planner-facing bridge that addresses this compactness-controllability dilemma by decoding a style-conditioned, four-channel semantic cost map from the latent representation. The cost map is conditioned on ego state and driving style and is consumed up-stream of the planning decision through two host-side interfaces: attention-level fusion for regression planners and reward-level fusion for anchor-score planners. We validate PLAN-S on two architecturally distinct hosts, ResWorld on nuScenes and WoTE on NAVSIM, while keeping the host backbones frozen to isolate the contribution of the proposed bridge. On nuScenes, PLAN-S reduces L2 at every horizon over the baseline, with 0.55 m average L2 and a 42% relative reduction in the 3 s collision rate. On NAVSIM, the rule-cost variant reaches 89.4 Predictive Driver Model Score (PDMS), while the learned cost variant provides complementary gains on baseline-challenging scenes. Ablations show that the cost pathway contributes most directly to safer trajectory selection. Qualitative results further show that PLAN-S can produce diverse cost maps, with spatially consistent variations aligned to different driving styles.",
      "short_abstract": "Latent world models (LWMs) have strengthened end-to-end autonomous driving by forecasting compact scene dynamics for downstream planning. However, existing LWM-based planners usually generate trajectories directly from entangled latent representations. This...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 7,
        "evidence": [
          "abstract: forecasting",
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.06014",
      "pdf_url": "https://arxiv.org/pdf/2606.06014.pdf"
    },
    {
      "id": "2606.05774",
      "anchor": "paper-2606-05774",
      "title": "LiAuto-GeoX: Efficient Grounded Driving Transformer",
      "authors": [
        "Jiawei Lian",
        "Haoyi Sun",
        "Yang Wu",
        "Lifu Mu",
        "Siyuan Wang",
        "Le Hui",
        "Ning Mao",
        "Tao Wei",
        "Pan Zhou",
        "Kun Zhan",
        "Jian Yang"
      ],
      "abstract": "Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an open challenge. Existing large-scale visual geometry models typically require substantial computational resources and lack the long-range geometric fidelity, surround-view consistency, and real-time efficiency demanded by dynamic driving environments. To bridge this gap, we present \\textbf{LiAuto-GeoX}, an efficient grounded driving transformer designed for deployable, ego-centric 3D scene understanding. Our approach begins by learning a high-capacity driving geometry model from large-scale surround-view data, utilizing sparse LiDAR priors to provide robust geometric grounding in distant, ambiguous, or structure-sparse regions. We then instantiate this capability into a highly compact 155M-parameter onboard model through a novel geometry-preserving distillation framework. This framework employs mask-guided depth-aware distillation to retain fine-grained metric structures by emphasizing geometrically informative regions, and relative-pose relational distillation to enforce cross-view spatial consistency through pose-induced geometric relations. Extensive evaluations reveal that \\textbf{LiAuto-GeoX} runs at 220 FPS on KITTI while maintaining high-fidelity dense reconstruction, enabling real-time deployment. The learned geometry transfers seamlessly to downstream autonomy tasks, achieving 90.6 PDMS in trajectory prediction, 24.63 mIoU in occupancy prediction, and 47.67 IoU in future-frame prediction. These all demonstrate that efficient dense 3D reconstruction can transcend its traditional role as a perception target to serve as a scalable, foundational geometric representation for next-generation autonomous driving.",
      "short_abstract": "Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an open challenge. Existing large-scale visual geometry models typically require...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 11,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: occupancy",
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Perception & Sensor Fusion: abstract: scene understanding",
          "Prediction & World Models: abstract: trajectory prediction",
          "Prediction & World Models: abstract: future-frame prediction",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.05774",
      "pdf_url": "https://arxiv.org/pdf/2606.05774.pdf"
    },
    {
      "id": "2606.05677",
      "anchor": "paper-2606-05677",
      "title": "LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video",
      "authors": [
        "Shiqiang Lang",
        "Jing Liu",
        "Haoyang He",
        "Peiwen Sun",
        "Yuanteng Chen",
        "Tao Liu",
        "Lan Yang",
        "Longteng Guo",
        "Honggang Zhang"
      ],
      "abstract": "Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving and robotic navigation require more than recognizing the current view, as models must remember and retrieve previously observed spatial layouts, routes, viewpoint changes, and object states. To evaluate this capability, we introduce LongSpace-Bench, a room-tour video benchmark for long-horizon spatial memory, covering scene perception, spatial relations, and spatial memory. In this work, we further propose LongSpace, a memory framework for long-video spatial reasoning. LongSpace models long videos as sequential chunks, incorporates 3D structural cues into early decoder layers, and constructs layer-aware memory for question-guided retrieval. Experiments on multiple spatial reasoning benchmarks show that LongSpace improves long-video spatial understanding, further demonstrating explicit spatial memory as a key capability for long-horizon video MLLMs.",
      "short_abstract": "Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving and robotic navigation require more than recognizing the current view, as...",
      "published": "2026-06-04",
      "updated": "2026-06-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 10,
        "evidence": [
          "title: perception",
          "abstract: perception"
        ]
      },
      "recency": "fresh",
      "age_days": 76,
      "arxiv_url": "https://arxiv.org/abs/2606.05677",
      "pdf_url": "https://arxiv.org/pdf/2606.05677.pdf"
    },
    {
      "id": "2606.07657",
      "anchor": "paper-2606-07657",
      "title": "QDS-SNN: Energy-efficient Quantum Deeply-Supervised Spiking Neural Network Algorithm for Traffic Sign Recognition",
      "authors": [
        "Zhiguo Qu",
        "Keqi Li",
        "Le Sun",
        "Wenjie Liu",
        "Yimin Yu",
        "Saif Al-Kuwari",
        "Ahmed Farouk"
      ],
      "abstract": "Traffic sign recognition is crucial for intelligent transportation and autonomous driving, as it can improve driving efficiency and ensure road safety. However, traditional recognition methods are based on large datasets and intensive computation, which limits their real-time applicability. Spiking Neural Networks (SNNs) offer a biologically inspired, energy-efficient alternative due to their spatiotemporal processing capabilities, but suffer from information loss and vanishing gradients during training. To overcome these limitations, this study proposes a Quantum Deep-supervised Spiking Neural Network (QDS-SNN) that integrates Quantum Neural Networks (QNNs) for efficient, low-power deep supervision. Using quantum superposition and entanglement, QNNs enable expressive representations and parallel computation, thereby enhancing performance without compromising energy efficiency. The proposed QDS-SNN incorporates a temporally and spatially adaptive LIF (TSA-LIF) neuron and a quantum-assisted classifier module (QACM) to mitigate gradient issues and improve training effectiveness. This study conducts experiments on the PennyLane quantum simulation platform, and the results show that QDS-SNN achieves 99.72\\% accuracy on the GTSRB dataset in only 6 time steps -- outperforming the MS-ResNet baseline by 1.32\\% while reducing energy consumption by 55.77\\%. In the TSRD dataset, it achieves 97.90\\% accuracy while reducing energy use to 52.68\\% of the baseline. These results demonstrate that QDS-SNN offers a high-performance, energy-efficient solution for traffic sign recognition in intelligent transportation systems.",
      "short_abstract": "Traffic sign recognition is crucial for intelligent transportation and autonomous driving, as it can improve driving efficiency and ensure road safety. However, traditional recognition methods are based on large datasets and intensive computation, which...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 1,
        "evidence": [
          "title: traffic sign recognition",
          "abstract: traffic sign recognition"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.07657",
      "pdf_url": "https://arxiv.org/pdf/2606.07657.pdf"
    },
    {
      "id": "2606.05645",
      "anchor": "paper-2606-05645",
      "title": "Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning",
      "authors": [
        "Ziyang Yao",
        "Haochen Liu",
        "Yuncheng Jiang",
        "Zeyu Zhu",
        "Zibin Guo",
        "Jingru Wang",
        "Tianle Liu",
        "Jianwei Cui",
        "Kuiyuan Yang",
        "Hongwei Xie",
        "Jingwei Zhao",
        "Guang Chen",
        "Hangjun Ye"
      ],
      "abstract": "Autonomous driving requires reasoning about how ego actions shape future world evolution, rather than merely mapping observations to actions. However, most end-to-end methods rely on direct state-to-action imitation, while existing world models often remain weakly aligned with downstream policy generation. We introduce Discrete-WAM, a unified discrete vision-action world-policy framework that represents visual observations, future states, high-level decisions, and ego actions within a shared token space. Built on this discrete alignment, Discrete-WAM jointly trains world modeling, world-policy modeling, and policy modeling through multi-task and multi-stage pretraining, allowing action-conditioned future prediction to directly support policy generation. For downstream planning, Discrete-WAM further decomposes policy generation into hierarchical decision prediction and parallel action-token editing, where the decision token provides a high-level planning skeleton and confidence-based scheduling refines dense future actions efficiently. Experiments on large-scale autonomous-driving benchmarks show that Discrete-WAM achieves strong planning performance while supporting controllable future generation, counterfactual evaluation, surprise-based world-model analysis, and efficient parallel policy decoding. These results suggest that discrete representation alignment, unified world-policy training, and hierarchical token editing provide a promising design paradigm for physical AI.",
      "short_abstract": "Autonomous driving requires reasoning about how ego actions shape future world evolution, rather than merely mapping observations to actions. However, most end-to-end methods rely on direct state-to-action imitation, while existing world models often remain...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 7,
        "evidence": [
          "title: law or policy",
          "abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.05645",
      "pdf_url": "https://arxiv.org/pdf/2606.05645.pdf"
    },
    {
      "id": "2606.05461",
      "anchor": "paper-2606-05461",
      "title": "Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety",
      "authors": [
        "Abhinaw Priyadershi",
        "Mandar Pitale",
        "Jelena Frtunikj",
        "Maria Spence"
      ],
      "abstract": "Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantified interventional effects, named root-cause variables), yet the XAI literature is organised by output type and technique family (saliency maps, feature attribution, counterfactuals, causal graphs, language traces). SHAP, the most-recommended ADS XAI method, returns a ranked feature list that no implementation effort can convert into a directed chain (Fig.1). We name this mismatch the evidence-type gap. From AMLAS, ISO 26262, ISO21448, ISO/PAS 8800 we derive 19 testable evidentiary criteria across 7 lifecycle stages with representative clause-cited derivations and score six XAI method classes structurally. Causal XAI emerges as structurally required to satisfy the derived criteria at three stages: hazard identification (+62% rubric gap), incident investigation (+50%), and data management (+50%); the verdict set is stable across thresholds T in (0%, 50%]$ and survives a worst-case single-cell flip down to T = 25%. At the remaining four stages, correlational or language-based methods are comparable or sufficient. The rubric identifies structural admissibility (necessary but not sufficient for compliance): an admissible method's specific output content may still be wrong, and validating that fidelity (the edges a fitted SCM produces, the cause a trace names) is the open assurance challenge. A single-VLA proof of concept on 1,996 real-world driving clips (79,840 rows, ten splits) is consistent with each method's observed output type matching its rubric prediction. XAI method selection for ADS safety assurance should be driven by lifecycle-stage evidence demand, not by method popularity.",
      "short_abstract": "Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantified interventional effects, named root-cause variables), yet the XAI literature is organised by output type...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 10,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.05461",
      "pdf_url": "https://arxiv.org/pdf/2606.05461.pdf"
    },
    {
      "id": "2606.05159",
      "anchor": "paper-2606-05159",
      "title": "X4Val: Learning Neural Surrogates for Variance-Reduced Policy Evaluation",
      "authors": [
        "Rachel Luo",
        "Michael Watson",
        "Apoorva Sharma",
        "Heng Yang",
        "Han Qi",
        "Edward Schmerling",
        "Sushant Veer",
        "Boris Ivanovic",
        "Marco Pavone"
      ],
      "abstract": "Rigorous evaluation of learning-based robotic systems is an essential prerequisite for deployment. However, real-world test data is expensive to gather; moreover, in a typical iterative development context, data gathered from the latest policy is necessarily limited in scale. This motivates evaluation methodologies that make use of heterogeneous data sources, including simulation, historical policy logs, and data collected from related platforms or environments. While such auxiliary data are abundant and inexpensive, they are generally not directly representative of real-world outcomes -- for example, performance in simulation may differ substantially from performance in the real world -- making their principled use for high-confidence performance estimation challenging. In this paper, we introduce X4Val, a general framework for variance-reduced real-world metric estimation in the presence of non-paired, multi-domain data. X4Val embeds samples from real and auxiliary domains into a shared representation space and learns a transferable predictor of real-world metrics; this learned predictor is then incorporated into a control-variates estimator, enabling variance reduction even when paired samples are unavailable. We provide theoretical analysis and empirical evaluations on autonomous driving and real-world robot manipulation tasks, domains across which X4Val achieves up to 38.4% variance reduction and demonstrates consistent improvements over strong baselines. These results show that non-paired, heterogeneous data can be leveraged to substantially improve the sample efficiency of rigorous robotic system validation.",
      "short_abstract": "Rigorous evaluation of learning-based robotic systems is an essential prerequisite for deployment. However, real-world test data is expensive to gather; moreover, in a typical iterative development context, data gathered from the latest policy is necessarily...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: law or policy",
          "abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.05159",
      "pdf_url": "https://arxiv.org/pdf/2606.05159.pdf"
    },
    {
      "id": "2606.04884",
      "anchor": "paper-2606-04884",
      "title": "D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving",
      "authors": [
        "Renju Feng",
        "Rukang Wang",
        "Ning Xi",
        "Jianguo Yu",
        "Liping Lu",
        "Pan Zhou",
        "Duanfeng Chu"
      ],
      "abstract": "Traditional end-to-end autonomous driving frameworks frequently suffer from the \"style-averaging\" dilemma when trained on high-variance human demonstrations, yielding homogenized, style-uncontrollable, and even kinematically unsafe policies. To overcome this limitation, we present D$^3$-MoE (Dual Disentangled Diffusion Mixture-of-Experts), which disentangles trajectory modeling along two complementary axes. On the behavioral axis, generation is decoupled from selection: a style-conditioned diffusion process synthesizes multi-style candidate trajectories in parallel within a single scene, allowing a downstream module to select the optimal trajectory based on user preference or an evaluation score. On the physical axis, decoupled longitudinal and lateral routers activate their respective experts during inference time, trained without manual labels using self-supervised targets from orthogonal ground-truth kinematics. These activated experts, architected as Diffusion Transformers (DiT) and equipped with style-conditioned AdaLN and asymmetric lateral-fusion cross-attention, independently predict their corresponding physical state before being reassembled into a unified, kinematically coherent trajectory. Extensive evaluations on the challenging NAVSIM benchmark demonstrate that D$^3$-MoE achieves state-of-the-art planning performance, reaching 88.2 PDMS and 84.3 EPDMS by default. Moreover, our Best-of-Three ensemble strategy effectively broadens the multi-modal solution space, raising performance to 91.3 PDMS and 87.5 EPDMS. Both quantitative and qualitative analyses jointly confirm the framework's advantages in planning quality and style controllability.",
      "short_abstract": "Traditional end-to-end autonomous driving frameworks frequently suffer from the \"style-averaging\" dilemma when trained on high-variance human demonstrations, yielding homogenized, style-uncontrollable, and even kinematically unsafe policies. To overcome this...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 33,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.04884",
      "pdf_url": "https://arxiv.org/pdf/2606.04884.pdf"
    },
    {
      "id": "2606.04871",
      "anchor": "paper-2606-04871",
      "title": "Recent Advances and Trends in Learning-based 3D Representations",
      "authors": [
        "Adrien Schockaert",
        "Hamid Laga",
        "Hazem Wannous",
        "Vincent Magnier",
        "Guillaume Dufaye",
        "Jean-françois Witz"
      ],
      "abstract": "The selection of an appropriate 3D representation is a fundamental design decision that dictates the efficiency, quality, and capabilities of modern computer vision and graphics pipelines for tasks such as 3D reconstruction, novel-view synthesis and rendering, shape and motion analysis, recognition, and generation. While traditional representations (\\eg meshes, point clouds, and volumetric grids) remain standard outputs of 3D sensors (\\eg LiDAR and 3D scanners) and are widely used in downstream applications (\\eg editing and simulation), recent neural and primitive-based representations (\\eg 3D Gaussian Splatting) offer compact and differentiable alternatives opening a wide range of opportunities in applications such as games, AR/VR, autonomous driving, robot navigation, and medical imaging, to name a few. The goal of this paper is to survey the main families of 3D representations from discrete explicit formats to continuous implicit fields based either on neural rendering or primitive splatting. For each type of representation, we present the general formulation and its variants, discuss its benefits and limitations, and highlight key applications. We conclude the paper by outlining the open challenges and potential directions for future research. Distinct from recent surveys that broadly cover 3D object and scene reconstruction, this paper provides a focused analysis on the evolution of 3D representations themselves. We specifically emphasize the paradigm shift toward implicit representations, offering a novel perspective on how these emerging formats fundamentally alter 3D/4D workflows.",
      "short_abstract": "The selection of an appropriate 3D representation is a fundamental design decision that dictates the efficiency, quality, and capabilities of modern computer vision and graphics pipelines for tasks such as 3D reconstruction, novel-view synthesis and...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 4,
        "evidence": [
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.04871",
      "pdf_url": "https://arxiv.org/pdf/2606.04871.pdf"
    },
    {
      "id": "2606.04656",
      "anchor": "paper-2606-04656",
      "title": "Instance-Level Post Hoc Uncertainty Quantification in Object Detection",
      "authors": [
        "Chongzhe Zhang",
        "Zifan Zeng",
        "Qunli Zhang",
        "Feng Liu",
        "Zheng Hu"
      ],
      "abstract": "Object detection is a safety-critical component of autonomous driving. It is essential to quantify the uncertainty in bounding-box predictions for safety assurance. Post hoc uncertainty quantification without retraining aligns with real-world deployment requirements; therefore, we employ the Laplace approximation. Because instance-level uncertainty is needed, linearized inference methods that require multiple backpropagations are not time-efficient, and sampling-based methods are not fully post hoc. We propose Monte-Carlo generalized linearized model (MC-GLM), which provides instance-level and approximately post hoc uncertainty quantification. The number of samples required in the Monte Carlo step is constant and independent of the number of output instances, so it can be parallelized. Experiments on the nuScenes dataset with the CenterPoint detector validate the effectiveness of our method, and the resulting uncertainties exhibit good quality.",
      "short_abstract": "Object detection is a safety-critical component of autonomous driving. It is essential to quantify the uncertainty in bounding-box predictions for safety assurance. Post hoc uncertainty quantification without retraining aligns with real-world deployment...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.04656",
      "pdf_url": "https://arxiv.org/pdf/2606.04656.pdf"
    },
    {
      "id": "2606.04513",
      "anchor": "paper-2606-04513",
      "title": "MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation",
      "authors": [
        "Deguo Xia",
        "Zihan Li",
        "Haochen Zhao",
        "Dong Xie",
        "Yuyao Kong",
        "Xiyan Liu",
        "Jizhou Huang",
        "Mengmeng Yang",
        "Diange Yang"
      ],
      "abstract": "Lane-level maps are critical infrastructure for autonomous driving and lane-level navigation, yet constructing and maintaining standardized lane networks for hundreds of cities remains highly labor-intensive. Recent end-to-end vectorized mapping methods can predict lane geometry and topology directly from sensor data, but they typically treat mapping specifications and traffic regulations as implicit, dataset-dependent supervision. Moreover, in complex scenes (e.g., worn or missing markings and occlusions), correct lane configurations are often under-determined by visual evidence alone, making specification violations a major source of human post-editing. We propose MapAgent, an industrial-grade agentic architecture that augments a vectorization backbone for specification-compliant lane-map production. Rather than merely adding an agent loop to map prediction, MapAgent couples backbone perception with explicit specification verification, constraint-aware reasoning, and deterministic map editing under a bounded, verification-driven Judge-Planner-Worker loop. A vision-language Judge diagnoses errors by jointly inspecting visual evidence and draft vectors, while a tool-calling Planner generates minimal corrective edits with post-edit re-validation. To remain scalable for city-scale production, MapAgent is selectively triggered only on tiles with low backbone confidence, adding modest overhead while preserving throughput. Experiments on real-world datasets show consistent gains over strong production baselines, especially in complex and long-tail scenarios. Additionally, MapAgent has been integrated into Baidu Maps, supporting lane-level map generation for over 360 cities nationwide and elevating the overall production automation to over 95%, demonstrating MapAgent's practicality and effectiveness for large-scale lane-level map generation.",
      "short_abstract": "Lane-level maps are critical infrastructure for autonomous driving and lane-level navigation, yet constructing and maintaining standardized lane networks for hundreds of cities remains highly labor-intensive. Recent end-to-end vectorized mapping methods can...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "abstract: verification",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.04513",
      "pdf_url": "https://arxiv.org/pdf/2606.04513.pdf"
    },
    {
      "id": "2606.04437",
      "anchor": "paper-2606-04437",
      "title": "INTACT: Ego-Guided Typed Sparse Evidence Retrieval for Heterogeneous Collaborative Perception",
      "authors": [
        "Chen Li",
        "Shengrong Yuan",
        "Jialong Zuo",
        "Xinzhong Zhu",
        "Nong Sang",
        "Changxin Gao"
      ],
      "abstract": "Collaborative perception extends the perceptual range of autonomous vehicles by sharing information across agents, but heterogeneous sensors and perception models make intermediate feature fusion difficult to deploy at scale. Existing heterogeneous collaboration methods typically follow a translation-first paradigm: collaborator features must be aligned, adapted, or projected into an ego-compatible space before fusion. Such feature-compatibility contracts improve fixed-system performance, but they couple deployment to collaborator-specific adaptation and make newly joined heterogeneous agents costly to integrate. To address this gap, we propose INTACT, an ego-guided typed sparse evidence retrieval framework for heterogeneous collaborative perception. Instead of translating an entire collaborator feature map, INTACT lets the ego vehicle issue typed evidence queries that express suspected objects and evidence-deficient regions. Collaborators respond only with local evidence at queried locations, and the ego selects useful responses through sparse per-query routing and injects them through gated residual write-back. This changes the compatibility requirement from global feature-map interpretability to local, typed response comparability under ego-issued queries, enabling a zero-training heterogeneous insertion protocol in which the ego interface is trained once and new collaborators join through checkpoint merging. Extensive experiments on simulated and real-world heterogeneous collaborative perception benchmarks validate the effectiveness and deployability of INTACT. On OPV2V-H, INTACT achieves 80.1 AP70 with only 0.52M additional parameters and 18.0 $\\log_2$ communication volume, corresponding to about 16$\\times$ compression over dense feature transmission. On DAIR-V2X, INTACT achieves 43.8 AP50 under challenging real-world conditions.",
      "short_abstract": "Collaborative perception extends the perceptual range of autonomous vehicles by sharing information across agents, but heterogeneous sensors and perception models make intermediate feature fusion difficult to deploy at scale. Existing heterogeneous...",
      "published": "2026-06-03",
      "updated": "2026-06-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 11,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 77,
      "arxiv_url": "https://arxiv.org/abs/2606.04437",
      "pdf_url": "https://arxiv.org/pdf/2606.04437.pdf"
    },
    {
      "id": "2606.20641",
      "anchor": "paper-2606-20641",
      "title": "MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning",
      "authors": [
        "Letian Chen",
        "Yiren Lu",
        "Justin Fu",
        "Yichen Xie",
        "Runsheng Xu",
        "Jyh-Jing Hwang",
        "Ben Sapp",
        "Drago Anguelov"
      ],
      "abstract": "Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in semantic understanding and common sense reasoning, making them promising candidates for solving planning problems in autonomous driving. However, the next-token text prediction objectives traditionally used in pre-training and supervised fine-tuning (SFT) of MLLMs may fall short of fulfilling the planning objectives for autonomous vehicles. The next-token prediction objective merely encourages per-token imitation in text, often irrespective of multi-step consequences and the alignment with crucial planning considerations such as giving space to other road actors. To overcome these limitations, we propose a reinforcement learning fine-tuning (RLFT) approach, MAGNIFIED, that aligns the MLLM-based driving agent with planning objectives by learning from token-level rewards. By mapping a sequence of predicted tokens to corresponding vehicle trajectories and learning from planning rewards, MAGNIFIED optimizes for the true planning objectives rather than focusing solely on token prediction accuracy, enabling the model to refine its understanding of the planning task beyond simple imitation. We validate our approach on the Waymo Open Motion Dataset with a novel setup incorporating rasterized birds-eye views and tokenized trajectories as inputs and planning-oriented outputs. An initial SFT phase establishes a strong baseline in outputting plan trajectories as sequences of X-Y coordinates in text, while subsequent RL fine-tuning substantially enhances planning performance relative to the SFT baseline (demonstrating over a 10.5% reduction in overlap rate and a 38.9% reduction in off-road rate), underscoring the potential of RLFT on MLLMs to achieve vehicle planning that is better aligned with compliant, comfortable, and efficient driving.",
      "short_abstract": "Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in semantic understanding and common sense reasoning, making them promising candidates for solving planning problems in autonomous driving. However, the next-token text...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 23,
        "evidence": [
          "title: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.20641",
      "pdf_url": "https://arxiv.org/pdf/2606.20641.pdf"
    },
    {
      "id": "2606.20640",
      "anchor": "paper-2606-20640",
      "title": "An LLM-Explainable DRL Framework for Passenger-Directed Autonomous Driving",
      "authors": [
        "Ouided Braoui",
        "Meriem Bouali",
        "Nadir Farhi"
      ],
      "abstract": "Autonomous vehicles offer the potential for safer and more efficient mobility, yet public trust remains limited due to the lack of transparency in their decision-making. This work addresses this issue by combining deep reinforcement learning (DRL) for adaptive driving control with large language model (LLM)-based explainability modules designed to communicate agent behavior to passengers. DRL agents were trained in simulation using a Dueling Double Deep Q-Network to follow distinct driving requests: \\textit{fast}, \\textit{comfort}, and \\textit{stop}. They demonstrated stable learning, safe compliance with traffic rules, and reliable switching between modes within a single trip. In parallel, LLM modules were introduced to interpret passenger requests, determine when explanations were needed, and generate concise, safety-oriented justifications. Results show that this framework, serving as a proof of concept for integrating RL decision-making and LLMs, balances safety, adaptability, and explainability, and is most effective when requests are delayed or overridden due to safety constraints.",
      "short_abstract": "Autonomous vehicles offer the potential for safer and more efficient mobility, yet public trust remains limited due to the lack of transparency in their decision-making. This work addresses this issue by combining deep reinforcement learning (DRL) for...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "abstract: social acceptance",
          "title: explainability",
          "abstract: explainability"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.20640",
      "pdf_url": "https://arxiv.org/pdf/2606.20640.pdf"
    },
    {
      "id": "2606.04271",
      "anchor": "paper-2606-04271",
      "title": "StandardE2E: A Unified Framework for End-to-End Autonomous Driving Datasets",
      "authors": [
        "Stepan Konev"
      ],
      "abstract": "Autonomous driving has shifted from modular perception-prediction-planning stacks toward end-to-end (E2E) models that map sensor inputs directly to vehicle control, often regularized by auxiliary tasks such as 3D detection, motion forecasting, and HD-map perception. Progress is driven by a fast-growing ecosystem of sensor-rich driving datasets, yet each ships its own file formats, APIs, coordinate conventions, and modality coverage, leaving cross-dataset experimentation and even basic per-dataset preprocessing to be re-implemented per project. We present StandardE2E, a framework that provides a single unified interface over E2E driving datasets. StandardE2E (i) standardizes per-dataset preprocessing under one shared data schema; (ii) combines multiple datasets in a single PyTorch DataLoader for cross-dataset pretraining, auxiliary-task supervision, and scenario-level filtering; and (iii) reduces adding a new dataset to a single per-dataset mapping from raw frames to the canonical schema, leaving the entire downstream pipeline unchanged. The framework supports six datasets out of the box: Waymo End-to-End, Waymo Perception, Argoverse 2 Sensor, Argoverse 2 LiDAR, NAVSIM (OpenScene-v1.1), and WayveScenes101, and is released as the open-source standard-e2e Python package, available at https://github.com/stepankonev/StandardE2E.",
      "short_abstract": "Autonomous driving has shifted from modular perception-prediction-planning stacks toward end-to-end (E2E) models that map sensor inputs directly to vehicle control, often regularized by auxiliary tasks such as 3D detection, motion forecasting, and HD-map...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 20,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.04271",
      "pdf_url": "https://arxiv.org/pdf/2606.04271.pdf"
    },
    {
      "id": "2606.03915",
      "anchor": "paper-2606-03915",
      "title": "PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion",
      "authors": [
        "Qingdong Xu",
        "Jiajun Zhu",
        "Shilin Zhu",
        "Xinjing He",
        "Chao Lu",
        "Huanran Wang",
        "Jiyao Zhang"
      ],
      "abstract": "We propose PatchScene, a novel diffusion-based framework for large-scale LiDAR scene completion. Unlike existing methods that rely on global latent representations or dense voxel grids, PatchScene adopts a patch-based voxel diffusion paradigm that explicitly generates fine-grained geometry within localized 3D regions. To ensure coherent reconstruction at both spatial and temporal scales, we introduce a confidence-guided spatio-temporal fusion mechanism that integrates overlapping patches and adjacent frames in a unified generative process. Furthermore, we design an Annular-Flow diffusion strategy that leverages the radial density pattern of LiDAR scans to progressively propagate high-fidelity information from near-range to far-range regions, enabling spatially unbounded scene completion. Extensive experiments on the SemanticKITTI benchmark demonstrate that PatchScene achieves state-of-the-art performance across all standard metrics, surpassing previous approaches in both geometric accuracy and temporal consistency. Remarkably, the model trained on 20 m LiDAR ranges generalizes effectively to 50 m scenes without retraining, highlighting its strong scalability and generalization capability for real-world autonomous driving applications.",
      "short_abstract": "We propose PatchScene, a novel diffusion-based framework for large-scale LiDAR scene completion. Unlike existing methods that rely on global latent representations or dense voxel grids, PatchScene adopts a patch-based voxel diffusion paradigm that explicitly...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Perception & Sensor Fusion: abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03915",
      "pdf_url": "https://arxiv.org/pdf/2606.03915.pdf"
    },
    {
      "id": "2606.03890",
      "anchor": "paper-2606-03890",
      "title": "OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs",
      "authors": [
        "Yifei Li",
        "Pengyiang Liu",
        "Yuhang Zang",
        "Zhongyue Shi",
        "Qi Fu",
        "Hongye Hao",
        "Jiwen Lu"
      ],
      "abstract": "Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Existing benchmarks either evaluate offline over full videos or target events rather than spatial structure. We introduce OVO-S-Bench, a fully human-annotated benchmark for streaming spatial intelligence, comprising 1,680 questions over 348 source videos. Annotation involves 12 trained annotators, each also serving as a blind cross-reviewer, across roughly 804 person-hours of multi-round quality assurance. Each question carries a query timestamp and an evidence interval, and at evaluation, the model sees only the prefix preceding the query. Questions span four levels of increasing abstraction: instantaneous egocentric perception, spatiotemporal context tracking, spatial simulation and reasoning, and allocentric mapping. Across 38 proprietary and open-source MLLMs, Gemini-3.1-Pro trails human experts by 27 points, 59.2 vs. 86.6, with allocentric mapping as the dominant bottleneck. Notably, streaming and spatially fine-tuned MLLMs underperform their own backbones. We further find that chain-of-thought reasoning amplifies spatial errors when ungrounded in the stream. By exposing these limitations, OVO-S-Bench establishes a demanding testbed for next-generation streaming spatial MLLMs.",
      "short_abstract": "Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Existing benchmarks either evaluate offline over full videos or target events...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: mapping",
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03890",
      "pdf_url": "https://arxiv.org/pdf/2606.03890.pdf"
    },
    {
      "id": "2606.03694",
      "anchor": "paper-2606-03694",
      "title": "Face versus Body Tracking for Human-Robot Interaction: An Egocentric Dataset",
      "authors": [
        "Jessica Wenninger",
        "Gabriel Skantze"
      ],
      "abstract": "Meaningful human-robot interaction (HRI) requires a robot to continuously assess user engagement through persistent user tracking. However, state-of-the-art Multi-Object Tracking models are heavily optimized for surveillance or autonomous driving. A social robot faces distinct egocentric challenges, such as humans moving in unpredictable nonlinear patterns, obstructing each other, or leaving and reentering the scene. These dynamics trigger frequent identity switches (IDSW), causing the robot to lose its footing mid-conversation. To address this, we introduce a focused, custom-annotated egocentric dataset collected via the Furhat robot. We present a systematic evaluation isolating detection errors from tracking logic, comparing face versus body tracking, and assessing the impact of extended memory and appearance re-identification (ReID). Results indicate that increasing temporal memory mitigates prolonged occlusions but fails on complex dynamic events. Integrating ReID resolves complex switches but exhibits opposing effects: it substantially improves body tracking stability, yet causes facial IDSW to spike due to profile angle sensitivity. Ultimately, our optimized pipeline reduces IDSW by 49% compared to a standard tracking-by-detection baseline, effectively mitigating interaction breakdowns. As standard benchmarks lack dense, close-quarter occlusions, this work highlights the critical need for natively captured social dynamics to truly validate HRI perception models.",
      "short_abstract": "Meaningful human-robot interaction (HRI) requires a robot to continuously assess user engagement through persistent user tracking. However, state-of-the-art Multi-Object Tracking models are heavily optimized for surveillance or autonomous driving. A social...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03694",
      "pdf_url": "https://arxiv.org/pdf/2606.03694.pdf"
    },
    {
      "id": "2606.03678",
      "anchor": "paper-2606-03678",
      "title": "EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents",
      "authors": [
        "Tong Nie",
        "Yuewen Mei",
        "Yihong Tang",
        "Junlin He",
        "Jie Deng",
        "Jian Sun",
        "Wei Ma"
      ],
      "abstract": "Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversariality to expose failures while preserving realism. Existing methods usually manage this trade-off with handcrafted heuristics, confining generation to known priors and overlooking underexplored patterns. While recent open-ended agentic evolution can push this limit, unconstrained general agents lack strict simulator grounding and tend to collapse the multi-objective tension into single-scalar maximization. Here we present EvoDrive, the first automated, LLM-based agentic evolution framework for multi-objective scenario generation. EvoDrive employs a simulator-grounded actor-critic architecture where a memory-driven actor iteratively proposes improvements to the generators and critics filter out implausible candidates, and a self-evolving world evaluator routes promising proposals to optimize simulation budgets. EvoDrive further maintains a Pareto archive of evaluated candidates to preserve diverse attack-realism trade-offs and guide future evolution via simulation feedback. Benchmark results on MetaDrive and CARLA show that EvoDrive not only significantly expands the Pareto frontier across various generators, but also produces valuable scenarios for policy training.",
      "short_abstract": "Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversariality to expose failures while preserving realism. Existing methods usually manage this trade-off with...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 18,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03678",
      "pdf_url": "https://arxiv.org/pdf/2606.03678.pdf"
    },
    {
      "id": "2606.03590",
      "anchor": "paper-2606-03590",
      "title": "CANMOT: Class-Aware Noise Modeling for Multi-Object Tracking in Autonomous Driving",
      "authors": [
        "Timo Osterburg",
        "Stefan Schütte",
        "Torsten Bertram"
      ],
      "abstract": "Kalman filter (KF)-based multi-object tracking (MOT) remains a strong baseline for autonomous driving due to its strong performance, computational efficiency and interpretability. In most practical systems, the process noise and measurement noise covariances are defined globally and shared across object classes, presuming identical uncertainty characteristics across heterogeneous traffic participants. This work revisits this assumption and proposes CANMOT, a class-aware and object-aligned noise modeling framework for KF-based 3D MOT. Class-specific diagonal process and measurement covariance matrices are introduced and optionally expressed in the object coordinate frame to preserve longitudinal-lateral anisotropy. Systematic experiments on the nuScenes benchmark show that class-aware and object-aligned noise modeling improves tracking performance and substantially reduces identity switches compared to state-of-the-art (SotA). In addition, the consistency of the estimated uncertainty is analyzed using the Average Normalized Estimation Error Squared (ANEES) and $χ^2$-based violation tests. The results reveal severe overconfidence in standard KF-based MOT baselines. While the proposed formulation improves calibration without modifying the underlying filtering framework, it still exhibits substantial inconsistency, highlighting the need for further research in this area. Code is available at https://github.com/rst-tu-dortmund/learned-3d-nms.",
      "short_abstract": "Kalman filter (KF)-based multi-object tracking (MOT) remains a strong baseline for autonomous driving due to its strong performance, computational efficiency and interpretability. In most practical systems, the process noise and measurement noise covariances...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03590",
      "pdf_url": "https://arxiv.org/pdf/2606.03590.pdf"
    },
    {
      "id": "2606.03581",
      "anchor": "paper-2606-03581",
      "title": "UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene via Rendering Fusion",
      "authors": [
        "Ye Wu",
        "Ruiqi Song",
        "Baiyong Ding",
        "Nanxin Zeng",
        "Junjie Cheng",
        "Yunfeng Ai"
      ],
      "abstract": "Unstructured scenes present unique challenges for autonomous driving, as irregular obstacles and sparse scene layouts undermine the effectiveness of traditional perception methods such as 3D object detection. 3D semantic occupancy prediction has emerged as a prominent focus due to its ability to provide dense spatial representations by assigning semantic labels to individual voxels in 3D space. However, directly applying 3D semantic occupancy prediction to unstructured scenes remains challenging because scene sparsity hinders effective cross-modal fusion and the more severe long-tail distribution in these scenarios further degrades prediction performance. To validate the effectiveness of our approach, we construct a dedicated dataset of unstructured scenes collected from open-pit mines. Based on this, we propose UnsOcc, a multi-modal 3D semantic occupancy prediction framework that improves robustness in unstructured environments. At its core, we introduce a rendering-based fusion module, RenderFusion, which enhances cross-modal feature alignment through bidirectional rendering supervision. Furthermore, we propose GSRefinement, a detail-aware auxiliary supervision method based on Gaussian Splatting that projects sparse 3D occupancy predictions into dense 2D semantic segmentation maps, enabling effective supervision for long-tail categories. Extensive experiments on both the open-pit mine dataset and the nuScenes dataset demonstrate that our method significantly outperforms existing state-of-the-art approaches.",
      "short_abstract": "Unstructured scenes present unique challenges for autonomous driving, as irregular obstacles and sparse scene layouts undermine the effectiveness of traditional perception methods such as 3D object detection. 3D semantic occupancy prediction has emerged as a...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 11,
        "evidence": [
          "abstract: perception",
          "abstract: object detection",
          "abstract: scene segmentation",
          "title: occupancy",
          "abstract: occupancy"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03581",
      "pdf_url": "https://arxiv.org/pdf/2606.03581.pdf"
    },
    {
      "id": "2606.03314",
      "anchor": "paper-2606-03314",
      "title": "TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing",
      "authors": [
        "Tim-Felix Faasch",
        "Jochen Kall",
        "Lucas Nunes",
        "Jens Behley",
        "Cyrill Stachniss"
      ],
      "abstract": "High-fidelity semantic 3D scene representations are crucial for numerous applications, including robotics, autonomous driving, and simulation. Beyond this, the ability to edit such representations enables developers to adapt these applications more easily to specific target scenarios. Current approaches provide limited support for controllable editing. We introduce TASE, a method that projects pretrained 2D semantic features into a truncation-aware embedding space to enable flexible 3D scene editing. Our method explicitly optimizes a feature space in which progressively reducing feature channels yields increasingly abstract semantic representations, while retaining more channels preserves fine-grained detail. Additionally, we improve multi-view consistency of the features using a scale- and translation-equivariance loss. The resulting truncation-aware embedding space enables text-driven edits to 3D scenes, providing explicit control over how strongly edits adhere to the original scene content and allowing more substantial modifications than prior methods. Moreover, we propose a finetuning stage for the editing diffusion model to mitigate artifacts caused by geometric changes. Experimental results demonstrate competitive performance in 3D scene editing, substantially outperforming prior methods on edits involving large geometric modifications.",
      "short_abstract": "High-fidelity semantic 3D scene representations are crucial for numerous applications, including robotics, autonomous driving, and simulation. Beyond this, the ability to edit such representations enables developers to adapt these applications more easily to...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 7,
        "evidence": [
          "title: scene understanding"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03314",
      "pdf_url": "https://arxiv.org/pdf/2606.03314.pdf"
    },
    {
      "id": "2606.03296",
      "anchor": "paper-2606-03296",
      "title": "Bridging Predictive Uncertainty and Safe Action: Sample-Conditioned Differentiable Planning for Autonomous Driving",
      "authors": [
        "Chengzhen Meng",
        "Pei Liu",
        "Zhiyu Huang",
        "Chen Lv",
        "Jun Ma"
      ],
      "abstract": "Complex, dynamic, and interactive driving environments pose significant challenges for autonomous driving, primarily due to the pervasive uncertainty of surrounding traffic. A fundamental bottleneck in current systems is the disconnect between highly expressive uncertainty modeling and interpretable, safe motion planning. In this paper, we propose a novel sample-conditioned differentiable planning framework that bridges this gap by explicitly incorporating diffusion-generated future trajectories into the optimization process. Rather than compressing predictions into a single deterministic future or relying on black-box end-to-end architectures, our approach leverages a conditional diffusion model to generate a diverse set of plausible future scenarios. Crucially, these samples are directly fed into a differentiable planner, which explicitly mitigates predictive uncertainty via an empirical Conditional Value-at-Risk (CVaR) tail-risk constraint. This allows the planner to optimize a physically interpretable trajectory that is robust to rare yet safety-critical interactions. Furthermore, we introduce a directed graph representation for scene context that yields substantial improvements in both predictive effectiveness and computational efficiency. Validated through extensive open-loop and closed-loop evaluations on the Waymo Open Motion and Argoverse 2 datasets, our framework significantly outperforms state-of-the-art baselines in safety, efficiency, and ride comfort.",
      "short_abstract": "Complex, dynamic, and interactive driving environments pose significant challenges for autonomous driving, primarily due to the pervasive uncertainty of surrounding traffic. A fundamental bottleneck in current systems is the disconnect between highly...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 3,
        "evidence": [
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03296",
      "pdf_url": "https://arxiv.org/pdf/2606.03296.pdf"
    },
    {
      "id": "2606.03160",
      "anchor": "paper-2606-03160",
      "title": "SRENet: Spectral Re-Entry Network for Point Cloud Action Recognition",
      "authors": [
        "Qiuxia Wu",
        "Jiarui Lan",
        "Wenxiong Kang",
        "Zhiyong Wang",
        "Kun Hu"
      ],
      "abstract": "Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer interaction. However, the irregular structure and temporal inconsistency of point clouds pose unique challenges for spatio-temporal representation learning, especially in capturing both global motion context and fine-grained temporal dynamics. We propose SRENet, a spectral-aware framework designed to explicitly learn both global context and fine-grained temporal dynamics of motion from a frequency perspective for action recognition. SRENet introduces a Spectral Decomposition Block (SDeBlock) that performs wavelet-based analysis along temporal and spatial axes, disentangling features into low- and high-frequency components with frequency-specific attention. To recover residual dynamics and re-align temporal frequency structures distorted during semantic fusion, a Spectral Re-entry Block (SReBlock) performs secondary temporal decomposition. Furthermore, a spectral-aware learning strategy is devised to enhance discriminability in both frequency subspaces via contrastive loss and a curriculum schedule that gradually shifts focus from low- to high-frequency spaces in line with coarse to detailed motion patterns. Extensive experiments on MSR-Action3D, NTU-RGBD and NTU-RGBD120 demonstrate that SRENet achieves state-of-the-art performance, validating the effectiveness of frequency modeling in point cloud-based action understanding.",
      "short_abstract": "Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer interaction. However, the irregular structure and temporal inconsistency of point clouds pose unique...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "title: point cloud",
          "abstract: point cloud"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03160",
      "pdf_url": "https://arxiv.org/pdf/2606.03160.pdf"
    },
    {
      "id": "2606.03159",
      "anchor": "paper-2606-03159",
      "title": "NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation",
      "authors": [
        "NVIDIA",
        ":",
        "Aarti Basant",
        "Amlan Kar",
        "Despoina Paschalidou",
        "Fangyin Wei",
        "Francesco Ferroni",
        "Guillermo Garcia Cobo",
        "Haithem Turki",
        "Huan Ling",
        "Jaewoo Seo",
        "James Lucas",
        "Jay Zhangjie Wu",
        "Jialiang Wang",
        "Jonathan Lorraine",
        "Jun Gao",
        "Kai He",
        "Katarina Tothova",
        "Kevin Xie",
        "Michał Tyszkiewicz",
        "Qi Wu",
        "Riccardo de Lutio",
        "Ruilong Li",
        "Sanja Fidler",
        "Seung Wook Kim"
      ],
      "abstract": "As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its actions dynamically update the simulator state and directly influence the next set of generated sensor observations. While recent reconstruction-based neural simulators offer photorealism, they are fundamentally constrained by their initial captured data and struggle to generalize to highly dynamic or novel scenes. To overcome these limitations, we introduce OmniDreams, a foundation generative world model mid- and post-trained from the Cosmos diffusion model to autoregressively generate action-conditioned videos in real time. By leveraging the rich visual priors of Cosmos and mid- and post-training on 21k hours of driving scenarios, OmniDreams synthesizes complex, unobserved phenomena that are hard for traditional simulators to capture, such as extreme weather and unpredictable dynamic agent behaviors. Crucially, it autoregressively conditions its photorealistic sensor generation on past frames, the current simulator state, and immediate driving actions. Deployed in a closed-loop system with the Alpamayo 1 policy model and AlpaSim orchestrator, OmniDreams acts as a highly responsive, reactive environment, providing a scalable and comprehensive solution for training and evaluating next-generation autonomous driving policies. We additionally show preliminary results indicating that a world-action model (WAM) post-trained from OmniDreams achieves strong performance on the Physical AI Autonomous Vehicles NuRec dataset, surpassing the VLA-based Alpamayo 1.5 research policy model while using only 1/5 the total parameters. These results highlight the potential for a real-time world model like OmniDreams to also serve as a backbone for policy architectures.",
      "short_abstract": "As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its actions...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "VLA",
        "World Model",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03159",
      "pdf_url": "https://arxiv.org/pdf/2606.03159.pdf"
    },
    {
      "id": "2606.03755",
      "anchor": "paper-2606-03755",
      "title": "LAP: An Agent-to-Instrument Protocol for Autonomous Science",
      "authors": [
        "Linwu Zhu",
        "Liqiang Gao",
        "Yan Chen",
        "Dan Zhu",
        "Jian Huang"
      ],
      "abstract": "Autonomous science is moving from demonstration to infrastructure. Large language model agents now plan experiments, and self-driving laboratories execute them. Yet every such system rebuilds the link between the reasoning agent and the physical instrument from scratch, against fragmented vendor SDKs and standards built for deterministic software clients rather than probabilistic, goal-directed agents. Recent agent-interoperability protocols clarify two of the three edges of an agentic ecosystem (Anthropic's Model Context Protocol (MCP) standardizes the agent-to-tool edge, and Google's Agent2Agent (A2A) the agent-to-agent edge), but neither models the agent-to-instrument edge, where operations are stateful, safety-critical, exclusively owned, physically embodied, and produce measurements with units, calibration, and uncertainty. We present the Lab Agent Protocol (LAP), a protocol design that fills this gap. LAP retains A2A's peer-to-peer, discovery-first, task-lifecycle structure and adds four physical-world primitives: (i) the InstrumentCard, a signed capability and physical-limit description; (ii) first-class reservation for exclusive instrument and sample locking; (iii) a safety-fence handshake with operator-confirmation tokens cryptographically bound to a specific task and its parameters, gating hazardous and irreversible operations; and (iv) a MeasurementResult schema that makes every result physically typed (QUDT/UCUM), calibration-anchored, uncertainty-bearing, and reproducible by construction. We specify roles, a six-layer architecture, the JSON-RPC method set, the task and safety state machines, the error model, and cross-laboratory federation, and walk a closed-loop autonomous campaign through the protocol end-to-end. LAP is transport-compatible with the A2A/MCP ecosystem and encapsulates rather than replaces existing device standards such as SiLA 2 and OPC-UA.",
      "short_abstract": "Autonomous science is moving from demonstration to infrastructure. Large language model agents now plan experiments, and self-driving laboratories execute them. Yet every such system rebuilds the link between the reasoning agent and the physical instrument...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03755",
      "pdf_url": "https://arxiv.org/pdf/2606.03755.pdf"
    },
    {
      "id": "2606.04072",
      "anchor": "paper-2606-04072",
      "title": "CADET: A Modular Platform for Evaluating Distributed Cooperative Autonomy in Connected Autonomous Vehicles",
      "authors": [
        "Pragya Sharma",
        "Brian Wang",
        "Mani Srivastava"
      ],
      "abstract": "Deep learning models are increasingly central to autonomous vehicle (AV) pipelines, yet their integration has traditionally followed a monolithic design where perception, planning, and control execute on a single onboard computer. This design overlooks the emerging paradigm of cooperative autonomy, where vehicles interact with roadside units (RSUs), edge servers, and cloud-hosted intelligence through vehicle-to-everything (V2X) connectivity. Cooperative perception and control improve safety and efficiency, but also introduce systems-level challenges: network latency, compute heterogeneity, and multi-tenant contention, all critically affect real-time decision-making. These challenges are further amplified by the increasing reliance on large foundation models, whose scale necessitates cloud deployment. We present CADET (Cooperative Autonomy through Distributed Experimentation Toolkit), a modular platform for systematic and reproducible evaluation of distributed cooperative autonomy systems under realistic deployment conditions. CADET decouples the AV stack into composable modules that can be flexibly deployed across vehicles, infrastructure, and edge/cloud tiers. The framework integrates state-of-the-art models, incorporates trace-driven network and workload emulation, and provides synchronized model-, system-, and task-level instrumentation. Through V2V and V2I experiments, we show that distributed deployment choices fundamentally shape safety, with V2V intent packets outperforming cloud-based perception and RSU-assisted perception sustaining safety until overloaded by concurrent requests. Although designed for AV pipelines, CADET also supports dataset-driven experimentation, enabling systems and ML researchers to benchmark distributed inference workloads independently of full vehicle simulation. CADET is open source, with code and demo available at https://nesl.github.io/cadet-web.",
      "short_abstract": "Deep learning models are increasingly central to autonomous vehicle (AV) pipelines, yet their integration has traditionally followed a monolithic design where perception, planning, and control execute on a single onboard computer. This design overlooks the...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 40,
        "margin": 33,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "title: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.04072",
      "pdf_url": "https://arxiv.org/pdf/2606.04072.pdf"
    },
    {
      "id": "2606.03662",
      "anchor": "paper-2606-03662",
      "title": "Agentic Generation and Evolution of Knowledge Models",
      "authors": [
        "Man Zhang",
        "Tao Yue",
        "Nazareno M. Aguirre",
        "Diego Garbervetsky",
        "Sebastian Uchitel"
      ],
      "abstract": "Complex software systems such as autonomous vehicles, robotics increasingly interact with dynamic physical, cyber, and social environments. Reasoning about their behavior, maintaining them under continuous change, and evolving them safely require trustworthy knowledge about the system, its assumptions, and its operating context. Knowledge models (KMs) provide a practical basis for such reasoning, but they may themselves become incomplete, inconsistent, or outdated as systems evolve. This paper presents TrustModel, a vision for the agentic generation and evolution of living KMs. TrustModel comprises three agentic subsystems: Modeling, for constructing and updating KMs; Conformance, for assessing their alignment with the system and its environment; and Evolution, for generating guidance to keep KMs synchronized with emerging changes. We demonstrate how TrustModel can be instantiated for model-based testing and discuss its potential for supporting other MDE activities, such as requirements and assumption monitoring, architectural drift tracking, and change impact assessment. Overall, TrustModel positions living KMs as a foundation for dependable engineering of continuously evolving software systems.",
      "short_abstract": "Complex software systems such as autonomous vehicles, robotics increasingly interact with dynamic physical, cyber, and social environments. Reasoning about their behavior, maintaining them under continuous change, and evolving them safely require trustworthy...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03662",
      "pdf_url": "https://arxiv.org/pdf/2606.03662.pdf"
    },
    {
      "id": "2606.03617",
      "anchor": "paper-2606-03617",
      "title": "SA-DTS: Semantic-Aware Digital Twin Synchronization over 6G Networks",
      "authors": [
        "Vincenzo Sammartino"
      ],
      "abstract": "Digital Twins (DTs) are emerging as a cornerstone of the 6G vision, enabling real-time cyber-physical mirroring for smart manufacturing, autonomous vehicles, and remote healthcare. However, maintaining high-fidelity synchronization at scale demands an enormous and sustained uplink bandwidth, threatening both the feasibility and the energy efficiency of large deployments. We propose a Semantic-Aware DT Synchronization (SA-DTS) framework that radically redefines the synchronization pipeline: instead of streaming raw sensor or video data, a lightweight neural semantic encoder at the physical-world source extracts only task-relevant features and transmits compact semantic descriptors over the 6G air interface. At the DT replica, a paired decoder coupled with a dynamic Knowledge Graph (KG) reconstructs the full contextual state. A hierarchical KG partitioning strategy with an adaptive partition count $G = \\lceil N / \\log_2 N \\rceil$ ensures that aggregate update overhead scales as $O(N \\log N)$ rather than $O(N^2)$, making the framework viable for deployments with hundreds of simultaneously twinned entities. Extensive simulations on three canonical DT workloads -- industrial robot control, patient-monitoring, and vehicular platooning -- demonstrate bandwidth savings of up to 94%, end-to-end synchronization latency reductions of 87%, and KG-assisted state-reconstruction accuracy exceeding 97%, all under realistic 6G channel conditions. Empirical correlation confirms that the proposed Semantic Fidelity Score tracks standard task metrics (collision accuracy, alarm F1, spacing deviation) with Pearson $r > 0.97$ (95% CI: [0.961, 0.982]). Our results reveal that semantic communication is not merely a compression tool but a fundamental enabler for truly real-time, scalable DT ecosystems.",
      "short_abstract": "Digital Twins (DTs) are emerging as a cornerstone of the 6G vision, enabling real-time cyber-physical mirroring for smart manufacturing, autonomous vehicles, and remote healthcare. However, maintaining high-fidelity synchronization at scale demands an...",
      "published": "2026-06-02",
      "updated": "2026-06-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 10,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 78,
      "arxiv_url": "https://arxiv.org/abs/2606.03617",
      "pdf_url": "https://arxiv.org/pdf/2606.03617.pdf"
    },
    {
      "id": "2606.02979",
      "anchor": "paper-2606-02979",
      "title": "Towards Compact Autonomous Driving Perception with Balanced Learning and Multi-sensor Fusion",
      "authors": [
        "Oskar Natan",
        "Jun Miura"
      ],
      "abstract": "We present a novel compact deep multi-task learning model to handle various autonomous driving perception tasks in one forward pass. The model performs multiple views of semantic segmentation, depth estimation, light detection and ranging (LiDAR) segmentation, and bird's eye view projection simultaneously without being supported by other models. We also provide an adaptive loss weighting algorithm to tackle the imbalanced learning issue that occurred due to plenty of given tasks. Through data pre-processing and intermediate sensor fusion techniques, the model can process and combine multiple input modalities retrieved from RGB cameras, dynamic vision sensors (DVS), and LiDAR placed at several positions on the ego vehicle. Therefore, a better understanding of a dynamically changing environment can be achieved. Based on the ablation study, the model variant trained with our proposed method achieves a better performance. Furthermore, a comparative study is also conducted to clarify its performance and effectiveness against the combination of some recent models. As a result, our model maintains better performance even with much fewer parameters. Hence, the model can inference faster with less GPU memory utilization. Moreover, the result tends to be consistent in 3 different CARLA simulation datasets and 1 real-world nuScenes-lidarseg dataset. To support future research, we share codes and other files publicly at https://github.com/oskarnatan/compact-perception.",
      "short_abstract": "We present a novel compact deep multi-task learning model to handle various autonomous driving perception tasks in one forward pass. The model performs multiple views of semantic segmentation, depth estimation, light detection and ranging (LiDAR)...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 42,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: scene segmentation",
          "title: sensor fusion",
          "abstract: sensor fusion",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: depth estimation"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.02979",
      "pdf_url": "https://arxiv.org/pdf/2606.02979.pdf"
    },
    {
      "id": "2606.02956",
      "anchor": "paper-2606-02956",
      "title": "The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset",
      "authors": [
        "Richard Schwarzkopf",
        "Fabian Immel",
        "Alexander Blumberg",
        "Jonas Merkert",
        "Nils Rack",
        "Kaiwen Wang",
        "Fabian Konstantinidis",
        "Julian Truetsch",
        "Carlos Fernandez",
        "Annika Bätz",
        "Kevin Rösch",
        "Marlon Steiner",
        "Willi Poh",
        "Yinzhe Shen",
        "Royden Wagner",
        "Felix Hauser",
        "Dominik Strutz",
        "Jaime Villa",
        "Gleb Stepanov",
        "Holger Caesar",
        "Ömer Şahin Taş",
        "Frank Bieder",
        "Jan-Hendrik Pauls",
        "Christoph Stiller"
      ],
      "abstract": "Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully synchronized sensor suite combines high-resolution global-shutter cameras, long-range lidar beyond 400m, 4D imaging radar, and redundant GNSS/INS localization. Our HD maps are, to our knowledge, the most complete of any sensor dataset, validated through autonomous driving trials on open-source software. For the first time in a public dataset, all driving-relevant traffic elements, such as traffic lights, are mapped in 3D to a reprojection-accurate level with full topological connectivity. Recorded in cities with irregular street layouts and mixed traffic modes, our dataset complements existing datasets by broadening the available geographic diversity. We also introduce four benchmarks, each advancing spatial learning for embodied AI: online HD map construction, long-range depth estimation, novel view synthesis, and end-to-end driving. Project page: https://kitscenes.com/",
      "short_abstract": "Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 5,
        "evidence": [
          "abstract: localization",
          "abstract: HD map",
          "abstract: mapping",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.02956",
      "pdf_url": "https://arxiv.org/pdf/2606.02956.pdf"
    },
    {
      "id": "2606.02924",
      "anchor": "paper-2606-02924",
      "title": "ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception",
      "authors": [
        "Mellon M. Zhang",
        "Siddhant Panse",
        "Zimo Fan",
        "Akshal Dhal",
        "Rishit Sarkar",
        "Glen Chou"
      ],
      "abstract": "Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and potentially adversarial sensor anomalies. This gap is especially critical for LiDAR, where external actors can physically manipulate the sensing process to induce black-box perception failures without accessing the model. Existing LiDAR benchmarks provide little visibility into this failure mode. Prior adversarial LiDAR studies have largely centered on attack hardware, geometric and algorithmic defenses, and early-generation detectors, leaving the robustness of modern perception systems unexplored. To address this evaluation gap, we introduce ATLAS (Adversarial Temporal LiDAR Attack Suite), the first large-scale, physically grounded evaluation benchmark for LiDAR perception models under black-box sensor attacks, simulating the two primary attack modes -- point injection and point removal -- across real driving sequences. Evaluating a broad cross-section of current state-of-the-art LiDAR perception models, ATLAS reveals a surprising robustness asymmetry: models with stronger performance on standard benchmarks tend to better withstand removal attacks, yet are actually more vulnerable to injection attacks than weaker models. We trace this vulnerability to standard object database sampling augmentations, revealing how current training practices can induce architecture-agnostic robustness failures, and study initial directions for mitigating both attack modes. We release the ATLAS generation code to support extensible, reproducible evaluations as attack capabilities evolve, helping make black-box sensor robustness an explicit consideration in future LiDAR perception development.",
      "short_abstract": "Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and potentially adversarial sensor anomalies. This gap is especially critical for LiDAR, where external actors can...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 24,
        "margin": 4,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.02924",
      "pdf_url": "https://arxiv.org/pdf/2606.02924.pdf"
    },
    {
      "id": "2606.02774",
      "anchor": "paper-2606-02774",
      "title": "GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving",
      "authors": [
        "Yingzi Ma",
        "Chaowei Xiao",
        "Ming Jiang"
      ],
      "abstract": "Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains underexplored, raising uncertainties about their deployment across diverse global settings. We therefore introduce GeoDrive-Bench, a novel benchmark that enables the systematic investigation of VLMs' geo-culturally grounded driving reasoning. We curated 5,053 human-validated multiple-choice QA pairs across six countries covering diverse driving cultures. Specifically, we emphasize four driving tasks: perception, prediction, planning, and region reasoning. Each question requires models to infer the correct driving behavior from visual evidence and local traffic conventions without explicit country labels. Beyond evaluation, we further design a distillation algorithm that injects region-specific traffic-rule knowledge into the internal representations of VLMs, enabling models to better align visual scene understanding with local driving policies. Experiments on nine state-of-the-art VLMs show substantial performance variations across geo-driving cultures for each task, while our proposed baseline models exhibit improved geo-cultural reasoning across regions. These results suggest that current VLMs still lack robust region-aware driving intelligence and highlight GeoDrive-Bench as a diagnostic and training-oriented testbed for deployable autonomous driving foundation models.",
      "short_abstract": "Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains underexplored, raising uncertainties about their deployment across diverse global settings. We therefore...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: perception",
          "abstract: scene understanding"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.02774",
      "pdf_url": "https://arxiv.org/pdf/2606.02774.pdf"
    },
    {
      "id": "2606.02482",
      "anchor": "paper-2606-02482",
      "title": "X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding",
      "authors": [
        "Peiwen Sun",
        "Xudong Lu",
        "Huadai Liu",
        "Yang Bo",
        "Dongming Wu",
        "Huankang Guan",
        "Minghong Cai",
        "Jinpeng Chen",
        "Xintong Guo",
        "Shuhan Li",
        "Fang Liu",
        "Rui Liu",
        "Xiangyu Yue"
      ],
      "abstract": "While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and multi-screen collaboration, inherently demand continuous, multi-stream interactions. However, existing benchmarks are confined to single-stream paradigms, leaving a critical gap in evaluating online, cross-stream reasoning. To bridge this, we introduce X-Stream, the first benchmark dedicated to multi-stream streaming understanding. Comprising 4,220 rigorously curated QA pairs across 932 videos, X-Stream evaluates 11 subtasks across multi-window, multi-view, and multi-device scenarios. Crucially, our dataset is constructed using a novel dual-verification pipeline that prevents over-reliance on a single stream. Furthermore, we pioneer the conceptualization of multi-modal large language models (MLLMs) as naive multiplexers, systematically evaluating their performance through the lens of Signal Multiplexing Theory. Our extensive online inference experiments reveal a stark reality: state-of-the-art MLLMs struggle significantly with concurrent streams, achieving only about 50% score and exhibiting poor proactive ability. Ultimately, X-Stream exposes the trade-off of current multiplexing schemes, providing both a practical evaluation protocol and empirical guidance for next-generation multi-stream agents.",
      "short_abstract": "While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and multi-screen collaboration, inherently demand continuous, multi-stream interactions. However, existing...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Safety, Security & Verification: abstract: verification"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.02482",
      "pdf_url": "https://arxiv.org/pdf/2606.02482.pdf"
    },
    {
      "id": "2606.02105",
      "anchor": "paper-2606-02105",
      "title": "Multimodal Action Diffusion for Robust End-to-End Autonomous Driving",
      "authors": [
        "Jorge Daniel Rodríguez-Vidal",
        "Diego Porres",
        "Gabriel Villalonga Pineda",
        "Antonio M. López Peña"
      ],
      "abstract": "End-to-End Autonomous Driving (E2E-AD) systems have largely converged on predicting intermediate trajectory waypoints, delegating final control to hand-crafted controllers with GPS access. Direct control-signal prediction (outputting throttle, steer and brake in an end-to-end fashion) remains underexplored, and critically, the role of action multimodality in such systems is not well understood. We argue that moving beyond deterministic, single-action outputs is not merely a modelling choice, but a key driver of driving performance, representational quality, and training stability. To validate this, we introduce the Action Diffusion Transformer (ADT), an anchor-free diffusion transformer trained with a MSE objective that natively models the multimodal distribution of plausible driving actions. Rather than committing to a single deterministic command, ADT generates K action candidates and selects the most suitable one at inference via Nearest Neighbour Matching (NNM). Beyond strong benchmark numbers, we show that action multimodality yields measurable benefits in learned representations and behavioral consistency, effects that deterministic architectures cannot replicate. ADT surpasses previous state-of-the-art on the challenging closed-loop Bench2Drive benchmark while achieving ten times lower latency, demonstrating that expressive, multimodal action modelling is both practically efficient and conceptually essential for robust end-to-end driving.",
      "short_abstract": "End-to-End Autonomous Driving (E2E-AD) systems have largely converged on predicting intermediate trajectory waypoints, delegating final control to hand-crafted controllers with GPS access. Direct control-signal prediction (outputting throttle, steer and brake...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 28,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.02105",
      "pdf_url": "https://arxiv.org/pdf/2606.02105.pdf"
    },
    {
      "id": "2606.01972",
      "anchor": "paper-2606-01972",
      "title": "AI-Based KPI Prediction Methods in Future 6G Networks: A Survey",
      "authors": [
        "Niloofar Mehrnia",
        "Gourav Prateek Sharma",
        "Samie Mostafavi",
        "Andreas Johnsson",
        "Sinem Coleri",
        "Carlo Fischione",
        "James Gross"
      ],
      "abstract": "The evolution from 5G to 5G-Advanced and the vision of 6G demand unprecedented levels of network performance, in which meeting stringent network Key Performance Indicators (KPIs), including capacity, latency, coverage, and reliability, is critical to supporting emerging applications such as autonomous driving, industrial automation, and immersive communications. Traditional reactive network management is insufficient in this context, driving the need for predictive, data-driven approaches. Machine Learning (ML) has emerged as a key enabler, enabling the forecasting of KPI trends from diverse data sources and thereby enabling proactive, AI-native automation in mobile networks. This survey provides the first comprehensive and systematic review of data-driven KPI prediction methods for future 6G networks. We introduce a multi-dimensional taxonomy that classifies prediction approaches by KPI type, data source, the network protocol stack at which the KPI is predicted, prediction horizon, model family, and prediction objective. Using this taxonomy, we analyze the state of the art across various KPIs, highlighting representative methods ranging from classical statistical models to deep learning and reinforcement learning. We further discuss enabling system aspects, including data collection and learning architectures, and examine deployment challenges, including data availability, scalability, privacy, and sustainability. Finally, we outline open research directions spanning new KPI definitions, probabilistic and explainable predictions. This survey aims to provide researchers and practitioners with a structured understanding of the KPI prediction landscape and a roadmap toward predictive network automation in future 6G systems.",
      "short_abstract": "The evolution from 5G to 5G-Advanced and the vision of 6G demand unprecedented levels of network performance, in which meeting stringent network Key Performance Indicators (KPIs), including capacity, latency, coverage, and reliability, is critical to...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Reinforcement Learning",
        "Survey",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 13,
        "margin": 2,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.01972",
      "pdf_url": "https://arxiv.org/pdf/2606.01972.pdf"
    },
    {
      "id": "2606.01935",
      "anchor": "paper-2606-01935",
      "title": "Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning",
      "authors": [
        "Ziyang Yao",
        "Zeyu Zhu",
        "YunCheng Jiang",
        "Zibin Guo",
        "Huijing Zhao"
      ],
      "abstract": "Discrete visual tokens should provide a compact representation for both token-based world modeling and planning in autonomous driving. However, most tokenizers are inherited from image generation and are optimized mainly for pixel reconstruction, which may leave a gap between what is easy to generate and what is useful to decode for driving decisions. We present a representation-guided and geometry-enhanced tokenizer that learns discrete tokens under joint supervision. The tokenizer aligns its discrete bottleneck with a frozen DINO feature space through feature decoding, while preserving appearance via RGB reconstruction with perceptual and adversarial losses. To inject geometric state-related cues, we add adjacent-frame depth and relative-pose supervision during training and stabilize joint objectives with multi-codebook quantization. We evaluate the same learned tokens with a lightweight planning readout and a GPT-style next-token world model. Experiments on NAVSIM show improved reconstruction fidelity and representation consistency, competitive planning performance under a fixed decoder, and better generative quality under matched settings.",
      "short_abstract": "Discrete visual tokens should provide a compact representation for both token-based world modeling and planning in autonomous driving. However, most tokenizers are inherited from image generation and are optimized mainly for pixel reconstruction, which may...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: title: planning",
          "Planning & Decision-Making: abstract: planning",
          "Planning & Decision-Making: abstract: decision-making",
          "Prediction & World Models: title: world model",
          "Prediction & World Models: abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.01935",
      "pdf_url": "https://arxiv.org/pdf/2606.01935.pdf"
    },
    {
      "id": "2606.01822",
      "anchor": "paper-2606-01822",
      "title": "Hierarchically Decoupled Mixture-of-Experts for Robust Traffic Sign Recognition in Complex Driving Scenarios",
      "authors": [
        "Mingxiao Wang",
        "Xiaozhen Qu",
        "Bolin Gao",
        "Tong Wang",
        "Lei He"
      ],
      "abstract": "Traffic sign detection is a fundamental component of environmental perception in autonomous driving and intelligent transportation systems. However, most existing detectors rely on static inference with globally shared parameters, limiting their ability to adapt to diverse and unstructured traffic scenarios. As a result, a single static model often struggles to simultaneously handle both clear near-range samples and challenging conditions such as distant small targets or adverse weather environments. To address this limitation, we propose CBDES MoE TSR, a hierarchically decoupled heterogeneous mixture-of-experts(MoE) framework for traffic sign recognition. The proposed framework departs from the conventional globally shared parameter paradigm by introducing a heterogeneous You Only Look Once (YOLO) expert pool together with a lightweight gating network, enabling an image-level dynamic routing mechanism. Based on the semantic characteristics of the input image, the gating module selectively activates the most suitable expert model from the expert pool, enabling a shift from fixed parameter fitting to on-demand dynamic representation. This design enhances feature extraction capability for specific scenarios while maintaining controlled inference overhead. Experimental results demonstrate that the proposed method achieves a remarkable balance between detection accuracy and efficiency on the composite traffic sign dataset. Specifically, our method attains an mAP50-95 of 76.8%, yielding a 2.3% improvement over the baseline method (74.5%) while simultaneously reducing computational overhead by approximately 39.4%. These findings robustly validate the effectiveness of the proposed approach.",
      "short_abstract": "Traffic sign detection is a fundamental component of environmental perception in autonomous driving and intelligent transportation systems. However, most existing detectors rely on static inference with globally shared parameters, limiting their ability to...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "title: traffic sign recognition",
          "abstract: traffic sign recognition"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.01822",
      "pdf_url": "https://arxiv.org/pdf/2606.01822.pdf"
    },
    {
      "id": "2606.01757",
      "anchor": "paper-2606-01757",
      "title": "PillarDETR: YOLO-Backbone and RT-DETR Head for Real-Time 3D Object Detection",
      "authors": [
        "Smit Kadvani",
        "Shriya Gumber",
        "Kriti Faujdar",
        "Harsh Dave"
      ],
      "abstract": "Real-time 3D object detection is a critical component for the safe operation of autonomous driving systems and robotics. While LiDAR point clouds provide accurate spatial information, processing them efficiently remains a significant challenge. Traditional methods rely on complex 3D convolutions or anchor-based paradigms that struggle to balance detection accuracy with inference speed. In this paper, we propose PillarDETR, a novel end-to-end 3D object detection architecture that combines the efficiency of pillar-based LiDAR encoding with the representational power of modern 2D vision models. Specifically, PillarDETR replaces standard convolutional backbones with a Cross Stage Partial (CSP) network derived from YOLOv8, enabling richer feature extraction from pseudoimages. Furthermore, we discard conventional anchor-based or center-based detection heads in favor of a Real-Time Detection Transformer (RT-DETR) decoder. This hybrid design allows the network to capture global context and directly predict 3D bounding boxes without relying on non-maximum suppression (NMS). Extensive experiments on the KITTI and nuScenes benchmarks demonstrate that PillarDETR achieves a compelling trade-off between mean Average Precision (mAP) and inference latency. Our ablation studies confirm that integrating the YOLOv8 backbone and RT-DETR head yields substantial improvements over the PointPillars baseline, establishing PillarDETR as a highly effective solution for real-time 3D perception.",
      "short_abstract": "Real-time 3D object detection is a critical component for the safe operation of autonomous driving systems and robotics. While LiDAR point clouds provide accurate spatial information, processing them efficiently remains a significant challenge. Traditional...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.01757",
      "pdf_url": "https://arxiv.org/pdf/2606.01757.pdf"
    },
    {
      "id": "2606.01818",
      "anchor": "paper-2606-01818",
      "title": "Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing",
      "authors": [
        "Jiahe Fan",
        "Shaolong Shu",
        "Mingjian Sun",
        "Tiehua Zhang",
        "Bohong Xiao",
        "Hanli Wang",
        "Rui Fan"
      ],
      "abstract": "Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting perception models to new deployment domains remains challenging because pixel-level annotations are expensive to obtain, while source-domain data are often inaccessible due to privacy, security, or ownership constraints. Existing source-free unsupervised domain adaptation methods typically rely on a single pre-trained source model, which makes the adapted perception system vulnerable to source-specific biases and limits its robustness under diverse road layouts, illumination conditions, weather patterns, and traffic conditions. This article presents an unsupervised collaborative domain adaptation (UCDA) framework for driving scene parsing in a source-free setting, which transfers complementary knowledge from multiple pre-trained source models to a unified target model without accessing any original source samples. To compare predictions from independently trained models, UCDA constructs a class-level prototype memory bank and estimates cross-model prediction reliability through prototype similarity, reducing the effect of inconsistent confidence scales across source models. Based on the resulting complementary supervision, UCDA adopts a two-stage transfer strategy: multiple source models are first refined on unlabeled target-domain driving data through collaborative optimization with positive and negative consistency constraints, and their validated expertise is then distilled into a single deployable target model. Comprehensive evaluations on public driving-scene datasets and real-world data collected from an autonomous vehicle platform demonstrate that UCDA effectively consolidates complementary multi-source knowledge, improving target-domain scene parsing reliability and generalization across diverse driving environments.",
      "short_abstract": "Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting perception models to new deployment domains remains challenging because pixel-level annotations are...",
      "published": "2026-06-01",
      "updated": "2026-06-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 8,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 79,
      "arxiv_url": "https://arxiv.org/abs/2606.01818",
      "pdf_url": "https://arxiv.org/pdf/2606.01818.pdf"
    },
    {
      "id": "2606.02641",
      "anchor": "paper-2606-02641",
      "title": "CARVE: Certified Affordable Repair of Vetoed Maneuvers via Envelopes for Interactive Driving",
      "authors": [
        "Yifan Wang"
      ],
      "abstract": "Interactive driving exposes a failure mode that is easy to miss in rule-aware autonomous-driving stacks: a hard-rule margin can be negative for an ego candidate even though a small lawful accommodation by a non-priority agent would restore feasibility. Existing rulebooks, shields, and reachability filters are strong at vetoing unsafe actions, while prediction-based planners model likely responses. Neither returns a runtime proof object that states which bounded multi-agent edit repairs the maneuver, who owns the edit, whether the request is right-of-way affordable, and what ego fallback remains if the request is not observed. We formulate this missing object as *interactive repair certification* and introduce *CARVE*, a prediction-free certificate layer over a finite lattice of ego-owned and agent-owned tactical operators. Agent-owned requests are admissible only inside \\(B_j(s) = β(π_j)α_j^{\\max}(s)\\), a cooperation envelope that separates kinematic reachability from normative priority. The resulting certificate records the binding rule, repair category, repair set, responsibility-weighted cost split, and fallback. On 589 Lanelet2-geometry-grounded INTERACTION replay episodes, CARVE-Greedy accepts 98.64% of initially vetoed maneuvers and recovers 370/378 human-resolved false vetoes, while preserving 589/589 right-of-way respect, zero priority-agent false positives, and 400/400 negative-stress vetoes. We prove certificate soundness, structural right-of-way respect, exact finite-lattice minimality, fallback contingency, and blame-consistency conditions. CARVE does not predict or require another driver's compliance; it certifies whether a proposed interaction is bounded, attributable, and normatively admissible under declared assumptions.",
      "short_abstract": "Interactive driving exposes a failure mode that is easy to miss in rule-aware autonomous-driving stacks: a hard-rule margin can be negative for an ego candidate even though a small lawful accommodation by a non-priority agent would restore feasibility....",
      "published": "2026-05-31",
      "updated": "2026-05-31",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 80,
      "arxiv_url": "https://arxiv.org/abs/2606.02641",
      "pdf_url": "https://arxiv.org/pdf/2606.02641.pdf"
    },
    {
      "id": "2606.01277",
      "anchor": "paper-2606-01277",
      "title": "DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance",
      "authors": [
        "Oskar Natan",
        "Andi Dharmawan",
        "Aufaclav Zatu Kusuma Frisky",
        "Jazi Eko Istiyanto",
        "Jun Miura"
      ],
      "abstract": "Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings. To address this critical safety vulnerability, we propose DeepIPCv3, a novel multi-modal autonomous navigation framework that synergizes the dense 3D spatial geometry of LiDAR point clouds with the microsecond-level asynchronous event streams of a Dynamic Vision Sensor (DVS). We introduce a Transformer-inspired cross-modal attention mechanism to dynamically correlate these distinct modalities, allowing the network to instantaneously prioritize high-speed dynamic updates without sacrificing structural scene awareness. The fused latent representations are then mapped to safe local waypoints and executable control commands via a hybrid policy network that blends heuristic trajectory tracking with direct neural predictions. Due to the severe physical risks associated with live testing of these sudden crossing scenarios, the framework is rigorously evaluated offline using a custom multi-modal dataset collected across both well-illuminated noon and challenging evening conditions. Extensive comparative and ablation studies demonstrate that DeepIPCv3 achieves state-of-the-art predictive performance. By effectively eliminating exposure failures and motion blur, the proposed LiDAR and DVS fusion yields the lowest trajectory and control command errors, enabling highly reactive, mathematically bounded evasive maneuvers regardless of ambient illumination. To support future research, we will release the codes to our GitHub repo at https://github.com/oskarnatan/DeepIPCv3.",
      "short_abstract": "Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings. To address this critical safety...",
      "published": "2026-05-31",
      "updated": "2026-05-31",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 12,
        "evidence": [
          "abstract: perception",
          "title: sensor fusion",
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 80,
      "arxiv_url": "https://arxiv.org/abs/2606.01277",
      "pdf_url": "https://arxiv.org/pdf/2606.01277.pdf"
    },
    {
      "id": "2606.01192",
      "anchor": "paper-2606-01192",
      "title": "PairedGTA: Generating Driving Datasets for Controlled Photometric Shift Analysis",
      "authors": [
        "Andrea Chianese",
        "Giulio Rossolini",
        "Alessandro Biondi",
        "Marco Cococcioni",
        "Giorgio Buttazzo"
      ],
      "abstract": "Evaluating the performance of visual perception systems for autonomous driving is essential to ensure reliable operation across diverse environmental scenarios. Ideally, a balanced and fair analysis across different adverse conditions would require perfectly paired images of the same scene under different weather or illumination changes. This would allow evaluating the effect of photometric shifts independently of geometry and semantic changes. Unfortunately, real-world datasets rarely provide images of the same scene under different environmental conditions, because, normally, camera pose, traffic, and locations of dynamic objects (vehicles, pedestrians, etc.) vary over time, thus yielding only coarsely paired data. To address this challenge, this work introduces a data generation framework based on a high-fidelity game engine for extracting perfectly paired images. By leveraging software APIs that communicate with the GTA game engine, the framework modifies illumination and weather conditions while preserving scene geometry, camera pose, and the identity and placement of dynamic objects. For each sampled location, it procedurally instantiates dynamic entities and renders pixel-aligned images under diverse adverse conditions. The benefit of the proposed generation framework in driving scenarios is demonstrated through a systematic analysis of semantic segmentation models, whose output degradation can be attributed more directly to photometric shifts rather than to uncontrolled semantic or geometric factors.",
      "short_abstract": "Evaluating the performance of visual perception systems for autonomous driving is essential to ensure reliable operation across diverse environmental scenarios. Ideally, a balanced and fair analysis across different adverse conditions would require perfectly...",
      "published": "2026-05-31",
      "updated": "2026-05-31",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 9,
        "evidence": [
          "abstract: perception",
          "abstract: scene segmentation",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 80,
      "arxiv_url": "https://arxiv.org/abs/2606.01192",
      "pdf_url": "https://arxiv.org/pdf/2606.01192.pdf"
    },
    {
      "id": "2606.01164",
      "anchor": "paper-2606-01164",
      "title": "Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends",
      "authors": [
        "Jiuming Liu",
        "Chaojun Ni",
        "Mengmeng Liu",
        "Chensheng Peng",
        "Fangjinhua Wang",
        "Sitian Shen",
        "Marc Pollefeys",
        "Masayoshi Tomizuka",
        "Ayush Tewari",
        "Per Ola Kristensson"
      ],
      "abstract": "With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, benefiting various downstream domains such as game engines, embodied AI, autonomous driving, etc. Through explicitly incorporating user actions into world state transition, recent literature empowers world modeling with interactivity in an action-conditioned video or 3D generation paradigm, further enhancing controllability over world evolutions and facilitating users to freely traverse, manipulate, navigate, and personalize the state evolution. In this paper, we aim to systematically review recent research trends, technical developments, evaluation benchmarks, and also propose future potential directions in interactive world modeling. Specifically, we first summarize recent efforts and trends in terms of application scenarios, world state evolution, and scene modality. Afterwards, we delve into three crucial technical challenges, including action-conditioned controllability, long-horizon interactions and memory, and action-following responsiveness for real-time interactivity. Furthermore, we also thoroughly compare existing benchmarks and metrics in four specific application fields: open-world exploration, game engine, autonomous driving, and robotics. Finally, we discuss several promising future directions in achieving next-generation interactive world modeling. The corresponding repository is publicly available at: https://github.com/liujiuming123/Awesome-Interactive-World-Model.",
      "short_abstract": "With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, benefiting various downstream domains such as game engines, embodied AI, autonomous driving, etc. Through...",
      "published": "2026-05-31",
      "updated": "2026-05-31",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "fresh",
      "age_days": 80,
      "arxiv_url": "https://arxiv.org/abs/2606.01164",
      "pdf_url": "https://arxiv.org/pdf/2606.01164.pdf"
    },
    {
      "id": "2606.01312",
      "anchor": "paper-2606-01312",
      "title": "A Communication-Centric 6G-LLM Architecture for Scalable Tactical Autonomous Defense Vehicle Networks",
      "authors": [
        "Kiran Khurshid",
        "Shumaila Javaid",
        "Nasir Saeed"
      ],
      "abstract": "The integration of Artificial Intelligence (AI) and emerging 6G networks introduces new opportunities for scalable coordination in tactical autonomous vehicle systems. This paper proposes a communication-centric hierarchical architecture for Tactical Autonomous Defense Vehicle Networks (TADVNs) that models the integration of edge-assisted Large Language Model (LLM) reasoning with 6G-enabled connectivity and semantic communication. The framework is designed to improve coordination efficiency, reduce communication overhead, and enhance latency resilience under increasing fleet-scale operation. Unlike conventional task-specific AI pipelines that rely on structured feature processing and rule-based coordination, the proposed approach incorporates semantic abstraction and context-aware decision support within a layered edge-cloud communication architecture. We evaluate communication and coordination performance via Monte Carlo simulations across fleet sizes of 5-30 vehicles under contested network conditions. Results indicate that at a 30-vehicle scale, the 6G-LLM configuration achieves 75.2% latency reduction (29.1 ms vs. 117.5 ms), a 68.7 percentage point increase in mission success rate (82.9% vs. 14.2%), and an 88.6% reduction in communication overhead compared to a 5G-based conventional AI baseline. These findings demonstrate measurable benefits in coordination and communication when semantic reasoning is combined with low-latency 6G connectivity.",
      "short_abstract": "The integration of Artificial Intelligence (AI) and emerging 6G networks introduces new opportunities for scalable coordination in tactical autonomous vehicle systems. This paper proposes a communication-centric hierarchical architecture for Tactical...",
      "published": "2026-05-31",
      "updated": "2026-05-31",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 10,
        "evidence": [
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 80,
      "arxiv_url": "https://arxiv.org/abs/2606.01312",
      "pdf_url": "https://arxiv.org/pdf/2606.01312.pdf"
    },
    {
      "id": "2606.07626",
      "anchor": "paper-2606-07626",
      "title": "Eyes All Around: Design and Analysis of 360-Degree LiDAR Perception Using Equivariant Feature Learning in Unstructured Traffic",
      "authors": [
        "Pranav Darshan",
        "Raghuveer Narayanan Rajesh",
        "M Uttara Kumari"
      ],
      "abstract": "Perception in dense, unstructured urban traffic remains a major challenge for autonomous driving because of the wide variety of road users, frequent occlusions, irregular motion patterns, and the lack of standardized road layouts. Although recent LiDAR based 3D object detectors have shown strong performance in structured driving scenarios, most are developed and evaluated for limited field of view settings, and their behavior under full surround 360-degree sensing is still not well understood. This paper studies a 360-degree LiDAR perception pipeline for autonomous driving, with particular attention to panoramic sensing, azimuthal sector wise spatial processing, and transformation equivariant feature extraction in complex urban scenes. The paper presents a practical 360-degree perception framework that combines sector wise panoramic processing with rotation equivariant sparse convolutions and evaluates its behavior on a custom Ouster OS0 LiDAR dataset collected across diverse Indian urban traffic conditions. The results show generally stable detection across several object classes, with the strongest performance for cars at 92.02/90.51, buses at 80.53/76.34, and trucks at 78.59/74.16, while lower scores for pedestrians at 67.45/61.02, cyclists at 73.21/69.54, and motorcyclists at 71.20/68.13 reflect the greater difficulty of detecting smaller and more variable road users in dense urban scenes.",
      "short_abstract": "Perception in dense, unstructured urban traffic remains a major challenge for autonomous driving because of the wide variety of road users, frequent occlusions, irregular motion patterns, and the lack of standardized road layouts. Although recent LiDAR based...",
      "published": "2026-05-30",
      "updated": "2026-05-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 24,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "fresh",
      "age_days": 81,
      "arxiv_url": "https://arxiv.org/abs/2606.07626",
      "pdf_url": "https://arxiv.org/pdf/2606.07626.pdf"
    },
    {
      "id": "2606.00857",
      "anchor": "paper-2606-00857",
      "title": "From Cues to Horizons: Dynamic Risk Horizon Profiling for Trajectory Prediction",
      "authors": [
        "Xinyi Ning",
        "Zilin Bian",
        "Dachuan Zuo",
        "Semiha Ergan",
        "Kaan Ozbay"
      ],
      "abstract": "Accurate and reliable vehicle trajectory prediction is essential for safe autonomous driving. Recent studies have incorporated safety risk into trajectory prediction to quantify dangers posed by surrounding agents. However, most risk-aware approaches use past risk information as a secondary signal to help guide decisions, overlooking its future evolution and uncertainty. In this paper, we propose a risk horizon profiling (RHP) module that incorporates a continuous, learnable potential field model for risk-aware trajectory prediction. The RHP module calculates the spatial-temporal proximity of surrounding objects to profile risk distributions across future horizons, which supports better trajectory prediction by adaptively identifying what human drivers perceive as critical moments. We evaluate our method on two datasets from different driving settings, highD for highway corridors and SHRP2 for urban streets, which cover diverse risk scenarios including safe, near-crash, and crash events. Compared to the baseline methods, our framework achieves a 25.0\\% reduction in 5s RMSE on the highD dataset and a 29.1\\% reduction in 5s minFDE on SHRP2. These results indicate strong performance for both short and long horizon prediction and robust generalization across highway and urban scenarios. The proposed method enables more realistic AV path planning and strategic selection, thereby supporting safer autonomous driving and more advanced driver-assistance systems. The source code for this work is available at: https://github.com/bilab-nyu/RHP",
      "short_abstract": "Accurate and reliable vehicle trajectory prediction is essential for safe autonomous driving. Recent studies have incorporated safety risk into trajectory prediction to quantify dangers posed by surrounding agents. However, most risk-aware approaches use past...",
      "published": "2026-05-30",
      "updated": "2026-05-30",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 16,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 81,
      "arxiv_url": "https://arxiv.org/abs/2606.00857",
      "pdf_url": "https://arxiv.org/pdf/2606.00857.pdf"
    },
    {
      "id": "2606.00688",
      "anchor": "paper-2606-00688",
      "title": "Shape-Prior-Based Point Cloud Completion for Single-Stage Fully Sparse 3D Object Detection",
      "authors": [
        "Kaizheng Wang",
        "Mingqian Ji",
        "Jian Yang",
        "Shanshan Zhang"
      ],
      "abstract": "Single-stage fully sparse 3D object detectors rely on point clouds data to detect objects in autonomous driving scenarios. However, the sparsity and incompleteness of point clouds significantly limit the performance of 3D object detection. To address this issue, this paper proposes a point clouds completion method specifically designed for single-stage fully sparse detectors. The entire shape-prior-based completion process consists of two consecutive steps. In the first step, we design a novel Instance Selection module, which is capable of identifying point clouds corresponding to foreground objects even when the baseline model does not generate proposals, while effectively ignoring the point clouds of background regions. In the second step, we introduce a novel Alignment-Based Point Completion module, which aligns the point clouds of foreground objects with prototypes in terms of both their centers and orientations. Subsequently, points are selected from the prototype to fill in the missing parts of the foreground object. We evaluated our method on two single-stage fully sparse detectors using the KITTI dataset. The experimental results demonstrate that the proposed method significantly improves the detection performance, confirming its effectiveness and generalizability.",
      "short_abstract": "Single-stage fully sparse 3D object detectors rely on point clouds data to detect objects in autonomous driving scenarios. However, the sparsity and incompleteness of point clouds significantly limit the performance of 3D object detection. To address this...",
      "published": "2026-05-30",
      "updated": "2026-05-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 28,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "title: point cloud",
          "abstract: point cloud"
        ]
      },
      "recency": "fresh",
      "age_days": 81,
      "arxiv_url": "https://arxiv.org/abs/2606.00688",
      "pdf_url": "https://arxiv.org/pdf/2606.00688.pdf"
    },
    {
      "id": "2606.00519",
      "anchor": "paper-2606-00519",
      "title": "DriveAnchor: Progressive Anchor-based Flow Learning for Autonomous Driving Planning",
      "authors": [
        "Limin Yan",
        "Haoyun Tang",
        "Yutao Qiu",
        "Hongqing Liu",
        "Haoyu Xu"
      ],
      "abstract": "We present DriveAnchor, a three-stage framework for autonomous driving planning that achieves behavioral diversity, controllability, and safety in a composable pipeline. Demonstration Flow Pretraining replaces the unstructured Gaussian prior with a vocabulary of 2,398 trajectory shapes constructed by farthest-point sampling, structurally grounding behavioral diversity in vocabulary coverage. Guided Flow Post-training jointly post-trains an Energy Field module with flow matching (FM), conditioning the Energy Field on static road geometry alone, to relocate anchors toward user-specified corridor polygons before flow generation, adding controllability without differentiable guidance; after Stage 2, new corridor presets require only Energy Field updates, not FM retraining. Reward-Refined Flow Fine-tuning applies zeroth-order reinforcement learning to align each anchor's output with collision-avoidance objectives: because the flow-matching model is a deterministic feedforward network in single-step mode, each anchor uniquely determines the output trajectory, reducing reward optimization to a direction search in anchor space without log-likelihood computation or ODE-to-SDE conversion. Evaluated on approximately 2 million held-out driving scenarios, DriveAnchor reduces near-range collision rates by 89% and improves mean reward by 32% without degradation in imitation accuracy, with 2.06 ms inference on NVIDIA Drive Orin. DriveAnchor has been validated through real-world vehicle testing, confirming its practicality for production deployment.",
      "short_abstract": "We present DriveAnchor, a three-stage framework for autonomous driving planning that achieves behavioral diversity, controllability, and safety in a composable pipeline. Demonstration Flow Pretraining replaces the unstructured Gaussian prior with a vocabulary...",
      "published": "2026-05-30",
      "updated": "2026-05-30",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 5,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 81,
      "arxiv_url": "https://arxiv.org/abs/2606.00519",
      "pdf_url": "https://arxiv.org/pdf/2606.00519.pdf"
    },
    {
      "id": "2606.00471",
      "anchor": "paper-2606-00471",
      "title": "MUSCLE-NET: Predicted-Multiscale-Aware Network for Pedestrian Trajectory Forecasting",
      "authors": [
        "Yu Liu",
        "Ming Huang",
        "Xiao Ren",
        "Zhijie Liu",
        "Youfu Li",
        "He Kong"
      ],
      "abstract": "Accurate pedestrian trajectory prediction is essential for safe navigation in autonomous driving and intelligent transportation systems. Despite substantial progress made by recent methods, most existing approaches are limited in fully exploiting diverse observations and often overlook the scale dependency of future motion, treating multiscale features uniformly regardless of underlying motion dynamics. This limits their robustness across diverse pedestrian behaviors. To address these challenges, we propose a Predicted-MUltiSCale-Aware Network (MUSCLE-NET) for Pedestrian Trajectory Forecasting that integrates complementary multimodal cues with scale-adaptive prediction mechanisms. The proposed framework is built upon a Multiscale Multimodal Feature Extraction (MMFE) module, which combines multiscale representation, modality-aware recalibration, and directional cross-modal fusion to construct semantically aligned representations from bounding boxes, velocities, and pose information. Building on these features, a Multiscale Enhanced Hierarchical Prediction (MEHP) module performs prediction-aware future-motion refinement via a probabilistic coarse predictor, scale-aligned fusion, and progressive refinement, adaptively selecting scale-relevant cues to mitigate spatial drift. Extensive experiments on the JAAD and PIE benchmarks demonstrate that the proposed MUSCLE-Net achieves competitive performance and consistent gains compared with state-of-the-art trajectory prediction methods.",
      "short_abstract": "Accurate pedestrian trajectory prediction is essential for safe navigation in autonomous driving and intelligent transportation systems. Despite substantial progress made by recent methods, most existing approaches are limited in fully exploiting diverse...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 11,
        "evidence": [
          "abstract: trajectory prediction",
          "title: forecasting",
          "abstract: forecasting",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2606.00471",
      "pdf_url": "https://arxiv.org/pdf/2606.00471.pdf"
    },
    {
      "id": "2606.00416",
      "anchor": "paper-2606-00416",
      "title": "4D Radar Meets LiDAR and Camera: Cooperative Perception under Adverse Weather",
      "authors": [
        "Melih Yazgan",
        "Iramm Hamdard",
        "Qiyuan Wu",
        "J. Marius Zoellner"
      ],
      "abstract": "Cooperative perception is important for autonomous driving but remains fragile when cameras and LiDAR degrade in adverse weather. We address this challenge by integrating 4D imaging radar as a weather-robust modality into collaborative perception and introducing a Doppler-guided spatial attention mechanism for multi-agent fusion. Our approach extends two representative backbones: a radar-camera pipeline where radar substitutes LiDAR, and a LiDAR-radar pipeline where radar complements LiDAR. To support evaluation, we release radar-augmented benchmarks, OPV2V-R and Adver-City-R, with physics-based LiDAR degradation. Experiments show strong robustness gains in fog and rain, including substantial improvements when radar replaces degraded LiDAR. Additional validation on MAN TruckScenes demonstrates transfer beyond simulation. Overall, our results highlight 4D imaging radar as a robust modality for all-weather collaborative perception. Dataset and code are available at: https://url.fzi.de/SlimComm.",
      "short_abstract": "Cooperative perception is important for autonomous driving but remains fragile when cameras and LiDAR degrade in adverse weather. We address this challenge by integrating 4D imaging radar as a weather-robust modality into collaborative perception and...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 44,
        "margin": 32,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "title: LiDAR",
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2606.00416",
      "pdf_url": "https://arxiv.org/pdf/2606.00416.pdf"
    },
    {
      "id": "2606.00267",
      "anchor": "paper-2606-00267",
      "title": "StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement",
      "authors": [
        "Junwon Seo",
        "Sushant Veer",
        "Ran Tian",
        "Wenhao Ding",
        "Apoorva Sharma",
        "Karen Leung",
        "Edward Schmerling",
        "Marco Pavone",
        "Andrea Bajcsy"
      ],
      "abstract": "Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over futures, policy evaluation and improvement typically rely on nominal imaginations, which can miss high-impact outcomes of robot actions unless prohibitively many samples are drawn. To enable robust policy evaluation and improvement over WM imaginations, we propose StressDream, which steers imaginations toward high-impact yet plausible outcomes specified at inference time by optimizing the initial noise of diffusion-based WMs. However, optimizing high-dimensional noise is challenging: the optimization must reason about nuanced, scene-dependent target events in generated videos while avoiding out-of-distribution (OOD) noise that yields implausible imaginations. We address this with two complementary objectives: a semantic objective with a Vision-Language Model that provides informative gradients by reasoning about the generated video, and a plausibility objective that prevents the optimized noise from drifting OOD. With state-of-the-art video world models for autonomous driving and robotic manipulation, we show that StressDream effectively steers imaginations toward high-impact yet plausible outcomes specified by text at inference time, such as task failures, enabling robust policy evaluation and improvement by identifying actions whose plausible futures include undesirable outcomes. Video results are available at https://junwon.me/StressDream/.",
      "short_abstract": "Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over futures, policy evaluation and improvement typically rely on...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: title: law or policy",
          "Human Factors & Policy: abstract: law or policy",
          "Prediction & World Models: title: world model",
          "Prediction & World Models: abstract: world model"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2606.00267",
      "pdf_url": "https://arxiv.org/pdf/2606.00267.pdf"
    },
    {
      "id": "2606.00191",
      "anchor": "paper-2606-00191",
      "title": "Safe2Drive: Evaluating Safe Driving Behaviors of E2E Autonomous Driving Models",
      "authors": [
        "Nishad Sahu",
        "Kalpana Panda",
        "Congyuan Yu",
        "Changzhong Qian",
        "Shounak Sural",
        "Ragunathan Rajkumar"
      ],
      "abstract": "Recent end-to-end (E2E) autonomous driving policies achieve high driving scores in closed-loop simulations. Yet it remains unclear whether these policies handle common safety-critical scenarios. We present Safe2Drive (S2D), a set of Bench2Drive-aligned scenario extensions focused on three frequent families of road hazards: work zones, pedestrian jaywalking, and occluded vulnerable road users (VRUs). Safe2Drive adds 100 common but challenging scenarios and introduces SafeDriving Score (SDS), a safety-centric metric that augments prior evaluators with pre-crash braking, work zone-object contact, lane centering, and smoothness checks. Evaluating two state-of-the-art policies (LEAD and SimLingo) on S2D, we find that their driving scores drop sharply relative to their reported Bench2Drive baselines (LEAD: from 94.70 DS on Bench2Drive to 39.95 DS on S2D; SimLingo: from 85.07 DS on Bench2Drive to 41.00 DS on S2D) and that SDS on S2D is low (11.85 for LEAD and 15.27 for Sim-Lingo). These results are consistent with brittle safe-driving behaviors such as poor work-zone understanding, red-light violations, and late or absent braking for pedestrians. This study highlights a lack of safe behavioral reasoning in E2E models even when tested on CARLA towns that are part of the training set. We plan to release the code and videos for all 100 S2D scenarios.",
      "short_abstract": "Recent end-to-end (E2E) autonomous driving policies achieve high driving scores in closed-loop simulations. Yet it remains unclear whether these policies handle common safety-critical scenarios. We present Safe2Drive (S2D), a set of Bench2Drive-aligned...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2606.00191",
      "pdf_url": "https://arxiv.org/pdf/2606.00191.pdf"
    },
    {
      "id": "2605.31572",
      "anchor": "paper-2605-31572",
      "title": "nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving",
      "authors": [
        "Zhiyu Huang",
        "Johnson Liu",
        "Rui Song",
        "Zewei Zhou",
        "Ruining Yang",
        "Yun Zhang",
        "Tianhui Cai",
        "Hanyin Zhang",
        "Mingxuan Gao",
        "Valeria Xu",
        "Jiali Chen",
        "Yishan Shen",
        "Yiluan Guo",
        "Tony",
        "Qi",
        "Jiaqi Ma"
      ],
      "abstract": "Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. However, existing AD datasets and benchmarks mainly target perception, prediction, or planning, and provide limited supervision for reasoning over realistic long-tail driving scenes. We introduce nuReasoning, a large-scale real-world dataset and benchmark for reasoning-centric AD. Following the lineage of nuScenes and nuPlan, nuReasoning advances real-world AD datasets and benchmarks toward reasoning in long-tail driving scenarios. The dataset contains 20,000 clips, each 20 seconds long, collected across multiple cities, with synchronized multi-camera images, LiDAR data, HD maps, object annotations, and human-verified reasoning annotations spanning Spatial Reasoning, Decision Reasoning, and Counterfactual Reasoning. Unlike prior datasets that focus primarily on visual question answering, nuReasoning supports both reasoning evaluation and planning evaluation, enabling a direct study of how reasoning supervision affects driving performance. Experiments show that fine-tuning VLMs on nuReasoning substantially improves driving-specific question answering, while incorporating reasoning supervision into VLA training improves planning performance even when textual reasoning outputs are disabled at inference time. These results establish nuReasoning as a foundation for evaluating and improving robust, interpretable, reasoning-driven AD systems in realistic long-tail settings.",
      "short_abstract": "Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. However, existing AD datasets and benchmarks mainly...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark",
        "VLA",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 2,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2605.31572",
      "pdf_url": "https://arxiv.org/pdf/2605.31572.pdf"
    },
    {
      "id": "2605.31476",
      "anchor": "paper-2605-31476",
      "title": "IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving",
      "authors": [
        "Chenghao Zhang",
        "Timin Li",
        "Dongmei Li"
      ],
      "abstract": "End-to-end autonomous driving has emerged as a compelling paradigm for learning planning directly from sensor observations, while recent world-model-based approaches further enrich this paradigm by enabling explicit reasoning about how the scene may evolve in the future. Yet future prediction alone does not guarantee better planning unless the predicted evolution can be converted into planning-relevant trajectory updates. Many current methods still forecast future scene states without explicitly decoding the motion implications hidden in state transitions. As a result, future reasoning often remains descriptively useful but only weakly coupled to executable motion generation. To address this limitation, we propose \\mathbf{IDOL}, an inverse-dynamics-guided future prediction framework for world-model-based end-to-end planning in latent BEV space, where inverse dynamics serves as the key bridge between future prediction and trajectory optimization. IDOL first predicts multiple future latent scene states with a BEV world model, then applies an inverse dynamics model to adjacent latent futures to decode transition-aware trajectory features and recover planning-relevant motion deltas that explain how the latent world evolves over time. These inverse-dynamics-derived signals are used to optimize the planned trajectory, turning future forecasting from passive scene anticipation into actionable planning guidance. A lightweight closed-loop refinement module further improves long-horizon consistency by reusing the optimized trajectory for another round of future-aware reasoning. By introducing inverse dynamics into latent future reasoning, IDOL tightens the coupling between world modeling and planning. Extensive experiments on the NAVSIM v1 and NAVSIM v2 benchmarks show that IDOL achieves state-of-the-art performance among comparable methods.",
      "short_abstract": "End-to-end autonomous driving has emerged as a compelling paradigm for learning planning directly from sensor observations, while recent world-model-based approaches further enrich this paradigm by enabling explicit reasoning about how the scene may evolve in...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 9,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2605.31476",
      "pdf_url": "https://arxiv.org/pdf/2605.31476.pdf"
    },
    {
      "id": "2605.31256",
      "anchor": "paper-2605-31256",
      "title": "Before Parc Fermé: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving",
      "authors": [
        "Luca Benfenati",
        "Ali Azimi",
        "Matteo Risso",
        "Fabio Carapellese",
        "Daniele Jahier Pagliari",
        "Alessio Burrello"
      ],
      "abstract": "Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but their memory and generation latency make real-time deployment difficult. Pruning can reduce these costs, but for controllers that undergo multiple pre- and post-training phases, the crucial question is not only how much to prune, but when pruning should occur. In this work, we propose Before Parc Fermé (BPF), a pruning strategy performed during RL that compresses embodied LLM controllers while they are still being optimized for closed-loop behavior. This allows pruning decisions to account for the task-specific supervision and closed-loop feedback that shape the final controller. We propose two variants: BPF-RL, which performs iterative pruning during RL by removing part of the model at predefined training intervals, and BPF-SFT/RL, which first prunes part of the model structure during SFT and then further compresses it during RL using the same iterative strategy as BPF-RL until the target pruning ratio is reached. We evaluate BPF on RobotxR1, an LLM-based autonomous-driving control pipeline, using an established LLM pruning framework (LLM-Pruner), and compare it against post-training pruning, post-training pruning with RL recovery, SFT-stage pruning, and smaller dense models from the same family. Our results show that BPF provides the best task-performance vs. memory and throughput trade-off among the considered pruning strategies. When compressing the larger RobotxR1 models, BPF-SFT/RL achieves a $1.69\\times$ better size-end-to-end performance trade-off than directly selecting a smaller dense model from the same family, measured as removed parameters per lost percentage point of control adaptability. On the Jetson AGX Orin mounted on the target robotic platform, the compact models improve decode throughput by up to $27\\%$.",
      "short_abstract": "Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but their memory and generation latency make real-time deployment difficult. Pruning can reduce these costs, but...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: controller",
          "Control & Vehicle Dynamics: abstract: control",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2605.31256",
      "pdf_url": "https://arxiv.org/pdf/2605.31256.pdf"
    },
    {
      "id": "2605.31116",
      "anchor": "paper-2605-31116",
      "title": "NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving",
      "authors": [
        "Jiahui Li",
        "Jiawei Sun",
        "Zixiang Ren",
        "Ming Liu",
        "Jiamin Shi",
        "Ruiteng Zhao",
        "Zhiyang Liu",
        "Liying Liu",
        "Zuoguan Wang",
        "Kaidi Yang"
      ],
      "abstract": "Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compact scene tokens for downstream trajectory generation and scoring. While these scene tokens form a compact visual bottleneck for the planner, they receive supervision solely from the planning objective, providing limited constraints on the encoded visual information. To address this limitation, we introduce Neural Token Reconstruction (NTR), a representation learning framework to directly constrain the compact scene-token bottleneck in perception-free driving. NTR introduces a self-distillation masked latent reconstruction objective that reconstructs masked patch-level latent features using only compact scene tokens as reconstruction memory. This forces reconstruction gradients to pass exclusively through the scene-token bottleneck, encouraging scene tokens to preserve richer and less redundant visual representations for planning. We further introduce semantic priors derived from foundation-model annotations as a weak semantic interface biasing reconstruction targets toward driving-related structures without introducing explicit perception heads. All auxiliary reconstruction components are removed at inference time, leaving the deployed planner unchanged. NTR achieves state-of-the-art performance on three public autonomous driving benchmarks, including 8.0461 RFS on Waymo E2E and 94.1 PDMS / 90.9 EPDMS on NavSim1&2. The learned scene tokens exhibit lower pairwise redundancy and higher effective rank, indicating that effective bottleneck supervision improves both compact visual representation learning and planning performance.",
      "short_abstract": "Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compact scene tokens for downstream trajectory generation and scoring. While these scene tokens form a compact...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2605.31116",
      "pdf_url": "https://arxiv.org/pdf/2605.31116.pdf"
    },
    {
      "id": "2605.31041",
      "anchor": "paper-2605-31041",
      "title": "Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?",
      "authors": [
        "Jingtao He",
        "Hongliang Lu",
        "Xiaoyun Qiu",
        "Yixuan Wang",
        "Xinhu Zheng"
      ],
      "abstract": "Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal architectures for jointly modeling perception and planning. However, how current VLA-based driving behavior is grounded in visual information remains poorly understood. Existing evaluation protocols mainly focus on aggregate performance metrics, lacking structured and practical diagnostics to quantify visual-behavior dependency. In this work, we introduce a structured multi-level visual perturbation framework to analyze visual-behavior dependency in VLA-based driving models systematically. The framework organizes controlled visual perturbations along three complementary dimensions: channellevel degradation, information-level disruption, and structurelevel modification. We apply it to VLA-based driving systems and evaluate behavioral responses under both open-loop trajectory prediction and interactive closed-loop safety evaluation. Experimental results reveal evaluation-dependent dependency patterns and uneven visual grounding across abstraction levels. These findings call for more structured analyses and principled design of VLA driving models to better understand how visual information shapes behavior and develop safer, more robust systems.",
      "short_abstract": "Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal architectures for jointly modeling perception and planning. However, how current VLA-based driving behavior is...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 17,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2605.31041",
      "pdf_url": "https://arxiv.org/pdf/2605.31041.pdf"
    },
    {
      "id": "2605.30983",
      "anchor": "paper-2605-30983",
      "title": "Can BEV Perception Gracefully Degrade under Sensor Failures?",
      "authors": [
        "Haifa Zhang",
        "Yijing Wang",
        "Haoyu Wang",
        "Zheng Li",
        "Zhiqiang Zuo"
      ],
      "abstract": "Despite the remarkable success of multi-modal bird's-eye view (BEV) perception in autonomous driving, current systems exhibit a critical vulnerability: existing fusion mechanisms are highly brittle to sensor corruptions, often causing catastrophic performance degradation. This vulnerability largely stems from the fact that standard fusion frameworks typically integrate multi-modal representations in a static manner, leading to a precipitous performance collapse under missing or corrupted modalities. In contrast, we show that graceful degradation is achievable through active modality reliability assessment. To this end, we present Grace-BEV, a lightweight and plug-and-play framework that enforces active reliability awareness during multi-modal fusion. Instead of relying on computationally expensive cross-modal interactions, Grace-BEV leverages the aligned BEV space to explicitly assess modality trustworthiness via a TrustGate Router and dynamically recalibrate feature integration using the FailSafe Fusion Block. Furthermore, we devise a Three-Phase Training strategy with Modality Dropout to prevent modality dominance and encourage balanced cross-modal learning under unreliable inputs. Extensive experiments on nuScenes-R and nuScenes-C show that Grace-BEV maintains robust performance across diverse corruption settings. Notably, under catastrophic LiDAR failures where standard baselines collapse to 0.0% mean Average Precision (mAP), Grace-BEV restores performance to as high as 34.7% mAP. Moreover, it improves clean accuracy by up to 1.4%, achieving a strong trade-off between robustness and efficiency.",
      "short_abstract": "Despite the remarkable success of multi-modal bird's-eye view (BEV) perception in autonomous driving, current systems exhibit a critical vulnerability: existing fusion mechanisms are highly brittle to sensor corruptions, often causing catastrophic performance...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 13,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: LiDAR",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2605.30983",
      "pdf_url": "https://arxiv.org/pdf/2605.30983.pdf"
    },
    {
      "id": "2605.30939",
      "anchor": "paper-2605-30939",
      "title": "IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road Segmentation",
      "authors": [
        "Bingtao Wang",
        "Daojie Peng",
        "Fulong Ma",
        "Jun Ma",
        "Liang Zhang"
      ],
      "abstract": "Semantic road segmentation is important for autonomous driving, but existing methods suffer severe performance degradation under low-light conditions. Many existing multi-modal fusion methods do not explicitly adapt to illumination-dependent changes in modality reliability, which can propagate degraded RGB features into the fused representation at night. We propose IAF-Net (Illumination-Adaptive Fusion Network), an end-to-end framework with illumination-adaptive fusion for robust road segmentation across different lighting conditions. It dynamically adjusts fusion weights of RGB and geometric features via the core Illumination-Adaptive Fusion (IAF) module, and enhances low-light feature selection with a brightness-modulated attention decoder. We also construct two dedicated datasets: nuScenes Nighttime Road Segmentation (nuScenes-NRS) and CARLA Multi-Weather Road Segmentation (CARLA-MWRS). Experiments on nuScenes-NRS show state-of-the-art overall performance among the compared methods, while CARLA-MWRS further validates robustness across adverse weather conditions. Ablation studies on a 40% training subset further highlight the importance of the IAF module, which provides the largest individual gain of 0.70% in MaxF.",
      "short_abstract": "Semantic road segmentation is important for autonomous driving, but existing methods suffer severe performance degradation under low-light conditions. Many existing multi-modal fusion methods do not explicitly adapt to illumination-dependent changes in...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2605.30939",
      "pdf_url": "https://arxiv.org/pdf/2605.30939.pdf"
    },
    {
      "id": "2606.00372",
      "anchor": "paper-2606-00372",
      "title": "LFA: Layer Feature Attention for Run-Time Introspection of 2D Object Detectors in Automated Driving",
      "authors": [
        "Mert Keser",
        "Alois Knoll"
      ],
      "abstract": "Reliable object detection is critical for automated driving, yet even state-of-the-art detectors inevitably make errors that can compromise safety. Introspection methods that predict detector failures enable safer deployment by triggering fallback mechanisms or alerting human operators. However, existing approaches rely solely on last-layer features or hand-crafted statistics, discarding valuable information from earlier layers that capture different levels of visual abstraction. We propose Layer Feature Attention (LFA), a lightweight introspection method that learns to aggregate features from multiple backbone layers through an attention mechanism. Our key insight is that detection errors manifest differently across feature hierarchies: low-level layers capture fine-grained details essential for detecting small or occluded objects, while high-level layers encode semantic information for scene understanding. LFA learns layer importance weights end-to-end, enabling both improved error prediction and interpretable analysis of which feature levels are most indicative of detector failures. Extensive experiments on KITTI and BDD100K demonstrate that LFA achieves state-of-the-art introspection performance, outperforming single-layer baselines across multiple detector architectures.",
      "short_abstract": "Reliable object detection is critical for automated driving, yet even state-of-the-art detectors inevitably make errors that can compromise safety. Introspection methods that predict detector failures enable safer deployment by triggering fallback mechanisms...",
      "published": "2026-05-29",
      "updated": "2026-05-29",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: object detection",
          "abstract: scene understanding"
        ]
      },
      "recency": "fresh",
      "age_days": 82,
      "arxiv_url": "https://arxiv.org/abs/2606.00372",
      "pdf_url": "https://arxiv.org/pdf/2606.00372.pdf"
    },
    {
      "id": "2606.00133",
      "anchor": "paper-2606-00133",
      "title": "World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications",
      "authors": [
        "Arif Hassan Zidan",
        "Yi Pan",
        "Hanqi Jiang",
        "Ruiyu Yan",
        "Wei Ruan",
        "Zihao Wu",
        "Lifeng Chen",
        "Weihang You",
        "Xinliang Li",
        "Bowen Chen",
        "Huawen Hu",
        "Peilong Wang",
        "Sizhuang Liu",
        "Jing Zhang",
        "Siyuan Li",
        "Zhengliang Liu",
        "Yu Bao",
        "Lin Zhao",
        "Lichao Sun",
        "Dajiang Zhu",
        "Xiang Li",
        "Jinglei Lv",
        "Quanzheng Li",
        "Wei Liu",
        "Tianming Liu"
      ],
      "abstract": "World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artificial general intelligence, enabling agents to predict, plan, and reason within learned representations. Despite rapid progress across reinforcement learning, robotics, autonomous driving, and video generation, the field lacks a unified framework integrating its diverse architectural choices, training methods, reasoning mechanisms, and application settings. This survey addresses that gap with a multi-axis taxonomy organized along four dimensions: (i) architecture, encompassing representation format, dynamics formulation, input modality, learning paradigm, and downstream application; (ii) methodological family, including state-space and recurrent approaches, transformer-based models, diffusion-based generators, physics-informed networks, and language-augmented multimodal systems; (iii) reasoning strategy, covering imagination-based planning, latent policy learning, counterfactual reasoning, and planning under uncertainty; and (iv) application domain, spanning robotics, autonomous driving, video prediction, multimodal agents, reinforcement learning, scientific modeling, medical imaging, educational measurement, and business and finance. Tracing the field from early cognitive-science foundations to milestone systems such as PlaNet, the Dreamer family, MuZero, Sora, Cosmos, and Genie, we examine how these dimensions interact and highlight the recent convergence of chain-of-thought reasoning with world-model imagination. We review evaluation protocols and benchmarks, identify persistent challenges such as compounding prediction errors, sim-to-real transfer, and fragmented evaluation, and outline future directions toward unified multimodal world models, foundation-scale interactive simulators, and safe deployment in safety-critical domains.",
      "short_abstract": "World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artificial general intelligence, enabling agents to predict, plan, and reason within learned representations....",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "World Model",
        "Reinforcement Learning",
        "Survey"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 12,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2606.00133",
      "pdf_url": "https://arxiv.org/pdf/2606.00133.pdf"
    },
    {
      "id": "2605.30576",
      "anchor": "paper-2605-30576",
      "title": "Uncertainty-Aware and Temporally Regulated Expert Advice in Reinforcement Learning for Autonomous Driving",
      "authors": [
        "Ahmed Abouelazm",
        "Felix Klingebiel",
        "Philip Schörner",
        "J. Marius Zöllner"
      ],
      "abstract": "Exploration in reinforcement learning for autonomous driving is inherently unsafe: agents must experience novel behaviors to learn, yet exploration can lead to collisions or off-road driving. We propose an uncertainty-aware framework that leverages expert advice to guide exploration while avoiding long-term dependence. Advice is triggered when epistemic or aleatoric uncertainty exceeds adaptive thresholds derived from rolling buffers, ensuring advice evolves with the agent's confidence. A commitment-cooldown strategy with a stochastic early-stop heuristic regulates the duration and frequency of guidance, exposing the agent to coherent maneuvers without exhausting the advice budget. Expert and agent experiences are combined in a shared replay buffer within an off-policy implicit quantile network (IQN) backbone, enabling efficient reuse of expert trajectories. Experiments in CARLA show that our method outperforms the IQN baseline, improving success by 5-7% and reducing failures, demonstrating that risk-sensitive uncertainty coupled with regulated expert integration enables safer and more efficient exploration for sensor-based RL policy learning in unsignalized intersection navigation.",
      "short_abstract": "Exploration in reinforcement learning for autonomous driving is inherently unsafe: agents must experience novel behaviors to learn, yet exploration can lead to collisions or off-road driving. We propose an uncertainty-aware framework that leverages expert...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 6,
        "evidence": [
          "title: uncertainty",
          "abstract: uncertainty",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.30576",
      "pdf_url": "https://arxiv.org/pdf/2605.30576.pdf"
    },
    {
      "id": "2605.30328",
      "anchor": "paper-2605-30328",
      "title": "Supercharging Thermal Gaussian Splatting with Depth Estimation",
      "authors": [
        "Manoj Biswanath",
        "Chenxin Cai",
        "Hannah Schieber",
        "Daniel Roth",
        "Benjamin Busam"
      ],
      "abstract": "Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content for 3D reconstruction, other modalities like thermal or depth can enable additional information on the environment. Lately, novel view synthesis methods like 3D Gaussian Splatting have started using multiple modalities to further boost their performance. But fusing or combining multimodal data can make the process slower and can bring in additional challenges. Therefore, our project aims to use single modality based on thermal infrared domain, by removing the reliance on visible light as much as possible. This single modality can be expected to be faster as it does not rely on multimodal data. We propose a method, Thermal-to-Depth Gaussian Splatting (TDg), that uses only thermal images and depth estimation in its architecture to derive the radiance fields. Our TDg method outperforms the MSMG (Multiple Single-Modal Gaussians) baseline in most cases on our test datasets, RGBT-Scenes and ThermalMix. On average, the rendering quality metrics such as learned perceptual image patch similarity (LPIPS), structural similarity index measure (SSIM), and peak signal-to-noise ratio (PSNR) of TDg are 1.12%, 0.034%, and 0.01% better than the baseline MSMG values. It also reduces the training time significantly, by 12 mins 47 secs (55% improvement). Overall, our method is successful in deriving these thermal radiance fields, which can ultimately have several applications, such as identifying heat sources critical in surveillance, search or rescue operations, and industrial inspections where temperature is widely used to monitor machines.",
      "short_abstract": "Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content for 3D reconstruction, other modalities like thermal or depth can enable additional information on the...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 10,
        "evidence": [
          "title: depth estimation",
          "abstract: depth estimation"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.30328",
      "pdf_url": "https://arxiv.org/pdf/2605.30328.pdf"
    },
    {
      "id": "2605.30239",
      "anchor": "paper-2605-30239",
      "title": "CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction",
      "authors": [
        "Xin Dong",
        "Weijian Deng",
        "Lihan Zhang",
        "Tianru Dai",
        "Wenfeng Deng",
        "Yansong Tang"
      ],
      "abstract": "Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static and holistic visual fidelity, with limited support for multi-object interaction. We argue that an interaction-ready reconstruction should anticipate potential scene changes and preserve geometric completeness, visual quality, multi-object spatial relationship, and physical plausibility under potential interactions. To this end, motivated by the causal intervention, we propose CA-World, an efficient framework that integrates counterfactual alignment learning into a decoupling-reintegration reconstruction pipeline. Specifically, we formulate foreground-background decoupling as a visual intervention, separate object generation and background inpainting as counterfactual generation, and scene reintegration as an inverse intervention. According to counterfactual consistency, reversing the intervention should recover the factual world, motivating three alignment objectives between the reintegrated and original scenes: appearance, spatial, and physical consistency. This formulates interaction-ready reconstruction as counterfactual alignment learning with direct supervision. Moreover, leveraging the locality of object-level interventions, CA-World constrains counterfactual states using the observed scene, enabling efficient and coherent reintegration without jointly optimizing all object states, thereby reducing computational cost and error accumulation. Experiments on object completeness, spatial accuracy, outdoor background completion, rendering quality, simulated dynamics, and downstream applications demonstrate the effectiveness of CA-World. Project page: https://chnxindong.github.io/ca-world/.",
      "short_abstract": "Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static and holistic visual fidelity, with limited support for multi-object...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.30239",
      "pdf_url": "https://arxiv.org/pdf/2605.30239.pdf"
    },
    {
      "id": "2605.29935",
      "anchor": "paper-2605-29935",
      "title": "CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving",
      "authors": [
        "Zezhong Qian",
        "Zhao Yang",
        "Lu Tan",
        "Zhihao Yan",
        "Weiyi Hong",
        "Haizhuang Liu",
        "Yawei Jueluo"
      ],
      "abstract": "Autonomous driving systems are commonly trained and evaluated within limited geographic regions, which hinders their scalability when deployed in new cities. However, significant domain shifts in appearance, road topology, and traffic patterns often cause severe performance degradation under cross-city deployment. Existing approaches based on domain adaptation, data augmentation, or synthetic data generation typically rely on labeled target data, city-specific annotations, or task-specific designs, limiting their scalability and effectiveness for holistic evaluation. In this paper, we introduce CityTransfer-Bench, a geographically disjoint benchmark for evaluating cross-city generalization across perception, segmentation, and planning, and propose CityGen, a diffusion-based generative framework that performs zero-label city adaptation via HD-map-conditioned synthesis guided by city-level visual prompts. Extensive experiments demonstrate that CityGen consistently improves cross-city robustness across multiple tasks, establishing a scalable and label-efficient foundation for generalizable autonomous driving.",
      "short_abstract": "Autonomous driving systems are commonly trained and evaluated within limited geographic regions, which hinders their scalability when deployed in new cities. However, significant domain shifts in appearance, road topology, and traffic patterns often cause...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.29935",
      "pdf_url": "https://arxiv.org/pdf/2605.29935.pdf"
    },
    {
      "id": "2605.29471",
      "anchor": "paper-2605-29471",
      "title": "V2VCrafter: Consistent Street-View Image Generation Across Vehicles",
      "authors": [
        "Yihang Tao",
        "Yu Guo",
        "Senkang Hu",
        "Yanan Ma",
        "Zihan Fang",
        "Sam Kwong",
        "Yuguang Fang"
      ],
      "abstract": "Connected and autonomous driving (CAD) systems leverage vehicle-to-vehicle (V2V) communication for multi-agent collaborative perception, yet remain constrained by scarce annotated real-world V2V datasets and limited generalization across diverse driving conditions. While image generation offers a feasible solution for data augmentation, existing single-vehicle multi-view generation frameworks face two key challenges in multi-agent settings: (1) the expanded learning objective degrades generation quality, and (2) dynamic inter-agent variation hinders consistency modeling for physical attributes (e.g., color, category) of jointly observed objects. To bridge this gap, we propose V2VCrafter, the first framework for generating controllable and realistic multi-view driving images across vehicles. For effective learning, we develop a progressive multi-agent diffusion model based on a single-agent backbone, using neighboring agents' latent states to progressively guide single-to-multi-agent generation. To address cross-vehicle inconsistency, we further propose a cross-agent attention module that leverages a collaboration view graph and learnable jointly observed object representations to model dynamic cross-vehicle camera view relationships. Experiments on real-world V2X-Real dataset show that V2VCrafter generates high-fidelity, controllable, and consistent street views across vehicles, thereby effectively enhancing downstream collaborative 3D object detection tasks.",
      "short_abstract": "Connected and autonomous driving (CAD) systems leverage vehicle-to-vehicle (V2V) communication for multi-agent collaborative perception, yet remain constrained by scarce annotated real-world V2V datasets and limited generalization across diverse driving...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 11,
        "margin": 2,
        "evidence": [
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.29471",
      "pdf_url": "https://arxiv.org/pdf/2605.29471.pdf"
    },
    {
      "id": "2605.30571",
      "anchor": "paper-2605-30571",
      "title": "Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode",
      "authors": [
        "Josef Chen"
      ],
      "abstract": "Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive decode, where one robot, camera feed or user session waits on the next token. This workload is usually described as memory-bandwidth-bound. Each decode step streams model weights and the active KV cache, so latency should scale with peak HBM bandwidth. We show that this account is true but incomplete. We measure batch-1 decode for three 7 to 8B-class GQA transformers across four NVIDIA GPUs: H100 SXM5, A100-80GB SXM4, L40S and L4. We evaluate context lengths from 2048 to 16384, producing 44 valid cells under a controlled bf16 SDPA setup. The achieved fraction of peak HBM bandwidth falls as peak bandwidth rises. On the headline Qwen-2.5-7B ctx=2048 cell, an L4 reaches roughly 81 percent of its analytic memory floor, while an H100 reaches only 27 percent. Physical-AI decode is memory-dominated, but faster memory does not translate into proportional latency gains. We test the missing term with a CUDA Graphs A/B experiment. On H100 at ctx=2048, CUDA Graphs improves decode latency by 1.259x across N=10 fresh sessions, with a 95 percent bootstrap confidence interval of 1.253 to 1.267. On L4, the same intervention gives only 1.028x. This isolates a launch-side overhead that becomes visible on fast GPUs but remains mostly hidden on slower, bandwidth-bound GPUs. The deployment implication is that memory savings matter only when the runtime realises them. On L4, bf16 decode sits close to the memory floor, but common quantised paths do not recover the expected 4x weight-traffic reduction: bnb-nf4 reaches 59.36 ms/step and AutoAWQ+Marlin reaches 45.24 ms/step from a 62.32 ms bf16 baseline. GPTQ+ExLlamaV2, with Ada-tuned int4 kernels, reaches 17.36 ms/step.",
      "short_abstract": "Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive decode, where one robot, camera feed or user session waits on...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "abstract: deployment or real-time",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.30571",
      "pdf_url": "https://arxiv.org/pdf/2605.30571.pdf"
    },
    {
      "id": "2605.29518",
      "anchor": "paper-2605-29518",
      "title": "Network Optimization Aspects of Autonomous Vehicles: Challenges and Future Directions",
      "authors": [
        "Rudolf Krecht",
        "Tamas Budai",
        "Erno Horvath",
        "Akos Kovacs",
        "Nobert Marko",
        "Miklos Unger"
      ],
      "abstract": "Global megatrends, such as urbanization, population growth, and emerging network solutions are accelerating the development of the Connected and Autonomous Vehicles (CAVs) industry. There are many truths, some misconceptions, and even some excitement about CAVs in the public's opinion. The main objective of the current article is to provide a comprehensive review, eliminate misconceptions, and outline the future of the network optimization aspects of autonomous vehicles by presenting various multidisciplinary methods, such as cooperative perception. Given our extensive experience with CAVs, we are aiming to share some of the insights and knowledge we have gained, along with relevant use-cases and experiment results.",
      "short_abstract": "Global megatrends, such as urbanization, population growth, and emerging network solutions are accelerating the development of the Connected and Autonomous Vehicles (CAVs) industry. There are many truths, some misconceptions, and even some excitement about...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 12,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.29518",
      "pdf_url": "https://arxiv.org/pdf/2605.29518.pdf"
    },
    {
      "id": "2605.29818",
      "anchor": "paper-2605-29818",
      "title": "Teleoperation Operational Design Domain based on Minimal Risk Maneuver Capability",
      "authors": [
        "Leon Johann Brettin",
        "Nayel Fabian Salem",
        "Ole Hans",
        "Markus Maurer"
      ],
      "abstract": "This article discusses the concept of an Operational Design Domain (ODD) designed specifically for teleoperated road vehicles. For this purpose, the ODD concept designed for automated driving is adapted for teleoperation. As teleoperation becomes more common in regular traffic, the question arises under which operating conditions such vehicles are able and allowed to drive. Currently, these conditions are selected primarily based on network performance. From a safety perspective, it is difficult to base such a selection on a reliable connection because it is almost impossible to guarantee sufficient reliability. With this in mind, the ODD concept designed for automated driving is adapted for teleoperation: A concept is proposed for basing the ODD for a teleoperation system on the capability of the teleoperated vehicle to perform a minimal risk maneuver using a dedicated system designed solely for this purpose. This concept is then demonstrated using a use case example.",
      "short_abstract": "This article discusses the concept of an Operational Design Domain (ODD) designed specifically for teleoperated road vehicles. For this purpose, the ODD concept designed for automated driving is adapted for teleoperation. As teleoperation becomes more common...",
      "published": "2026-05-28",
      "updated": "2026-05-28",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 18,
        "evidence": [
          "title: teleoperation",
          "abstract: teleoperation",
          "abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 83,
      "arxiv_url": "https://arxiv.org/abs/2605.29818",
      "pdf_url": "https://arxiv.org/pdf/2605.29818.pdf"
    },
    {
      "id": "2605.29138",
      "anchor": "paper-2605-29138",
      "title": "Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff in Autonomous Driving",
      "authors": [
        "Qitao Weng",
        "Heechul Yun"
      ],
      "abstract": "Latency-accuracy tradeoffs are fundamental in real-time applications of deep neural networks (DNNs) for cyber-physical systems. In autonomous driving, in particular, safety depends on both prediction quality and the end-to-end delay from sensing to actuation. We observe that (1) when latency is accounted for, the latency-optimal network configuration varies with scene context and compute availability; and (2) a single fixed-resolution model becomes suboptimal as conditions change. We present a multi-resolution, end-to-end deep neural network for the CARLA urban driving challenge using monocular camera input. Our approach employs a convolutional neural network (CNN) that supports multiple input resolutions through per-resolution batch normalization, enabling runtime selection of an ideal input scale under a latency budget, as well as resolution retargeting, which allows multi-resolution training without access to the original training dataset. We implement and evaluate our multi-resolution end-to-end CNN in CARLA to explore the latency-safety frontier. Results show consistent improvements in per-route safety metrics - lane invasions, red-light infractions, and collisions - relative to fixed-resolution baselines.",
      "short_abstract": "Latency-accuracy tradeoffs are fundamental in real-time applications of deep neural networks (DNNs) for cyber-physical systems. In autonomous driving, in particular, safety depends on both prediction quality and the end-to-end delay from sensing to actuation....",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 7,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.29138",
      "pdf_url": "https://arxiv.org/pdf/2605.29138.pdf"
    },
    {
      "id": "2605.29114",
      "anchor": "paper-2605-29114",
      "title": "ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving",
      "authors": [
        "Mohammadreza Teymoorianfard",
        "Jean-Philippe Monteuuis",
        "Jonathan Petit",
        "Amir Houmansadr"
      ],
      "abstract": "Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework capturing both semantic and structural aspects of reasoning, along with safety-centric measures. We also introduce a benchmark for evaluating attacks and defenses on reasoning-trajectory interactions in autonomous driving. Our results highlight the need for rigorous evaluation and improved defenses to ensure the safety of reasoning-enabled VLA systems in autonomous driving.",
      "short_abstract": "Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Benchmark",
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 23,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.29114",
      "pdf_url": "https://arxiv.org/pdf/2605.29114.pdf"
    },
    {
      "id": "2605.28587",
      "anchor": "paper-2605-28587",
      "title": "Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation",
      "authors": [
        "Yang Gao",
        "Wuyang Li",
        "Po-Chien Luan",
        "Alexandre Alahi"
      ],
      "abstract": "Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. However, existing weakly supervised occupancy prediction frameworks predominantly assume rigid-body motion and rely on simple frame-to-frame offsets, limiting their ability to capture fine-grained deformations and maintain temporal coherence. To address this issue, we propose DeGO, a deformable Gaussian occupancy framework that unifies decoupled Gaussian deformation with factorized 4D foundation-model distillation. DeGO disentangles rigid and nonrigid motion, enabling each Gaussian primitive to evolve through both deformation and offset-based updates. In parallel, a factorized 4D distillation strategy transfers cross-camera and cross-frame knowledge from the VGGT foundation model, producing foundation-aligned features that enhance temporal consistency. Experiments on the Occ3D-NuScenes benchmark demonstrate that our method achieves state-of-the-art performance under weak supervision, delivering 13.5% gains on human-centric instances and 10.9% overall improvements. These results highlight the effectiveness of deformation-aware and foundation-guided occupancy modeling for dynamic scene understanding. The code is publicly available: https://github.com/vita-epfl/DeGO",
      "short_abstract": "Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. However, existing weakly supervised occupancy prediction frameworks predominantly assume rigid-body motion and...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 11,
        "evidence": [
          "title: occupancy",
          "abstract: occupancy",
          "abstract: camera",
          "abstract: scene understanding"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.28587",
      "pdf_url": "https://arxiv.org/pdf/2605.28587.pdf"
    },
    {
      "id": "2605.28583",
      "anchor": "paper-2605-28583",
      "title": "SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving",
      "authors": [
        "Kangyu Wu",
        "Peng Cui",
        "Guoxi Chen",
        "Ya Zhang"
      ],
      "abstract": "Ensuring both safety and efficiency in decision-making for autonomous driving systems remains a fundamental challenge. Traditional Deep Reinforcement Learning (DRL) suffers from unsafe random exploration and slow convergence, while Large Language Models (LLMs) demonstrate inherent latency in real-time inference operations. To address these limitations, this paper proposes SARAD, a novel safety-aware hybrid framework that synergizes LLMs and DRL for autonomous driving. SARAD substitutes the random exploration of DRL with Retrieval-Augmented Generation (RAG)-enhanced, LLM-guided decisions sourced from a dynamic expert knowledge repository. An attention discriminator is proposed to integrate the prior knowledge of LLMs into DRL policy optimization. A collision predictor module, fine-tuned with historical collision data, is further designed to improve vehicle safety. Extensive experiments show that SARAD achieves significant performance improvements in the Highway-Env simulator, validating the effectiveness of the proposed model in autonomous driving.",
      "short_abstract": "Ensuring both safety and efficiency in decision-making for autonomous driving systems remains a fundamental challenge. Traditional Deep Reinforcement Learning (DRL) suffers from unsafe random exploration and slow convergence, while Large Language Models...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 10,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.28583",
      "pdf_url": "https://arxiv.org/pdf/2605.28583.pdf"
    },
    {
      "id": "2605.28544",
      "anchor": "paper-2605-28544",
      "title": "DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving",
      "authors": [
        "Chen Shi",
        "Jinrui Xu",
        "Shaoshuai Shi",
        "Kehua Sheng",
        "Bo Zhang",
        "Li Jiang"
      ],
      "abstract": "Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture temporal dynamics and motion priors that are naturally suited for driving. We present DriveWAM, a driving world-action model that adapts a pretrained video diffusion transformer into an autoregressive video-action policy. DriveWAM organizes video and action streams into a unified temporal token sequence and trains them under a joint flow-matching objective, preserving the pretrained video-generation architecture while adapting its large-scale video priors to action generation. To incorporate high-level scene understanding, we introduce scene-evolving driving guidance, where a frozen VLM produces chunk-specific semantic intent to guide video-action generation. To keep long-horizon rollout bounded, we further introduce selective KV memory, which maintains bounded modality-aware video and action memory pools through relevance-redundancy cache selection at inference time. Experiments on NAVSIM and the PhysicalAI-Autonomous-Vehicles benchmark show that DriveWAM achieves strong planning performance, and a data-scaling study from 4k to 100k driving clips further confirms the scaling potential of world-action modeling for end-to-end autonomous driving.",
      "short_abstract": "Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture temporal dynamics and motion priors that are...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 5,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.28544",
      "pdf_url": "https://arxiv.org/pdf/2605.28544.pdf"
    },
    {
      "id": "2605.28180",
      "anchor": "paper-2605-28180",
      "title": "Tensor Train Decomposition Based Noise Reduction and Enhanced Parameter Estimation for FMCW MIMO Radar Systems",
      "authors": [
        "Luoyan Zhu",
        "Sergiy A. Vorobyov",
        "Jie Wang",
        "Yinsheng Liu",
        "Zhangdui Zhong"
      ],
      "abstract": "Frequency modulated continuous wave (FMCW) radar is widely used in autonomous driving and industrial inspection due to its high-resolution target location and velocity estimation capability. However, the plethora of connected devices in automotive applications introduces electromagnetic interference and brings challenges to location-aware services, primarily due to the issue of low signal-to-noise ratio (SNR) caused by mixed noise contamination. Conventional matrix-based signal processing methods exhibit performance deterioration when handling higher-order signals under low SNR conditions. To address this challenge, this paper proposes a tensor decomposition-based framework that jointly performs noise reduction and parameter estimation for four-dimensional signals in FMCW multiple-input multiple-output (MIMO) radar systems. Specifically, the framework exploits the inherent low-rank structure and multidimensional correlations of the received signals through tensor train decomposition to effectively separate noise subspace. A data smoothing processor then reconstructs an augmented signal tensor to resolve rank deficiency caused by coherent signals. Finally, an enhanced rotational subspace algorithm is employed to jointly decouple the distance, velocity, and angle parameters by exploiting the structural fitting to the restored signal. Both simulation and field experiments under real-world noise demonstrate that our proposed framework achieves significant noise reduction while improving target SNR and parameter estimation accuracy. These advancements make the proposed framework a robust solution for high-precision MIMO FMCW radar applications in dynamic, noise-polluted environments.",
      "short_abstract": "Frequency modulated continuous wave (FMCW) radar is widely used in autonomous driving and industrial inspection due to its high-resolution target location and velocity estimation capability. However, the plethora of connected devices in automotive...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 10,
        "evidence": [
          "title: radar",
          "abstract: radar"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.28180",
      "pdf_url": "https://arxiv.org/pdf/2605.28180.pdf"
    },
    {
      "id": "2605.28136",
      "anchor": "paper-2605-28136",
      "title": "SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving",
      "authors": [
        "Toomas Tahves",
        "Mauro Bellone",
        "Junyi Gu",
        "Raivo Sell"
      ],
      "abstract": "Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (ZOD) provides rich multi-sensor data but only bounding-box labels, limiting its use for segmentation research. Our primary contribution is a Segment Anything Model (SAM)-based annotation pipeline that produces dense, pixel-level annotations for ZOD by converting bounding boxes into semantic masks. In this pilot study, we process over 100,000 frames and manually curate a 2,300-frame subset (36% acceptance rate) to establish a reliable baseline. Using these annotations, we evaluate transformer-based CLFT and CNN-based DeepLabV3+ architectures across diverse weather conditions, achieving up to 48.1% mIoU with CLFT-Hybrid. To address extreme class imbalance, where pedestrians, cyclists, and signs constitute less than 1% of pixels, we explore specialized models targeting rare classes. We further validate the pipeline on the Iseauto autonomous-vehicle platform, achieving 77.5% mIoU, and show that SAM-derived representations transfer effectively across sensor configurations via bidirectional transfer learning. All code and annotations are released to support reproducible research.",
      "short_abstract": "Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (ZOD) provides rich multi-sensor data but only bounding-box labels, limiting its use for segmentation...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Perception & Sensor Fusion: abstract: scene segmentation"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.28136",
      "pdf_url": "https://arxiv.org/pdf/2605.28136.pdf"
    },
    {
      "id": "2605.27964",
      "anchor": "paper-2605-27964",
      "title": "DRIFT: Driving Risk Inference via Field Transmission for Human-like Autonomous Driving",
      "authors": [
        "Zian Wang",
        "Yiming Shu",
        "Zejian Deng",
        "Chen Sun"
      ],
      "abstract": "Risk fields offer spatially structured alternatives to scalar safety metrics. However, hand-crafted static risk field models struggle with occlusion and topology-driven propagation. We present DRIFT, a spatiotemporal risk field governed by an advection-diffusion-reaction partial differential equation (PDE), with an optional telegrapher term. DRIFT draws on three sources: anisotropic Gaussian kernels to capture velocity-induced risk, occlusion-aware latent hazards behind large vehicles, and topology-coupled merge-zone conflict pressure. We further introduce field-centric evaluation metrics to complement the existing Surrogate Safety Measures (SSMs), including Lane-Change Risk Differential, Temporal Anticipation Index, Occlusion Sensitivity Index, and Occlusion Response Latency. Experiments on real-world traffic datasets show that DRIFT reduces occlusion response latency and lowers the near-collision rate under occlusion compared with selected baselines in synthetic scenarios.",
      "short_abstract": "Risk fields offer spatially structured alternatives to scalar safety metrics. However, hand-crafted static risk field models struggle with occlusion and topology-driven propagation. We present DRIFT, a spatiotemporal risk field governed by an...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.27964",
      "pdf_url": "https://arxiv.org/pdf/2605.27964.pdf"
    },
    {
      "id": "2606.00119",
      "anchor": "paper-2606-00119",
      "title": "V2I Work Zone Geometry Reconstruction with Pose-Conditioned UWB Range Denoising",
      "authors": [
        "Jiaxi Liu",
        "Hangyu Li",
        "Yang Cheng",
        "Rui Gana",
        "Junwei You",
        "Weizhe Tang",
        "Peng Zhang",
        "Steven T. Parker",
        "Xiaopeng Li",
        "Bin Ran"
      ],
      "abstract": "Reliable work zone mapping is important for connected and autonomous vehicles (CAVs) to navigate safely and smoothly through work zone areas. Cone-mounted ultra-wideband (UWB) roadside units (RSU) offer a cost-effective way for work zone layout inference, as roadside anchors and vehicle tags provide direct vehicle-to-infrastructure (V2I) range constraints for work zone geometry reconstruction. However, UWB range estimation is degraded by bursty outliers, non-line-of-sight (NLOS) errors, arbitrary anchor-ordering issues, and vehicle pose uncertainties in practical field deployments. To address these challenges, this study proposes a pose-conditioned, permutation-equivariant predictive denoiser for multi-anchor UWB ranging. The model employs shared anchor-wise temporal prediction to capture range dynamics, symmetric set aggregation to handle unordered and missing anchors, and pose-conditioned residual decoding to incorporate vehicle motion as a geometric prior. A two-stage training strategy first learns prediction from observed ranges, and then fine-tunes the denoiser with NLOS-weighted supervision. The method is evaluated on rare real-world V2I UWB field data collected with a CAV, as well as on controlled large-scale simulation benchmarks for ablative insights. Results show that the proposed method substantially improves range accuracy, cone localization, and work zone geometry reconstruction in challenging NLOS-dominated regimes, remains robust to anchor re-indexing and moderate anchor dropout, and reduces measurement-weighted field MSE by 66.9% relative to the raw input.",
      "short_abstract": "Reliable work zone mapping is important for connected and autonomous vehicles (CAVs) to navigate safely and smoothly through work zone areas. Cone-mounted ultra-wideband (UWB) roadside units (RSU) offer a cost-effective way for work zone layout inference, as...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 4,
        "evidence": [
          "abstract: V2X",
          "abstract: connected vehicles",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2606.00119",
      "pdf_url": "https://arxiv.org/pdf/2606.00119.pdf"
    },
    {
      "id": "2605.29231",
      "anchor": "paper-2605-29231",
      "title": "$α$-stability of Differentially Flat Systems with Application to Newton-Raphson Tracking Control for Vehicle Dynamics",
      "authors": [
        "Aadila Ali Sabry",
        "Gennaro Notomista"
      ],
      "abstract": "This paper studies the $α$-stability property of differentially flat nonlinear dynamical systems. The results build off the recently introduced notion of $α$-stability, which is particularly amenable to characterize the ability of a system to track dynamic output reference signals. We consider systems controlled using the Newton-Raphson tracking controller, which results in closed-form control policies, therefore it is computationally efficient, and it has been shown to be effective to control a large variety of mobile robots, including autonomous vehicles. The main results of the paper consist in sufficient conditions for the $α$-stability of differentially flat systems and for the equivalence between the proposed control algorithm and the Newton-Raphson tracking controller applied directly to the nonlinear dynamics. We demonstrate the behavior of the proposed controller applied to the kinematic unicycle and dynamic bicycle models.",
      "short_abstract": "This paper studies the $α$-stability property of differentially flat nonlinear dynamical systems. The results build off the recently introduced notion of $α$-stability, which is particularly amenable to characterize the ability of a system to track dynamic...",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 26,
        "evidence": [
          "title: vehicle dynamics",
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.29231",
      "pdf_url": "https://arxiv.org/pdf/2605.29231.pdf"
    },
    {
      "id": "2605.28552",
      "anchor": "paper-2605-28552",
      "title": "Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning",
      "authors": [
        "Qingwen Pu",
        "Kun Xie",
        "Hong Yang",
        "Di Yang",
        "Junqing Wang"
      ],
      "abstract": "As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle types in safety-critical interactions is essential for the safe deployment of automated driving technologies. This study extracts safety-critical pedestrian-vehicle interactions from the Argoverse 2 dataset to capture real-world crash avoidance behaviors in encounters involving AVs and HDVs. To model vehicle-type-specific pedestrian crash avoidance behavior, we develop a Smooth-Mamba Deep Deterministic Policy Gradient framework, termed SMamba-DDPG, which integrates smooth action constraints with efficient temporal representation learning. To quantify pedestrian behavioral differences, the framework trains separate crash avoidance policies for pedestrian interactions with AVs and HDVs. Results show that SMamba-DDPG outperforms baseline reinforcement learning and supervised learning models in reproducing pedestrian crash avoidance behaviors. Reconstructed trajectories demonstrate strong behavioral realism, accurately reproducing crash avoidance kinematics in both AV and HDV scenarios. Reaction time analysis shows that the model captures human-like response delays and reveals that pedestrians respond more quickly to AVs than to HDVs. Counterfactual analysis further indicates that pedestrians adopt lower crossing speeds when interacting with AVs. Large-scale safety analysis of model-generated data revealed that pedestrian-AV interactions consistently yielded lower conflict rates and higher pedestrian yielding rates compared to pedestrian-HDV interactions. The findings highlight the importance of incorporating vehicle-type-specific pedestrian behavioral models for safer automated driving system design and more realistic traffic simulations in mixed-traffic environments.",
      "short_abstract": "As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle types in safety-critical interactions is essential for the safe deployment of automated driving technologies....",
      "published": "2026-05-27",
      "updated": "2026-05-27",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 84,
      "arxiv_url": "https://arxiv.org/abs/2605.28552",
      "pdf_url": "https://arxiv.org/pdf/2605.28552.pdf"
    },
    {
      "id": "2605.27038",
      "anchor": "paper-2605-27038",
      "title": "TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving",
      "authors": [
        "Jiaxiang Li",
        "Yumao Liu",
        "Ke Ma"
      ],
      "abstract": "Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial forecasting remains a critical challenge. Existing representation strategies generally follow two paths: text-aligned methods flatten continuous spatial states into symbols, which compromises geometric structure and induces \"spatial hallucinations\"; dense visual methods preserve spatial topology but overwhelm standard tokenizers with redundant background textures, leading to \"representation interference\". To address these limitations, we introduce TPS-Drive, a novel framework centered on Task-Guided Representation Purification that empowers VLMs to Think in Purified Space. At its core, an Agent-Centric Tokenizer utilizes a task-guided vector quantization mechanism supervised by a frozen 3D detection head, which explicitly reallocates limited codebook capacity from pervasive static backgrounds to critical dynamic agents and effectively isolates spatial redundancy. Leveraging this purified spatial vocabulary, TPS-Drive employs a decoupled reasoning pipeline that sequentially performs scene understanding, future forecasting, and action generation. The framework is optimized via a progressive three-stage training paradigm, culminating in reward-driven refinement that surpasses pure imitation learning. Extensive experiments validate our approach: TPS-Drive achieves accurate agent spatial state forecasting and reduces collision rates in open-loop nuScenes evaluations, while establishing new safety records on the rigorous closed-loop NAVSIMv1 and NAVSIMv2 benchmarks.",
      "short_abstract": "Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial forecasting remains a critical challenge. Existing representation strategies generally follow two paths:...",
      "published": "2026-05-26",
      "updated": "2026-05-26",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Perception & Sensor Fusion: abstract: scene understanding"
        ]
      },
      "recency": "fresh",
      "age_days": 85,
      "arxiv_url": "https://arxiv.org/abs/2605.27038",
      "pdf_url": "https://arxiv.org/pdf/2605.27038.pdf"
    },
    {
      "id": "2605.26577",
      "anchor": "paper-2605-26577",
      "title": "Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial",
      "authors": [
        "Haoyu Li",
        "Xiangru Zhong",
        "Hao Cheng",
        "Bin Hu",
        "Huan Zhang"
      ],
      "abstract": "Learning-based methods for synthesizing controllers have gained popularity due to their high expressiveness and strong empirical performance. However, in safety-critical scenarios such as autonomous driving, robotics, and power systems, empirical performance alone is insufficient, and formal verification of controller properties such as stability and safety is highly desirable. Unfortunately, many prior verification approaches are either tied to specific structural assumptions on the system or the certificate, making them difficult to transfer across settings, or suffer from poor scalability on higher-dimensional neural network systems. In this tutorial, we present a unified framework that aims to mitigate this gap via bridging control with the state-of-the-art neural network verifier $α,\\!β$-CROWN (alpha-beta-CROWN). At its core, $α,\\!β$-CROWN is a general-purpose bounding engine for nonlinear functions represented as computation graphs: given an input domain, it can produce certified bounds and explicit linear relaxation of the nonlinear function. These certified bounds are useful on their own for tasks such as reachability analysis, and they also provide the foundation for more complex routines that perform satisfiability checking and optimization. More specifically, many control problems reduce to verifying real-valued inequalities over a state domain (e.g., Lyapunov theory). Consequently, $α,\\!β$-CROWN enables scalable verification of such conditions by computing tight bounds and recursively partitioning and pruning subdomains based on the bounds. Thanks to GPU parallelization, this pipeline demonstrates superior scalability on verification and optimization problems that are challenging for traditional approaches. In this tutorial, we discuss the basics of $α,\\!β$-CROWN and introduce its application to various control-related tasks.",
      "short_abstract": "Learning-based methods for synthesizing controllers have gained popularity due to their high expressiveness and strong empirical performance. However, in safety-critical scenarios such as autonomous driving, robotics, and power systems, empirical performance...",
      "published": "2026-05-26",
      "updated": "2026-05-26",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Survey"
      ],
      "classification": {
        "confidence": "low",
        "score": 11,
        "margin": 2,
        "evidence": [
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 85,
      "arxiv_url": "https://arxiv.org/abs/2605.26577",
      "pdf_url": "https://arxiv.org/pdf/2605.26577.pdf"
    },
    {
      "id": "2605.27643",
      "anchor": "paper-2605-27643",
      "title": "Agentic Language-to-Objective Synthesis for Optofluidic Assembly",
      "authors": [
        "Ivan Saraev",
        "Elena Erben",
        "Weida Liao",
        "Fan Nan",
        "Gerhard Neumann",
        "Eric Lauga",
        "Moritz Kreysing"
      ],
      "abstract": "Light-based advanced manufacturing increasingly requires programmable, closed-loop tools that translate human design intent into executable operations at small length scales. Yet a key bottleneck persists across robotic and manufacturing modalities: turning user intent into machine-readable objectives that are reliably executable. While micro-robotics offers versatile manipulation via optical actuation of fluids, mathematically tractable goal specification remains manual and hard to reuse. Here, we introduce Speak-to-Objective, a modular agentic pipeline that uses a conditioned Large Language Model (LLM) to translate spoken or written commands into fully differentiable objective functions for assembling microparticles in a constraint-aware inverse solver (SLSQP) and on an experimental optofluidic platform. The approach employs a compact loop - perceive -> compose -> propose -> act -> report & learn - that treats the objective as the interface between intent and actuation, separating what to assemble or pattern from how to actuate, while learning from user feedback. The pipeline composes geometry, spacing, and assignment/topology terms to generate robust descriptive objectives that assemble from partial traces and recover after perturbations, as well as explicit objectives for precise placement, all in an actuator-agnostic fashion. Using laser-induced thermoviscous flows as the physical actuation modality, we demonstrate natural-language-programmable, light-based microscale assembly of particle patterns in a microfluidic environment. Beyond its immediate impact on programmable microassembly, and using laser-induced optofluidic actuation as a reduced-complexity experimental platform, our work points toward self-driving, AI-assisted optical manufacturing platforms in which natural language, differentiable objectives, and laser-based actuation are coupled into a reusable digital workflow.",
      "short_abstract": "Light-based advanced manufacturing increasingly requires programmable, closed-loop tools that translate human design intent into executable operations at small length scales. Yet a key bottleneck persists across robotic and manufacturing modalities: turning...",
      "published": "2026-05-26",
      "updated": "2026-05-26",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 85,
      "arxiv_url": "https://arxiv.org/abs/2605.27643",
      "pdf_url": "https://arxiv.org/pdf/2605.27643.pdf"
    },
    {
      "id": "2605.27659",
      "anchor": "paper-2605-27659",
      "title": "Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment",
      "authors": [
        "Gengyue Han",
        "Yiheng Feng"
      ],
      "abstract": "Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicles) are first trained in simulators. However, when deployed in real world environments, they often suffer from performance degradation or safety violations because of the inevitable Sim2Real gap. Existing zero-shot approaches, such as robust safe RL and domain randomization, mitigate this issue but typically at the cost of degraded performance or residual safety risks when experiencing unmodeled system dynamics. To address these limitations, we propose a novel reinforcement learning framework that enables safe and efficient policy transfer via probabilistic latent embeddings and dynamic policy adaptation. We consider a family of Constrained Markov Decision Processes (CMDPs) under different environment contexts. By leveraging latent context variable in meta-RL, the proposed framework infers the latent representation of the environment from simulated experiences. Furthermore, it incorporates a distributional RL formulation, which allows risk levels of the deployed policy to be adjusted dynamically, based on the estimation accuracy of the latent context variable. This strategy promotes safety at the early deployment stage and improves efficiency through fast policy adaptation under the Sim2Real gap.",
      "short_abstract": "Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicles) are first trained in simulators. However, when deployed in real world environments, they often suffer from...",
      "published": "2026-05-26",
      "updated": "2026-05-26",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: law or policy",
          "abstract: law or policy"
        ]
      },
      "recency": "fresh",
      "age_days": 85,
      "arxiv_url": "https://arxiv.org/abs/2605.27659",
      "pdf_url": "https://arxiv.org/pdf/2605.27659.pdf"
    },
    {
      "id": "2605.26627",
      "anchor": "paper-2605-26627",
      "title": "Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty",
      "authors": [
        "Chayan Banerjee",
        "Ethan Goan"
      ],
      "abstract": "Deploying reinforcement learning in safety critical domains, from autonomous vehicles to medical decision support, is constrained by failures arising when systems encounter unfamiliar conditions. We argue that the fundamental bottleneck is not individual challenges like changing dynamics or incomplete observations, but their synergistic interaction, which we term the Epistemic Trap: agents cannot estimate their state without knowing system dynamics, nor learn dynamics without accurate state information. Proof-of-concept experiments in simulated locomotion reveal that combining these uncertainties causes failures far worse than either challenge alone, a 77% observed degradation against the 46% additive prediction, demonstrating that compounding failure modes can emerge and, when they do, far exceed what additive reasoning would predict. Conventional approaches typically adopt a passive epistemic stance that cannot resolve this coupled uncertainty. We propose reframing safety as an information problem. We introduce an Adaptive Safety Architecture built around three contributions. First, the Compound Uncertainty Coefficient ($κ$), a mutual-information based metric that quantifies how tightly state and dynamics uncertainties are coupled. Second, information-seeking policies governed by a MaxInfoRL objective that actively probe system dynamics rather than waiting for the environment to reveal itself passively. Third, regime adaptive safety constraints that tighten automatically as epistemic coupling rises. Together, these constitute a paradigm shift from passive robustness to active perception, offering a principled path toward decision making systems that operate under uncertainty, recognize their own ignorance, and act strategically to resolve it.",
      "short_abstract": "Deploying reinforcement learning in safety critical domains, from autonomous vehicles to medical decision support, is constrained by failures arising when systems encounter unfamiliar conditions. We argue that the fundamental bottleneck is not individual...",
      "published": "2026-05-26",
      "updated": "2026-05-26",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: robustness",
          "title: uncertainty",
          "abstract: uncertainty",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 85,
      "arxiv_url": "https://arxiv.org/abs/2605.26627",
      "pdf_url": "https://arxiv.org/pdf/2605.26627.pdf"
    },
    {
      "id": "2605.26501",
      "anchor": "paper-2605-26501",
      "title": "Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization",
      "authors": [
        "Xiang Fang",
        "Wanlong Fang",
        "Changshuo Wang"
      ],
      "abstract": "Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks, particularly those exploiting both modalities, remains underexplored, posing risks to critical applications like autonomous driving and content moderation. Existing attacks focus on single modalities or require impractical white-box access, limiting their real-world relevance. In this paper, we introduce Multi-Modal Adversarial Synergy, a groundbreaking framework that crafts universal, black-box multi-modal attacks against LVLMs. MMAS simultaneously generates a texture scale-constrained universal adversarial perturbation for images and a learnable prompt perturbation for text, optimized jointly using only model queries. The image perturbation leverages wavelet-based texture constraints to ensure imperceptibility and robustness across diverse visual inputs. The text perturbation, constrained by an L-norm in the embedding space, maintains semantic coherence while steering outputs toward a target. A novel cross-modal regularization term aligns the perturbations' gradient directions, enhancing their synergistic impact and transferability across tasks and models. Extensive experiments show the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs.",
      "short_abstract": "Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks,...",
      "published": "2026-05-25",
      "updated": "2026-05-25",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 16,
        "evidence": [
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 86,
      "arxiv_url": "https://arxiv.org/abs/2605.26501",
      "pdf_url": "https://arxiv.org/pdf/2605.26501.pdf"
    },
    {
      "id": "2605.26477",
      "anchor": "paper-2605-26477",
      "title": "Variational Inference for Evidential Deep Learning",
      "authors": [
        "Jiawei Tang",
        "Xinyan Du",
        "Hui Liu",
        "Junhui Hou",
        "Yuheng Jia"
      ],
      "abstract": "While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mitigates this by formulating predictions as a Dirichlet distribution over class probabilities to explicitly quantify epistemic uncertainty. However, we found that the conventional EDL suffers from two fundamental limitations: a Kullback-Leibler (KL) penalty that only suppresses the evidence of negative classes, producing excessively high evidence therefore decreasing the model's ability to quantify uncertainty, and an absence in theoretical guarantee of setting Dirichlet parameter $α=e+1$. In this paper, we propose a mathematically principled framework, Variational Inference Evidential Deep Learning (VI-EDL). By reformulating evidential learning through the lens of variational inference, we derive an Evidence Lower Bound (ELBO), which prevents the evidence from growing excessively. Theoretically, we rigorously establish a generalization bound and reveal how the predicted uncertainty, feature and network complexity affect this bound, and why setting $\\boldsymbolα = \\mathbf{e} + \\mathbf{1}$ can minimize it. Extensive experiments on standard visual and medical datasets demonstrate that VI-EDL achieves state-of-the-art performance, showing excellent performance in out-of-distribution detection, noise detection and autonomous driving scenario. The code is available in https://github.com/seutjw/VI-EDL.",
      "short_abstract": "While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mitigates this by formulating predictions as a Dirichlet distribution over class probabilities to explicitly...",
      "published": "2026-05-25",
      "updated": "2026-05-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Safety, Security & Verification: abstract: uncertainty",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 86,
      "arxiv_url": "https://arxiv.org/abs/2605.26477",
      "pdf_url": "https://arxiv.org/pdf/2605.26477.pdf"
    },
    {
      "id": "2605.26113",
      "anchor": "paper-2605-26113",
      "title": "AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond",
      "authors": [
        "Haiming Zhang",
        "Junfei Zhou",
        "Feng Jiang",
        "Jingzhong Li",
        "Zhenglong Guo",
        "Penglin Dai",
        "Jifeng Dai",
        "Yan Xie",
        "Benjin Zhu"
      ],
      "abstract": "Generating high-fidelity and controllable synthetic data is critical for advancing end-to-end autonomous driving, particularly for addressing the long tail of rare safety-critical scenarios. Existing occupancy-guided methods typically rely on shallow conditioning mechanisms and reference-frame-dependent video synthesis, which limits fine-grained controllability from arbitrary BEV layouts and restricts their applicability for scalable simulation. In this paper, we propose AnyScene, a unified occupancy-centric framework for driving scene generation. AnyScene generates semantic occupancy sequences from BEV layouts through a Spatial-Temporal Occupancy Diffusion Transformer that jointly tokenizes BEV and occupancy features in an autoregressive manner. This design enables precise controllability from cross-dataset and user-defined BEV inputs while naturally supporting long-horizon generation. Building upon the generated occupancy, a Geometry-Grounded View Expansion module treats occupancy as the canonical spatial representation and synthesizes temporally consistent multi-view driving videos in a reference-free and autoregressive fashion, supporting flexible camera configurations at inference time. Extensive experiments demonstrate that AnyScene achieves state-of-the-art performance in both occupancy and video generation. It exhibits strong generalization to unseen and customized layouts, and provides measurable benefits for downstream tasks such as sparse-view 3D reconstruction.",
      "short_abstract": "Generating high-fidelity and controllable synthetic data is critical for advancing end-to-end autonomous driving, particularly for addressing the long tail of rare safety-critical scenarios. Existing occupancy-guided methods typically rely on shallow...",
      "published": "2026-05-25",
      "updated": "2026-05-25",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 3,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 86,
      "arxiv_url": "https://arxiv.org/abs/2605.26113",
      "pdf_url": "https://arxiv.org/pdf/2605.26113.pdf"
    },
    {
      "id": "2605.25947",
      "anchor": "paper-2605-25947",
      "title": "A Pedestrian-Vehicle Interaction Benchmark and Annotation Framework for Unstructured Scenes via Uncalibrated Cameras",
      "authors": [
        "Haoyang Peng",
        "Qian Hu",
        "Songan Zhang",
        "Ming Yang"
      ],
      "abstract": "Predicting the interaction between pedestrian and vehicle is essential for autonomous driving safety in unstructured and semi-structured scenarios; however, this task is severely hindered by the scarcity of public datasets that feature dense pedestrian-vehicle interactions. Most current studies rely on structured road data, leaving the complex, heterogeneous interactions found in unstructured environments insufficiently represented and researched. In this paper, we propose a dataset annotation framework based on video data from uncalibrated surveillance cameras and present PINNS (Pedestrian-vehicle Interaction dataset from uNcalibrated cameras in uNstructured Scenes). The dataset covers multiple countries and regions, includes diverse typical traffic scenarios, and considers variations in seasons, lighting conditions, and weather. It focuses on complex scenes with dense pedestrian-vehicle interactions and is designed to be easily extensible. The dataset is constructed and annotated according to the standard issued by the Chinese Association of Automation, providing both trajectory data and corresponding scene-level information. Furthermore, this paper analyzes current challenges and research directions in heterogeneous agent trajectory prediction, shows the necessity and usefulness of the proposed dataset. We hope our framework and dataset will facilitate research on trajectory prediction and autonomous driving in complex mixed traffic scenarios. PINNS is publicly available at https://github.com/Songan-Lab.",
      "short_abstract": "Predicting the interaction between pedestrian and vehicle is essential for autonomous driving safety in unstructured and semi-structured scenarios; however, this task is severely hindered by the scarcity of public datasets that feature dense...",
      "published": "2026-05-25",
      "updated": "2026-05-25",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 1,
        "evidence": [
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 86,
      "arxiv_url": "https://arxiv.org/abs/2605.25947",
      "pdf_url": "https://arxiv.org/pdf/2605.25947.pdf"
    },
    {
      "id": "2605.25617",
      "anchor": "paper-2605-25617",
      "title": "Justice-informed Planning of Intermodal Autonomous Mobility-on-Demand Systems under Operational Constraints",
      "authors": [
        "Giacomo Ganassoli",
        "Francesco Mazzeo",
        "Cecilia Pasquale",
        "Silvia Siri",
        "Mauro Salazar"
      ],
      "abstract": "To date, most of the research on transport planning has focused on optimizing revenues or utilitarian metrics such as average travel times, which often ends up penalizing the worst-off for the sake of profit or efficiency. At the same time, most of the research in transport justice has focused on assessing injustices, without being able to prescribe operational solutions. This paper contributes to bridging this gap and presents optimization models for justice-informed operational planning of intermodal mobility systems that explicitly account for the budget and safety limitations of users, and for infrastructural capacity constraints. Specifically, we first focus on an intermodal Autonomous Mobility-on-Demand (AMoD) system -- where self-driving robotaxis provide on-demand mobility jointly with public transit and active modes -- and characterize its operations from a mesoscopic planning perspective via network flow models. Second, we leverage these models to optimize system operations through both utilitarian efficiency and justice-informed objectives. We showcase our framework in a real-world case-study for Manhattan, New York. Our results show that monetary budgets significantly limit the social justice potential of AMoD systems if they are to be deployed as transportation network companies. At the same time, granting free public transit can result in sufficiency levels very close to a completely free intermodal AMoD system, where justice-informed operations can be achieved without compromising standard efficiency metrics, ultimately highlighting the strong potential of social policies.",
      "short_abstract": "To date, most of the research on transport planning has focused on optimizing revenues or utilitarian metrics such as average travel times, which often ends up penalizing the worst-off for the sake of profit or efficiency. At the same time, most of the...",
      "published": "2026-05-25",
      "updated": "2026-05-25",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 86,
      "arxiv_url": "https://arxiv.org/abs/2605.25617",
      "pdf_url": "https://arxiv.org/pdf/2605.25617.pdf"
    },
    {
      "id": "2605.26061",
      "anchor": "paper-2605-26061",
      "title": "Neuronal Stochastic Attention Circuit (NSAC) for Probabilistic Representation Learning",
      "authors": [
        "Waleed Razzaq",
        "Yun-Bo Zhao"
      ],
      "abstract": "Reliable uncertainty quantification in continuous-time (CT) representation learning remains nascent, particularly within CT attention literature. We introduce the Neuronal Stochastic Attention Circuit (NSAC), a novel biologically-inspired CT attention architecture that reformulates attention logit computation as the solution of an Ornstein-Uhlenbeck stochastic differential equation modulated by input-dependent, nonlinear interlinked gates derived from repurposed C. elegans Neuronal Circuit Policies (NCPs) wiring mechanism. It induces a Gaussian distribution over logits that propagates principled stochasticity through a logistic-normal distribution over attention weights to yield probabilistic output. A two-term objective function combining Gaussian negative log-likelihood with an epistemic-separation regularizer enforces higher predictive variance under distributional shifts and enables joint quantification of aleatoric and epistemic uncertainty. Theoretically, we provide: (i) state stability bounds; (ii) closed-form guarantees; and (iii) frozen-coefficient error approximation. Empirically, we implement NSAC in a diverse set of learning tasks including: (i) irregular CT function approximation; (ii) multivariate regression; (iii) long-range forecasting; (iv) Industry 4.0; and (v) lane-keeping of autonomous vehicles. We observe that NSAC remains competitive against several baselines in terms of accuracy and produces informative uncertainty estimates while being interpretable at the neuronal cell level.",
      "short_abstract": "Reliable uncertainty quantification in continuous-time (CT) representation learning remains nascent, particularly within CT attention literature. We introduce the Neuronal Stochastic Attention Circuit (NSAC), a novel biologically-inspired CT attention...",
      "published": "2026-05-25",
      "updated": "2026-05-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Prediction & World Models: abstract: forecasting",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 86,
      "arxiv_url": "https://arxiv.org/abs/2605.26061",
      "pdf_url": "https://arxiv.org/pdf/2605.26061.pdf"
    },
    {
      "id": "2605.25431",
      "anchor": "paper-2605-25431",
      "title": "Mode 0: Architecture, Risk Taxonomy, and Standardization Pathway for RCU-Assisted V2X Safety Communication",
      "authors": [
        "Dewei Jiang",
        "Xiang Gu"
      ],
      "abstract": "The 3GPP V2X resource allocation framework offers two entity classes for coordinating vehicle communication --- base-station scheduling and autonomous vehicle selection --- a design space we show to be structurally incomplete for safety-critical coordination. We propose Mode 0, a third entity class centered on the Roadside Computing Unit (RCU): infrastructure integrating elevated sensing, sidelink communication, and local computation, owned by traffic management authorities rather than network operators. We formalize Mode 0's subfamily taxonomy (Mode 0a through 0c), its traffic risk taxonomy and mandatory-escalation boundary, its cross-domain coexistence with Modes 1-4, and its standardization pathway. Of the three failure modes motivating this proposal --- scheduling saturation, occlusion-driven information gaps, and an institutional authority gap during severe hazards --- only scheduling is evaluated by simulation; information and authority are argued architecturally and evidenced by independent deployment convergence across Chinese national standards, Chinese operator infrastructure, and European and US C-V2X programs, properties no channel-level simulation can meaningfully instantiate. A fifteen-run multi-agent reinforcement learning simulation shows shared-actor policies empirically converge to the analytical coordination floor for independent uniform selection. Pairing per-vehicle actors with demand separation achieves strict Pareto improvement for both traffic classes when the safety pool is correctly sized, and --- in the undersized regime where collision-free assignment is combinatorially impossible --- still lifts worst-TTI delivery reliability (5th-percentile intra-episode M0 packet delivery ratio) from 0.113 to 0.601, the only tested configuration whose delivery protection holds structurally rather than on average. We call for a 3GPP study item to formalize Mode 0.",
      "short_abstract": "The 3GPP V2X resource allocation framework offers two entity classes for coordinating vehicle communication --- base-station scheduling and autonomous vehicle selection --- a design space we show to be structurally incomplete for safety-critical coordination....",
      "published": "2026-05-25",
      "updated": "2026-05-25",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 40,
        "margin": 22,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "fresh",
      "age_days": 86,
      "arxiv_url": "https://arxiv.org/abs/2605.25431",
      "pdf_url": "https://arxiv.org/pdf/2605.25431.pdf"
    },
    {
      "id": "2605.26155",
      "anchor": "paper-2605-26155",
      "title": "When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability",
      "authors": [
        "Mehmet Haklidir"
      ],
      "abstract": "Guided Soft Actor-Critic (GSAC) distills knowledge from a privileged full-state teacher to a partial-observation student for autonomous driving, but uses a fixed distillation coefficient lambda regardless of the agent's uncertainty. We present Belief-Aware GSAC (BA-GSAC), which modulates lambda via ensemble disagreement, and use it as a testbed for a systematic empirical study asking: when does adaptive guidance actually help? Evaluating five strategies (fixed lambda in {0.01, 0.1}, adaptive, linear decay, and vanilla SAC) across three POMDP difficulty levels on Highway-Env, we find that preliminary single-seed runs suggest benefits under mild and moderate partial observability, but under severe occlusion (evaluated with 3 seeds for all methods) the adaptive coefficient collapses to lambda_min within about 3K steps. We trace this to an observability blindness phenomenon: because the ensemble predicts partial observations, it achieves low disagreement even under heavy occlusion, modeling what is visible but unable to detect what is missing. We diagnose the root cause and propose an architectural fix (training the ensemble on full-state predictions using the guiding actor's privileged access); while not validated here, we show that even with current limitations, the warmup phase provides measurable stabilization (CV=13.3% vs. 29.8% for constant lambda=0.01). In fact, a simple deterministic linear decay schedule achieves the best severe-POMDP performance across all metrics (mean 116.5, CV=8.9%), suggesting that the scheduling effect, not the ensemble, drives the stability benefit. These findings provide practical guidance for designing uncertainty-aware teacher-student frameworks and highlight ensemble prediction targets as an important design choice.",
      "short_abstract": "Guided Soft Actor-Critic (GSAC) distills knowledge from a privileged full-state teacher to a partial-observation student for autonomous driving, but uses a fixed distillation coefficient lambda regardless of the agent's uncertainty. We present Belief-Aware...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Prediction & World Models: abstract: prediction",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.26155",
      "pdf_url": "https://arxiv.org/pdf/2605.26155.pdf"
    },
    {
      "id": "2605.25393",
      "anchor": "paper-2605-25393",
      "title": "Decision-Making with Lightweight Confidence-Aware Language Model for Autonomous Driving",
      "authors": [
        "Ruoyu Yao",
        "Ruiguo Zhong",
        "Pei Liu",
        "Mingxing Peng",
        "Rui Yang",
        "Jun Ma"
      ],
      "abstract": "Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning and open-world generalization. However, the excessive computational overhead and high inference latency of these massive models severely hinder their deployment in resource-constrained AD systems. To address this challenge, we propose a novel decision-making framework utilizing a lightweight confidence-aware language model, which bridges the gap between complex multimodal intention reasoning and efficient inference. Specifically, we design a multi-agent collaborative workflow, comprising action voting, confidence assessment, and summarization agents, to generate high-quality, confidence-annotated decision demonstrations via explicit Chain-of-Thought (CoT) reasoning. These demonstrations are then distilled into a lightweight language model featuring a dual-head architecture, enabling the joint prediction of decision probabilities and the generation of textual rationales. The distillation is realized via a confidence-aware fine-tuning strategy coupled with Retrieval Augmented Generation (RAG) to enhance the model's adaptability and data efficiency. Comprehensive closed-loop experiments on the nuPlan benchmark demonstrate that our approach achieves state-of-the-art (SOTA) success rates in both regular and long-tail scenarios while maintaining low inference latency.",
      "short_abstract": "Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning and open-world generalization. However, the excessive computational overhead and high inference latency of...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: decision-making",
          "abstract: decision-making"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.25393",
      "pdf_url": "https://arxiv.org/pdf/2605.25393.pdf"
    },
    {
      "id": "2605.25373",
      "anchor": "paper-2605-25373",
      "title": "Physics-Aware 3D Gaussian Editing for Driving Scene Generation",
      "authors": [
        "Feng Zhou",
        "Jian Zhang",
        "Yuhang Sun",
        "He Wang",
        "Qiong Wen",
        "Debao Kong",
        "Tieru Wu",
        "Rui Ma"
      ],
      "abstract": "3D Gaussian Splatting (3DGS) has shown great potential in autonomous driving simulation and data generation, enabling photorealistic reconstruction and flexible scene manipulation. However, existing 3DGS scene editing methods have limited support for road geometry editing (e.g., inserting speed humps or sunken roads), and generally do not couple such edits with plausible vehicle-road interaction dynamics. Such editing is essential for generating training data under extreme driving scenarios or evaluating system reliability under these road irregularities. Moreover, many optimization-based methods require minutes of per-edit refinement, while existing efficient alternatives mainly focus on appearance-level or object-level manipulation rather than physics-aware road irregularity editing. To address these limitations, we propose RoVES, a Road-and-Vehicle Editing System for physics-aware 3D Gaussian editing in driving scenes. RoVES enables single-image-driven road geometry insertion and couples the edited road profile with a 4-DOF half-car vehicle dynamics model to achieve physics-aware vehicle pose correction in vertical displacement and pitch. RoVES inserts road elements in a one-shot, optimization-free pipeline (1.84s), and the full pipeline (including color transfer and vehicle-dynamics-based pose correction) completes in 6.24s; it edits dynamic vehicles via pose editing and corrects poses frame-by-frame to approximate dynamics-consistent vertical displacement and pitch responses. Experiments on the Waymo dataset show that RoVES provides practical efficiency and competitive visual consistency for physics-aware driving scene generation.",
      "short_abstract": "3D Gaussian Splatting (3DGS) has shown great potential in autonomous driving simulation and data generation, enabling photorealistic reconstruction and flexible scene manipulation. However, existing 3DGS scene editing methods have limited support for road...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: vehicle dynamics"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.25373",
      "pdf_url": "https://arxiv.org/pdf/2605.25373.pdf"
    },
    {
      "id": "2605.25308",
      "anchor": "paper-2605-25308",
      "title": "Stabilizing Streaming Video Geometry via Dynamic Feature Normalization",
      "authors": [
        "Xiaoyang Lyu",
        "Muxin Liu",
        "Xiaoshan Wu",
        "Ruicheng Wang",
        "Yi-Hua Huang",
        "Yang-Tian Sun",
        "Shaoshuai Shi",
        "Xiaojuan Qi"
      ],
      "abstract": "Consistent 3D geometry estimation from streaming RGB input is crucial for real-world applications such as autonomous driving, embodied AI, and large-scale reconstruction. While modern monocular geometry foundation models achieve strong single-image accuracy, they exhibit severe temporal inconsistency on continuous input, notably dominated by scale--shift drifting. Through targeted empirical analysis, we trace this instability to its root cause: fluctuations in latent feature statistics, whose mean and variance directly determine the predicted depth's scale and shift. Building on this insight, we introduce Dynamic Feature Normalization (DyFN), a lightweight, causal recurrent module that dynamically and robustly modulates feature statistics to maintain stable geometry over time. We adapt powerful pretrained monocular geometry models for streaming by finetuning only DyFN, a mere 2\\% additional parameters, while keeping the backbone frozen, thereby achieving temporal consistency without compromising single-image accuracy. Extensive experiments across four benchmarks show that DyFN effectively eliminates temporal artifacts such as disjointed layering and positional jitter, and achieves state-of-the-art temporal stability, improving over prior streaming methods by up to 14\\% and even outperforming heavier non-causal video baselines. Project Page: https://shawlyu.github.io/DyFN",
      "short_abstract": "Consistent 3D geometry estimation from streaming RGB input is crucial for real-world applications such as autonomous driving, embodied AI, and large-scale reconstruction. While modern monocular geometry foundation models achieve strong single-image accuracy,...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.25308",
      "pdf_url": "https://arxiv.org/pdf/2605.25308.pdf"
    },
    {
      "id": "2605.25293",
      "anchor": "paper-2605-25293",
      "title": "Neuromorphic LiDAR-based Bird's Eye View Object Detection using Energy-efficient Spiking Neural Networks",
      "authors": [
        "Sambit Mohapatra",
        "Senthil Yogamani",
        "Heinrich Gotzig",
        "Patrick Mader"
      ],
      "abstract": "Autonomous driving perception demands accurate and efficient processing of three-dimensional sensor data under strict power constraints. Traditional convolutional neural networks achieve strong detection accuracy but are computationally intensive, limiting their suitability for deployment on resource-constrained neuromorphic platforms. Spiking neural networks offer a compelling alternative through event-driven sparse computation, yet their application to complex real-world perception tasks such as three-dimensional object detection remains limited. In this work, we propose an end-to-end spiking encoder-decoder network for object detection in bird's eye view representations of LiDAR point clouds, trained using surrogate gradient backpropagation. We train two variants: a membrane potential variant that reads continuous neuron state at the output stage for maximum accuracy, achieving $92.05$/$87.04$/$86.51$ AP at $\\mathrm{IoU}\\!=\\!0.5$ (Easy/Moderate/Hard), and, a fully binary spiking variant that operates exclusively on spike trains at every layer for direct neuromorphic deployment. We evaluate four input spike encoding strategies and demonstrate that allowing the network to learn spike representations directly from data outperforms hand-crafted Poisson, latency, and z-axis encoding schemes on the KITTI benchmark, where sequential frames are unavailable and the BEV input is presented repeatedly across timesteps as a proxy for temporal streaming. A block-wise energy analysis demonstrates a $3.33\\times$ reduction in synaptic operation energy over an equivalent CNN under conservative loop-based operation. Together, these results demonstrate the viability of spiking neural networks for accurate and energy-efficient neuromorphic perception in autonomous driving.",
      "short_abstract": "Autonomous driving perception demands accurate and efficient processing of three-dimensional sensor data under strict power constraints. Traditional convolutional neural networks achieve strong detection accuracy but are computationally intensive, limiting...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 29,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.25293",
      "pdf_url": "https://arxiv.org/pdf/2605.25293.pdf"
    },
    {
      "id": "2605.25262",
      "anchor": "paper-2605-25262",
      "title": "Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection",
      "authors": [
        "Prabuddhi Wariyapperuma",
        "Rajitha de Silva",
        "Marc Hanheide",
        "Thomas Bohné",
        "Leonardo Guevara"
      ],
      "abstract": "Accurate 3D bird's-eye view (BEV) object detection is essential for autonomous driving, and depends strongly on effective multimodal representations from complementary sensors such as cameras and LiDAR. Multimodal masked autoencoders have shown strong potential for learning such representations for downstream 3D BEV object detection. However, existing methods typically apply uniform random masking to camera and LiDAR inputs, treating all regions equally, and learn representations only through masked reconstruction. We propose a semantics-guided multimodal masked autoencoder framework that introduces semantic information during pretraining through two separate components: (i) semantics-guided LiDAR voxel masking, which preserves semantically important LiDAR regions more strongly, and (ii) an auxiliary point-wise LiDAR semantic decoder branch that injects semantic guidance in addition to reconstruction. On BEVFusion 3D object detection, our semantics-guided pretraining strategy improves performance on the nuScenes mini validation set compared to the standard UniM2AE baseline: semantics-guided LiDAR voxel masking yields +1.49% mean Average Precision (mAP) and +1.66% nuScenes Detection Score (NDS), while decoder-side point semantic supervision yields +1.39% mAP and +3.22% NDS over the baseline.",
      "short_abstract": "Accurate 3D bird's-eye view (BEV) object detection is essential for autonomous driving, and depends strongly on effective multimodal representations from complementary sensors such as cameras and LiDAR. Multimodal masked autoencoders have shown strong...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 26,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: LiDAR",
          "abstract: camera",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.25262",
      "pdf_url": "https://arxiv.org/pdf/2605.25262.pdf"
    },
    {
      "id": "2605.25095",
      "anchor": "paper-2605-25095",
      "title": "RECTOR: Priority-Aware Rule-Based Reranking for Compliance-Aware Autonomous Driving Trajectory Selection",
      "authors": [
        "Hadi Hajieghrary",
        "Benedikt Walter",
        "Chaitanya Shinde",
        "Paul Schmitt",
        "Miguel Hurtado"
      ],
      "abstract": "Autonomous driving stacks must pick one trajectory from a multi-modal candidate set; choosing by model confidence ignores safety, traffic-law, and comfort constraints. We present \\textsc{RECTOR} (Rule-Enforced Constrained Trajectory Orchestrator), a post-generation reranking layer that scores candidates against a tiered rulebook (Safety~$\\succ$~Legal~$\\succ$~Road~$\\succ$~Comfort) via differentiable proxies and a scene-conditioned applicability mechanism, then selects with a deterministic $\\varepsilon$-lexicographic rule that preserves cross-tier priority by construction -- without retraining the predictor. On the Waymo Open Motion Dataset \\texttt{validation\\_interactive} split (43{,}219 augmented instances, $K{=}6$), under Protocol~B (28-rule proxy catalog, oracle applicability) rule-aware selection cuts Safety+Legal violations from 28.58\\% to 20.42\\% and Total from 40.32\\% to 32.41\\% versus confidence-only on the same candidates. A uniform-weight weighted-sum baseline matches binary compliance on this benchmark -- the empirical lift comes from rule-aware ranking, while the lexicographic guarantee is the structural differentiator no weight calibration can replicate. Under adversarial confidence corruption, confidence-only selection fails in 100\\% of scenarios while both rule-aware selectors reject the injected mode in $\\sim$96\\%. All figures are proxy-evaluator results (not a safety certificate), open-loop, 5\\,s horizon, U.S.\\ rules, validation split.",
      "short_abstract": "Autonomous driving stacks must pick one trajectory from a multi-modal candidate set; choosing by model confidence ignores safety, traffic-law, and comfort constraints. We present \\textsc{RECTOR} (Rule-Enforced Constrained Trajectory Orchestrator), a...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: attack or threat"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.25095",
      "pdf_url": "https://arxiv.org/pdf/2605.25095.pdf"
    },
    {
      "id": "2605.24950",
      "anchor": "paper-2605-24950",
      "title": "ARCANE-PedSynth: Synthetic Multi-Pedestrian Datasets with Behavioural Crossing Annotations",
      "authors": [
        "Muhammad Naveed Riaz",
        "Maciej Wielgosz",
        "Antonio M. López Peña"
      ],
      "abstract": "We present ARCANE-PedSynth, an open-source CARLA-based software framework for generating synthetic multi-pedestrian datasets with dense behavioural annotations for pedestrian crossing prediction in autonomous driving. The framework overcomes CARLA's native 9% crossing rate through a hybrid AI-manual pedestrian control architecture, enabling configurable target rates up to 75%. A 12-state behavioural finite state machine with five character archetypes produces diverse crossing behaviours. The framework generates synchronised RGB, LiDAR, and DVS data with per-frame crossing labels, behavioural states, and estimated 2D pose keypoints. We demonstrate ARCANE-PedSynth through PedSynth++, an example dataset generated with the framework, comprising 533 multi-pedestrian clips across 12 weather conditions with RGB, LiDAR, and DVS streams. ARCANE-PedSynth is fully reproducible via CLI parameterisation and Docker containerisation.",
      "short_abstract": "We present ARCANE-PedSynth, an open-source CARLA-based software framework for generating synthetic multi-pedestrian datasets with dense behavioural annotations for pedestrian crossing prediction in autonomous driving. The framework overcomes CARLA's native 9%...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.24950",
      "pdf_url": "https://arxiv.org/pdf/2605.24950.pdf"
    },
    {
      "id": "2605.25021",
      "anchor": "paper-2605-25021",
      "title": "Highway Readiness Assessment for SAE Levels of Automation and V2X Notification",
      "authors": [
        "Lorenzo Italiano",
        "Federico Marino",
        "Mattia Brambilla",
        "Giovanni Megna",
        "Benedetto Carambia",
        "Monica Nicoli"
      ],
      "abstract": "While highway automation is advancing rapidly, road operators still lack practical methods to assess the readiness of their infrastructures for supporting automated driving systems. This work proposes a quantitative Highway Readiness Index (HRI) that maps static Operational Design Domain (ODD) infrastructure conditions into measurable attributes and weights them through an expert survey to evaluate readiness across Society of Automotive Engineers (SAE) automation levels. A real corridor case study shows how HRI scores can be computed, interpreted, and used to identify infrastructure gaps that limit higher automation. Finally, we outline how these indicators can be integrated into a standardized Cooperative Intelligent Transport System (C-ITS) message, i.e., Infrastructure-to-Vehicle Information Message (IVIM), to communicate segment-level automation guidance to connected vehicles.",
      "short_abstract": "While highway automation is advancing rapidly, road operators still lack practical methods to assess the readiness of their infrastructures for supporting automated driving systems. This work proposes a quantitative Highway Readiness Index (HRI) that maps...",
      "published": "2026-05-24",
      "updated": "2026-05-24",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 25,
        "evidence": [
          "title: V2X",
          "abstract: cooperative systems",
          "abstract: connected vehicles"
        ]
      },
      "recency": "fresh",
      "age_days": 87,
      "arxiv_url": "https://arxiv.org/abs/2605.25021",
      "pdf_url": "https://arxiv.org/pdf/2605.25021.pdf"
    },
    {
      "id": "2605.24562",
      "anchor": "paper-2605-24562",
      "title": "PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction",
      "authors": [
        "Naman Mishra",
        "Shankar Gangisetty",
        "C. V. Jawahar"
      ],
      "abstract": "Pedestrian intention and trajectory prediction are critical for the safe deployment of autonomous driving systems, directly influencing navigation decisions in complex traffic environments. Recent advances in large vision-language models offer a powerful new paradigm for these tasks by combining high-capacity visual understanding with flexible natural language reasoning. In this work, we introduce PedestrianQA, a large-scale video-based dataset that formulates pedestrian intention and trajectory prediction as question-answering tasks augmented with structured rationales. PedestrianQA expresses richly annotated pedestrian sequences, in natural language, enabling VLMs to learn from visual dynamics, contextual cues, and interactions among traffic agents while generating concise explanations of their predictions without needing specialized architectures tailored for each task. Empirical evaluations across PIE, JAAD, TITAN, and IDD-PeD show that finetuning state-of-the-art VLMs on PedestrianQA significantly improves intention classification, trajectory forecasting accuracy, and the quality of explanatory rationales, demonstrating the strong potential of VLMs as a unified and explainable framework for safety-critical pedestrian behavior modeling.",
      "short_abstract": "Pedestrian intention and trajectory prediction are critical for the safe deployment of autonomous driving systems, directly influencing navigation decisions in complex traffic environments. Recent advances in large vision-language models offer a powerful new...",
      "published": "2026-05-23",
      "updated": "2026-05-23",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Dataset",
        "Benchmark",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 27,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "fresh",
      "age_days": 88,
      "arxiv_url": "https://arxiv.org/abs/2605.24562",
      "pdf_url": "https://arxiv.org/pdf/2605.24562.pdf"
    },
    {
      "id": "2605.24354",
      "anchor": "paper-2605-24354",
      "title": "SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation",
      "authors": [
        "Ruoyu Wang",
        "Jingke Wang",
        "Yukai Ma",
        "Yuehao Huang",
        "Shuangming Lei",
        "Guanglin Xu",
        "Aixue Ye",
        "Yong Liu"
      ],
      "abstract": "Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, existing driving world models are typically built upon dense scene representations, causing high computational costs and redundant information. In this paper, we present SparseWorld, a lightweight world model that focuses on predicting only the critical layout of the scene, enabling efficient future forecasting for end-to-end driving systems. SparseWorld first performs autoregressive rollout to forecast future map elements and surrounding agents, enabling the model to learn how driving scenarios evolve over time. It then leverages these predicted futures to refine downstream motion prediction and trajectory planning. Specifically, we propose a Sparse Dreamer that anticipates future instances in the latent space through joint temporal and spatial attention. By interacting with predicted future instances, the motion planner captures more accurate motion patterns and generates more informed and safety-aware trajectories. Extensive experiments demonstrate that SparseWorld significantly reduces collision risk and achieves state-of-the-art performance on the open-loop planning metrics of the nuScenes dataset with a collision rate of 0.05\\%. Moreover, it substantially outperforms the baseline method in closed-loop planning metrics on the Bench2Drive benchmark. Supplementary material is available at the project page: https://wryzju.github.io/SparseWorld/.",
      "short_abstract": "Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, existing driving world models are typically built upon dense scene...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 10,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.24354",
      "pdf_url": "https://arxiv.org/pdf/2605.24354.pdf"
    },
    {
      "id": "2605.24098",
      "anchor": "paper-2605-24098",
      "title": "D2-V2X: Depth-Driven Cooperative V2X Reasoning for Autonomous Driving",
      "authors": [
        "Kevin Richard",
        "Alphin Varghese",
        "Colin Pham",
        "David Oh",
        "Srijan Das"
      ],
      "abstract": "Single-vehicle Vision-Language Models (VLMs) are fundamentally constrained by sensor occlusions. While Vehicle-to-Everything (V2X) systems mitigate this, current benchmarks lack the cooperative reasoning required for resolving ambiguities in complex environments. We introduce D2-V2X, a spatially-aware Question-Rationale-Answer (QRA) benchmark featuring 8,500 triplets derived from multimodal vehicle and infrastructure sensors. We additionally establish a baseline that aligns 3D LiDAR features with the VLM's latent space. By enforcing natural language Chain-of-Thought rationales prior to structured JSON outputs, our model is forced to explicitly articulate spatial relations. Our experiments demonstrate that grounding VLMs in cooperative LiDAR achieves 24.4% recall in identifying occluded hazards compared to near-zero in zero-shot models and reduces spatial estimation error for visible objects by 77% compared to the zero-shot baseline. While the model achieves a functional decision-making F1-score of 53.5, we identify 3D-to-2D projection as a fundamental bottleneck in current VLM architectures, establishing a new baseline for future innovation. Data, code, and trained models available at https://github.com/KevinRichard1/D2-V2X",
      "short_abstract": "Single-vehicle Vision-Language Models (VLMs) are fundamentally constrained by sensor occlusions. While Vehicle-to-Everything (V2X) systems mitigate this, current benchmarks lack the cooperative reasoning required for resolving ambiguities in complex...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Benchmark",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 32,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.24098",
      "pdf_url": "https://arxiv.org/pdf/2605.24098.pdf"
    },
    {
      "id": "2605.23630",
      "anchor": "paper-2605-23630",
      "title": "To Overlay or to Customize? Revisiting Architectural Choices in Heterogeneous Systems",
      "authors": [
        "Xingzhen Chen",
        "Shixin Ji",
        "Zheng Dong",
        "Peipei Zhou"
      ],
      "abstract": "In this work, we present a systematic study of this trade-off from a deployment-centric perspective, focusing on an autonomous driving scenario. Instead of treating overlay and customized acceleration as isolated design points, we analyze when each approach is preferable under practical conditions, including workload variation, architectural design, reconfiguration latency, and switching frequency. Our analysis shows that overlay-based architecture is more suitable for highly frequent model switching under the state-of-the-art architecture. However, as bitstream reload overhead continues to reduce, customized architectures may become increasingly attractive, especially for workloads with efficiency requirements. Conversely, if overlay architectures become more capable and flexible, they may further expand their advantage over customized architectures. These observations provide design insights for future architectural design, and the optimal deployment strategy will be flipped according to the technique development.",
      "short_abstract": "In this work, we present a systematic study of this trade-off from a deployment-centric perspective, focusing on an autonomous driving scenario. Instead of treating overlay and customized acceleration as isolated design points, we analyze when each approach...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: systems constraints",
          "Control & Vehicle Dynamics: abstract: steering or braking"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.23630",
      "pdf_url": "https://arxiv.org/pdf/2605.23630.pdf"
    },
    {
      "id": "2605.23406",
      "anchor": "paper-2605-23406",
      "title": "RS2AD-LiDAR: End-to-End Autonomous Driving LiDAR Data Generation from Roadside Sensor Observations",
      "authors": [
        "Runyi Huang",
        "Ni Ding",
        "Ruidan Xing",
        "Yuheng Shi",
        "Lei He",
        "Keqiang Li"
      ],
      "abstract": "End-to-end autonomous driving solutions, which directly process multimodal sensory data and output fine-grained control commands, have gradually become a mainstream direction with the development of autonomous driving technology. However, current methods in this category rely on single-vehicle data collection for model training and optimization, which suffers from high acquisition and annotation costs, scarcity of valuable scenarios, and data silos. To address these challenges, we propose RS2AD-LiDAR, a novel framework for reconstructing and generating vehicle-mounted LiDAR data from roadside sensor observations. Since no public dataset currently provides highly overlapping perception coverage between roadside and vehicle-mounted LiDAR sensors, which is essential for studying roadside-to-vehicle data generation, we constructed a dedicated dataset named R2V-LiDAR which is used solely for evaluation in this work. Specifically, our method transforms roadside LiDAR point clouds into the vehicle-mounted LiDAR coordinate system, and synthesizes high-fidelity vehicle-mounted data via virtual LiDAR modeling and point cloud resampling techniques. To the best of our knowledge, this is the first approach to reconstruct vehicle-mounted LiDAR data from roadside sensor inputs. Extensive experimental comparisons demonstrate the semantic similarity between the generated data and real data. Furthermore, object detection experiments show that incorporating the generated data into real data for model training improves both Bird's Eye View (BEV) and 3D detection accuracy, thereby validating the effectiveness of the proposed method.",
      "short_abstract": "End-to-end autonomous driving solutions, which directly process multimodal sensory data and output fine-grained control commands, have gradually become a mainstream direction with the development of autonomous driving technology. However, current methods in...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 12,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.23406",
      "pdf_url": "https://arxiv.org/pdf/2605.23406.pdf"
    },
    {
      "id": "2605.23327",
      "anchor": "paper-2605-23327",
      "title": "GFSR: Geometric Fidelity and Spatial Refinement for Reliable Lane Detection",
      "authors": [
        "Tiancheng Wang",
        "Zhaolu Ding",
        "Richeng Xu",
        "Tianhui Zheng",
        "Hui Liu",
        "Hanyu Xuan",
        "Zhiliang Wu",
        "Guanghui Yue"
      ],
      "abstract": "Lane detection stands as a crucial perception task in autonomous driving and advanced driver assistance systems. However, existing methods still degrade in complex real scenarios due to two major limitations. First, classification confidence only characterizes the categorical existence of lane priors and has no strong correlation with geometric quality. If threshold filtering and NMS are conducted merely based on this confidence, the model tends to retain lane priors with high confidence while eliminating those with lower confidence but superior geometric representation. Secondly, the regression modules in existing methods weaken correlations among sampling points, hindering fine-grained optimization of distant, high-curvature and complex-topology lanes and causing underfitting. To address these issues, we propose Geometric Fidelity and Spatial Refinement (GFSR), a framework consisting of LaneIoU-guided Confidence Calibration (LCC) and Adaptive Gated Location Refinement (AGLR). Specifically, LCC adopts LaneIoU as soft supervision to explicitly estimate the geometric fidelity of lane priors, which is further fused with classification confidence to construct the Collaborative Reliability Index (CRI). This index guides lane prior filtering, effectively retaining those with high classification confidence and favorable geometric quality. Meanwhile, cooperating with regression heads in each refinement stage, AGLR predicts sampling point lateral offsets and adopts a gating mechanism to adaptively regulate correction magnitude, strengthen inter-point correlations and boost model adaptability as well as robustness toward complex lane scenarios. Extensive experiments on CULane and CurveLanes demonstrate that our GFSR achieves state-of-the-art performance on CULane, with F1_50 and F1_75 scores of 81.46% and 65.01%, and reaches 87.35% F1_50 on CurveLanes.",
      "short_abstract": "Lane detection stands as a crucial perception task in autonomous driving and advanced driver assistance systems. However, existing methods still degrade in complex real scenarios due to two major limitations. First, classification confidence only...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 12,
        "evidence": [
          "abstract: perception",
          "title: lane detection",
          "abstract: lane detection"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.23327",
      "pdf_url": "https://arxiv.org/pdf/2605.23327.pdf"
    },
    {
      "id": "2605.23270",
      "anchor": "paper-2605-23270",
      "title": "ChainFlow-VLA: Causal Flow Planning with Vision-Language Models",
      "authors": [
        "Xiyang Wang",
        "Xinlin Wang",
        "Tingguang Zhou",
        "Gong Chen",
        "Xingtai Gui",
        "Zhi Xu",
        "Xiaolei Wu",
        "Feiyang Tan",
        "Hangning Zhou",
        "Mu Yang"
      ],
      "abstract": "Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but their step-wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints, making them unreliable in interactive and safety-critical scenarios. This dichotomy reveals a deeper issue: existing methods treat causal modeling and global optimization as separate paradigms, without a principled way to unify them within a single trajectory distribution. To address this, we propose ChainFlow-VLA, which unifies causal generation and global refinement within a unified probabilistic framework. We formulate planning as a mixture over AR-induced modes and learn Vision-Language Model (VLM)-conditioned residual distributions over these modes. An autoregressive generator (Chain) produces a discrete set of causal trajectory modes, followed by a diffusion-based refiner (Flow) that leverages VLM hidden states as semantic priors to perform mode-conditioned correction in residual space while preserving causal structure. This straightforward conditioning seamlessly injects high-level scene understanding into fine-grained trajectory adjustments. Experiments demonstrate that ChainFlow-VLA achieves robust planning in ambiguous and long-tail scenarios, achieving a state-of-the-art score of 94.85 on the NAVSIM v1 leaderboard, matching human-level performance (94.8). Code will be available at https://github.com/AFARI-Research/ChainFlow-VLA.",
      "short_abstract": "Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 21,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.23270",
      "pdf_url": "https://arxiv.org/pdf/2605.23270.pdf"
    },
    {
      "id": "2605.23816",
      "anchor": "paper-2605-23816",
      "title": "SDNator is Not Another SDN Controller: Enabling Extensible Data-Driven Control in Cyber-Physical Systems",
      "authors": [
        "Y. Lin",
        "R. Zhang",
        "E. Balta",
        "X. Zhu",
        "J. Zhang",
        "K. Barton",
        "D. Tilbury",
        "Z. Mao"
      ],
      "abstract": "An SDN-like centralized control architecture is increasingly popular and has been widely explored in cyber-physical systems (CPS) such as manufacturing, internet-of-things, and autonomous vehicle systems for higher flexibility, programmability and scalability. However, no existing frameworks can offer domain-agnostic, easily extensible support for data-driven CPS applications. In this work, we design, implement, and open-source \\textit{SDNator}, the first framework to enable extensible, data-driven control in CPS. SDNator embraces an application- and data-driven design where applications function as data consumers and producers to collectively define the workflows of the controller. SDNator also incorporates two data store backends to support both event-driven and data-driven programming patterns. Benchmarks show that SDNator is highly scalable, and delivers comparable performance to Ryu, a widely used SDN controller. Moreover, we demonstrate the capabilities and usability of SDNator through our case studies of manufacturing and networking systems. By integrating applications from respective domains, we build different ``controllers'' for different scenarios. Most notably, we leverage SDNator to implement the first digital-twin-equipped central controller for additive manufacturing fleets. We show through extensive and realistic simulations that SDNator-based scheduling can (1) significantly shorten production time and improve reliability in the presence of anomalies compared to decentralized approaches, and (2) flexibly adjust and optimize production plans upon urgent requests such as producing Personal Protective Equipment during the COVID-19 pandemic.",
      "short_abstract": "An SDN-like centralized control architecture is increasingly popular and has been widely explored in cyber-physical systems (CPS) such as manufacturing, internet-of-things, and autonomous vehicle systems for higher flexibility, programmability and...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 16,
        "evidence": [
          "title: controller",
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.23816",
      "pdf_url": "https://arxiv.org/pdf/2605.23816.pdf"
    },
    {
      "id": "2605.23782",
      "anchor": "paper-2605-23782",
      "title": "Routing Equilibrium in Mixed-Autonomy Traffic Networks with Altruistic Autonomous Agents",
      "authors": [
        "Lihui Yi",
        "Ermin Wei"
      ],
      "abstract": "Recent advancements in vehicle autonomy have drawn interest in understanding the impact of autonomous vehicles on traffic systems. In this paper, we study a traffic assignment problem in a mixed-autonomy setting where both human-driven and autonomous vehicles coexist. We model the interaction as a simultaneous routing game where human drivers are self-interested and aim to minimize their own travel times, while autonomous agents are altruistic and aim to minimize the total social cost. The standard nonatomic congestion game analysis establishes the existence of equilibrium to this game under convex cost functions, and does not have any implication of its uniqueness. In this work, we formulate the equilibrium as a variational inequality (VI), which enables us to establish the equilibrium existence without convexity assumption, and guarantees the uniqueness of the aggregated link flow and social cost at equilibrium under a specific class of cost functions. Leveraging this VI framework, we provide sufficient conditions under which including autonomous agents improves, deteriorates, or has no effect on social cost. While the possibility of deterioration has been established in prior work, our results complement existing worst-case bounds by explicitly characterizing sufficient conditions under which each outcome occurs, thereby providing a deeper understanding of mixed-autonomy traffic systems. Furthermore, we consider a centralized scenario where a social planner optimizes the routing of autonomous agents, and show that the same equilibrium is achieved as in the decentralized scenario when assuming convex costs.Finally, we conduct numerical experiments that illustrate how social cost changes with the amount of autonomous vehicles under different system parameters.",
      "short_abstract": "Recent advancements in vehicle autonomy have drawn interest in understanding the impact of autonomous vehicles on traffic systems. In this paper, we study a traffic assignment problem in a mixed-autonomy setting where both human-driven and autonomous vehicles...",
      "published": "2026-05-22",
      "updated": "2026-05-22",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "title: communications"
        ]
      },
      "recency": "fresh",
      "age_days": 89,
      "arxiv_url": "https://arxiv.org/abs/2605.23782",
      "pdf_url": "https://arxiv.org/pdf/2605.23782.pdf"
    },
    {
      "id": "2605.23203",
      "anchor": "paper-2605-23203",
      "title": "Lipschitz Optimization for Formal Verification of Homographies",
      "authors": [
        "Jean-Guillaume Durand",
        "Panagiotis Kouvaros",
        "Maxime Gariel",
        "Alessio Lomuscio"
      ],
      "abstract": "The adoption of vision neural networks in regulated industries requires formal robustness guarantees, especially in safety-critical domains such as healthcare, autonomous vehicles, and aerospace. However, current approaches are confined to incomplete statistical verification or robustness to $\\ell_p$-norm and affine transforms, which cover only a narrow subset of perturbations to the image formation process. In particular, robustness to camera motion remains an open problem despite being key to deploy many vision applications. We present a formal verification approach that targets robustness against 3D motion perturbations of the capturing camera. We first establish a closed-form mapping from camera pose to pixel values. By analyzing the continuity properties of the resulting homographies, we show that recent work on Lipschitz optimization and piecewise continuity can be extended to derive tight linear bounds on perturbed pixel values. Our approach applies to scenes with predominantly planar structure, such as ground planes in augmented reality, road markings and traffic signs in autonomous driving, or planar workspaces in robotic manipulation. This enables the first formal verification of projective geometry transforms, without complex simulation, surrogate networks, or explicit image-formation models. We validate our implementation and show up to 89% speedup and 7% tighter bounds over prior work. We then evaluate our method on the VNN-COMP benchmark and reveal systematic weaknesses to projective perturbations. Finally, we demonstrate a real-world case study on a safety-critical runway classifier, highlighting practical vulnerabilities to camera motion, and addressing a key challenge in the certification of learned models. Data and code are publicly available at https://github.com/jeangud/homography-verification .",
      "short_abstract": "The adoption of vision neural networks in regulated industries requires formal robustness guarantees, especially in safety-critical domains such as healthcare, autonomous vehicles, and aerospace. However, current approaches are confined to incomplete...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 22,
        "evidence": [
          "abstract: safety",
          "title: verification",
          "abstract: verification",
          "abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.23203",
      "pdf_url": "https://arxiv.org/pdf/2605.23203.pdf"
    },
    {
      "id": "2605.23176",
      "anchor": "paper-2605-23176",
      "title": "DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving",
      "authors": [
        "Hao Vo",
        "Khoa Vo",
        "Phu Loc Nguyen",
        "Sieu Tran",
        "Duc Minh Nguyen",
        "Ngo Xuan Cuong",
        "Gladys Gawugah",
        "Sreevenkata Anjani Tishita Godavarthi",
        "Chase Rainwater",
        "Nghi D. Q. Bui",
        "Anh Nguyen",
        "Duy Minh Ho Nguyen",
        "Ngan Le"
      ],
      "abstract": "Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations, interactions, and future dynamics. However, existing AD vision-language benchmarks largely focus on single-view, static, ego-centric, or single-source question answering, leaving it unclear whether current Vision-Language Models (VLMs) can truly construct and reason over dynamic driving scenes. We introduce DriveSpatial, a benchmark of 15.6K human-verified QA pairs across 20 tasks from five large-scale AD datasets. DriveSpatial evaluates four abilities: Cognitive Scene Construction, Multi-view Relational Understanding, Temporal Reasoning, and Generalization. Unlike prior benchmarks, DriveSpatial is generated from a dynamic multi-relational scene graph that encodes object states, spatial relations, interactions, camera visibility, and temporal correspondences, enabling QA pairs that enforce genuine cross-view and spatiotemporal reasoning. Evaluating 15 representative VLMs reveals a substantial human-model gap: the strongest model trails humans by 28.4 points, with Cognitive Scene Construction emerging as the key bottleneck. Further diagnostics show that language-only prompting is insufficient, while explicit BEV grounding consistently improves performance. These results suggest that current VLMs lack the scene-construction ability needed for reliable spatiotemporal driving intelligence. DriveSpatial and its construction pipeline will be released to support future research.",
      "short_abstract": "Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations, interactions, and...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera",
          "Perception & Sensor Fusion: abstract: bird's-eye view"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.23176",
      "pdf_url": "https://arxiv.org/pdf/2605.23176.pdf"
    },
    {
      "id": "2605.23163",
      "anchor": "paper-2605-23163",
      "title": "Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving",
      "authors": [
        "Kewei Zhang",
        "Jin Wang",
        "Sensen Gao",
        "Chengyue Wu",
        "Yulong Cao",
        "Songyang Han",
        "Boris Ivanovic",
        "Langechuan Liu",
        "Marco Pavone",
        "Song Han",
        "Daquan Zhou",
        "Enze Xie"
      ],
      "abstract": "End-to-end autonomous driving via Vision-Language-Action (VLA) models demands a precarious balance between high-fidelity trajectory planning and efficient inference. Existing paradigms typically fall short: autoregressive (AR) VLAs are memory-bandwidth-bound on edge hardware and prone to exposure-bias drift, while full-sequence diffusion models preclude KV-cache reuse and suffer from \"logical leakage\" that violates the fundamental perceive-then-plan causality. We present Fast-dDrive, a block-diffusion VLA that performs bidirectional refinement within semantic units while enforcing strict causal ordering across them. Leveraging the observation that driving VLAs often emit structured JSON-like outputs, Fast-dDrive freezes structural tokens into a section scaffold and employs a section-aware training recipe that prioritizes safety-critical planning. We further introduce Scaffold Speculative Decoding to achieve AR-equivalent quality at significantly higher throughput. Finally, we propose a low-overhead test-time scaling scheme: by forking $N$ stochastic trajectory rollouts from a single shared-prefix KV cache and averaging them, we effectively suppress prediction variance at a fractional computational cost. Empirical results demonstrate that Fast-dDrive redefines the speed-accuracy frontier for driving agents. On the WOD-E2E test set, Fast-dDrive achieves SOTA ADE@3s and ADE@5s, alongside the highest RFS among diffusion-based VLAs; on nuScenes, it reduces average L2 error to $0.32$m (a $22\\%$ improvement). When integrated with SGLang, our framework delivers $12\\times$ throughput speedup over the AR baseline, narrowing the gap between high-capacity VLAs and the efficiency demands of real-time on-vehicle deployment.",
      "short_abstract": "End-to-end autonomous driving via Vision-Language-Action (VLA) models demands a precarious balance between high-fidelity trajectory planning and efficient inference. Existing paradigms typically fall short: autoregressive (AR) VLAs are memory-bandwidth-bound...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 7,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.23163",
      "pdf_url": "https://arxiv.org/pdf/2605.23163.pdf"
    },
    {
      "id": "2605.22997",
      "anchor": "paper-2605-22997",
      "title": "Scene Reconstruction as Mapping Priors for 3D Detection",
      "authors": [
        "Yang Fu",
        "Yuliang Zou",
        "Hao Xiang",
        "Xin Huang",
        "Yijing Bai",
        "Chen Song",
        "Weijing Shi",
        "Govind Thattai",
        "Dragomir Anguelov",
        "Mingxing Tan",
        "Yingwei Li"
      ],
      "abstract": "In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection. Maps can provide robust structural priors of the static environment, helping resolve ambiguities and correct for sensor data sparsity or noise, especially for distant objects or under adverse weather conditions. However, conventional High-Definition (HD) maps are resource-intensive to obtain and maintain, which presents a challenge for efficient, large-scale deployment. In this paper, we propose a scalable solution to systematically leverage mapping to improve 3D detection by overcoming two primary challenges. First, we introduce a pipeline to automatically build dense mapping priors from aggregated sensor data, eliminating the need for human labeling. Second, we design a novel Mapping Priors Augmented 3D Detection (MPA3D) framework to effectively integrate mapping priors with different sensor modalities. Extensive experiments on the Waymo Open Dataset demonstrate that our approach achieves new state-of-the-art results, proving the effectiveness of scalable reconstructed scene priors for enhancing 3D detection.",
      "short_abstract": "In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection. Maps can provide robust structural priors of the static environment, helping resolve ambiguities and...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22997",
      "pdf_url": "https://arxiv.org/pdf/2605.22997.pdf"
    },
    {
      "id": "2605.22809",
      "anchor": "paper-2605-22809",
      "title": "Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving",
      "authors": [
        "Jiahao Wang",
        "Bo Sun",
        "Yijing Bai",
        "Vincent Casser",
        "Songyou Peng",
        "Zehao Zhu",
        "Meng-Li Shih",
        "Xander Masotto",
        "Shih-Yang Su",
        "Kanaad V Parvate",
        "Tiancheng Ge",
        "Linn Bieske",
        "Dragomir Anguelov",
        "Mingxing Tan",
        "Chiyu Max Jiang"
      ],
      "abstract": "Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (AV) fleets, while high-fidelity, are limited in scale, diversity of sensor configurations, as well as geographic and long-tail-behavioral coverage. In contrast, in-the-wild data from sources like dashcams offers immense scale and diversity, capturing critical long-tail scenarios and novel environments. However, this unstructured, in-the-wild video data is incompatible with ADS expecting structured, multi-modal sensor inputs for validation and training. To bridge this data gap, we propose Sensor2Sensor, a novel generative modeling paradigm that translates in-the-wild monocular dashcam videos into a high-fidelity, multi-modal sensor suite (AV logs) comprising multi-view camera images and LiDAR point clouds. A core challenge is the lack of paired training data. We address this by converting real AV logs into dashcam-style videos via 4D Gaussian Splatting (4DGS) reconstruction and novel-view rendering. Sensor2Sensor then utilizes a diffusion architecture to perform the generative conversion. We perform comprehensive quantitative evaluations on the fidelity and realism of the generated sensor data. We demonstrate Sensor2Sensor's practical utility by converting challenging in-the-wild internet and dashcam footage into realistic, multi-modal data formats, further unlocking vast external data sources for AV development.",
      "short_abstract": "Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (AV) fleets, while high-fidelity, are limited in scale, diversity of sensor configurations, as well as...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 3,
        "evidence": [
          "abstract: point cloud",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22809",
      "pdf_url": "https://arxiv.org/pdf/2605.22809.pdf"
    },
    {
      "id": "2605.22600",
      "anchor": "paper-2605-22600",
      "title": "Branch-Stochastic Model Predictive Control for Motion Planning under Multi-Modal Uncertainty with Scenario Clustering",
      "authors": [
        "Zekun Xing",
        "Ramkrishna Chaudhari",
        "Marion Leibold",
        "Dirk Wollherr",
        "Martin Buss"
      ],
      "abstract": "Motion planning for autonomous driving must account for multi-modal uncertainty in both the intentions and trajectories of surrounding vehicles. Handling uncertainty in a worst-case manner guarantees robustness but often leads to excessive conservatism. Stochastic Model Predictive Control (SMPC) reduces trajectory-level conservatism through chance constraints, yet remains conservative with respect to intention uncertainty since constraints must hold across all intentions. We present a novel combination of SMPC and the branching structure, enabling the planner to generate distinct trajectories for different possible intentions while maintaining safety under trajectory uncertainty. A novel scenario clustering is proposed to merge prediction scenarios based on high-level decision similarity, thereby ensuring real-time tractability. Furthermore, an adaptive branching-time computation postpones commitment to separate plans until intention uncertainty is sufficiently reduced. Simulation studies in challenging highway scenarios demonstrate that the proposed method improves safety, reduces conservatism, and achieves real-time computational performance.",
      "short_abstract": "Motion planning for autonomous driving must account for multi-modal uncertainty in both the intentions and trajectories of surrounding vehicles. Handling uncertainty in a worst-case manner guarantees robustness but often leads to excessive conservatism....",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 32,
        "margin": 4,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22600",
      "pdf_url": "https://arxiv.org/pdf/2605.22600.pdf"
    },
    {
      "id": "2605.22578",
      "anchor": "paper-2605-22578",
      "title": "Beyond Chamfer Distance: Granular Order-aware Evaluation Metric For Online Mapping",
      "authors": [
        "Chouaib Bencheikh Lehocine",
        "Adam Lilja",
        "Junsheng Fu",
        "Lars Hammarstrand"
      ],
      "abstract": "Online map estimation is a crucial component of autonomous driving systems that reduces the reliance on costly high-definition maps. State-of-the-art (SOTA) methods commonly predict map elements as ordered sequences of points that form polylines and polygons. The evaluation of these methods relies predominantly on mean average precision (mAP) based on thresholded Chamfer distance (CD). This framework lacks sensitivity to point ordering and provides limited granularity in assessing geometric quality, making it difficult to distinguish which methods truly excel over others. In this work, we address these limitations on two fronts. For the single-instance similarity measure, we introduce sequence optimal sub-pattern assignment (SOSPA), an order-aware metric that enables fine-grained evaluation of individual geometries while satisfying all metric axioms. For the multi-instance evaluation framework, we propose polyline localisation and detection (PLD), a soft metric that jointly captures detection quality and geometric accuracy, replacing the hard thresholding of mAP with a principled soft assignment. Through evaluations on nuScenes, we demonstrate that PLD effectively ranks SOTA online mapping methods (MapTRv2, StreamMapNet, MapTracker) while providing a decomposed error analysis. This analysis identifies detection capability as the dominant bottleneck in current methods, revealing a performance trend that mAP fails to capture. Code for evaluation using our metrics will be released.",
      "short_abstract": "Online map estimation is a crucial component of autonomous driving systems that reduces the reliance on costly high-definition maps. State-of-the-art (SOTA) methods commonly predict map elements as ordered sequences of points that form polylines and polygons....",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 26,
        "evidence": [
          "abstract: localization",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22578",
      "pdf_url": "https://arxiv.org/pdf/2605.22578.pdf"
    },
    {
      "id": "2605.22455",
      "anchor": "paper-2605-22455",
      "title": "Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light",
      "authors": [
        "Valeria Pais",
        "Malena Mendilaharzu",
        "Daniele Faccio",
        "Luis Oala",
        "Christoph Clausen",
        "Bruno Sanguinetti"
      ],
      "abstract": "Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse and uneven: long-tailed or unbalanced distributions hinder generalization, and the low number of samples in low density regions makes it hard to run evaluations. Synthetic data can fill these gaps, providing us with a way to sample the input space more continuously and improve data coverage for benchmarks. Focusing on the autonomous driving safety-critical case of pedestrian detection in the dark, we show how synthetic low-light samples can be used to better characterize the performance of a state-of-the-art object detection model as a function of the scene illumination. We use a synthetic RAW image augmentation technique to generate low-light samples that match the noise model of the camera sensor. Performance metrics on real and synthetic low-light data are similar, indicating that the AI model finds it hard to distinguish between them.",
      "short_abstract": "Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse and uneven: long-tailed or unbalanced distributions hinder generalization, and the low number of samples in low...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22455",
      "pdf_url": "https://arxiv.org/pdf/2605.22455.pdf"
    },
    {
      "id": "2605.22420",
      "anchor": "paper-2605-22420",
      "title": "Diffusion-guided Generalizable Enhancer for Urban Scene Reconstruction",
      "authors": [
        "Henry Che",
        "Jingkang Wang",
        "Yun Chen",
        "Ze Yang",
        "Sivabalan Manivasagam",
        "Raquel Urtasun"
      ],
      "abstract": "Urban scene reconstruction from real-world observations has emerged as a powerful tool for self-driving development and testing. While current neural rendering approaches achieve high-fidelity rendering along the recorded trajectories, their quality degrades significantly under large viewpoint shifts, limiting the applicability for closed-loop simulation. Recent works have shown promising results in using diffusion models to enhance quality at these challenging viewpoints and distill improvements back into 3D representations. However, they often require costly per-scene optimization, and the distilled representations remain fragile and fail to generalize beyond limited synthesized views. To address these limitations, we propose GenRe, a novel diffusion-guided generalizable enhancer for urban scene reconstruction. GenRe takes as input any pretrained 3D Gaussian representation and fixes the deficiencies within a few minutes. By learning to distill generative priors across diverse scenes, GenRe produces robust and high-fidelity representation efficiently that generalizes reliably to challenging unseen viewpoints (e.g., lane change). Experiments show that GenRe outperforms existing methods in both quality and efficiency and benefits various downstream tasks, enabling robust and scalable sensor simulation for autonomous driving.",
      "short_abstract": "Urban scene reconstruction from real-world observations has emerged as a powerful tool for self-driving development and testing. While current neural rendering approaches achieve high-fidelity rendering along the recorded trajectories, their quality degrades...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22420",
      "pdf_url": "https://arxiv.org/pdf/2605.22420.pdf"
    },
    {
      "id": "2605.22189",
      "anchor": "paper-2605-22189",
      "title": "Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments",
      "authors": [
        "Jie Jia",
        "Yaofeng Su",
        "Zeyu Bao",
        "Yun Hong",
        "Bingzhao Gao",
        "Zhongxue Gan",
        "Wenchao Ding"
      ],
      "abstract": "Occlusion-aware prediction remains a critical challenge in autonomous driving due to the inherent uncertainty of unobserved regions. Existing approaches either overestimate risk based on reachable states or struggle to predict accurate trajectories under high occlusion uncertainty. To address these limitations, we propose a unified risk map modeling and learning framework for partially observable environments. Our method integrates traffic flow risk and collision risk through spatiotemporal modeling, enabling fine-grained assessment of occlusion-induced hazards. To address the scarcity of scenarios involving occluded interactions, we introduce a diffusion-based scenario generation framework that produces realistic yet adversarial scenarios. We integrate the modeling and learning of a unified risk map into a framework that supports risk-aware planning under partial observability. Experiments on the Waymo Open Motion Dataset show that our method significantly outperforms the state-of-the-art occlusion-aware baseline, improving minimum time-to-collision by 0.78 times and average time-to-collision by 1.67 times. The proposed framework offers a comprehensive and practical solution for risk-aware planning in partially observable environments.",
      "short_abstract": "Occlusion-aware prediction remains a critical challenge in autonomous driving due to the inherent uncertainty of unobserved regions. Existing approaches either overestimate risk based on reachable states or struggle to predict accurate trajectories under high...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 7,
        "evidence": [
          "abstract: attack or threat",
          "abstract: risk assessment",
          "abstract: uncertainty"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22189",
      "pdf_url": "https://arxiv.org/pdf/2605.22189.pdf"
    },
    {
      "id": "2605.22089",
      "anchor": "paper-2605-22089",
      "title": "LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model",
      "authors": [
        "Xiaodong Mei",
        "Diankun Zhang",
        "Hongwei Xie",
        "Guang Chen",
        "Hangjun Ye",
        "Dan Xu"
      ],
      "abstract": "Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and reasoning capabilities. Recent attempts to incorporate dense visual supervision via world modeling often overemphasize pixel-level image reconstruction, neglecting semantically meaningful scene representation learning. In this work, we propose LVDrive, a Latent Visual representation enhanced VLA framework for autonomous driving. LVDrive introduces a future scene prediction task into the VLA paradigm, where future representations are learned entirely in a high-level latent space under auxiliary supervision from a pretrained vision backbone. Departing from inefficient autoregressive generation, we jointly model future scene and motion prediction within a unified embedding space, processed in a single forward pass to conduct the future-aware reasoning. We further design a two-stage trajectory decoding strategy that explicitly leverages the learned latent future representations to refine trajectory generation. Extensive experiments on the challenging Bench2Drive benchmark demonstrate that LVDrive achieves significant improvements in closed-loop driving performance, outperforming both action supervised methods and image-reconstruction-based world model approaches.",
      "short_abstract": "Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and reasoning capabilities....",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 19,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22089",
      "pdf_url": "https://arxiv.org/pdf/2605.22089.pdf"
    },
    {
      "id": "2605.22018",
      "anchor": "paper-2605-22018",
      "title": "FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments",
      "authors": [
        "Connor Malone",
        "Sebastien Demmel",
        "Sebastien Glaser"
      ],
      "abstract": "The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collection of data from scenarios involving water hazards on the road. The dataset contains images from a 2.3 MP FLIR Blackfly USB3 camera, 64-beam 360 degree point clouds from an Ouster OS1-64 LiDAR, and data from an iXblue ATLANS-C IMU corrected by a Geoflex RTK GNSS, from five separate locations captured both during and after flooding events. The data has been released in two formats: a KITTI-style format for easy integration with existing data tools, and the RTMaps format for direct replay of the vehicle's data capture. We provide semantic labels to enable the training and evaluation of both single-sensor and sensor-fusion methods for water hazard detection. Position and velocity, as well as data captured under dry conditions, are provided to enable the development of location-based detection methods that may incorporate maps, and to evaluate other tasks such as localisation and SLAM.",
      "short_abstract": "The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collection of data from scenarios involving water hazards on the road. The dataset contains images from a 2.3 MP...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 5,
        "evidence": [
          "abstract: localization",
          "abstract: SLAM",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22018",
      "pdf_url": "https://arxiv.org/pdf/2605.22018.pdf"
    },
    {
      "id": "2605.22659",
      "anchor": "paper-2605-22659",
      "title": "A Metalens-based Bicycle Safety Reflector for Autonomous Vehicle Radars",
      "authors": [
        "Sepideh Ghasemi",
        "Jimmy Hester",
        "Aline Eid"
      ],
      "abstract": "With the rising number of interactions between autonomous or sensor-assisted vehicles -- especially in poor weather conditions -- come the need and opportunity for a new class of bicycle safety reflectors designed to enhance cyclist visibility to radars. To this effect, the first retrodirective planar metalens-based tag operating in the millimeter-wave automotive frequency range is proposed. The compact, lightweight ($0.61~\\mathrm{g}$) design consists of two layers: a metalens layer and a patch antenna pixel layer. The metalens focuses incoming plane waves from different incidence angles onto corresponding patch antenna pixels on the second layer, which re-radiate the signal back through the metalens, enabling retrodirective operation. The proposed tag was thoroughly evaluated, demonstrating reliable detection beyond 70 m and a peak monostatic radar cross section (RCS) of $3.54~\\mathrm{dBsm}$ with stable retrodirectivity over $\\pm 40^\\circ$, providing an average gain improvement of $7.58~\\mathrm{dB}$ and an RCS enhancement of $15.16~\\mathrm{dB}$ relative to a lens-less reference. A realistic deployment scenario on a metallic bicycle demonstrated up to a 110x improvement in its detectability at broadside. These results highlight the potential of the proposed passive tag to operate as a low-cost, lightweight, and easily integrable bicycle safety reflector for next-generation autonomous vehicle radar systems.",
      "short_abstract": "With the rising number of interactions between autonomous or sensor-assisted vehicles -- especially in poor weather conditions -- come the need and opportunity for a new class of bicycle safety reflectors designed to enhance cyclist visibility to radars. To...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22659",
      "pdf_url": "https://arxiv.org/pdf/2605.22659.pdf"
    },
    {
      "id": "2605.22122",
      "anchor": "paper-2605-22122",
      "title": "Adversarial Trust Poisoning in Vehicular Collaborative Perception",
      "authors": [
        "Yutong Liu",
        "Chenyi Wang",
        "Ming F. Li",
        "Qingzhao Zhang"
      ],
      "abstract": "Collaborative perception (CP) enables connected and autonomous vehicles to share sensor data and jointly reason about their environment. To defend against adversaries that fabricate or manipulate shared data, existing systems employ cross-vehicle inconsistency detection and trust estimation, penalizing vehicles whose observations conflict with the majority. In this work, we show that these defenses themselves introduce a new attack surface. We present TrustFlip, a novel attack that weaponizes consistency-based defenses to poison the trust assigned to benign vehicles. Instead of injecting false data into the collaboration pipeline, it deploys physical adversarial objects that are genuine but induce inconsistent observations among benign vehicles. The resulting inconsistencies are misattributed by the defense to the targeted vehicle, causing its trust score to degrade and eventually leading to its downweighting or exclusion from collaboration. Consequently, the system loses reliable sensing contributors, degrading perception capability and potentially inducing safety-critical failures. We evaluate TrustFlip across multiple collaborative perception architectures and defense mechanisms. Our results show that state-of-the-art defenses can be significantly affected: the attack removes the targeted benign vehicle from collaboration in up to 87.7% of scenarios and drops Average Precision (AP) by up to 13%. As an initial mitigation, we introduce TrustReflect, a lightweight self-reflection mechanism that marks disputed regions as uncertain and excludes them from trust evaluation, reducing the attack success rate by 35-100%.",
      "short_abstract": "Collaborative perception (CP) enables connected and autonomous vehicles to share sensor data and jointly reason about their environment. To defend against adversaries that fabricate or manipulate shared data, existing systems employ cross-vehicle...",
      "published": "2026-05-21",
      "updated": "2026-05-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 22,
        "margin": 6,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "fresh",
      "age_days": 90,
      "arxiv_url": "https://arxiv.org/abs/2605.22122",
      "pdf_url": "https://arxiv.org/pdf/2605.22122.pdf"
    },
    {
      "id": "2605.21446",
      "anchor": "paper-2605-21446",
      "title": "Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs",
      "authors": [
        "Abhinaw Priyadershi",
        "Jelena Frtunikj"
      ],
      "abstract": "Interpretable autonomous driving planners depend not only on generating explanations, but also on those explanations remaining reliable under real-world sensor degradation. In this paper we present a controlled perturbation study of Vision-Language-Action (VLA) robustness in autonomous driving, evaluating Alpamayo R1 (10B parameters) across 1,996 scenarios under eight sensor perturbations (Gaussian noise at four intensities, two lighting extremes, and two fog levels; ${\\sim}18{,}000$ inference trials). We find that reasoning consistency is a high-fidelity indicator of trajectory reliability: when Chain-of-Causation (CoC) explanations change after perturbation, trajectory deviation spikes $5.3{\\times}$ (21.8m vs 4.1m), with $r\\!=\\!0.99$ across attack types and $r_{pb}\\!=\\!0.53$ per-sample (Cohen's $d\\!=\\!1.12$). A controlled ablation provides evidence that enabling CoC generation is associated with improved trajectory accuracy (11.8% on average across conditions; $p < 0.0001$) under matched inference settings. Over the tested noise range ($σ\\in \\{10, 30, 50, 70\\}$), degradation is approximately linear ($R^2\\!=\\!0.957$), while standard input preprocessing defenses provide only marginal relief. Together, these results establish CoC consistency as a quantitative proxy for planning safety and motivate reasoning-based runtime monitoring for safer VLA deployment.",
      "short_abstract": "Interpretable autonomous driving planners depend not only on generating explanations, but also on those explanations remaining reliable under real-world sensor degradation. In this paper we present a controlled perturbation study of Vision-Language-Action...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "VLA",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21446",
      "pdf_url": "https://arxiv.org/pdf/2605.21446.pdf"
    },
    {
      "id": "2605.21395",
      "anchor": "paper-2605-21395",
      "title": "Towards Resilient and Autonomous Networks: A BlueSky Vision on AI-Native 6G",
      "authors": [
        "Liang Wu",
        "Kelly Wan",
        "Mayank Darbari",
        "Liangjie Hong"
      ],
      "abstract": "The proliferation of emerging applications, such as autonomous driving and immersive experiences, demands cellular networks that are not only faster, but fundamentally more resilient and autonomous. This paper presents a BlueSky vision on how Artificial Intelligence will be natively integrated into 6G, shifting the paradigm from \\underline{Network for AI} to \\underline{AI for Network}. We envision that, unlike 5G's reliance on scattered, ad-hoc models each trained for a single task, native AI in the 6G era will be anchored by a foundation model and and orchestrated via collaborative multi-agent systems, framing network management as a unified, multi-modal, multi-task optimization problem. Built on this vision, we outline two transformative directions. The first focuses on developing a 6G foundation model as a unified backbone, with task-specific knowledge distilled into compact models suited for diverse edge deployments. The second advances multi-agent systems designed to autonomously diagnose, maintain, and recover networks with minimal human intervention. These directions chart a roadmap for 6G to evolve into an intelligent, self-sustaining communication infrastructure.",
      "short_abstract": "The proliferation of emerging applications, such as autonomous driving and immersive experiences, demands cellular networks that are not only faster, but fundamentally more resilient and autonomous. This paper presents a BlueSky vision on how Artificial...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 11,
        "evidence": [
          "abstract: cooperative systems",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21395",
      "pdf_url": "https://arxiv.org/pdf/2605.21395.pdf"
    },
    {
      "id": "2605.21372",
      "anchor": "paper-2605-21372",
      "title": "Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training",
      "authors": [
        "Hongzhi Ruan",
        "Pei Liu",
        "Weiliang Ma",
        "Zhengning Li",
        "Xueyang Zhang",
        "Jun Ma",
        "Dan Xu",
        "Kun Zhan"
      ],
      "abstract": "Data scaling is fundamental to modern deep learning, and grows increasingly critical as autonomous driving shifts to end-to-end learning. Real-world driving data is expensive to annotate and scene-biased, making real-synthetic co-training with near-infinite synthetic data a promising direction. However, naively incorporating all available synthetic data is inefficient and leads to distribution shifts, and optimizing data mixture under practical training budgets remains a critical yet under-explored problem. In this sense, we claim that the mixture of training data requires clear guidance in terms of scene types and quantities. Particularly in this work, we conceptualize the data mixture approximately as a dynamic optimization process that iteratively adjusts the training data mixture to maximize model performance, guided by closed-loop evaluation feedback, and propose AutoScale, a fully automated closed-loop data engine unifying scene representation, data mixture optimization and retrieval, as well as model training and evaluation. Specifically, we propose Graph Regularized AutoEncoder (Graph-RAE) for driving scene representations, introduce Cluster-aware Gradient Ascent (Cluster-GA) for cluster-wise importance estimation and reweighting, and perform cluster-guided vector retrieval to select high-value samples. Experiments on NavSim demonstrate that AutoScale outperforms vanilla co-training and cross-domain baselines, achieving better performance with fewer synthetic samples under constrained budgets.",
      "short_abstract": "Data scaling is fundamental to modern deep learning, and grows increasingly critical as autonomous driving shifts to end-to-end learning. Real-world driving data is expensive to annotate and scene-biased, making real-synthetic co-training with near-infinite...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21372",
      "pdf_url": "https://arxiv.org/pdf/2605.21372.pdf"
    },
    {
      "id": "2605.21309",
      "anchor": "paper-2605-21309",
      "title": "Hyper-V2X: Hypernetworks for Estimating Epistemic and Aleatoric Uncertainty in Cooperative Bird's-Eye-View Semantic Segmentation",
      "authors": [
        "Abhishek Dinkar Jagtap",
        "Sanath Tiptur Sadashivaiah",
        "Andreas Festag"
      ],
      "abstract": "Cooperative perception enabled by Vehicle-to-Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental representation through shared sensory data. While recent works have advanced multi-agent fusion for improved perception, uncertainty quantification in such cooperative frameworks remains largely unexplored. This paper introduces Hyper-V2X, a hypernetwork-based framework for estimating both epistemic and aleatoric uncertainties in V2X-based perception. Specifically, we propose a partial weight generation scheme and V2X context embedding module that conditions a Bayesian hypernetwork on fused multi-agent features to generate weight distributions for stochastic Bird's-Eye-View (BEV) segmentation. Unlike existing deterministic BEV models, Hyper-V2X enables efficient uncertainty estimation with little computation overhead. Our approach is architecture-agnostic, and can be seamlessly integrating with modern cooperative backbones such as CoBEVT. Experiments on the OPV2V benchmark demonstrate that Hyper-V2X provides accurate, well-calibrated uncertainty estimates and improves overall perception reliability. Our code and benchmark are publicly available under an open-source license: https://github.com/abhishekjagtap1/Hyper-V2X",
      "short_abstract": "Cooperative perception enabled by Vehicle-to-Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental representation through shared sensory data. While recent works have advanced multi-agent fusion for improved...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 38,
        "margin": 21,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21309",
      "pdf_url": "https://arxiv.org/pdf/2605.21309.pdf"
    },
    {
      "id": "2605.21168",
      "anchor": "paper-2605-21168",
      "title": "ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving",
      "authors": [
        "Qiyu Ruan",
        "Yuxuan Wang",
        "He Li",
        "Zhenning Li",
        "Cheng-zhong Xu"
      ],
      "abstract": "Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress testing indispensable. Most scenario generation methods treat surrounding agents as adversaries, but they either (i) induce failures without explicitly modeling vehicle-road physical limits, yielding visually extreme yet physically unsolvable crashes, or (ii) enforce physical feasibility or policy feasibility in isolation, which can over-focus on aggressive maneuvers or remain tied to a controller-dependent capability boundary. We propose ScenePilot, a feasibility-guided, boundary-driven framework that targets the boundary band: scenarios that are physically solvable in principle yet still cause the deployed autonomy stack to fail. We formulate generation as constrained multi-objective reinforcement learning, combining an RSS-derived physical-feasibility score $σ$ with an online-learned AV-risk predictor $Φ$, and introduce step-level feasibility-aware shielding to keep exploration near the feasibility boundary while avoiding infeasible artifacts. Experiments on SafeBench with multiple planners show that ScenePilot yields substantially higher collision rates (+6.2 percentage points) while preserving physical validity, and that adversarial fine-tuning on these boundary-band scenarios consistently reduces downstream crash rates. The code is available at https://github.com/QiyuRuan/ScenePilot.",
      "short_abstract": "Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress testing indispensable. Most scenario generation methods treat surrounding agents as adversaries, but they...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 9,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21168",
      "pdf_url": "https://arxiv.org/pdf/2605.21168.pdf"
    },
    {
      "id": "2605.21139",
      "anchor": "paper-2605-21139",
      "title": "Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving",
      "authors": [
        "Yang Wu",
        "Qiang Meng",
        "Zhaojiang Liu",
        "Youquan Liu",
        "Jian Yang",
        "Jin Xie"
      ],
      "abstract": "Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement learning offers a path to smarter autonomy, it demands two missing pieces of infrastructure: (1) a cognitive foundation that understands traffic semantics and driving intent, and (2) a foresighted physical environment that can anticipate the consequences of candidate actions. To this end, we propose CoPhy, a CognitivePhysical reinforcement learning framework for autonomous driving. To distill to think, we distill VLM knowledge into the BEV encoder and then discard the VLM entirely, retaining cognitive ability at zero inference cost while releasing the cognitive channel as a pluggable interface for optional human language commands. To foresee to act, we build an auto-regressive BEV world model that explicitly predicts future semantic maps conditioned on candidate actions, serving as an interpretable physical sandbox from which safety metrics are directly derived. Built upon this dual infrastructure, we optimize the driving policy via GRPO with a novel dual-reward mechanism: a physical reward derived from BEV rollouts enforces hard safety constraints, while a cognitive reward from a language-aligned scorer ensures intent compliance. Extensive experiments demonstrate that CoPhy not only achieves state-of-the-art results on NAVSIM v1 and v2 benchmarks, but also enables safer driving via cognitively informed scene compliance and flexible intent control through user-defined language instructions.",
      "short_abstract": "Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement learning offers a path to smarter autonomy, it demands two missing pieces of infrastructure: (1) a...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model",
        "Reinforcement Learning",
        "Imitation Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end",
          "abstract: driving policy"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21139",
      "pdf_url": "https://arxiv.org/pdf/2605.21139.pdf"
    },
    {
      "id": "2605.21112",
      "anchor": "paper-2605-21112",
      "title": "RCGDet3D: Rethinking 4D Radar-Camera Fusion-based 3D Object Detection with Enhanced Radar Feature Encoding",
      "authors": [
        "Weiyi Xiong",
        "Bing Zhu"
      ],
      "abstract": "4D automotive radar is indispensable for autonomous driving due to its low cost and robustness, yet its point cloud sparsity challenges 3D object detection. Existing 4D radar-camera fusion methods focus on complex fusion strategies, trading inference speed for marginal gains. This trade-off hinders real-time deployment due to heavy computation on dense feature maps. In contrast, feature extraction from sparse radar points is less time-consuming but remains under-explored. This work uncovers that simply enhancing radar feature extraction can achieve comparable or even higher performance than elaborate fusion modules, while maintaining real-time performance. Based on this finding, we propose RCGDet3D, which centers on radar feature encoding and simplifies multi-modal fusion. Its encoder inherits from the efficient Gaussian Splatting-based Point Gaussian Encoder (PGE) in RadarGaussianDet3D with two key improvements. First, the Ray-centric PGE (R-PGE) predicts Gaussian attributes in ray-aligned coordinate systems before unifying them to Bird's-Eye View (BEV) space, significantly improving geometric consistency and reducing learning difficulty by decoupling the coordinate transformation from representation learning. Second, a Semantic Injection (SI) module incorporates visual cues from images, producing more geometrically accurate and semantically enriched radar features. Experiments on View-of-Delft (VoD) and TJ4DRadSet show that RCGDet3D outperforms state-of-the-art methods in both accuracy and speed, setting a new benchmark for real-time deployment.",
      "short_abstract": "4D automotive radar is indispensable for autonomous driving due to its low cost and robustness, yet its point cloud sparsity challenges 3D object detection. Existing 4D radar-camera fusion methods focus on complex fusion strategies, trading inference speed...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 38,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: point cloud",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21112",
      "pdf_url": "https://arxiv.org/pdf/2605.21112.pdf"
    },
    {
      "id": "2605.21032",
      "anchor": "paper-2605-21032",
      "title": "Towards Physically Consistent 4D Scene Reconstruction for Closed-loop Autonomous Driving Simulation",
      "authors": [
        "Bowyn Tan",
        "Yutong Xie",
        "Bai Huang",
        "Fan Luo",
        "Xiao Li",
        "Naizheng Wang",
        "Yang Guan",
        "Shengbo Eben Li"
      ],
      "abstract": "High-fidelity street scene reconstruction is pivotal for end-to-end autonomous driving simulation, where novel-view synthesis (NVS) and time-varying information modeling are two fundamental capabilities to facilitate closed-loop training. However, existing 3DGS methods and their 4D extensions fail to simultaneously achieve both. To bridge this gap, we establish an information-geometric diagnostic framework, revealing that this limitation stems from a credit assignment dilemma between spatial and temporal parameters. Specifically, the deterministic coupling between viewpoint and time in single-source observation creates a low-rank structure that induces massive null-space ambiguity between static view-dependent and dynamic time-varying components. Temporal information overshadows spatial cues, causing the estimation variance of spatial parameters to diverge. To address this issue, we propose Orthogonal Projected Gradient (OPG), a hierarchical training method designed to restore spatial identifiability. OPG prioritizes the integrity of spatial representations by securing them in an initial stage, then restricts temporal updates to the spatial null space, enabling proactive credit assignment. While OPG isolates temporal updates algebraically, Temporal Regularization Strategy is proposed to further refine the temporal solution space by imposing a smoothness constraint based on the physical prior of consistent appearance evolution, ensuring that the reconstructed scene remains physically consistent in closed-loop simulation. Extensive experiments demonstrate that our method not only maintains stable NVS capabilities but also demonstrates superior performance in traditional observation-reproducing metrics, which indirectly reflect the capability of modeling temporal dynamics.",
      "short_abstract": "High-fidelity street scene reconstruction is pivotal for end-to-end autonomous driving simulation, where novel-view synthesis (NVS) and time-varying information modeling are two fundamental capabilities to facilitate closed-loop training. However, existing...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 9,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21032",
      "pdf_url": "https://arxiv.org/pdf/2605.21032.pdf"
    },
    {
      "id": "2605.21007",
      "anchor": "paper-2605-21007",
      "title": "LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation",
      "authors": [
        "Daojie Peng",
        "Bingtao Wang",
        "Fulong Ma",
        "Liang Zhang",
        "Jun Ma"
      ],
      "abstract": "Road segmentation is a fundamental perception task for autonomous driving and intelligent robotic systems, requiring both high accuracy and real-time inference, especially for deployment on resource-constrained edge devices. Existing multi-modal road segmentation methods often rely on heavy transformer-based encoders to achieve state-of-the-art performance, but their enormous computational cost prohibits real-time deployment on embedded platforms. To address this dilemma, we propose LiteViLNet, a lightweight multi-modal network that fuses RGB texture information and LiDAR geometric information for efficient road segmentation. Specifically, we design a dual-stream lightweight encoder and depth-wise separable convolutions to extract hierarchical features from both modalities with minimal parameters. We further propose a Multi-Scale Feature Fusion Module (MSFM) to facilitate cross-modal interaction at different levels, and a large-kernel-bridge module to capture long-range dependencies with linear complexity. Extensive experiments on the KITTI Road dataset and real-world applications demonstrate that LiteViLNet achieves a promising balance between accuracy and efficiency. Notably, with only 14.04M parameters, our model attains a 96.36% MaxF score, ranking the best among all CNN-based methods and being comparable to larger transformer-based models, and runs at 163.79 FPS in model-only inference on RTX 4060 Ti (22.18 FPS on Jetson Orin NX). It outperforms numerous heavy-weight methods in inference speed while maintaining highly competitive accuracy, fully validating the potential of LiteViLNet for real-time embedded deployment in autonomous driving and intelligent robotics.",
      "short_abstract": "Road segmentation is a fundamental perception task for autonomous driving and intelligent robotic systems, requiring both high accuracy and real-time inference, especially for deployment on resource-constrained edge devices. Existing multi-modal road...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 4,
        "evidence": [
          "abstract: perception",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21007",
      "pdf_url": "https://arxiv.org/pdf/2605.21007.pdf"
    },
    {
      "id": "2605.20942",
      "anchor": "paper-2605-20942",
      "title": "Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding",
      "authors": [
        "Lena Wild",
        "Katie Z Luo",
        "Marco Pavone"
      ],
      "abstract": "Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer promising semantic flexibility, they lack the geometric and relational grounding required for precise road reasoning. Conversely, traditional modular systems, e.g., HD maps and topological road graphs, provide structural precision but remain semantically rigid. To bridge this gap, we introduce the Combined Road Substrate (CRS), a graph-grounded framework that makes geometric road structure and open-vocabulary semantics jointly executable in a single representation. CRS enables the automatic generation of compositionally complex and linguistically varied question-answer pairs via recursive graph queries, augmented with a \"grounding for free\" mechanism that ensures logical traceability to specific map elements, and procedurally extracted chain-of-thought supervision traces. We demonstrate that state-of-the-art VLMs - including large, closed-source models - struggle significantly with structured road reasoning, yet training a small 2- or 4-billion-parameter model with as few as 20 to 80 CRS-enriched scenes yields stable gains in compositional reasoning tasks of varying depth. Analysis of model behavior via verifiable reasoning traces reveals a systematic shift in failure modes: whereas baseline models fail at relational scene understanding, CRS-trained models reduce failures to attribute recognition, suggesting that the primary bottleneck in road understanding is not model scale, but the absence of structured supervision.",
      "short_abstract": "Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer promising semantic flexibility, they lack the geometric and relational grounding...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Mapping & Localization: abstract: HD map",
          "Perception & Sensor Fusion: abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.20942",
      "pdf_url": "https://arxiv.org/pdf/2605.20942.pdf"
    },
    {
      "id": "2605.21820",
      "anchor": "paper-2605-21820",
      "title": "Beyond Scalar Objectives: Expert-Feedback-Driven Autonomous Experimentation for Scientific Discovery at the Nanoscale",
      "authors": [
        "Ralph Bulanadi",
        "Jefferey Baxter",
        "Arpan Biswas",
        "Hiroshi Funakubo",
        "Dennis Meier",
        "Jan Schultheiß",
        "Rama Vasudevan",
        "Yongtao Liu"
      ],
      "abstract": "Self-driving laboratories or autonomous experimentation are emerging as transformative platforms for accelerating scientific discovery. Bayesian optimization (BO) is among the most widely used machine learning frameworks for these purposes, but these BO-based frameworks rely on predefined scalar descriptors to guide experimentation. In many situations, the determination of an appropriate scalar descriptor can be challenging, and may fail to capture subtle yet scientifically important phenomena apparent to experts with interdisciplinary insight. To overcome this limitation, here we develop deep-kernel pairwise learning (DKPL), an approach for autonomous microscopy experiments which incorporates human expertise and interdisciplinary scientific knowledge into an active learning loop. Instead of relying on explicit scalar objectives, DKPL enables experts to directly evaluate which experimental output is more promising using interdisciplinary knowledge. DKPL then learns a latent utility function from these expert judgements to guide subsequent autonomous microscopy experiments. We demonstrate DKPL's performance in learning physically meaningful nanoscale structures while effectively prioritizing high-information measurement regions using an experimental model dataset with known ground truth. We further apply DKPL to analyze the character of ferroelectric domain walls, where we find DKPL capable of distinguishing between high and low characteristic domain-wall angles in bismuth ferrite, and able to discover both head-to-head and tail-to-tail domain-wall character in erbium manganite. This development establishes an approach to integrate expert knowledge into autonomous microscopy experiments and demonstrates a pathway toward expert-guided self-driving laboratories capable of addressing scientific problems beyond the limits of scalar-metrics-driven learning.",
      "short_abstract": "Self-driving laboratories or autonomous experimentation are emerging as transformative platforms for accelerating scientific discovery. Bayesian optimization (BO) is among the most widely used machine learning frameworks for these purposes, but these BO-based...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21820",
      "pdf_url": "https://arxiv.org/pdf/2605.21820.pdf"
    },
    {
      "id": "2605.21747",
      "anchor": "paper-2605-21747",
      "title": "Improving 3D Labeling in Self-Driving by Inferring Vehicle Information using Vision Language Models",
      "authors": [
        "Steven Chen",
        "Shivesh Khaitan",
        "Nemanja Djuric"
      ],
      "abstract": "We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicle Make and Model Recognition (VMMR) methods. The proposed approach utilizes a Vision Language Model (VLM) to both infer a vehicle's make, model, and generation from image crops, and output accurate 3D bounding box dimensions to seed manual labeling. We evaluate the impact of iterative prompt engineering and the choice of different VLMs on both vehicle bounding box inference and make/model/generation recognition. When compared to strong baselines, the proposed approach not only shows high accuracy, but also excels in mitigating specific failure modes where VLMs provide better dimensions than initial lidar-aided human annotated labels (e.g., in cases of significant vehicle occlusion). Experiments on both public and proprietary data strongly suggest that our conclusions are generalizable across different labelers and datasets. The results demonstrate that integrating VLMs into the labeling process can reduce manual labeling time while increasing label quality.",
      "short_abstract": "We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicle Make and Model Recognition (VMMR) methods. The proposed approach utilizes a Vision Language Model (VLM)...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Safety, Security & Verification: abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21747",
      "pdf_url": "https://arxiv.org/pdf/2605.21747.pdf"
    },
    {
      "id": "2606.00069",
      "anchor": "paper-2606-00069",
      "title": "Invascal: Inverse-Vacuity Self-Calibration for Uncertainty-Aware LiDAR Range-View Semantic Segmentation",
      "authors": [
        "Kerim Turacan",
        "Hannes Reichert",
        "Andrei Bolandut",
        "Konrad Doll"
      ],
      "abstract": "LiDAR semantic segmentation is a core perception capability for autonomous vehicles and mobile robots. However, safe operation also depends on knowing when predictions are unreliable. Existing approaches typically rely on softmax confidence, which is often miscalibrated and overconfident, while stronger uncertainty estimates from Monte Carlo dropout or ensembles are often computationally expensive for real-time use. To this end, we introduce a novel, architecture-agnostic uncertainty-aware Adapter Head. It decomposes the prediction into a Preference Head for class ranking and a Strength Head that refines uncertainty assessment, thereby enabling a principled construction of evidential Dirichlet representations. Building on this design, we propose our inverse-vacuity self-calibration objective (Invascal), which directly supervises the strength signal to produce reliable and well-calibrated uncertainty estimates while preventing runaway evidence growth. We evaluate our framework across multiple LiDAR datasets and backbone architectures. We compare against deterministic training, Monte Carlo dropout and ensembles, and prior evidential methods. Our approach consistently improves uncertainty calibration over traditional deterministic methods with minimal computational overhead. At the same time, it preserves competitive segmentation accuracy, where prior evidential methods often suffer performance degradation.",
      "short_abstract": "LiDAR semantic segmentation is a core perception capability for autonomous vehicles and mobile robots. However, safe operation also depends on knowing when predictions are unreliable. Existing approaches typically rely on softmax confidence, which is often...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 23,
        "evidence": [
          "abstract: perception",
          "title: scene segmentation",
          "abstract: scene segmentation",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2606.00069",
      "pdf_url": "https://arxiv.org/pdf/2606.00069.pdf"
    },
    {
      "id": "2605.21676",
      "anchor": "paper-2605-21676",
      "title": "SENTIL: A Runtime Verification Tool for Probabilistic Temporal Logic",
      "authors": [
        "Paapa Kwesi Quansah",
        "Ernest Bonnah"
      ],
      "abstract": "Stochastic cyber-physical systems (CPS) permeate critical infrastructure, from autonomous vehicles to medical devices. Yet, tools for runtime verification of such systems capturing the probabilistic dynamics in stochastic systems remain generally absent despite theoretical foundations established nearly a decade ago. In this paper, we present SENTIL, a novel runtime verification tool with provable statistical guarantees for the runtime monitoring of requirements expressed as Probabilistic Signal Temporal Logic (PrSTL). SENTIL combines an efficient Rust core with universal ecosystem integration, delivering performance exceeding existing deterministic monitors while providing rigorous probabilistic guarantees through statistical model checking, sequential probability ratio testing, and adaptive rare event estimation. SENTIL employs streaming algorithms for incremental robustness computation, parallel Monte Carlo sampling, and a language-agnostic C-ABI enabling seamless deployment across ROS, Apollo, MATLAB Simulink, and AUTOSAR platforms, and direct integration in C, C++, Python, and Java. To validate the effectiveness of the proposed tool, we validate SENTIL across various scenarios spanning autonomous vehicle monitoring, medical device validation, and biological networks, demonstrating 10-1,000$\\times$ performance improvements over existing tools while maintaining provable confidence intervals. SENTIL is open source (\\href{https://github.com/sedislab/SENTIL}{\\texttt{sedislab/SENTIL}}) and it positions probabilistic runtime verification as a deployable infrastructure for all real-world safety-critical stochastic systems.",
      "short_abstract": "Stochastic cyber-physical systems (CPS) permeate critical infrastructure, from autonomous vehicles to medical devices. Yet, tools for runtime verification of such systems capturing the probabilistic dynamics in stochastic systems remain generally absent...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 29,
        "margin": 24,
        "evidence": [
          "abstract: safety",
          "title: verification",
          "abstract: verification",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21676",
      "pdf_url": "https://arxiv.org/pdf/2605.21676.pdf"
    },
    {
      "id": "2605.21145",
      "anchor": "paper-2605-21145",
      "title": "Cloud-Native Operation of Roadside Infrastructure Enabling Demand-Driven Collective Perception via V2X",
      "authors": [
        "Lukas Zanger",
        "Fabian Thomsen",
        "Guido Linden",
        "Jean-Pierre Busch",
        "Lennart Reiher",
        "Lutz Eckstein"
      ],
      "abstract": "Intelligent roadside infrastructure is a key enabler for cooperative intelligent transport systems (C-ITS), supporting vehicles equipped with automated driving systems (ADS), e.g., through enhanced environment perception. With a growing number and an expanding functional scope of roadside units, scalable and efficient operation becomes a challenge. This paper presents a cloud-native architecture for the operation of distributed roadside infrastructure based on a Kubernetes cluster spanning roadside units and a cloud server. Building on this architecture, a demand-driven orchestration approach is implemented to dynamically deploy resource-intensive services only when required. As a representative use case, a V2X-based collective perception application is deployed on-demand when a connected vehicle is nearby. The approach is validated in a real-world experiment in our test field in Aachen, demonstrating that the collective perception application starts in time for the vehicle to benefit from it. Without any demand, the application remains inactive, reducing energy consumption, channel congestion, and hardware wear. Beyond the primary evaluation, V2X recordings from the test field are analyzed to estimate the energy-saving potential of demand-driven operation. In summary, the results demonstrate the practical feasibility of cloud-native, demand-driven operation of roadside infrastructure and indicate its potential to improve scalability and (energy) efficiency in future C-ITS deployments.",
      "short_abstract": "Intelligent roadside infrastructure is a key enabler for cooperative intelligent transport systems (C-ITS), supporting vehicles equipped with automated driving systems (ADS), e.g., through enhanced environment perception. With a growing number and an...",
      "published": "2026-05-20",
      "updated": "2026-05-20",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 46,
        "margin": 34,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: hardware",
          "title: roadside infrastructure",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "archive",
      "age_days": 91,
      "arxiv_url": "https://arxiv.org/abs/2605.21145",
      "pdf_url": "https://arxiv.org/pdf/2605.21145.pdf"
    },
    {
      "id": "2605.24004",
      "anchor": "paper-2605-24004",
      "title": "Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving",
      "authors": [
        "Zhengqi Sun",
        "Yiwen Sun",
        "Boxuan Liu",
        "Tailai Chen",
        "Tianxu Guo",
        "Jiabin Liu"
      ],
      "abstract": "Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic traffic. Existing methods either perform online language reasoning without explicit dynamics verification or use world models mainly in offline pipelines, leaving a gap between semantic intent and physical feasibility at decision time. We propose Reason--Imagine--Act (RIA), a closed-loop framework that couples an LLM reasoner with an action-conditioned world model for online safety verification. At each step, the LLM proposes an action template and candidate sub-actions, the world model performs short-horizon rollouts, and a safety scorer selects the safest executable action with feedback to the next reasoning step. Under a unified CARLA point-goal protocol (1000 episodes), RIA achieves 80.05% route completion, 51.10% arrival rate, and 0.20% collision rate. Under the same closed-loop interface, RIA consistently outperforms training-free baselines, including CARLA TM and MADA, on core closed-loop metrics. For reproducibility, code is available at https://github.com/pku-smart-city/source_code/tree/main/RIA.",
      "short_abstract": "Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic traffic. Existing methods either perform online language reasoning without explicit dynamics verification...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.24004",
      "pdf_url": "https://arxiv.org/pdf/2605.24004.pdf"
    },
    {
      "id": "2605.20666",
      "anchor": "paper-2605-20666",
      "title": "A Semantic and Occlusion-Aware GM-PHD Filter",
      "authors": [
        "Jovan Menezes",
        "Mark Campbell"
      ],
      "abstract": "This paper proposes a new birth model including semantic information derived from deep learning to create an occlusion-aware Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter. Unlike prior approaches that rely on simplistic or uniform assumptions, the proposed Semantic-Occlusion Aware (S-OA) birth model defines initialization terms by explicitly considering regions of occlusion and by leveraging semantic information about the environment. This enables the filter to accurately represent where new objects are more likely to appear, thereby improving tracking performance in complex and high-density driving scenarios. The method is evaluated through Monte Carlo simulations and experiments on the KITTI dataset. Performance is assessed by measuring the latency between first detection and track initiation, along with the mean absolute cardinality error and the Optimal Subpattern Assignment (OSPA) metric. Results demonstrate that the S-OA birth model reduces initialization delay in occlusion-heavy settings, matching or outperforming the strongest baseline in approximately 70% of cases. A sensitivity analysis of birth model weights is also provided. Overall, the findings underscore the benefits of integrating occlusion reasoning and semantic priors into Bayesian tracking frameworks for autonomous driving.",
      "short_abstract": "This paper proposes a new birth model including semantic information derived from deep learning to create an occlusion-aware Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter. Unlike prior approaches that rely on simplistic or uniform...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.20666",
      "pdf_url": "https://arxiv.org/pdf/2605.20666.pdf"
    },
    {
      "id": "2605.20390",
      "anchor": "paper-2605-20390",
      "title": "STELLAR: Scaling 3D Perception Large Models for Autonomous Driving",
      "authors": [
        "Yingwei Li",
        "Xin Huang",
        "Yang Liu",
        "Yang Fu",
        "Alex Zihao Zhu",
        "Chen Song",
        "Junwen Yao",
        "Anant Subramanian",
        "Hao Xiang",
        "Weijing Shi",
        "Yuliang Zou",
        "Tom Hoddes",
        "Zhaoqi Leng",
        "Govind Thattai",
        "Dragomir Anguelov",
        "Mingxing Tan"
      ],
      "abstract": "Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing heterogeneous sensor data and the need for sophisticated 3D spatial understanding. To bridge this gap, we present a comprehensive study on systematically analyzing the impact of scale on these systems. We develop our STELLAR model based on Sparse Window Transformer, by extending the input modalities to include LiDAR, radar, camera, and map prior. We train the model on a large-scale dataset of 50 million driving examples with up to 500 million parameters. Our large-scale experiments reveal empirical scaling trends that connect model performance to model size, data, and compute. The resulting model establishes a new state-of-the-art on the Waymo Open Dataset challenge, outperforming prior arts by a large margin. Our work demonstrates that large-scale training is a highly promising path for advancing the capabilities of perception models for autonomous driving.",
      "short_abstract": "Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 20,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: radar",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.20390",
      "pdf_url": "https://arxiv.org/pdf/2605.20390.pdf"
    },
    {
      "id": "2605.20301",
      "anchor": "paper-2605-20301",
      "title": "Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection",
      "authors": [
        "Wenxuan Li",
        "Qin Zou",
        "Shoubing Chen",
        "Chi Chen",
        "Yingyi Yang",
        "Qingxiang Meng"
      ],
      "abstract": "In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion often induce cross-frame spatiotemporal inconsistencies in BEV-based detectors, leading to temporal BEV feature misalignment and degraded spatiotemporal consistency. To address these challenges, we propose Co-Fusion4D, a unified framework that explicitly preserves cross-frame spatiotemporal consistency and suppresses temporal feature drift. Co-Fusion4D adopts a current-frame-centric strategy, treating the current frame as the primary source of information while selectively incorporating historical frames after spatiotemporal filtering and alignment. This dominant-complementary mechanism effectively mitigates cumulative alignment errors, suppresses noisy feature propagation, and exploits reliable temporal cues for a more consistent BEV representation. In addition, Co-Fusion4D integrates a Dual Attention Fusion (DAF) module to further enhance spatiotemporal feature interaction. DAF jointly leverages intra-frame spatial attention and inter-frame temporal attention to adaptively align and fuse multi-frame features, emphasizing motion-consistent regions while suppressing spurious correlations. By departing from conventional uniform fusion paradigms, this design substantially improves the temporal stability and discriminative capability of BEV representations. Extensive experiments on the nuScenes benchmark demonstrate that Co-Fusion4D achieves state-of-the-art performance, with 74.9% mAP and 75.6% NDS, without relying on test-time augmentation or external data.",
      "short_abstract": "In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion often induce cross-frame spatiotemporal inconsistencies in BEV-based detectors, leading to temporal BEV feature...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 12,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.20301",
      "pdf_url": "https://arxiv.org/pdf/2605.20301.pdf"
    },
    {
      "id": "2605.20082",
      "anchor": "paper-2605-20082",
      "title": "VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving",
      "authors": [
        "Zhefan Xu",
        "Ghassen Jerfel",
        "Marina Haliem",
        "Qi Zhao",
        "Jeonhyung Kang",
        "Khaled S. Refaat"
      ],
      "abstract": "The rapid growth of autonomous driving datasets has enabled the scaling of powerful motion forecasting models. While large-scale pretraining provides strong performance, the standard imitation objective may not fully capture the complex nuances of human driving preferences. Meanwhile, recent advances in vision-language models (VLMs) have demonstrated impressive reasoning and commonsense understanding. Building on these capabilities, this paper presents VL-DPO, a vision-language-guided framework that aligns ego-vehicle motion forecasting models with human preferences. Our approach leverages a VLM as a zero-shot reasoner to automatically generate preference pairs from a pretrained model's rollouts, which are then used to finetune the model via Direct Preference Optimization (DPO). We finetune our models on the Waymo Open End-to-End Driving Dataset (WOD-E2E) and evaluate performance against held-out human preference annotations using rater feedback score (RFS) and average displacement error (ADE). Our experiments confirm that the VLM's trajectory selection is a high-quality proxy for human preference. Our final model, VL-DPO, yields an 11.94% increase in RFS and a 10.01% reduction in ADE over the pretrained model.",
      "short_abstract": "The rapid growth of autonomous driving datasets has enabled the scaling of powerful motion forecasting models. While large-scale pretraining provides strong performance, the standard imitation objective may not fully capture the complex nuances of human...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.20082",
      "pdf_url": "https://arxiv.org/pdf/2605.20082.pdf"
    },
    {
      "id": "2605.19881",
      "anchor": "paper-2605-19881",
      "title": "Trajectory Planning and Control near the Limits: an Open Experimental Benchmark on the RoboRacer Platform",
      "authors": [
        "Mattia Piccinini",
        "Patrick Zambiasi",
        "Aniello Mungiello",
        "Mattia Piazza",
        "Felix Jahncke",
        "Johannnes Betz"
      ],
      "abstract": "We present a modular framework to benchmark new and existing methods for trajectory planning and control in high-acceleration maneuvers that push autonomous driving to the limits. Our framework includes time-optimal raceline generation, online time-optimal velocity replanning, geometric path tracking controllers, and a new model-structured neural network (MS-NN) to learn the inverse dynamics for steering control. We deploy our framework on a 1:10-scale RoboRacer platform, using two circuits. Through several ablations with cautious and aggressive racelines, we study the performance of single modules and their combinations. We show that our MS-NN significantly improves tracking accuracy, decreases steering oscillations, and is physically interpretable. Moreover, online velocity replanning improves lap times by compensating for execution errors, and enables the vehicle to safely reach higher speeds and accelerations. To support future research, our code, datasets, videos and results are publicly available at https://roboracer-benchmark.github.io/planning_control_benchmark/.",
      "short_abstract": "We present a modular framework to benchmark new and existing methods for trajectory planning and control in high-acceleration maneuvers that push autonomous driving to the limits. Our framework includes time-optimal raceline generation, online time-optimal...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Benchmark",
        "Cooperative / V2X",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 15,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19881",
      "pdf_url": "https://arxiv.org/pdf/2605.19881.pdf"
    },
    {
      "id": "2605.19837",
      "anchor": "paper-2605-19837",
      "title": "CADENet: Condition-Adaptive Asynchronous Dual-Stream Enhancement Network for Adverse Weather Perception in Autonomous Driving",
      "authors": [
        "Sherif Khairy",
        "Catherine M. Elias"
      ],
      "abstract": "Adverse weather (rain, fog, sand, and snow) degrades camera-based object detection in autonomous vehicles. Existing enhancement-then-detect approaches stall the safety-critical perception loop, violating hard real-time requirements. Progress on this problem is also constrained by an under-recognized evaluation ceiling: ground truth annotated on degraded images cannot credit a detector that recovers objects the annotators themselves could not see, so a genuinely useful enhancement can register as a near-flat F1 gain. This paper presents CADENet (Condition-Adaptive Asynchronous Dual-stream Enhancement Network), a training-free three-thread system: Thread S (YOLOv11n) delivers detections at full frame rate with zero added latency; Thread Q applies condition-adaptive enhancement (CAPE) and fuses results via entropy-guided NMS (EG-NMS) without blocking Thread S; Thread E provides CLIP zero-shot weather classification, so new weather categories require only a new text prompt, with no labeled data and no retraining. Evaluated on 1327 DAWN images (YOLOv11m, IoU = 0.5, confidence = 0.25), CADENet achieves Recall = 0.0103 (micro), F1 = 0.0230 on snow, and F1 = 0.0038 on rain. We formalize the annotation completeness bias on DAWN-class data, so the reported F1 values are lower bounds on the true gain; recall is the annotation-gap-immune headline metric. Thread S sustains approximately 44 FPS regardless of enhancement load. No model retraining or additional sensor hardware is required.",
      "short_abstract": "Adverse weather (rain, fog, sand, and snow) degrades camera-based object detection in autonomous vehicles. Existing enhancement-then-detect approaches stall the safety-critical perception loop, violating hard real-time requirements. Progress on this problem...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 18,
        "margin": 2,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: object detection",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19837",
      "pdf_url": "https://arxiv.org/pdf/2605.19837.pdf"
    },
    {
      "id": "2605.19771",
      "anchor": "paper-2605-19771",
      "title": "Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives",
      "authors": [
        "Junli Wang",
        "Zhihua Hua",
        "Xueyi Liu",
        "Zebin Xing",
        "Haochen Tian",
        "Kun Ma",
        "Hangjun Ye",
        "Guang Chen",
        "Long Chen",
        "Qichao Zhang"
      ],
      "abstract": "Existing imitation learning methods for end-to-end autonomous driving predominantly learn from successful demonstrations by minimizing geometric deviations from expert trajectories. This paradigm implicitly assumes that spatial proximity implies behavioral safety, leading to a critical objective mismatch: trajectories with nearly identical imitation losses may exhibit drastically different safety outcomes, where one remains recoverable while the other results in collision. To address this limitation, we propose BeyondDrive, a failure-aware imitation learning framework that jointly learns from successful and failed driving behaviors. First, we introduce a flow matching-based negative trajectory generator that synthesizes safety-critical yet expert-proximate trajectories, enabling explicit modeling of safety asymmetry. Second, we develop a diversity-aware sampling strategy that mitigates mode collapse and improves coverage of diverse failure modes during negative trajectory generation. Third, we propose a Repulsive Distance Loss that simultaneously attracts predictions toward expert demonstrations while repelling them from hard negative trajectories, thereby establishing discriminative safety boundaries in trajectory space. Applied to the uni-modal baseline Latent TransFuser, BeyondDrive achieves 89.7 PDMS on the NAVSIMv1 closed-loop benchmark, outperforming prior state-of-the-art methods. Moreover, BeyondDrive generalizes effectively across different autonomous driving architectures, including multi-modal planners, and further demonstrates strong zero-shot transferability on the HUGSIM benchmark.",
      "short_abstract": "Existing imitation learning methods for end-to-end autonomous driving predominantly learn from successful demonstrations by minimizing geometric deviations from expert trajectories. This paradigm implicitly assumes that spatial proximity implies behavioral...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 30,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19771",
      "pdf_url": "https://arxiv.org/pdf/2605.19771.pdf"
    },
    {
      "id": "2605.19744",
      "anchor": "paper-2605-19744",
      "title": "Real-World On-Vehicle Evaluation of Embedding-Based Anomaly Detection",
      "authors": [
        "Albert Schotschneider",
        "Daniel Bogdoll",
        "Svetlana Pavlitska",
        "Ahmed Abouelazm",
        "Johann Marius Zoellner"
      ],
      "abstract": "Detecting anomalies in traffic scenes is crucial for ensuring safety in autonomous driving, yet collecting representative anomalous data remains challenging. Existing anomaly detection methods are highly specialized and rely on normality as defined by the abstract semantic Cityscapes classes, making it difficult to adapt to diverse real-world scenarios. We propose an adaptable real-time anomaly detection method that leverages foundation models in the form of pretrained vision transformer embeddings to detect deviations via nearest-neighbor similarity in the latent semantic feature space. Based on patch-wise processing, the algorithm produces dense anomaly masks, allowing for the localization of detected anomalies. The method robustly models normality through a single reference image. This formulation avoids explicit supervision and dataset-specific training, making it suitable for real-world deployment. We evaluate the method on standard benchmarks and on an automated vehicle in real-world scenarios. Despite its simplicity, the method achieves good performance on the Road Anomaly benchmark and demonstrates consistent qualitative behavior in practice, successfully highlighting semantically unusual objects in diverse scenes. These results suggest that simple, reference-based methods can provide useful anomaly signals under realistic operating conditions.",
      "short_abstract": "Detecting anomalies in traffic scenes is crucial for ensuring safety in autonomous driving, yet collecting representative anomalous data remains challenging. Existing anomaly detection methods are highly specialized and rely on normality as defined by the...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: localization",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19744",
      "pdf_url": "https://arxiv.org/pdf/2605.19744.pdf"
    },
    {
      "id": "2605.19631",
      "anchor": "paper-2605-19631",
      "title": "HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models",
      "authors": [
        "Hoonhee Cho",
        "Giwon Lee",
        "Jae-Young Kang",
        "Hyemin Yang",
        "Heejun Park",
        "Kuk-Jin Yoon"
      ],
      "abstract": "End-to-end autonomous driving has emerged as a compelling alternative to traditional modular pipelines by directly mapping raw sensor data to driving actions. While recent approaches achieve strong performance on single-domain datasets, their performance degrades significantly when trained jointly across multiple heterogeneous domains. In practice, however, autonomous systems must operate across diverse environments with heterogeneous distributions, including different cities, sensor configurations, and traffic patterns, without domain-specific retraining. This gap highlights a key challenge in multi-domain learning: domain-specific variations across heterogeneous domains introduce conflicting learning signals, driving models toward compromised solutions that are suboptimal across domains. To address this, we propose a trajectory-driven learning paradigm that organizes training around planning trajectories, enabling the model to capture domain-invariant representations of driving intent. Furthermore, we incorporate a world model that predicts future latent features conditioned on ego actions, improving feature consistency and mitigating domain-induced biases. We evaluate our approach on three benchmarks, nuScenes, NAVSIM, and the Waymo end-to-end dataset, and show substantial improvements over existing methods across all domains. Our results demonstrate that a single unified model can be trained on heterogeneous datasets while maintaining strong performance within each domain, highlighting a step toward scalable real-world deployment. We will make our code publicly available.",
      "short_abstract": "End-to-end autonomous driving has emerged as a compelling alternative to traditional modular pipelines by directly mapping raw sensor data to driving actions. While recent approaches achieve strong performance on single-domain datasets, their performance...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 20,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19631",
      "pdf_url": "https://arxiv.org/pdf/2605.19631.pdf"
    },
    {
      "id": "2605.19620",
      "anchor": "paper-2605-19620",
      "title": "Bézier Degradation Modeling for LiDAR-based Human Motion Capture",
      "authors": [
        "Xiaoqi An",
        "Lin Zhao",
        "Jun Li",
        "Chen Gong",
        "Jian Yang"
      ],
      "abstract": "LiDAR-based 3D human motion capture has broad applications in fields such as autonomous driving and robotics, where accurate motion reconstruction is crucial. However, existing methods often struggle with unstable inputs and severe occlusions, leading to jittery or even failed pose predictions. To address these challenges, we propose BMLiCap, a coarse-to-fine framework that models motion using temporally compressible Bézier curves. By reducing control points through a trajectory-preserving strategy, we obtain a coherent and learning-friendly motion representation. To reconstruct human actions from LiDAR point-cloud cues, we design a progressive motion-reconstruction module. Specifically, a Time-scale Motion Transformer (TMT) is introduced to predict motion curves at multiple temporal scales, and a Multi-level Motion Aggregator (MMA) is utilized to adaptively fuse the multi-scale curves to recover detailed, temporally coherent poses, effectively bridging observation gaps caused by occlusions and noise. Across four mainstream benchmarks LiDARHuman26M, FreeMotion, NoiseMotion, and SLOPER4D, BMLiCap achieves state-of-the-art accuracy and temporal continuity in complex scenes, demonstrating its ability to compensate for severe occlusions and reduce prediction jitter.",
      "short_abstract": "LiDAR-based 3D human motion capture has broad applications in fields such as autonomous driving and robotics, where accurate motion reconstruction is crucial. However, existing methods often struggle with unstable inputs and severe occlusions, leading to...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 10,
        "evidence": [
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19620",
      "pdf_url": "https://arxiv.org/pdf/2605.19620.pdf"
    },
    {
      "id": "2605.19592",
      "anchor": "paper-2605-19592",
      "title": "Implicit Action Chunking for Smooth Continuous Control",
      "authors": [
        "Bosun Liang",
        "Shuo Pei",
        "Zirui Chen",
        "Chuanzhi Fan",
        "Chen Sun",
        "Yuankai Wu",
        "Huachun Tan",
        "Yong Wang"
      ],
      "abstract": "Reinforcement learning often produces high-frequency oscillatory control signals that undermine the safety and stability required for physical deployment. Explicit action chunking addresses this by predicting fixed-horizon trajectories but scales the policy output dimension proportionally with the horizon length, leading to optimization difficulties and incompatibility with standard step-wise interaction. To overcome these challenges, this paper proposes Dual-Window Smoothing (DWS), an implicit action chunking framework for smooth continuous control. Unlike explicit methods, DWS enforces temporal coherence without expanding the action space. It uses a dual-window design: an execution window that ensures physical smoothness through deterministic modulation, and a value window that aligns temporal-difference targets over the horizon to correct critic bias caused by open-loop execution. DWS also includes a lightweight actor-side temporal regularizer based on first-order action differences to promote global continuity. This design effectively bridges the gap between temporal abstraction and reactive step-wise control. Experiments on benchmarks including the DeepMind Control Suite and industrial energy management tasks show that DWS outperforms state-of-the-art (SOTA) baselines. In complex vision-based autonomous driving tasks, DWS achieves smoother control, safer behavior with reduced jitter, and attains a 100% success rate.",
      "short_abstract": "Reinforcement learning often produces high-frequency oscillatory control signals that undermine the safety and stability required for physical deployment. Explicit action chunking addresses this by predicting fixed-horizon trajectories but scales the policy...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19592",
      "pdf_url": "https://arxiv.org/pdf/2605.19592.pdf"
    },
    {
      "id": "2605.19524",
      "anchor": "paper-2605-19524",
      "title": "SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving",
      "authors": [
        "Kefei Tian",
        "Yuansheng Lian",
        "Kai Yang",
        "Xiangdong Chen",
        "Shen Li"
      ],
      "abstract": "End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models are promising due to their strong reasoning capabilities. However, most VLA-based approaches rely on positive expert demonstrations, rarely exploiting negative samples, leading to insufficient understanding of risky behaviors and safety boundaries. To address this limitation, we propose SafeAlign-VLA, a unified negative-enhanced safe alignment framework that incorporates negative data into supervised learning and reinforcement learning. First, we develop a counterfactual safety pairing paradigm to generate structured safety labels and counterfactual positive trajectories from risky scenarios via counterfactual reasoning. Then, a two-stage training strategy is adopted: negative-enhanced supervised fine-tuning for failure feedback and trajectory correction, followed by anchor-based group relative policy optimization that uses positive and negative trajectories as contrastive anchors to steer sampling and penalize high-risk behaviors via group-relative advantages. Experiments on NAVSIM and DeepAccident validate the proposed framework. SafeAlign-VLA achieves 89.1 PDMS on the NAVSIM v1 testset, improving over the baseline without negative data by 1.3%. On DeepAccident, it reduces the collision rate to 3.36%, while achieving 84.2% language accuracy and 85.8% risk prediction accuracy. These results demonstrate the effectiveness of the proposed negative-enhanced safe alignment framework for safe and robust autonomous driving.",
      "short_abstract": "End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models are promising due to their strong reasoning capabilities. However, most VLA-based approaches rely on...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 13,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19524",
      "pdf_url": "https://arxiv.org/pdf/2605.19524.pdf"
    },
    {
      "id": "2605.22874",
      "anchor": "paper-2605-22874",
      "title": "NeuroNL2LTL: A Neurosymbolic Framework for Natural Language Translation of Linear Temporal Logic",
      "authors": [
        "Paapa Kwesi Quansah",
        "Ernest Bonnah"
      ],
      "abstract": "Effectively translating between natural language (NL) and formal logics like Linear Temporal Logic (LTL) requires expertise that limits formal verification's reach in safety-critical development. Template-based approaches sacrifice expressiveness for reliability; neural methods achieve fluency but provide no correctness guarantees. We present NeuroNL2LTL, a neurosymbolic architecture unifying learned translation with formal verification. NeuroNL2LTL routes translation through an intermediate representation whose mapping to LTL is structure-preserving by construction. Generated specifications undergo satisfiability and non-triviality checking; a minimal-edit repair mechanism corrects near-miss outputs before they reach downstream tools. The central innovation is verifier-in-the-loop training: verification outcomes serve as reward signals for reinforcement learning, producing neural components that optimize directly for formal correctness. On 200,000+ requirements spanning aerospace, robotics, autonomous vehicles, and ten additional domains, NeuroNL2LTL achieves 28\\% semantic equivalence with reference specifications while ensuring 86\\% of outputs are verified satisfiable. The system also generates contextually grounded explanations from LTL, enabling domain experts to validate specifications without specialized training. This work demonstrates that formal verification can function as both training objective and runtime filter for neural specification systems, allowing us to build neural-based tools whose reliability derives from logical guarantees rather than statistical confidence.",
      "short_abstract": "Effectively translating between natural language (NL) and formal logics like Linear Temporal Logic (LTL) requires expertise that limits formal verification's reach in safety-critical development. Template-based approaches sacrifice expressiveness for...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: verification"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.22874",
      "pdf_url": "https://arxiv.org/pdf/2605.22874.pdf"
    },
    {
      "id": "2605.19824",
      "anchor": "paper-2605-19824",
      "title": "From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning",
      "authors": [
        "Ahmed Y. Gado",
        "Omar Y. Goba",
        "Alaa Hassanein",
        "Catherine M. Elias",
        "Ahmed Hussein"
      ],
      "abstract": "Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and Large Multimodal Models (LMMs) continue to treat time as a secondary property. This lack of temporal grounding leads to inconsistencies in reasoning about continuous actions, undermining both safety and interpretability. This work explores whether temporal conditioning within inter-agent communication can preserve or enhance coherence without introducing degradation in semantic or logical consistency. To investigate this, we introduce three planner architectures with progressively increasing temporal integration and evaluate them on curated subsets of the BDD-X dataset using semantic, syntactic, and logical metrics. Results show that while temporal conditioning reshapes reasoning style, it yields no statistically significant improvements in standard NLP-based correctness metrics. However, qualitative analysis reveals predictive hazard reasoning, stable corrective behavior, and strategic divergence in the Sentinel. These findings clarify the limits of prompt-based temporal grounding and establish the first empirical benchmark for temporal scene-to-plan reasoning.",
      "short_abstract": "Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and Large Multimodal Models (LMMs) continue to treat time as a secondary property. This lack of temporal...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19824",
      "pdf_url": "https://arxiv.org/pdf/2605.19824.pdf"
    },
    {
      "id": "2605.19328",
      "anchor": "paper-2605-19328",
      "title": "RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents",
      "authors": [
        "Doguhuan Yeke",
        "Yanming Zhou",
        "Leo Y. Lin",
        "Hongyu Cai",
        "Antonio Bianchi",
        "Z. Berkay Celik"
      ],
      "abstract": "Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical platforms, e.g. robots and autonomous vehicles, to interpret visual scenes and execute natural language commands in diverse environments. Previous research has introduced jailbreak attacks and defenses for embodied AI. Their evaluations, however, rely on ad-hoc datasets, limited metrics, and emphasize attack success while neglecting the trade-off between security and the ability to follow benign commands. Existing benchmarks and evaluation frameworks either target traditional chat-based models or focus on non-adversarial safety evaluation for embodied AI; neither captures the adversarial risks, inputs, consequences, and evaluation criteria necessary for jailbreak attacks in embodied AI systems. In this paper, we address this gap with RoboJailBench, which consists of three core components. We establish a security taxonomy derived from ISO standards, regulatory rules, and documented incidents. This effort yields 18 categories of security violation consequences for embodied AI. We introduce an intent contrast dataset pipeline that augments existing datasets with paired adversarial and benign goals to measure both security and utility. Lastly, we provide an evolving repository with standardized metrics and a unified process for assessing and integrating new attacks and defenses. With this benchmark, we construct a new taxonomy-balanced dataset and augment five existing datasets. We integrate four attacks and two defenses to evaluate their performance on leading embodied VLMs. This benchmark provides the first standardized evaluation framework for jailbreak attacks in embodied AI and supports future research. We release our code, datasets, and artifacts, and maintain a leaderboard at https://purseclab.github.io/benchmark-for-robotics-security.",
      "short_abstract": "Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical platforms, e.g. robots and autonomous vehicles, to interpret visual scenes and execute natural language commands in...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 25,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19328",
      "pdf_url": "https://arxiv.org/pdf/2605.19328.pdf"
    },
    {
      "id": "2605.19655",
      "anchor": "paper-2605-19655",
      "title": "Equalized Coverage in Motion Control Performance Prediction for Self-Adaptive Road Vehicles",
      "authors": [
        "Ole Reuter",
        "Richard Schubert",
        "Marvin Loba",
        "Markus Maurer"
      ],
      "abstract": "Automated driving systems require monitoring mechanisms to ensure operation as intended, especially when system elements degrade and/or fail. Hence, capability monitoring is crucial in order to evaluate the system's remaining performance and implement capability-based behavior. In this paper, we investigate the dynamics of a highly over-actuated automated vehicle under actuator degradations and failures, affecting the vehicle's motion control capabilities. We propose a lightweight prediction model based on conformalized quantile regression that predicts whether an automated vehicle can be controlled with sufficiently low lateral deviation from a planned trajectory under nominal, degraded, and failed actuator conditions. We recognize that statistical guarantees should hold not only across all data (marginal coverage) but also for different regimes within the data (conditional coverage). We therefore employ equalized coverage methods to address this challenge. During runtime behavior generation our predictor can provide a heuristic for determining the admissible action space. Its application and limitations are discussed in this paper.",
      "short_abstract": "Automated driving systems require monitoring mechanisms to ensure operation as intended, especially when system elements degrade and/or fail. Hence, capability monitoring is crucial in order to evaluate the system's remaining performance and implement...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: title: control",
          "Control & Vehicle Dynamics: abstract: control",
          "Prediction & World Models: title: prediction",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19655",
      "pdf_url": "https://arxiv.org/pdf/2605.19655.pdf"
    },
    {
      "id": "2605.19490",
      "anchor": "paper-2605-19490",
      "title": "Closed-Loop Hybrid Digital Twin Platform for Connected and Automated Vehicle Validation",
      "authors": [
        "Kanglong Quan",
        "Zhebing Xia",
        "Linfeng Jiang",
        "Hao Yu",
        "Ziheng Qiao",
        "Dapeng Dong",
        "Dongyao Jia"
      ],
      "abstract": "Comprehensive and efficient validation of connected and automated vehicles (CAVs) is critical prior to real-world deployment. While simulation-based testing offers scalability, existing approaches often lack seamless integration with real vehicles and field data, limiting their fidelity in capturing dynamic, real-world interactions. To bridge this gap, this paper proposes a novel real-time hybrid digital twin platform. Its core innovation lies in the tight coupling of a high-fidelity CARLA-SUMO co-simulation with a physical test site and vehicle via a low-latency Vehicle-to-Everything (V2X) communication link. A custom-developed middleware serves as the critical bridge, synchronizing a real CAV's kinematic state as a shadow vehicle in the simulation and translating virtual control commands into chassis-actuating Controller Area Network (CAN) messages for closed-loop control. Detailed implementation includes using photogrammetry for full-scale asset reconstruction and a cloud-edge collaborative architecture for scalable, multi-user operation. Experimental results demonstrate stable synchronization and effective closed-loop control with low latency, confirming the platform's practicality for multi-scenario CAV verification.",
      "short_abstract": "Comprehensive and efficient validation of connected and automated vehicles (CAVs) is critical prior to real-world deployment. While simulation-based testing offers scalability, existing approaches often lack seamless integration with real vehicles and field...",
      "published": "2026-05-19",
      "updated": "2026-05-19",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 15,
        "evidence": [
          "abstract: V2X",
          "abstract: cooperative systems",
          "title: connected vehicles",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 92,
      "arxiv_url": "https://arxiv.org/abs/2605.19490",
      "pdf_url": "https://arxiv.org/pdf/2605.19490.pdf"
    },
    {
      "id": "2605.20255",
      "anchor": "paper-2605-20255",
      "title": "Multi-Agent Reinforcement Learning for Safe Autonomous Driving Under Pedestrian Behavioral Uncertainty",
      "authors": [
        "Prakash Aryan",
        "Kaushik Raghupathruni",
        "Timo Kehrer",
        "Sebastiano Panichella"
      ],
      "abstract": "Simulation-based testing of self-driving cars (SDCs) typically relies on scripted pedestrian models that do not capture the heterogeneity and uncertainty of real crossing behavior, limiting the realism of safety assessments, especially for jaywalking, which is governed by latent personality traits the vehicle cannot observe. We hypothesize that jointly training pedestrians and the SDC with multi-agent reinforcement learning (MARL) yields more realistic interaction scenarios than training against fixed pedestrian policies, and that the behavior gap between predictable and unpredictable crossings can be measured directly from trajectories. We co-train an SDC and 12 pedestrians using Multi-Agent Proximal Policy Optimization (MAPPO): pedestrian locomotion follows scripted Dijkstra pathfinding while an RL policy controls high-level go/wait decisions, and jaywalking probability depends on a per-pedestrian trait sampled at episode start and hidden from the SDC. In 500-episode evaluations, the co-trained SDC reached 78% of goals with a 14% collision rate, versus 35%/33% for the best rule-based baseline. A speed differential metric shows the SDC traveled 2.65 m/s faster near jaywalkers than near crosswalk users at close range (0-3 m), indicating jaywalking encounters were not anticipated. Jaywalking was 13% of crossing events but 62% of collisions, and co-training reduced collisions by 30% relative to single-agent RL as pedestrians learned to wait when the SDC approached at speed.",
      "short_abstract": "Simulation-based testing of self-driving cars (SDCs) typically relies on scripted pedestrian models that do not capture the heterogeneity and uncertainty of real crossing behavior, limiting the realism of safety assessments, especially for jaywalking, which...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 11,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "title: uncertainty",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.20255",
      "pdf_url": "https://arxiv.org/pdf/2605.20255.pdf"
    },
    {
      "id": "2605.19038",
      "anchor": "paper-2605-19038",
      "title": "Guiding Neuro-Symbolic Scenario Generation with Spatio-Temporal Logic",
      "authors": [
        "Lorenzo Bonin",
        "Francesco Giacomarra",
        "Luca Bortolussi",
        "Jyotirmoy V. Deshmukh",
        "Francesca Cairoli"
      ],
      "abstract": "The rapid advancement of autonomous driving (AD) technologies has outpaced the development of robust safety evaluation methods. Conventional testing relies on exposing AD systems to vast numbers of real-world traffic scenes -- a brute-force approach that is prohibitively expensive and statistically ineffective at capturing the rare, safety-critical edge cases essential for validating real-world robustness. To address this fundamental limitation, we introduce STRELGen, a scalable framework for the targeted generation of safety-critical driving scenarios. STRELGen synergistically combines a multi-agent trajectory-generation diffusion model (DM) with Spatio-Temporal Logic (STREL) specifications that encode complex safety and realism properties through a highly interpretable formalism. Crucially, monitoring satisfaction levels of these specifications is differentiable, enabling gradient-based search. At inference time, we optimize directly over the DM latent space to maximize STREL formula satisfaction. The result is efficient generation of highly plausible yet safety-critical multi-agent scenarios that lie within the learned data distribution. STRELGen thus provides a flexible, interpretable, and powerful tool for stress-testing autonomous driving systems, moving beyond the limitations of brute-force data collection.",
      "short_abstract": "The rapid advancement of autonomous driving (AD) technologies has outpaced the development of robust safety evaluation methods. Conventional testing relies on exposing AD systems to vast numbers of real-world traffic scenes -- a brute-force approach that is...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 9,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.19038",
      "pdf_url": "https://arxiv.org/pdf/2605.19038.pdf"
    },
    {
      "id": "2605.18137",
      "anchor": "paper-2605-18137",
      "title": "Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving",
      "authors": [
        "Lijun Zhou",
        "Hongcheng Luo",
        "Zhenxin Zhu",
        "Cheng Chi",
        "Mingfei Tu",
        "Kaixin Xiong",
        "Lei Gong",
        "Zhanqian Wu",
        "Zehan Zhang",
        "Fangzhen Li",
        "Hao Li",
        "Yingying Shen",
        "Jiale He",
        "Haohui Zhu",
        "Shan Zhao",
        "Kai Wang",
        "Zhiwei Zhan",
        "Yuechuan Pu",
        "Kaiyuan Tan",
        "Ruiling Yang",
        "Xianqi Wang",
        "Tianyi Yan",
        "Jiawei Zhou",
        "Lei Zhang",
        "Jingyang Zhao"
      ],
      "abstract": "This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture driven by sparse scene queries. WorldRec initializes structured queries in 3D space, leveraging them to aggregate cross-view, cross-temporal features, thereby naturally enforcing spatial consistency across frames and yielding compact yet high-fidelity 3D Gaussian scene representations. For world generation, we propose WorldGen, a two-stage training framework of bidirectional pretraining followed by causal fine-tuning through three progressive stages (Teacher Forcing, ODE distillation, and DMD), enabling high-quality online causal video generation in as few as 4 denoising steps. Building on both modules, we further introduce the JWM, which deeply integrates WorldRec and WorldGen to achieve synergistic gains in generation stability, cross-frame consistency, and visual fidelity, providing a solid foundation for closed-loop simulation, data synthesis, and end-to-end training in autonomous driving.",
      "short_abstract": "This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.18137",
      "pdf_url": "https://arxiv.org/pdf/2605.18137.pdf"
    },
    {
      "id": "2605.18096",
      "anchor": "paper-2605-18096",
      "title": "A regularization method for planar offset curves and bi-offset recognition",
      "authors": [
        "Rosanna Campagna",
        "Salvatore Mondrone",
        "Tomas Sauer"
      ],
      "abstract": "Offset curves for planar trajectories are interesting in the generation of tool paths for numerically controlled industrial machines and in trajectory planning methods for autonomous driving systems. Theoretical offset curves may exhibit peculiar singularities, including self-intersections, which limit their use in practical applications. Existing approaches address these issue through geometric filtering techniques to detect and remove undesirable features but the computation of accurate and well-behaved offset curves remains a challenging task. We assume a first stage of functional approximation of trajectories by penalized Hermite spline regression enabling the simultaneous fitting of positions and tangents. The regularization is imposed on the second derivatives, effectively mitigating the jerk effect, which is particularly relevant in motion planning and path smoothing applications. Then, taking into account the geometrical pointwise properties of the resulting curve, we design two offset curves through the simultaneous approximation of function values and derivatives. Then, a mathematical model to obtain the so-called bi-offset as most fitting as with the original generator curve is proposed, also relating the offset range and pointwise curvature values. The adaptive reconstruction of the center line from the external boundaries is a topic of interest and is the main focus of our work. Numerical experiments confirm the reliability of our approach at every stage of the resolution process.",
      "short_abstract": "Offset curves for planar trajectories are interesting in the generation of tool paths for numerically controlled industrial machines and in trajectory planning methods for autonomous driving systems. Theoretical offset curves may exhibit peculiar...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "abstract: motion planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.18096",
      "pdf_url": "https://arxiv.org/pdf/2605.18096.pdf"
    },
    {
      "id": "2605.18074",
      "anchor": "paper-2605-18074",
      "title": "4DLidarOpen: An Open 4D FMCW Lidar Dataset for Motion-Aware Autonomous Driving",
      "authors": [
        "Kane Qian",
        "Xin Zhao",
        "Yining Shi",
        "Rujun Yan",
        "Zhengqing Pan",
        "Kaojin Zhu",
        "Mengmeng Yang",
        "Kai Sun",
        "Diange Yang",
        "Kun Jiang"
      ],
      "abstract": "We present 4DLidarOpen, a large-scale open multi-modal dataset for autonomous driving, centered on 4D frequency-modulated continuous-wave (FMCW) Lidar sensing. Unlike conventional time-of-flight Lidar datasets that mainly provide geometric measurements, 4DLidarOpen includes point-wise radial velocity measurements from a forward-facing 4D FMCW Lidar, together with multiple Lidars of different types, including rotating, solid-state, and blind-spot variants, surround-view cameras, and 6-DOF ego-vehicle poses. The dataset was collected in complex urban environments in Beijing and covers dense pedestrian interactions, congested traffic, high-speed driving, and unprotected maneuvers. 4DLidarOpen provides synchronized multi-sensor data and 3D bounding-box annotations with persistent track IDs across five object categories. A hybrid annotation strategy is adopted, where large-scale auto-labeled data support scalable training and human experts refine annotations for the human-annotated training and validation sets. Based on this dataset, we establish benchmarks for 3D object detection, birds-eye view (BEV) segmentation and flow prediction, and motion forecasting with planning. Extensive experiments show that direct velocity measurements from 4D FMCW Lidar provide complementary motion cues for dynamic-scene understanding. Compared with geometric-only sensing, the velocity-aware representation improves motion-related perception and downstream forecasting and planning, especially in scenarios involving vulnerable road users and fast-moving objects. These results indicate that 4D FMCW Lidar is a promising sensing modality for motion-aware autonomous driving. The dataset and evaluation toolkit are publicly released to support research on 4D scene understanding, multi-Lidar fusion, and velocity-aware perception and planning.",
      "short_abstract": "We present 4DLidarOpen, a large-scale open multi-modal dataset for autonomous driving, centered on 4D frequency-modulated continuous-wave (FMCW) Lidar sensing. Unlike conventional time-of-flight Lidar datasets that mainly provide geometric measurements,...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 16,
        "evidence": [
          "abstract: perception",
          "abstract: object detection",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: scene understanding",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.18074",
      "pdf_url": "https://arxiv.org/pdf/2605.18074.pdf"
    },
    {
      "id": "2605.18059",
      "anchor": "paper-2605-18059",
      "title": "Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations",
      "authors": [
        "Zhiyuan Zhang",
        "Zhenghao Jin",
        "Yanlun Peng",
        "Xianda Guo",
        "Haoran Liu",
        "Shaofeng Zhang",
        "Xingjun Ma",
        "Zuxuan Wu",
        "Junchi Yan",
        "Xiaosong Jia",
        "Yu-Gang Jiang"
      ],
      "abstract": "Robustness is a critical requirement for deploying autonomous driving systems in the real world. Existing robustness benchmarks for autonomous driving have made important progress in studying the effects of image-level corruptions, such as adverse weather or camera degradation, on perception modules and open-loop planning outputs. However, deployment can also involve system-level imperfections, such as inference latency and ego-state estimation errors, which remain less studied in closed-loop E2E-AD evaluation. These imperfections can accumulate through the feedback loop and destabilize control. In this work, we present Bench2Drive-Robust, to our knowledge the first device-centric robustness benchmark for closed-loop end-to-end autonomous driving under realistic deployment perturbations. We systematically evaluate deployment-oriented perturbations arising from three major sources: camera-stream failures (frame drop, partial observation), ego-state estimation errors (GPS noise, and speed or odometry errors), and compute-induced control delay (model inference delay). We evaluate representative end-to-end driving methods and analyze their robustness under different perturbation severities. Our results show that these deployment-related perturbations can substantially degrade closed-loop driving performance, revealing robustness challenges that are not fully captured by conventional image-level corruption evaluations. By establishing a closed-loop evaluation protocol and demonstrating the substantial impact of these deployment-oriented perturbations, Bench2Drive-Robust defines practical robustness problems for end-to-end autonomous driving and encourages further research on deployment-aware robust driving systems.",
      "short_abstract": "Robustness is a critical requirement for deploying autonomous driving systems in the real world. Existing robustness benchmarks for autonomous driving have made important progress in studying the effects of image-level corruptions, such as adverse weather or...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Benchmark",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 4,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.18059",
      "pdf_url": "https://arxiv.org/pdf/2605.18059.pdf"
    },
    {
      "id": "2605.17984",
      "anchor": "paper-2605-17984",
      "title": "See Silhouettes in Motion with Neuromorphic Vision",
      "authors": [
        "Pei Zhang",
        "Shijie Lin",
        "Zhou Ge",
        "Jinpeng Chen",
        "Wei Pu"
      ],
      "abstract": "Quasi-bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to clear silhouettes, binarization uses a minimal language to convey essential vision cues for maximum downstream efficiency, especially for tasks that require simple geometric, topological reasoning rather than heavy appearance modeling. The catch is that frame-based imaging often struggles on mobile platforms like drones, self-driving cars, and underwater vehicles, in which rapid motion causes severe motion blur and harsh lighting washes out scene details. To overcome these physical limits, neuromorphic vision via event cameras, featuring microsecond time resolution and high dynamic range, steps in as a natural solution. Building upon this event-driven paradigm, we propose a simple yet effective dual-modal approach that harnesses the synergy between frames and events for training-free, real-time, high-frame-rate binarization on CPU-only devices. Extensive evaluations show that it earns competitive performance against leading techniques in reducing blur artifacts and delivers impressive improvements under challenging illumination at a lower computational cost. Besides, its asynchronous nature bypasses long-standing event-scarcity issues that break traditional time-binning reconstruction at fixed time slots, maintaining clear target shapes even at extreme kilohertz frame rates. Its binary results further serve as reliable representations to facilitate a range of downstream tasks. This work paves the way towards lightweight perception and interaction in embodied intelligence on resource-constrained edge platforms.",
      "short_abstract": "Quasi-bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to clear silhouettes, binarization uses a minimal language to convey essential vision cues for maximum downstream...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: camera",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.17984",
      "pdf_url": "https://arxiv.org/pdf/2605.17984.pdf"
    },
    {
      "id": "2605.27418",
      "anchor": "paper-2605-27418",
      "title": "Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections",
      "authors": [
        "Wenzhe Song",
        "Hao Zhang"
      ],
      "abstract": "The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transportation systems. This paper addresses the complex problem of coordinating heterogeneous agents with disparate dynamics at unregulated intersections. We introduce a novel framework, differentiable model predictive safety (DMPS), which embeds the foresight of model-predictive control into a data-driven, end-to-end reinforcement learning architecture. DMPS agents learn a latent dynamics model to predict future trajectories contingent on their actions. A learned, differentiable safety critic then evaluates the risk of these trajectories. Crucially, by leveraging backpropagation through the entire unrolled predictive model, agents can efficiently compute the gradient of future safety with respect to their current action, enabling a minimal and precise online safety correction. Integrated into a multi-agent training scheme, DMPS virtually eliminates collisions to less than 5.6% in high-density, mixed vehicle-robot traffic simulations, demonstrating state-of-the-art safety without compromising energy and traffic efficiency.",
      "short_abstract": "The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transportation systems. This paper addresses the complex problem of coordinating heterogeneous agents with...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.27418",
      "pdf_url": "https://arxiv.org/pdf/2605.27418.pdf"
    },
    {
      "id": "2605.30368",
      "anchor": "paper-2605-30368",
      "title": "Reinterpreting Safety Thresholds as Neuron Spiking Thresholds",
      "authors": [
        "Enrico Del Re",
        "Mohamed Sabry",
        "Cristina Olaverri-Monreal"
      ],
      "abstract": "Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SSM-based evaluations employ fixed thresholds that fail to capture the human response to sustained borderline conditions or the reaction to brief, high-risk peaks. The present work proposes a biologically inspired reinterpretation of SSM thresholds. This is modelled as spiking thresholds of leaky integrate-and-fire (LIF) neurons, with multiple SSM inputs combined into a spiking neural network (SNN). The SNN is trained to emit spikes that are aligned with human braking onsets. The training data was recorded in a controlled car-following experiment using the 3D-CoAutoSim platform with CARLA/Unreal and a 6-DOF motion platform, where induced critical events were generated. The results demonstrate that the learned spiking activity qualitatively aligns with braking behaviour across scenarios and captures reactions that are not consistently explained by threshold crossings alone. Analysis across participants further indicates that learned input thresholds remain relatively consistent, while learned decay factors encode different temporal sensitivities for the SSMs. The findings of this study indicate that spiking dynamics may serve as a mechanism to facilitate the convergence of objective SSMs with subjective human safety perception.",
      "short_abstract": "Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SSM-based evaluations employ fixed thresholds that fail to capture the human response to sustained borderline...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.30368",
      "pdf_url": "https://arxiv.org/pdf/2605.30368.pdf"
    },
    {
      "id": "2605.18921",
      "anchor": "paper-2605-18921",
      "title": "Geo-Data-Driven HD Map Generation Workflow with Integrated Reference-Free Constraint-Based Verification",
      "authors": [
        "Ruidi He",
        "Vaibhav Tiwari",
        "Mohanad Al-Ghobari",
        "Meng Zhang",
        "Andreas Rausch"
      ],
      "abstract": "High-definition (HD) maps are core artifacts for automated driving systems, but their generation commonly relies on sensor-intensive mobile mapping campaigns, while quality assessment often depends on high-precision reference data. These dependencies make HD map engineering costly and difficult to apply in settings where specialised measurement data or independently measured reference maps are unavailable. This paper presents an engineering-oriented geo-data-driven workflow for HD map generation with integrated representation-level verification. The workflow uses openly available geo-engineering datasets as the primary input source and transforms them into lane-level HD map representations of existing road environments through explicit intermediate representations and processing stages. To assess the generated representations without external reference maps, the workflow integrates executable constraint-based verification into the engineering process. Selected constraints are derived from specifications relevant to automated driving and road-design guidelines. They are evaluated directly on the generated lanelet-based representation to detect geometric, topological, and elevation-related inconsistencies. The workflow is evaluated using real-world shapefile-based road-network data from four cities in Lower Saxony, Germany, and controlled defect-injection scenarios. The real-world evaluation shows that the generated map representations satisfy the selected constraints in the evaluated scenarios, while the defect-injection study demonstrates complete detection of the considered defect types without observed false positives. The results indicate that geo-data-driven HD map generation with integrated executable verification can provide a modular and inspectable complement to sensor-intensive mapping workflows under reduced sensing and reference-data availability.",
      "short_abstract": "High-definition (HD) maps are core artifacts for automated driving systems, but their generation commonly relies on sensor-intensive mobile mapping campaigns, while quality assessment often depends on high-precision reference data. These dependencies make HD...",
      "published": "2026-05-18",
      "updated": "2026-05-18",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 24,
        "margin": 4,
        "evidence": [
          "title: HD map",
          "abstract: HD map",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 93,
      "arxiv_url": "https://arxiv.org/abs/2605.18921",
      "pdf_url": "https://arxiv.org/pdf/2605.18921.pdf"
    },
    {
      "id": "2605.18895",
      "anchor": "paper-2605-18895",
      "title": "KG-ASG: Collision-Knowledge-Guided Closed-Loop Adversarial Scenario Generation With Primary-Support Attribution",
      "authors": [
        "Cheng Wang",
        "Chen Xiong",
        "Ziwen Wang",
        "Yuchen Zhou",
        "Qiang Liu"
      ],
      "abstract": "Safety validation of autonomous driving systems requires high-risk scenario coverage, clear collision semantics, executable trajectories, and attributable multi-vehicle interactions. Existing safety-critical scenario generation methods often rely on low-level trajectory perturbations, collision-proxy optimization, or single-adversary search, which may produce adversarial samples with ambiguous collision causes or uncontrolled multi-vehicle collisions. This paper proposes KG-ASG, a collision-knowledge-guided closed-loop adversarial scenario generation framework with primary-support attribution. KG-ASG constructs a structured collision knowledge base and trains a lightweight Collision Expert to infer the target collision mode, the unique primary adversary, support vehicles, and their interaction roles. Guided by this semantic prior, multi-vehicle adversarial generation is formulated as a primary-support process, where the primary adversary induces the main conflict and support vehicles shape the surrounding risk structure without becoming additional colliders. Rule, physical, interaction-safety, and single-collider constraints are imposed as hard gates to filter non-executable samples. To handle reactive ego behaviors, planner-controller feedback is further used for failure diagnosis, candidate re-ranking, and terminal refinement. Experiments on WOMD scenarios reconstructed in MetaDrive show that KG-ASG achieves strong adversarial effectiveness while improving Valid Primary Attack, reducing multi-collision, and obtaining closed-loop recovery gains under IDM, Cruise, and Expert controllers. These results demonstrate that collision-knowledge guidance and primary-support single-collider reasoning improve adversarial effectiveness, interpretability, and executability for autonomous driving safety validation.",
      "short_abstract": "Safety validation of autonomous driving systems requires high-risk scenario coverage, clear collision semantics, executable trajectories, and attributable multi-vehicle interactions. Existing safety-critical scenario generation methods often rely on low-level...",
      "published": "2026-05-17",
      "updated": "2026-05-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 22,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 94,
      "arxiv_url": "https://arxiv.org/abs/2605.18895",
      "pdf_url": "https://arxiv.org/pdf/2605.18895.pdf"
    },
    {
      "id": "2605.17822",
      "anchor": "paper-2605-17822",
      "title": "Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection",
      "authors": [
        "Yixing Yong",
        "Jian Wang",
        "Ming Lei",
        "Lijun He",
        "Fan Li"
      ],
      "abstract": "Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlike in the RGB domain, where attacks rely on color texture, infrared attacks must manipulate thermal signatures, making the geometry shape of heat-blocking materials the primary adversarial information carrier. Current shape-based methods suffer from a fundamental trade-off between representational capability and optimization power, limiting their attack effectiveness.In this work, we overcome this dilemma by introducing learnable Fourier shapes to the infrared domain. We utilize an end-to-end differentiable framework where a compact set of Fourier coefficients, defining the shape boundary, is analytically mapped to a pixel-space mask via the winding number theorem. This enables efficient gradient-based optimization to generate potent shapes that cause human targets to evade detection. Extensive digital and physical experiments provide a comprehensive evaluation and validate our superior performance. Our resulting physical patch achieves striking robustness, successfully evading detectors across diverse distances, angles, poses, and individuals, and achieves over 88% attack success rate at distances greater than 25m (conf.=0.5). Code is available at https://github.com/Yongyx99/Fourier-shape-attack.",
      "short_abstract": "Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlike in the RGB domain, where attacks rely on color texture, infrared attacks must manipulate thermal...",
      "published": "2026-05-17",
      "updated": "2026-05-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 13,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 94,
      "arxiv_url": "https://arxiv.org/abs/2605.17822",
      "pdf_url": "https://arxiv.org/pdf/2605.17822.pdf"
    },
    {
      "id": "2605.17682",
      "anchor": "paper-2605-17682",
      "title": "GEM: Gaussian Evolution Model for Occupancy Forecasting and Motion Planning",
      "authors": [
        "Cheng Chen",
        "Hao Huang",
        "Saurabh Bagchi"
      ],
      "abstract": "Future 3D semantic occupancy forecasting and motion planning are central to autonomous driving, as they require models to reason about how surrounding scenes evolve and how the ego vehicle should act. Existing occupancy world models commonly discretize scenes into latent embeddings, volumetric features, or quantized tokens, and forecast future states through fixed-step autoregressive generation. This limits temporal flexibility, obscures scene evolution, accumulates errors over long horizons, and poorly matches the continuous-time dynamics of real driving scenes. We propose GEM, a Gaussian Evolution Model for non-autoregressive occupancy world modeling, where driving scenes are represented as explicit continuous 4D Gaussian primitives with learned dynamics. Instead of rolling out future occupancy states step by step, GEM directly queries the Gaussian world representation at arbitrary timestamps and splats the corresponding conditional 3D Gaussians into semantic occupancy volumes. This enables efficient forecasting over the full horizon while retaining a compact and interpretable scene representation. By decoupling spatial geometry, temporal support, and primitive motion, GEM makes the predicted world easier to inspect, as each primitive's evolution can be followed continuously over time. The same representation also supports motion planning by predicting future ego trajectories from the learned Gaussian world. Extensive experiments show that GEM achieves state-of-the-art future semantic occupancy forecasting and strong motion planning performance, while providing flexible temporal querying.",
      "short_abstract": "Future 3D semantic occupancy forecasting and motion planning are central to autonomous driving, as they require models to reason about how surrounding scenes evolve and how the ego vehicle should act. Existing occupancy world models commonly discretize scenes...",
      "published": "2026-05-17",
      "updated": "2026-05-17",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "World Model",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 16,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 94,
      "arxiv_url": "https://arxiv.org/abs/2605.17682",
      "pdf_url": "https://arxiv.org/pdf/2605.17682.pdf"
    },
    {
      "id": "2605.17421",
      "anchor": "paper-2605-17421",
      "title": "MUSE: Multimodal Uncertainty Quantification of State Estimation",
      "authors": [
        "Minkyung Kim",
        "Henry Che",
        "Bhargav Chandaka",
        "Bhumsitt Pramuanpornsatid",
        "Chengyu Yang",
        "Sheng Cheng",
        "Xiaofeng Wang",
        "Naira Hovakimyan",
        "Shenlong Wang"
      ],
      "abstract": "Accurate visual state estimation has been a central topic in robotics with a wide range of applications in robot navigation, autonomous driving, and autonomous flight. Recent advances in robot perception have led to significant improvements in the accuracy and robustness of state estimation, yet a fundamental challenge remains in how to quantify and calibrate its precision, i.e., how confident we are in an estimate and whether failures can be detected. This issue is particularly pronounced in visual-inertial odometry (VIO), where the heteroscedastic and multimodal nature of the problem makes uncertainty quantification especially difficult. This paper introduces MUSE (Multimodal Uncertainty Quantification of State Estimation), a novel real-time learning-based framework that leverages the strong and efficient sequential modeling capacity of Mamba to estimate localization uncertainty from multiple asynchronous sensor streams. Experiments on both public and in-house datasets demonstrate that MUSE achieves superior reliability and robustness compared to existing uncertainty quantification methods, and ablation studies justify the benefits of its key design choices.",
      "short_abstract": "Accurate visual state estimation has been a central topic in robotics with a wide range of applications in robot navigation, autonomous driving, and autonomous flight. Recent advances in robot perception have led to significant improvements in the accuracy...",
      "published": "2026-05-17",
      "updated": "2026-05-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "abstract: robustness",
          "title: uncertainty",
          "abstract: uncertainty",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 94,
      "arxiv_url": "https://arxiv.org/abs/2605.17421",
      "pdf_url": "https://arxiv.org/pdf/2605.17421.pdf"
    },
    {
      "id": "2605.17284",
      "anchor": "paper-2605-17284",
      "title": "CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving",
      "authors": [
        "Ruiyang Zhu",
        "Yuehan He",
        "Boyuan Zheng",
        "Zesen Zhao",
        "Ahmad Chalhoub",
        "Qingzhao Zhang",
        "Z. Morley Mao"
      ],
      "abstract": "End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remain brittle in rare but safety-critical long-tail situations such as active construction zones and complex yielding geometries. In this paper, we present a method that addresses the long-tail challenging scenes beyond data scaling and model training. We introduce CLAP (Contrastive Latent-space Prompt optimization), a location-aware adaptation framework that augments a frozen VLA driving model with per-roadblock soft prompts, optimized from crowdsourced data and retrieved on demand via Vehicle-to-Everything (V2X) communication. Our approach rests on two observations from VLAs' latent space: (i) at the VLA's hidden-state layer, scenarios from the same roadblock cluster tightly and occupy compact regions of the latent space; and (ii) within a single roadblock, long-tail and normal frames are heavily intermixed in the latent representation, making it difficult to improve one without disturbing the other. CLAP addresses this via a two-stage pipeline: supervised contrastive learning to discover a roadblock-specific hard-scene direction, followed by directionally regularized prompt optimization that selectively improves challenging frames while preserving normal frame performance. On the NAVSIM benchmark with various state-of-the-art VLA backbones, CLAP reduces challenging scenario planning error by 24% with no regression on normal frames, significantly improving planning performance.",
      "short_abstract": "End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remain brittle in rare but safety-critical long-tail situations such as active construction zones and complex...",
      "published": "2026-05-17",
      "updated": "2026-05-17",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Cooperative / V2X",
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 34,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 94,
      "arxiv_url": "https://arxiv.org/abs/2605.17284",
      "pdf_url": "https://arxiv.org/pdf/2605.17284.pdf"
    },
    {
      "id": "2605.17268",
      "anchor": "paper-2605-17268",
      "title": "Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models",
      "authors": [
        "Nicanor Mayumu",
        "Xiaoheng Deng",
        "Patrick Mukala"
      ],
      "abstract": "We present the first systematic study of faithfulness in Vision-Language-Action (VLA) driving models, analyzing 300 Alpamayo-R1-10B inferences across 100 diverse PhysicalAI-AV scenarios. Our main finding is that output natural-language rationales with trajectories may be significantly unfaithful: (i) overall reasoning fidelity is only 42.5%, with Chain-of-Causation matching scene reality less than half the time; (ii) 94 missed pedestrians in one-third of pedestrian-relevant scenes; (iii) 97.7% trajectory fragility under mild visual perturbations; and (iv) only 48.3% mean reasoning-action consistency, with 53.3% of inferences exhibiting low consistency, including 37.9% of stop-claimed cases where the model continues instead. We formalize faithfulness information-theoretically, define entity and action fidelity with verification criteria, and outline a four-component safety architecture aligned with these results.",
      "short_abstract": "We present the first systematic study of faithfulness in Vision-Language-Action (VLA) driving models, analyzing 300 Alpamayo-R1-10B inferences across 100 diverse PhysicalAI-AV scenarios. Our main finding is that output natural-language rationales with...",
      "published": "2026-05-17",
      "updated": "2026-05-17",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 24,
        "margin": 3,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "archive",
      "age_days": 94,
      "arxiv_url": "https://arxiv.org/abs/2605.17268",
      "pdf_url": "https://arxiv.org/pdf/2605.17268.pdf"
    },
    {
      "id": "2606.20592",
      "anchor": "paper-2606-20592",
      "title": "Empowering Embodied AI in 6G Networks: Architecture, Enablers, and Open Challenges",
      "authors": [
        "Junaid Sajid",
        "Sheikh Salman Hassan",
        "Wenshuai Liu",
        "Yan Kyaw Tun",
        "Yaru Fu",
        "Nguyen H. Tran",
        "Zhu Han",
        "Cedomir Stefanovic",
        "Tharmalingam Ratnarajah",
        "Muhammad Mahtab Alam"
      ],
      "abstract": "Embodied artificial intelligence (AI) is emerging as a key driver of the sixth-generation (6G) wireless networks by enabling agents that continuously perceive, communicate, and act in dynamic physical environments. Unlike conventional AI systems that process disembodied data, embodied agents such as robots, autonomous vehicles, and extended reality (XR) devices operate through closed-loop perception-communication-action (PCA) interactions, where communication performance directly affects physical behavior, control stability, and task success. However, existing AI-native wireless architectures remain largely connectivity-centric and are not designed to support task-driven embodied intelligence at large scale. Therefore, we present a holistic framework for embodied AI-native 6G systems, in which communication, sensing, computation, and control are jointly designed as a unified closed-loop infrastructure. We introduce a system-level PCA architecture, discuss key enabling technologies and representative applications, and highlight major open challenges in multimodal intelligence, edge-aware deployment, evaluation, trustworthiness, and practical implementation. Our central argument is that future 6G systems must evolve from intelligent communication platforms into active enablers of embodied physical intelligence.",
      "short_abstract": "Embodied artificial intelligence (AI) is emerging as a key driver of the sixth-generation (6G) wireless networks by enabling agents that continuously perceive, communicate, and act in dynamic physical environments. Unlike conventional AI systems that process...",
      "published": "2026-05-17",
      "updated": "2026-05-17",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 8,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 94,
      "arxiv_url": "https://arxiv.org/abs/2606.20592",
      "pdf_url": "https://arxiv.org/pdf/2606.20592.pdf"
    },
    {
      "id": "2605.16925",
      "anchor": "paper-2605-16925",
      "title": "P2GS: Physical Prior-guided Gaussian Splatting for Photometrically Consistent Urban Reconstruction",
      "authors": [
        "Kota Shimomura",
        "Hidehisa Arai",
        "Tsubasa Takahashi",
        "Takayoshi Yamashita",
        "Hironobu Fujiyoshi"
      ],
      "abstract": "3D Gaussian Splatting (3DGS) has recently emerged as a powerful explicit representation enabling fast, high-fidelity rendering, making it a promising foundation for closed-loop simulators and perception models in autonomous driving. However, conventional 3DGS implicitly assumes consistent exposure and tone mapping across views. Real driving data violates this assumption due to heterogeneous camera pipelines and dynamic outdoor illumination, baking exposure discrepancies and sensor noise into the radiance field and producing artifacts and inconsistent illumination especially in static backgrounds crucial for realistic simulation. These issues are amplified in autonomous driving, where sparse viewpoints, varying exposures, and outdoor lighting interact, while prior work mainly targets dynamic-object reconstruction and overlooks cross-view photometric consistency. To address this limitation, we introduce P2GS, a physically consistent Gaussian Splatting framework that jointly decomposes a view-invariant linear HDR radiance field, per-view exposure scales, and tone-mapping functions from only LDR images without HDR supervision. P2GS employs a unified optimization strategy grounded in the physical image-formation process, enforcing relative-exposure consistency and HDR-domain radiance regularization. This yields a radiance field robust to inter-camera illumination differences while preserving the real-time efficiency of standard 3DGS. Experiments across real and simulated driving environments show that P2GS matches or surpasses prior methods in LDR reconstruction while providing substantially improved photometric consistency, reliable exposure normalization, and physically coherent illumination across diverse scenes.",
      "short_abstract": "3D Gaussian Splatting (3DGS) has recently emerged as a powerful explicit representation enabling fast, high-fidelity rendering, making it a promising foundation for closed-loop simulators and perception models in autonomous driving. However, conventional 3DGS...",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: camera",
          "Mapping & Localization: abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.16925",
      "pdf_url": "https://arxiv.org/pdf/2605.16925.pdf"
    },
    {
      "id": "2605.16922",
      "anchor": "paper-2605-16922",
      "title": "Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation",
      "authors": [
        "Youngdong Jang",
        "Gyeongrok Oh",
        "Jong Wook Kim",
        "Hyunju Ryu",
        "Hyung-gun Chi",
        "SeungHyeon Kim",
        "Seungryong Kim",
        "Jonghyun Choi",
        "Sangpil Kim"
      ],
      "abstract": "LiDAR scene flow estimation is essential for autonomous driving, as it provides 3D motion for each point. Self-supervised approaches use static-dynamic classification to mitigate the imbalance between static and dynamic points, deriving targeted supervision. However, existing methods rely on sparse geometric observations for this classification, making them vulnerable to data sparsity and occlusions. The resulting noisy labels provide incorrect motion guidance and degrade scene flow learning. To address this, we introduce TrackCue, a tracking-guided framework for improving dynamic object representation in LiDAR scene flow estimation. In particular, TrackCue repurposes point tracking to obtain dense image-space trajectories anchored to LiDAR points, providing motion cues beyond sparse geometric observations. Furthermore, we present a visually consistent motion compensation strategy that compares the tracked trajectories with ego-induced rigid trajectories in the image plane, effectively isolating true object motion from ego-induced apparent motion. To transfer these isolated motion cues back to the LiDAR domain, we perform visual motion cue lifting, which associates ego-compensated image trajectories with LiDAR points for static-dynamic label refinement. As a result, TrackCue produces more accurate static-dynamic classification and provides more reliable supervision for scene flow learning. Experimental results show that TrackCue significantly improves the precision and F1 score of dynamic labels, leading to performance gains in self-supervised scene flow estimation.",
      "short_abstract": "LiDAR scene flow estimation is essential for autonomous driving, as it provides 3D motion for each point. Self-supervised approaches use static-dynamic classification to mitigate the imbalance between static and dynamic points, deriving targeted supervision....",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.16922",
      "pdf_url": "https://arxiv.org/pdf/2605.16922.pdf"
    },
    {
      "id": "2605.16859",
      "anchor": "paper-2605-16859",
      "title": "VGGT-CD: Training-Free Robust Registration for 3D Change Detection",
      "authors": [
        "Wei Zhang",
        "Songhua Li",
        "Yihang Wu",
        "Qiang Li",
        "Qi Wang"
      ],
      "abstract": "3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods predominantly operate in the 2D domain, where viewpoint variations are mistaken for physical changes and depth is unavailable. While visual geometry foundation models like VGGT rapidly produce dense point clouds from unposed images, independent per-epoch reconstruction encounters fundamental obstacles: unpredictable inter-epoch scale ambiguity, registration-change paradox where scene changes corrupt alignment, and pervasive edge-flying noise. To address these challenges, we present VGGT-CD, a training-free pipeline decoupling cross-temporal registration from dynamic-change interference. In the Coarse Stage, sparse keyframe joint inference establishes a unified metric space and yields an initial Sim(3) prior. In the Fine Stage, dense reconstructions are purified by isolating static-background correspondences. A closed-form centroid alignment refines the translation while locking scale and rotation, using a residual self-check to mathematically guarantee non-degradation. Evaluated on an 11-scene benchmark from the World Across Time dataset, VGGT-CD reduces Absolute Trajectory Error by 44% outdoors and 59% indoors. It completes registration over 6 times faster, producing high-purity 3D change maps without task-specific training.",
      "short_abstract": "3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods predominantly operate in the 2D domain, where viewpoint variations are mistaken for physical changes and depth...",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "title: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.16859",
      "pdf_url": "https://arxiv.org/pdf/2605.16859.pdf"
    },
    {
      "id": "2605.16801",
      "anchor": "paper-2605-16801",
      "title": "Knapsack-based Online Sensor Selection for Vehicle State Estimation",
      "authors": [
        "Jehyeop Han",
        "Minhee Kang",
        "Alessandro Colombo",
        "Marcello Farina",
        "Heejin Ahn"
      ],
      "abstract": "As connected and autonomous driving technologies advance, vehicles increasingly rely on data from external sensors. Although this information can enhance state estimation, processing all available streams imposes significant communication and computational costs. To address this challenge, we introduce a Sensor Management Center (SMC) that selects a low-cost subset of external sensors in real time while satisfying chance-constrained error bounds derived from an Extended Kalman Filter (EKF) covariance. We formulate the selection problem as a multidimensional minimum knapsack problem and adopt a deficiency-weighted greedy algorithm as an approximate yet efficient solution. The proposed approach is validated through MATLAB simulations and experiments on a 1:15-scale cooperative driving testbed.",
      "short_abstract": "As connected and autonomous driving technologies advance, vehicles increasingly rely on data from external sensors. Although this information can enhance state estimation, processing all available streams imposes significant communication and computational...",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.16801",
      "pdf_url": "https://arxiv.org/pdf/2605.16801.pdf"
    },
    {
      "id": "2605.16892",
      "anchor": "paper-2605-16892",
      "title": "DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios",
      "authors": [
        "Sainithin Artham",
        "Shankar Gangisetty",
        "Avijit Dasgupta",
        "C. V. Jawahar"
      ],
      "abstract": "Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification and mitigation of potential risks. Although recent Multimodal Large Language Models (MLLMs) have shown promise on general vision-language tasks, our findings indicate that zero-shot MLLMs still underperform compared to domain-specific methods in fine-grained, spatially grounded risk assessment. To address this gap, we propose DriveSafe, a framework for risk-aware scene understanding that leverages structured natural language descriptions. Specifically, our method first generates spatially grounded captions enriched with multimodal context, including motion, spatial, and depth cues. These captions are then used for downstream risk assessment, explicitly identifying hazardous objects, their locations, and the unsafe behaviors they imply, followed by actionable safety suggestions. To further improve performance, we employ caption-risk pairings to fine-tune a lightweight adapter module, efficiently injecting domain-specific knowledge into the base LLM. By conditioning risk assessment on explicit language-based scene representations, DriveSafe achieves significant gains over both zero-shot MLLMs and prior domain-specific baselines. Exhaustive experiments on the DRAMA benchmark demonstrate state-of-the-art performance, while ablation studies validate the effectiveness of our key design choices. Project page: https://cvit.iiit.ac.in/ research/projects/cvit-projects/drivesafe",
      "short_abstract": "Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification and mitigation of potential risks. Although recent Multimodal Large Language Models (MLLMs) have shown promise...",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 17,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: risk assessment"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.16892",
      "pdf_url": "https://arxiv.org/pdf/2605.16892.pdf"
    },
    {
      "id": "2605.16858",
      "anchor": "paper-2605-16858",
      "title": "Pedestrian-Aware LLM-Driven Behavioral Planning for Autonomous Vehicles",
      "authors": [
        "Aidana Baimbetova",
        "Haruki Yonekura",
        "Hamada Rizk",
        "Hirozumi Yamaguchi"
      ],
      "abstract": "Autonomous Vehicles (AVs) must make reliable decisions in dense urban environments where pedestrian behavior is variable, sometimes abnormal, and often unseen during training. Reinforcement learning (RL)-based AV control systems perform well in structured traffic but struggle to generalize to unpredictable pedestrian interactions and out-of-distribution scenarios. Their reliance on handcrafted rewards and opaque decisions further limits their suitability for safety-critical, pedestrian-rich environments. To address these limitations, we introduce a Large Language Model (LLM)-based decision-making framework for pedestrian-aware behavioral planning. The system converts structured scene observations into natural-language reasoning prompts, enabling the LLM to infer pedestrian intent, anticipate risk, and generate cautious tactical driving decisions. These decisions are executed by a motion planner that ensures smooth, kinematically feasible control. We evaluate the framework in SUMO across multiple pedestrian-interaction scenarios, including unexpected jaywalking, turn-back crossing, hesitation, and bidirectional crossing. In zero-shot evaluation, the LLM-based agent achieves a 68% collision-free success rate, substantially outperforming deep RL baselines (17.7%). With few-shot episodic memory in a single-pedestrian scenario, performance increases to 96.0%, exceeding a custom DQN controller (82.0%). Cross-behavior evaluation further shows that memory derived from turn-back interactions transfers to unseen hesitation and bidirectional crossing scenarios, achieving 82.0% and 90.0% success, respectively. The system consistently initiates earlier responses, maintains wider safety buffers, and produces interpretable, human-aligned decisions.",
      "short_abstract": "Autonomous Vehicles (AVs) must make reliable decisions in dense urban environments where pedestrian behavior is variable, sometimes abnormal, and often unseen during training. Reinforcement learning (RL)-based AV control systems perform well in structured...",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 11,
        "evidence": [
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.16858",
      "pdf_url": "https://arxiv.org/pdf/2605.16858.pdf"
    },
    {
      "id": "2605.17229",
      "anchor": "paper-2605-17229",
      "title": "Generating Realistic Safety-Critical Scenarios for Vehicle-Pedestrian Interactions",
      "authors": [
        "Qingwen Pu",
        "Kun Xie",
        "Yuan Zhu",
        "Guocong Zhai"
      ],
      "abstract": "Automated driving system deployment requires rigorous validation across safety-critical vehicle-pedestrian interactions, yet real-world datasets rarely capture high-risk scenarios while simulation platforms lack realistic behavior. In response, this study proposes a three-stage framework that combines real-world grounding with adaptive simulation to generate behaviorally realistic safety-critical scenarios at scale. Stage 1 pre-trains multi-agent state-space Transformer-enhanced DDPG (MA-SST-DDPG) agents on real-world safety-critical data to learn human-like interactive evasive behaviors through data-driven learning. Stage 2 deploys pre-trained multi-agents in CARLA for online reinforcement learning to generalize across diverse scenarios, integrating real-world knowledge with simulation experience to produce a refined MA-SST-DDPG model. Stage 3 uses CARLA with the refined model to generate over 198,000 high-resolution interaction episodes from eight intersection scenarios, culminating in the Vehicle-Pedestrian Safety-Critical Interaction (VPSCI) dataset. The Refined MA-SST-DDPG model outperformed baseline methods in reproducing realistic evasive behaviors, achieving the lowest trajectory errors (ADE = 0.072 m, FDE = 0.142 m). Statistical comparison confirmed distributional equivalence between the generated and real-world data in both conflict severity and behavioral response. A Turing test confirmed that the three-stage framework generated evasive behaviors were indistinguishable from real-world interactions. These results demonstrate the framework's effectiveness in producing high-fidelity safety-critical data, offering valuable sources for the development of ADS and simulation-based safety evaluations.",
      "short_abstract": "Automated driving system deployment requires rigorous validation across safety-critical vehicle-pedestrian interactions, yet real-world datasets rarely capture high-risk scenarios while simulation platforms lack realistic behavior. In response, this study...",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 16,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.17229",
      "pdf_url": "https://arxiv.org/pdf/2605.17229.pdf"
    },
    {
      "id": "2605.16894",
      "anchor": "paper-2605-16894",
      "title": "Beyond Safety Filtering: Control Barrier Function-Informed Reinforcement Learning for Connected and Automated Vehicles",
      "authors": [
        "Jianye Xu",
        "Bassam Alrifaee"
      ],
      "abstract": "Reinforcement Learning (RL) uses rewards to guide learning, yet reward design is typically hand-crafted using heuristics that can be difficult to tune. We propose a Control Barrier Function (CBF)-informed reward design for Multi-Agent RL (MARL) that converts CBF constraint values under joint MARL actions into a reward signal that explicitly guides safe learning. We compare against two heuristic reward baselines in a four-way multi-lane intersection with connected and automated vehicles. Results show that our method achieves the highest task performance and is less sensitive to reward hyperparameters, yielding consistently strong performance across the tested hyperparameter range. Code for reproducing the experimental results and a video demonstration are available at https://github.com/bassamlab/SigmaRL.",
      "short_abstract": "Reinforcement Learning (RL) uses rewards to guide learning, yet reward design is typically hand-crafted using heuristics that can be difficult to tune. We propose a Control Barrier Function (CBF)-informed reward design for Multi-Agent RL (MARL) that converts...",
      "published": "2026-05-16",
      "updated": "2026-05-16",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: connected vehicles",
          "abstract: connected vehicles"
        ]
      },
      "recency": "archive",
      "age_days": 95,
      "arxiv_url": "https://arxiv.org/abs/2605.16894",
      "pdf_url": "https://arxiv.org/pdf/2605.16894.pdf"
    },
    {
      "id": "2605.16737",
      "anchor": "paper-2605-16737",
      "title": "DriveSafer: End-to-End Autonomous Driving with Safety Guidance",
      "authors": [
        "Shounak Sural",
        "Raj Rajkumar"
      ],
      "abstract": "End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging benchmarks. However, modern generative E2E planners still suffer from a substantial number of catastrophic failures in safety-critical scenarios. We find that many such failures arise from violations of physical constraints and safety requirements, leading to unsafe behavior. Motivated by this finding, in this paper, we focus on improving safety outcomes in generative end-to-end driving with a targeted reduction of catastrophic planning failures, instead of enhancing average planning quality. Towards this end, we propose DriveSafer, a failure-aware safety framework for end-to-end planners. DriveSafer explicitly steers generative planners towards safe behaviors leveraging both training-time safety constraints and inference-time safety guidance. Compared to the state-of-the-art DiffusionDrive model, on the NAVSIM benchmark, DriveSafer reduces the number of catastrophic failures (PDMS=0) by 48%, with over 65% reduction in drivable-area compliance failures.",
      "short_abstract": "End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging benchmarks. However, modern generative E2E planners still suffer from a substantial number of catastrophic...",
      "published": "2026-05-15",
      "updated": "2026-05-15",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 18,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 96,
      "arxiv_url": "https://arxiv.org/abs/2605.16737",
      "pdf_url": "https://arxiv.org/pdf/2605.16737.pdf"
    },
    {
      "id": "2605.16087",
      "anchor": "paper-2605-16087",
      "title": "Towards Trustworthy and Explainable AI for Perception Models: From Concept to Prototype Vehicle Deployment",
      "authors": [
        "Till Beemelmanns",
        "Shayan Sharifi",
        "Manas Mehrotra",
        "Ayushman Choudhuri",
        "Lutz Eckstein"
      ],
      "abstract": "Deep Neural Networks have become the dominant solution for Autonomous Driving perception, but their opacity conflicts with emerging Trustworthy AI guidelines and complicates safety assurance, debugging, and human oversight. While theoretical frameworks for safe and Explainable AI (XAI) exist, concrete implementations of Trustworthy AI for 3D scene understanding remain scarce. We address this gap by proposing a Trustworthy AI perception module that is remarkably robust, integrates faithful explainability, and calibrated uncertainty estimates. Building on a transformer-based detector, we derive explanation from the attention mechanism at inference time and validate their faithfulness using perturbation-based consistency tests. We further integrate an uncertainty estimation and calibration module, and apply robustness-enhancing training methods. Experiments show faithful saliency behavior, improved robustness, and well-calibrated uncertainty estimates. Finally, we deploy these Trustworthy AI elements in a prototype vehicle and provide an XAI Interface that visualizes documentation artifacts, model uncertainty state, and saliency maps, demonstrating the feasibility of trustworthy perception monitoring in real time. Supplementary materials are available at https://tillbeemelmanns.github.io/trustworthy_ai/ .",
      "short_abstract": "Deep Neural Networks have become the dominant solution for Autonomous Driving perception, but their opacity conflicts with emerging Trustworthy AI guidelines and complicates safety assurance, debugging, and human oversight. While theoretical frameworks for...",
      "published": "2026-05-15",
      "updated": "2026-05-15",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 15,
        "margin": 1,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 96,
      "arxiv_url": "https://arxiv.org/abs/2605.16087",
      "pdf_url": "https://arxiv.org/pdf/2605.16087.pdf"
    },
    {
      "id": "2605.15789",
      "anchor": "paper-2605-15789",
      "title": "Learning Context-conditioned Gaussian Overbounds for Convolution-Based Uncertainty Propagation",
      "authors": [
        "Ruirui Liu",
        "Xuejie Hou",
        "Yiping Jiang",
        "Hui Ren"
      ],
      "abstract": "Uncertainty quantification is essential in safety-critical settings--from autonomous driving to aviation, finance, and health--where decisions must rely on conservative bounds rather than point estimates. Predictor-level intervals (e.g., from quantile regression, conformal prediction, variance networks, or Bayesian models) generally do not compose: adding two per-variable intervals need not yield a valid interval for their sum or preserve coverage. In aviation, Gaussian overbounding replaces complex error distributions with a conservative Gaussian whose tails dominate the truth, so conservatism propagates through linear operations. Yet classical overbounds are global, often overly conservative, and hard to adapt to feature-conditioned errors. We propose a unified learning framework that trains neural networks to produce context-aware Gaussian overbounds--mean and scale--with provable conservatism on a finite quantile grid and, under three explicit regularity assumptions, continuous-tail conservatism on a certified interval. Our overbounding loss enforces conservativeness at selected quantiles while penalizing distributional distance with a Wasserstein-style term. The learned bounds support conservative linear-combination and convolution analysis on the enforced grid, and on the certified interval when assumptions hold, while being less redundant than traditional methods. We provide a scoped analysis of discrete-to-continuous conservatism and compact-domain objective regularity, and validate on synthetic data and real-world datasets, including multipath, ionospheric, and tropospheric residual errors. Across these settings, the method yields tighter bounds while maintaining conservatism on the enforced grid and in experiments. The framework is modality-agnostic and applicable to learning systems that require conservative, feature-conditioned uncertainty estimates in dynamic environments.",
      "short_abstract": "Uncertainty quantification is essential in safety-critical settings--from autonomous driving to aviation, finance, and health--where decisions must rely on conservative bounds rather than point estimates. Predictor-level intervals (e.g., from quantile...",
      "published": "2026-05-15",
      "updated": "2026-05-15",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 10,
        "evidence": [
          "abstract: safety",
          "title: uncertainty",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 96,
      "arxiv_url": "https://arxiv.org/abs/2605.15789",
      "pdf_url": "https://arxiv.org/pdf/2605.15789.pdf"
    },
    {
      "id": "2605.15654",
      "anchor": "paper-2605-15654",
      "title": "PCASim: Promptable Closed-loop Adversarial Simulation for Urban Traffic Environment",
      "authors": [
        "Chuancheng Zhang",
        "Zhenhao Wang",
        "Kaizheng Li",
        "Yaran Lin",
        "Qiang Guo",
        "Bin Jiang"
      ],
      "abstract": "Real-world autonomous driving, particularly in urban environments with numerous corner cases, requires rigorous testing to ensure product safety and robustness. However, few studies have explored integrating adversarial scenario generation with the training of safety agents in closed-loop testing, enabling efficient co-evolution and mutual enhancement of both. To address this challenge, an adversarial behavior knowledge repository is constructed by applying rule-based filtering to an open-source dataset, combined with knowledge retrieval modules tailored for simulation environments. A large language model (LLM) is employed to integrate knowledge-, data-, and adversarial-driven approaches, generating safety-critical traffic scenarios customized to user needs. Additionally, while evaluating the generated scenarios, we employ reinforcement learning models to train the behaviors of different types of vehicles, thereby enriching scenario diversity beyond existing datasets while preserving realism. Experimental results demonstrate that the proposed framework improves the accuracy of domain-specific language generation by 12\\%. Moreover, the success rate of newly generated scenario transformations increases by 8\\%, while obstacle-avoidance capability is enhanced by 30\\%. For the complete manuscript, please refer to: https://zhenhaooo.github.io/PCASim.github.io/",
      "short_abstract": "Real-world autonomous driving, particularly in urban environments with numerous corner cases, requires rigorous testing to ensure product safety and robustness. However, few studies have explored integrating adversarial scenario generation with the training...",
      "published": "2026-05-15",
      "updated": "2026-05-15",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset",
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 25,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 96,
      "arxiv_url": "https://arxiv.org/abs/2605.15654",
      "pdf_url": "https://arxiv.org/pdf/2605.15654.pdf"
    },
    {
      "id": "2605.16754",
      "anchor": "paper-2605-16754",
      "title": "Stable Fiber-Koopman Residual Dynamics for Environment-Constrained Robust Control",
      "authors": [
        "Syed Pouladi"
      ],
      "abstract": "Learning-based dynamical models face a persistent tension between expressiveness and formal guarantees: richer model classes improve predictive accuracy, but their stability properties are typically verified only empirically, if at all. This paper proposes \\emph{Stable Fiber-Koopman Residual Dynamics} (SFKD), a unified framework that simultaneously addresses environment-aware geometric consistency, latent-space stability certification, and bounded residual perturbation propagation. Concretely, SFKD constructs a fiber bundle latent manifold whose fibers encode environment-specific dynamics; an environment-conditioned Koopman operator governs the dominant linear evolution on each fiber; and a contraction-constrained residual neural network captures unmodeled nonlinear effects while admitting an explicit input-to-state stability (ISS) certificate. The resulting model is embedded in a sampling-based MPPI controller for autonomous vehicle path tracking under variable surface conditions and wind disturbances. Theoretical analysis establishes ISS of the latent dynamics and a finite ultimate bound on tracking error. Numerical experiments against five baselines -- Koopman MPC, Neural ODE, ICODE, ControlSynth, and ICODE-MPPI -- demonstrate a 31\\% reduction in tracking RMSE, a 44\\% improvement in control smoothness, and near-zero latent stability violation rate across environment-switching scenarios.",
      "short_abstract": "Learning-based dynamical models face a persistent tension between expressiveness and formal guarantees: richer model classes improve predictive accuracy, but their stability properties are typically verified only empirically, if at all. This paper proposes...",
      "published": "2026-05-15",
      "updated": "2026-05-15",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 10,
        "evidence": [
          "abstract: model predictive control",
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 96,
      "arxiv_url": "https://arxiv.org/abs/2605.16754",
      "pdf_url": "https://arxiv.org/pdf/2605.16754.pdf"
    },
    {
      "id": "2605.15120",
      "anchor": "paper-2605-15120",
      "title": "CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning",
      "authors": [
        "Sining Ang",
        "Yuguang Yang",
        "Canyu Chen",
        "Yan Wang"
      ],
      "abstract": "End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule-based planning metrics that measure safety, feasibility, progress, and comfort. This creates a training--evaluation mismatch: trajectories close to the logged path may violate planning rules, while alternatives farther from the demonstration can remain valid and high-scoring. The mismatch is especially limiting for proposal-selection planners, whose performance depends on candidate-set coverage and scorer ranking quality. We propose CLOVER, a Closed-LOop Value Estimation and Ranking framework for end-to-end autonomous driving planning. CLOVER follows a lightweight generator--scorer formulation: a generator produces diverse candidate trajectories, and a scorer predicts planning-metric sub-scores to rank them at inference time. To expand proposal support beyond single-trajectory imitation, CLOVER constructs evaluator-filtered pseudo-expert trajectories and trains the generator with set-level coverage supervision. It then performs conservative closed-loop self-distillation: the scorer is fitted to true evaluator sub-scores on generated proposals, while the generator is refined toward teacher-selected top-$k$ and vector-Pareto targets with stability regularization. We analyze when an imperfect scorer can improve the generator, showing that scorer-mediated refinement is reliable when scorer-selected targets are enriched under the true evaluator and updates remain conservative. On NAVSIM, CLOVER achieves 94.5 PDMS and 90.4 EPDMS, establishing a new state of the art. On the more challenging NavHard split, it obtains 48.3 EPDMS, matching the strongest reported result. On supplementary nuScenes open-loop evaluation, CLOVER achieves the lowest L2 error and collision rate among compared methods. Code data will be released at https://github.com/WilliamXuanYu/CLOVER.",
      "short_abstract": "End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule-based planning metrics that measure safety, feasibility, progress, and comfort. This creates a training--evaluation mismatch:...",
      "published": "2026-05-14",
      "updated": "2026-05-14",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 97,
      "arxiv_url": "https://arxiv.org/abs/2605.15120",
      "pdf_url": "https://arxiv.org/pdf/2605.15120.pdf"
    },
    {
      "id": "2605.15116",
      "anchor": "paper-2605-15116",
      "title": "DriveCtrl: Conditioned Sim-to-Real Driving Video Generation",
      "authors": [
        "Haonan Zhao",
        "Yiting Wang",
        "Jingkun Chen",
        "Valentina Donzella",
        "Thomas Bashford-Rogers",
        "Kurt Debattista"
      ],
      "abstract": "Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated data, the domain gap between synthetic and real-world driving videos significantly limits its utility for downstream deployment. Existing video generation methods are not well-suited for this task, as they fail to simultaneously preserve scene structure, object dynamics, temporal consistency, and visual realism, all of which are critical for maintaining annotation validity in generated data. In this paper, we present DriveCtrl, a depth-conditioned controllable sim-to-real video generation framework for realistic driving video synthesis. Built upon a pretrained video foundation model, DriveCtrl introduces a structure-aware adapter that enables depth-guided generation while preserving the scene layout and motion patterns of the source simulation, producing temporally coherent driving videos that remain aligned with the original simulated sequences. We further introduce a scalable data generation pipeline that transforms simulator videos into realistic driving footage matching the visual style of a target real-world dataset. The pipeline supports three conditioning signals: structural depth, reference-dataset style, and text prompts, while preserving frame-level annotations for downstream perception tasks. To better assess this task, we propose a driving-domain-specific knowledge-informed evaluation metric called Driving Video Realism Score (DVRS) that assesses the realism of generated videos. Experiments demonstrate that DriveCtrl consistently outperforms the base model and competing alternatives in realism, temporal quality, and perception task performance, substantially narrowing the sim-to-real gap for driving video generation.",
      "short_abstract": "Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated data, the domain gap between synthetic and real-world driving videos significantly limits its utility for...",
      "published": "2026-05-14",
      "updated": "2026-05-14",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "archive",
      "age_days": 97,
      "arxiv_url": "https://arxiv.org/abs/2605.15116",
      "pdf_url": "https://arxiv.org/pdf/2605.15116.pdf"
    },
    {
      "id": "2605.14832",
      "anchor": "paper-2605-14832",
      "title": "Learning Direct Control Policies with Flow Matching for Autonomous Driving",
      "authors": [
        "Marcello Ceresini",
        "Federico Pirazzoli",
        "Andrea Bertogalli",
        "Lorenzo Cipelli",
        "Filippo D'Addeo",
        "Anthony Dell'Eva",
        "Alessandro Paolo Capasso",
        "Alberto Broggi"
      ],
      "abstract": "We present a flow-matching planner for autonomous driving that directly outputs actionable control trajectories defined by acceleration and curvature profiles. The model is conditioned on a bird's-eye-view (BEV) raster of the surrounding scene and generates control sequences in a small number of Ordinary Differential Equations (ODE) integration steps, enabling low-latency inference suitable for real-time closed-loop re-planning. We train exclusively on urban scenarios (real urban city streets, intersections and roundabouts of the city of Parma, Italy) collected from a 2D traffic simulator with reactive agents, and evaluate in closed-loop on both in-distribution and markedly out-of-distribution environments, including multi-lane highways and unseen urban scenarios. Our results show that the model generalizes reliably to these unseen conditions, maintaining stable closed-loop control and successfully completing scenarios that differ substantially from the training distribution. We attribute this to the BEV representation, which provides a geometry-centric view of the scene that is inherently less sensitive to distributional shifts, and to the flow-matching formulation, which learns a smooth vector field that degrades gracefully under distribution shift. We provide video demonstrations of closed-loop behavior at https://marcelloceresini.github.io/DirectControlFlowMatching.",
      "short_abstract": "We present a flow-matching planner for autonomous driving that directly outputs actionable control trajectories defined by acceleration and curvature profiles. The model is conditioned on a bird's-eye-view (BEV) raster of the surrounding scene and generates...",
      "published": "2026-05-14",
      "updated": "2026-05-14",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 5,
        "evidence": [
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 97,
      "arxiv_url": "https://arxiv.org/abs/2605.14832",
      "pdf_url": "https://arxiv.org/pdf/2605.14832.pdf"
    },
    {
      "id": "2605.14696",
      "anchor": "paper-2605-14696",
      "title": "EponaV2: Driving World Model with Comprehensive Future Reasoning",
      "authors": [
        "Jiawei Xu",
        "Zhizhou Zhong",
        "Zhijian Shu",
        "Mingkai Jia",
        "Mingxiao Li",
        "Jia-Wang Bian",
        "Qian Zhang",
        "Kaicheng Zhang",
        "Jin Xie",
        "Jian Yang",
        "Wei Yin"
      ],
      "abstract": "Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive manual annotations to supervise trajectory planning, which severely limits its scalability. Conversely, although existing perception-free driving world models achieve impressive driving performance, their real-world reasoning ability for planning is solely built on next frame image forecasting. Due to the lack of enough supervision, these models often struggle with comprehensive scene understanding, resulting in unsatisfactory trajectory planning. In this paper, we propose EponaV2, a novel paradigm of driving world models, which achieves high-quality planning with comprehensive future reasoning. Inspired by how human drivers anticipate 3D geometry and semantics, we train our model to forecast more comprehensive future representations, which can be additionally decoded to future geometry and semantic maps. Extracting the 3D and semantic modalities enables our model to deeply understand the surrounding environment, and the future prediction task significantly enhances the real-world reasoning capabilities of EponaV2, ultimately leading to improved trajectory planning. Moreover, inspired by the training recipe of Large Language Models (LLMs), we introduce a flow matching group relative policy optimization mechanism to further improve planning accuracy. The state-of-the-art (SOTA) performances of EponaV2 among perception-free models on three NAVSIM benchmarks (+1.3PDMS, +5.5EPDMS) demonstrate the effectiveness of our methods.",
      "short_abstract": "Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive manual annotations to supervise trajectory planning, which severely limits its...",
      "published": "2026-05-14",
      "updated": "2026-05-14",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 16,
        "evidence": [
          "abstract: forecasting",
          "title: world model",
          "abstract: world model",
          "abstract: future prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 97,
      "arxiv_url": "https://arxiv.org/abs/2605.14696",
      "pdf_url": "https://arxiv.org/pdf/2605.14696.pdf"
    },
    {
      "id": "2605.14396",
      "anchor": "paper-2605-14396",
      "title": "Systematic Discovery of Semantic Attacks in Online Map Construction through Conditional Diffusion",
      "authors": [
        "Chenyi Wang",
        "Ruoyu Song",
        "Raymond Muller",
        "Jean-Philippe Monteuuis",
        "Jonathan Petit",
        "Z. Berkay Celik",
        "Ryan Gerdes",
        "Ming F. Li"
      ],
      "abstract": "Autonomous vehicles depend on online HD map construction to perceive lane boundaries, dividers, and pedestrian crossings -- safety-critical road elements that directly govern motion planning. While existing pixel perturbation attacks can disrupt the mapping, they can be neutralized by standard adversarial defenses. We present MIRAGE, a framework for systematic discovery of semantic attacks that bypass adversarial defenses and degrade mapping predictions by finding plausible environmental variation (e.g. shadows, wet roads). MIRAGE exploits the latent manifold of real-world data learned by diffusion models, and searches for semantically mutated scenes neighboring the ground truth with the same road topology yet mislead the mapping predictions. We evaluate MIRAGE on nuScenes and demonstrate two attacks: (1) boundary removal, suppressing 57.7% of detections and corrupting 96% of planned trajectories; and (2) boundary injection, the only method that successfully injects fictitious boundaries, while pixel PGD and AdvPatch fail entirely. Both attacks remain potent under various adversarial defenses. We use two independent VLM judges to quantify realism, where MIRAGE passes as realistic 80--84% of the time (vs. 97--99% for clean nuScenes), while AdvPatch only 0--9%. Our findings expose a categorical gap in current adversarial defenses: semantic-level perturbations that manifest as legitimate environmental variation are substantially harder to mitigate than pixel-level perturbations.",
      "short_abstract": "Autonomous vehicles depend on online HD map construction to perceive lane boundaries, dividers, and pedestrian crossings -- safety-critical road elements that directly govern motion planning. While existing pixel perturbation attacks can disrupt the mapping,...",
      "published": "2026-05-14",
      "updated": "2026-05-14",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 21,
        "margin": 1,
        "evidence": [
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 97,
      "arxiv_url": "https://arxiv.org/abs/2605.14396",
      "pdf_url": "https://arxiv.org/pdf/2605.14396.pdf"
    },
    {
      "id": "2605.14201",
      "anchor": "paper-2605-14201",
      "title": "MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving",
      "authors": [
        "Rajeev Yasarla",
        "Deepti Hegde",
        "Hsin-Pai Cheng",
        "Shizhong Han",
        "Yunxiao Shi",
        "Meysam Sadeghigooghari",
        "Hanno Ackermann",
        "Litian Liu",
        "Pranav Desai",
        "Fatih Porikli",
        "Mohammad Ghavamzadeh",
        "Hong Cai"
      ],
      "abstract": "Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to being trained under traditional imitation learning framework. Existing closed-loop supervision approaches lack scalability and fail to completely model a reactive environment. We propose MAPLE, a novel framework for reactive, multi-agent rollout of a dynamic driving scenario in the latent space of the VLA model. The ego vehicle and nearby traffic agents are independently controlled over multi-step horizons, while being reactive to other agents in the scene, enabling closed-loop training. MAPLE consists of two training stages: (1) supervised fine-tuning on the latent rollouts based on ground-truth trajectories, followed by (2) reinforcement learning with global and agent -specific rewards that encourage safety, progress, and interaction realism. We further propose diversity rewards that encourage the model to generate planning behaviors that may not be present in logged driving data. Notably, our closed-loop training framework is scalable and does not require external simulators, which can be computationally expensive to run and have limited visual fidelity to the real-world. MAPLE achieves state-of-the-art driving performance on Bench2Drive and demonstrates scalable, closed-loop multi-agent play for robust E2E autonomous driving systems.",
      "short_abstract": "Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to being trained under traditional imitation learning framework. Existing closed-loop supervision approaches lack...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation",
        "VLA",
        "Reinforcement Learning",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 30,
        "evidence": [
          "title: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.14201",
      "pdf_url": "https://arxiv.org/pdf/2605.14201.pdf"
    },
    {
      "id": "2605.14199",
      "anchor": "paper-2605-14199",
      "title": "Motion Planning for Autonomous Vehicles using Optimization over Graphs of Convex Sets",
      "authors": [
        "Matheus Wagner",
        "Antônio Augusto Fröhlich"
      ],
      "abstract": "Motion planning for autonomous vehicles requires generating collision-free and dynamically feasible trajectories in complex environments under real-time constraints. While nonlinear optimal control formulations provide high-fidelity solutions, they are computationally demanding and sensitive to initialization, whereas geometric planning methods scale well but often decouple path selection from trajectory optimization. This paper studies the extent to which optimization over Graphs of Convex Sets (GCS) can approximate solutions of nonlinear optimal control problems in the context of autonomous driving. The free space is represented as a finite union of convex regions organized as a directed graph, allowing nonconvex geometry to be handled through discrete connectivity decisions while maintaining convex trajectory constraints within each region. Vehicle motion is parameterized using Bezier curves for the spatial path and a polynomial time-scaling function for temporal evolution. Under small-slip and linear tire assumptions, a simplified dynamic bicycle model enables approximate enforcement of dynamic feasibility through convex constraints on trajectory derivatives. The approach is evaluated in CommonRoad scenarios involving static obstacle avoidance and lane-changing maneuvers, and is compared against a nonlinear discrete-time optimal control formulation. The results indicate that the GCS-based method generates collision-free and dynamically consistent trajectories that closely match those obtained from the nonlinear program, while exhibiting improved computational efficiency and reduced sensitivity to initialization. These findings suggest that GCS provides a structured approximation of nonlinear motion planning problems, capturing dominant geometric and dynamic effects while preserving convexity in the continuous relaxation.",
      "short_abstract": "Motion planning for autonomous vehicles requires generating collision-free and dynamically feasible trajectories in complex environments under real-time constraints. While nonlinear optimal control formulations provide high-fidelity solutions, they are...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 27,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.14199",
      "pdf_url": "https://arxiv.org/pdf/2605.14199.pdf"
    },
    {
      "id": "2605.13755",
      "anchor": "paper-2605-13755",
      "title": "Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception",
      "authors": [
        "Arka Bhowmick",
        "Enes Ozeren",
        "Ahmed Abdullah",
        "Oliver Wasenmuller"
      ],
      "abstract": "In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critical scenarios. Real world datasets struggle to meet this demand as require ments continuously evolve and large-scale annotated data collection remains costly and time-consuming making syn thetic data a scalable, practical and controllable alterna tive. Pedestrian detection is among the most safety-critical tasks in autonomous driving. In this paper, we propose a simple yet effective method for scaling variability in 3D pedestrian assets for synthetic scene generation. Starting from a single 3D base asset, we generate multiple distinct pedestrian instances by synthesizing diverse facial textures and identity-level appearance variations using StyleGAN2 and automatically mapping them onto 3D meshes. This ap proach enables scalable appearance-level asset diversifica tion without requiring the design of new geometries for each instance. Using the assets, we construct synthetic datasets and study the impact of mixing real and synthetic data for RGB-based object detection. Through complementary ex periments, we analyze geometry-driven distribution shifts in point cloud perception for 3D object detection. Our findings demonstrate that controlled synthetic diversifica tion improves robustness in 2D detection while revealing the sensitivity of 3D perception models to geometric domain gaps. Overall, this work highlights how generative AI en ables scalable, simulation-ready pedestrian diversification through controlled facial texture synthesis, along with the benefits and limitations of cross-domain training strategies in autonomous driving pipelines.",
      "short_abstract": "In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critical scenarios. Real world datasets struggle to meet this demand as require ments continuously evolve and...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 7,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: object detection",
          "abstract: point cloud"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13755",
      "pdf_url": "https://arxiv.org/pdf/2605.13755.pdf"
    },
    {
      "id": "2605.13751",
      "anchor": "paper-2605-13751",
      "title": "Learning Responsibility-Attributed Adversarial Scenarios for Testing Autonomous Vehicles",
      "authors": [
        "Yizhuo Xiao",
        "Haotian Yan",
        "Ying Wang",
        "Zhongpan Zhu",
        "Yuxin Zhang",
        "Xintao Yan",
        "Mustafa Suphi Erden",
        "Cheng Wang"
      ],
      "abstract": "Establishing trustworthy safety assurance for autonomous driving systems (ADSs) requires evidence that failures arise from avoidable system deficiencies rather than unavoidable traffic conflicts. Current adversarial simulation methods can efficiently expose collisions, but generally lack mechanisms to distinguish these fundamentally different failure modes. Here we present CARS (Context-Aware, Responsibility-attributed Scenario generation), a framework that integrates responsibility attribution directly into adversarial scenario generation. CARS combines context-aware adversary selection with a generative adversarial policy optimized in closed-loop simulation to construct collision scenarios that are both physically feasible and diagnostically attributable. Across benchmark datasets spanning heterogeneous national traffic environments, CARS consistently discovers feasible collision scenarios with high attribution rates under multiple regulation-prescribed careful and competent driver models. By coupling adversarial generation with normative responsibility assessment, CARS moves simulation testing beyond collision discovery toward the construction of interpretable, regulation-aligned safety evidence for scalable ADS validation.",
      "short_abstract": "Establishing trustworthy safety assurance for autonomous driving systems (ADSs) requires evidence that failures arise from avoidable system deficiencies rather than unavoidable traffic conflicts. Current adversarial simulation methods can efficiently expose...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 34,
        "margin": 30,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13751",
      "pdf_url": "https://arxiv.org/pdf/2605.13751.pdf"
    },
    {
      "id": "2605.13646",
      "anchor": "paper-2605-13646",
      "title": "Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling",
      "authors": [
        "Seokha Moon",
        "Minseung Lee",
        "Joon Seo",
        "Jinkyu Kim",
        "Jungbeom Lee"
      ],
      "abstract": "End-to-end autonomous driving, which bypasses traditional modular pipelines by directly predicting future trajectories from sensor inputs, has recently achieved substantial progress. However, existing methods often overlook the causal inter-dependencies in ego-vehicle planning, ignoring the reciprocal relations between the ego vehicle and surrounding agents. This causal oversight leads to inconsistent and unreliable trajectory predictions, especially in interaction-critical scenarios where ego decisions and neighboring agent behaviors must be reasoned about jointly. To address this limitation, we propose CaAD, a Causality-aware end-to-end Autonomous Driving framework that captures these dependencies within a shared latent scene representation. First, we propose an ego-centric joint-causal modeling module that builds on the marginal prediction branch, and learns causal dependencies between the ego vehicle and interaction-relevant agents. Second, we employ a causality-aware policy alignment stage implemented with joint-mode embeddings to align the stochastic ego policy with planning-oriented closed-loop feedback computed from surrounding traffic and map context. On the Bench2Drive and NAVSIM benchmarks, CaAD demonstrates strong closed-loop planning performance, achieving a Driving Score of 87.53 and Success Rate of 71.81 on Bench2Drive, and a PDMS of 91.1 on NAVSIM. The project page is available at https://moonseokha.github.io/CaAD/.",
      "short_abstract": "End-to-end autonomous driving, which bypasses traditional modular pipelines by directly predicting future trajectories from sensor inputs, has recently achieved substantial progress. However, existing methods often overlook the causal inter-dependencies in...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 32,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13646",
      "pdf_url": "https://arxiv.org/pdf/2605.13646.pdf"
    },
    {
      "id": "2605.13591",
      "anchor": "paper-2605-13591",
      "title": "Real2Sim: A Physics-driven and Editable Gaussian Splatting Framework for Autonomous Driving Scenes",
      "authors": [
        "Kaicong Huang",
        "Talha Azfar",
        "Weisong Shi",
        "Ruimin Ke"
      ],
      "abstract": "Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recent generative frameworks and radiance-field methods improve visual fidelity, they still struggle with temporal and spatial consistency and cannot ensure physics-aware behavior, limiting their applicability to driving scenario generation. To address these challenges, we propose Real2Sim, an unified framework that combines 4D Gaussian Splatting (4DGS) with a differentiable Material Point Method (MPM) solver. Real2Sim explicitly reconstructs dynamic driving scenes as temporally continuous Gaussian primitives, supports instance-level editing, and simulates realistic object-object and object-environment interactions. This framework enables physics-aware, high-fidelity synthesis of diverse, editable scenarios, including challenging corner cases such as collisions and post-impact trajectories. Experiments on the Waymo Open Dataset validate Real2Sim's capabilities in rendering, reconstruction, editing, and physics simulation, demonstrating its potential as a scalable tool for data generation in downstream tasks such as perception, tracking, trajectory prediction, and end-to-end policy learning.",
      "short_abstract": "Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recent generative frameworks and...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 3,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13591",
      "pdf_url": "https://arxiv.org/pdf/2605.13591.pdf"
    },
    {
      "id": "2605.13177",
      "anchor": "paper-2605-13177",
      "title": "Volumetric Optical Scattering Neural Networks",
      "authors": [
        "Xuhao Luo",
        "Qiang Song",
        "Weiwei Cai",
        "Lei Chen",
        "Enbo Yang",
        "Hao Wang",
        "Zhipei Sun",
        "Yueqiang Hu",
        "Joel K. W. Yang",
        "Huigao Duan"
      ],
      "abstract": "Optical neural networks offer a route to low-latency and energy-efficient inference by encoding computation in light propagation. However, most existing implementations rely on planar photonic circuits or discretely spaced diffractive layers, restricting volumetric integration and imposing stringent alignment requirements. Here we demonstrate a volumetric optical scattering neural network (OSNN) in which densely packed weak scatterers form a three-dimensional, locally connected optical computing medium. In contrast to fully connected diffractive architectures, the OSNN uses near-field scattering interactions, described under the first-Born approximation, to compress optical interconnections into a monolithic volume. We implement this concept using resilient inverse design and two-photon nanolithography, yielding OSNN devices with a volume of ~$3.8*10^{-4}mm^{3}$ and a record-breaking neuron density of $1.0*10^{9}/mm^{3}$. Experimentally, the fabricated classifier achieves $94.8\\%$ blind-test accuracy on MNIST, while the imager performs optical compressed imaging with a $1-μm$ effective resolution and average FSIM values of $0.93$ on Fashion-MNIST and $0.91$ on VesselMNIST3D. OSNN paves the way for ultra-dense, ultra-compact, and efficient optical computing, creating a universal platform for embedded optical intelligence and promising widespread application in AI fields ranging from autonomous driving to medical diagnosis.",
      "short_abstract": "Optical neural networks offer a route to low-latency and energy-efficient inference by encoding computation in light propagation. However, most existing implementations rely on planar photonic circuits or discretely spaced diffractive layers, restricting...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 10,
        "evidence": [
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13177",
      "pdf_url": "https://arxiv.org/pdf/2605.13177.pdf"
    },
    {
      "id": "2605.13168",
      "anchor": "paper-2605-13168",
      "title": "Variance-Aware Estimation and Inference for Michaelis--Menten Models with Heteroscedastic Errors and Clustered Measurements",
      "authors": [
        "Mijeong Kim",
        "Minkyoung Cha",
        "Ah Young Jeong"
      ],
      "abstract": "Michaelis--Menten analysis is often conducted by nonlinear least squares under a constant-variance assumption, even though enzyme-kinetic data frequently display concentration-dependent heteroscedasticity and often include repeated or clustered measurements. We develop a variance-aware procedure for Michaelis--Menten estimation and inference that is motivated by conditional moment restrictions and implemented through simple conditionally Gaussian working models. For single curves, the method reduces to one-dimensional root finding for $K_m$ followed by closed-form plug-in updates for $V_{\\max}$ and a variance scale parameter; the same score logic yields a cluster-level extension through a random-effect-induced working covariance. In simulation, modeling heteroscedasticity improved variance recovery and interval efficiency relative to homoscedastic nonlinear least squares, while cluster-aware semiparametric and NLME fits restored fixed-effect coverage far more effectively than pooled analyses that ignored clustering. In self-driving laboratory and soil exoenzyme data, heteroscedastic models achieved lower information criteria than homoscedastic nonlinear least squares, with the square-root variance function giving the most stable empirical fit among the prespecified working models. We implement the workflow in the companion \\texttt{inferMM} package for single-curve, grouped, and clustered Michaelis--Menten analysis. These results show that simple variance-function and covariance modeling can stabilize original-scale Michaelis--Menten inference when variability changes with substrate concentration or measurements are clustered.",
      "short_abstract": "Michaelis--Menten analysis is often conducted by nonlinear least squares under a constant-variance assumption, even though enzyme-kinetic data frequently display concentration-dependent heteroscedasticity and often include repeated or clustered measurements....",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13168",
      "pdf_url": "https://arxiv.org/pdf/2605.13168.pdf"
    },
    {
      "id": "2605.14204",
      "anchor": "paper-2605-14204",
      "title": "Day-to-Day Traffic Network Modeling under Route-Guidance Misinformation: Endogenous Trust and Resilience in CAV Environments",
      "authors": [
        "Eunhan Ka",
        "Satish V. Ukkusuri"
      ],
      "abstract": "Connected and autonomous vehicles and smart mobility services increasingly use digital route guidance as an operational input to traffic network management. When this information becomes unreliable or adversarial, day-to-day traffic models must represent not only flow adaptation but also the evolution of user trust in the information source. This paper develops a coupled day-to-day traffic assignment and trust-evolution framework for route-guidance misinformation. Within-day congestion is represented by Lighthill-Whitham-Richards network loading, while day-to-day route choice follows bounded-rationality logit learning with trust-dependent reliance on external guidance. Trust is modeled as an aggregate class-level behavioral reliance state encoded by a Beta evidence model and updated from repeated guidance errors. Theoretical analysis establishes stationary equilibria, a conservative stability guide, a weighted compliance index for population-level vulnerability, and an asymmetric recovery law that explains post-attack trust hysteresis. Numerical experiments on Sioux Falls, with an Anaheim robustness check, show that endogenous trust creates a threshold-based resilience mechanism. Below the trust-activation threshold, the attack remains behaviorally stealthy and dynamic trust provides almost no attenuation. Above the threshold, trust erosion reduces the impact of the fixed-trust attack by about 91 percent in Sioux Falls and 85 percent in Anaheim. The experiments also show that CAV penetration increases fixed-trust vulnerability while preserving dynamic attenuation, and that traffic performance can recover before trust, resulting in a 77-day hidden vulnerability window. The results provide a trust-aware modeling basis for resilience analysis in CAV-enabled traffic networks.",
      "short_abstract": "Connected and autonomous vehicles and smart mobility services increasingly use digital route guidance as an operational input to traffic network management. When this information becomes unreliable or adversarial, day-to-day traffic models must represent not...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.14204",
      "pdf_url": "https://arxiv.org/pdf/2605.14204.pdf"
    },
    {
      "id": "2605.13699",
      "anchor": "paper-2605-13699",
      "title": "Memristor Technologies for Dynamic Vision Sensors: A Critical Assessment and Research Roadmap",
      "authors": [
        "Mohamad Yazan Sadoun",
        "Edris Zaman Farsa",
        "Sarah Sharif",
        "Yaser Mike Banad"
      ],
      "abstract": "Edge-AI deployment is bottlenecked by data-movement energy; pairing event-driven vision sensors with in-memory analog compute could lift that ceiling by orders of magnitude. Both technologies are individually mature; the framework distinguishing fabricated demonstrations from projected systems is missing. Of six application domains surveyed (robotics, autonomous vehicles, AR/VR, surveillance, medical imaging, IoT), half rest entirely on projection, and existing hardware sits at Technology Readiness Levels 2-5. This evidence-graded review applies a three-paradigm architectural taxonomy and benchmarks the gap against current digital neuromorphic alternatives. It identifies an end-to-end integrated DVS-memristor system as the field's open challenge, with testable accuracy and power targets.",
      "short_abstract": "Edge-AI deployment is bottlenecked by data-movement energy; pairing event-driven vision sensors with in-memory analog compute could lift that ceiling by orders of magnitude. Both technologies are individually mature; the framework distinguishing fabricated...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13699",
      "pdf_url": "https://arxiv.org/pdf/2605.13699.pdf"
    },
    {
      "id": "2605.13220",
      "anchor": "paper-2605-13220",
      "title": "Real-time Gaussian Process based Approximate Model Predictive Trajectory Tracking Control for Autonomous Vehicles",
      "authors": [
        "Alexander Rose",
        "Lukas Theiner",
        "Rolf Findeisen"
      ],
      "abstract": "Applying model predictive control on embedded systems remains challenging due to the high computational cost of solving optimal control problems. To address this limitation, computationally efficient Gaussian process approximations of the implicit model predictive control law can be employed. However, for trajectory-tracking applications, the large amount of training data required for successful generalization across distinct reference trajectories poses a significant challenge. To improve data efficiency, we propose to transform the model into curvilinear coordinates around the reference trajectory. Secondly, we use a nominal feedforward component, allowing the Gaussian process to learn only the residual control input, making the approximation of a trajectory-tracking controller feasible. To underline the applicability of the approach, we deploy the controller on a Raspberry Pi in a small-scale vehicle and validate it experimentally. Compared to a model predictive control implementation using real-time iterations, the Gaussian process based approximation computes control inputs about five times faster while achieving similar closed-loop tracking performance.",
      "short_abstract": "Applying model predictive control on embedded systems remains challenging due to the high computational cost of solving optimal control problems. To address this limitation, computationally efficient Gaussian process approximations of the implicit model...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 16,
        "evidence": [
          "abstract: model predictive control",
          "abstract: controller",
          "title: control",
          "abstract: control",
          "title: trajectory tracking"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13220",
      "pdf_url": "https://arxiv.org/pdf/2605.13220.pdf"
    },
    {
      "id": "2605.13539",
      "anchor": "paper-2605-13539",
      "title": "Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles",
      "authors": [
        "Christian Geller",
        "Daniel Becker",
        "Jobst Beckmann",
        "Lutz Eckstein"
      ],
      "abstract": "Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are reliable, the real world must be modeled with sufficient fidelity, including not only the static environment but also the surrounding traffic of a vehicle under test. Thus, the availability of traffic agent models is of common interest to model naturalistic and parameterizable behavior, similar to human drivers. The interchangeability of agent models across different simulation environments represents a major challenge and necessitates harmonization and standardization. To address this challenge, we present a standardized and modular simulation integration architecture that enables the tool-independent integration of traffic agent models. The architecture builds upon the Open Simulation Interface (OSI) as a structured message format and the Functional Mock-up Interface (FMI) for dynamic model exchange. Rather than introducing yet another model or simulation tool, we provide a reusable reference implementation that translates these standards into a practical integration blueprint, including clear interfaces, data mappings, and execution semantics. The generic nature of the architecture is demonstrated by integrating an exemplary agent model into three widely used simulation environments: OpenPASS, CARLA, and CarMaker. As part of the evaluation, we show that the model yields consistent behavior in all simulation platforms, thereby validating the interoperability, modularity, and standard compliance of the proposed architecture. The reference implementation lowers integration barriers, serves as a foundation for future research, and is made publicly available at github.com/ika-rwth-aachen/agent-model-integration",
      "short_abstract": "Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are reliable, the real world must be modeled with sufficient fidelity, including not only the static environment...",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13539",
      "pdf_url": "https://arxiv.org/pdf/2605.13539.pdf"
    },
    {
      "id": "2605.13525",
      "anchor": "paper-2605-13525",
      "title": "Beyond VMAF: Towards Application-Specific Metrics for Teleoperation Video",
      "authors": [
        "Ines Trautmannsheimer",
        "Richard Grauberger",
        "Frank Diermeyer"
      ],
      "abstract": "Automated driving has made remarkable progress, yet situations still arise where human intervention is necessary. Teleoperation provides a scalable solution to address such cases, enabling remote operators to support vehicles without being physically present. In this context, video transmission forms the operator's primary source of situational awareness, making video quality a decisive factor for both safety and task performance. In an online study, participants rated compressed video sequences from the Zenseact Dataset and provided subjective quality ratings. These ratings were then used to retrain the Video Multi-Method Assessment Fusion (VMAF) model, yielding an adapted variant tailored to teleoperation. The retrained model demonstrated improved alignment with human ratings compared to the original 4K VMAF. In particular, RMSE decreased from 10.36 to 8.83, and MAD from 8.71 to 6.38, corresponding to improvements of 15% and 27%, respectively. These results highlight that incorporating domain-specific data can enhance the predictive power of established quality metrics in safety-critical applications. At the same time, Outlier cases emerged in which videos received high objective scores despite noticeable degradations in regions critical for the driving task.",
      "short_abstract": "Automated driving has made remarkable progress, yet situations still arise where human intervention is necessary. Teleoperation provides a scalable solution to address such cases, enabling remote operators to support vehicles without being physically present....",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 16,
        "evidence": [
          "title: teleoperation",
          "abstract: teleoperation"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13525",
      "pdf_url": "https://arxiv.org/pdf/2605.13525.pdf"
    },
    {
      "id": "2605.13125",
      "anchor": "paper-2605-13125",
      "title": "MoCCA: A Movable Circle Probability of Collision Approximation",
      "authors": [
        "Tobias Kern",
        "Christian Birkner"
      ],
      "abstract": "In automated driving, crash mitigation is crucial to ensure passenger safety. Accurate avoidance requires precise knowledge of the object's position and orientation. However, sensor noise and occlusions often result in tracking and prediction uncertainties. To account for these uncertainties, estimating the Probability of Collision (POC) is a critical requirement. While Monte Carlo sampling is a common estimation technique, its high computational demand and stochastic nature often render it unsuitable for real-time applications. Analytical POC calculations are simplified by approximating vehicle geometries using circular bounds. While multi-circle approximations offer higher fidelity than a single circumscribed circle, they significantly increase computational complexity. This paper proposes a shape approximation algorithm, MoCCA, which utilizes a single circle for each vehicle, optimized to minimize the relative distance between them. MoCCA maintains a computational efficiency comparable to standard single-circle techniques while reducing over-conservatism. To address the potential underestimation of POC inherent in partial coverage, we establish an upper bound for the approximation error, demonstrating that it depends primarily on inter-vehicle distance and orientation variance. Furthermore, we introduce a safety distance margin that can be calibrated solely based on orientation variance.",
      "short_abstract": "In automated driving, crash mitigation is crucial to ensure passenger safety. Accurate avoidance requires precise knowledge of the object's position and orientation. However, sensor noise and occlusions often result in tracking and prediction uncertainties....",
      "published": "2026-05-13",
      "updated": "2026-05-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 98,
      "arxiv_url": "https://arxiv.org/abs/2605.13125",
      "pdf_url": "https://arxiv.org/pdf/2605.13125.pdf"
    },
    {
      "id": "2605.12957",
      "anchor": "paper-2605-12957",
      "title": "GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion",
      "authors": [
        "Hanxin Zhu",
        "Cong Wang",
        "Peiyan Tu",
        "Jiayi Luo",
        "Tianyu He",
        "Xin Jin",
        "Zhibo Chen"
      ],
      "abstract": "Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains including spatial intelligence, embodied intelligence, and autonomous driving. While achieving remarkable progress, existing approaches to 3D world generation typically prioritize appearance prediction with limited modeling of the underlying geometry, leading to issues such as unreliable scene structure estimation and degraded cross-view consistency. To address these limitations, motivated by the coarse-to-fine nature of human visual perception, we propose GTA, a novel image-to-3D world generation method following a Geometry-Then-Appearance paradigm. Specifically, given a single input image, to improve the structural fidelity of synthesized 3D scenes, GTA adopts a two-stage framework with two dedicated video diffusion models, which first generate coarse geometric structure from novel viewpoints and then synthesize fine-grained appearance conditioned on the predicted geometry. To further enhance cross-view appearance consistency, we introduce a random latent shuffle strategy during the training process, along with a test-time scaling scheme that improves perceptual quality without compromising quantitative performance. Extensive experiments have demonstrated that our proposed method consistently outperforms existing approaches in terms of fidelity, visual quality, and geometric accuracy. Moreover, GTA is shown to be effective as a general enhancement module that further improves the generation quality of existing image-to-3D world pipelines, as well as supporting multiple downstream applications and exhibiting favorable data efficiency during model training, highlighting its versatility and broad applicability. Project page: https://hanxinzhu-lab.github.io/GTA/.",
      "short_abstract": "Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains including spatial intelligence, embodied intelligence, and autonomous driving. While achieving remarkable...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12957",
      "pdf_url": "https://arxiv.org/pdf/2605.12957.pdf"
    },
    {
      "id": "2605.12789",
      "anchor": "paper-2605-12789",
      "title": "Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention",
      "authors": [
        "Hamza Ahmed Durrani",
        "Rafay Suleman Durrani"
      ],
      "abstract": "Large language-vision models (LVLMs) such as CLIP, Flamingo, and BLIP have revolutionized AI by enabling understanding across textual and visual modalities. These models excel at tasks like image captioning, visual question answering, and cross-modal retrieval. However, they face catastrophic forgetting when learning new tasks sequentially, particularly challenging in multi-modal settings where preserving cross-modal alignments adds complexity to the learning process. This paper presents a comprehensive continual learning framework for LVLMs that combines enhanced Elastic Weight Consolidation (EWC) with parameter-efficient fine-tuning techniques. We integrate multi-modal Fisher Information Matrix calculation, consistency preservation across modalities, and adaptive regularization that considers dependencies across visual and textual encoders. The framework achieves a 78% reduction in forgetting rates relative to naive sequential training approaches through extensive evaluation testing. The framework also preserves alignment between modalities during sequential learning with only 15% additional computational cost. This work advances the state of the art in lifelong learning for multi-modal AI systems, with direct applications to autonomous driving, intelligent robotic assistants, and adaptive robotic systems that must continuously learn in dynamic real-world environments.",
      "short_abstract": "Large language-vision models (LVLMs) such as CLIP, Flamingo, and BLIP have revolutionized AI by enabling understanding across textual and visual modalities. These models excel at tasks like image captioning, visual question answering, and cross-modal...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12789",
      "pdf_url": "https://arxiv.org/pdf/2605.12789.pdf"
    },
    {
      "id": "2605.12743",
      "anchor": "paper-2605-12743",
      "title": "Still Camouflage, Moving Illusion: View-Induced Trajectory Manipulation in Autonomous Driving",
      "authors": [
        "Shuo Ju",
        "Qingzhao Zhang",
        "Huashan Chen",
        "Xuheng Wang",
        "Haotang Li",
        "Wanqian Zhang",
        "Feng Liu",
        "Kebin Peng",
        "Sen He"
      ],
      "abstract": "Existing physical adversarial attacks on vision-based autonomous driving induce time-evolving perception errors, including biased object tracking or trajectory prediction, through (i) sophisticated physical patch inducing detection box drift when entering the view distance, or (ii) dynamically changing patches that cause different perception errors at different time. In both cases, viewing-angle variation is treated as a challenge, requiring adversarial patches to remain effective across frames under varying views, leading to complex multi-view optimization. In contrast, we show that viewing-angle variation itself can be turned into an attack tool. We design a new attack paradigm where a static, passive adversarial camouflage is mounted on a vehicle whose view-dependent appearance naturally evolves with relative motion, inducing consistent feature drift across frames. This causes the system to infer a physically plausible but incorrect trajectory, such as a false cut-in, which propagates to downstream decision-making and triggers unnecessary braking. Unlike prior approaches that require multi-view robustness or active intervention, our attack emerges from normal driving dynamics and is easy to deploy: a parked vehicle with a natural camouflage can induce hard braking in passing autonomous vehicles. We demonstrate the novel attack on nuScenes dataset, showing the effectiveness with an end-to-end success rate of up to 87.5%, measured by hard-braking events, and robustness across different scene backgrounds, victim vehicle speeds, and perception models.",
      "short_abstract": "Existing physical adversarial attacks on vision-based autonomous driving induce time-evolving perception errors, including biased object tracking or trajectory prediction, through (i) sophisticated physical patch inducing detection box drift when entering the...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12743",
      "pdf_url": "https://arxiv.org/pdf/2605.12743.pdf"
    },
    {
      "id": "2605.12674",
      "anchor": "paper-2605-12674",
      "title": "Revealing Interpretable Failure Modes of VLMs",
      "authors": [
        "Isha Chaudhary",
        "Vedaant V Jain",
        "Kavya Sachdeva",
        "Sayan Ranu",
        "Gagandeep Singh"
      ],
      "abstract": "Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to generalize with minimal task-specific engineering. Despite these advantages, they can exhibit catastrophic failures in specific real-world situations, constituting failure modes. We introduce REVELIO, a framework for systematically uncovering interpretable failure modes in VLMs. We define a failure mode as a composition of interpretable, domain-relevant concepts-such as pedestrian proximity or adverse weather conditions-under which a target VLM consistently behaves incorrectly. Identifying such failures requires searching over an exponentially large discrete combinatorial space. To address this challenge, REVELIO combines two search procedures: a diversity-aware beam search that efficiently maps the failure landscape, and a Gaussian-process Thompson Sampling strategy that enables broader exploration of complex failure modes. We apply REVELIO to autonomous driving and indoor robotics domains, uncovering previously unreported vulnerabilities in state-of-the-art VLMs. In driving environments, the models often demonstrate weak spatial grounding and fail to account for major obstructions, leading to recommendations that would result in simulated crashes. In indoor robotics tasks, VLMs either miss safety hazards or behave excessively conservatively, producing false alarms and reducing operational efficiency. By identifying structured and interpretable failure modes, REVELIO offers actionable insights that can support targeted VLM safety improvements.",
      "short_abstract": "Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to generalize with minimal task-specific engineering. Despite these advantages, they can exhibit catastrophic...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12674",
      "pdf_url": "https://arxiv.org/pdf/2605.12674.pdf"
    },
    {
      "id": "2605.12624",
      "anchor": "paper-2605-12624",
      "title": "MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving",
      "authors": [
        "Yuzhou Huang",
        "Benjin Zhu",
        "Hengtong Lu",
        "Victor Shea-Jay Huang",
        "Haiming Zhang",
        "Wei Chen",
        "Jifeng Dai",
        "Yan Xie",
        "Hongsheng Li"
      ],
      "abstract": "Autonomous driving has progressed from modular pipelines toward end-to-end unification, and Vision-Language-Action (VLA) models are a natural extension of this journey beyond Vision-to-Action (VA). In practice, driving VLAs have often trailed VA on planning quality, suggesting that the difficulty is not simply model scale but the interface through which semantic reasoning, temporal context, and continuous control are combined. We argue that this gap reflects how VLA has been built -- as isolated subtask improvements that fail to compose coherent driving capabilities -- rather than what VLA is. We present MindVLA-U1, the first unified streaming VLA architecture for autonomous driving. A unified VLM backbone produces AR language tokens (optional) and flow-matching continuous action trajectories in a single forward pass over one shared representation, preserving the natural output form of each modality. A full streaming design processes the driving video framewise rather than as fixed video-action chunks under costly temporal VLM modeling. Planned trajectories evolve smoothly across frames while a learned streaming memory channel carries temporal context and updates. The unified architecture enables fast/slow systems on dense & sparse MoT backbones via flexible self-attention context management, and exposes a measurable language-control path for action: language-predicted driving intents steers the action diffusion via classifier-free guidance (CFG), turning language-side intent into control signals for continuous action planning. On the long-tail WOD-E2E benchmark, MindVLA-U1 surpasses experienced human drivers for the first time (8.20 RFS vs. 8.13 GT RFS) with 2 diffusion steps, achieves state-of-the-art planning ADEs over prior VA/VLA by large margins, and matches VA latency (16 FPS vs. RAP's 18 FPS at 1B scale) while preserving natural language interfaces for human-vehicle interaction.",
      "short_abstract": "Autonomous driving has progressed from modular pipelines toward end-to-end unification, and Vision-Language-Action (VLA) models are a natural extension of this journey beyond Vision-to-Action (VA). In practice, driving VLAs have often trailed VA on planning...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 24,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12624",
      "pdf_url": "https://arxiv.org/pdf/2605.12624.pdf"
    },
    {
      "id": "2605.12622",
      "anchor": "paper-2605-12622",
      "title": "Action Emergence from Streaming Intent",
      "authors": [
        "Pengfei Jing",
        "Victor Shea-Jay Huang",
        "Hengtong Lu",
        "Jifeng Dai",
        "Yan Xie",
        "Benjin Zhu"
      ],
      "abstract": "We formalize action emergence as a target capability for end-to-end autonomous driving: the ability to generate physically feasible, semantically appropriate, and safety-compliant actions in arbitrary, long-tail traffic scenes through scene-conditioned reasoning rather than retrieval or interpolation of learned scene-action mappings. We show that previous paradigms cannot deliver action emergence: autoregressive trajectory decoders collapse the inherently multimodal future into a single averaged output, while diffusion and flow-matching generators express multimodality but are not steerable by reasoned intent. We propose Streaming Intent as a concrete way to approach action emergence: a mechanism that makes driving intent (i) semantically streamed through a continuous chain-of-thought that causally derives the intent from scene understanding, and (ii) temporally streamed across clips so that intent commitments remain coherent along the driving horizon. We realize Streaming Intent in a VLA model we call SI (Streaming Intent). SI autoregressively decodes a four-step chain-of-thought and emits an intent token; the decoded intent then drives classifier-free guidance (CFG) on a flow-matching action head, requiring only two denoising steps to generate the final trajectory. On the Waymo End-to-End benchmark, SI achieves competitive aggregate performance, with an RFS score of 7.96 on the validation set and 7.74 on the test set. Beyond aggregate metrics, the model demonstrates -- to our knowledge for the first time in a fully end-to-end VLA -- intent-faithful controllability: for a fixed scene, varying the intent class at inference yields qualitatively distinct yet consistently high-quality plans, arising purely from data-driven learning without any pre-built trajectory bank or hand-coded post-hoc selector.",
      "short_abstract": "We formalize action emergence as a target capability for end-to-end autonomous driving: the ability to generate physically feasible, semantically appropriate, and safety-compliant actions in arbitrary, long-tail traffic scenes through scene-conditioned...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 8,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12622",
      "pdf_url": "https://arxiv.org/pdf/2605.12622.pdf"
    },
    {
      "id": "2605.11987",
      "anchor": "paper-2605-11987",
      "title": "Random-Set Graph Neural Networks",
      "authors": [
        "Tommy Woodley",
        "Shireen Kudukkil Manchingal",
        "Matteo Tolloso",
        "Davide Bacciu",
        "Fabio Cuzzolin"
      ],
      "abstract": "Uncertainty quantification has become an important factor in understanding the data representations produced by Graph Neural Networks (GNNs). Despite their predictive capabilities being ever useful across industrial workspaces, the inherent uncertainty induced by the nature of the data is a huge mitigating factor to GNN performance. While aleatoric uncertainty is the result of noisy and incomplete stochastic data such as missing edges or over-smoothing, epistemic uncertainty arises from lack of knowledge about a system or model (e.g., a graph's topology or node feature representation), which can be reduced by gathering more data and information. In this paper, we propose an original new framework in which node-level epistemic uncertainty is modelled in a belief function (finite random set) formalism. The resulting Random-Set Graph Neural Networks have a belief-function head predicting a random set over the list of classes, from which both a precise probability prediction and a measure of epistemic uncertainty can be obtained. Extensive experiments on 9 different graph learning datasets, including real-world autonomous driving benchmarks as such Nuscene and ROAD, demonstrate RS-GNN's superior uncertainty quantification capabilities",
      "short_abstract": "Uncertainty quantification has become an important factor in understanding the data representations produced by Graph Neural Networks (GNNs). Despite their predictive capabilities being ever useful across industrial workspaces, the inherent uncertainty...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 6,
        "evidence": [
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11987",
      "pdf_url": "https://arxiv.org/pdf/2605.11987.pdf"
    },
    {
      "id": "2605.11678",
      "anchor": "paper-2605-11678",
      "title": "OOM-Free Alpamayo via CPU-GPU Memory Swapping for Vision-Language-Action Models",
      "authors": [
        "Seungwoo Roh",
        "Huiyeong Kim",
        "Jong-Chan Kim"
      ],
      "abstract": "End-to-end Vision-Language-Action (VLA) models for autonomous driving unify perception, reasoning, and control in a single neural network, achieving strong driving performance but requiring 20-60GB of GPU memory-far exceeding the 12-16GB available on commodity GPUs. We present a framework, which enables memory-efficient VLA inference on VRAM-constrained GPUs through system-level optimization alone, without model modification. Our work proceeds in three stages: (1) Sequential Demand Layering reduces VRAM usage from model-level to layer-level granularity; (2) Pipelined Demand Layering hides parameter transfer time within layer execution time via transfer--compute overlap; and (3) a GPU-Resident Layer Decision Policy, informed by per-module residency benefit analysis, eliminates the residual transfer overhead that pipelining cannot hide. We further propose a performance prediction model that determines the optimal configuration-both the number and placement of resident layers-from a single profiling run with less than 1.3% prediction error across all configurations. Applied to NVIDIA's Alpamayo-R1-10B (21.52GB) on an RTX 5070Ti (16GB), our work achieves up to 3.55x speedup over Accelerate offloading while maintaining full BF16 precision.",
      "short_abstract": "End-to-end Vision-Language-Action (VLA) models for autonomous driving unify perception, reasoning, and control in a single neural network, achieving strong driving performance but requiring 20-60GB of GPU memory-far exceeding the 12-16GB available on...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 23,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11678",
      "pdf_url": "https://arxiv.org/pdf/2605.11678.pdf"
    },
    {
      "id": "2605.11594",
      "anchor": "paper-2605-11594",
      "title": "PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations",
      "authors": [
        "Cheng Chi",
        "Xianqi Wang",
        "Hongcheng Luo",
        "Mingfei Tu",
        "Gangwei Xu",
        "Zehan Zhang",
        "Bing Wang",
        "Guang Chen",
        "Hangjun Ye",
        "Sida Peng",
        "Xin Yang",
        "Haiyang Sun"
      ],
      "abstract": "High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their per-pixel Gaussian prediction paradigm often suffers from multi-view inconsistency and layering artifacts. Moreover, existing methods often model dynamic instances via dense flow prediction, which lacks explicit cross-view correspondence and instance-level consistency. In this paper, we propose PointForward, a feedforward driving reconstruction framework through point-aligned representations. Unlike pixel-aligned methods, we initialize sparse 3D queries in world space and aggregate multi-view image information via spatial-temporal fusion onto these queries, enforcing explicit cross-view consistency in a single feedforward pass. To handle scene dynamics, we introduce scene graphs that explicitly organize moving instances during reconstruction. By leveraging 3D bounding boxes, our method enables instance-level motion propagation and temporally consistent dynamic representations. Extensive experiments demonstrate that PointForward achieves state-of-the-art performance on large-scale driving benchmarks. The code will be available upon the publication of the paper.",
      "short_abstract": "High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their per-pixel Gaussian prediction paradigm often suffers from multi-view...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11594",
      "pdf_url": "https://arxiv.org/pdf/2605.11594.pdf"
    },
    {
      "id": "2605.11550",
      "anchor": "paper-2605-11550",
      "title": "The DAWN of World-Action Interactive Models",
      "authors": [
        "Hongbo Lu",
        "Liang Yao",
        "Chenghao He",
        "Haoyu Wang",
        "Xiang Gu",
        "Xianfei Li",
        "Wenlong Liao",
        "Tao He",
        "Pai Peng"
      ],
      "abstract": "A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action Models (WAMs) largely miss this reciprocity, treating world prediction and action generation as either isolated parallel branches or rigid predict-then-plan pipelines. We formalize this perspective as World-Action Interactive Models (WAIMs), and instantiate it in autonomous driving with \\textbf{DAWN} (\\textbf{D}enoising \\textbf{A}ctions and \\textbf{W}orld i\\textbf{N}teractive model), a simple yet strong latent generative baseline. DAWN operates in a compact semantic latent space and couples a \\emph{World Predictor} with a \\emph{World-Conditioned Action Denoiser}: the predicted world hypothesis conditions action denoising, while the denoised action hypothesis is fed back to update the world prediction, so that both are recursively refined during inference. Rather than eliminating test-time world evolution altogether or rolling out the full future in pixel space, DAWN performs a short explicit latent rollout that is sufficient to support long-horizon trajectory generation in complex interactive scenes. Experiments show that DAWN achieves strong planning performance and favorable safety-related results across multiple autonomous driving benchmarks. More broadly, our results suggest that interactive world-action generation is a principled path toward truly actionable world models.",
      "short_abstract": "A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action Models (WAMs) largely miss this reciprocity, treating world prediction and action generation as either...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: abstract: planning",
          "Planning & Decision-Making: abstract: trajectory generation",
          "Prediction & World Models: abstract: world model",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11550",
      "pdf_url": "https://arxiv.org/pdf/2605.11550.pdf"
    },
    {
      "id": "2605.11521",
      "anchor": "paper-2605-11521",
      "title": "XWOD: A Real-World Benchmark for Object Detection under Extreme Weather Conditions",
      "authors": [
        "Chih-Hsin Chen",
        "Yu-Tung Liu",
        "Amar Fadillah",
        "Kuan-Ting Lai",
        "Dong Liu"
      ],
      "abstract": "Autonomous driving and intelligent transportation systems remain vulnerable under extreme weather. The U.S. Federal Highway Administration reports that roughly 745,000 crashes and 3,800 fatalities per year are weather-related, and recent regulatory investigations have examined failures of Level-2/3 driving systems under reduced-visibility conditions. However, datasets commonly used to evaluate weather robustness remain limited in scale, diversity, and realism. In this paper, we introduce XWOD (Extreme Weather Object Detection), a large-scale real-world traffic-object detection benchmark containing 10,010 images and 42,924 bounding boxes across seven extreme weather conditions: rain, snow, fog, haze/sand/dust, flooding, tornado, and wildfire. The dataset covers six traffic-object categories, including car, person, truck, motorcycle, bicycle, and bus. XWOD extends the weather taxonomy from one to seven conditions, and is the first to cover the emerging class of climate-amplified hazards, such as flooding, tornado, and wildfire. To evaluate the quality of our data, we train standard YOLO-family detectors on XWOD and test them zero-shot on external weather benchmarks, achieving mAP$_{50}$ scores of 63.00% on RTTS, 59.94% on DAWN, and 61.12% on WEDGE, compared with the corresponding published YOLO-based baselines of 40.37%, 32.75%, and 45.41%, respectively, representing relative improvements of 56%, 83%, and 35%. These cross-dataset results show that XWOD provides a strong source domain for learning weather-robust traffic perception. We release the dataset, splits, baseline weights, and reproducible evaluation code under a research-use license.",
      "short_abstract": "Autonomous driving and intelligent transportation systems remain vulnerable under extreme weather. The U.S. Federal Highway Administration reports that roughly 745,000 crashes and 3,800 fatalities per year are weather-related, and recent regulatory...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 15,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11521",
      "pdf_url": "https://arxiv.org/pdf/2605.11521.pdf"
    },
    {
      "id": "2605.11520",
      "anchor": "paper-2605-11520",
      "title": "PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting",
      "authors": [
        "Yixiao Song",
        "Qingyong Li",
        "Wen Wang",
        "Zhicheng Yan"
      ],
      "abstract": "Unsupervised point cloud segmentation is critical for embodied artificial intelligence and autonomous driving, as it mitigates the prohibitive cost of dense point-level annotations required by fully supervised methods. While integrating 2D pre-trained models such as the Segment Anything Model (SAM) to supplement semantic information is a natural choice, this approach faces a fundamental mismatch between discrete 3D points and continuous 2D images. This mismatch leads to inevitable projection overlap and complex modality alignment, resulting in compromised semantic consistency across 2D-3D transfer. To address these limitations, this paper proposes PointGS, a simple yet effective pipeline for unsupervised 3D point cloud segmentation. PointGS leverages 3D Gaussian Splatting as a unified intermediate representation to bridge the discrete-continuous domain gap. Input sparse point clouds are first reconstructed into dense 3D Gaussian spaces via multi-view observations, filling spatial gaps and encoding occlusion relationships to eliminate projection-induced semantic conflation. Multi-view dense images are rendered from the Gaussian space, with 2D semantic masks extracted via SAM, and semantics are distilled to 3D Gaussian primitives through contrastive learning to ensure consistent semantic assignments across different views. The Gaussian space is aligned with the original point cloud via two-step registration, and point semantics are assigned through nearest-neighbor search on labeled Gaussians. Experiments demonstrate that PointGS outperforms state-of-the-art unsupervised methods, achieving +0.9% mIoU on ScanNet-V2 and +2.8% mIoU on S3DIS.",
      "short_abstract": "Unsupervised point cloud segmentation is critical for embodied artificial intelligence and autonomous driving, as it mitigates the prohibitive cost of dense point-level annotations required by fully supervised methods. While integrating 2D pre-trained models...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "title: point cloud",
          "abstract: point cloud"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11520",
      "pdf_url": "https://arxiv.org/pdf/2605.11520.pdf"
    },
    {
      "id": "2605.12628",
      "anchor": "paper-2605-12628",
      "title": "Multistep Belief Space Dynamics Learning For Risk-Aware Control",
      "authors": [
        "Jason Gibson",
        "Bogdan Vlahov",
        "Patrick Spieler",
        "Evangelos A. Theodorou"
      ],
      "abstract": "As autonomous vehicles move from a simplified research setting to practical use, there exists a large gap between the dynamic behavior of a human driving and an autonomous system. Risk-aware behavior needs to naturally develop in order to scale to the demands of the real world. A major issue for risk-aware planning and control has been predicting how dynamical uncertainty evolves through time and optimizing plans that account for this without being overly conservative. Here, we present a learning framework to predict distributional dynamics that can be optimized in real time for Model Predictive Control (MPC). We explore the importance of structure when learning distributional dynamics for use in MPC. A rigorous ablation study is conducted on a large dataset of real world off-road driving that shows the impact of deviations from our proposed structure. Furthermore, we deploy our learned model and planning stack on a full sized vehicle in challenging off-road conditions. Our planning architecture is able to naturally regulate the speed of the vehicle based on the environment and consistently demonstrates intelligent behavior over miles of diverse terrain.",
      "short_abstract": "As autonomous vehicles move from a simplified research setting to practical use, there exists a large gap between the dynamic behavior of a human driving and an autonomous system. Risk-aware behavior needs to naturally develop in order to scale to the demands...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 18,
        "margin": 5,
        "evidence": [
          "title: risk assessment",
          "abstract: risk assessment",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12628",
      "pdf_url": "https://arxiv.org/pdf/2605.12628.pdf"
    },
    {
      "id": "2605.12608",
      "anchor": "paper-2605-12608",
      "title": "A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline",
      "authors": [
        "Mohamed Ahmed Mohamed",
        "Xiaowei Huang"
      ],
      "abstract": "Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains a significant bottleneck. In this paper, we propose Clear2Fog (C2F), an end-to-end, physics-based pipeline that simulates fog for clear-weather datasets under a unified camera and LiDAR framework. C2F combines monocular depth estimation with a novel atmospheric light estimation method to improve the physical consistency of synthetic fog generation while reducing structural artefacts and chromatic biases observed in existing frameworks. Utilising a training set of 270,000 images from the Waymo Open Dataset, we conduct an extensive data efficiency study to investigate whether environmental diversity can reduce dataset scale requirements and improve model generalisation under varying fog conditions. Our findings reveal that models trained on mixed-density fog datasets at 75% scale achieve performance that is not statistically different from those trained on fixed-density datasets at 100% scale, reducing synthetic training data requirements by 25%. We observe that this efficiency trend is consistent across two representative detector architectures. Furthermore, we investigate the sim-to-real transfer by using C2F-generated data as a pre-training foundation before fine-tuning on real-world fog data. We demonstrate that, within the evaluated settings, a relative 10x increase in the default fine-tuning learning rate reduces the negative transfer caused by standard fine-tuning, resulting in an absolute increase of up to 0.0117 mAP beyond the real-only baseline. Overall, this work demonstrates the value of diverse synthetic fog as a pre-training tool for real-world adaptation.",
      "short_abstract": "Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains a significant bottleneck. In this paper, we propose Clear2Fog (C2F), an end-to-end, physics-based pipeline...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 20,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12608",
      "pdf_url": "https://arxiv.org/pdf/2605.12608.pdf"
    },
    {
      "id": "2605.11940",
      "anchor": "paper-2605-11940",
      "title": "Lane-Aware Graph Attention Network for Multi-Vehicle Trajectory Prediction in Expressway Merge Zones",
      "authors": [
        "Eni Solomon Laughter"
      ],
      "abstract": "Accurate multi-vehicle trajectory prediction in expressway merge and diverge areas is fundamental to the decision-making frameworks of autonomous vehicle systems. However, the majority of existing graph-based prediction models are developed and validated on mainline freeway segments and do not address the geometrically distinct interaction structures that characterize merge zones. Furthermore, standard evaluation protocols rely exclusively on displacement error metrics, leaving the safety consequences of predicted trajectories unquantified. This paper proposes a Lane-Aware Graph Attention Network (LA-GAT) that encodes vehicle interaction within dynamic scene graphs, augmented with a trainable lane-relationship attention bias that prioritizes merge-conflict interactions from the outset of training. The model is pre-trained on the raw NGSIM US-101 and I-80 datasets and subsequently fine-tuned on UAV-captured UTE SQM-W-1 trajectory data from a Chinese expressway merge area, with final evaluation on the held-out SQM-W-2 dataset. Evaluation spans both displacement metrics (ADE, FDE at 1s, 3s, 5s horizons) and surrogate safety measures (TTC violation rate, DRAC exceedance rate, collision rate). Fine-tuned results on SQM-W-2 yield ADE of 0.865 m at 1s and 2.518 m at 3s, demonstrating that drone-informed fine-tuning substantially reduces the cross-dataset transfer gap. The deliberate use of unfiltered NGSIM data is shown to characterize raw-condition generalization limits, with the performance degradation attributed to the well-documented measurement errors in that dataset.",
      "short_abstract": "Accurate multi-vehicle trajectory prediction in expressway merge and diverge areas is fundamental to the decision-making frameworks of autonomous vehicle systems. However, the majority of existing graph-based prediction models are developed and validated on...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 20,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11940",
      "pdf_url": "https://arxiv.org/pdf/2605.11940.pdf"
    },
    {
      "id": "2605.11799",
      "anchor": "paper-2605-11799",
      "title": "SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions",
      "authors": [
        "Markus Essl",
        "Marta Moscati",
        "Mubashir Noman",
        "Muhammad Zaigham Zaheer",
        "Usman Naseem",
        "Shah Nawaz",
        "Markus Schedl"
      ],
      "abstract": "Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods transform multimodal data from independent sensors, such as camera and LiDAR, into a unified bird's-eye view (BEV) representation for fusion. Although effective in ideal conditions, this strategy suffers from substantial performance deterioration when camera or LiDAR data are missing, corrupted, or noisy. To address this vulnerability, we develop a framework-agnostic fusion module for camera and LiDAR data that allows for handling cases when one of the two modalities is missing or corrupted. To demonstrate the effectiveness of our module, we instantiate it in BEVFusion [1], a well-established framework to combine camera and LiDAR data for 3D object detection. By means of quantitative experiments on the MultiCorrupt dataset, we demonstrate that our module achieves favorable performance improvements under scenarios of missing and corrupted modalities, substantially outperforming existing unified representation approaches across a wide range of sensor deterioration scenarios and reaching state-of-the-art performance in scenarios of corrupted modality due to extreme weather conditions and sensor failure.",
      "short_abstract": "Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods transform multimodal data from independent sensors, such as camera and LiDAR,...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 7,
        "evidence": [
          "abstract: object detection",
          "abstract: sensor fusion",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11799",
      "pdf_url": "https://arxiv.org/pdf/2605.11799.pdf"
    },
    {
      "id": "2605.11798",
      "anchor": "paper-2605-11798",
      "title": "Advancing Dynamic Ride-Pooling Simulation -- A Highly Scalable Dispatcher",
      "authors": [
        "Moritz Laupichler",
        "Robin Andre",
        "Kim Kandler",
        "Peter Sanders",
        "Peter Vortisch"
      ],
      "abstract": "In ride-pooling, a fleet of vehicles is dynamically dispatched to bring travelers from A to B, trying to pool riders with similar itineraries to improve the use of resources compared to taxis or private cars. Ride-pooling is considered a core building block of future transport systems with autonomous vehicles. In this paper, we introduce Mt-KaRRi, a novel dispatcher for dynamic ride-pooling that leverages state-of-the-art shortest-path algorithms to process millions of travelers per hour. We add a simple mode choice model and use realistic travel demand in three different urban areas for extensive experiments. We find that our dispatcher scales well with a response time per request of around 1ms even for our largest instances. We show how this scalability can be used to conduct ride-pooling studies at unprecedented scale. For instance, we determine how the quality of rides and usage of vehicle resources develop for tens of thousands of vehicles and millions of travelers. We envision Mt-KaRRi as a tool for future ride-pooling simulation studies at scale.",
      "short_abstract": "In ride-pooling, a fleet of vehicles is dynamically dispatched to bring travelers from A to B, trying to pool riders with similar itineraries to improve the use of resources compared to taxis or private cars. Ride-pooling is considered a core building block...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.11798",
      "pdf_url": "https://arxiv.org/pdf/2605.11798.pdf"
    },
    {
      "id": "2605.12719",
      "anchor": "paper-2605-12719",
      "title": "A Five-Layer MLOps Architecture for Connected Automated Driving",
      "authors": [
        "Bastian Lampe",
        "Lutz Eckstein"
      ],
      "abstract": "The continual assurance of safety and performance of automated driving systems (ADSs) poses significant challenges. ADSs operate in complex, dynamic, open-world environments allowing a wide range of scenarios, including ones that are rare or not foreseen during initial development. While the incorporation of artificial intelligence (AI) and machine learning (ML) technology allows ADSs to learn from data gathered during operation and thus enables them to adapt over time, these approaches come with their own challenges. A key advantage of ADSs compared to human drivers is their greater ability to gather data collectively across a fleet of vehicles, or even across multiple fleets operated by different entities, and to learn from this data collectively. Vehicles can share and combine their data to identify additional learning opportunities otherwise missed by individual vehicles. This creates new opportunities to tackle the challenges of continual assurance of safety and performance, but requires the implementation of architectures that leverage the collective learning potential. Based on established MLOps principles and existing work in the field of connected automated driving, this paper presents a five-layer architecture for collective learning-enabled MLOps processes for ADSs. The goal of this architecture is to provide a conceptual blueprint for the design and implementation of MLOps processes by fleet operators and other relevant stakeholders. The paper describes the main responsibilities of each layer, their interactions, and how multi-level self-assessments enabled by the architecture can support the detection and reduction of edge cases including black swan events.",
      "short_abstract": "The continual assurance of safety and performance of automated driving systems (ADSs) poses significant challenges. ADSs operate in complex, dynamic, open-world environments allowing a wide range of scenarios, including ones that are rare or not foreseen...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Human Factors & Policy: abstract: human driver"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12719",
      "pdf_url": "https://arxiv.org/pdf/2605.12719.pdf"
    },
    {
      "id": "2605.12710",
      "anchor": "paper-2605-12710",
      "title": "Belief-Space Residual Risk for Automated Driving under Localization Uncertainty",
      "authors": [
        "Nijinshan Karunainayagam",
        "Nils Gehrke",
        "Frank Diermeyer"
      ],
      "abstract": "Residual risk metrics have recently been introduced to assess the safety implications of automated driving systems. Existing approaches typically assume a deterministic ego pose and concentrate mainly on perception errors related to surrounding objects and latency effects. In practice, however, automated vehicles operate under considerable localization uncertainty, especially in complex urban settings and in adverse weather conditions. This work extends the spatial residual risk formulation to the belief space by explicitly modeling ego pose uncertainty as a Gaussian distribution. Residual risk is reformulated as the expected degradation-induced risk over the ego pose belief distribution. Within a particle-based risk estimation framework, localization uncertainty is incorporated into the computation of collision probabilities through covariance fusion of ego and object uncertainties.",
      "short_abstract": "Residual risk metrics have recently been introduced to assess the safety implications of automated driving systems. Existing approaches typically assume a deterministic ego pose and concentrate mainly on perception errors related to surrounding objects and...",
      "published": "2026-05-12",
      "updated": "2026-05-12",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 8,
        "evidence": [
          "title: localization",
          "abstract: localization"
        ]
      },
      "recency": "archive",
      "age_days": 99,
      "arxiv_url": "https://arxiv.org/abs/2605.12710",
      "pdf_url": "https://arxiv.org/pdf/2605.12710.pdf"
    },
    {
      "id": "2605.10904",
      "anchor": "paper-2605-10904",
      "title": "MDrive: Benchmarking Closed-Loop Cooperative Driving for End-to-End Multi-agent Systems",
      "authors": [
        "Marco Coscoy",
        "Zewei Zhou",
        "Seth Z. Zhao",
        "Henry Wei",
        "Angela Magtoto",
        "Johnson Liu",
        "Rui Song",
        "Walter Zimmer",
        "Zhiyu Huang",
        "Chen Tang",
        "Bolei Zhou",
        "Jiaqi Ma"
      ],
      "abstract": "Vehicle-to-Everything (V2X) communication has emerged as a promising paradigm for autonomous driving, enabling connected agents to share complementary perception information and negotiate with each other to benefit the final planning. Existing V2X benchmarks, however, fall short in two ways: (i) open-loop evaluations fail to capture the inherently closed-loop nature of driving, leading to evaluation gaps, and (ii) current closed-loop evaluations lack behavioral and interactive diversity to reflect real-world driving. Thus, it is still unclear the extent of benefits of multi-agent systems for closed-loop driving. In this paper, we introduce MDrive, a closed-loop cooperative driving benchmark comprising 225 scenarios grounded in both NHTSA pre-crash typologies and real-world V2X datasets. Our benchmark results demonstrate that multi-agent systems are generally better than single-agent counterparts. However, current multi-agent systems still face two important challenges: (i) perception sharing enhances perceptions, but doesn't always translate to better planning; (ii) negotiation improves planning performance but harms it in complex and dense traffic scenarios. MDrive further provides an open-source toolbox for scenario generation, Real2Sim conversion, and human-in-the-loop simulation. Together, MDrive establishes a reproducible foundation for evaluating and improving the generalization and robustness of cooperative driving systems.",
      "short_abstract": "Vehicle-to-Everything (V2X) communication has emerged as a promising paradigm for autonomous driving, enabling connected agents to share complementary perception information and negotiate with each other to benefit the final planning. Existing V2X benchmarks,...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Benchmark",
        "Cooperative / V2X",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 11,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10904",
      "pdf_url": "https://arxiv.org/pdf/2605.10904.pdf"
    },
    {
      "id": "2605.10744",
      "anchor": "paper-2605-10744",
      "title": "C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving",
      "authors": [
        "Kefei Tian",
        "Yuansheng Lian",
        "Kai Yang",
        "Xiangdong Chen",
        "Shen Li"
      ],
      "abstract": "Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing methods, whether rule-based or data-driven, frequently struggle to capture complex scene semantics, infer potential risks, and make reliable decisions in rare, high-risk situations. While vision-language models (VLMs) offer promising approaches for safe decision-making in these environments, most current approaches lack reflective and causal reasoning, thereby limiting their overall robustness. To address this, we propose a counterfactual chain-of-thought (C-CoT) framework that leverages VLMs to decompose driving decisions into five sequential stages: scene description, critical object identification, risk prediction, counterfactual risk reasoning, and final action planning. Within the counterfactual reasoning stage, we introduce a structured meta-action evaluation tree to explicitly assess the potential consequences of alternative action combinations. This self-reflective reasoning establishes causal links between action choices and safety outcomes, improving robustness in long-tail and out-of-distribution scenarios. To validate our approach, we construct the DeepAccident-CCoT dataset based on the DeepAccident benchmark and fine-tune a Qwen2.5-VL (7B) model using low-rank adaptation. Our model achieves a risk prediction recall of 81.9%, reduces the collision rate to 3.52%, and lowers L2 error to 1.98 m. Ablation studies further confirm the critical role of counterfactual reasoning and the meta-action evaluation tree in enhancing safety and interpretability.",
      "short_abstract": "Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing methods, whether rule-based or data-driven, frequently struggle to capture complex scene semantics, infer...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Dataset",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 1,
        "evidence": [
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10744",
      "pdf_url": "https://arxiv.org/pdf/2605.10744.pdf"
    },
    {
      "id": "2605.10564",
      "anchor": "paper-2605-10564",
      "title": "DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving",
      "authors": [
        "Lingjun Zhang",
        "Changjie Wu",
        "Linzhe Shi",
        "Jiangyang Li",
        "Jiaxin Liu",
        "Lei Yang",
        "Hang Zhang",
        "Mu Xu",
        "Hong Wang"
      ],
      "abstract": "End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptations from general domains, lacking in-depth exploration tailored to autonomous driving scenarios, particularly within visual reasoning modules. In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird's-eye-view (BEV) space, thereby enabling long-horizon modeling of future world states. We also introduce an efficient and adaptive text reasoning mechanism that utilizes additional social knowledge and reasoning capabilities to further improve driving performance in challenging long-tail scenarios. We present a novel, efficient, and effective approach that achieves state-of-the-art (SOTA) results on the closed-loop Bench2drive benchmark. Codes are available at: https://github.com/hotdogcheesewhite/DeepSight.",
      "short_abstract": "End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10564",
      "pdf_url": "https://arxiv.org/pdf/2605.10564.pdf"
    },
    {
      "id": "2605.10426",
      "anchor": "paper-2605-10426",
      "title": "CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving",
      "authors": [
        "Minqing Huang",
        "Yujiao Xiang",
        "Zihan Liang",
        "Jiajie Huang",
        "Jingqi Wang",
        "Zhi Xu",
        "Feiyang Tan",
        "Hangning Zhou",
        "Mu Yang",
        "Gong Che"
      ],
      "abstract": "Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide planning-oriented intermediate representations: textual Chain-of-Thought (CoT) fails to preserve continuous spatiotemporal structure, while latent world reasoning remains difficult to use as a direct condition for action generation. In this paper, we propose CoWorld-VLA, a multi-expert world reasoning framework for autonomous driving, where world representations serve as explicit conditions to guide action planning. CoWorld-VLA extracts complementary world information through multi-source supervision and encodes it into expert tokens within the VLA, thereby providing planner-accessible conditioning signals. Specifically, we construct four types of tokens: semantic interaction, geometric structure, dynamic evolution, and ego trajectory tokens, which respectively model interaction intent, spatial structure, future temporal dynamics, and behavioral goals. During action generation, CoWorld-VLA employs a diffusion-based hierarchical multi-expert fusion planner, which is coupled with scene context throughout the joint denoising process to generate continuous ego trajectories. Experiments show that CoWorld-VLA achieves competitive results in both future scene generation and planning on the NAVSIM v1 benchmark, demonstrating strong performance in collision avoidance and trajectory accuracy. Ablation studies further validate the complementarity of expert tokens and their effectiveness as planning conditions for action generation. Code will be available at https://github.com/AFARI-Research/CoWorld-VLA.",
      "short_abstract": "Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide planning-oriented intermediate representations: textual Chain-of-Thought (CoT) fails...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Synthetic Data",
        "VLA",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 21,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10426",
      "pdf_url": "https://arxiv.org/pdf/2605.10426.pdf"
    },
    {
      "id": "2605.10388",
      "anchor": "paper-2605-10388",
      "title": "Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction",
      "authors": [
        "Yumao Liu",
        "Tao Liu",
        "Xiangyu Li",
        "Jiaxiang Li",
        "Ke Ma"
      ],
      "abstract": "End to end (E2E) autonomous driving trajectory prediction is often trained with camera frames sampled at the highest available temporal frequency, assuming that denser sampling improves performance. We question this assumption by treating temporal sampling frequency as an explicit training set design variable. Starting from high frequency E2E driving datasets, we construct frequency sweep training sets by temporally subsampling camera frames along each trajectory. For each model dataset pair, we train and evaluate the same model under a fixed protocol, so the frequency response reflects how prediction performance changes with sampling frequency. We analyze this response from a capacity aware perspective. Sparse sampling may miss driving relevant cues, while dense sampling may add redundant visual content and off manifold noise. For finite capacity models, this can create a driving irrelevant capacity burden. We evaluate three smaller E2E models and a larger VLA style AutoVLA model on Waymo, nuScenes, and PAVE. Results show model and dataset dependent frequency responses. Smaller E2E models often show non monotonic or near plateau trends and achieve their best 3 second ADE at lower or intermediate frequencies. In contrast, AutoVLA achieves its best 3 second ADE and FDE at the highest evaluated frequency on all three datasets. Iteration matched controls suggest that the advantage of lower or intermediate frequencies for smaller models is not explained only by unequal training update counts. These findings show that temporal sampling frequency should be reported and tuned, rather than fixed to the highest available value.",
      "short_abstract": "End to end (E2E) autonomous driving trajectory prediction is often trained with camera frames sampled at the highest available temporal frequency, assuming that denser sampling improves performance. We question this assumption by treating temporal sampling...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 8,
        "evidence": [
          "title: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10388",
      "pdf_url": "https://arxiv.org/pdf/2605.10388.pdf"
    },
    {
      "id": "2605.10386",
      "anchor": "paper-2605-10386",
      "title": "GuardAD: Safeguarding Autonomous Driving MLLMs via Markovian Safety Logic",
      "authors": [
        "Tianyuan Zhang",
        "Peng Yue",
        "Zihao Peng",
        "Jiangfan Liu",
        "Zonghao Ying",
        "Jiakai Wang",
        "Tianlin Li",
        "Jian Yang",
        "Yaodong Yang",
        "Aishan Liu",
        "Xianglong Liu"
      ],
      "abstract": "Multimodal large language models (MLLMs) are increasingly integrated into autonomous driving (AD) systems; however, they remain vulnerable to diverse safety threats, particularly in accident-prone scenarios. Recent safeguard mechanisms have shown promise by incorporating logical constraints, yet most rely on static formulations that lack temporally grounded safety reasoning over evolving traffic interactions, resulting in limited robustness in dynamic driving environments. To address these limitations, we propose GuardAD, a model-agnostic safeguard that formulates AD safety as an evolving Markovian logical state. GuardAD introduces Neuro-Symbolic Logic Formalization, which represents safety predicates over heterogeneous traffic participants and continuously induces them via n-th order Markovian Logic Induction. This design enables the inference of emerging and latent hazards beyond single-step observations. Rather than simply vetoing unsafe actions, GuardAD performs Logic-Driven Action Revision, where inferred safety states actively guide action refinement without modifying the underlying MLLM. Extensive experiments on multiple benchmarks and AD-MLLMs demonstrate that GuardAD substantially reduces accident rates (-32.07%) while slightly improving task performance (+6.85%). Moreover, closed-loop simulation evaluations, together with physical-world vehicle studies, further validate the effectiveness and potential of GuardAD.",
      "short_abstract": "Multimodal large language models (MLLMs) are increasingly integrated into autonomous driving (AD) systems; however, they remain vulnerable to diverse safety threats, particularly in accident-prone scenarios. Recent safeguard mechanisms have shown promise by...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 22,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10386",
      "pdf_url": "https://arxiv.org/pdf/2605.10386.pdf"
    },
    {
      "id": "2605.10177",
      "anchor": "paper-2605-10177",
      "title": "MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning",
      "authors": [
        "Guangli Chen",
        "Dianzhao Li",
        "Wenjian Zhong",
        "Bangquan Xie",
        "Ostap Okhrin"
      ],
      "abstract": "Robust urban autonomous driving requires reliable 3D scene understanding and stable decision-making under dense interactions. However, existing end-to-end models lack interpretability, while modular pipelines suffer from error propagation across brittle interfaces. This paper proposes MTA-RL, the first framework that bridges perception and control through Multi-modal Transformer-based 3D Affordances and Reinforcement Learning (RL). Unlike previous fusion models that directly regress actions, RGB images and LiDAR point clouds are fused using a transformer architecture to predict explicit, geometry-aware affordance representations. These structured representations serve as a compact observation space, enabling the RL policy to operate purely on predicted driving semantics, which significantly improves sample efficiency and stability. Extensive evaluations in CARLA Town01-03 across varying densities (20-60 background vehicles) show that MTA-RL consistently outperforms state-of-the-art baselines. Trained solely on Town03, our method demonstrates superior zero-shot generalization in unseen towns, achieving up to a 9.0% increase in Route Completion, an 11.0% increase in Total Distance, and an 83.7% improvement in Distance Per Violation. Furthermore, ablation studies confirm that our multi-modal fusion and reward shaping are critical, significantly outperforming image-only and unshaped variants, demonstrating the effectiveness of MTA-RL for robust urban autonomous driving.",
      "short_abstract": "Robust urban autonomous driving requires reliable 3D scene understanding and stable decision-making under dense interactions. However, existing end-to-end models lack interpretability, while modular pipelines suffer from error propagation across brittle...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "abstract: perception",
          "abstract: point cloud",
          "abstract: LiDAR",
          "abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10177",
      "pdf_url": "https://arxiv.org/pdf/2605.10177.pdf"
    },
    {
      "id": "2605.10117",
      "anchor": "paper-2605-10117",
      "title": "Think as Needed: Geometry-Driven Adaptive Perception for Autonomous Driving",
      "authors": [
        "Donghyun Kim",
        "Jaehyoung Park"
      ],
      "abstract": "Autonomous driving scenes range from empty highways to dense intersections with dozens of interacting road users, yet current 3D detection models apply a fixed computation budget to every frame, wasting resources on simple scenes while lacking capacity for complex ones. Existing approaches compound this problem: Transformer-based interaction models scale quadratically with the number of detected objects, and frame-by-frame processing causes the system to immediately forget objects the moment they become occluded. We propose Enhanced HOPE, an adaptive perception architecture that measures the geometric complexity of each incoming LiDAR frame using an unsupervised statistical estimator and routes it through a shallow or deep processing path accordingly, requiring no manual scene labels. To keep interaction modeling efficient, we replace quadratic pairwise attention with a linear-time subspace-based network that groups nearby objects into clusters and processes them jointly. The computational savings from these two mechanisms free up resources for a persistent temporal memory module that retains previously detected objects and traffic rules across frames, enabling the system to recall occluded objects seconds after they disappear from view. On the nuScenes and CARLA benchmarks, Enhanced HOPE reduces latency by 38% on simple scenes with no accuracy loss, improves mean Average Precision by 2.7 points on rare long-tail scenarios, and tracks objects through occlusions lasting over 5 seconds, where all tested baselines fail.",
      "short_abstract": "Autonomous driving scenes range from empty highways to dense intersections with dozens of interacting road users, yet current 3D detection models apply a fixed computation budget to every frame, wasting resources on simple scenes while lacking capacity for...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 11,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10117",
      "pdf_url": "https://arxiv.org/pdf/2605.10117.pdf"
    },
    {
      "id": "2605.10034",
      "anchor": "paper-2605-10034",
      "title": "Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving",
      "authors": [
        "Aron Distelzweig",
        "Faris Janjoš",
        "Andreas Look",
        "Anna Rothenhäusler",
        "Daniel Jost",
        "Oliver Scheel",
        "Raghu Rajan",
        "Daphne Cornelisse",
        "Eugene Vinitsky",
        "Joschka Boedecker"
      ],
      "abstract": "Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driving policies. Yet such policies remain disconnected from established benchmarks, leaving the performance of large-scale RL for driving on standardized evaluations unknown. We present BehaviorBench -- a comprehensive test suite that closes this gap along three axes: Evaluation, Complexity, and Behavior Diversity. In terms of Evaluation, we provide an interface connecting PufferDrive to nuPlan, which, for the first time, enables policies trained via RL at scale to be evaluated on an established planning benchmark for autonomous driving. Complementarily, we offer an evaluation framework that allows planners to be benchmarked directly inside the PufferDrive simulation, at a fraction of the time. Regarding Complexity, we observe that today's standardized benchmarks are so simple that near-perfect scores are achievable by straight lane following with collision checking. We extract a meaningful, interaction-rich split from the Waymo Open Motion Dataset (WOMD) on which strong performance is impossible without multi-agent reasoning. Lastly, we address Behavior Diversity. Existing benchmarks commonly evaluate planners against a single rule-based traffic model, the Intelligent Driver Model (IDM). We provide a diverse suite of interactive traffic agents to stress-test policies under heterogeneous behaviors, beyond just using IDM. Overall, our benchmarking analysis uncovers the following insight: despite learning interactive behaviors in an emergent manner, policies trained via pure self-play under standard reward functions overfit to their training opponents and fail to generalize to other traffic agent behaviors. Building on this observation, we propose a hybrid planner that combines a PPO policy with a rule-based planner.",
      "short_abstract": "Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driving policies. Yet such policies remain disconnected from established benchmarks, leaving the performance of...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10034",
      "pdf_url": "https://arxiv.org/pdf/2605.10034.pdf"
    },
    {
      "id": "2605.10026",
      "anchor": "paper-2605-10026",
      "title": "MUSDA: Multi-source Multi-modality Unsupervised Domain Adaptive 3D Object Detection for Autonomous Driving",
      "authors": [
        "Xiaohu Lu",
        "Hamed Khatounabadi",
        "Hayder Radha"
      ],
      "abstract": "With the advancement of autonomous driving, numerous annotated multi-modality datasets have become available. This presents an opportunity to develop domain-adaptive 3D object detectors for new environments without relying on labor-intensive manual annotations. However, traditional domain adaptation methods typically focus on a single source domain or a single modality, limiting their effectiveness in multi-source, multi-modality scenarios. In this paper, we propose a novel framework for multi-source, multi-modality unsupervised domain adaptation in 3D object detection for autonomous driving. Given multiple labeled source domains and one unlabeled target domain, our framework first introduces hierarchical spatially-conditioned (HSC) domain classifiers, which jointly align features from both camera and LiDAR modalities at two distinct levels for each source-target domain pair. To effectively leverage information from multiple source domains, we construct a prototype graph between each pair of domains. Based on this, we develop a prototype graph weighted (PGW) multi-source fusion strategy to aggregate predictions from multiple source detection heads. Experimental results on three widely used 3D object detection datasets - Waymo, nuScenes, and Lyft - demonstrate that our proposed framework effectively integrates information across both modalities and source domains, consistently outperforming state-of-the-art methods.",
      "short_abstract": "With the advancement of autonomous driving, numerous annotated multi-modality datasets have become available. This presents an opportunity to develop domain-adaptive 3D object detectors for new environments without relying on labor-intensive manual...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 21,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10026",
      "pdf_url": "https://arxiv.org/pdf/2605.10026.pdf"
    },
    {
      "id": "2605.09972",
      "anchor": "paper-2605-09972",
      "title": "HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving",
      "authors": [
        "Zhongyu Xia",
        "Guanyu Zhu",
        "Guo Tang",
        "Wenhao Chen",
        "Yongtao Wang"
      ],
      "abstract": "End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving near-perfect scores on widely used open-loop and closed-loop benchmarks. This saturation does not mean that the problem has been solved; instead, it reveals that current benchmarks remain limited in scenario diversity, object variety, and the breadth of driving capabilities they evaluate. In particular, they lack sufficient long-tail scenarios involving rare but safety-critical objects and fail to assess advanced decision-making such as legal compliance, ethical reasoning, and emergency response. To address these gaps, we propose HiDrive, a new closed-loop benchmark for end-to-end autonomous driving that emphasizes long-tail scenarios and a richer evaluation of driving capabilities. HiDrive introduces a diverse set of rare objects and uncommon traffic situations, and expands evaluation from basic driving skills to more advanced capabilities, including rule compliance, moral reasoning, and context-dependent emergency maneuvers. Correspondingly, we extend previous collision-avoidance-centered metrics into a comprehensive evaluation system that encompasses collision and braking, traffic-rule compliance, and moral-reasoning indicators. Built on a more advanced physics engine, HiDrive provides physically realistic lighting and high-fidelity visual rendering, offering a more challenging and realistic testbed for assessing whether autonomous driving systems can handle the complexity of real-world deployment. The HiDrive software, source code, digital assets, and documentation are available at https://github.com/VDIGPKU/HiDrive.",
      "short_abstract": "End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving near-perfect scores on widely used open-loop and closed-loop benchmarks. This saturation does not mean that...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.09972",
      "pdf_url": "https://arxiv.org/pdf/2605.09972.pdf"
    },
    {
      "id": "2605.11205",
      "anchor": "paper-2605-11205",
      "title": "The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains",
      "authors": [
        "Jung Min Kang"
      ],
      "abstract": "Benchmark evaluation across AI and safety-critical domains overwhelmingly relies on simple averaging. We demonstrate that this practice produces substantially misleading rankings when two conditions co-occur: (1) the evaluation matrix is sparse and (2) items vary substantially in difficulty. Through controlled simulation experiments across four domains -- NLP (GLUE), clinical drug trials, autonomous vehicle safety, and cybersecurity -- we show that Spearman rank correlation $ρ$ between simple-average rankings and ground-truth rankings degrades from $ρ= 1.000$ at 100% coverage to $ρ= 0.809$ at 67% coverage with high difficulty heterogeneity (mean over 20 seeds). A standard two-parameter logistic (2PL) Item Response Theory (IRT) model maintains $ρ\\geq 0.996$ across all conditions. A 150-condition grid sweep over sparsity $S \\in [0, 0.70]$ and difficulty gap $D \\in [0.5, 5.0]$ confirms that ranking error forms a failure surface with a strong $S \\times D$ interaction ($γ_3 = +0.20$, $t = 13.05$), while IRT maintains $ρ\\geq 0.993$ throughout. We discuss implications for Physical AI benchmarking, where evaluation matrices are often incomplete and difficulty gaps are extreme.",
      "short_abstract": "Benchmark evaluation across AI and safety-critical domains overwhelmingly relies on simple averaging. We demonstrate that this practice produces substantially misleading rankings when two conditions co-occur: (1) the evaluation matrix is sparse and (2) items...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 17,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.11205",
      "pdf_url": "https://arxiv.org/pdf/2605.11205.pdf"
    },
    {
      "id": "2605.10653",
      "anchor": "paper-2605-10653",
      "title": "Embodied AI in Action: Insights from SAE World Congress 2026 on Safety, Trust, Robotics, and Real-World Deployment",
      "authors": [
        "Jan-Mou Li",
        "Paul Schmitt",
        "Wei Tong",
        "Majed Mohammed",
        "Akshay Chalana",
        "Arpan Kusari",
        "Edward Griffor"
      ],
      "abstract": "Embodied artificial intelligence is rapidly moving from research into real-world systems such as autonomous vehicles, mobile robots, and industrial machines. As these systems become more capable of perceiving, deciding, and acting in dynamic environments, they also introduce new challenges in safety, trust, governance, and operational reliability. This white paper summarizes key insights from the SAE World Congress 2026 panel session \\textit{Embodied AI in Action}, which brought together experts from automotive, robotics, artificial intelligence, and safety engineering. The discussion highlighted the need to treat embodied AI as a systems challenge requiring engineering rigor, lifecycle governance, human-centered design, and evolving standards. The paper provides practical perspectives for executives, policymakers, and technical leaders seeking to adopt embodied AI responsibly. The panel reached broad agreement that long-term success will depend not only on advances in AI capability, but equally on safe and trustworthy deployment.",
      "short_abstract": "Embodied artificial intelligence is rapidly moving from research into real-world systems such as autonomous vehicles, mobile robots, and industrial machines. As these systems become more capable of perceiving, deciding, and acting in dynamic environments,...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10653",
      "pdf_url": "https://arxiv.org/pdf/2605.10653.pdf"
    },
    {
      "id": "2605.10505",
      "anchor": "paper-2605-10505",
      "title": "A Theory of Multilevel Interactive Equilibrium in NeuroAI",
      "authors": [
        "Zhe Sage Chen",
        "Quanyan Zhu"
      ],
      "abstract": "We propose a game-theoretic framework for adaptive multi-agent intelligent systems. Unlike classical game theory, which often treats strategies as primitive objects chosen by perfectly rational agents, the proposed framework provides a mathematical foundation for studying equilibrium in NeuroAI and can be viewed as an extension of game theory under relaxed assumptions, including partial observability, bounded computation, and uncertainty. At its core, Multilevel Interactive Equilibrium (MIE) generalizes the classical Nash equilibrium to intelligent systems with internal computation. Rather than being defined solely at the level of observable behavior, equilibrium emerges when neural learning dynamics, cognitive representations, and behavioral strategies mutually stabilize between interacting agents. This framework applies uniformly to interactions between two biological brains, two artificial agents, or hybrid human-AI systems. We discuss applications of multilevel game theory to human-autonomous vehicle driving, human-machine interaction, human-large language model (LLM) interaction, and computational psychiatry. We also outline experimental strategies and computational methods for estimating MIE and discuss challenges and prospects for future research.",
      "short_abstract": "We propose a game-theoretic framework for adaptive multi-agent intelligent systems. Unlike classical game theory, which often treats strategies as primitive objects chosen by perfectly rational agents, the proposed framework provides a mathematical foundation...",
      "published": "2026-05-11",
      "updated": "2026-05-11",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Human Factors & Policy: abstract: human-machine interaction",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 100,
      "arxiv_url": "https://arxiv.org/abs/2605.10505",
      "pdf_url": "https://arxiv.org/pdf/2605.10505.pdf"
    },
    {
      "id": "2605.09902",
      "anchor": "paper-2605-09902",
      "title": "Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment",
      "authors": [
        "Haobo Wang",
        "Xiaorong Ma",
        "Weiqi Luo",
        "Xiaojun Jia",
        "Jiwu Huang"
      ],
      "abstract": "Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risks in safety-critical scenarios such as autonomous driving and medical diagnosis. This makes transfer-based targeted attacks crucial for understanding and improving black-box MLLM robustness. Existing transfer-based targeted attack methods typically rely on the final global features of the surrogate encoder and anchor optimization to original-resolution target crops, leading to their limited transferability and robustness. To address these challenges, we propose Progressive Resolution Processing and Adaptive Feature Alignment (PRAF-Attack), a targeted transfer-based attack framework that integrates multi-scale global semantic guidance with robust intermediate-layer local alignment. Unlike prior methods that align only the surrogate encoder's final layer, we design an adaptive feature alignment strategy that leverages intermediate representations to enhance transferability. Specifically, we introduce an adaptive intermediate layer selection mechanism to identify transferable hierarchical features across surrogate ensembles via gradient consistency, along with an adaptive patch-level optimization strategy that preserves highly correlated local regions through efficient patch filtering. To overcome the reliance on fixed original-resolution target crops, we propose a progressive resolution processing strategy that gradually refines optimization from coarse to fine, enabling the attack to better exploit target information at multiple scales and achieve stronger transferability. We evaluate PRAF-Attack on a diverse suite of black-box MLLMs, including six open-source models and six closed-source commercial APIs. Compared with seven state-of-the-art targeted attack baselines, the proposed PRAF-Attack consistently achieves superior transferability.",
      "short_abstract": "Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risks in safety-critical scenarios such as autonomous driving and medical diagnosis. This makes transfer-based...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 22,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09902",
      "pdf_url": "https://arxiv.org/pdf/2605.09902.pdf"
    },
    {
      "id": "2605.09774",
      "anchor": "paper-2605-09774",
      "title": "DRIVE-C: A Controlled Corruption Dataset for Autonomous Driving",
      "authors": [
        "Shiva Aher"
      ],
      "abstract": "DRIVE-C is a controlled corruption dataset designed to evaluate visual perception robustness in autonomous driving systems. It is built from real-world forward-facing driving videos collected across daytime, nighttime, urban, rural, freeway, and parking environments. Clean clips are anonymized via localized face and license plate blurring, then transformed with physics-inspired synthetic degradations. The dataset contains 10 clean clips and 600 corrupted clips spanning 12 camera degradation types across five severity levels, with per-clip metadata and Global Sensor Health Index (GSHI) annotations. DRIVE-C supports robustness benchmarking, degradation-aware modeling, uncertainty estimation, out-of-distribution (OOD) detection, and sensor health monitoring for Advanced Driver Assistance Systems (ADAS). By providing pixel-aligned clean and degraded video clips with fully reproducible corruption parameters, DRIVE-C offers a structured testbed for studying perception reliability under controlled camera degradation.",
      "short_abstract": "DRIVE-C is a controlled corruption dataset designed to evaluate visual perception robustness in autonomous driving systems. It is built from real-world forward-facing driving videos collected across daytime, nighttime, urban, rural, freeway, and parking...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: camera",
          "Safety, Security & Verification: abstract: robustness",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09774",
      "pdf_url": "https://arxiv.org/pdf/2605.09774.pdf"
    },
    {
      "id": "2605.09701",
      "anchor": "paper-2605-09701",
      "title": "DriveFuture: Future-Aware Latent World Models for Autonomous Driving",
      "authors": [
        "Yufeng Hong",
        "Xiaotian Zhou",
        "Yingyan Li",
        "Xiangpo Zhou",
        "Lin Liu",
        "Yadan Luo",
        "Shaoqing Xu",
        "Lei Yang",
        "Ziying Song"
      ],
      "abstract": "Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat future latent states as prediction targets or auxiliary signals, rather than directly conditioning trajectory planning. This can entangle current and future features in latent space. In this work, we propose DriveFuture, a future-aware latent world modeling framework for autonomous driving that explicitly learns planning-oriented foresight by conditioning the current latent state modeling process on future world states. Specifically, during training, the model first predicts future latent world states from the current latent state and ego action, and then refines the prediction against the ground-truth future latent state via cross-attention. The resulting future-aware latent serves as an explicit condition for a diffusion-based trajectory planner. During inference, DriveFuture conditions on the predicted future latent state instead of the ground-truth future state. DriveFuture achieves SOTA performance on the public NAVSIM benchmarks, reaching \\textbf{55.5} EPDMS on NAVSIM-v2 {\\textcolor{blue}{\\textit{navhard}}}, \\textbf{89.9} EPDMS on NAVSIM-v2 {\\textcolor{blue}{\\textit{navtest}}}, and \\textbf{90.7} PDMS on NAVSIM-v1 {\\textcolor{blue}{\\textit{navtest}}}, respectively. These results suggest that the key to latent world modeling lies not merely in simulating future states, but more importantly in conditioning current decision-making on future states. Notably, as of April 2026, DriveFuture ranks \\textbf{1st} on the \\href{https://huggingface.co/spaces/AGC2025/e2e-driving-navhard}{NAVSIM-v2 {\\textcolor{blue}{\\textit{navhard}}}} leaderboard and achieves SOTA performance on \\href{https://huggingface.co/spaces/AGC2024-P/e2e-driving-navtest}{NAVSIM-v1 {\\textcolor{blue}{\\textit{navtest}}}}.",
      "short_abstract": "Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat future latent states as prediction targets or auxiliary signals, rather than directly conditioning...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 18,
        "margin": 6,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09701",
      "pdf_url": "https://arxiv.org/pdf/2605.09701.pdf"
    },
    {
      "id": "2605.09619",
      "anchor": "paper-2605-09619",
      "title": "GSMap: 2D Gaussians for Online HD Mapping",
      "authors": [
        "Zhenxuan Zeng",
        "Lingxuan Wang",
        "Sheng Yang",
        "Yanan He",
        "Mingxia Chen",
        "Wei Suo",
        "Peng Wang"
      ],
      "abstract": "Accurate High-Definition (HD) map construction is critical for autonomous driving, yet existing methods face a fundamental trade-off: vectorization-based approaches preserve topology but struggle with geometric fidelity, while rasterization-based approaches enable precise geometric supervision but produce unstructured outputs. To bridge this gap, we propose GSMap, a novel framework that unifies both paradigms via a learnable 2D Gaussian representation. Each map element is modeled as an ordered sequence of 2D Gaussians, whose centers correspond to the vertices of the vectorized polyline/polygon. This formulation enables simultaneous optimization through: (1) Differentiable rasterization that enforces pixel-level geometric constraints, and (2) Topology-aware vectorization that maintains structural regularity. Experiments on both nuScenes and Argoverse2 demonstrate that our Gaussian-based representation effectively unifies geometric and topological learning, achieving significant performance improvements and demonstrating strong compatibility with existing HD mapping architectures. Code will be available at https://github.com/peakpang/GSMap",
      "short_abstract": "Accurate High-Definition (HD) map construction is critical for autonomous driving, yet existing methods face a fundamental trade-off: vectorization-based approaches preserve topology but struggle with geometric fidelity, while rasterization-based approaches...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 16,
        "evidence": [
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09619",
      "pdf_url": "https://arxiv.org/pdf/2605.09619.pdf"
    },
    {
      "id": "2605.09425",
      "anchor": "paper-2605-09425",
      "title": "AtteConDA: Attention-Based Conflict Suppression in Multi-Condition Diffusion Models and Synthetic Data Augmentation",
      "authors": [
        "Shogo Noguchi"
      ],
      "abstract": "Recent conditional image generation methods can improve controllability by generating images that are faithful to conditions such as sketches, human poses, segmentation maps, and depth. By applying these techniques to image augmentation while preserving annotations, generated images can be used as additional training data and can improve recognition performance. However, for high-level driving tasks such as traffic-rule extraction and driving-behavior understanding, simply using annotations as conditions is insufficient. Instead, images must be augmented while preserving the detailed high-level structure of the original scene. One possible solution is to use multiple conditions so that generated images retain diverse structural cues after generation. However, when multiple conditions are used, conflicts among conditions can prevent reliable structure preservation. In this work, we input semantic segmentation, depth, and edges extracted from the original image into a multi-condition image generation model, thereby providing rich structural information as conditions. We further propose a modeling approach for handling conflicts among multiple conditions and show that it enables image generation with stronger structural preservation. We also build a generation framework and evaluation protocol for driving tasks, establishing a basis for comparison with prior and future models. As a result, this work contributes to image generation research by addressing condition conflicts in multi-condition generation and provides an important step toward mitigating data scarcity in high-level autonomous-driving tasks.",
      "short_abstract": "Recent conditional image generation methods can improve controllability by generating images that are faithful to conditions such as sketches, human poses, segmentation maps, and depth. By applying these techniques to image augmentation while preserving...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Perception & Sensor Fusion: abstract: scene segmentation"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09425",
      "pdf_url": "https://arxiv.org/pdf/2605.09425.pdf"
    },
    {
      "id": "2605.09910",
      "anchor": "paper-2605-09910",
      "title": "CloudEmu: A Trace-Driven Cloud-Native Emulation Testbed for Vehicle Video Uplink over Cellular Networks",
      "authors": [
        "Takashi Torii",
        "Soto Anno",
        "Masaki Okada",
        "Takuma Tsubaki",
        "Nobuhiro Azuma",
        "Takuya Tojo"
      ],
      "abstract": "We present CloudEmu, a trace-driven, cloud-native cellular-emulation testbed for vehicle video uplink communication. Reliable, low-latency video uplink over cellular networks is essential for remote monitoring of autonomous vehicles. However, existing testbeds fall into two extremes. Physical-vehicle platforms provide realism but are costly and make validation under identical network conditions difficult, whereas simulations are inexpensive and reproducible but generally cannot replay field-measured end-to-end performance dynamics without substantial calibration or readily run production video-uplink stacks. A software-defined, cloud-native emulation approach can combine the fidelity of trace-driven replay with the agility and scalability that network softwarization principles offer. To this end, we propose CloudEmu that replays time-synchronized cellular and position traces, collected once from vehicles, on commodity Linux-based virtual vehicle and video-receiver nodes. A Linux-based emulation framework couples traffic replay with position replay, tying network dynamics to each point along the route and enabling repeatable, route-aware experiments without repeated on-road trials. Our demo deploys a production-grade video-uplink stack on CloudEmu, allowing attendees to experience low-cost, repeatable trials and controlled comparisons under identical replayed network conditions.",
      "short_abstract": "We present CloudEmu, a trace-driven, cloud-native cellular-emulation testbed for vehicle video uplink communication. Reliable, low-latency video uplink over cellular networks is essential for remote monitoring of autonomous vehicles. However, existing...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 7,
        "evidence": [
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09910",
      "pdf_url": "https://arxiv.org/pdf/2605.09910.pdf"
    },
    {
      "id": "2605.09481",
      "anchor": "paper-2605-09481",
      "title": "TSNBench: Benchmarking LLM Proficiency in Time-Sensitive Networking",
      "authors": [
        "Rubi Debnath",
        "Daniel Bujosa Mateu",
        "Luxi Zhao",
        "Silviu S. Craciunas",
        "Paul Pop",
        "Sebastian Steinhorst"
      ],
      "abstract": "We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802.1 standards for deterministic communication with bounded latency in safety-critical domains such as autonomous vehicles, aviation, defense, and industrial automation. While LLMs have been extensively evaluated on general knowledge tasks, their capabilities in safety-critical networking domains remain largely unexplored. TSNBench comprises 939 expert-validated multiple-choice questions (MCQs) covering diverse TSN mechanisms, along with 100 open-ended Worst-Case Delay (WCD) computation tasks for Credit-Based Shaper (CBS) and Cyclic Queuing and Forwarding (CQF) across varying network topologies and traffic conditions. MCQ answers are validated by domain experts, and open-ended ground truth WCD values are computed using a verified Network Calculus (NC) solver for CBS and closed-form mathematical upper bounds for CQF. We evaluate 16 LLMs and find that although models achieve 67 to 95% accuracy on MCQs, they fail substantially on open-ended WCD computation. For CBS, only GPT-5 achieves a Mean Absolute Percentage Error (MAPE) of 36.2%, meaning its predicted WCD deviates by 36.2% of the actual TSN flow delay on average, while most models exceed 80%. For CQF, the best model achieves 41.8% MAPE, with most models clustering between 80% and 100%. Such errors are large relative to TSN latency budgets and can lead to violations of real-time constraints and unsafe configurations. TSNBench demonstrates that MCQ benchmarks may overestimate LLM capabilities in safety-critical networking domains.",
      "short_abstract": "We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802.1 standards for deterministic communication with bounded latency in safety-critical domains such as...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 9,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09481",
      "pdf_url": "https://arxiv.org/pdf/2605.09481.pdf"
    },
    {
      "id": "2605.09376",
      "anchor": "paper-2605-09376",
      "title": "Mismatch-Aware Adaptive Constraint Tightening for Bicycle-Model Trajectory Optimization",
      "authors": [
        "Lingxue Lyu",
        "Zihui Liu"
      ],
      "abstract": "Trajectory optimization for autonomous vehicles usually relies on the kinematic bicycle model because of its computational simplicity. However, when the planned trajectory is executed under the true vehicle dynamics, which include lateral slip, tire stiffness and yaw-lateral coupling, safety constraints can be violated owing to the model mismatch. In this paper, we make three theoretical contributions. First, we derive a characteristic speed $v_c=\\sqrt{C_αL/M}$ which separates two different mismatch regimes: below $v_c$ the dynamic bicycle initially oversteers inward (safe); above $v_c$ it understeers outward (safety-critical). Second, we prove that the peak outward deviation $\\varepsilon^*$ follows a $T^2$ horizon scaling whose coefficient transitions between a transient bound $\\frac{1}{2}(v^2-v_c^2)κ$ and a steady-state bound. Third, we obtain a simulation-free analytical coefficient $a_2^{\\mathrm{anal}}=\\frac{1}{2}(1-v_c^2/v_{\\max}^2)T^2$ that is computable from vehicle parameters and the planning horizon alone. Putting these together, we propose Mismatch-Aware Adaptive Constraint Tightening (MACT), $ε(v,κ)=a_2 v^2|κ|$, which replaces a fixed worst-case margin by a state-dependent one that is large at high speed/curvature but nearly zero on gentle paths. Eight numerical experiments confirm the scaling laws. MACT reaches 100% safety with 84% less wasted margin than a fixed-margin baseline on the 2-DOF vehicle, extends to a nonlinear leaning bicycle, and in a closed-loop direct-shooting MPC comparison it cuts the applied margin by 34% compared with tube MPC while keeping the same safety.",
      "short_abstract": "Trajectory optimization for autonomous vehicles usually relies on the kinematic bicycle model because of its computational simplicity. However, when the planned trajectory is executed under the true vehicle dynamics, which include lateral slip, tire stiffness...",
      "published": "2026-05-10",
      "updated": "2026-05-10",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 6,
        "evidence": [
          "abstract: model predictive control",
          "abstract: vehicle dynamics"
        ]
      },
      "recency": "archive",
      "age_days": 101,
      "arxiv_url": "https://arxiv.org/abs/2605.09376",
      "pdf_url": "https://arxiv.org/pdf/2605.09376.pdf"
    },
    {
      "id": "2605.08975",
      "anchor": "paper-2605-08975",
      "title": "Latency Analysis and Optimization of Alpamayo 1 via Efficient Trajectory Generation",
      "authors": [
        "Yunseong Jeon",
        "Namcheol Lee",
        "Yoonsu Lee",
        "Jangwoon Park",
        "Sol Ahn",
        "Jong-Chan Kim",
        "Seongsoo Hong"
      ],
      "abstract": "Reasoning-based end-to-end (E2E) autonomous driving has recently emerged as a promising approach to improving the interpretability of driving decisions as it can generate human-readable reasoning together with predicted trajectories. Such approaches commonly generate multiple trajectories to capture diverse future behaviors, and they fall into two categories: (1) multi-reasoning, where one reasoning sequence is generated per trajectory, and (2) single-reasoning, where a single reasoning is shared across all trajectories. The former offers richer diversity at the cost of redundant computation, while the latter is more efficient but is often assumed to sacrifice diversity. Alpamayo 1, a representative system, adopts the multi-reasoning approach and achieves competitive trajectory prediction performance. However, the efficiency of this design remains largely unexplored, making it a well-motivated subject for investigation. In this paper, we systematically analyze and improve Alpamayo 1 in two ways. First, we reduce inference latency while preserving trajectory diversity by redesigning Alpamayo 1 into a single-reasoning system. Through extensive experiments, we find that replacing multi-reasoning with single-reasoning does not meaningfully degrade trajectory diversity. Second, we accelerate diffusion-based action generation by eliminating inter-block overhead arising from unnecessary copy operations and inefficient kernel execution. Through closed-loop and open-loop experiments, we validate both optimizations, demonstrating a 69.23% reduction in inference latency while maintaining trajectory diversity and prediction quality. These results highlight the importance of jointly analyzing system architecture and runtime execution to improve the efficiency of reasoning-based E2E AD systems.",
      "short_abstract": "Reasoning-based end-to-end (E2E) autonomous driving has recently emerged as a promising approach to improving the interpretability of driving decisions as it can generate human-readable reasoning together with predicted trajectories. Such approaches commonly...",
      "published": "2026-05-09",
      "updated": "2026-05-09",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Hardware / Real-Time",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 13,
        "margin": 1,
        "evidence": [
          "abstract: decision-making",
          "title: trajectory generation"
        ]
      },
      "recency": "archive",
      "age_days": 102,
      "arxiv_url": "https://arxiv.org/abs/2605.08975",
      "pdf_url": "https://arxiv.org/pdf/2605.08975.pdf"
    },
    {
      "id": "2605.08839",
      "anchor": "paper-2605-08839",
      "title": "Cross-Sample Relational Fusion: Unifying Domain Generalization and Class-Incremental Learning",
      "authors": [
        "Zhen-Hao Xie",
        "Yan Wang",
        "Hao Sun",
        "Han-Jia Ye",
        "De-Chuan Zhan",
        "Da-Wei Zhou"
      ],
      "abstract": "Class-Incremental Learning (CIL) requires a learning system to learn new classes while retaining previously learned knowledge. However, in real-world scenarios such as autonomous driving, a system trained on urban roads in sunny weather may later need to operate in rural or highway environments with different traffic patterns and weather conditions. This requires the model not only to overcome catastrophic forgetting, but also to effectively handle domain shifts. In this paper, we propose CrOss-sample Relational Fusion (CORF), a unified framework to address domain shift and catastrophic forgetting simultaneously. To enhance generalizability, we perform selective refinement of training samples by leveraging spatial contribution maps to highlight semantically informative regions. Furthermore, we incorporate predictive confidence to adaptively weigh samples, thereby facilitating the learning of domain-agnostic representations. To alleviate forgetting, we propose a cascaded distillation framework that captures cross-sample relational dependencies across multiple feature hierarchies, enabling multi-grained knowledge transfer from previous tasks. CORF can be seamlessly integrated into existing CIL algorithms to enhance their generalizability, achieving competitive performance across various benchmark datasets. Code is available at https://github.com/LAMDA-CL/TMM26-CORF .",
      "short_abstract": "Class-Incremental Learning (CIL) requires a learning system to learn new classes while retaining previously learned knowledge. However, in real-world scenarios such as autonomous driving, a system trained on urban roads in sunny weather may later need to...",
      "published": "2026-05-09",
      "updated": "2026-05-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 102,
      "arxiv_url": "https://arxiv.org/abs/2605.08839",
      "pdf_url": "https://arxiv.org/pdf/2605.08839.pdf"
    },
    {
      "id": "2605.08830",
      "anchor": "paper-2605-08830",
      "title": "VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving",
      "authors": [
        "Rui Zhao",
        "Jianlin Yu",
        "Zhenhai Gao",
        "Jiaqiao Liu",
        "Fei Gao"
      ],
      "abstract": "End-to-end autonomous driving requires models to understand traffic scenes, infer driving intent, and generate executable motion plans. Recent vision-language-action (VLA) models inherit semantic priors from large-scale vision-language pretraining, yet still face a coupling trade-off: fully shared backbones preserve multimodal interaction but may entangle language reasoning and trajectory prediction, whereas decou pled reasoning-action pipelines reduce task conflict but weaken semantic-motion coupling. We propose VECTOR-DRIVE, a tightly coupled VLA framework built on Qwen2.5-VL-3B. VECTOR-DRIVE keeps all tokens coupled through shared self attention and routes feed-forward computation according to token semantics. Vision and language tokens are processed by a Vision-Language Expert to preserve semantic priors, while target-point, ego-state, and noisy action tokens are routed to a Trajectory Expert for motion-specific computation. On the action-token pathway, a flow-matching planner refines noisy action tokens into future waypoints and speed profiles. This design couples semantic reasoning and motion planning within a single multimodal Transformer while separating task-specific FFN computation. On Bench2Drive, VECTOR-DRIVE achieves 88.91 Driving Score and outperforms representative end-to end and VLA-based baselines. Qualitative results and ablations further validate the benefits of shared attention, semantic-aware expert routing, progressive training, and flow-based action de coding.",
      "short_abstract": "End-to-end autonomous driving requires models to understand traffic scenes, infer driving intent, and generate executable motion plans. Recent vision-language-action (VLA) models inherit semantic priors from large-scale vision-language pretraining, yet still...",
      "published": "2026-05-09",
      "updated": "2026-05-09",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 34,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 102,
      "arxiv_url": "https://arxiv.org/abs/2605.08830",
      "pdf_url": "https://arxiv.org/pdf/2605.08830.pdf"
    },
    {
      "id": "2605.08911",
      "anchor": "paper-2605-08911",
      "title": "Unified Modeling of Lane and Lane Topology for Driving Scene Reasoning",
      "authors": [
        "Han Li",
        "Yulu Gao",
        "Si Liu",
        "Yuhang Wang",
        "Bo Liu",
        "Beipeng Mu"
      ],
      "abstract": "Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements like lane centerlines and their topology. Existing lane topology reasoning methods typically follow a reasoning-by-detection paradigm, where lane topological relationships are primarily derived from lane detection results. In this paper, we propose an innovative method called Unified Modeling of Lane and Lane Topology (UniTopo), which represents the topological relationships between lanes as connected lanes, encompassing predecessor lanes, successor lanes, and their interconnections. This unified representation of lanes and lane topology allows us to simultaneously obtain both the positions and topological information of lanes within a shared perception pipeline, establishing a new paradigm for directly perceiving lane topology from original image features. We validate our method on the driving scene reasoning benchmark OpenLane-V2, which consists of two subsets, built based on Argoverse2 and nuScenes, respectively. Our method achieves TOP_ll of 30.1% and 31.8% on the two subsets, significantly surpassing the existing state-of-the-art method T^2SG by 6.0% and 8.6%.",
      "short_abstract": "Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements like lane centerlines and their topology. Existing lane topology reasoning methods typically follow a...",
      "published": "2026-05-09",
      "updated": "2026-05-09",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 6,
        "evidence": [
          "abstract: perception",
          "abstract: lane detection"
        ]
      },
      "recency": "archive",
      "age_days": 102,
      "arxiv_url": "https://arxiv.org/abs/2605.08911",
      "pdf_url": "https://arxiv.org/pdf/2605.08911.pdf"
    },
    {
      "id": "2605.08626",
      "anchor": "paper-2605-08626",
      "title": "Large Language Models over Networks: Collaborative Intelligence under Resource Constraints",
      "authors": [
        "Liangqi Yuan",
        "Wenzhi Fang",
        "Shiqiang Wang",
        "H. Vincent Poor",
        "Christopher G. Brinton"
      ],
      "abstract": "Large language models (LLMs) are transforming society, powering applications from smartphone assistants to autonomous driving. Yet cloud-based LLM services alone cannot serve a growing class of applications, including those operating under intermittent connectivity, sub-second latency budgets, data-residency constraints, or sustained high-volume inference. On-device deployment is in turn constrained by limited computation and memory. No single endpoint can deliver high-quality service across this spectrum. This article focuses on collaborative intelligence, a paradigm in which multiple independent LLMs distributed across device and cloud endpoints collaborate at the task level through natural language or structured messages. Such collaboration strives for superior response quality under heterogeneous resource constraints spanning computation, memory, communication, and cost across network tiers. We present collaborative inference along two complementary and composable dimensions: vertical device-cloud collaboration and horizontal multi-agent collaboration, which can be combined into hybrid topologies in practice. We then examine learning to collaborate, addressing the training of routing policies and the development of cooperative capabilities among LLMs. Finally, we identify open research challenges including scaling under resource heterogeneity and trustworthy collaborative intelligence.",
      "short_abstract": "Large language models (LLMs) are transforming society, powering applications from smartphone assistants to autonomous driving. Yet cloud-based LLM services alone cannot serve a growing class of applications, including those operating under intermittent...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 25,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.08626",
      "pdf_url": "https://arxiv.org/pdf/2605.08626.pdf"
    },
    {
      "id": "2605.08572",
      "anchor": "paper-2605-08572",
      "title": "ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction",
      "authors": [
        "Alen Mrdovic",
        "Qingze",
        "Liu",
        "Danrui Li",
        "Mathew Schwartz",
        "Kaidong Hu",
        "Sejong Yoon",
        "Mubbasir Kapadia",
        "Vladimir Pavlovic"
      ],
      "abstract": "Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time-critical settings like autonomous driving. Fast-sampling variants using DDIM and informed initial noise distributions partially alleviate this issue, but they either fail to achieve true single-step generation or are constrained by the chosen noise distribution. Consistency Models (CMs) offer high-quality one-step generation by mapping noise directly to data, but are difficult to train from scratch. We propose ECTraj, an enhanced CM pipeline with improved training and conditional generation for trajectory prediction. Our framework extends the student-teacher consistency training scheme: the student produces standard outputs, while the teacher explicitly fuses its predictions with parts of the ground truth to give stronger supervision. We also exploit CMs' direct denoising for top-K multi-shot generation during training. Combining conditional generation with this enhanced consistency objective yields faster inference and improved prediction accuracy, establishing competitive new benchmarks on the large-scale Argoverse 2 dataset.",
      "short_abstract": "Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time-critical settings like autonomous driving. Fast-sampling variants using DDIM and informed initial noise...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Benchmark",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 24,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.08572",
      "pdf_url": "https://arxiv.org/pdf/2605.08572.pdf"
    },
    {
      "id": "2605.08528",
      "anchor": "paper-2605-08528",
      "title": "SceneFactory: GPU-Accelerated Multi-Agent Driving Simulation with Physics-Based Vehicle Dynamics",
      "authors": [
        "Yicheng Zhu",
        "Yang Chen",
        "Tao Li",
        "Zilin Bian"
      ],
      "abstract": "Autonomous-driving simulators typically trade physical fidelity for scalable parallelism. Physics-based platforms such as CARLA and MetaDrive provide articulated vehicle dynamics and contact, but their non-vectorized interfaces make batched training difficult. GPU-batched systems such as Waymax and GPUDrive scale to hundreds of scenarios by replacing rigid-body physics with simplified kinematic models, omitting tire--road interaction, suspension, contact dynamics, and road-condition-dependent friction. We introduce SceneFactory, a GPU-vectorized platform for procedural scene construction, physics-based multi-agent simulation, and RL in autonomous-driving environments. Built on NVIDIA Isaac Sim + Isaac Lab, SceneFactory represents worlds and agents as batched tensors: control, observations, rewards, resets, and policy inference run as GPU tensor operations over the Isaac Lab tensor API. SceneFactory converts Waymo Open Motion Dataset road topologies into simulation-ready USD worlds, runs many worlds concurrently on one GPU, populates each with multiple articulated PhysX vehicles, and maps precipitation and road-surface type to PhysX material friction coefficients. With GPU vectorization, SceneFactory achieves up to 127$\\times$ higher throughput than a non-vectorized PhysX baseline on the same GPU and physics solver, reaching 19,250 controlled-agent simulation steps per second at 256 worlds $\\times$ 16 agents. Cross-simulator transfer reveals an asymmetric dynamics gap: physics-grounded RL policies transfer to a simplified kinematic bicycle model with 99.5% success, whereas reverse transfer drops to 47.3%. Under wet-road friction, friction-aware policies reduce mean peak DRAC from 58.7 to 27.8,m/s$^2$ without sacrificing goal reach. SceneFactory shows that scalable autonomous-driving training need not discard articulated rigid-body dynamics or physically grounded road-condition variation.",
      "short_abstract": "Autonomous-driving simulators typically trade physical fidelity for scalable parallelism. Physics-based platforms such as CARLA and MetaDrive provide articulated vehicle dynamics and contact, but their non-vectorized interfaces make batched training...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 18,
        "evidence": [
          "title: vehicle dynamics",
          "abstract: vehicle dynamics",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.08528",
      "pdf_url": "https://arxiv.org/pdf/2605.08528.pdf"
    },
    {
      "id": "2605.08293",
      "anchor": "paper-2605-08293",
      "title": "Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion",
      "authors": [
        "Yijing Wang",
        "Ruonan Li",
        "Qilin Wang",
        "Rongqiang Zhao",
        "Jie Liu"
      ],
      "abstract": "LiDAR-based semantic segmentation is essential for autonomous-driving perception, yet dense point-wise annotations are costly, and long-tailed outdoor scenes make small safety-critical objects difficult to discover without supervision. Existing unsupervised methods face three key challenges: they struggle to preserve small and sparsely observed objects under substantial scale variation, have difficulty enforcing intra-region consistency and inter-region discrimination during cross-modal transfer, and lack an efficient feature-preserving mechanism for contextual propagation over superpoint graphs. We therefore propose DDS, an unsupervised 3D semantic segmentation framework. First, a coarse-to-fine multi-granularity mask cascade provides complementary 3D region cues for objects across different scales, improving the preservation of small and sparsely observed objects. Second, region-guided multi-level distillation transfers self-supervised visual knowledge through point-level alignment, mask-level prototype alignment, and prototype-level contrastive learning, enhancing intra-region consistency and inter-region discrimination. Third, restart-based graph diffusion efficiently propagates contextual information among superpoints while anchoring the refined representation to the initial distilled features and avoiding explicit graph eigendecomposition. Experiments on real-world driving datasets show that DDS outperforms representative unsupervised baselines, improving oAcc, mAcc, and mIoU by up to 2.9%, 9.7%, and 4.1%, respectively. These results demonstrate the effectiveness and transferability of DDS for unsupervised 3D scene understanding in autonomous-driving scenarios.",
      "short_abstract": "LiDAR-based semantic segmentation is essential for autonomous-driving perception, yet dense point-wise annotations are costly, and long-tailed outdoor scenes make small safety-critical objects difficult to discover without supervision. Existing unsupervised...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 21,
        "evidence": [
          "abstract: perception",
          "title: scene segmentation",
          "abstract: scene segmentation",
          "abstract: LiDAR",
          "abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.08293",
      "pdf_url": "https://arxiv.org/pdf/2605.08293.pdf"
    },
    {
      "id": "2605.08084",
      "anchor": "paper-2605-08084",
      "title": "123D: Unifying Multi-Modal Autonomous Driving Data at Scale",
      "authors": [
        "Daniel Dauner",
        "Valentin Charraut",
        "Bastian Berle",
        "Tianyu Li",
        "Long Nguyen",
        "Jiabao Wang",
        "Changhui Jing",
        "Maximilian Igl",
        "Holger Caesar",
        "Boris Ivanovic",
        "Yiyi Liao",
        "Andreas Geiger",
        "Kashyap Chitta"
      ],
      "abstract": "The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain largely untapped. Each dataset adopts different 2D and 3D modalities, such as cameras, lidar, ego states, annotations, traffic lights, and HD maps, with different rates and synchronization schemes. They come in fragmented formats requiring complex dependencies that cannot natively coexist in the same development environment. Further, major inconsistencies in annotation conventions prevent training or measuring generalization across multiple datasets. We present 123D, an open-source framework that unifies such multi-modal driving data through a single API. To handle synchronization, we store each modality as an independent timestamped event stream with no prescribed rate, enabling synchronous or asynchronous access across arbitrary datasets. Using 123D, we consolidate eight real-world driving datasets spanning 3,300 hours and 90,000 kilometers, together with a synthetic dataset with configurable collection scripts, and provide tools for data analysis and visualization. We conduct a systematic study comparing annotation statistics and assessing each dataset's pose and calibration accuracy. Further, we showcase two applications 123D enables: cross-dataset 3D object detection transfer and reinforcement learning for planning, and offer recommendations for future directions. Code and documentation are available at https://github.com/kesai-labs/py123d.",
      "short_abstract": "The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain largely untapped. Each dataset adopts different 2D and 3D modalities, such as cameras, lidar, ego states,...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 4,
        "evidence": [
          "abstract: object detection",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.08084",
      "pdf_url": "https://arxiv.org/pdf/2605.08084.pdf"
    },
    {
      "id": "2605.07910",
      "anchor": "paper-2605-07910",
      "title": "One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction",
      "authors": [
        "Yulong Chen",
        "Xiaoyun Dong",
        "Haoyu Zhang",
        "Zongxian Yang",
        "Lewei Xie",
        "Xinke Li",
        "Yifan Zhang",
        "Kai Wang",
        "Jianping Wang"
      ],
      "abstract": "Reconstructing dynamic scenes from Vehicle-to-Infrastructure Cooperative Autonomous Driving (VICAD) data is fundamentally complicated by temporal asynchrony: vehicle and infrastructure cameras operate on independent clocks, capturing the same dynamic agent such as cars and pedestrians at different physical times. Existing Gaussian Scene Graph methods implicitly assume synchronized observations and assign a single pose per agent per frame, which is an assumption that breaks in cooperative settings, where the resulting gradient conflicts cause severe ghosting on dynamic agents. We identify this as a representation-level failure, not an optimization artifact: we prove that any single-timeline formulation incurs an irreducible photometric loss scaling quadratically with agent velocity and cross-source time offset. To resolve this, we propose Dust (DecoUpled Spatio-Temporal) Gaussian Scene Graph for 4D Cooperative Driving Reconstruction. DUST Gaussian Scene Graph shares a canonical Gaussian set per agent for appearance consistency, while maintaining decouple pose trajectories aligned to each source's true capture timestamps. We prove that this decoupling enables the pose-gradient kernel block-diagonal, eliminating cross-source interference entirely. To make Dust practical, we further introduce a static anchor-based pose correction pipeline that corrects spatio misalignment between vehicle and infrastructure annotations, and a pose-regularized joint optimization scheme that prevents trajectory jitter and drift during early training. On 26 sequences from V2X-Seq, DUST achieves state-of-the-art performance, improving dynamic-area PSNR by 3.2 dB over the strongest baseline and reducing Fréchet Video Distance by 37.7%, with keeping robustness under larger temporal asynchrony.",
      "short_abstract": "Reconstructing dynamic scenes from Vehicle-to-Infrastructure Cooperative Autonomous Driving (VICAD) data is fundamentally complicated by temporal asynchrony: vehicle and infrastructure cameras operate on independent clocks, capturing the same dynamic agent...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 14,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.07910",
      "pdf_url": "https://arxiv.org/pdf/2605.07910.pdf"
    },
    {
      "id": "2605.07549",
      "anchor": "paper-2605-07549",
      "title": "Probabilistic Object Detection with Conformal Prediction",
      "authors": [
        "Christopher Ries",
        "Moussa Kassem Sbeyti",
        "Nicolas Bianco",
        "Nadja Klein"
      ],
      "abstract": "Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety-critical object detection. However, object detection introduces structured multi-output predictions, complicating the application of classical CP theory developed for single outputs. In addition, standard, unscaled CP produces fixed-width prediction intervals across inputs, leading to unnecessary width for low-uncertainty predictions. While scaled CP addresses this by adapting the interval width to an input-dependent uncertainty estimate, prior work has neither systematically compared unscaled and scaled CP for multi-class object detection, nor integrated CP with a complementary uncertainty quantification method in this setting. We fill this gap by: (i) applying CP coordinate-wise to bounding box corners with a Bonferroni correction for box-level guarantees; (ii) scaling the resulting intervals using per-prediction aleatoric uncertainty estimates derived from a probabilistic object detector trained with loss attenuation, evaluated in uncalibrated and two calibrated variants; (iii) extending to a two-step pipeline that constructs prediction sets for the class using RAPS and conditions the conformalized bounding boxes on the predicted class set. Across three autonomous driving datasets (KITTI, BDD, CODA), including a cross-domain setting under distribution shift, scaled CP consistently improves interval sharpness over unscaled CP, achieving up to 19% higher IoU and 39% lower interval scores, without sacrificing coverage. Class-wise calibration further improves coverage for both variants with a negligible effect on sharpness. Together, these improvements yield more actionable uncertainty estimates for real-time, real-world object detection.",
      "short_abstract": "Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety-critical object detection. However,...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.07549",
      "pdf_url": "https://arxiv.org/pdf/2605.07549.pdf"
    },
    {
      "id": "2605.07370",
      "anchor": "paper-2605-07370",
      "title": "MORPH-U: Multi-Objective Resilient Motion Planning for V2X-Enabled Autonomous Driving in High-Uncertainty Environments via Simulation",
      "authors": [
        "Shih-Yu Lai"
      ],
      "abstract": "V2X can warn an autonomous vehicle about hazards beyond line-of-sight, but it also brings uncertainty: messages may be delayed, dropped, or even forged. Meanwhile, map knowledge may change during a trip, forcing the vehicle to replan under tight real-time budgets. This paper studies how to make motion planning and low-level control robust to such uncertain, event-driven updates. We present MORPH-U, a CARLA-based closed-loop stack that fuses LiDAR/radar/camera with V2X (CAM/DENM) into a Local Dynamic Map (LDM) and triggers Hybrid-A* replanning when validated hazards or map changes affect the planned route. We expose the planning/control trade-offs via a multi-objective formulation over tracking error, safety margin (minimum TTC), responsiveness, and smoothness, and select operating points using Pareto-frontier analysis. To avoid unsafe replanning from faulty V2X triggers, MORPH-U adds a lightweight Byzantine-inspired acceptance gate that combines a quorum rule with an on-board sensor veto. Experiments in dynamic CARLA scenarios show that V2X-augmented LDM improves downstream safety, Pareto tuning provides controllable accuracy-comfort trade-offs, and the gate prevents replanning under saturated false-DENM injection ($p_{\\text{attack}}=1.0$).",
      "short_abstract": "V2X can warn an autonomous vehicle about hazards beyond line-of-sight, but it also brings uncertainty: messages may be delayed, dropped, or even forged. Meanwhile, map knowledge may change during a trip, forcing the vehicle to replan under tight real-time...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 32,
        "margin": 5,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.07370",
      "pdf_url": "https://arxiv.org/pdf/2605.07370.pdf"
    },
    {
      "id": "2605.07356",
      "anchor": "paper-2605-07356",
      "title": "UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition",
      "authors": [
        "Shuai Zhang",
        "Zhecheng Shi",
        "Zhuxiao Li",
        "Jing Ou",
        "Tengxi Wang",
        "Yuan Liu",
        "Wufan Zhao"
      ],
      "abstract": "Semantic segmentation of large-scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse sampling pattern of LiDAR and the view-dependent geometric distortion in image observations complicate cross-modal alignment and hinder stable fusion. Inspired by the fact that 2D images captured by cameras are representations of the 3D world, we recognize that the features learned from 2D and 3D segmentation share some common semantics, while other aspects remain modality-specific. This insight motivates a unified multimodal framework for joint 2D-3D semantic segmentation. We combine a SAM-based vision encoder with a SPTNet-based geometric encoder to extract complementary semantic and geometric representations. The resulting features from both modalities are explicitly decomposed into shared and private subspaces, where the shared components summarize semantic factors common to both domains, and the private components preserve properties that are unique to each modality. A lightweight attention-based fusion module aggregates the shared features into a consistent cross-modal representation, and a regularized training objective ensures both semantic alignment and subspace independence. Experiments on the SemanticKITTI and nuScenes benchmarks demonstrate consistent improvements in segmentation accuracy over representative multimodal baselines, accompanied by competitive computational efficiency. Cross-domain evaluation on nuScenes USA-Singapore shows stable performance under distribution shifts, demonstrating strong generalization. The implementation code is publicly available at: https://github.com/shuaizhang69/UniD-Shift.",
      "short_abstract": "Semantic segmentation of large-scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse sampling pattern of LiDAR and the view-dependent geometric distortion in image observations complicate...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 24,
        "evidence": [
          "title: scene segmentation",
          "abstract: scene segmentation",
          "abstract: point cloud",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.07356",
      "pdf_url": "https://arxiv.org/pdf/2605.07356.pdf"
    },
    {
      "id": "2605.07326",
      "anchor": "paper-2605-07326",
      "title": "GEM: Generating LiDAR World Model via Deformable Mamba",
      "authors": [
        "Yang Wu",
        "Zhaojiang Liu",
        "Qiang Meng",
        "Youquan Liu",
        "Renliang Weng",
        "Jianjun Qian",
        "Jian Yang",
        "Jin Xie"
      ],
      "abstract": "World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, progress in LiDAR-based world models has lagged behind those built on camera videos or occupancy data, primarily due to two core challenges: the inherent disorder of LiDAR point clouds and the difficulty of distinguishing dynamic objects from static structures. To address these issues, we propose GEM: a Generative LiDAR world model that leverages deformable mamba architecture, significantly improving fidelity and imaginative capability. Specifically, leveraging the structural similarity between sequential laser scanning and Mamba's processing mechanism, we first tokenize LiDAR sweeps into compact representations via a custom LiDAR scene tokenizer. After unsupervised disentanglement of tokenized features via a dynamic-static separator, a tri-path deformable Mamba is introduced to perform selective scanning and adaptive gating fusion over the disentangled features, leading to enhanced spatial-temporal understanding of the world evolution. Optionally, a planner and a BEV layout controller can be integrated to explore the model's capability for autonomous rollout and its potential to generate ``what-if\" scenarios. Extensive experiments show that GEM achieves state-of-the-art performances across diverse benchmarks and evaluation settings, demonstrating its superiority and effectiveness. Project page: https://github.com/wuyang98/GEM.",
      "short_abstract": "World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, progress in LiDAR-based world models has lagged behind those built on camera videos or occupancy data,...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 21,
        "margin": 5,
        "evidence": [
          "abstract: occupancy",
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.07326",
      "pdf_url": "https://arxiv.org/pdf/2605.07326.pdf"
    },
    {
      "id": "2605.21500",
      "anchor": "paper-2605-21500",
      "title": "A Task-Agnostic Algebraic Integrity Metric for Event-Camera Streams Toward SOTIF-Compliant Perception using Pearson Correlation Coefficient",
      "authors": [
        "Arthur de Miranda Neto"
      ],
      "abstract": "Event cameras have emerged as a high-bandwidth, low-latency sensing modality for safety-critical perception in automated driving systems (ADS), offering microsecond temporal resolution, 120-140 dB dynamic range, and intrinsic absence of motion blur. However, no task-agnostic quality metric currently operates directly on the asynchronous event stream: state-of-the-art proxies require a downstream task (e.g., detection accuracy, tracking error) to assess stream integrity, which is incompatible with the certification requirements of ISO 21448 (SOTIF) and ISO/PAS 8800:2024. The recent BiasBench benchmark (CVPR 2025) explicitly identifies this gap. This work proposes a unified algebraic framework that lifts the Pearson Correlation Coefficient (PCC), historically used in two prior works for redundancy filtering and ROI selection on frame-based images, to the three standard event representations: Time Surface, Event Frame, and Voxel Grid. The framework yields three metrics: (i) r-TS for stream integrity monitoring against an ego-motion-predicted Time Surface, (ii) r2-EF for adaptive ROI selection requiring only integer comparisons, and (iii) r-VG for temporal redundancy gating. A structural isomorphism is established between the contrast-threshold mechanism of the event camera (|Delta L| >= C) and the PCC-based change criterion, the three lifted metrics are formalized, and pipeline latency and information loss are analyzed symmetrically against the raw stream. Illustrative behavior of each metric is demonstrated on a procedural-synthetic event stream, generated by direct simulation of the emission model rather than drawn from any real or video-derived dataset, including a tunnel-dip integrity-anomaly scenario in which r_C drops from 0.93 (coherent flow) to below 0 (alarm). An explicit epistemic convention ([ESTABLISHED], [SOLID], [HYPOTH.], [OPEN]) delineates the status of every contribution.",
      "short_abstract": "Event cameras have emerged as a high-bandwidth, low-latency sensing modality for safety-critical perception in automated driving systems (ADS), offering microsecond temporal resolution, 120-140 dB dynamic range, and intrinsic absence of motion blur. However,...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 8,
        "evidence": [
          "abstract: safety",
          "title: SOTIF",
          "abstract: SOTIF"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.21500",
      "pdf_url": "https://arxiv.org/pdf/2605.21500.pdf"
    },
    {
      "id": "2605.07649",
      "anchor": "paper-2605-07649",
      "title": "Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models",
      "authors": [
        "Berkehan Ünal",
        "Hauke Dierend",
        "Dren Fazlija",
        "Christopher Plachetka"
      ],
      "abstract": "Over the last few years, research on autonomous systems has matured to such a degree that the field is increasingly well-positioned to translate research into practical, stakeholder-driven use cases across well-defined domains. However, for a wide-scale practical adoption of autonomous systems, adherence to safety regulations is crucial. Many regulations are influenced by the Operational Design Domain (ODD), which defines the specific conditions in which an autonomous agent can function. This is especially relevant for Automated Driving Systems (ADS), as a dependable perception of ODD elements is essential for safe implementation and auditing. Vision-language models (VLMs) integrate visual recognition and language reasoning, functioning without task-specific training data, which makes them suitable for adaptable ODD perception. To assess whether VLMs can function as zero-shot \"ODD sensors\" that adapt to evolving definitions, we contribute (i) an empirical study of zero-shot ODD classification and detection using four VLMs on a custom dataset and Mapillary Vistas, along with failure analyses; (ii) an ablation of zero-shot optimization strategies with a cost-performance overview; and (iii) a suite of reusable prompting templates with guidance for adaptation. Our findings indicate that definition-anchored chain-of-thought prompting with persona decomposition performs best, while other methods may result in reduced recall. Overall, our results pave the way for transparent and effective ODD-based perception in safety-critical applications.",
      "short_abstract": "Over the last few years, research on autonomous systems has matured to such a degree that the field is increasingly well-positioned to translate research into practical, stakeholder-driven use cases across well-defined domains. However, for a wide-scale...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "title: perception",
          "abstract: perception"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.07649",
      "pdf_url": "https://arxiv.org/pdf/2605.07649.pdf"
    },
    {
      "id": "2605.07743",
      "anchor": "paper-2605-07743",
      "title": "Efficient MILP-based Urban Network Traffic Control in Mixed Autonomy with Dynamic Saturation Rates",
      "authors": [
        "Muhammad Haris",
        "Claudio Roncoli"
      ],
      "abstract": "This paper introduces a novel control strategy to optimize urban network traffic in mixed autonomy settings, featuring Connected and Automated Vehicles (CAVs) alongside Human-Driven Vehicles (HDVs). Unlike previous control strategies, where the impact of driver behaviour of CAVs and HDVs is not explicitly considered, we propose a dynamic, queue-responsive saturation rate to account for autonomy-driven variations in traffic flow characteristics. The proposed method is based on an extended multi-commodity store-and-forward model to a mixed autonomy environment, integrating optimized routing for CAVs via infrastructure-linked connectivity, and signal timings at every signalized intersection. The problem is formulated as a Non-Convex Quadratic Program (NQP), which accounts for queue evolution, spillback, green time allocation, and CAVs routing. To enable computational efficiency for real-time applications, we transform the NQP into a sequence of convex subproblems, leveraging under- and over-estimators to reformulate it as a Mixed Integer Linear Program (MILP). Experimental results via microscopic simulations validate the efficiency and robustness of the proposed methodology. The results reflect that the proposed model outperforms the existing multi-commodity approach, thus demonstrating its potential for real-time traffic optimization in future urban mobility systems.",
      "short_abstract": "This paper introduces a novel control strategy to optimize urban network traffic in mixed autonomy settings, featuring Connected and Automated Vehicles (CAVs) alongside Human-Driven Vehicles (HDVs). Unlike previous control strategies, where the impact of...",
      "published": "2026-05-08",
      "updated": "2026-05-08",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 7,
        "evidence": [
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 103,
      "arxiv_url": "https://arxiv.org/abs/2605.07743",
      "pdf_url": "https://arxiv.org/pdf/2605.07743.pdf"
    },
    {
      "id": "2605.07195",
      "anchor": "paper-2605-07195",
      "title": "See Tomorrow, Act Today: Foresight-Driven Autonomous Driving",
      "authors": [
        "Bozhou Zhang",
        "Nan Song",
        "Yuang Wang",
        "Jiankang Deng",
        "Xiatian Zhu",
        "Li Zhang"
      ],
      "abstract": "Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous agents should instead imagine future scenes before deciding, just as human drivers mentally simulate ``what will happen next\" before acting. We introduce ForeSight, a foundation world model centric planning framework that reframes autonomous driving as anticipatory decision-making. Rather than treating world models as auxiliary components, ForeSight makes future scene imagination the primary driver of action prediction. Our approach operates in two stages: (1) generating plausible future visual worlds via a pretrained world model, and (2) planning actions conditioned on these imagined futures. This paradigm shift from ``what should I do now?\" to ``what will happen, and how should I respond?\" enables genuinely anticipatory rather than reactive planning. By grounding decisions in anticipated contexts rather than present observations alone, ForeSight navigates dynamic, interactive scenarios more effectively. Extensive experiments on NAVSIM and nuScenes demonstrate that explicit future imagination significantly outperforms previous state-of-the-art alternatives, validating our foresight-driven approach.",
      "short_abstract": "Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous agents should instead imagine future scenes before deciding, just as human...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 2,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.07195",
      "pdf_url": "https://arxiv.org/pdf/2605.07195.pdf"
    },
    {
      "id": "2605.06264",
      "anchor": "paper-2605-06264",
      "title": "Can Attribution Predict Risk? From Multi-View Attribution to Planning Risk Signals in End-to-End Autonomous Driving",
      "authors": [
        "Le Yang",
        "Haijun Liu",
        "Jiawei Liang",
        "ShangQuan Sun",
        "Xiaochun Cao"
      ],
      "abstract": "End-to-end autonomous driving models generate future trajectories from multi-view inputs, improving system integration but introducing opaque decisions and hard-to-localize risks. Existing methods either rely on auxiliary monitoring models or generate textual explanations, but are decoupled from the planning process and fail to reveal the visual evidence underlying trajectory generation. While attribution offers a direct alternative, planning differs from image classification by taking six-view camera images as input and predicting continuous multi-step trajectories, requiring attribution to capture both critical views and regions and their influence on outputs. Moreover, whether attribution maps can support risk identification remains underexplored. To address this, we propose a hierarchical attribution framework for end-to-end planning. Specifically, using L2 consistency with the original trajectory as the objective, we design a coarse-to-fine region attribution strategy that searches candidate regions across the full six-view input and refines attribution within them. We further extract three attribution statistics as predictive signals for planning risk, including attribution entropy to measure how concentrated the planner's reliance is over the joint visual space, within-camera spatial variance to characterize how spread out the attribution is within each view, and cross-camera Gini coefficient to quantify how unevenly attribution is distributed across the six cameras. Experiments on BridgeAD, UniAD, and GenAD show that these statistics correlate with planning risk, achieving Spearman correlations of $0.30 \\pm 0.07$ with trajectory error and AUROC of $0.77 \\pm 0.04$ for collision detection. The signal generalizes to held-out scenes with negligible degradation and remains stable under an alternative attribution baseline.",
      "short_abstract": "End-to-end autonomous driving models generate future trajectories from multi-view inputs, improving system integration but introducing opaque decisions and hard-to-localize risks. Existing methods either rely on auxiliary monitoring models or generate textual...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 21,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.06264",
      "pdf_url": "https://arxiv.org/pdf/2605.06264.pdf"
    },
    {
      "id": "2605.06022",
      "anchor": "paper-2605-06022",
      "title": "Lattice fermion formulation via Physics-Informed Neural Networks: Ginsparg-Wilson relation and Overlap fermions",
      "authors": [
        "Tatsuhiro Misumi"
      ],
      "abstract": "We propose a novel, machine-learning-based framework for constructing lattice fermions using Physics-Informed Neural Networks (PINNs). Our approach treats the formulation of the Dirac operator as an optimization problem guided by physical requirements, such as symmetries, locality and doubler-decoupling conditions. We first demonstrate that, when trained to satisfy the Ginsparg-Wilson (GW) relation as a soft constraint, a neural network reproduces the overlap fermion operator to high numerical accuracy and learns an effective sign-function mapping without explicitly using a prescribed polynomial or rational approximation. Secondly, we extend the framework from operator construction to machine-assisted algebraic discovery. Within a generalized polynomial ansatz, the network autonomously drives higher-order terms to zero and recovers the standard Ginsparg-Wilson relation. Remarkably, by changing the initial search bias, the same framework also finds a distinct solution corresponding to a Fujikawa-type generalized GW relation.",
      "short_abstract": "We propose a novel, machine-learning-based framework for constructing lattice fermions using Physics-Informed Neural Networks (PINNs). Our approach treats the formulation of the Dirac operator as an optimization problem guided by physical requirements, such...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.06022",
      "pdf_url": "https://arxiv.org/pdf/2605.06022.pdf"
    },
    {
      "id": "2605.06010",
      "anchor": "paper-2605-06010",
      "title": "Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models",
      "authors": [
        "Yuchen Guo",
        "Junli Gong",
        "Wenjun Dong",
        "Yiuming Cheung",
        "Weifeng Su"
      ],
      "abstract": "Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared imaging captures heat-emitting sources and provides critical complementary information, but existing high-fidelity fusion methods suffer from prohibitive latency, rendering them impractical for real-time edge deployment. To address this, we propose FusionProxy, a real-time image fusion module designed as a fully independent, plug-and-play component with diffusion level quality. FusionProxy exploits two complementary statistics of a teacher sample ensemble: per-pixel variance in raw image space, used to weight pixel-level supervision, and per-pixel variance inside frozen foundation backbones, used to route feature-level alignment spatially. Once trained, FusionProxy can be directly integrated into any visual perception system without joint optimization. Extensive experiments demonstrate that our method achieves superior performance on static recognition tasks and significantly enhances robustness in dynamic tasks, including closed-loop autonomous driving. Crucially, FusionProxy achieves real-time inference speeds on diverse platforms, from high-end GPUs to commodity hardware, providing a flexible and generalizable solution for all-day perception.",
      "short_abstract": "Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared imaging captures heat-emitting sources and provides critical complementary...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 11,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: hardware",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.06010",
      "pdf_url": "https://arxiv.org/pdf/2605.06010.pdf"
    },
    {
      "id": "2605.05964",
      "anchor": "paper-2605-05964",
      "title": "Uncertainty Estimation via Hyperspherical Confidence Mapping",
      "authors": [
        "Eunseo Choi",
        "Ho-Yeon Kim",
        "Jaewon Lee",
        "Taeyong jo",
        "Myungjun lee",
        "Heejin Ahn"
      ],
      "abstract": "Quantifying uncertainty in neural network predictions is essential for high-stakes domains such as autonomous driving, healthcare, and manufacturing. While existing approaches often depend on costly sampling or restrictive distributional assumptions, we propose Hyperspherical Confidence Mapping (HCM), a simple yet principled framework for sampling-free and distribution-free uncertainty estimation. HCM decomposes outputs into a magnitude and a normalized direction vector constrained to lie on the unit hypersphere, enabling a novel interpretation of uncertainty as the degree of violation of this geometric constraint. This yields deterministic and interpretable estimates applicable to both regression and classification. Experiments across diverse benchmarks and real-world industrial tasks demonstrate that HCM matches or surpasses ensemble and evidential approaches, with far lower inference cost and stronger confidence-error alignment. Our results highlight the power of geometric structure in uncertainty estimation and position HCM as a versatile alternative to conventional techniques.",
      "short_abstract": "Quantifying uncertainty in neural network predictions is essential for high-stakes domains such as autonomous driving, healthcare, and manufacturing. While existing approaches often depend on costly sampling or restrictive distributional assumptions, we...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.05964",
      "pdf_url": "https://arxiv.org/pdf/2605.05964.pdf"
    },
    {
      "id": "2605.05843",
      "anchor": "paper-2605-05843",
      "title": "Comparative Analysis of Direct-to-Cell (D2C) and 3GPP Non-Terrestrial Networks (NTN) for Global Connectivity",
      "authors": [
        "Donglin Wang",
        "Anjie Qiu",
        "Qiuheng Zhou",
        "Hans D. Schotten"
      ],
      "abstract": "The quest for ubiquitous mobile coverage has catalyzed two fundamentally distinct architectural paradigms: Direct-to-Cell (D2C) and standardized 3GPP Non-Terrestrial Networks (NTN). D2C, pioneered by SpaceX Starlink and AST SpaceMobile, leverages existing terrestrial spectrum and unmodified consumer handsets to provide emergency connectivity as a market-driven overlay. In contrast, 3GPP NTN, standardized across Releases 17-19, offers a systematic satellite-native framework designed for long-term scalability, high-throughput broadband, and deep integration with terrestrial 5G/6G networks. This paper presents a comprehensive technical comparison of these approaches, analyzing their standardization trajectories, network architectures, physical-layer innovations, security postures, and operational trade-offs. We further examine their implications for emerging 6G use cases, particularly autonomous driving, where safety-critical redundancy motivates a hybrid tri-link architecture combining terrestrial 5G, NTN broadband, and D2C emergency fallback. Our analysis shows that, although D2C enables rapid market entry through legacy-device compatibility, NTN provides superior performance, security, and scalability, positioning it as the foundational framework for 6G satellite-terrestrial convergence. A hybrid model that combines the strengths of both paradigms is identified as the most practical path toward truly global connectivity.",
      "short_abstract": "The quest for ubiquitous mobile coverage has catalyzed two fundamentally distinct architectural paradigms: Direct-to-Cell (D2C) and standardized 3GPP Non-Terrestrial Networks (NTN). D2C, pioneered by SpaceX Starlink and AST SpaceMobile, leverages existing...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: security"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.05843",
      "pdf_url": "https://arxiv.org/pdf/2605.05843.pdf"
    },
    {
      "id": "2605.07022",
      "anchor": "paper-2605-07022",
      "title": "Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale",
      "authors": [
        "Haydn Jones",
        "Yimeng Zeng",
        "Alden Rose",
        "Li S. Yifei",
        "Yining Huang",
        "Kaiwen Wu",
        "Jiaming Liang",
        "Maggie Ziyu Huan",
        "Yoseph Barash",
        "Cesar de la Fuente-Nunez",
        "Osbert Bastani",
        "Zachary Ives",
        "Mark Yatskar",
        "Jacob R. Gardner"
      ],
      "abstract": "Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and coverage. We show that PubMed itself can be autonomously and cost-effectively turned into structured datasets that are larger, more nuanced, and more accurate than the curated databases they replace. We present three coupled contributions: (1) an LLM-based entity-tagging pipeline, grounded in nine biomedical ontologies, that tags 4.5B entities across 19 categories in a 22.5M-paper, 2.5T-token PubMed corpus; (2) hybrid sparse-dense retrieval supporting entity-filtered semantic queries over the tagged corpus; and (3) Starling, a multi-agent deep research system that, given only a natural-language task description, designs precision- and recall-targeted retrieval filters, induces an extraction schema, and emits structured records with nuance-rich fields and supporting passages. Across six tasks -- blood-brain barrier permeability, oral bioavailability, acute toxicity (LD50), gene-disease associations, protein subcellular localization, and chemical reactions -- Starling produces ~6.3M records (91K-3M per task); several are, to our knowledge, the largest public datasets for their property. Frontier-model rejection of our extractions is 0.6-7.7% across tasks, far below error rates we measure on widely used curated counterparts (e.g., 16.5% on BBB_Martins, 7.3% on Bioavailability_Ma). Beyond scale and accuracy, the supporting passages carry nuance tabular databases discard -- e.g., oral bioavailability may depend on fed vs. fasted state. Together, the corpus, retrieval, and agent establish a foundation for AI-driven therapeutic design. Code and datasets: https://github.com/starling-labs/starling.",
      "short_abstract": "Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and coverage. We show...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Mapping & Localization: abstract: localization"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.07022",
      "pdf_url": "https://arxiv.org/pdf/2605.07022.pdf"
    },
    {
      "id": "2605.06966",
      "anchor": "paper-2605-06966",
      "title": "Traffic Scenario Orchestration from Language via Constraint Satisfaction",
      "authors": [
        "Frieda Rong",
        "Chris Zhang",
        "Kelvin Wong",
        "Raquel Urtasun"
      ],
      "abstract": "Autonomous vehicles (AVs) require extensive testing in simulation, but test case generation for driving scenarios is laborious. The desired scenarios are often out-of-distribution and have precise requirements on interactions with the AV policy under test. Manually programming scenarios allows for precise controllability but is difficult to scale. On the other hand, statistical models can leverage compute and data, but struggle with precise controllability when out-of-distribution. We cast scenario orchestration as a constraint-solving problem and present a language-in, simulation-out scenario orchestrator for closed-loop testing AVs. Our approach leverages foundation model reasoning to translate general, natural language descriptions into a set of constraints as a scenario representation. This then allows us to leverage off the shelf solvers to solve for actor behaviors which meet precise testing intentions in closed-loop. Under a benchmark of carefully crafted and diverse scenario descriptions, our approach greatly outperforms our baselines in orchestration success rate. We further show that our closed-loop approach is especially important for scenarios which require ego-reactive specifications.",
      "short_abstract": "Autonomous vehicles (AVs) require extensive testing in simulation, but test case generation for driving scenarios is laborious. The desired scenarios are often out-of-distribution and have precise requirements on interactions with the AV policy under test....",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.06966",
      "pdf_url": "https://arxiv.org/pdf/2605.06966.pdf"
    },
    {
      "id": "2605.06439",
      "anchor": "paper-2605-06439",
      "title": "From Review to Design: Ethical Multimodal Driver Monitoring Systems for Risk Mitigation, Incident Response, and Accountability in Automated Vehicles",
      "authors": [
        "Bilal Khana",
        "Waseem Shariff",
        "Rory Coyne",
        "Muhammad Ali Farooq",
        "Peter Corcoran"
      ],
      "abstract": "As vehicles transition toward higher levels of automation, Driver Monitoring Systems (DMS) have become essential for ensuring human oversight, safety, and regulatory compliance in a vehicle. These systems rely on multimodal sensing and AI-driven inference to assess driver attention, cognitive state, and readiness to take control. While technologically promising, their deployment introduces a complex set of ethical and legal challenges - ranging from privacy and consent to data ownership and algorithmic fairness. While overarching frameworks such as the GDPR, EU AI Act, and IEEE standards offer important guidance, they lack the specificity required for addressing the unique risks posed by in-cabin sensing technologies. This paper adopts a review-to-design perspective, critically examining existing regulatory instruments and ethical frameworks -- such as the GDPR, the EU AI Act, and IEEE guidelines -- and identifying gaps in their applicability to the distinctive risks posed by multimodal, AI-enabled in-cabin monitoring. Building on this review, we propose a modular ethical design framework tailored specifically to Driver Monitoring Systems. The framework translates high-level principles into actionable design and deployment guidance, including user-configurable consent mechanisms, fairness-aware model development, transparency and explainability tools, and safeguards for driver emotional well-being. Finally, the paper outlines a risk analysis and failure mitigation strategy, emphasizing proactive incident response and accountability mechanisms tailored to the DMS context. Together, these contributions aim to inform the development of transparent, trustworthy, and human-centered driver monitoring systems for next-generation autonomous vehicles.",
      "short_abstract": "As vehicles transition toward higher levels of automation, Driver Monitoring Systems (DMS) have become essential for ensuring human oversight, safety, and regulatory compliance in a vehicle. These systems rely on multimodal sensing and AI-driven inference to...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Survey",
        "Explainability",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 43,
        "margin": 37,
        "evidence": [
          "title: driver monitoring",
          "abstract: driver monitoring",
          "title: ethics",
          "abstract: ethics",
          "abstract: law or policy",
          "abstract: explainability"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.06439",
      "pdf_url": "https://arxiv.org/pdf/2605.06439.pdf"
    },
    {
      "id": "2605.05717",
      "anchor": "paper-2605-05717",
      "title": "Space-Time Diversity in Observability and Estimation on Product Lie Groups",
      "authors": [
        "Somasundhar Venkatasubramanian",
        "Anirudh Venkat",
        "Advaidh Venkat"
      ],
      "abstract": "Robust state estimation in coupled dynamical systems depends critically not only on sensor quality but on the structural alignment between observation channels and the system's intrinsic dynamics. This paper develops a rigorous framework for analyzing spatial and temporal diversity in dynamical state estimation on product Lie groups, drawing structural parallels to diversity gains in space-time coding. Three main results are established: (i) coupling-based necessary and sufficient conditions for cross-factor observability, showing that a sensor local to one group factor renders another factor observable if and only if the dynamics propagate error directions across the corresponding Lie algebra components; (ii) a spatial diversity saturation theorem identifying precisely when additional observation channels fail to expand the propagated observation subspace and thus provide no structural benefit; and (iii) a time-space diversity decomposition that exactly separates instantaneous spatial information from accumulated temporal information in the estimation error covariance. The framework is applied to planar SE(2) and spatial SE(3) navigation, yielding exact observability guarantees for redundant and non redundant sensor architectures in modern robotics and autonomous vehicles. These results extend classical observability theory beyond Euclidean state spaces, exposing structural constraints invisible to standard rank-based analysis that fundamentally govern robust inference in coupled dynamical systems.",
      "short_abstract": "Robust state estimation in coupled dynamical systems depends critically not only on sensor quality but on the structural alignment between observation channels and the system's intrinsic dynamics. This paper develops a rigorous framework for analyzing spatial...",
      "published": "2026-05-07",
      "updated": "2026-05-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: abstract: navigation",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 104,
      "arxiv_url": "https://arxiv.org/abs/2605.05717",
      "pdf_url": "https://arxiv.org/pdf/2605.05717.pdf"
    },
    {
      "id": "2605.05077",
      "anchor": "paper-2605-05077",
      "title": "FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching",
      "authors": [
        "Andranik Sargsyan",
        "Shant Navasardyan"
      ],
      "abstract": "Accurate image segmentation is essential for modern computer vision applications such as image editing, autonomous driving, and medical image analysis. In recent years, Dichotomous Image Segmentation (DIS) has become a standard task for training and evaluating highly accurate segmentation models. Existing DIS approaches often fail to preserve fine-grained details or fully capture the semantic structure of the foreground. To address these challenges, we present FlowDIS, a novel dichotomous image segmentation method built on the flow matching framework, which learns a time-dependent vector field to transport the image distribution to the corresponding mask distribution, optionally conditioned on a text prompt. Moreover, with our Position-Aware Instance Pairing (PAIP) training strategy, FlowDIS offers strong controllability through text prompts, enabling precise, pixel-level object segmentation. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches both with and without language guidance. Compared with the best prior DIS method, FlowDIS achieves a 5.5% higher $F_β^ω$ measure and 43% lower MAE ($\\mathcal{M}$) on the DIS-TE test set. The code is available at: https://github.com/Picsart-AI-Research/FlowDIS",
      "short_abstract": "Accurate image segmentation is essential for modern computer vision applications such as image editing, autonomous driving, and medical image analysis. In recent years, Dichotomous Image Segmentation (DIS) has become a standard task for training and...",
      "published": "2026-05-06",
      "updated": "2026-05-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 105,
      "arxiv_url": "https://arxiv.org/abs/2605.05077",
      "pdf_url": "https://arxiv.org/pdf/2605.05077.pdf"
    },
    {
      "id": "2605.05014",
      "anchor": "paper-2605-05014",
      "title": "CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography",
      "authors": [
        "Gasser Elazab",
        "Frank Neuhaus",
        "Tilman Koß",
        "Malte Splietker",
        "Aditya Date",
        "Michael Unterreiner",
        "Maximilian Jansen",
        "Olaf Hellwich"
      ],
      "abstract": "Autonomous driving must operate across diverse surfaces to enable safe mobility. However, most driving datasets are captured on well-paved flat roads. Moreover, recent driving datasets primarily provide sparse LiDAR ground truth for images, which is insufficient for assessing fine-grained geometry in depth estimation and completion. To address these gaps, we introduce CARD, a multi-modal driving dataset that delivers quasi-dense 3D ground truth across continuous sequences rich in speed bumps, potholes, irregular surfaces and off-road segments. Our sensor suite includes synchronized global-shutter stereo cameras, front and rear LiDARs, 6-DoF poses from LiDAR-inertial odometry, per-wheel motion traces, and full calibration. Notably, our multi-LiDAR fusion yields ~500K valid depth pixels per frame, about 6.5x more than KITTI Depth Completion and 10x more on average than other public driving datasets. The dataset spans ~110 km and 4.7 hours across Germany and Italy. In addition, CARD provides 2D bounding boxes targeting road-topography irregularities, enabling accurate benchmarking for both geometry and perception tasks. Furthermore, we establish a standardized evaluation protocol for road surface irregularities on CARD and benchmark state-of-the-art depth estimation models to provide strong baselines. The CARD dataset is hosted on https://huggingface.co/CARD-Data.",
      "short_abstract": "Autonomous driving must operate across diverse surfaces to enable safe mobility. However, most driving datasets are captured on well-paved flat roads. Moreover, recent driving datasets primarily provide sparse LiDAR ground truth for images, which is...",
      "published": "2026-05-06",
      "updated": "2026-05-06",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 11,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 105,
      "arxiv_url": "https://arxiv.org/abs/2605.05014",
      "pdf_url": "https://arxiv.org/pdf/2605.05014.pdf"
    },
    {
      "id": "2605.04963",
      "anchor": "paper-2605-04963",
      "title": "Spiral metasurface enables tunable directional edge enhancement",
      "authors": [
        "Wenli Wang",
        "Yao Hu",
        "Qiusheng Huang",
        "Dan Chen",
        "Qun Hao"
      ],
      "abstract": "Tunable directional edge enhancement facilitates the acquisition of distinct morphological features from objects, a capability that plays a vital role in enhancing the reliability and safety of autonomous driving systems. However, building a simple, miniature, and switchable directional edge enhancement system remains an urgent challenge. To address this, we propose a compact spiral metasurface capable of achieving tunable vertical and horizontal edge enhancement imaging without requiring a conventional 4f system, relying instead on a single integrated device. This functionality is realized by engineering the metasurface's spiral phase profile and its polarization-dependent response, where the edge enhancement direction is controlled by varying the incident beam's polarization, enabling switching between vertical and horizontal enhancement modes. Simulations demonstrate the metasurface's capability for switchable edge detection of lane markings and barrier contours. Broadband operation is also confirmed through simulation results. Owing to its compactness, switchable functionality, and broadband performance, the proposed spiral metasurface shows significant potential for applications in optical analog computing and autonomous driving.",
      "short_abstract": "Tunable directional edge enhancement facilitates the acquisition of distinct morphological features from objects, a capability that plays a vital role in enhancing the reliability and safety of autonomous driving systems. However, building a simple,...",
      "published": "2026-05-06",
      "updated": "2026-05-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 105,
      "arxiv_url": "https://arxiv.org/abs/2605.04963",
      "pdf_url": "https://arxiv.org/pdf/2605.04963.pdf"
    },
    {
      "id": "2605.04675",
      "anchor": "paper-2605-04675",
      "title": "Physical Adversarial Clothing Evades Visible-Thermal Detectors via Non-Overlapping RGB-T Pattern",
      "authors": [
        "Xiaopei Zhu",
        "Guanning Zeng",
        "Zhanhao Hu",
        "Jun Zhu",
        "Xiaolin Hu"
      ],
      "abstract": "Visible-thermal (RGB-T) object detection is a crucial technology for applications such as autonomous driving, where multimodal fusion enhances performance in challenging conditions like low light. However, the security of RGB-T detectors, particularly in the physical world, has been largely overlooked. This paper proposes a novel approach to RGB-T physical attacks using adversarial clothing with a non-overlapping RGB-T pattern (NORP). To simulate full-view (0$^{\\circ}$--360$^{\\circ}$) RGB-T attacks, we construct 3D RGB-T models for human and adversarial clothing. NORP is a new adversarial pattern design using distinct visible and thermal materials without overlap, avoiding the light reduction in overlapping RGB-T patterns (ORP). To optimize the NORP on adversarial clothing, we propose a spatial discrete-continuous optimization (SDCO) method. We systematically evaluated our method on RGB-T detectors with different fusion architectures, demonstrating high attack success rates both in the digital and physical worlds. Additionally, we introduce a fusion-stage ensemble method that enhances the transferability of adversarial attacks across unseen RGB-T detectors with different fusion architectures.",
      "short_abstract": "Visible-thermal (RGB-T) object detection is a crucial technology for applications such as autonomous driving, where multimodal fusion enhances performance in challenging conditions like low light. However, the security of RGB-T detectors, particularly in the...",
      "published": "2026-05-06",
      "updated": "2026-05-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 17,
        "evidence": [
          "abstract: security",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 105,
      "arxiv_url": "https://arxiv.org/abs/2605.04675",
      "pdf_url": "https://arxiv.org/pdf/2605.04675.pdf"
    },
    {
      "id": "2605.04647",
      "anchor": "paper-2605-04647",
      "title": "ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving",
      "authors": [
        "Huimin Wang",
        "Yue Wang",
        "Bihao Cui",
        "Pengxiang Li",
        "Ben Lu",
        "Mingqian Wang",
        "Tong Wang",
        "Chuan Tang",
        "Teng Zhang",
        "Kun Zhan"
      ],
      "abstract": "We introduce ReflectDrive-2, a masked discrete diffusion planner with separate action expert for autonomous driving that represents plans as discrete trajectory tokens and generates them through parallel masked decoding. This discrete token space enables in-place trajectory revision: AutoEdit rewrites selected tokens using the same model, without requiring an auxiliary refinement network. To train this capability, we use a two-stage procedure. First, we construct structure-aware perturbations of expert trajectories along longitudinal progress and lateral heading directions and supervise the model to recover the original expert trajectory. We then fine-tune the full decision--draft--reflect rollout with reinforcement learning (RL), assigning terminal driving reward to the final post-edit trajectory and propagating policy-gradient credit through full-rollout transitions. Full-rollout RL proves crucial for coupling drafting and editing: under supervised training alone, inference-time AutoEdit improves PDMS by at most $0.3$, whereas RL increases its gain to $1.9$. We also co-design an efficient reflective decoding stack for the decision--draft--reflect pipeline, combining shared-prefix KV reuse, Alternating Step Decode, and fused on-device unmasking. On NAVSIM, ReflectDrive-2 achieves $91.0$ PDMS with camera-only input and $94.8$ PDMS in a best-of-6 oracle setting, while running at $31.8$ ms average latency on NVIDIA Thor.",
      "short_abstract": "We introduce ReflectDrive-2, a masked discrete diffusion planner with separate action expert for autonomous driving that represents plans as discrete trajectory tokens and generates them through parallel masked decoding. This discrete token space enables...",
      "published": "2026-05-06",
      "updated": "2026-05-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Systems, Deployment & Connectivity: abstract: communications",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 105,
      "arxiv_url": "https://arxiv.org/abs/2605.04647",
      "pdf_url": "https://arxiv.org/pdf/2605.04647.pdf"
    },
    {
      "id": "2605.08211",
      "anchor": "paper-2605-08211",
      "title": "Learning the Channel Gain from Anywhere to Anywhere via Cross-environment Transformer Estimators",
      "authors": [
        "Prasenjit Dhara",
        "Daniel Romero"
      ],
      "abstract": "Channel-gain maps provide the channel gain between any two locations in a geographical region. They find numerous applications, from resource allocation and interference control to path planning for autonomous vehicles. Channel-gain map estimation (CGME) is considerably more challenging than conventional radio map estimation (RME) because channel-gain maps are functions over a 6-dimensional input space. This calls for specialized methods, which currently rely on the (inaccurate) radio tomographic model or require a prohibitively large number of measurements since they do not exploit any spatial structure. This paper overcomes this issue by leveraging spatial patterns that channel-gain maps exhibit across environments, as dictated by the laws of physics and typical environmental characteristics (e.g. building materials and layouts). Adopting a metalearning perspective, a transformer-based estimator is proposed to implicitly learn this common structure from measurements collected in multiple environments. This enables CGME in new environments from significantly fewer measurements (five times less in our experiments). To maximize learning efficiency, the transformer is composed with a feature map that enforces the invariances of CGME, such as those following from reciprocity. Numerical experiments corroborate the merits of the proposed estimator relative to existing methods.",
      "short_abstract": "Channel-gain maps provide the channel gain between any two locations in a geographical region. They find numerous applications, from resource allocation and interference control to path planning for autonomous vehicles. Channel-gain map estimation (CGME) is...",
      "published": "2026-05-06",
      "updated": "2026-05-06",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 6,
        "evidence": [
          "abstract: motion planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 105,
      "arxiv_url": "https://arxiv.org/abs/2605.08211",
      "pdf_url": "https://arxiv.org/pdf/2605.08211.pdf"
    },
    {
      "id": "2605.05050",
      "anchor": "paper-2605-05050",
      "title": "Kinematic Discriminants of Deceleration Behavior Modes in Car-Following: Evidence from NGSIM Trajectory Data",
      "authors": [
        "Eni Solomon Laughter"
      ],
      "abstract": "Gap-closing rate and visual looming swap discriminative dominance depending on deceleration intensity - a finding that reconciles a long-standing conflict in the car-following literature and challenges spacing-centered assumptions in traditional driver behavior models. This study presents a two-stage analytical framework that distinguishes between information availability (kinematic variables measurable in the environment) and information utilization (variables that demonstrably separate driver behavioral patterns), applied to 1,060,119 valid car-following observations from the NGSIM trajectory dataset (2,932 vehicles). Six kinematic features are extracted, and deceleration events are detected under two threshold conditions (-0.5 m/s^2 and -0.3 m/s^2). K-means clustering identifies behavioral modes, and one-way ANOVA with eta-squared effect sizes ranks each feature's discriminative power. Three key findings emerge: (1) threshold selection fundamentally shapes behavioral inference - the stricter threshold yields three interpretable modes while the permissive threshold collapses these to two; (2) hard braking prioritizes gap-closing rate (eta^2 = 0.715) while moderate braking emphasizes visual looming (eta^2 = 0.574); and (3) spacing headway is negligible (eta^2 <= 0.014) across both thresholds. These findings provide empirically grounded candidates for perceptual cue prioritization and have direct implications for ADAS warning system design and autonomous vehicle control.",
      "short_abstract": "Gap-closing rate and visual looming swap discriminative dominance depending on deceleration intensity - a finding that reconciles a long-standing conflict in the car-following literature and challenges spacing-centered assumptions in traditional driver...",
      "published": "2026-05-06",
      "updated": "2026-05-06",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 3,
        "evidence": [
          "abstract: vehicle control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 105,
      "arxiv_url": "https://arxiv.org/abs/2605.05050",
      "pdf_url": "https://arxiv.org/pdf/2605.05050.pdf"
    },
    {
      "id": "2605.04475",
      "anchor": "paper-2605-04475",
      "title": "Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding",
      "authors": [
        "Shuo Liu",
        "Lei Shi",
        "Haowen Liu",
        "Jing Xu",
        "Yufei Gao",
        "Yucheng Shi"
      ],
      "abstract": "Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning stage. However, many recent LLM-driven driving systems attach the language model as a post-processor and force it to reason over redundant or conflicting perception outputs, which can amplify hallucinated entities and unsafe conclusions. This paper proposes InfoCoordiBridge, a BEV-centric neuro-symbolic architecture that inserts an explicit coordination bridge between perception and language reasoning. InfoCoordiBridge comprises (i) a unified multi-agent perception layer that outputs typed structured facts together with modality-focused synopses, (ii) an ICA module that aligns and fuses multi-source outputs into a single SceneSummary, and (iii) an SSRE module that performs SceneSummary-grounded reasoning with verification. Experiments on nuScenes and Waymo show that ICA preserves competitive 3D detection accuracy while substantially improving fusion consistency, reducing redundancy to below 1% and achieving about 98% attribute agreement. On NuScenes-QA and a template-aligned Waymo-QA benchmark, SSRE improves factual grounding and reduces hallucinated entity mentions compared with representative VLM and agentic baselines. Overall, by coordinating multi-sensor outputs into a single conflict-aware SceneSummary before prompting, InfoCoordiBridge prevents redundant and cross-modally inconsistent perception evidence from propagating into high-level reasoning.",
      "short_abstract": "Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning stage. However, many recent LLM-driven driving systems attach the language model as a post-processor and...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 12,
        "evidence": [
          "abstract: perception",
          "title: scene understanding",
          "abstract: scene understanding",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.04475",
      "pdf_url": "https://arxiv.org/pdf/2605.04475.pdf"
    },
    {
      "id": "2605.04470",
      "anchor": "paper-2605-04470",
      "title": "CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies",
      "authors": [
        "Keyu Chen",
        "Nanfei Ye",
        "Yida Wang",
        "Wenchao Sun",
        "Danqi Zhao",
        "Hao Cheng",
        "Sifa Zheng"
      ],
      "abstract": "Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL fine-tuning provides grounded feedback from executed actions but is constrained by the sparsity of informative events, whereas counterfactual fine-tuning provides dense supervision over candidate futures but inherits bias from imperfect future estimates. We introduce Counterfactual-to-Interactive Reinforcement Fine-Tuning (CRAFT), an on-policy framework that formulates closed-loop post-training as proxy-residual optimization. CRAFT uses group-normalized counterfactual advantages as a dense proxy for real closed-loop advantages and aligns this proxy with the closed-loop world through grounded residual correction from interaction-critical events. To stabilize adaptation, CRAFT regularizes the online policy toward an EMA teacher via asymmetric KL self-distillation. Theoretically, CRAFT decomposes the real closed-loop policy gradient into proxy and residual terms under the same visited-state distribution, reducing residual variance with an aligned proxy while mitigating proxy bias through grounded residual approximation. Empirically, CRAFT achieves the strongest closed-loop gains on Bench2Drive across hierarchical planning, vision-language-action, and vocabulary-scoring architectures. Ablations, scaling behavior, stability analyses, and transfer results further validate the complementary roles of dense counterfactual proxy and grounded residual correction. Project page: https://currychen77.github.io/CRAFT.",
      "short_abstract": "Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "abstract: vision-language-action",
          "abstract: driving policy"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.04470",
      "pdf_url": "https://arxiv.org/pdf/2605.04470.pdf"
    },
    {
      "id": "2605.04435",
      "anchor": "paper-2605-04435",
      "title": "Ground4D: Spatially-Grounded Feedforward 4D Reconstruction for Unstructured Off-Road Scenes",
      "authors": [
        "Shuo Wang",
        "Jilin Mei",
        "Fuyang Liu",
        "Wenfei Guan",
        "Fanjie Kong",
        "Zhihua Zhao",
        "Shuai Wang",
        "Chen Min",
        "Yu Hu"
      ],
      "abstract": "Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-road scenes, its performance degrades due to high-frequency geometry, ego-motion jitter, and increased non-rigid dynamics. These factors introduce conflicting Gaussian observations across timestamps, leading to either over-smoothed renderings or structural artifacts. To address this issue, we propose Ground4D, a spatially-grounded 4D feedforward framework for pose-free off-road reconstruction. The key idea is to resolve temporal conflicts through spatially localized conditioning. Specifically, we introduce voxel-grounded temporal Gaussian aggregation, which partitions the canonical Gaussian space into spatial voxels and performs query-conditioned temporal attention within each voxel. Intra-voxel softmax normalization ensures that temporal selectivity and spatial occupancy become mutually reinforcing rather than conflicting. We furthermore introduce surface normal cues as auxiliary geometric guidance to regularize the geometry of Gaussian primitives. Extensive experiments on ORAD-3D and RELLIS-3D demonstrate that Ground4D consistently outperforms existing feedforward methods in reconstruction quality and generalizes zero-shot to unseen off-road domains. Project page and code:https://github.com/wsnbws/Ground4D.",
      "short_abstract": "Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-road scenes, its performance degrades due to high-frequency geometry, ego-motion jitter, and increased...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Perception & Sensor Fusion: abstract: occupancy"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.04435",
      "pdf_url": "https://arxiv.org/pdf/2605.04435.pdf"
    },
    {
      "id": "2605.04355",
      "anchor": "paper-2605-04355",
      "title": "InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making",
      "authors": [
        "Mustafa Sakhaia",
        "Kaung Sithua",
        "Min Khant Soe Okea",
        "Maciej Wielgosza"
      ],
      "abstract": "Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often struggle in high-dynamic- range scenes or high-speed scenarios due to motion blur and latency. Dynamic Vision Sensors (DVS), or event cameras, offer a paradigm shift by capturing asynchronous brightness changes with microsecond temporal resolution and high dynamic range. In this paper, we propose an extended architecture of the state-of-the-art InterFuser model, integrating DVS as an additional modality to enhance perception reliability. We introduce a novel token-based fusion strategy that incorporates accumulated event frames into the transformer-based backbone of InterFuser. Our method leverages the complementary nature of RGB, LiDAR, and DVS data. We evaluate our approach on the Car Learning to Act (CARLA) Leaderboard benchmarks, demonstrating that the inclusion of DVS improves the robustness of the driving agent, achieving a competitive Driving Score of 77.2 and a superior Route Completion of 100%. The results indicate that event-based vision is a promising direction for improving safety and performance in adverse lighting and dynamic conditions.",
      "short_abstract": "Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often struggle in high-dynamic- range scenes or high-speed scenarios due to motion blur and latency. Dynamic...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 12,
        "evidence": [
          "abstract: perception",
          "title: sensor fusion",
          "abstract: sensor fusion",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.04355",
      "pdf_url": "https://arxiv.org/pdf/2605.04355.pdf"
    },
    {
      "id": "2605.04299",
      "anchor": "paper-2605-04299",
      "title": "Beyond Fixed Thresholds and Domain-Specific Benchmarks for Explainable Multi-Task Classification in Autonomous Vehicles",
      "authors": [
        "Maryam Sadat Hosseini Azad",
        "Shahriar Baradaran Shokouhi"
      ],
      "abstract": "Scene understanding is a vital part of autonomous driving systems, which requires the use of deep learning models. Deep learning methods are intrinsically black box models, which lack transparency and safety in autonomous driving. To make these systems transparent, multi-task visual understanding has become crucial for explainable autonomous driving perception systems, where simultaneous prediction of multiple driving behaviors and their underlying explanations is essential for safe navigation and human trust in autonomous vehicles. In order to design an accurate and cross-cultural explainable autonomous driving system, we introduce a comprehensive confidence threshold sensitivity analysis that evaluates various threshold values to identify optimal decision boundaries for different tasks. Our analysis demonstrates that traditional fixed threshold approaches are suboptimal for multi-task scenarios. Through extensive evaluation, we demonstrate that our adaptive threshold selection methodology improves F1-scores across different tasks. In addition, we introduce IUST-XAI-AD, a novel dataset consisting of 958 images with human annotations for driving decisions and corresponding reasoning. This dataset addresses the critical gap in domain-specific evaluation benchmarks for distinct driving contexts and provides a more challenging test environment compared to existing datasets. Experimental results demonstrate that confidence threshold sensitivity analysis can significantly improve model performance, while the introduction of the IUST-XAI-AD dataset reveals important insights about cross-cultural driving behavior patterns. The combined contributions of this work provide both methodological advances and practical evaluation tools that can accelerate the development of more reliable, explainable, and culturally-adaptive autonomous driving systems for global deployment.",
      "short_abstract": "Scene understanding is a vital part of autonomous driving systems, which requires the use of deep learning models. Deep learning methods are intrinsically black box models, which lack transparency and safety in autonomous driving. To make these systems...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Dataset",
        "Benchmark",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "title: explainability",
          "abstract: explainability"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.04299",
      "pdf_url": "https://arxiv.org/pdf/2605.04299.pdf"
    },
    {
      "id": "2605.03491",
      "anchor": "paper-2605-03491",
      "title": "Real-Time Evaluation of Autonomous Systems under Adversarial Attacks",
      "authors": [
        "Adithya Mohan",
        "Xujun Xie",
        "Venkatesh Thirugnana Sambandham",
        "Torsten Schön"
      ],
      "abstract": "Most evaluations of autonomous driving policies under adversarial conditions are conducted in simulation, due to cost efficiency and the absence of physical risk. However, purely virtual testing fails to capture structural inconsistencies, supervision constraints, and state-representation effects that arise in real-world data and fundamentally shape policy robustness. This work presents an offline trajectory-learning and adversarial robustness evaluation framework grounded in real-world intersection driving data. Within a controlled data contract, we train and compare three trajectory-learning paradigms: Multi-Layer Perceptron (MLP)-based Behavior Cloning (BC), Transformer-based object-tokenized BC, and inverse reinforcement learning (IRL) formulated within a Generative Adversarial Imitation Learning (GAIL) framework. Models are evaluated using Average Displacement Error (ADE) and Final Displacement Error (FDE). Inference-time robustness is assessed by subjecting trained policies to gradient-based adversarial perturbations across multiple intersection scenarios, yielding a structured robustness evaluation matrix. Results show that state-structure design and architectural inductive biases critically influence adversarial stability, leading to markedly different robustness profiles despite comparable nominal prediction accuracy (ADE < 0.08). Inference-time Projected Gradient Descent (PGD) attacks induce final displacement errors of up to approximately 8 meters. The proposed framework establishes a scalable benchmark for studying offline trajectory learning and adversarial robustness in real-world autonomous driving settings.",
      "short_abstract": "Most evaluations of autonomous driving policies under adversarial conditions are conducted in simulation, due to cost efficiency and the absence of physical risk. However, purely virtual testing fails to capture structural inconsistencies, supervision...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark",
        "Reinforcement Learning",
        "Imitation Learning",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 12,
        "evidence": [
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.03491",
      "pdf_url": "https://arxiv.org/pdf/2605.03491.pdf"
    },
    {
      "id": "2605.03365",
      "anchor": "paper-2605-03365",
      "title": "Dual-Foundation Models for Unsupervised Domain Adaptation",
      "authors": [
        "Yerin Cheon",
        "Aruna Balasubramanian",
        "Francois Rameau"
      ],
      "abstract": "Semantic segmentation provides pixel-level scene understanding essential for autonomous driving and fine-grained perception tasks. However, training segmentation models requires costly, labor-intensive annotations on real-world datasets. Unsupervised Domain Adaptation (UDA) addresses this by training models on labeled synthetic data and adapting them to unlabeled real images. While conceptually simple, adaptation is challenging due to the domain gap, i.e., differences in visual appearance and scene structure between synthetic and real data. Prior approaches bridge this gap through pixel-level mixing or feature-level contrastive learning. Yet, these techniques suffer from two major limitations: (1) reliance on high-confidence pseudo-labels restricts learning to a subset of the target domain, and (2) prototype-based contrastive methods initialize class prototypes from source-trained models, yielding biased and unstable anchors during adaptation. To address these issues, we propose a dual-foundation UDA framework that leverages two complementary foundation models. First, we employ the Segment Anything Model (SAM) with superpixel-guided prompting to enable learning from a broader range of target pixels beyond high-confidence predictions. Second, we incorporate DINOv3 to construct stable, domain-invariant class prototypes through its robust representation learning. Our method achieves consistent improvements of +1.3% and +1.4% mIoU over strong UDA baselines on GTA-to-Cityscapes and SYNTHIA-to-Cityscapes, respectively.",
      "short_abstract": "Semantic segmentation provides pixel-level scene understanding essential for autonomous driving and fine-grained perception tasks. However, training segmentation models requires costly, labor-intensive annotations on real-world datasets. Unsupervised Domain...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 8,
        "evidence": [
          "abstract: perception",
          "abstract: scene segmentation",
          "abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.03365",
      "pdf_url": "https://arxiv.org/pdf/2605.03365.pdf"
    },
    {
      "id": "2605.16327",
      "anchor": "paper-2605-16327",
      "title": "Differentiable Optimization Layered Safety-Critical Control for Risk-Aware Navigation via Conformal Prediction",
      "authors": [
        "Jinyang Dong",
        "Shizhen Wu",
        "Yongchun Fang"
      ],
      "abstract": "Risk-aware navigation in unknown environments is a fundamental challenge for autonomous vehicles operating in complex urban systems. To address this issue, this paper presents a differentiable optimization layered safety-critical control method based on conformal prediction. First, to handle uncertainties arising from sensor noise, the conformal prediction method is employed to generate risk-aware obstacle ellipsoids around an elliptical-shaped robot. Second, two nested differentiable optimization layers are introduced to build the control barrier functions for obstacle avoidance and feasibility guarantee, respectively. Then, a quadratic program based safety-critical control law is proposed to integrate the above control barrier function constraints as well as input constraints. In the end, the effectiveness of the proposed framework is demonstrated through numerical simulations.",
      "short_abstract": "Risk-aware navigation in unknown environments is a fundamental challenge for autonomous vehicles operating in complex urban systems. To address this issue, this paper presents a differentiable optimization layered safety-critical control method based on...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 34,
        "margin": 26,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: risk assessment",
          "abstract: risk assessment",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.16327",
      "pdf_url": "https://arxiv.org/pdf/2605.16327.pdf"
    },
    {
      "id": "2605.08190",
      "anchor": "paper-2605-08190",
      "title": "Synergistic Simplex: Cooperative Runtime Assurance for Safety-Critical Autonomous Systems",
      "authors": [
        "Ayoosh Bansal",
        "Mikael Yeghiazaryan",
        "Artyom Khachatryan",
        "Tianyi Zhu",
        "Hunmin Kim",
        "Naira Hovakimyan",
        "Lui Sha"
      ],
      "abstract": "Autonomous systems increasingly rely on machine-learning (ML) components for safety-critical tasks such as perception and control in autonomous vehicles (AVs). While ML enables essential capabilities, it inevitably exhibits long-tail faults that make it unsuitable for safety-critical tasks. Runtime assurance (RTA) mitigates this issue by pairing ML components with verifiable safety monitors, e.g., Control Simplex and Perception Simplex architectures. However, the limited performance of safety monitors remains a major bottleneck. The Synergistic Simplex (SS) architecture improves system performance by enabling bidirectional integration between ML components and safety monitors while preserving formal safety guarantees. The key innovation here is allowing safety monitors to use ML outputs, which is typically prohibited in RTA systems. We formally derive conditions under which this integration preserves safety and demonstrate the performance benefits. We present the design, analysis, and evaluation of SS for AV obstacle detection.",
      "short_abstract": "Autonomous systems increasingly rely on machine-learning (ML) components for safety-critical tasks such as perception and control in autonomous vehicles (AVs). While ML enables essential capabilities, it inevitably exhibits long-tail faults that make it...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 7,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.08190",
      "pdf_url": "https://arxiv.org/pdf/2605.08190.pdf"
    },
    {
      "id": "2605.04421",
      "anchor": "paper-2605-04421",
      "title": "FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning",
      "authors": [
        "Waleed Razzaq",
        "Yun-Bo Zhao"
      ],
      "abstract": "Continuous-time (CT) Transformers improve irregular and long-range modeling over CT-RNNs by exploiting inputs or outputs embeddings with continuous dynamics. However, the core scaled-dot-product-attention (SDPA) mechanism remains inherently discrete. We propose FLUID (Flexible Unified Information Dynamics), a CT Transformer that incorporates continuous dynamics directly into the attention computation by replacing it with Liquid Attention Network (LAN). LAN reinterprets attention logits as continuous dynamical system and reformulates them as the solution to a linear ODE modulated by input-dependent nonlinear recurrent gates. Theoretically, we establish stability guarantees for LAN dynamics and show that it serves as an interpolating middle ground between SDPA and CT-RNNs, recovering each as special case under well-defined parameterization of its gating functions. LAN also introduces an explicit attention-sink gate to eliminate disproportionate attention mass on uninformative nodes. FLUID replaces standard residual connections with input-dependent Liquid Hyper-Connections to adaptively regulate interlayer information flow. Empirically, we evaluate FLUID on a broad set of learning tasks, including (i) irregular time-series, (ii) long-range modeling, (iii) lane-keeping control of autonomous vehicles, and (iv) learning physical dynamics under a scarce data regime. Across all the tasks, FLUID consistently matches or outperforms CT baselines, achieving improvements of up to 47% in certain scenarios and enhancing generalization under distributional shifts. Additionally, FLUID demonstrates superior noise robustness and a self-correcting inductive bias in autonomous vehicle control. We also provide a detailed analysis of key hyperparameters to guide tuning and show that FLUID occupies an intermediate position among competing approaches in terms of runtime and memory efficiency.",
      "short_abstract": "Continuous-time (CT) Transformers improve irregular and long-range modeling over CT-RNNs by exploiting inputs or outputs embeddings with continuous dynamics. However, the core scaled-dot-product-attention (SDPA) mechanism remains inherently discrete. We...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 4,
        "evidence": [
          "abstract: vehicle control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.04421",
      "pdf_url": "https://arxiv.org/pdf/2605.04421.pdf"
    },
    {
      "id": "2605.04366",
      "anchor": "paper-2605-04366",
      "title": "Conditional Flow-VAE for Safety-Critical Traffic Scenario Generation",
      "authors": [
        "Zimu Gong",
        "Brian Zhaoning Zhang",
        "Chris Zhang",
        "Kelvin Wong",
        "Raquel Urtasun"
      ],
      "abstract": "Safety-critical scenarios are essential for the development of autonomous vehicles (AVs) but are rare in real-world driving data. While simulation offers a way to generate such scenarios, manually designed test cases lack scalability, and adversarial optimization often produces unrealistic behaviors. In this work, we introduce a conditional latent flow matching approach for scalable and realistic safety-critical scenario generation. Our method uses distribution matching to transform nominal scenes into safety-critical rollouts. Furthermore, we demonstrate that incorporating both simulation and real-world data enables our framework to efficiently generate diverse, data-driven scenarios. Experimental results highlight that our approach is able to more consistently and realistically generate novel safety-critical scenarios, making it a valuable tool for training and benchmarking AV systems.",
      "short_abstract": "Safety-critical scenarios are essential for the development of autonomous vehicles (AVs) but are rare in real-world driving data. While simulation offers a way to generate such scenarios, manually designed test cases lack scalability, and adversarial...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 20,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.04366",
      "pdf_url": "https://arxiv.org/pdf/2605.04366.pdf"
    },
    {
      "id": "2605.03856",
      "anchor": "paper-2605-03856",
      "title": "Nested array design of extended coprime sets for DOA estimation of non-circular signals",
      "authors": [
        "Dongqi Chen",
        "Kun Ye",
        "Chuanxi Xing",
        "Waqas Khalid",
        "Huiping Huang"
      ],
      "abstract": "In recent years, direction of arrival estimation utilizing non-circular signals has become a focal point for scholarly research. To enhance the degrees of freedom (DOF) in receiver arrays specifically for non-circular signal DOA estimation, this study introduces a novel array configuration. This design leverages an extended coprime framework, applying a sliding translation technique to optimize sensor placement. Crucially, this rearranged structure preserves the continuity of the difference co-array (DCA). Furthermore, the sum co-array (SCA) is shifted to merge seamlessly with the DCA, eliminating redundancy and substantially expanding both the virtual aperture array (VAA) and the DOF. Consequently, the proposed array demonstrates superior performance in practical DOA estimation tasks involving non-circular signals. Simulation results and comparative analyses confirm that, relative to traditional Nested Arrays (NA), Extended Sliding Nested Array (ESNA), and other benchmark structures, the proposed array achieves better DOF and VAA, leading to enhanced estimation accuracy in practical scenarios.",
      "short_abstract": "In recent years, direction of arrival estimation utilizing non-circular signals has become a focal point for scholarly research. To enhance the degrees of freedom (DOF) in receiver arrays specifically for non-circular signal DOA estimation, this study...",
      "published": "2026-05-05",
      "updated": "2026-05-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 106,
      "arxiv_url": "https://arxiv.org/abs/2605.03856",
      "pdf_url": "https://arxiv.org/pdf/2605.03856.pdf"
    },
    {
      "id": "2605.02762",
      "anchor": "paper-2605-02762",
      "title": "Unified Map Prior Encoder for Mapping and Planning",
      "authors": [
        "Zongzheng Zhang",
        "Sizhe Zou",
        "Guantian Zheng",
        "Zhenxin Zhu",
        "Yu Gao",
        "Guoxuan Chi",
        "Shuo Wang",
        "Yuwen Heng",
        "Zhigang Sun",
        "Yiru Wang",
        "Hao Sun",
        "Chao Ma",
        "Zhen Li",
        "Anqing Jiang",
        "Hao Zhao"
      ],
      "abstract": "Online mapping and end-to-end (E2E) planning in autonomous driving remain largely sensor-centric, leaving rich map priors, including HD/SD vector maps, rasterized SD maps, and satellite imagery, underused because of heterogeneity, pose drift, and inconsistent availability at test time. We present UMPE, a Unified Map Prior Encoder that can ingest any subset of four priors and fuse them with BEV features for both mapping and planning. UMPE has two branches. The vector encoder pre-aligns HD/SD polylines with a frame-wise SE(2) correction, encodes points via multi-frequency sinusoidal features, and produces polyline tokens with confidence scores. BEV queries then apply cross-attention with confidence bias, followed by normalized channel-wise gating to avoid length imbalance and softly down-weight uncertain sources. The raster encoder shares a ResNet-18 backbone conditioned by FiLM with scaling and shift at every stage, performs SE(2) micro-alignment, and injects priors through zero-initialized residual fusion, so the network starts from a do-no-harm baseline and learns to add only useful prior evidence. A vector-then-raster fusion order reflects the inductive bias of geometry first, appearance second. On nuScenes mapping, UMPE lifts MapTRv2 from 61.5 to 67.4 mAP (+5.9) and MapQR from 66.4 to 71.7 mAP (+5.3). On Argoverse2, UMPE adds +4.1 mAP over strong baselines. UMPE is compositional: when trained with all priors, it outperforms single-prior models even when only one prior is available at test time, demonstrating powerset robustness. For E2E planning with the VAD backbone on nuScenes, UMPE reduces trajectory error from 0.72 to 0.42 m L2 on average (-0.30 m) and collision rate from 0.22% to 0.12% (-0.10%), surpassing recent prior-injection methods. These results show that a unified, alignment-aware treatment of heterogeneous map priors yields better mapping and better planning.",
      "short_abstract": "Online mapping and end-to-end (E2E) planning in autonomous driving remain largely sensor-centric, leaving rich map priors, including HD/SD vector maps, rasterized SD maps, and satellite imagery, underused because of heterogeneity, pose drift, and inconsistent...",
      "published": "2026-05-04",
      "updated": "2026-05-04",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 107,
      "arxiv_url": "https://arxiv.org/abs/2605.02762",
      "pdf_url": "https://arxiv.org/pdf/2605.02762.pdf"
    },
    {
      "id": "2605.02716",
      "anchor": "paper-2605-02716",
      "title": "Parking Assistance for Trailer-Truck Transport Vehicles Using Sensor Fusion and Motion Planning",
      "authors": [
        "George Alenchery",
        "Thomas Jeske",
        "Tova Quinones",
        "Lentz Fortune",
        "Tristan Lindo-Slones",
        "Amber Jones",
        "Jordan Fletcher"
      ],
      "abstract": "Autonomous driving technology has rapidly evolved over the past decade, offering significant improvements in transportation efficiency, safety, and cost reduction. While much of the progress has focused on highway driving and obstacle avoidance, low-speed maneuvers such as parking remain among the most difficult challenges for autonomous systems. This challenge is especially pronounced in trailer-truck transport vehicles due to their articulated motion and environmental constraints. This paper presents a proposed framework for autonomous truck parking that integrates perception, motion planning, control systems, and infrastructure awareness. By combining sensor fusion, Hybrid A* path planning, nonlinear model predictive control (NMPC), and data-driven parking systems, this work highlights the importance of system-level coordination for reliable and scalable autonomous parking solutions. As a proof-of-concept implementation, we adapted an open-source A* path planning simulation to incorporate a tractor-trailer kinematic model, demonstrating articulated vehicle path planning within a command-line simulation environment, with jackknife prevention identified as an area requiring further development.",
      "short_abstract": "Autonomous driving technology has rapidly evolved over the past decade, offering significant improvements in transportation efficiency, safety, and cost reduction. While much of the progress has focused on highway driving and obstacle avoidance, low-speed...",
      "published": "2026-05-04",
      "updated": "2026-05-04",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 13,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 107,
      "arxiv_url": "https://arxiv.org/abs/2605.02716",
      "pdf_url": "https://arxiv.org/pdf/2605.02716.pdf"
    },
    {
      "id": "2605.02580",
      "anchor": "paper-2605-02580",
      "title": "Hyp2Former: Hierarchy-Aware Hyperbolic Embeddings for Open-Set Panoptic Segmentation",
      "authors": [
        "Yao Lu",
        "Rohit Mohan",
        "Florian Drews",
        "Yakov Miron",
        "Abhinav Valada"
      ],
      "abstract": "Recognizing unknown objects is crucial for safety-critical applications such as autonomous driving and robotics. Open-Set Panoptic Segmentation (OPS) aims to segment known thing and stuff classes while identifying valid unknown objects as separate instances. Prior OPS approaches largely treat known categories as a flat label set, ignoring the semantic hierarchy that provides valuable structural priors for distinguishing unknown objects from in-distribution classes. In this work, we propose Hyp2Former, an end-to-end framework for OPS that does not require explicit modeling of unknowns during training, and instead learns hierarchical semantic similarities continuously in hyperbolic space. By explicitly encoding hierarchical relationships among known categories, the model learns a structured embedding space that captures multiple levels of semantic abstraction. As a result, unknown objects that cannot be confidently classified as known categories still remain in close proximity to higher-level concepts (e.g., an unknown animal remains closer to \"animal\" or \"object\" than to unrelated concepts such as \"electronics\" or \"stuff\") and can therefore be reliably detected, even if their fine-grained category was not represented during training. Empirical evaluations across multiple public datasets such as MS COCO, Cityscapes, and Lost&Found demonstrate that Hyp2Former outperforms existing methods on OPS, achieving the best balance between unknown object discovery and in-distribution robustness.",
      "short_abstract": "Recognizing unknown objects is crucial for safety-critical applications such as autonomous driving and robotics. Open-Set Panoptic Segmentation (OPS) aims to segment known thing and stuff classes while identifying valid unknown objects as separate instances....",
      "published": "2026-05-04",
      "updated": "2026-05-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 10,
        "evidence": [
          "title: scene segmentation",
          "abstract: scene segmentation"
        ]
      },
      "recency": "archive",
      "age_days": 107,
      "arxiv_url": "https://arxiv.org/abs/2605.02580",
      "pdf_url": "https://arxiv.org/pdf/2605.02580.pdf"
    },
    {
      "id": "2605.02357",
      "anchor": "paper-2605-02357",
      "title": "Channel-Level Relation to Attentive Aggregation with Neighborhood-Homogeneity Constraint for Point Cloud Analysis",
      "authors": [
        "Jiaqi Shi",
        "Jin Xiao",
        "Xiaoguang Hu",
        "Wenxuan Ji",
        "Zichong Jia",
        "Zifan Long",
        "Tianyou Chen",
        "Baochang Zhang"
      ],
      "abstract": "In 3D point cloud understanding, the core challenge lies in accurately capturing discriminative features within complex neighborhoods, which directly affects the execution precision of downstream tasks such as embodied AI and autonomous driving. Existing methods explore feature correlation discrimination but are limited to point-level spatial distribution or channel responses, enabling only coarse-grained level evaluation. For modern multi-scale point cloud networks, such coarse-grained metrics inevitably incur significant information loss in deeper layers. To address this, we propose PointCRA, a novel network with a channel-level metric-based enhancement mechanism. Our core idea is to introduce temporal trend variation as a new evaluation dimension to avoid the information loss caused by weight dimension collapse in existing spatial and channel attention mechanisms. On this basis, we construct a multi-level calibration framework guided by neighborhood homogeneity for weight calibration, and design a dedicated loss function to enhance channel discriminability.PointCRA leverages intrinsic feature priors to adaptively correct feature aggregation, offering interpretability with low parameter overhead. Our method is transferable, interpretable, and efficient. We validate the proposed method on diverse datasets and benchmark models, and further demonstrate its rationality through extensive analytical experiments. Our PointCRA achieves 77.5\\% mIoU on the S3DIS dataset, 90.4\\% OA on the ScanObjectNN dataset, and 87.4\\% instance mIoU on the ShapeNetPart dataset. The code and pretrained weights are publicly available on GitHub: https://github.com/AGENT9717/PointCRA",
      "short_abstract": "In 3D point cloud understanding, the core challenge lies in accurately capturing discriminative features within complex neighborhoods, which directly affects the execution precision of downstream tasks such as embodied AI and autonomous driving. Existing...",
      "published": "2026-05-04",
      "updated": "2026-05-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 10,
        "evidence": [
          "title: point cloud",
          "abstract: point cloud"
        ]
      },
      "recency": "archive",
      "age_days": 107,
      "arxiv_url": "https://arxiv.org/abs/2605.02357",
      "pdf_url": "https://arxiv.org/pdf/2605.02357.pdf"
    },
    {
      "id": "2605.02284",
      "anchor": "paper-2605-02284",
      "title": "Beyond Known Objects: A Novel Framework for Open-Set Object Detection using Negative-Aware Norm",
      "authors": [
        "Yuchen Zhang",
        "Yao Lu",
        "Johannes Betz"
      ],
      "abstract": "Open-Set Object Detection (OSOD) is crucial for autonomous driving, where perception systems must recognize and localize both known and previously unseen objects in complex, dynamic environments. While recent approaches deliver promising results, they often require retraining the detector extensively to learn objectness, which describes the likelihood that a bounding box tightly encloses a valid object, regardless of whether its category was learned during training. Deviating from existing work, we hypothesize that standard off-the-shelf detectors may already contain helpful cues for objectness, owing to their training on numerous and diverse known categories. Building on this idea, we propose NAN-SPOT, a training-light framework that does not require to retrain the base object detector and estimates objectness by leveraging a hidden layer metric called Negative-Aware Norm (NAN), requiring only minutes of training on just hundreds of images. To support comprehensive evaluation, we introduce COCO-Open, an expanded version of the existing COCO-Mixed dataset, increasing unknown object annotations from 433 to 1853, making it the most exhaustively labeled dataset for OSOD to the best of our knowledge. Experimental results demonstrate that NAN-SPOT achieves even better performance on unknown object detection than methods requiring heavy training, without compromising performance on known objects. This efficiency and robustness make NAN-SPOT a promising step towards open-world perception in autonomous driving.",
      "short_abstract": "Open-Set Object Detection (OSOD) is crucial for autonomous driving, where perception systems must recognize and localize both known and previously unseen objects in complex, dynamic environments. While recent approaches deliver promising results, they often...",
      "published": "2026-05-04",
      "updated": "2026-05-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 17,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 107,
      "arxiv_url": "https://arxiv.org/abs/2605.02284",
      "pdf_url": "https://arxiv.org/pdf/2605.02284.pdf"
    },
    {
      "id": "2605.02235",
      "anchor": "paper-2605-02235",
      "title": "Distributed Observer-based Fault Detection over Intelligent Networked Multi-Vehicle Systems",
      "authors": [
        "Mohammadreza Doostmohammadian",
        "Hamid R. Rabiee"
      ],
      "abstract": "Decentralized strategies are of interest for local decision-making over multi-vehicle networks. This paper studies mixed traffic networks of human-driven and autonomous vehicles with partial sensor measurements. The idea is to enable the group of connected autonomous vehicles (CAVs) to track the state of a group of human-driven vehicles (HDVs) via distributed consensus-based observers/estimators. Particularly, we make no assumption that the group of HDVs is locally observable in the direct neighborhood of any CAV. Then, the main contribution is to design local residual-based fault detection and isolation (FDI) at every CAV to detect possible faults/attacks in the sensor measurements. This distributed detection strategy enables every CAV to locally find possible anomalies in its taken sensor measurement with no need for a central processing unit. Two FDI logics are proposed with and without considering the history of the residuals. These FDI techniques are based on probabilistic threshold design on the residuals (in contrast to the existing deterministic threshold FDI techniques) with no assumption that the noise is of bounded support. This is more realistic in real-world multi-vehicle transportation systems.",
      "short_abstract": "Decentralized strategies are of interest for local decision-making over multi-vehicle networks. This paper studies mixed traffic networks of human-driven and autonomous vehicles with partial sensor measurements. The idea is to enable the group of connected...",
      "published": "2026-05-04",
      "updated": "2026-05-04",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 2,
        "evidence": [
          "abstract: connected vehicles",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 107,
      "arxiv_url": "https://arxiv.org/abs/2605.02235",
      "pdf_url": "https://arxiv.org/pdf/2605.02235.pdf"
    },
    {
      "id": "2605.01995",
      "anchor": "paper-2605-01995",
      "title": "From Concept to Capability: Evaluating 3D Gaussian Splatting for Synthetic Scene Editing in Autonomous Driving",
      "authors": [
        "Ali Nouri",
        "Yifei Zhang",
        "Yifan Zhang",
        "Tayssir Bouraffa",
        "Zhennan Fei",
        "Zijian Han",
        "Håkan Sivencrona",
        "Anders Heyden"
      ],
      "abstract": "The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while operating in the environment. Field data collection lacks completeness with respect to the list of rare but still possible safety-related scenarios needed for the development, verification, and validation of the ADS. 3D Gaussian Splatting (3DGS) has shown promising capabilities for the reconstruction and editing of scenes based on data collected by cameras and LiDAR sensors. However, the industrial fidelity evaluation of reconstructions is underexplored, which is crucial when employing such methods in safety-related systems, especially for ADS. This becomes more challenging as ADS operates in a dynamic, uncontrolled environment with limited viewpoints and often partially occluded objects. This paper addresses this gap by proposing and implementing a framework (Fig. 1) to systematically analyze the capabilities and limitations of 3DGS for use in the reconstruction of safety-related scenes. It focuses on the quality of reconstruction for vehicles and pedestrians, which are the two most critical object classes for ADS. Our findings provide industry insights into the fidelity degradation of reconstructions from multiple novel viewpoints, both lateral and longitudinal, enabling the integration of these methods into real-world industrial AD software development and testing pipelines.",
      "short_abstract": "The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while operating in the environment. Field data collection lacks completeness with respect to the list of rare but...",
      "published": "2026-05-03",
      "updated": "2026-05-03",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 108,
      "arxiv_url": "https://arxiv.org/abs/2605.01995",
      "pdf_url": "https://arxiv.org/pdf/2605.01995.pdf"
    },
    {
      "id": "2605.01924",
      "anchor": "paper-2605-01924",
      "title": "SimPB++: Simultaneously Detecting 2D and 3D Objects from Multiple Cameras",
      "authors": [
        "Yingqi Tang",
        "Zhaotie Meng",
        "Erkang Cheng",
        "Haibin Ling"
      ],
      "abstract": "Simultaneous perception of 2D objects in perspective view and 3D objects in Bird's Eye View (BEV) is challenging for multi-camera autonomous driving. Existing two-stage pipelines use 2D results only as a one-time cue for 3D detection. We propose SimPB++, which simultaneously detects 2D objects in perspective and 3D objects in BEV from multiple cameras. It unifies both tasks into an end-to-end model with a hybrid decoder architecture, coupling multi-view 2D and 3D decoders interactively. Two novel modules enable deep interaction: Dynamic Query Allocation adaptively assigns 2D queries to 3D candidates, and Adaptive Query Aggregation refines 3D representations using multi-view 2D features, forming a cyclic 3D-2D-3D refinement. For multi-view 2D detection, we use Query-group Attention for intra-group communication. We also design a Crop-and-Scale strategy for long-range perception and a Propagating Denoising strategy with an auxiliary RoI detector. SimPB++ supports mixed supervision with 2D-only and fully annotated data, reducing reliance on expensive 3D labels. Experiments show state-of-the-art performance on nuScenes for both tasks and strong long-range detection (up to 150m) on Argoverse2.",
      "short_abstract": "Simultaneous perception of 2D objects in perspective view and 3D objects in Bird's Eye View (BEV) is challenging for multi-camera autonomous driving. Existing two-stage pipelines use 2D results only as a one-time cue for 3D detection. We propose SimPB++,...",
      "published": "2026-05-03",
      "updated": "2026-05-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 10,
        "evidence": [
          "abstract: perception",
          "title: camera",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 108,
      "arxiv_url": "https://arxiv.org/abs/2605.01924",
      "pdf_url": "https://arxiv.org/pdf/2605.01924.pdf"
    },
    {
      "id": "2605.01916",
      "anchor": "paper-2605-01916",
      "title": "EAPFusion: Intrinsic Evolving Auxiliary Prior Guidance for Infrared and Visible Image Fusion",
      "authors": [
        "Zhenyu Sun",
        "Luobin Zhang",
        "Axi Niu",
        "Haishen Wang",
        "Qingsen Yan"
      ],
      "abstract": "Infrared-visible image fusion aims to create an information-rich fused image by integrating the complementary thermal saliency from infrared sensing and fine textures from visible imaging. Such accurate fusion is essential for real-world perception applications in complex scenes, including nighttime autonomous driving, search and rescue, and surveillance, and can further benefit downstream tasks such as semantic segmentation. However, most existing fusion methods rely upon static trained weights that cannot adapt to scene-specific content at inference time, and often suffer from a granularity mismatch when coarse auxiliary semantics are injected, which makes it difficult to simultaneously highlight targets and preserve details. In this work, we propose EAPFusion to address these issues by using self-evolving intrinsic priors instead of relying on external auxiliary models. Concretely, EAPFusion maintains a compact set of intrinsic priors and progressively updates them across scales. These evolved priors are utilized to dynamically generate convolutional kernels, shifting the paradigm from fixed, pre-trained filters to instance-adaptive parameters via prior-conditioned dynamic convolution. Furthermore, we design a channel-level fusion module that shuffles and interleaves infrared and visible channels, applying local channel mixing to boost cross-modal complementarity. Experiments on different datasets, including cross-dataset evaluation and semantic segmentation, show that the proposed method achieves state-of-the-art quantitative and qualitative fusion results, and consistently boosts downstream performance. Code is coming soon.",
      "short_abstract": "Infrared-visible image fusion aims to create an information-rich fused image by integrating the complementary thermal saliency from infrared sensing and fine textures from visible imaging. Such accurate fusion is essential for real-world perception...",
      "published": "2026-05-03",
      "updated": "2026-05-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 7,
        "evidence": [
          "abstract: perception",
          "abstract: scene segmentation"
        ]
      },
      "recency": "archive",
      "age_days": 108,
      "arxiv_url": "https://arxiv.org/abs/2605.01916",
      "pdf_url": "https://arxiv.org/pdf/2605.01916.pdf"
    },
    {
      "id": "2605.01860",
      "anchor": "paper-2605-01860",
      "title": "Optimizing Trajectory-Trees in Belief Space: An Application from Model Predictive Control to Task and Motion Planning",
      "authors": [
        "Camille Phiquepal",
        "Marc Toussaint"
      ],
      "abstract": "This paper explores the benefits of computing arborescent trajectories (trajectory-trees) instead of commonly used sequential trajectories for partially observable robotic planning problems. In such environments, a robot infers knowledge from observations, and the optimal course of action depends on these observations. Trajectory-trees, optimized in belief space, naturally capture this dependency by branching where the belief state is expected to evolve into multiple distinct scenarios, such as upon receiving an observation. Unlike sequential trajectories, which model a single forward evolution of the system, trajectory-trees capture multiple possible contingencies. First, we focus on Model Predictive Control (MPC) and demonstrate the benefits of planning tree-like trajectories. We formulate the control problem as the optimization of a tree with a single branching (PO-MPC). This improves performance by reducing control costs through more informed planning. To satisfy the real-time constraints of MPC, we develop an optimization algorithm called Distributed Augmented Lagrangian (D-AuLa), which leverages the decomposability of the PO-MPC formulation to parallelize and accelerate the optimization. We apply the method to both linear and non-linear MPC problems using autonomous driving examples. Second, we address Task And Motion Planning (TAMP), and introduce a planner (PO-LGP) reasoning on decision trees at task level, and trajectory-trees at motion-planning level. This approach builds upon the Logic-Geometric-Programming Framework (LGP) and extends it to partially observable problems. The experiments show the method's applicability to problems with a small belief state size, and scales to larger problems by optimizing explorative policies, which are used as macro-actions in an overarching task plan.",
      "short_abstract": "This paper explores the benefits of computing arborescent trajectories (trajectory-trees) instead of commonly used sequential trajectories for partially observable robotic planning problems. In such environments, a robot infers knowledge from observations,...",
      "published": "2026-05-03",
      "updated": "2026-05-03",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 32,
        "margin": 4,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 108,
      "arxiv_url": "https://arxiv.org/abs/2605.01860",
      "pdf_url": "https://arxiv.org/pdf/2605.01860.pdf"
    },
    {
      "id": "2605.01898",
      "anchor": "paper-2605-01898",
      "title": "Fast Newton methods for linear-quadratic dynamic games with application to autonomous vehicle platooning and intersection crossing",
      "authors": [
        "Reza Rahimi Baghbadorani",
        "Sergio Grammatico"
      ],
      "abstract": "We consider constrained linear-quadratic dynamic games arising in autonomous vehicle platooning, intersection crossing and other cooperative driving scenarios. Infinite-horizon Nash equilibria are reformulated as receding-horizon affine variational inequalities with special structure. Exploiting this formulation, we design Newton-type algorithms with local quadratic convergence. The resulting methods achieve extremely fast convergence, making them well suited for real-time and embedded receding-horizon control in safety-critical traffic applications. Simulations of platooning and intersection crossing demonstrate substantial performance gains over first-order and operator-splitting approaches, hence high application potential.",
      "short_abstract": "We consider constrained linear-quadratic dynamic games arising in autonomous vehicle platooning, intersection crossing and other cooperative driving scenarios. Infinite-horizon Nash equilibria are reformulated as receding-horizon affine variational...",
      "published": "2026-05-03",
      "updated": "2026-05-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 2,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "archive",
      "age_days": 108,
      "arxiv_url": "https://arxiv.org/abs/2605.01898",
      "pdf_url": "https://arxiv.org/pdf/2605.01898.pdf"
    },
    {
      "id": "2605.01888",
      "anchor": "paper-2605-01888",
      "title": "AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments",
      "authors": [
        "Xi Zhou",
        "Tao Huang",
        "Qing-Long Han",
        "Rana Abbas",
        "Mostafa Rahimi Azghadi"
      ],
      "abstract": "Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything (V2X) communication to share perceptual data, enhances detection but is vulnerable to channel impairments, such as noise, fading, and interference. To strengthen the reliability of intelligent transportation systems, this work improves the robustness of V2X cooperative perception under communication conditions that reflect common channel impairments. This paper proposes an Adaptive Feature Fusion Transformer (AFFormer), a Transformer-based framework that mitigates the adverse effects of corrupted features by modeling temporal, inter-agent, and spatial correlations. AFFormer introduces three key modules: Multi-Agent and Temporal Aggregation for context-aware fusion across agents and over time, Dual Spatial Attention for efficient modeling of spatial dependencies, and Uncertainty-Guided Fusion for entropy-driven refinement of fused features. A teacher-student knowledge distillation strategy further enhances robustness by aligning fused features with reliable early-collaboration supervision. AFFormer is validated on the V2XSet and DAIR-V2X datasets, where it consistently outperforms existing methods under both ideal and impaired communication conditions, demonstrating improved robustness to communication-induced feature degradation while maintaining a competitive efficiency-accuracy trade-off.",
      "short_abstract": "Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything (V2X) communication to share perceptual data, enhances detection but is vulnerable to channel impairments,...",
      "published": "2026-05-03",
      "updated": "2026-05-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 38,
        "margin": 22,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 108,
      "arxiv_url": "https://arxiv.org/abs/2605.01888",
      "pdf_url": "https://arxiv.org/pdf/2605.01888.pdf"
    },
    {
      "id": "2605.01731",
      "anchor": "paper-2605-01731",
      "title": "Lateral String Stability for Vehicle Platoons: Formulation, Definition, and Analysis",
      "authors": [
        "Sixu Li",
        "Swaroop Darbha",
        "Yang Zhou"
      ],
      "abstract": "Platooning of connected and automated vehicles provides significant benefits in terms of energy efficiency, traffic throughput, and, most critically, safety. These safety benefits depend on string stability, which dictates how disturbances propagate along a vehicle string. Although longitudinal string stability has been extensively examined, lateral string stability, which governs the propagation of path-tracking errors that can lead to unsafe deviations from the desired path, remains underexplored. Its importance is growing as autonomous vehicles increasingly depend on onboard sensing and map-free navigation, where sensor occlusions and tight formations amplify safety risks. This paper presents a framework for lateral string stability that focuses directly on safety-critical, path-relative tracking errors and enables consistent comparison across vehicles that follow the same planned path. The key element of the framework is an arc-length (Eulerian) viewpoint, a departure from traditional analyses, that clarifies how tracking errors at a given point on the path propagate from one vehicle to the next. Building on this foundation, we propose the definition of L2 lateral string stability along with two control strategies: a feedback-feedforward strategy that relies solely on onboard sensing, and a novel learn-from-predecessor strategy that makes use of vehicle-to-vehicle communication. Both strategies are analyzed for lateral string stability with respect to two error measures: tracking error vector and lateral (cross-track) error. Our results show that onboard sensing alone cannot guarantee attenuation of path-tracking errors, imposing a fundamental safety limitation, while V2V communication enables true error attenuation. The analysis further identifies structural controller requirements, showing that nonzero feedback on specific measurements is essential for guaranteeing stability.",
      "short_abstract": "Platooning of connected and automated vehicles provides significant benefits in terms of energy efficiency, traffic throughput, and, most critically, safety. These safety benefits depend on string stability, which dictates how disturbances propagate along a...",
      "published": "2026-05-03",
      "updated": "2026-05-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "abstract: V2X",
          "abstract: connected vehicles",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 108,
      "arxiv_url": "https://arxiv.org/abs/2605.01731",
      "pdf_url": "https://arxiv.org/pdf/2605.01731.pdf"
    },
    {
      "id": "2605.01516",
      "anchor": "paper-2605-01516",
      "title": "Dynamics Distillation for Efficient and Transferable Control Learning",
      "authors": [
        "Xunjiang Gu",
        "Kashyap Chitta",
        "Mahsa Golchoubian",
        "Vladimir Suplin",
        "Igor Gilitschenski"
      ],
      "abstract": "Robust control policy learning for autonomous driving requires training environments to be both physically realistic and computationally scalable, properties that existing simulators provide only in isolation. We introduce Sim2Sim2Sim, a framework that bridges high-fidelity vehicle simulation and scalable reinforcement learning by distilling simulator dynamics into a highly parallelizable learned dynamics model. By training control policies purely within this distilled environment and deploying them back into the high-fidelity source simulator, we demonstrate more efficient policy optimization and reliable transfer under challenging dynamics. We further show that predictive accuracy alone does not fully characterize a learned dynamics model's suitability as a reinforcement learning training environment, which should also be assessed by the quality of the policies it enables.",
      "short_abstract": "Robust control policy learning for autonomous driving requires training environments to be both physically realistic and computationally scalable, properties that existing simulators provide only in isolation. We introduce Sim2Sim2Sim, a framework that...",
      "published": "2026-05-02",
      "updated": "2026-05-02",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 109,
      "arxiv_url": "https://arxiv.org/abs/2605.01516",
      "pdf_url": "https://arxiv.org/pdf/2605.01516.pdf"
    },
    {
      "id": "2605.01478",
      "anchor": "paper-2605-01478",
      "title": "LIE: LiDAR-only HD Map Construction with Intensity Enhancement via Online Knowledge Distillation",
      "authors": [
        "Kanak Mazumder",
        "Fabian B. Flohr"
      ],
      "abstract": "Online High-Definition (HD) map construction is a key component of autonomous driving. Recent methods rely on multi-view camera images for cost-effective HD map segmentation, but cameras lack depth information for accurate scene geometry. In contrast, LiDAR provides precise 3D measurements but lacks dense semantic cues. In this work, we propose LIE, LiDAR-only semantic map construction method that employ Knowledge Distillation (KD) to handle the lack of dense semantic and texture cues. Specifically, the teacher branch fuses student LiDAR features and the corresponding 2D intensity map tile to provide dense supervision for segmenting map elements using online distillation scheme. Experimental results show that our method outperforms all single-modality approaches, achieving 8.2% higher mIoU than the state-of-the-art camera-based model on nuScenes. LIE is robust over long ranges and under challenging weather and lighting, and efficiently adapts to Argoverse2 with only 10% fine-tuning, surpassing camera-based models trained on the full dataset. Source code will be available \\href{https://iv.ee.hm.edu/lie/}{here}.",
      "short_abstract": "Online High-Definition (HD) map construction is a key component of autonomous driving. Recent methods rely on multi-view camera images for cost-effective HD map segmentation, but cameras lack depth information for accurate scene geometry. In contrast, LiDAR...",
      "published": "2026-05-02",
      "updated": "2026-05-02",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 22,
        "evidence": [
          "title: HD map",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 109,
      "arxiv_url": "https://arxiv.org/abs/2605.01478",
      "pdf_url": "https://arxiv.org/pdf/2605.01478.pdf"
    },
    {
      "id": "2606.20576",
      "anchor": "paper-2606-20576",
      "title": "Exploring LLM in Semantic Communication for V2X Networks",
      "authors": [
        "Sihem Bakri",
        "Navdeep Singh",
        "Nicola Marchetti"
      ],
      "abstract": "The rapid growth of connected and autonomous vehicles has created a demand for more efficient and intelligent communication systems. Traditional Vehicle-to-Everything (V2X) networks rely on transmitting raw sensor data, leading to high bandwidth usage and redundant information exchange. To address this, we propose a semantic communication framework that integrates a Large Language Model (LLM) with graph-based knowledge representation, to transmit only high-level, meaningful messages instead of raw data. Within this framework, the LLM performs semantic transformation, converting structured sensor inputs into concise natural language messages that describe context and intent. It also generates high-level control decisions based on shared situational awareness across the V2X network. A multilane traffic simulation was developed to compare semantic and non-semantic modes in terms of bandwidth usage. Results show an average 33.54% reduction in data transmission and illustrate context-aware coordination in representative scenarios.",
      "short_abstract": "The rapid growth of connected and autonomous vehicles has created a demand for more efficient and intelligent communication systems. Traditional Vehicle-to-Everything (V2X) networks rely on transmitting raw sensor data, leading to high bandwidth usage and...",
      "published": "2026-05-02",
      "updated": "2026-05-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 38,
        "margin": 36,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 109,
      "arxiv_url": "https://arxiv.org/abs/2606.20576",
      "pdf_url": "https://arxiv.org/pdf/2606.20576.pdf"
    },
    {
      "id": "2605.01485",
      "anchor": "paper-2605-01485",
      "title": "Cut-In Gap Acceptance Toward Autonomous vs. Human-Driven Vehicles: Evidence from the Waymo Open Motion Dataset",
      "authors": [
        "Abdulaziz Alhuraish",
        "Yuhang Wang",
        "Hao Zhou"
      ],
      "abstract": "Autonomous vehicles (AVs) are widely known to follow conservative, rule-based motion policies that surrounding drivers can learn to anticipate. A direct consequence is that human drivers may accept shorter longitudinal gaps when cutting in front of an AV than when targeting another human-driven vehicle (HDV). We test this hypothesis using the Waymo Open Motion Dataset (WOMD), which provides 25,906 real-world highway scenarios at 10 hertz. An eight-criterion lane-change detector extracts 706 HDV-to-AV and 3,172 HDV-to-HDV cut-in events from the same traffic environment. The median accepted gap in front of the Waymo AV is 7.58 meters versus 9.57 meters for HDV targets, a 1.99 meter reduction that is statistically significant (p equals 5.76 times 10 to the negative eighth power, d equals negative 0.224) and persists under speed-matched resampling. Cut-in speeds toward the AV are 37 percent higher (51.7 versus 37.7 kilometers per hour, d equals 0.502), and 68.0 percent of AV-targeted cut-ins occur below the 10 meter gap boundary versus 51.8 percent of HDV-targeted events (chi-squared equals 60.5, p is less than 10 to the negative thirteenth power). These results reveal a systematic and safety-relevant asymmetry in human gap-acceptance behavior that warrants AV-specific calibration of both motion-planning safety envelopes and traffic simulation models.",
      "short_abstract": "Autonomous vehicles (AVs) are widely known to follow conservative, rule-based motion policies that surrounding drivers can learn to anticipate. A direct consequence is that human drivers may accept shorter longitudinal gaps when cutting in front of an AV than...",
      "published": "2026-05-02",
      "updated": "2026-05-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Human Factors & Policy: abstract: human driver"
        ]
      },
      "recency": "archive",
      "age_days": 109,
      "arxiv_url": "https://arxiv.org/abs/2605.01485",
      "pdf_url": "https://arxiv.org/pdf/2605.01485.pdf"
    },
    {
      "id": "2605.01301",
      "anchor": "paper-2605-01301",
      "title": "From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception",
      "authors": [
        "Qingzhao Zhang",
        "Runting Zhang",
        "Z. Morley Mao"
      ],
      "abstract": "Collaborative perception allows connected and autonomous vehicles (CAVs) to improve perception by sharing sensory data, but it also introduces security risks from manipulated inputs. Prior work shows that attackers can spoof or remove objects by fabricating shared data, yet the practicality of such attacks in real-world driving remains unclear. Existing attacks are often detectable or evaluated in manually constructed scenarios, leaving open whether they can induce safety-critical outcomes in dynamic environments. To bridge this gap, we present a stealthy, scenario-realistic data fabrication attack that induces unsafe driving behaviors through end-to-end system effects. Instead of creating large, easily detectable anomalies, our attack subtly manipulates the poses of existing objects in shared perception results, keeping perturbations below detection thresholds. These small errors are then propagated through downstream modules, including object tracking and trajectory prediction, leading to significant deviations in predicted behaviors and ultimately unsafe driving decisions. We further design an online, scenario-aware attack framework that adapts to dynamic traffic conditions and optimizes attack strategies at runtime. Experiments on OPV2V and V2X-Real demonstrate that the attack achieves over 90% success in inducing detection errors and triggers safety-critical behaviors, such as unnecessary hard braking, in up to 50% of scenarios, while largely evading state-of-the-art defenses. We also propose a mitigation that focuses on detecting anomalies in localized, safety-critical regions, achieving an 80% detection rate on the small pose perturbation compared to 11% for the best existing methods.",
      "short_abstract": "Collaborative perception allows connected and autonomous vehicles (CAVs) to improve perception by sharing sensory data, but it also introduces security risks from manipulated inputs. Prior work shows that attackers can spoof or remove objects by fabricating...",
      "published": "2026-05-02",
      "updated": "2026-05-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 25,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 109,
      "arxiv_url": "https://arxiv.org/abs/2605.01301",
      "pdf_url": "https://arxiv.org/pdf/2605.01301.pdf"
    },
    {
      "id": "2605.08133",
      "anchor": "paper-2605-08133",
      "title": "VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving",
      "authors": [
        "Rui Zhao",
        "Haofeng Hu",
        "Zhenhai Gao",
        "Jiaqiao Liu",
        "Gao Fei"
      ],
      "abstract": "Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalization in long-tail scenarios. While Retrieval-Augmented Generation (RAG) offers a solution by accessing external expert priors, standard visual retrieval suffers from high latency and semantic ambiguity. To address these challenges, we propose \\textbf{VLADriver-RAG}, a framework that grounds planning in explicit, structure-aware historical knowledge. Specifically, we abstract sensory inputs into spatiotemporal semantic graphs via a \\textit{Visual-to-Scenario} mechanism, effectively filtering visual noise. To ensure retrieval relevance, we employ a \\textit{Scenario-Aligned Embedding Model} that utilizes Graph-DTW metric alignment to prioritize intrinsic topological consistency over superficial visual similarity. These retrieved priors are then fused within a query-based VLA backbone to synthesize precise, disentangled trajectories. Extensive experiments on the Bench2Drive benchmark establish a new state-of-the-art, achieving a Driving Score of 89.12.",
      "short_abstract": "Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalization in long-tail scenarios. While Retrieval-Augmented Generation (RAG) offers a...",
      "published": "2026-05-01",
      "updated": "2026-05-01",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 30,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 110,
      "arxiv_url": "https://arxiv.org/abs/2605.08133",
      "pdf_url": "https://arxiv.org/pdf/2605.08133.pdf"
    },
    {
      "id": "2605.01081",
      "anchor": "paper-2605-01081",
      "title": "WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather",
      "authors": [
        "Hamed Khatounabadi",
        "Xiaohu Lu",
        "Hayder Radha"
      ],
      "abstract": "The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for autonomous vehicles. Recent efforts address this problem by relying on synthetic data to train the object detectors, which limits their real-world applicability. Meanwhile, pseudo-labeling is widely used for cross-dataset domain adaptation problems. However, these methods have not been exploited by weather-based domain adaptation approaches due to the noisy nature of such labels generated under harsh weather conditions. In this paper, we propose two new approaches to mitigate this weather-induced domain shift. First, we propose a Weather-Induced pseudo Label Denoising (WILD) framework that filters noisy pseudo labels generated by real data captured under adverse weather conditions. Second, we develop a novel hybrid training methodology, WILD SAM, that exploits both pseudo-label denoising and simulation-based training solutions while using real-data from the target harsh-weather domain. We validate both proposed approaches, WILD and WILD SAM, on the recently released Four Seasons dataset across rainy and snowy scenarios. Experiments show that the proposed frameworks improve Average Precision (AP) up to 13\\% and significantly reduce the weather-induced performance gap relative to the baseline. The code is available at: https://github.com/Kh-Hamed/WILD-SAM",
      "short_abstract": "The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for autonomous vehicles. Recent efforts address this problem by relying on synthetic data to train the object...",
      "published": "2026-05-01",
      "updated": "2026-05-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 5,
        "evidence": [
          "title: perception"
        ]
      },
      "recency": "archive",
      "age_days": 110,
      "arxiv_url": "https://arxiv.org/abs/2605.01081",
      "pdf_url": "https://arxiv.org/pdf/2605.01081.pdf"
    },
    {
      "id": "2605.00781",
      "anchor": "paper-2605-00781",
      "title": "Map2World: Segment Map Conditioned Text to 3D World Generation",
      "authors": [
        "Jaeyoung Chung",
        "Suyoung Lee",
        "Jianfeng Xiang",
        "Jiaolong Yang",
        "Kyoung Mu Lee"
      ],
      "abstract": "3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user-defined segment maps of arbitrary shapes and scales, ensuring global-scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine-grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user-controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.",
      "short_abstract": "3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from...",
      "published": "2026-05-01",
      "updated": "2026-05-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 110,
      "arxiv_url": "https://arxiv.org/abs/2605.00781",
      "pdf_url": "https://arxiv.org/pdf/2605.00781.pdf"
    },
    {
      "id": "2605.00531",
      "anchor": "paper-2605-00531",
      "title": "From Research to Practice: An Interactive Rapid Review of Autonomous Driving System Testing in Industry",
      "authors": [
        "Qunying Song",
        "Ali Nouri",
        "Håkan Sivencrona",
        "Federica Sarro"
      ],
      "abstract": "Autonomous driving systems (ADS) are increasingly deployed in real traffic, yet testing remains fundamentally challenging due to open environments, complex scenarios, and the lack of established processes and metrics. Despite extensive research, a gap persists between academic advances and their applicability in industrial practice. To address this, we conduct an interactive rapid review in collaboration with 21 practitioners from a leading automotive company. Practitioners identified 12 key challenges in ADS testing, and prioritised two as the most critical issues, namely approaches to and completeness of testing for End-to-End (E2E) ADS. We analyzed 17 research studies relevant to these two challenges, most of which focus on generating critical testing scenarios, and subsequently assessed their relevance and applicability in practice. Our study provides the first practitioner-driven review and evaluation of current ADS testing research, reveals practical challenges in ADS testing, offers rapid insights for practitioners, and highlights the need for more context-aware, industry-relevant solutions to bridge the gap between research and practice.",
      "short_abstract": "Autonomous driving systems (ADS) are increasingly deployed in real traffic, yet testing remains fundamentally challenging due to open environments, complex scenarios, and the lack of established processes and metrics. Despite extensive research, a gap...",
      "published": "2026-05-01",
      "updated": "2026-05-01",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Survey"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 9,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 110,
      "arxiv_url": "https://arxiv.org/abs/2605.00531",
      "pdf_url": "https://arxiv.org/pdf/2605.00531.pdf"
    },
    {
      "id": "2605.00412",
      "anchor": "paper-2605-00412",
      "title": "Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling",
      "authors": [
        "Sen Cui",
        "Jingheng Ma"
      ],
      "abstract": "World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world model research is often dominated by three partially separated routes: 2D video-generative models that emphasize visual future synthesis, 3D scene-centric models that emphasize spatial reconstruction, and JEPA-like latent models that emphasize abstract predictive representations. While each route has made important progress, they still struggle to provide physically reliable, action-controllable, and long-horizon stable predictions for embodied decision making. In this paper, we argue that the bottleneck of world models is no longer only whether they can generate realistic futures, but whether those futures are physically meaningful and useful for action. We propose \\emph{Hamiltonian World Models} as a physically grounded perspective on world modeling. The key idea is to encode observations into a structured latent phase space, evolve the latent state through Hamiltonian-inspired dynamics with control, dissipation, and residual terms, decode the predicted trajectory into future observations, and use the resulting rollouts for planning. We discuss how Hamiltonian structure may improve interpretability, data efficiency, and long-horizon stability, while also noting practical challenges in real-world robotic scenes involving friction, contact, non-conservative forces, and deformable objects.",
      "short_abstract": "World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world model research is often dominated by three partially separated routes: 2D...",
      "published": "2026-05-01",
      "updated": "2026-05-01",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 9,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "archive",
      "age_days": 110,
      "arxiv_url": "https://arxiv.org/abs/2605.00412",
      "pdf_url": "https://arxiv.org/pdf/2605.00412.pdf"
    },
    {
      "id": "2605.00595",
      "anchor": "paper-2605-00595",
      "title": "Robust Fusion of Object-Level V2X for Learned 3D Object Detection",
      "authors": [
        "Lukas Ostendorf",
        "Lennart Reiher",
        "Onn Haran",
        "Lutz Eckstein"
      ],
      "abstract": "Perception for automated driving is largely based on onboard environmental sensors, such as cameras and radar, which are cost-effective but limited by line-of-sight and field-of-view constraints. These inherent limitations may cause onboard perception to fail under occlusions or poor visibility conditions. In parallel, cooperative awareness via vehicle-to-everything (V2X) communication is becoming increasingly available, enabling vehicles and infrastructure to share their own state as object-level information that complements onboard perception. In this work, we study how such V2X information can be integrated into 3D object detection and how robust the resulting system is to realistic V2X imperfections. Using the nuScenes dataset, we emulate object-level cooperative awareness messages from ground truth, injecting controlled noise and object dropout to mimic real-world conditions such as latency, localization errors, and low V2X penetration rates. We convert these messages into a dedicated bird's-eye view (BEV) input and fuse them into a BEVFusion-style detector. Our results demonstrate that while object-level cooperative information can substantially improve detection performance, achieving an NDS of 0.80 under favorable conditions, models trained on idealized data become fragile and over-reliant on V2X. Conversely, our proposed noise-aware training strategy, coupled with explicit confidence encoding, enhances robustness, maintaining performance gains even under severe noise and reduced V2X penetration.",
      "short_abstract": "Perception for automated driving is largely based on onboard environmental sensors, such as cameras and radar, which are cost-effective but limited by line-of-sight and field-of-view constraints. These inherent limitations may cause onboard perception to fail...",
      "published": "2026-05-01",
      "updated": "2026-05-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 31,
        "margin": 5,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 110,
      "arxiv_url": "https://arxiv.org/abs/2605.00595",
      "pdf_url": "https://arxiv.org/pdf/2605.00595.pdf"
    },
    {
      "id": "2605.00556",
      "anchor": "paper-2605-00556",
      "title": "Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving",
      "authors": [
        "Ashwin George",
        "Lucas Elbert Suryana",
        "Lorenzo Flipse",
        "Bart van Arem",
        "David A. Abbink",
        "Simeon Craig Calvert",
        "Luciano Cavalcante Siebert",
        "Arkady Zgonnikov"
      ],
      "abstract": "Partial driving automation creates a tension: drivers remain legally responsible for vehicle behaviour, yet their active control is significantly reduced. This reduction undermines the engagement and sense of agency needed to intervene safely. Meaningful human control (MHC) has been proposed as a normative framework to address this tension. However, empirical methods for evaluating whether existing systems actually provide MHC remain underdeveloped. In this study, we investigated the extent to which drivers experience MHC when interacting with partially automated driving systems. Twenty-four drivers completed a simulator study involving silent automation failures under two modes - haptic shared control (HSC) and traded control (TC). We derived behavioural metrics from telemetry data, subjective perception scores from post-trial surveys and used them to test hypothesised relations between them derived from the properties of systems under MHC. The confirmatory analysis showed a significant negative correlation between the perception of the automated vehicle (AV) understanding the driver and conflict in steering torques. An exploratory analysis also revealed a surprising positive correlation between reaction times and the perception of sufficient control. Qualitative feedback from open-ended post-experiment questionnaires revealed that mismatches in intentions between the driver and automation, lack of safety, and resistance to driver inputs contribute to the reduction of perceived MHC, while subtle haptic guidance aligned with driver intent had a positive effect. These findings suggest that future designs should prioritise effortless driver interventions, transparent communication of automation intent, and context-sensitive authority allocation to strengthen meaningful human control in partially automated driving.",
      "short_abstract": "Partial driving automation creates a tension: drivers remain legally responsible for vehicle behaviour, yet their active control is significantly reduced. This reduction undermines the engagement and sense of agency needed to intervene safely. Meaningful...",
      "published": "2026-05-01",
      "updated": "2026-05-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 2,
        "evidence": [
          "title: perception",
          "abstract: perception"
        ]
      },
      "recency": "archive",
      "age_days": 110,
      "arxiv_url": "https://arxiv.org/abs/2605.00556",
      "pdf_url": "https://arxiv.org/pdf/2605.00556.pdf"
    },
    {
      "id": "2605.00362",
      "anchor": "paper-2605-00362",
      "title": "Time-series Meets Complex Motion Modeling: Robust and Computational-effective Motion Predictor for Multi-object Tracking",
      "authors": [
        "Nhat-Tan Do",
        "Le-Huy Tu",
        "Nhi Ngoc-Yen Nguyen",
        "Dieu-Phuong Nguyen",
        "Trong-Hop Do"
      ],
      "abstract": "Multi-object tracking (MOT) is critical in numerous real-world applications, including surveillance, autonomous driving, and robotics. Accurately predicting object motion is fundamental to MOT, but current methods struggle with the complexities of real-world, non-linear motion (e.g., sudden stops, sharp turns). While recent research has gravitated towards increasingly complex and computationally expensive generative models to tackle this problem, their practical utility is often constrained. This paper challenges that paradigm, arguing that such complexity is not only unnecessary but can be outperformed by a more efficient, purpose-built approach. We introduce the Temporal Convolutional Motion Predictor (TCMP), a novel framework for MOT that leverages a modified Temporal Convolutional Network (TCN) featuring dilated convolutions and a regression head. This design allows for effective motion prediction across arbitrary temporal context lengths. Experimental results demonstrate that our approach achieves state-of-the-art performance, specifically improves upon the previous best method in several key metrics: HOTA (a measure of overall tracking accuracy) increases from 62.3% to 63.4%, IDF1 (a measure of identity preservation) rises from 63.0% to 65.0%, and AssA (a measure of association accuracy) improves from 47.2% to 49.1%. Significantly, TCMP achieves this performance while being highly efficient; it has only 0.014 times the parameters and requires only 0.05 times the computational cost (FLOPs) compared to the SOTA method. while is only 0.014 times the size (in terms of parameters) and requires only 0.05 times the computational cost (in terms of FLOPs). These findings highlight the robustness of our method to advance MOT systems by ensuring adaptability, accuracy, and efficiency in complex tracking environments.",
      "short_abstract": "Multi-object tracking (MOT) is critical in numerous real-world applications, including surveillance, autonomous driving, and robotics. Accurately predicting object motion is fundamental to MOT, but current methods struggle with the complexities of real-world,...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 1,
        "evidence": [
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2605.00362",
      "pdf_url": "https://arxiv.org/pdf/2605.00362.pdf"
    },
    {
      "id": "2605.00291",
      "anchor": "paper-2605-00291",
      "title": "An End-to-End Decision-Aware Multi-Scale Attention-Based Model for Explainable Autonomous Driving",
      "authors": [
        "Maryam Sadat Hosseini Azad",
        "Shahriar Baradaran Shokouhi",
        "Amir Abbas Hamidi Imani",
        "Shahin Atakishiyev",
        "Randy Goebel"
      ],
      "abstract": "The application of computer vision is gradually increasing across various domains. They employ deep learning models with a black-box nature. Without the ability to explain the behavior of neural networks, especially their decision-making processes, it is not possible to recognize their efficiency, predict system failures, or effectively implement them in real-world applications. Due to the inevitable use of deep learning in fully automated driving systems, many methods have been proposed to explain their behavior; however, they suffer from flawed reasoning and unreliable metrics, which have prevented a comprehensive understanding of complex models in autonomous vehicles and hindered the development of truly reliable systems. In this study, we propose a multi-scale attention-based model in which driving decisions are fed into the reasoning component to provide case-specific explanations for each decision simultaneously. For quantitative evaluation of our model's performance, we employ the F1-score metric, and also proposed a new metric called the Joint F1 score to demonstrate the accurate and reliable performance of the model in terms of Explainable Artificial Intelligence (XAI). In addition to the BDD-OIA dataset, the nu-AR dataset is utilized to further validate the generalization capability and robustness of the proposed network. The results demonstrate the superiority of our reasoning network over the classic and state-of-the-art models.",
      "short_abstract": "The application of computer vision is gradually increasing across various domains. They employ deep learning models with a black-box nature. Without the ability to explain the behavior of neural networks, especially their decision-making processes, it is not...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 3,
        "evidence": [
          "title: explainability",
          "abstract: explainability"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2605.00291",
      "pdf_url": "https://arxiv.org/pdf/2605.00291.pdf"
    },
    {
      "id": "2605.00066",
      "anchor": "paper-2605-00066",
      "title": "Do Open-Loop Metrics Predict Closed-Loop Driving? A Cross-Benchmark Correlation Study of NAVSIM and Bench2Drive",
      "authors": [
        "Yiru Wang",
        "Anqing Jiang",
        "Shuo Wang",
        "Yuwen Heng",
        "Hai Yang",
        "Yang Chen",
        "Hao Sun"
      ],
      "abstract": "Open-loop evaluation offers fast, reproducible assessment of autonomous driving planners, but its ability to predict real closed-loop driving performance remains questionable. Prior work has shown that traditional open-loop metrics such as Average Displacement Error (ADE) and Final Displacement Error (FDE) exhibit no reliable correlation with closed-loop Driving Score. In this paper, we ask whether the more recent, safety-aware open-loop metrics introduced by NAVSIM~v2 can bridge this gap. By systematically cross-referencing published results from 15 state-of-the-art methods across NAVSIM (open-loop) and Bench2Drive (closed-loop), we compile a paired dataset of open-loop sub-metrics and closed-loop performance, yielding 8 methods with complete paired data. Our analysis reveals three key findings: (1) the aggregate NAVSIM PDM Score shows a strong positive but non-monotonic correlation with Bench2Drive Driving Score, with clear ranking inversions; (2) among individual NAVSIM sub-metrics, Ego Progress (EP) is the strongest single predictor of closed-loop success, substantially exceeding the safety-critical collision metric NC; (3) the safety-progress trade-off manifests differently in open-loop and closed-loop: methods that maximize safety at the expense of progress rank highly in NAVSIM but underperform in closed-loop due to timeout and slow-driving penalties. We further demonstrate that a much simpler 3-metric formula matches the predictive power of the full 5-metric PDMS at the same Spearman $ρ{=}0.90$ on our paired sample of $n{=}8$ methods, suggesting that within current state-of-the-art methods -- where TTC and Comfort approach saturation -- these two sub-metrics add little marginal information for closed-loop ranking. Additionally, we identify the snowball effect -- where small open-loop deviations compound into closed-loop failures -- as a candidate mechanism for the residual gap.",
      "short_abstract": "Open-loop evaluation offers fast, reproducible assessment of autonomous driving planners, but its ability to predict real closed-loop driving performance remains questionable. Prior work has shown that traditional open-loop metrics such as Average...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 6,
        "evidence": [
          "abstract: safety",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2605.00066",
      "pdf_url": "https://arxiv.org/pdf/2605.00066.pdf"
    },
    {
      "id": "2604.28196",
      "anchor": "paper-2604-28196",
      "title": "HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation",
      "authors": [
        "Xin Zhou",
        "Dingkang Liang",
        "Xiwu Chen",
        "Feiyang Tan",
        "Dingyuan Zhang",
        "Hengshuang Zhao",
        "Xiang Bai"
      ],
      "abstract": "Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene generation, often overlooking comprehensive 3D scene understanding. Conversely, while Large Language Models (LLMs) demonstrate impressive reasoning capabilities, they lack the capacity to predict future geometric evolution, creating a significant disparity between semantic interpretation and physical simulation. To bridge this gap, we propose HERMES++, a unified driving world model that integrates 3D scene understanding and future geometry prediction within a single framework. Our approach addresses the distinct requirements of these tasks through synergistic designs. First, a BEV representation consolidates multi-view spatial information into a structure compatible with LLMs. Second, we introduce LLM-enhanced world queries to facilitate knowledge transfer from the understanding branch. Third, a Current-to-Future Link is designed to bridge the temporal gap, conditioning geometric evolution on semantic context. Finally, to enforce structural integrity, we employ a Joint Geometric Optimization strategy that integrates explicit geometric constraints with implicit latent regularization to align internal representations with geometry-aware priors. Extensive evaluations on multiple benchmarks validate the effectiveness of our method. HERMES++ achieves strong performance, outperforming specialist approaches in both future point cloud prediction and 3D scene understanding tasks. The model and code will be publicly released at https://github.com/H-EmbodVis/HERMESV2.",
      "short_abstract": "Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene generation, often overlooking comprehensive 3D scene understanding. Conversely,...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Synthetic Data",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 18,
        "margin": 1,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2604.28196",
      "pdf_url": "https://arxiv.org/pdf/2604.28196.pdf"
    },
    {
      "id": "2604.28111",
      "anchor": "paper-2604-28111",
      "title": "GSDrive: Reinforcing Driving Policies by Multi-mode Future Trajectory Probing with 3D Gaussian Splatting Environment",
      "authors": [
        "Ziang Guo",
        "Chen Min",
        "Xuefeng Zhang",
        "Yixiao Zhou",
        "Shuo Wang",
        "Sifa Zheng",
        "Dzmitry Tsetserukou",
        "Zufeng Zhang"
      ],
      "abstract": "End-to-end (E2E) autonomous driving aims to directly map sensory observations to driving actions, but its real-world deployment is hindered by evolving data distributions and the high cost of continual annotation. While combining imitation learning (IL) and reinforcement learning (RL) is a common strategy for policy improvement, conventional RL training relies on delayed, event-based rewards, where policies learn only from catastrophic outcomes such as collisions, leading to premature convergence to suboptimal behaviors. To address these limitations, we propose GSDrive, a framework that uses a differentiable 3D Gaussian Splatting (3DGS) environment for future-aware trajectory probing and reward shaping in E2E driving. GSDrive first learns a multi-mode trajectory probe via IL and then uses RL to evaluate multiple candidate futures in the 3DGS environment, converting their simulated returns into dense shaping rewards for policy optimization. This yields a cyclic hybrid IL-RL training loop, where IL supplies structured future priors and RL provides interactive feedback for iterative refinement. Evaluated on the reconstructed nuScenes dataset, our method outperforms other simulation-based RL approaches in closed-loop experiments. Code is available at https://github.com/ZionGo6/GSDrive.",
      "short_abstract": "End-to-end (E2E) autonomous driving aims to directly map sensory observations to driving actions, but its real-world deployment is hindered by evolving data distributions and the high cost of continual annotation. While combining imitation learning (IL) and...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2604.28111",
      "pdf_url": "https://arxiv.org/pdf/2604.28111.pdf"
    },
    {
      "id": "2604.28087",
      "anchor": "paper-2604-28087",
      "title": "Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles",
      "authors": [
        "Zainab Rehan",
        "Christian Medeiros Adriano",
        "Sona Ghahremani",
        "Holger Giese"
      ],
      "abstract": "Rule-based systems remain central in safety-critical domains but often struggle with scalability, brittleness, and goal misspecification. These limitations can lead to reward hacking and failures in formal verification, as AI systems tend to optimize for narrow objectives. In previous research, we developed a neuro-symbolic causal framework that integrates first-order logic abduction trees, structural causal models, and deep reinforcement learning within a MAPE-K loop to provide explainable adaptations under distribution shifts. In this paper, we extend that framework by introducing a meta-level layer designed to mitigate goal misspecification and support scalable rule maintenance. This layer consists of a Goal/Rule Synthesizer and a Rule Verification Engine, which iteratively refine a formal rule theory from high-level natural-language goals and principles provided by human experts. The synthesis pipeline employs large language models (LLMs) to: (1) decompose goals into candidate causes, (2) consolidate semantics to remove redundancies, (3) translate them into candidate first-order rules, and (4) compose necessary and sufficient causal sets. The verification pipeline then performs (1) syntax and schema validation, (2) logical consistency analysis, and (3) safety and invariant checks before integrating verified rules into the knowledge base. We evaluated our approach with a proof-of-concept implementation in two autonomous driving scenarios. Results indicate that, given human-specified goals and principles, the pipeline can successfully derive minimal necessary and sufficient rule sets and formalize them as logical constraints. These findings suggest that the pipeline supports incremental, modular, and traceable rule synthesis grounded in established legal and safety principles.",
      "short_abstract": "Rule-based systems remain central in safety-critical domains but often struggle with scalability, brittleness, and goal misspecification. These limitations can lead to reward hacking and failures in formal verification, as AI systems tend to optimize for...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 22,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: verification",
          "abstract: verification",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2604.28087",
      "pdf_url": "https://arxiv.org/pdf/2604.28087.pdf"
    },
    {
      "id": "2604.27499",
      "anchor": "paper-2604-27499",
      "title": "Towards All-Day Perception for Off-Road Driving: A Large-Scale Multispectral Dataset and Comprehensive Benchmark",
      "authors": [
        "Shuo Wang",
        "Jilin Mei",
        "Wenfei Guan",
        "Shuai Wang",
        "Yan Xing",
        "Chen Min",
        "Yu Hu"
      ],
      "abstract": "Off-road nighttime autonomous driving suffers from unreliable visible-light perception, making infrared modality crucial for accurate freespace detection. However, progress remains limited due to the scarcity of annotated infrared off-road datasets and the inter-frame inconsistencies inherent to current single-frame methods. To address these gaps, we present the IRON dataset, which, to our knowledge, is the first large-scale infrared dataset for off-road temporal freespace detection under all-day conditions, with strong support for nighttime perception. The dataset comprises 24,314 densely annotated infrared images with synchronized RGB images in diverse scenes and different light conditions. Building upon this dataset, we propose IRONet, a novel flow-free framework for temporal freespace detection that addresses inter-frame inconsistencies by aggregating historical context via a memory-attention mechanism and a carefully designed mask decoder. On our IRON dataset, IRONet achieves state-of-the-art performance, reaching 82.93%(+1.19%) IoU and 90.66%(+0.71%) F1 score at real-time inference. Remarkably, IRONet also exhibits robust generalization to RGB modalities on ORFD and Rellis datasets. Overall, our work establishes a foundation for reliable all-day off-road autonomous driving and future research in infrared temporal perception. The code and IRON dataset are available at https://github.com/wsnbws/IRON.",
      "short_abstract": "Off-road nighttime autonomous driving suffers from unreliable visible-light perception, making infrared modality crucial for accurate freespace detection. However, progress remains limited due to the scarcity of annotated infrared off-road datasets and the...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 9,
        "evidence": [
          "title: perception",
          "abstract: perception"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2604.27499",
      "pdf_url": "https://arxiv.org/pdf/2604.27499.pdf"
    },
    {
      "id": "2604.27414",
      "anchor": "paper-2604-27414",
      "title": "Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis",
      "authors": [
        "David Fernandez",
        "Pedram MohajerAnsari",
        "Amir Salarpour",
        "Mert D. Pese"
      ],
      "abstract": "Vision-language models (VLMs) are increasingly used in autonomous driving because they combine visual perception with language-based reasoning, supporting more interpretable decision-making, yet their robustness to physical adversarial attacks, especially whether such attacks transfer across different VLM architectures, is not well understood and poses a practical risk when attackers do not know which model a vehicle uses. We address this gap with a systematic cross-architecture study of adversarial transferability in VLM-based driving, evaluating three representative architectures (Dolphins, OmniDrive, and LeapVAD) using physically realizable patches placed on roadside infrastructure in both crosswalk and highway scenarios. Our transfer-matrix evaluation shows high cross-architecture effectiveness, with transfer rates of 73-91% (mean TR = 0.815 for crosswalk and 0.833 for highway) and sustained frame-level manipulation over 64.7-79.4% of the critical decision window even when patches are not optimized for the target model.",
      "short_abstract": "Vision-language models (VLMs) are increasingly used in autonomous driving because they combine visual perception with language-based reasoning, supporting more interpretable decision-making, yet their robustness to physical adversarial attacks, especially...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 14,
        "evidence": [
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2604.27414",
      "pdf_url": "https://arxiv.org/pdf/2604.27414.pdf"
    },
    {
      "id": "2604.28084",
      "anchor": "paper-2604-28084",
      "title": "Intelligent Self-tuning Active EMI Filtering for Electrified Automotive Power Systems Using Reinforcement Learning",
      "authors": [
        "Mahuizi Lu",
        "Kelin Jia",
        "Rajib Goswami",
        "Yukun Hu"
      ],
      "abstract": "The rapid electrification and intelligence of modern transportation systems place stringent demands on the electromagnetic compatibility, reliability, and adaptability of automotive power electronics. In electric and autonomous vehicles, electromagnetic interference (EMI) generated by high-frequency switching power converters can compromise safety-critical functions, in-vehicle communications, and system efficiency under dynamic operating conditions. Conventional passive EMI filters, while robust, are often oversized and lack adaptability, leading to increased weight, volume, and energy losses. This paper proposes an intelligent self-tuning active EMI filtering approach for electrified automotive power systems based on reinforcement learning (RL). The EMI mitigation problem is formulated as a Markov decision process, enabling an RL agent to continuously adapt filter parameters in response to time-varying interference characteristics. To improve robustness and generalisation under complex and non-stationary conditions, a variational autoencoder is employed for compact state representation, while a noise-based exploration mechanism enhances learning efficiency and prevents suboptimal convergence. The proposed method is evaluated using experimentally measured EMI spectra from an automotive electric drive unit within a MATLAB/Simulink co-simulation framework. Results demonstrate consistent EMI attenuation improvements of 25-30 dB across a wide frequency range compared with conventional control strategies and passive filtering solutions. By reducing reliance on oversized passive components and enabling adaptive EMI suppression, the proposed framework supports lightweight, energy-efficient, and reliable power-electronic systems for intelligent and green transportation applications.",
      "short_abstract": "The rapid electrification and intelligence of modern transportation systems place stringent demands on the electromagnetic compatibility, reliability, and adaptability of automotive power electronics. In electric and autonomous vehicles, electromagnetic...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2604.28084",
      "pdf_url": "https://arxiv.org/pdf/2604.28084.pdf"
    },
    {
      "id": "2604.28057",
      "anchor": "paper-2604-28057",
      "title": "Framework for Collaborative Operation of Autonomous Delivery Vehicles Within a Marshaling Yard",
      "authors": [
        "James O'Hara",
        "Karl Wunderlich",
        "Gregory Stevens"
      ],
      "abstract": "As autonomous vehicles slowly deploy into urban roads for limited use cases with significant edge case issues, closed facilities like marshaling yards provide a ripe case for combining lower-level vehicle autonomy with fixed infrastructure to create full autonomy without similar edge case concerns. Within a delivery marshaling yard, electric fleet vehicles complete a set of sequential tasks (charging, inspection, cleaning, and loading) before exiting the yard with their new load of deliveries. Hybrid automation of the vehicles and infrastructure can allow these vehicles to reach full autonomy and navigate the facility without the need of a driver, allowing for quicker movement between tasks increasing vehicle throughput. However, isolated autonomous operations based on static rules are prone to gridlock causing facility failures that temporarily shut down operations. Our orchestrated autonomy solution uses decentralized, dynamic priority scoring of vehicles based on the current status of the marshaling yard to optimally assign vehicles to tasks to increase vehicle throughput. Using a simulated facility with three marshaling yard sizes (small, medium, and large) and three demand levels (low, medium, high), we demonstrated that our orchestration solution increases vehicle throughput above static, isolated autonomy for all combinations of yard size and demand, while reducing facility failures at high demand levels.",
      "short_abstract": "As autonomous vehicles slowly deploy into urban roads for limited use cases with significant edge case issues, closed facilities like marshaling yards provide a ripe case for combining lower-level vehicle autonomy with fixed infrastructure to create full...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 7,
        "evidence": [
          "title: cooperative systems"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2604.28057",
      "pdf_url": "https://arxiv.org/pdf/2604.28057.pdf"
    },
    {
      "id": "2605.00377",
      "anchor": "paper-2605-00377",
      "title": "An eHMI Presenting Request-to-Intervene and Takeover Status of Level 3 Automated Vehicles to Support Surrounding Traffic Safety",
      "authors": [
        "Hailong Liu",
        "Masaki Kuge",
        "Toshihiro Hiraoka",
        "Takahiro Wada"
      ],
      "abstract": "Level 3 automated vehicles (AVs) issue a request to intervene (RtI) when the automated driving system approaches its system limitations. Although this takeover transition is safety-critical, it is usually invisible to surrounding manually driven vehicle (MV) drivers. This study proposes an external human-machine interface (eHMI) called eHMI C+O that externalizes the RtI-related takeover status of a Level~3 AV using cyan and orange light bars. A driving-simulator experiment with 40 participants examined whether the proposed eHMI supports surrounding MV drivers during AV takeover scenarios. The results showed that, compared with the ADS-status-only eHMI condition, which is similar to ``Automated Driving Marker Lights,'' and the no-eHMI condition, the proposed eHMI C+O significantly improved participants' understanding of the AV's driving intention, their prediction of its behavior, and their perceived sufficiency of the information presented by the AV. It also reduced hesitation, increased confidence, and promoted earlier and larger increases in time headway after the RtI was issued. In the AV accident scenario, eHMI C+O significantly reduced the odds of accident involvement for the following MV compared with the no-eHMI condition, corresponding to a 76.8% reduction in accident odds. Exploratory path analysis suggested that the safety benefit of the proposed eHMI C+O may be associated with improved situation awareness and earlier defensive driving responses. These findings indicate that externalizing RtI-related takeover status can help surrounding drivers better understand Level 3 AVs and respond more safely during safety-critical takeover transitions.",
      "short_abstract": "Level 3 automated vehicles (AVs) issue a request to intervene (RtI) when the automated driving system approaches its system limitations. Although this takeover transition is safety-critical, it is usually invisible to surrounding manually driven vehicle (MV)...",
      "published": "2026-04-30",
      "updated": "2026-04-30",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 9,
        "evidence": [
          "abstract: human-machine interaction",
          "title: takeover",
          "abstract: takeover"
        ]
      },
      "recency": "archive",
      "age_days": 111,
      "arxiv_url": "https://arxiv.org/abs/2605.00377",
      "pdf_url": "https://arxiv.org/pdf/2605.00377.pdf"
    },
    {
      "id": "2605.00907",
      "anchor": "paper-2605-00907",
      "title": "TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation",
      "authors": [
        "Han Gong",
        "Zhen Zhou",
        "Yunyang Shi",
        "Yan Tan",
        "Jinbiao Huo",
        "Qi Hong",
        "Zhiyuan Liu"
      ],
      "abstract": "Large language models (LLMs) and multimodal large models (MLLMs) are increasingly used for transportation tasks such as regulation question answering, traffic management support, engineering review, and autonomous-driving scene reasoning. Yet transportation workflows are rule-intensive, computation-intensive, safety-critical, and inherently multimodal. Existing general benchmarks provide limited evidence of whether a model can apply regulations correctly, perform verifiable engineering calculations, or interpret traffic scenes reliably, while the small number of public transportation benchmarks remain narrow in scope and rarely support fine-grained diagnosis across text, images, and point-cloud data. To address this gap, we present TRIP-Evaluate, an open multimodal benchmark for large models in transportation. The benchmark organizes 837 items using a role-task-knowledge taxonomy that covers vehicle, traffic-management, traveler, and planning-and-design functions. Each item is annotated with capability, modality, and difficulty labels, enabling diagnosis from overall accuracy down to specific failure modes. The current release includes 596 text items, 198 image items, and 43 point-cloud items. TRIP-Evaluate also standardizes item construction, quality control, prompting, decoding, and scoring to improve cross-model comparability. Results on a diverse panel of models show that text-based performance is improving, but substantial weaknesses remain in multi-step engineering calculation, rule-constrained reasoning, multimodal scene understanding, and point-cloud understanding. Overall, TRIP-Evaluate provides a reproducible, diagnosable, and engineering-aligned evaluation baseline for model selection, regression testing, and safer deployment in transportation applications.",
      "short_abstract": "Large language models (LLMs) and multimodal large models (MLLMs) are increasingly used for transportation tasks such as regulation question answering, traffic management support, engineering review, and autonomous-driving scene reasoning. Yet transportation...",
      "published": "2026-04-29",
      "updated": "2026-04-29",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 112,
      "arxiv_url": "https://arxiv.org/abs/2605.00907",
      "pdf_url": "https://arxiv.org/pdf/2605.00907.pdf"
    },
    {
      "id": "2605.00051",
      "anchor": "paper-2605-00051",
      "title": "Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation",
      "authors": [
        "Yanchen Guan",
        "Haicheng Liao",
        "Chengyue Wang",
        "Xingcheng Liu",
        "Jiaxun Zhang",
        "Keqiang Li",
        "Zhenning Li"
      ],
      "abstract": "Anticipating traffic accidents is a critical yet unresolved problem for autonomous driving, hindered by the inherent complexity of modeling interactions between road users and the limited availability of diverse, large-scale datasets. To address these issues, we propose a dual-path framework. On the one hand, we employ a video synthesis pipeline that, guided by structured prompts, derives feature distributions from existing corpora and produces high-fidelity synthetic driving scenes consistent with the statistical patterns of real data. On the other hand, we design a graph neural network enriched with semantic cues, enabling dynamic reasoning over both spatial and semantic relations among participants. To validate the effectiveness of our approach, we release a new benchmark dataset containing standardized, finely annotated video sequences that cover a broad spectrum of regions, weather, and traffic conditions. Evaluations across existing datasets and our new benchmark confirm notable gains in both accuracy and anticipation lead time, highlighting the capacity of the proposed framework to mitigate current data bottlenecks and enhance the reliability of autonomous driving systems.",
      "short_abstract": "Anticipating traffic accidents is a critical yet unresolved problem for autonomous driving, hindered by the inherent complexity of modeling interactions between road users and the limited availability of diverse, large-scale datasets. To address these issues,...",
      "published": "2026-04-29",
      "updated": "2026-04-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 112,
      "arxiv_url": "https://arxiv.org/abs/2605.00051",
      "pdf_url": "https://arxiv.org/pdf/2605.00051.pdf"
    },
    {
      "id": "2605.00050",
      "anchor": "paper-2605-00050",
      "title": "Learning physically grounded traffic accident reconstruction from public accident reports",
      "authors": [
        "Yanchen Guan",
        "Haicheng Liao",
        "Chengyue Wang",
        "Zhenning Li"
      ],
      "abstract": "Traffic accidents are routinely documented in textual reports, yet physically grounded accident reconstruction remains difficult because detailed scene measurements and expert reconstructions are scarce, costly and hard to scale. Here we formulate accident reconstruction from publicly accessible reports and scene measurements as a parameterized multimodal learning problem. We construct CISS-REC, a dataset of 6,217 real-world accident cases curated from the NHTSA Crash Investigation Sampling System, and develop a reconstruction framework that grounds report semantics to road topology and participant attributes, reconstructs lane consistent pre-impact motion, and refines collision relevant interactions through localized geometric reasoning and temporal allocation. Our method outperforms representative baselines on CISS-REC, achieving the strongest overall reconstruction fidelity, including improved accident point accuracy and collision consistency. These results show that public accident reports can serve as scalable computational substrates for quantitatively verifiable accident reconstruction, with potential value for traffic safety analysis, simulation and autonomous driving research.",
      "short_abstract": "Traffic accidents are routinely documented in textual reports, yet physically grounded accident reconstruction remains difficult because detailed scene measurements and expert reconstructions are scarce, costly and hard to scale. Here we formulate accident...",
      "published": "2026-04-29",
      "updated": "2026-04-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 112,
      "arxiv_url": "https://arxiv.org/abs/2605.00050",
      "pdf_url": "https://arxiv.org/pdf/2605.00050.pdf"
    },
    {
      "id": "2604.27366",
      "anchor": "paper-2604-27366",
      "title": "Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving",
      "authors": [
        "Lijin Yang",
        "Jianing Huang",
        "Zhongzhan Huang",
        "Shu Liu",
        "Hao Yang"
      ],
      "abstract": "Recent advances in vision language action (VLA) models have shown remarkable potential for autonomous driving by directly mapping multimodal inputs to control signals. However, previous VLA-based methods have not explicitly exploited the critic capability of VLAs to refine driving decisions, even though such capability has been well demonstrated in other LLM-based domains, thereby limiting their performance in complex closed-loop scenarios. In this work, we present a theoretically inspired two-stage framework, CriticVLA, which extends the role of VLAs from acting to judging. CriticVLA first generates a rough trajectory and then refines it through multimodal evaluation and single-step optimization guided by a VLA-based critic, yielding higher-quality driving behaviors. To support this process, we construct a large-scale synthetic dataset of 12.9 million annotated trajectories covering diverse driving scenarios, which enhances the critic's reasoning and refinement abilities. Extensive closed-loop experiments on the Bench2Drive benchmark show that CriticVLA significantly surpasses state-of-the-art baselines, achieving a 73.33% total success rate and delivering about 30% improvement in challenging scenarios.",
      "short_abstract": "Recent advances in vision language action (VLA) models have shown remarkable potential for autonomous driving by directly mapping multimodal inputs to control signals. However, previous VLA-based methods have not explicitly exploited the critic capability of...",
      "published": "2026-04-29",
      "updated": "2026-04-29",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 20,
        "evidence": [
          "title: vision-language-action",
          "abstract: vision-language-action"
        ]
      },
      "recency": "archive",
      "age_days": 112,
      "arxiv_url": "https://arxiv.org/abs/2604.27366",
      "pdf_url": "https://arxiv.org/pdf/2604.27366.pdf"
    },
    {
      "id": "2604.27266",
      "anchor": "paper-2604-27266",
      "title": "AutoREC: A software platform for developing reinforcement learning agents for equivalent circuit model generation from electrochemical impedance spectroscopy data",
      "authors": [
        "Ali Jaberi",
        "Yonatan Kurniawan",
        "Robert Black",
        "Shayan Mousavi M.",
        "Kabir Verma",
        "Zoya Sadighi",
        "Santiago Miret",
        "Jason Hattrick-Simpers"
      ],
      "abstract": "This paper introduces AutoREC, an open-source Python package for developing reinforcement learning (RL) agents to automatically generate equivalent circuit models (ECMs) from electrochemical impedance spectroscopy (EIS) data. While ECMs are a standard framework for interpreting EIS data, traditional identification is typically based on manual trial-and-error, which requires domain experts and limits scalability, particularly in autonomous experimental pipelines such as self-driving laboratories. AutoREC addresses this challenge by formulating ECM construction as a sequential decision-making problem within a Markov Decision Process framework. It implements a Double Deep Q-Network with prioritized experience replay, along with a dedicated dead-loop mitigation strategy, to efficiently explore a complex action space for circuit generation. To demonstrate the capabilities of the platform, we trained an RL agent using AutoREC and evaluated its strengths and limitations across diverse datasets, while also discussing possible strategies to mitigate these limitations in future agent designs. The trained agent achieved a success rate exceeding $99.6\\%$ on synthetic datasets and demonstrated strong generalization to unseen experimental EIS data from batteries, corrosion, oxygen evolution reaction, and CO$_2$ reduction systems. These results position AutoREC as a promising platform for adaptive and data-driven ECM generation, with potential for integration into automated electrochemical workflows.",
      "short_abstract": "This paper introduces AutoREC, an open-source Python package for developing reinforcement learning (RL) agents to automatically generate equivalent circuit models (ECMs) from electrochemical impedance spectroscopy (EIS) data. While ECMs are a standard...",
      "published": "2026-04-29",
      "updated": "2026-04-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Planning & Decision-Making: abstract: decision-making",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 112,
      "arxiv_url": "https://arxiv.org/abs/2604.27266",
      "pdf_url": "https://arxiv.org/pdf/2604.27266.pdf"
    },
    {
      "id": "2604.27118",
      "anchor": "paper-2604-27118",
      "title": "PALCAS: A Priority-Aware Intelligent Lane Change Advisory System for Autonomous Vehicles using Federated Reinforcement Learning",
      "authors": [
        "Yassine Ibork",
        "Nhat Ha Nguyen",
        "Myounggyu Won",
        "Lokesh Das"
      ],
      "abstract": "We present a priority-aware intelligent lane change advisory system based on multi-agent federated reinforcement learning, namely PALCAS, for autonomous vehicles (AVs). While existing lane-change approaches typically focus on single-agent systems or centralized multi-agent systems, we introduce a federated reinforcement learning-based multi-agent lane change system prioritizing lane changing based on vehicle destination urgency. PALCAS incorporates a novel priority-aware safe lane-change reward function to enable judicious lane-change decisions in both mandatory and discretionary scenarios. PALCAS leverages the parameterized deep Q-network (PDQN) algorithm to facilitate effective cooperation among agents, enabling both lateral and longitudinal motion controls of AVs. Extensive simulations conducted using the SUMO traffic simulator and Mosaic V2X communication framework demonstrate that PALCAS significantly improves traffic efficiency, driving safety, comfort, destination arrival rates, and merging success rates compared to baseline methods.",
      "short_abstract": "We present a priority-aware intelligent lane change advisory system based on multi-agent federated reinforcement learning, namely PALCAS, for autonomous vehicles (AVs). While existing lane-change approaches typically focus on single-agent systems or...",
      "published": "2026-04-29",
      "updated": "2026-04-29",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "abstract: V2X",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 112,
      "arxiv_url": "https://arxiv.org/abs/2604.27118",
      "pdf_url": "https://arxiv.org/pdf/2604.27118.pdf"
    },
    {
      "id": "2604.26479",
      "anchor": "paper-2604-26479",
      "title": "Recipes for Calibration Checks in Safety-Critical Applications",
      "authors": [
        "Romeo Valentin"
      ],
      "abstract": "Safety-critical prediction systems, such as autonomous vehicles, weather forecasters, and medical monitors, commonly rely on probabilistic forecasters. These forecasters make predictions about possible future outcomes, and their quality and robustness needs to be validated and certified. Often, only accuracy -- the mean of the predictions -- is evaluated against true outcomes. However, for safety-critical scenarios and decision making under uncertainty, the full distributional properties of the forecasts should be checked: do the observed prediction errors actually follow the forecasted probability distributions? To this end, we introduce a framework for calibration checks: statistical tests that validate distributional properties of forecasts when measured over many samples. In order to support ease-of-use in real-world operations, these checks produce a single accept/reject decision for data collected from a forecaster. This contrasts typical calibration calculations which produce one or multiple continuous calibration scores and require expertise to implement in a validation workflow. We further support operationalization by introducing modifications to calibration testing that (a) reject only overconfident predictions, allowing for pessimistic or cautious predictions in safety-critical settings, and (b) tolerate small, operationally acceptable deviations even for large numbers of validation samples. We organize the calibration checking process into a modular pipeline comprising four steps: (i) the data model, (ii) the chosen metric, (iii) the hypothesis formulation, and (iv) the testing procedure. Each step consists of independently swappable components, thereby supporting a large variety of possible use-cases and trade-offs. We demonstrate the applicability of the framework on two complementary example problems, weather forecasting and robot pose estimation.",
      "short_abstract": "Safety-critical prediction systems, such as autonomous vehicles, weather forecasters, and medical monitors, commonly rely on probabilistic forecasters. These forecasters make predictions about possible future outcomes, and their quality and robustness needs...",
      "published": "2026-04-29",
      "updated": "2026-04-29",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 18,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 112,
      "arxiv_url": "https://arxiv.org/abs/2604.26479",
      "pdf_url": "https://arxiv.org/pdf/2604.26479.pdf"
    },
    {
      "id": "2604.26181",
      "anchor": "paper-2604-26181",
      "title": "SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations",
      "authors": [
        "Jason Wu",
        "Shir-Kang Scott Jin",
        "Yuyang Yuan",
        "Maggie Wigness",
        "Lance M. Kaplan",
        "Hang Qiu",
        "Mani Srivastava"
      ],
      "abstract": "Multimodal deep neural networks deployed in realistic environments must contend with runtime variations: changes in modality quality, overall input complexity, and available platform resources. Current networks struggle with such fluctuations -- adaptive networks cannot adhere to a strict compute budget, controller-based networks neglect to consider input complexity, and statically provisioned networks fail at all the above. Consequently, they do not extract maximum utility from the expended computational resources. We present SWAN (Sample and World-Aware Multimodal Network), the first adaptive multimodal network that accomplishes all three goals. SWAN employs a quality-aware controller to assign resources among modalities according to a variable user-specified maximum budget. Within this budget, an adaptive gating module further optimizes efficiency by scaling layer utilization according to sample complexity. For further gains, SWAN also employs a token dropping module that masks semantically irrelevant multimodal features before performing detections. We evaluate SWAN in the domain of autonomous driving with complex multi-object 3D detection, reducing FLOPs by up to 49% with minimal degradation.",
      "short_abstract": "Multimodal deep neural networks deployed in realistic environments must contend with runtime variations: changes in modality quality, overall input complexity, and available platform resources. Current networks struggle with such fluctuations -- adaptive...",
      "published": "2026-04-28",
      "updated": "2026-04-28",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 5,
        "evidence": [
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 113,
      "arxiv_url": "https://arxiv.org/abs/2604.26181",
      "pdf_url": "https://arxiv.org/pdf/2604.26181.pdf"
    },
    {
      "id": "2604.25574",
      "anchor": "paper-2604-25574",
      "title": "Control Your Queries: Heterogeneous Query Interaction for Camera-Radar Fusion",
      "authors": [
        "Jialong Wu",
        "Yihan Wang",
        "Matthias Rottmann"
      ],
      "abstract": "In autonomous driving, camera-radar fusion offers complementary sensing and low deployment cost. Existing methods perform fusion through input mixing, feature map mixing, or query-based feature sampling. We propose a new fusion paradigm, termed heterogeneous query interaction, and present ConFusion, a camera-radar 3D object detector. ConFusion combines image queries, radar queries, and learnable world queries distributed in 3D space to improve query initialization and object coverage. To encourage cross-type interaction among heterogeneous queries, we introduce heterogeneous query mixing (QMix), which performs dedicated cross-type attention after feature sampling to consolidate complementary object evidence. We further propose interactive query swap sampling (QSwap), which improves feature sampling by allowing related queries to exchange informative feature tokens under attention and geometric constraints. Experiments on the nuScenes dataset show that ConFusion achieves state-of-the-art performance, reaching 59.1 mAP and 65.6 NDS on the validation set, and 61.6 mAP and 67.9 NDS on the test set.",
      "short_abstract": "In autonomous driving, camera-radar fusion offers complementary sensing and low deployment cost. Existing methods perform fusion through input mixing, feature map mixing, or query-based feature sampling. We propose a new fusion paradigm, termed heterogeneous...",
      "published": "2026-04-28",
      "updated": "2026-04-28",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 14,
        "evidence": [
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 113,
      "arxiv_url": "https://arxiv.org/abs/2604.25574",
      "pdf_url": "https://arxiv.org/pdf/2604.25574.pdf"
    },
    {
      "id": "2604.25405",
      "anchor": "paper-2604-25405",
      "title": "Leveraging Previous-Traversal Point Cloud Map Priors for Camera-Based 3D Object Detection and Tracking",
      "authors": [
        "Markus Käppeler",
        "Özgün Çiçek",
        "Yakov Miron",
        "Abhinav Valada"
      ],
      "abstract": "Camera-based 3D object detection and tracking are central to autonomous driving, yet precise 3D object localization remains fundamentally constrained by depth ambiguity when no expensive, depth-rich online LiDAR is available at inference. In many deployments, however, vehicles repeatedly traverse the same environments, making static point cloud maps from prior traversals a practical source of geometric priors. We propose DualViewMapDet, a camera-only inference framework that retrieves such map priors online and leverages them to mitigate the absence of a LiDAR sensor during deployment. The key idea is a dual-space camera-map fusion strategy that avoids one-sided view conversion. Specifically, we (i) project the map into perspective view (PV) and encode multi-channel geometric cues to enrich image features and support BEV lifting, and (ii) encode the map directly in bird's-eye view (BEV) with a sparse voxel backbone and fuse it with lifted camera features in a shared metric space. Extensive evaluations on nuScenes and Argoverse 2 demonstrate consistent improvements over strong camera-only baselines, with particularly strong gains in object localization. Ablations further validate the contributions of PV/BEV fusion and prior-map coverage. We make the code and pre-trained models available at https://dualviewmapdet.cs.uni-freiburg.de .",
      "short_abstract": "Camera-based 3D object detection and tracking are central to autonomous driving, yet precise 3D object localization remains fundamentally constrained by depth ambiguity when no expensive, depth-rich online LiDAR is available at inference. In many deployments,...",
      "published": "2026-04-28",
      "updated": "2026-04-28",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 36,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "title: point cloud",
          "abstract: point cloud",
          "abstract: LiDAR",
          "title: camera",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 113,
      "arxiv_url": "https://arxiv.org/abs/2604.25405",
      "pdf_url": "https://arxiv.org/pdf/2604.25405.pdf"
    },
    {
      "id": "2604.25329",
      "anchor": "paper-2604-25329",
      "title": "ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution",
      "authors": [
        "Chuyao Fu",
        "Shengzhe Gan",
        "Zhuoli Ouyang",
        "Yuhan Rui",
        "Xiaowei Chi",
        "Sirui Han",
        "Jiankun Wang",
        "Hong Zhang"
      ],
      "abstract": "End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, and such reactive planning cannot anticipate future scene evolution, often leading to myopic decisions and safety-critical failures. We propose ProDrive, a world-model-based proactive planning framework that enables ego-environment co-evolution for autonomous driving. ProDrive jointly trains a query-centric trajectory planner and a bird's-eye-view (BEV) world model end-to-end: the planner generates diverse candidate trajectories and planning-aware ego tokens, while the world model predicts future scene evolution conditioned on them. By injecting planner features into the world model and evaluating all candidates in parallel, ProDrive preserves end-to-end gradient flow and allows future outcome assessment to directly shape planning. This bidirectional coupling enables proactive planning beyond current-observation-driven decision-making. Experiments on NAVSIM v1 show that ProDrive outperforms strong baselines in both safety and planning efficiency, while ablations validate the effectiveness of the proposed ego-environment coupling design.",
      "short_abstract": "End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, and such reactive planning cannot anticipate future scene evolution, often leading to myopic decisions and...",
      "published": "2026-04-28",
      "updated": "2026-04-28",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 7,
        "evidence": [
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 113,
      "arxiv_url": "https://arxiv.org/abs/2604.25329",
      "pdf_url": "https://arxiv.org/pdf/2604.25329.pdf"
    },
    {
      "id": "2604.26050",
      "anchor": "paper-2604-26050",
      "title": "Risk Assessments for Evasive Emergency Maneuvers in Autonomous Vehicles",
      "authors": [
        "Aliasghar Arab",
        "Milad Khaleghi",
        "Koorosh Aslansefat"
      ],
      "abstract": "This paper presents a systematic verification and validation (V\\&V) framework for the Evasive Minimum Risk Maneuver (EMRM) feature in autonomous vehicles, addressing a critical gap in existing safety assessment methods. We introduce the first formally integrated pipeline that unifies Hazard Analysis and Risk Assessment (HARA), System-Theoretic Process Analysis (STPA), and Finite State Machine (FSM) modeling into a single traceable workflow specifically designed for EMRM V\\&V. HARA and STPA are combined through a structured hazard-loss mapping to identify hazards and unsafe control actions; an FSM layer captures hazard-to-loss state transitions that neither method models individually; and the unified framework drives automated scenario generation with measurable parameter-space coverage. Applied to a T-junction EMRM case study, the framework guides 1{,}880 RRT-based simulations spanning ego speed, time-to-collision (TTC), and road friction, uncovering a key physical result: the T-junction geometry gives nearly equal difficulty to stopping and to navigating, so the intermediate mitigation mode occupies only 1.9\\% of the feasible parameter space. EMRM steering strategies achieve 81\\% collision-avoidance rate and reduce mean residual impact speed from 18.9~km/h to 9.0~km/h compared with emergency braking alone, while the framework attains 100\\% hazard, UCA, and parameter-space coverage versus $\\leq$1\\% for traditional methods. These results demonstrate that the integrated HARA-STPA-FSM framework enables high-resolution, traceable EMRM V\\&V that is not achievable with any single method in isolation.",
      "short_abstract": "This paper presents a systematic verification and validation (V\\&V) framework for the Evasive Minimum Risk Maneuver (EMRM) feature in autonomous vehicles, addressing a critical gap in existing safety assessment methods. We introduce the first formally...",
      "published": "2026-04-28",
      "updated": "2026-04-28",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "abstract: verification",
          "abstract: validation or testing",
          "abstract: risk assessment"
        ]
      },
      "recency": "archive",
      "age_days": 113,
      "arxiv_url": "https://arxiv.org/abs/2604.26050",
      "pdf_url": "https://arxiv.org/pdf/2604.26050.pdf"
    },
    {
      "id": "2604.24934",
      "anchor": "paper-2604-24934",
      "title": "TEACar: An Open-Source Autonomous Driving Platform",
      "authors": [
        "Zhongzheng Zhang",
        "Maxwell Ruyle",
        "Andrew Kappes",
        "Tyler Ruble",
        "William Shaoul",
        "Dana Moreno",
        "Jack Penn",
        "Ivan Ruchkin"
      ],
      "abstract": "Intelligent Transportation Systems (ITS) increasingly rely on vision-based perception and learning-based control, necessitating experimental platforms that support realistic hardware-in-the-loop validation. Small-scale platforms for autonomous racing offer a practical path to hardware validation, but often suffer from limited modularity, high integration complexity, or restricted extensibility. This paper presents TEACAR, a 1/14- to 1/16-scale autonomous driving platform designed with modular mechanical architecture, hardware abstraction, and ROS 2-based software. The system adopts a four-layer deck structure that physically decouples sensing, computation, actuation, and power subsystems, improving structural rigidity while simplifying reconfiguration. We constructed and comprehensively evaluated the prototype of TEACAR. Its mechanical stability, structural characteristics, and software performance were quantified based on three CNN-based steering controllers. Inference latency, power consumption, and system operating time were measured to evaluate computational capability and robustness. Our experiments demonstrated that TEACAR offers a scalable, modular, and cost-effective testbed for ITS research, education, and development. Our project repository is available on GitHub.",
      "short_abstract": "Intelligent Transportation Systems (ITS) increasingly rely on vision-based perception and learning-based control, necessitating experimental platforms that support realistic hardware-in-the-loop validation. Small-scale platforms for autonomous racing offer a...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 2,
        "evidence": [
          "abstract: controller",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24934",
      "pdf_url": "https://arxiv.org/pdf/2604.24934.pdf"
    },
    {
      "id": "2604.24562",
      "anchor": "paper-2604-24562",
      "title": "Towards Lawful Autonomous Driving: Deriving Scenario-Aware Driving Requirements from Traffic Laws and Regulations",
      "authors": [
        "Bowen Jian",
        "Rongjie Yu",
        "Hong Wang",
        "Liqiang Wang",
        "Zihang Zou"
      ],
      "abstract": "Driving in compliance with traffic laws and regulations is a basic requirement for human drivers, yet autonomous vehicles (AVs) can violate these requirements in diverse real-world scenarios. To encode law compliance into AV systems, conventional approaches use formal logic languages to explicitly specify behavioral constraints, but this process is labor-intensive, hard to scale, and costly to maintain. With recent advances in artificial intelligence, it is promising to leverage large language models (LLMs) to derive legal requirements from traffic laws and regulations. However, without explicitly grounding and reasoning in structured traffic scenarios, LLMs often retrieve irrelevant provisions or miss applicable ones, yielding imprecise requirements. To address this, we propose a novel pipeline that grounds LLM reasoning in a traffic scenario taxonomy through node-wise anchors that encode hierarchical semantics. On Chinese traffic laws and OnSite dataset (5,897 scenarios), our method improves law-scenario matching by 29.1\\% and increases the accuracy of derived mandatory and prohibitive requirements by 36.9\\% and 38.2\\%, respectively. We further demonstrate real-world applicability by constructing a law-compliance layer for AV navigation and developing an onboard, real-time compliance monitor for in-field testing, providing a solid foundation for future AV development, deployment, and regulatory oversight.",
      "short_abstract": "Driving in compliance with traffic laws and regulations is a basic requirement for human drivers, yet autonomous vehicles (AVs) can violate these requirements in diverse real-world scenarios. To encode law compliance into AV systems, conventional approaches...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 16,
        "evidence": [
          "title: law or policy",
          "abstract: law or policy",
          "abstract: human driver"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24562",
      "pdf_url": "https://arxiv.org/pdf/2604.24562.pdf"
    },
    {
      "id": "2604.24353",
      "anchor": "paper-2604-24353",
      "title": "ARETE: Attention-based Rasterized Encoding for Topology Estimation using HSV-transformed Crowdsourced Vehicle Fleet Data",
      "authors": [
        "Daniel Fritz",
        "Dimitrios Lagamtzis",
        "Michael Mink",
        "Markus Enzweiler",
        "Steffen Schober"
      ],
      "abstract": "The continuous advancement of autonomous driving (AD) introduces challenges across multiple disciplines to ensure safe and efficient driving. One such challenge is the generation of High-Definition (HD) maps, which must remain up to date and highly accurate for downstream automotive tasks. One promising approach is the use of crowdsourced data from a vehicle fleet, representing road topology and lane-level features. This work focuses on the generation of centerlines and lane dividers from crowdsourced vehicle trajectories. We adopt a Detection Transformer (DETR)-based approach, where a rasterized representation of vehicle trajectories is used as input to predict vectorized lane representations. Each lane consists of a centerline with an associated direction and corresponding lane dividers that are geometrically constrained by the centerline. Our method includes the extraction of local tiles, from which crowdsourced vehicle trajectories are aggregated. Each tile undergoes a transformation into a rasterized representation encoding both the presence and direction of each trajectory, enabling the prediction of vectorized directed lanes. Experiments are conducted on an internal dataset as well as on the public datasets nuScenes and nuPlan.",
      "short_abstract": "The continuous advancement of autonomous driving (AD) introduces challenges across multiple disciplines to ensure safe and efficient driving. One such challenge is the generation of High-Definition (HD) maps, which must remain up to date and highly accurate...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24353",
      "pdf_url": "https://arxiv.org/pdf/2604.24353.pdf"
    },
    {
      "id": "2604.24295",
      "anchor": "paper-2604-24295",
      "title": "Projected Attainable Speed Space: A Driving Efficiency Metric Connecting Instantaneous Evaluation to Travel Time",
      "authors": [
        "Xiaohua Zhao",
        "Zhaowei Huang",
        "Chen Chen",
        "Haiyi Yang"
      ],
      "abstract": "Inefficient driving behaviors, such as overly conservative yielding, remain a key obstacle to deployment of autonomous vehicles (AVs). Instantaneous driving efficiency metrics are crucial for self-driving decision-making because they affect real-time performance evaluation and control optimization. However, commonly used indicators, including speed, relative speed, and inter-vehicle distance, are limited in capturing traffic context and in ensuring consistency between instantaneous outputs and travel-level outcomes. This study proposes the Projected Attainable Speed Space (PASS) model, a unified framework for driving efficiency assessment across instantaneous and travel-level analyses by integrating kinematic and spatial traffic information. PASS characterizes instantaneous driving efficiency with two coupled elements: potential for speed improvement (available acceleration space) and response to that potential (utilization of available acceleration space). Available acceleration space is referenced to projected attainable speed, derived from an idealized catch-up maneuver using relative speed and spacing to the leading vehicle; utilization is represented by the temporal change in available acceleration space. To ensure cross-scale consistency, time-aggregated PASS is defined as a travel-level efficiency metric. Trajectory data from a driving simulation experiment are used for parameter calibration to maximize agreement between time-aggregated PASS and observed travel times. Across 10 lane-change events, results show strong consistency, with an average coefficient of determination of 0.913, validating PASS for consistent efficiency evaluation across instantaneous and travel-level temporal scales. This study provides a unified, physically grounded framework that supports real-time decision-making and long-term performance analysis in autonomous driving.",
      "short_abstract": "Inefficient driving behaviors, such as overly conservative yielding, remain a key obstacle to deployment of autonomous vehicles (AVs). Instantaneous driving efficiency metrics are crucial for self-driving decision-making because they affect real-time...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control",
          "Control & Vehicle Dynamics: abstract: steering or braking",
          "Planning & Decision-Making: abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24295",
      "pdf_url": "https://arxiv.org/pdf/2604.24295.pdf"
    },
    {
      "id": "2604.24119",
      "anchor": "paper-2604-24119",
      "title": "TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relations",
      "authors": [
        "Yifeng Bai",
        "Zhirong Chen",
        "Bo Song",
        "Erkang Cheng",
        "Haibin Ling"
      ],
      "abstract": "Topology reasoning is crucial for autonomous driving. Current methods primarily focus on instance-level learning for centerline detection, followed by a sequential module for topology reasoning that relies on simplified MLP layers. Moreover, they often neglect the importance of \\textit{point-to-instance} (P2I) relationships in topology reasoning. To address these limitations, we present TopoHR (Topological Hierarchical Representation), a novel end-to-end framework that establishes cyclic interaction between centerline detection and topology reasoning, allowing them to iteratively enhance each other. Specifically, we introduce a hierarchical centerline representation including point queries, instance queries, and semantic representations. These multi-level features are seamlessly integrated and fused within a hierarchical centerline decoder. Furthermore, we design a hierarchical topology reasoning module that captures both fine-grained P2I relationships and global instance-to-instance (I2I) connections within a unified architecture. With these novel components, TopoHR ensures accurate and robust topology reasoning. On the OpenLane-V2 benchmark, TopoHR refreshes state-of-the-art performance with significant improvements. Notably, compared with previous best results, TopoHR achieves +3.8 in $\\mathrm{DET}_{\\text{l}}$, +5.4 in $\\mathrm{TOP}_{\\text{ll}}$ on $\\text{subset_A}$ and +11.0 in $\\mathrm{DET}_{\\text{l}}$, +7.9 in $\\mathrm{TOP}_{\\text{ll}}$ on $\\text{subset_B}$, validating the effectiveness of the proposed components. The code will be shared publicly at https://github.com/Yifeng-Bai/TopoHR.git.",
      "short_abstract": "Topology reasoning is crucial for autonomous driving. Current methods primarily focus on instance-level learning for centerline detection, followed by a sequential module for topology reasoning that relies on simplified MLP layers. Moreover, they often...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end",
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24119",
      "pdf_url": "https://arxiv.org/pdf/2604.24119.pdf"
    },
    {
      "id": "2604.24044",
      "anchor": "paper-2604-24044",
      "title": "CLLAP: Contrastive Learning-based LiDAR-Augmented Pretraining for Enhanced Radar-Camera Fusion",
      "authors": [
        "Bingyi Liu",
        "Chuanhui Zhu",
        "Hongfei Xue",
        "Jian Teng",
        "Jipeng Liu",
        "Enshu Wang",
        "Penglin Dai",
        "Pu Wang"
      ],
      "abstract": "Accurate 3D object detection is critical for autonomous driving, necessitating reliable, cost-effective sensors capable of operating in adverse weather conditions. Camera and millimeter-wave radar fusion has emerged as a promising solution; however, these methods often rely on finely annotated radar data, which is scarce and labor-intensive to produce. To address this challenge, we present CLLAP, a Contrastive Learning-based LiDAR-Augmented Pretraining framework that enhances the performance of existing radar-camera fusion methods for 3D object detection. CLLAP leverages abundant LiDAR data to generate pseudo-radar data using the proposed L2R (LiDAR-to-Radar) Sampling method. Then, it incorporates this data into a novel dual-stage, dual-modality contrastive learning strategy, enabling effective self-supervised learning from paired pseudo-radar and image data. This approach facilitates effective pretraining of existing radar-camera fusion models in a plug-and-play manner, enhancing their feature extraction capabilities and improving detection accuracy and robustness. Experimental results using NuScenes and Lyft Level 5 datasets demonstrate significant performance improvements across three baseline models, highlighting CLLAP's effectiveness in advancing radar-camera fusion for autonomous driving applications.",
      "short_abstract": "Accurate 3D object detection is critical for autonomous driving, necessitating reliable, cost-effective sensors capable of operating in adverse weather conditions. Camera and millimeter-wave radar fusion has emerged as a promising solution; however, these...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 34,
        "evidence": [
          "abstract: object detection",
          "title: LiDAR",
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24044",
      "pdf_url": "https://arxiv.org/pdf/2604.24044.pdf"
    },
    {
      "id": "2604.24242",
      "anchor": "paper-2604-24242",
      "title": "OpenPodcar2: a robust, ROS2 vehicle for self-driving research",
      "authors": [
        "Rakshit Soni",
        "Chris Waltham",
        "Md Umar Ibrahim",
        "Mark Crampton",
        "Charles Fox"
      ],
      "abstract": "OpenPodcar2 is a robust, ROS2-interfaced, low-cost, open source hardware and software, autonomous vehicle platform based on an off-the-shelf, hard-canopy, mobility scooter donor vehicle. It is a modification of the previous OpenPodcar design, which extends it with robust electronics and ROS2 interfacing, to enable both research and also potential deployment use cases. The platform consists of (a) hardware components: documented as a bill of materials and build instructions; (b) integration to the general purpose OSH R4 mechatronics board and a Gazebo simulation of the vehicle, both presenting a common ROS2 interface (c) higher-level ROS2 software implementations and configurations of standard robot autonomous planning and control, including the nav2 stack which performs SLAM and enacts commands to drive the vehicle from a current to a desired pose around obstacles. OpenPodcar2 can transport a human passenger or similar load at speeds up to 15km/h, for example for use as a last-mile autonomous taxi service or to transport delivery containers similarly around a city center. It is small and safe enough to be parked in a standard research lab robust enough for some deployment cases. Total build cost was around 7,000USD from new components, or 2,000USD with a used Donor Vehicle. OpenPodcar2 thus provides a research balance between real world utility, safety, cost and robustness.",
      "short_abstract": "OpenPodcar2 is a robust, ROS2-interfaced, low-cost, open source hardware and software, autonomous vehicle platform based on an off-the-shelf, hard-canopy, mobility scooter donor vehicle. It is a modification of the previous OpenPodcar design, which extends it...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 6,
        "evidence": [
          "abstract: safety",
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24242",
      "pdf_url": "https://arxiv.org/pdf/2604.24242.pdf"
    },
    {
      "id": "2604.24384",
      "anchor": "paper-2604-24384",
      "title": "Pedestrians play chicken with an autonomous vehicle",
      "authors": [
        "Rakshit Soni",
        "Charles Fox"
      ],
      "abstract": "Automated vehicles (AVs) are commonly programmed to yield unconditionally to pedestrians in the interest of safety. However, this design choice can give rise to the Freezing Robot Problem in which pedestrians learn to assert priority at every interaction, causing vehicles to stall and make no progress. The game theoretic Sequential Chicken model has shown that, like human drivers, AVs can resolve this problem by trading credible threats of very small risks of collision or larger risks of less severe invasion of personal space against the value of time due to yielding delays. This paper presents the first demonstration and evaluation of this approach using a real AV with human subjects and shows that pedestrian behavior under experimentally constrained safety conditions can be well fitted by Sequential Chicken, with a low time value of collision, suggestive of their planning to avoid proxemic personal space penalties as well as actual collisions.",
      "short_abstract": "Automated vehicles (AVs) are commonly programmed to yield unconditionally to pedestrians in the interest of safety. However, this design choice can give rise to the Freezing Robot Problem in which pedestrians learn to assert priority at every interaction,...",
      "published": "2026-04-27",
      "updated": "2026-04-27",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 114,
      "arxiv_url": "https://arxiv.org/abs/2604.24384",
      "pdf_url": "https://arxiv.org/pdf/2604.24384.pdf"
    },
    {
      "id": "2605.00880",
      "anchor": "paper-2605-00880",
      "title": "Adversarial Flow Matching for Imperceptible Attacks on End-to-End Autonomous Driving",
      "authors": [
        "Xinyu Zeng",
        "Xiangkun He",
        "Lei Tao",
        "Chen Lv",
        "Hong Cheng"
      ],
      "abstract": "Autonomous driving (AD) is evolving towards end-to-end (E2E) frameworks through two primary paradigms: monolithic models exemplified by Vision-Language-Action (VLA), and specialized modular architectures. Despite their divergent designs, both paradigms increasingly rely on Transformer backbones for complex reasoning, potentially causing a shared vulnerability: visually imperceptible perturbations can manipulate E2E AD models into hazardous maneuvers by targeting the Transformer module. Most existing adversarial attack approaches against AD systems operate under white-box or black-box settings; yet, they typically necessitate full model transparency, or suffer from either prohibitive query latency or limited attack transferability. In this paper, we propose Adversarial Flow Matching (AFM), a novel gray-box attack framework that exploits Transformer structural vulnerabilities in E2E AD models. AFM enables efficient one-step generation of adversarial examples via a neural average velocity field. Additionally, the proposed technique yields effective and visually imperceptible attacks by synergistically perturbing the generative latent space and the neural average velocity field. Extensive experiments demonstrate that AFM achieves a superior trade-off between attack effectiveness and imperceptibility: it substantially degrades the performance of both VLA and modular AD agents across various scenarios compared to baselines, while maintaining state-of-the-art visual imperceptibility. Furthermore, adversarial examples generated by AFM exhibit robust cross-model transferability, indicating that AFM closely approximates a black-box attack setting while requiring only the prior knowledge that the target AD model incorporates a Transformer-based module.",
      "short_abstract": "Autonomous driving (AD) is evolving towards end-to-end (E2E) frameworks through two primary paradigms: monolithic models exemplified by Vision-Language-Action (VLA), and specialized modular architectures. Despite their divergent designs, both paradigms...",
      "published": "2026-04-26",
      "updated": "2026-04-26",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 115,
      "arxiv_url": "https://arxiv.org/abs/2605.00880",
      "pdf_url": "https://arxiv.org/pdf/2605.00880.pdf"
    },
    {
      "id": "2604.23934",
      "anchor": "paper-2604-23934",
      "title": "VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions",
      "authors": [
        "Qingwen Pu",
        "Kun Xie",
        "Yuxiang Liu"
      ],
      "abstract": "Autonomous driving systems often infer pedestrian yielding behavior from geometric and kinematic cues alone, limiting their ability to reason about visual scene context and age-dependent behavioral variability. This limitation can produce delayed interventions in safety-critical encounters and unnecessary braking in benign interactions. This work introduces Vision-Language Model-based Vehicle-Pedestrian Interaction (VLM-VPI), a multimodal reasoning framework for pedestrian intent understanding and yielding-aware control in autonomous driving. The system combines three components: a multimodal perception layer that captures visual and kinematic observations, a reasoning layer that uses Qwen3-VL 8B for visual scene understanding and GPT-OSS 20B for few-shot intent reasoning, and a tiered safety controller that applies age-specific braking margins for children, adults, and seniors. In 112 CARLA scenarios, VLM-VPI achieves 92.3% intent classification accuracy, outperforming a rule-based baseline (78.4%), supervised trajectory models (73.5-82.4%), and a zero-shot LLM configuration (88.4%). Validation on 24 real-world PIE scenarios yields 87.5% accuracy, indicating functional sim-to-real transferability. Across 200 simulation cases, VLM-VPI reduces the false-alarm rate from 7.4% to 2.8% and mean intersection traversal time from 13.5 s to 11.8 s. Conflict occurrences decrease from 124 to 33, while mean minimum time-to-collision improves from 1.92 s to 4.47 s. Demographic-adaptive control further reduces conflicts by 60% for children and 54.5% for seniors compared with uniform control. These results show that an explicit vision-language reasoning layer can improve both safety and efficiency by linking pedestrian intent, demographic context, and vehicle control decisions.",
      "short_abstract": "Autonomous driving systems often infer pedestrian yielding behavior from geometric and kinematic cues alone, limiting their ability to reason about visual scene context and age-dependent behavioral variability. This limitation can produce delayed...",
      "published": "2026-04-26",
      "updated": "2026-04-26",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 4,
        "evidence": [
          "abstract: vehicle control",
          "abstract: controller",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 115,
      "arxiv_url": "https://arxiv.org/abs/2604.23934",
      "pdf_url": "https://arxiv.org/pdf/2604.23934.pdf"
    },
    {
      "id": "2604.23728",
      "anchor": "paper-2604-23728",
      "title": "ESIA: An Energy-Based Spatiotemporal Interaction-Aware Framework for Pedestrian Intention Prediction",
      "authors": [
        "Yanping Wu",
        "Meiting Dang",
        "Lin Wu",
        "Edmond S. L. Ho",
        "Zhenghua Chen",
        "Chongfeng Wei"
      ],
      "abstract": "Recent advances in autonomous driving have motivated research on pedestrian intention prediction, which aims to infer future crossing decisions and actions by modeling temporal dynamics, social interactions, and environmental context. However, existing studies remain constrained by oversimplified multi-agent interaction patterns, opaque reasoning logic, and a lack of global consistency in behavioral predictions, which compromise both robustness and interpretability. In this work, we propose ESIA (Energy-based Spatiotemporal Interaction-Aware framework), a novel Conditional Random Field (CRF)-based paradigm. We cast the intention prediction task as a structured prediction problem over a unified graph-based representation, treating pedestrians and the environment as spatiotemporal nodes. To characterize their distinct roles, we assign unary potentials to nodes to capture individual intentions, and pairwise potentials to edges to encode social and environmental interactions. These potentials are integrated into a unified global energy function to ensure scene-level consistency across behavioral predictions. To further constrain inference without ground-truth supervision, we introduce structural consistency terms to penalize logical contradictions. This optimization is efficiently solved via a novel Unary-Seeded Simulated Annealing (U-SSA) algorithm, which leverages high-confidence unary priors to rapidly converge to a high-quality solution. Extensive experiments on standard benchmarks demonstrate that ESIA achieves state-of-the-art performance with improved interpretability over existing methods.",
      "short_abstract": "Recent advances in autonomous driving have motivated research on pedestrian intention prediction, which aims to infer future crossing decisions and actions by modeling temporal dynamics, social interactions, and environmental context. However, existing...",
      "published": "2026-04-26",
      "updated": "2026-04-26",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 22,
        "evidence": [
          "title: behavior prediction",
          "abstract: behavior prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 115,
      "arxiv_url": "https://arxiv.org/abs/2604.23728",
      "pdf_url": "https://arxiv.org/pdf/2604.23728.pdf"
    },
    {
      "id": "2604.23604",
      "anchor": "paper-2604-23604",
      "title": "Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation",
      "authors": [
        "Simone Mosco",
        "Daniel Fusaro",
        "Alberto Pretto"
      ],
      "abstract": "Understanding the surrounding environment is fundamental in autonomous driving and robotic perception. Distinguishing between known classes and previously unseen objects is crucial in real-world environments, as done in Anomaly Segmentation. However, research in the 3D field remains limited, with most existing approaches applying post-processing techniques from 2D vision. To cover this lack, we propose a new efficient approach that directly operates in the feature space, modeling the feature distribution of inlier classes to constrain anomalous samples. Moreover, the only publicly available 3D LiDAR anomaly segmentation dataset contains simple scenarios, with few anomaly instances, and exhibits a severe domain gap due to its sensor resolution. To bridge this gap, we introduce a set of mixed real-synthetic datasets for 3D LiDAR anomaly segmentation, built upon established semantic segmentation benchmarks, with multiple out-of-distribution objects and diverse, complex environments. Extensive experiments demonstrate that our approach achieves state-of-the-art and competitive results on the existing real-world dataset and the newly introduced mixed datasets, respectively, validating the effectiveness of our method and the utility of the proposed datasets. Code and datasets are available at https://simom0.github.io/lido-page/.",
      "short_abstract": "Understanding the surrounding environment is fundamental in autonomous driving and robotic perception. Distinguishing between known classes and previously unseen objects is crucial in real-world environments, as done in Anomaly Segmentation. However, research...",
      "published": "2026-04-26",
      "updated": "2026-04-26",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 19,
        "evidence": [
          "abstract: perception",
          "abstract: scene segmentation",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 115,
      "arxiv_url": "https://arxiv.org/abs/2604.23604",
      "pdf_url": "https://arxiv.org/pdf/2604.23604.pdf"
    },
    {
      "id": "2604.23523",
      "anchor": "paper-2604-23523",
      "title": "Grammar-Constrained Refinement of Safety Operational Rules Using Language in the Loop: What Could Go Wrong",
      "authors": [
        "Khouloud Gaaloul",
        "Zaid Ghazal",
        "Madhu Latha Pulimi",
        "Sam Emmanuel Kathiravan"
      ],
      "abstract": "Safety specifications in cyber-physical systems (CPS) capture the operational conditions the system must satisfy to operate safely within its intended environment. As operating environments evolve, operational rules must be continuously refined to preserve consistency with observed system behavior during simulation-based verification and validation. Revising inconsistent rules is challenging because the changes must remain syntactically correct under a domain-specific grammar. Language-in-the-loop refinement further raises safety concerns beyond syntactic violations, as it can produce semantically unjustified refinements that overfit to the observed outcomes. We introduce a framework that combines counterfactual reasoning with a grammar-constrained refinement loop to refine operational rules, aligning them with the observed system behavior. Applied to an autonomous driving control system, our approach successfully resolved the inconsistencies in an operational rule inferred by a conventional baseline while remaining grammar compliant. An empirical large language model (LLM) study further revealed model-dependent refinement quality and safety lessons, which motivate rigorous grammar enforcement, stronger semantic validation, and broader evaluation in future work.",
      "short_abstract": "Safety specifications in cyber-physical systems (CPS) capture the operational conditions the system must satisfy to operate safely within its intended environment. As operating environments evolve, operational rules must be continuously refined to preserve...",
      "published": "2026-04-26",
      "updated": "2026-04-26",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 22,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: verification",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 115,
      "arxiv_url": "https://arxiv.org/abs/2604.23523",
      "pdf_url": "https://arxiv.org/pdf/2604.23523.pdf"
    },
    {
      "id": "2604.23685",
      "anchor": "paper-2604-23685",
      "title": "Reading in the Dark: Low-light Scene Text Recognition",
      "authors": [
        "Xuanshuo Fu",
        "Lei Kang",
        "Ernest Valveny",
        "Dimosthenis Karatzas",
        "Javier Vazquez-Corral"
      ],
      "abstract": "Accurate text recognition in low-light environments is essential for intelligent systems in applications ranging from autonomous vehicles to smart surveillance. However, challenges such as poor illumination and noise interference remain underexplored. To address this gap, we introduce LSTR, a large-scale Low-light Scene Text Recognition dataset comprising 11,273 low-light images generated from well-lit datasets (ICDAR2015, IIIT5K, and WordArt), along with ESTR, which includes 60 real nighttime street-scene images in English and Spanish for exclusive evaluation. We explore two solution strategies: (1) employing Optical Character Recognition (OCR) models with fine-tuning and LoRA-based fine-tuning and (2) a joint training strategy that integrates a low-light image enhancement (LLIE) module with an OCR model. In particular, we propose a novel re-render LLIE (RLLIE) module, which demonstrates improved performance on real-world data. Through extensive experimentation, we analyze various training strategies and address a key research question: \\emph{How bright is bright enough for effective scene text recognition?} Our results indicate that standalone LLIE or OCR models perform inadequately under low-light conditions, highlighting the advantages of specialized, jointly trained text-centric approaches. Additionally, we provide a comprehensive benchmark to support future research in robust low-light scene text recognition. https://huggingface.co/datasets/lumimusta/Low-light_Scene_Text_Dataset.",
      "short_abstract": "Accurate text recognition in low-light environments is essential for intelligent systems in applications ranging from autonomous vehicles to smart surveillance. However, challenges such as poor illumination and noise interference remain underexplored. To...",
      "published": "2026-04-26",
      "updated": "2026-04-26",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 115,
      "arxiv_url": "https://arxiv.org/abs/2604.23685",
      "pdf_url": "https://arxiv.org/pdf/2604.23685.pdf"
    },
    {
      "id": "2604.23513",
      "anchor": "paper-2604-23513",
      "title": "Large Language Model based Interactive Decision-Making for Autonomous Driving",
      "authors": [
        "Xinwei Dong",
        "Jiyang Li",
        "Jiabin Xie",
        "Yang Yi",
        "Tianshang Jia",
        "Shiyu Fang",
        "Ye Tian",
        "Peng Hang"
      ],
      "abstract": "In high-conflict mixed-traffic scenarios involving human-driven and autonomous vehicles, most existing autonomous driving systems default to overly conservative behaviors, lack proactive interaction, and consequently suffer from limited public acceptance. To mitigate intent misunderstandings and decision failures, we present a Large Language Model based interactive decision-making framework that augments scene understanding and intent-aware interaction to jointly improve safety and efficiency. The approach uses Object-Process Methodology to semantically model complex multi-vehicle scenes, abstracting low-level perceptual data into objects, processes, and relations, thereby streamlining reasoning over latent causal structure. Building on this representation, the Large Language Model parses both explicit and implicit intents of surrounding agents and, under jointly enforced safety and efficiency constraints, selects candidate maneuvers. We further generate perturbed trajectory candidates via Monte Carlo sampling and evaluate them to obtain an optimized executable trajectory. To foster transparency and coordination with nearby road users, the final decision is translated by the Large Language Model into concise natural-language messages and broadcast through an external Human-Machine Interface, completing a closed loop from scene understanding to action to language. Experiments in a cluster driving simulator demonstrate that the proposed method outperforms traditional baselines across safety, comfort, and efficiency metrics, while a Turing-test-style evaluation indicates a high degree of human-likeness in decision making. Besides, these results suggest that coupling semantic scene abstraction with Large Language Model mediated intent reasoning and language-based eHMI communication offers a practical pathway toward interactive, trustworthy autonomous driving in dense mixed traffic.",
      "short_abstract": "In high-conflict mixed-traffic scenarios involving human-driven and autonomous vehicles, most existing autonomous driving systems default to overly conservative behaviors, lack proactive interaction, and consequently suffer from limited public acceptance. To...",
      "published": "2026-04-25",
      "updated": "2026-04-25",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Simulation",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 10,
        "evidence": [
          "title: decision-making",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 116,
      "arxiv_url": "https://arxiv.org/abs/2604.23513",
      "pdf_url": "https://arxiv.org/pdf/2604.23513.pdf"
    },
    {
      "id": "2604.23362",
      "anchor": "paper-2604-23362",
      "title": "UniAda: Universal Adaptive Multi-objective Adversarial Attack for End-to-End Autonomous Driving Systems",
      "authors": [
        "Jingyu Zhang",
        "Jacky Wai Keung",
        "Yan Xiao",
        "Yihan Liao",
        "Yishu Li",
        "Xiaoxue Ma"
      ],
      "abstract": "Adversarial attacks play a pivotal role in testing and improving the reliability of deep learning (DL) systems. Existing literature has demonstrated that subtle perturbations to the input can elicit erroneous outcomes, thereby substantially compromising the security of DL systems. This has emerged as a critical concern in the development of DL-based safety-critical systems like Autonomous Driving Systems (ADSs). The focus of existing adversarial attack methods on End-to-End (E2E) ADSs has predominantly centered on misbehaviors of steering angle, which overlooks speed-related controls or imperceptible perturbations. To address these challenges, we introduce UniAda, a multi-objective white-box attack technique with a core function that revolves around crafting an image-agnostic adversarial perturbation capable of simultaneously influencing both steering and speed controls. UniAda capitalizes on an intricately designed multi-objective optimization function with the Adaptive Weighting Scheme (AWS), enabling the concurrent optimization of diverse objectives. Validated with both simulated and real-world driving data, UniAda outperforms five benchmarks across two metrics, inducing steering and speed deviations from 3.54 degrees to 29 degrees and 11 km per hour to 22 km per hour on average. This systematic approach establishes UniAda as a proven technique for adversarial attacks on modern DL-based E2E ADSs.",
      "short_abstract": "Adversarial attacks play a pivotal role in testing and improving the reliability of deep learning (DL) systems. Existing literature has demonstrated that subtle perturbations to the input can elicit erroneous outcomes, thereby substantially compromising the...",
      "published": "2026-04-25",
      "updated": "2026-04-25",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 30,
        "margin": 2,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 116,
      "arxiv_url": "https://arxiv.org/abs/2604.23362",
      "pdf_url": "https://arxiv.org/pdf/2604.23362.pdf"
    },
    {
      "id": "2604.23342",
      "anchor": "paper-2604-23342",
      "title": "Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts",
      "authors": [
        "Jingyu Zhang",
        "Fan Wang",
        "Jacky Keung",
        "Yihan Liao",
        "Yan Xiao",
        "Lei Ma"
      ],
      "abstract": "Deep learning (DL)-based systems can exhibit unexpected behavior when exposed to out-of-distribution (OOD) scenarios, posing serious risks in safety-critical domains such as malware detection and autonomous driving. This underscores the importance of thoroughly testing such systems before deployment. To this end, researchers have proposed a wide range of test selection metrics designed to effectively select inputs. However, prior evaluations of metrics reveal three key limitations: (1) narrow testing objectives, for example, many studies assess metrics only for fault detection, leaving their effectiveness for performance estimation unclear; (2) limited coverage of OOD scenarios, with natural and label shifts are rarely considered; (3) Biased dataset selection, where most work focuses on image data while other modalities remain underexplored. Consequently, a unified benchmark that examines how these metrics perform under multiple testing objectives, diverse OOD scenarios, and different data modalities is still lacking. This leaves practitioners uncertain about which test selection metrics are most suitable for their specific objectives and contexts. To address this gap, we conduct an extensive empirical study of 15 existing metrics, evaluating them under three testing objectives (fault detection, performance estimation, and retraining guidance), five types of OOD scenarios (corrupted, adversarial, temporal, natural, and label shifts), three data modalities (image, text, and Android packages), and 13 DL models. In total, our study encompasses 1,640 experimental scenarios, offering a comprehensive evaluation and statistical analysis.",
      "short_abstract": "Deep learning (DL)-based systems can exhibit unexpected behavior when exposed to out-of-distribution (OOD) scenarios, posing serious risks in safety-critical domains such as malware detection and autonomous driving. This underscores the importance of...",
      "published": "2026-04-25",
      "updated": "2026-04-25",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 17,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 116,
      "arxiv_url": "https://arxiv.org/abs/2604.23342",
      "pdf_url": "https://arxiv.org/pdf/2604.23342.pdf"
    },
    {
      "id": "2604.23262",
      "anchor": "paper-2604-23262",
      "title": "An Interactive Graphical Tool to Check the Coarray Continuity of Two-Fold Redundant Sparse Arrays (TFRSAs) Under Single Sensor Failures",
      "authors": [
        "Namya Malik",
        "Ashish Patwari",
        "Sangeetha N"
      ],
      "abstract": "Two-fold redundant sparse arrays possess inbuilt redundancy to tackle single-element failures. This property enables them to perform accurate direction of arrival (DOA) estimation even during single sensor faults. However, recent literature suggests that some TFRSAs suffer from hidden dependencies whereby a single sensor fault at peculiar positions within the array cause discontinuities (holes) in the difference coarray (DCA). This violates the very idea of providing two-fold redundancy. Such hidden dependencies could prove catastrophic in many critical applications such as defense, autonomous driving, and biomedical imaging. Despite this issue, no formal tools or techniques exist to ascertain whether a given array configuration is truly twofold redundant or not. To address this gap, we provide a comprehensive framework and a first-ever graphical user interface (GUI). The GUI has been built using the features in MATLAB app designer and tested with known examples available in sparse array literature. Several numerical examples have been discussed to check the tool's response in each scenario. We conclude that the GUI is functionally accurate and can be an indispensable tool for sparse array designers in making informed choices about array configurations prior to real deployment.",
      "short_abstract": "Two-fold redundant sparse arrays possess inbuilt redundancy to tackle single-element failures. This property enables them to perform accurate direction of arrival (DOA) estimation even during single sensor faults. However, recent literature suggests that some...",
      "published": "2026-04-25",
      "updated": "2026-04-25",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 5,
        "evidence": [
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 116,
      "arxiv_url": "https://arxiv.org/abs/2604.23262",
      "pdf_url": "https://arxiv.org/pdf/2604.23262.pdf"
    },
    {
      "id": "2604.23105",
      "anchor": "paper-2604-23105",
      "title": "Transferable Physical-World Adversarial Patches Against Object Detection in Autonomous Driving",
      "authors": [
        "Zihui Zhu",
        "Ziqi Zhou",
        "Yichen Wang",
        "Lulu Xue",
        "Minghui Li",
        "Shengshan Hu"
      ],
      "abstract": "Deep learning drives major advances in autonomous driving (AD), where object detectors are central to perception. However, adversarial attacks pose significant threats to the reliability and safety of these systems, with physical adversarial patches representing a particularly potent form of attack. Physical adversarial patch attacks pose severe risks but are usually crafted for a single model, yielding poor transferability to unseen detectors. We propose AdvAD, a transfer-based physical attack against object detection in autonomous driving. Instead of targeting a specific detector, AdvAD optimizes adversarial patches over multiple detection models in a unified framework, encouraging the learned perturbations to capture shared vulnerabilities across architectures. The optimization process adaptively balances model contributions and enforces robustness to physical variations. It further employs data augmentation and geometric transformations to maintain patch effectiveness under diverse physical conditions. Experiments in both digital and real-world settings show that AdvAD consistently outperforms state-of-the-art (SOTA) attacks in performance and transferability.",
      "short_abstract": "Deep learning drives major advances in autonomous driving (AD), where object detectors are central to perception. However, adversarial attacks pose significant threats to the reliability and safety of these systems, with physical adversarial patches...",
      "published": "2026-04-24",
      "updated": "2026-04-24",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 22,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 117,
      "arxiv_url": "https://arxiv.org/abs/2604.23105",
      "pdf_url": "https://arxiv.org/pdf/2604.23105.pdf"
    },
    {
      "id": "2604.22560",
      "anchor": "paper-2604-22560",
      "title": "Cross-Stage Coherence in Hierarchical Driving VQA: Explicit Baselines and Learned Gated Context Projectors",
      "authors": [
        "Gautam Kumar Jain",
        "Carsten Markgraf",
        "Julian Stähler"
      ],
      "abstract": "Graph Visual Question Answering (GVQA) for autonomous driving organizes reasoning into ordered stages, namely Perception, Prediction, and Planning, where planning decisions should remain consistent with the model's own perception. We present a comparative study of cross-stage context passing on DriveLM-nuScenes using two complementary mechanisms. The explicit variant evaluates three prompt-based conditioning strategies on a domain-adapted 4B VLM (Mini-InternVL2-4B-DA-DriveLM) without additional training, reducing NLI contradiction by up to 42.6% and establishing a strong zero-training baseline. The implicit variant introduces gated context projectors, which extract a hidden-state vector from one stage and inject a normalized, gated projection into the next stage's input embeddings. These projectors are jointly trained with stage-specific QLoRA adapters on a general-purpose 8B VLM (InternVL3-8B-Instruct) while updating only approximately 0.5% of parameters. The implicit variant achieves a statistically significant 34% reduction in planning-stage NLI contradiction (bootstrap 95% CIs, p < 0.05) and increases cross-stage entailment by 50%, evaluated with a multilingual NLI classifier to account for mixed-language outputs. Planning language quality also improves (CIDEr +30.3%), but lexical overlap and structural consistency degrade due to the absence of driving-domain pretraining. Since the two variants use different base models, we present them as complementary case studies: explicit context passing provides a strong training-free baseline for surface consistency, while implicit gated projection delivers significant planning-stage semantic gains, suggesting domain adaptation as a plausible next ingredient for full-spectrum improvement.",
      "short_abstract": "Graph Visual Question Answering (GVQA) for autonomous driving organizes reasoning into ordered stages, namely Perception, Prediction, and Planning, where planning decisions should remain consistent with the model's own perception. We present a comparative...",
      "published": "2026-04-24",
      "updated": "2026-04-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 117,
      "arxiv_url": "https://arxiv.org/abs/2604.22560",
      "pdf_url": "https://arxiv.org/pdf/2604.22560.pdf"
    },
    {
      "id": "2604.22552",
      "anchor": "paper-2604-22552",
      "title": "Transferable Physical-World Adversarial Patches Against Pedestrian Detection Models",
      "authors": [
        "Shihui Yan",
        "Ziqi Zhou",
        "Yufei Song",
        "Yifan Hu",
        "Minghui Li",
        "Shengshan Hu"
      ],
      "abstract": "Physical adversarial patch attacks critically threaten pedestrian detection, causing surveillance and autonomous driving systems to miss pedestrians and creating severe safety risks. Despite their effectiveness in controlled settings, existing physical attacks face two major limitations in practice: they lack systematic disruption of the multi-stage decision pipeline, enabling residual modules to offset perturbations, and they fail to model complex physical variations, leading to poor robustness. To overcome these limitations, we propose a novel pedestrian adversarial patch generation method that combines multi-stage collaborative attacks with robustness enhancement under physical diversity, called TriPatch. Specifically, we design a triplet loss consisting of detection confidence suppression, bounding-box offset amplification, and non-maximum suppression (NMS) disruption, which jointly act across different stages of the detection pipeline. In addition, we introduce an appearance consistency loss to constrain the color distribution of the patch, thereby improving its adaptability under diverse imaging conditions, and incorporate data augmentation to further enhance robustness against complex physical perturbations. Extensive experiments demonstrate that TriPatch achieves a higher attack success rate across multiple detector models compared to existing approaches.",
      "short_abstract": "Physical adversarial patch attacks critically threaten pedestrian detection, causing surveillance and autonomous driving systems to miss pedestrians and creating severe safety risks. Despite their effectiveness in controlled settings, existing physical...",
      "published": "2026-04-24",
      "updated": "2026-04-24",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 19,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 117,
      "arxiv_url": "https://arxiv.org/abs/2604.22552",
      "pdf_url": "https://arxiv.org/pdf/2604.22552.pdf"
    },
    {
      "id": "2604.22260",
      "anchor": "paper-2604-22260",
      "title": "Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset",
      "authors": [
        "Wenhui Huang",
        "Songyan Zhang",
        "Collister Chua",
        "Yang Liang",
        "Zhiqi Mao",
        "Heng Yang",
        "Chen Lv"
      ],
      "abstract": "Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While recent advances in foundation models and large-scale multimodal datasets have strengthened perception and reasoning in intelligent transportation systems (ITS), existing research remains largely centered on microscopic autonomous driving (AD), with limited attention to city-scale traffic analysis. In particular, open-ended safety-oriented visual question answering (VQA) and corresponding foundation models for reasoning over heterogeneous roadside camera observations remain underexplored. To address this gap, we introduce the Land Transportation Dataset (LTD), a large-scale open-source vision-language dataset for open-ended reasoning in urban traffic environments. LTD contains 11.6K high-quality VQA pairs collected from heterogeneous roadside cameras, spanning diverse road geometries, traffic participants, illumination conditions, and adverse weather. The dataset integrates three complementary tasks: fine-grained multi-object grounding, multi-image camera selection, and multi-image risk analysis, requiring joint reasoning over minimally correlated views to infer hazardous objects, contributing factors, and risky road directions. To ensure annotation fidelity, we combine multi-model vision-language generation with cross-validation and human-in-the-loop refinement. Building upon LTD, we further propose UniVLT, a transportation foundation model trained via curriculum-based knowledge transfer to unify microscopic AD reasoning and macroscopic traffic analysis within a single architecture. Extensive experiments on LTD and multiple AD benchmarks demonstrate that UniVLT achieves SOTA performance on open-ended reasoning tasks across diverse domains, while exposing limitations of existing foundation models in complex multi-view traffic scenarios.",
      "short_abstract": "Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While recent advances in foundation models and large-scale multimodal datasets have strengthened perception and...",
      "published": "2026-04-24",
      "updated": "2026-04-24",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 2,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 117,
      "arxiv_url": "https://arxiv.org/abs/2604.22260",
      "pdf_url": "https://arxiv.org/pdf/2604.22260.pdf"
    },
    {
      "id": "2604.22240",
      "anchor": "paper-2604-22240",
      "title": "OccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy Space",
      "authors": [
        "Zhuding Liang",
        "Tianyi Yan",
        "Dubing Chen",
        "Jiasen Zheng",
        "Huan Zheng",
        "Cheng-zhong Xu",
        "Yida Wang",
        "Kun Zhan",
        "Jianbing Shen"
      ],
      "abstract": "Generative world models increasingly rely on 4D occupancy for realistic autonomous driving simulation. However, existing generation frameworks depend on rigid geometric conditions (e.g., explicit trajectories) or simplistic attribute-level text, failing to orchestrate complex, sequential multi-agent interactions. To address this semantic-spatiotemporal gap, we propose OccDirector, a pioneering framework that generates 4D occupancy dynamics conditioned solely on natural language. Operating as a ``scenario director'', OccDirector maps natural language scripts into physically plausible voxel dynamics without requiring geometric priors. Technically, it employs a VLM-driven Spatio-Temporal MMDiT equipped with a history-prefix anchoring strategy to ensure long-horizon interaction consistency. Furthermore, we introduce OccInteract-85k, a novel dataset uniquely annotated with multi-level language instructions: ranging from static layouts to intricate multi-agent behaviors, alongside a novel VLM-based evaluation benchmark. Extensive experiments demonstrate that OccDirector achieves state-of-the-art generation quality and unprecedented instruction-following capabilities, successfully shifting the paradigm from appearance synthesis to language-driven behavior orchestration.",
      "short_abstract": "Generative world models increasingly rely on 4D occupancy for realistic autonomous driving simulation. However, existing generation frameworks depend on rigid geometric conditions (e.g., explicit trajectories) or simplistic attribute-level text, failing to...",
      "published": "2026-04-24",
      "updated": "2026-04-24",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: occupancy",
          "abstract: occupancy"
        ]
      },
      "recency": "archive",
      "age_days": 117,
      "arxiv_url": "https://arxiv.org/abs/2604.22240",
      "pdf_url": "https://arxiv.org/pdf/2604.22240.pdf"
    },
    {
      "id": "2604.22564",
      "anchor": "paper-2604-22564",
      "title": "Relational Archetypes: A Comparative Analysis of AV-Human and Agent-Human Interactions",
      "authors": [
        "Antoni Lorente",
        "Amin Oueslati",
        "Robin Staes-Polet"
      ],
      "abstract": "Over the last couple of years, AI Agents have gained significant traction due to substantial progress in the capabilities of underlying General Purpose AI (GPAI) models, enhanced scaffolding techniques, and the promise to drive societal transformation. Companies, researchers, and policy makers have started to consider the different effects that AI agents may have across different dimensions of our lives. However, the literature exploring the broader effects of human-agent interactions is still underdeveloped. In this paper, we review the problem of traffic modulation by autonomous vehicles (AVs) in mixed traffic flows and extrapolate the learnings to the different modes of interaction between humans and AVs to the pair humans-AI agents. In doing so, we propose a preliminary taxonomy of relational archetypes based on literature on Human-Computer Interaction (HCI) and AV-human interaction and tentatively explore how the resulting framework may lead to new questions regarding human-agent interactions. Our effort is aimed at strengthening existing bridges between these two research communities, which share similar traits: autonomy, fast adoption, high impact, and great potential for economic transformation. Building on previous analogies between AI Agents and AVs (e.g., regarding autonomy levels), we anticipate this paper to spark scholarly debate on the different types of impact that agents may have on our societies, while inviting other researchers to expand the scope of their comparative analysis regarding AI Agents.",
      "short_abstract": "Over the last couple of years, AI Agents have gained significant traction due to substantial progress in the capabilities of underlying General Purpose AI (GPAI) models, enhanced scaffolding techniques, and the promise to drive societal transformation....",
      "published": "2026-04-24",
      "updated": "2026-04-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy"
        ]
      },
      "recency": "archive",
      "age_days": 117,
      "arxiv_url": "https://arxiv.org/abs/2604.22564",
      "pdf_url": "https://arxiv.org/pdf/2604.22564.pdf"
    },
    {
      "id": "2604.22872",
      "anchor": "paper-2604-22872",
      "title": "Vision-Based Lane Following and Traffic Sign Recognition for Resource-Constrained Autonomous Vehicles",
      "authors": [
        "Md Tanjemul Islam",
        "Md Rafiul Kabir"
      ],
      "abstract": "Autonomous vehicles (AVs) rely on real-time perception systems to understand road environments and ensure safe navigation. However, implementing reliable perception algorithms on resource-constrained embedded platforms remains challenging due to limited computational resources. This paper presents a lightweight vision-based framework that integrates lane detection, lane tracking, and traffic sign recognition for embedded autonomous vehicles. A computationally efficient threshold-based lane segmentation method combined with perspective transformation and histogram-based curvature estimation is used for robust lane tracking under varying illumination conditions. A rule-based steering controller generates steering commands to maintain stable vehicle navigation. For traffic sign recognition, two lightweight convolutional neural networks (CNNs), EfficientNet-B0 and MobileNetV2, are evaluated using a custom dataset captured from the vehicle's onboard camera. Experimental results show that the system achieves real-time performance while maintaining accurate lane tracking with only 3.16% maximum offset RMSE. EfficientNet-B0 achieves a high offline classification accuracy of 98.77% on the test dataset, while achieving 90% accuracy during real-time on-device deployment, outperforming MobileNetV2 in both settings. MobileNetV2, however, offers slightly faster inference and lower computational cost. These results highlight the effectiveness of lightweight vision-based perception pipelines for resource-constrained autonomous driving applications.",
      "short_abstract": "Autonomous vehicles (AVs) rely on real-time perception systems to understand road environments and ensure safe navigation. However, implementing reliable perception algorithms on resource-constrained embedded platforms remains challenging due to limited...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 15,
        "evidence": [
          "abstract: perception",
          "abstract: camera",
          "abstract: lane detection",
          "title: traffic sign recognition",
          "abstract: traffic sign recognition"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.22872",
      "pdf_url": "https://arxiv.org/pdf/2604.22872.pdf"
    },
    {
      "id": "2604.22126",
      "anchor": "paper-2604-22126",
      "title": "GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems",
      "authors": [
        "Baodi Shan",
        "Mauricio Araya-Polo",
        "Barbara Chapman"
      ],
      "abstract": "Distributed GPU applications increasingly rely on kernel-level, cross-node coordination to reduce launch overheads and improve compute-communication overlap, but such support is lacking. On OFI-based interconnects such as HPE Slingshot, which powers six of the top ten systems in the November 2025 Top500, including the top three, GPU kernels cannot autonomously drive distributed coordination: existing runtimes rely on host-driven progress and lack a bounded mechanism for recycling pre-staged NIC work across repeated GPU-triggered operations. On InfiniBand, GPU-initiated communication is possible, but current implementations incur unnecessary synchronization and locking overheads. This paper presents GICC, a framework that enables GPU kernels to directly trigger NIC-level operations without host involvement on the fast path. In stencils, GPU threads initiate halo exchanges as soon as boundary regions are computed, enabling fine-grained overlap between interior computation and boundary transfer. GICC decouples coordination semantics from data movement and introduces asynchronous resource reclamation: the NIC signals completion to both GPU and host memory, letting a lightweight host thread recycle NIC resources concurrently with GPU execution without injecting latency into the coordination path. This sustains GPU-driven coordination under finite NIC state, absent from existing OFI-based runtimes. We implement GICC on NVIDIA and AMD GPUs over InfiniBand and Slingshot. On Slingshot, GICC reduces per-coordination latency by up to 229x and improves weak scaling efficiency by up to 25%. On InfiniBand, it achieves up to 1.95x lower put latency than NVSHMEM by eliminating unnecessary locking and synchronization. On an industrial stencil proxy on 64 AMD MI250X GCDs, GPU-aware MPI incurs over 52% higher communication time than GICC, which achieves 42% parallel efficiency versus MPI's 35.4%.",
      "short_abstract": "Distributed GPU applications increasingly rely on kernel-level, cross-node coordination to reduce launch overheads and improve compute-communication overlap, but such support is lacking. On OFI-based interconnects such as HPE Slingshot, which powers six of...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 10,
        "evidence": [
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.22126",
      "pdf_url": "https://arxiv.org/pdf/2604.22126.pdf"
    },
    {
      "id": "2604.21489",
      "anchor": "paper-2604-21489",
      "title": "MISTY: High-Throughput Motion Planning via Mixer-based Single-step Drifting",
      "authors": [
        "Yining Xing",
        "Zehong Ke",
        "Yiqian Tu",
        "Zhiyuan Liu",
        "Wenhao Yu",
        "Jianqiang Wang"
      ],
      "abstract": "Multi-modal trajectory generation is essential for safe autonomous driving, yet existing diffusion-based planners suffer from high inference latency due to iterative neural function evaluations. This paper presents MISTY (Mixer-based Inference for Single-step Trajectory-drifting Yield), a high-throughput generative motion planner that achieves state-of-the-art closed-loop performance with pure single-step inference. MISTY integrates a vectorized Sub-Graph encoder to capture environment context, a Variational Autoencoder to structure expert trajectories into a compact 32-dimensional latent manifold, and an ultra-lightweight MLP-Mixer decoder to eliminate quadratic attention complexity. Importantly, we introduce a latent-space drifting loss that shifts the complex distribution evolution entirely to the training phase. By formulating explicit attractive and repulsive forces, this mechanism empowers the model to synthesize novel, proactive maneuvers, such as active overtaking, that are virtually absent from the raw expert demonstrations. Extensive evaluations on the nuPlan benchmark demonstrate that MISTY achieves state-of-the-art results on the challenging Test14-hard split, with comprehensive scores of 80.32 and 82.21 in non-reactive and reactive settings, respectively. Operating at over 99 FPS with an end-to-end latency of 10.1 ms, MISTY offers an order-of-magnitude speedup over iterative diffusion planners while while achieving significantly robust generation.",
      "short_abstract": "Multi-modal trajectory generation is essential for safe autonomous driving, yet existing diffusion-based planners suffer from high inference latency due to iterative neural function evaluations. This paper presents MISTY (Mixer-based Inference for Single-step...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 24,
        "evidence": [
          "title: motion planning",
          "title: planning",
          "abstract: trajectory generation"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.21489",
      "pdf_url": "https://arxiv.org/pdf/2604.21489.pdf"
    },
    {
      "id": "2604.21479",
      "anchor": "paper-2604-21479",
      "title": "Frozen LLMs as Map-Aware Spatio-Temporal Reasoners for Vehicle Trajectory Prediction",
      "authors": [
        "Yanjiao Liu",
        "Jiawei Liu",
        "Xun Gong",
        "Zifei Nie"
      ],
      "abstract": "Large language models (LLMs) have recently demonstrated strong reasoning capabilities and attracted increasing research attention in the field of autonomous driving (AD). However, safe application of LLMs on AD perception and prediction still requires a thorough understanding of both the dynamic traffic agents and the static road infrastructure. To this end, this study introduces a framework to evaluate the capability of LLMs in understanding the behaviors of dynamic traffic agents and the topology of road networks. The framework leverages frozen LLMs as the reasoning engine, employing a traffic encoder to extract spatial-level scene features from observed trajectories of agents, while a lightweight Convolutional Neural Network (CNN) encodes the local high-definition (HD) maps. To assess the intrinsic reasoning ability of LLMs, the extracted scene features are then transformed into LLM-compatible tokens via a reprogramming adapter. By residing the prediction burden with the LLMs, a simpler linear decoder is applied to output future trajectories. The framework enables a quantitative analysis of the influence of multi-modal information, especially the impact of map semantics on trajectory prediction accuracy, and allows seamless integration of frozen LLMs with minimal adaptation, thereby demonstrating strong generalizability across diverse LLM architectures and providing a unified platform for model evaluation.",
      "short_abstract": "Large language models (LLMs) have recently demonstrated strong reasoning capabilities and attracted increasing research attention in the field of autonomous driving (AD). However, safe application of LLMs on AD perception and prediction still requires a...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 25,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.21479",
      "pdf_url": "https://arxiv.org/pdf/2604.21479.pdf"
    },
    {
      "id": "2604.22196",
      "anchor": "paper-2604-22196",
      "title": "V-STC: A Time-Efficient Multi-Vehicle Coordinated Trajectory Planning Approach",
      "authors": [
        "Pengfei Liu",
        "Jialing Zhou",
        "Yuezu Lv",
        "Guanghui Wen",
        "Tingwen Huang"
      ],
      "abstract": "Coordinating the motions of multiple autonomous vehicles (AVs) requires planning frameworks that ensure safety while making efficient use of space and time. This paper presents a new approach, termed variable-time-step spatio-temporal corridor (V-STC), that enhances the temporal efficiency of multi-vehicle coordination. An optimization model is formulated to construct a V-STC for each AV, in which both the spatial configuration of the corridor cubes and their time durations are treated as decision variables. By allowing the corridor's spatial position and time step to vary, the constructed V-STC reduces the overall temporal occupancy of each AV while maintaining collision-free separation in the spatio-temporal domain. Based on the generated V-STC, a dynamically feasible trajectory is then planned independently for each AV. Simulation studies demonstrate that the proposed method achieves safe multi-vehicle coordination and yields more time-efficient motion compared with existing STC approaches.",
      "short_abstract": "Coordinating the motions of multiple autonomous vehicles (AVs) requires planning frameworks that ensure safety while making efficient use of space and time. This paper presents a new approach, termed variable-time-step spatio-temporal corridor (V-STC), that...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 23,
        "evidence": [
          "title: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.22196",
      "pdf_url": "https://arxiv.org/pdf/2604.22196.pdf"
    },
    {
      "id": "2604.22068",
      "anchor": "paper-2604-22068",
      "title": "TRACE: Topology-aware Reconstruction of Accidents in CARLA for AV Evaluation",
      "authors": [
        "Nahian Salsabil",
        "Sebastian Elbaum"
      ],
      "abstract": "Validating Autonomous Vehicles (AVs) requires exposure to rare, safety-critical scenarios, infrequent in routine driving data. Existing benchmarks address this by generating synthetic conflicts or mapping accident descriptions to abstract road geometries, failing to capture the topological complexity of real-world crashes. We introduce TRACE , a pipeline that automates the reconstruction of NHTSA crash reports into high-fidelity CARLA simulations by (1) retrieving site-specific OpenStreetMap data to preserve exact road topology, (2) leveraging Large Language Models to infer vehicles' initial state from road geometry and pre-crash maneuvers, and (3) generating simulation trajectories from semi-structured report data. Using this pipeline, we curated a benchmark of 52 diverse accident scenarios covering varied collision types, road topologies, and pre-crash maneuvers, providing a challenging open source resource for testing AV systems against real-world failures.",
      "short_abstract": "Validating Autonomous Vehicles (AVs) requires exposure to rare, safety-critical scenarios, infrequent in routine driving data. Existing benchmarks address this by generating synthetic conflicts or mapping accident descriptions to abstract road geometries,...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.22068",
      "pdf_url": "https://arxiv.org/pdf/2604.22068.pdf"
    },
    {
      "id": "2604.22040",
      "anchor": "paper-2604-22040",
      "title": "Robust Localization for Autonomous Vehicles in Highway Scenes",
      "authors": [
        "Daqian Cheng",
        "Xuchu Ding",
        "Yujia Wu",
        "Xiang Zhang",
        "Lei Wang"
      ],
      "abstract": "Localization for autonomous vehicles on highways remains under-explored compared to urban roads, and state-of-the-art methods for urban scenes degrade when directly applied to highways. We identify key challenges including environment changes under information homogeneity, heavy occlusion, degraded GNSS signals, and stringent downstream requirements on accuracy and latency. We propose a robust localization system to address highway challenges, which uses a dual-likelihood LiDAR front end that decouples 3D geometric structures and 2D road-texture cues to handle environment changes; a Control-EKF further leverages steering and acceleration commands to reduce lag and improve closed-loop behavior. An automated offline mapping and ground-truth pipeline keep maps fresh at high cadence for optimal localization performance. To catalyze progress, we release a public dataset covering both urban roads and highways while focusing on representative challenging highway clips, totaling 163 km; benchmarking is standardized using product-oriented accuracy metrics and certified ground truth. Compared to Apollo and Autoware, our system performs similarly on urban roads but shows superior robustness on challenging highway scenarios. The system has been validated by more than one million kilometers of road testing.",
      "short_abstract": "Localization for autonomous vehicles on highways remains under-explored compared to urban roads, and state-of-the-art methods for urban scenes degrade when directly applied to highways. We identify key challenges including environment changes under...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 15,
        "evidence": [
          "title: localization",
          "abstract: localization",
          "abstract: mapping",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.22040",
      "pdf_url": "https://arxiv.org/pdf/2604.22040.pdf"
    },
    {
      "id": "2604.21854",
      "anchor": "paper-2604-21854",
      "title": "Bounding the Black Box: A Statistical Certification Framework for AI Risk Regulation",
      "authors": [
        "Natan Levy",
        "Gadi Perl"
      ],
      "abstract": "Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Governments have responded: the EU AI Act, the NIST Risk Management Framework, and the Council of Europe Convention all demand that high-risk systems demonstrate safety before deployment. Yet beneath this regulatory consensus lies a critical vacuum: none specifies what ``acceptable risk'' means in quantitative terms, and none provides a technical method for verifying that a deployed system actually meets such a threshold. The regulatory architecture is in place; the verification instrument is not. This gap is not theoretical. As the EU AI Act moves into full enforcement, developers face mandatory conformity assessments without established methodologies for producing quantitative safety evidence - and the systems most in need of oversight are opaque statistical inference engines that resist white-box scrutiny. This paper provides the missing instrument. Drawing on the aviation certification paradigm, we propose a two-stage framework that transforms AI risk regulation into engineering practice. In Stage One, a competent authority formally fixes an acceptable failure probability $δ$ and an operational input domain $\\varepsilon$ - a normative act with direct civil liability implications. In Stage Two, the RoMA and gRoMA statistical verification tools compute a definitive, auditable upper bound on the system's true failure rate, requiring no access to model internals and scaling to arbitrary architectures. We demonstrate how this certificate satisfies existing regulatory obligations, shifts accountability upstream to developers, and integrates with the legal frameworks that exist today.",
      "short_abstract": "Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Governments have responded: the EU AI Act, the NIST Risk Management Framework, and the Council of Europe...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 5,
        "evidence": [
          "title: law or policy",
          "abstract: law or policy"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.21854",
      "pdf_url": "https://arxiv.org/pdf/2604.21854.pdf"
    },
    {
      "id": "2604.21841",
      "anchor": "paper-2604-21841",
      "title": "Cross-Modal Phantom: Coordinated Camera-LiDAR Spoofing Against Multi-Sensor Fusion in Autonomous Vehicles",
      "authors": [
        "Shahriar Rahman Khan",
        "Raiful Hasan"
      ],
      "abstract": "Autonomous Vehicles (AVs) increasingly depend on Multi-Sensor Fusion (MSF) to combine complementary modalities such as cameras and LiDAR for robust perception. While this redundancy is intended to safeguard against single-sensor failures, the fusion process itself introduces a subtle and underexplored vulnerability. In this work, we investigate whether an attacker can bypass MSF's redundancy by fabricating cross-sensor consistency, making multiple sensors agree on the same false object. We design a coordinated, data-level (early-fusion) attack that emulates the outcome of two synchronized physical spoofing sources: an infrared (IR) projection that induces a false camera detection and a LiDAR signal injection that produces a matching 3D point cluster. Rather than implementing the physical attack hardware, we simulate its sensor-level outcomes by inserting perspective-aware image patches and synthetic LiDAR point clusters aligned in 3D space. This approach preserves the perceptual effects that real IR and IEMI-based spoofing would create at the sensor output. Using 400 KITTI scenes, our large-scale evaluation shows that the coordinated spoofing deceives a state-of-the-art perception model with an 85.5% successful attack rate. These findings provide the first quantitative evidence that malicious cross-modal consistency can compromise MSF-based perception, revealing a critical vulnerability in the core data-fusion logic of modern autonomous vehicle systems.",
      "short_abstract": "Autonomous Vehicles (AVs) increasingly depend on Multi-Sensor Fusion (MSF) to combine complementary modalities such as cameras and LiDAR for robust perception. While this redundancy is intended to safeguard against single-sensor failures, the fusion process...",
      "published": "2026-04-23",
      "updated": "2026-04-23",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 39,
        "margin": 19,
        "evidence": [
          "abstract: perception",
          "title: sensor fusion",
          "abstract: sensor fusion",
          "title: LiDAR",
          "abstract: LiDAR",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 118,
      "arxiv_url": "https://arxiv.org/abs/2604.21841",
      "pdf_url": "https://arxiv.org/pdf/2604.21841.pdf"
    },
    {
      "id": "2604.22856",
      "anchor": "paper-2604-22856",
      "title": "Attention-Augmented YOLOv8 with Ghost Convolution for Real-Time Vehicle Detection in Intelligent Transportation Systems",
      "authors": [
        "Syed Sajid Ullah",
        "Muhammad Zunair Zamir",
        "Ahsan Ishfaq",
        "Salman Khan"
      ],
      "abstract": "Accurate vehicle detection is a critical component of autonomous driving, traffic surveillance, and intelligent transportation systems. This paper presents an enhanced YOLOv8n-based model that integrates the Ghost Module, Convolutional Block Attention Module (CBAM), and Deformable Convolutional Networks v2 (DCNv2) to improve detection performance. The Ghost Module reduces feature redundancy through efficient feature generation, CBAM refines feature representation via channel and spatial attention, and DCNv2 enhances adaptability to geometric variations in vehicle structures. Evaluated on the KITTI dataset, the proposed model achieves 95.4% mAP@0.5, representing an 8.97% improvement over the baseline YOLOv8n, along with 96.2% precision, 93.7% recall, and a 94.93% F1-score. Comparative analysis against seven state-of-the-art detectors demonstrates consistent superiority across key performance metrics, while ablation studies validate the individual and combined contributions of the integrated modules. By addressing feature redundancy, attention refinement, and spatial adaptability, the proposed approach offers a robust and computationally efficient solution for vehicle detection in diverse and complex traffic environments.",
      "short_abstract": "Accurate vehicle detection is a critical component of autonomous driving, traffic surveillance, and intelligent transportation systems. This paper presents an enhanced YOLOv8n-based model that integrates the Ghost Module, Convolutional Block Attention Module...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "title: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.22856",
      "pdf_url": "https://arxiv.org/pdf/2604.22856.pdf"
    },
    {
      "id": "2604.22852",
      "anchor": "paper-2604-22852",
      "title": "SwarmDrive: Semantic V2V Coordination for Latency-Constrained Cooperative Autonomous Driving",
      "authors": [
        "Anjie Qiu",
        "Donglin Wang",
        "Zexin Fang",
        "Sanket Partani",
        "Hans D. Schotten"
      ],
      "abstract": "Cloud-hosted LLM inference for autonomous driving adds round-trip delay and depends on stable connectivity, while purely local edge models struggle under occlusion. We present SwarmDrive, a semantic Vehicle-to-Vehicle (V2V) coordination framework in which nearby vehicles run local Small Language Models (SLMs), share compact intent distributions only when uncertainty is high, and fuse them through event-triggered consensus. We evaluate SwarmDrive in a 5-seed executable study built around one occluded intersection case, combining matched operating-point comparisons with robustness sweeps. In that setting, SwarmDrive under its 6G communication setting (\"Swarm 6G\") raises success from 68.9% to 94.1% over a single local SLM while reducing latency from a 510 ms cloud reference to 151.4 ms. However, an increased number of participating vehicles leads to higher communication overhead and packet loss. SwarmDrive also evaluates the impact of swarm-size, packet-loss, and entropy-threshold sweeps and shows that the cooperative gain holds across ablations and is best balanced near an active swarm size of 4 vehicles and an entropy trigger threshold of 0.65 in the current prototype. These results show that semantic edge cooperation can work under tight latency constraints in the targeted intersection case, but they are not a deployment-grade validation of a real 6G stack.",
      "short_abstract": "Cloud-hosted LLM inference for autonomous driving adds round-trip delay and depends on stable connectivity, while purely local edge models struggle under occlusion. We present SwarmDrive, a semantic Vehicle-to-Vehicle (V2V) coordination framework in which...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 24,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.22852",
      "pdf_url": "https://arxiv.org/pdf/2604.22852.pdf"
    },
    {
      "id": "2604.22851",
      "anchor": "paper-2604-22851",
      "title": "EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving",
      "authors": [
        "Finn Rasmus Schäfer",
        "Yuan Gao",
        "Dingrui Wang",
        "Thomas Stauner",
        "Stephan Günnemann",
        "Mattia Piccinini",
        "Sebastian Schmidt",
        "Johannes Betz"
      ],
      "abstract": "While Vision-Language Models (VLMs) have advanced high-level reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench [Project page: (https://tum-avs.github.io/EgoDyn-Bench-Website/), Code: (https://github.com/TUM-AVS/EgoDyn-Bench), Dataset: (https://huggingface.co/datasets/fnc1901/EgoDyn-Bench)], a diagnostic benchmark for evaluating the semantic ego-motion understanding of vision-centric foundation models. By mapping continuous vehicle kinematics to discrete motion concepts via a deterministic oracle, we decouple a model's internal physical logic from its visual perception. Our large-scale empirical audit spanning 20$+$ models, including closed-source MLLMs, open-source VLMs across multiple scales, and specialized VLAs, identifies a significant Perception Bottleneck: while models exhibit logical physical concepts, they consistently fail to accurately align them with visual observations, frequently underperforming classical non-learned geometric baselines. This failure persists across model scales and domain-specific training, indicating a structural deficit in how current architectures couple visual perception with physical reasoning. We demonstrate that providing explicit trajectory encodings substantially restores physical consistency across all evaluated models, revealing a functional disentanglement between vision and language: ego-motion logic is derived almost exclusively from the language modality, while visual observations contribute negligible temporal signal. This structural finding provides a standardized diagnostic framework and a practical pathway toward physically aligned embodied AI. Ego-motion - Physical Reasoning - Foundation Models",
      "short_abstract": "While Vision-Language Models (VLMs) have advanced high-level reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench [Project page:...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: mapping",
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.22851",
      "pdf_url": "https://arxiv.org/pdf/2604.22851.pdf"
    },
    {
      "id": "2604.21249",
      "anchor": "paper-2604-21249",
      "title": "Reasoning About Traversability: Language-Guided Off-Road 3D Trajectory Planning",
      "authors": [
        "Byounggun Park",
        "Soonmin Hwang"
      ],
      "abstract": "While Vision-Language Models (VLMs) enable high-level semantic reasoning for end-to-end autonomous driving, particularly in unstructured environments, existing off-road datasets suffer from language annotations that are weakly aligned with vehicle actions and terrain geometry. To address this misalignment, we propose a language refinement framework that restructures annotations into action-aligned pairs, enabling a VLM to generate refined scene descriptions and 3D future trajectories directly from a single image. To further encourage terrain-aware planning, we introduce a preference optimization strategy that constructs geometry-aware hard negatives and explicitly penalizes trajectories inconsistent with local elevation profiles. Furthermore, we propose off-road-specific metrics to quantify traversability compliance and elevation consistency, addressing the limitations of conventional on-road evaluation. Experiments on the ORAD-3D benchmark demonstrate that our approach reduces average trajectory error from 1.01m to 0.97m, improves traversability compliance from 0.621 to 0.644, and decreases elevation inconsistency from 0.428 to 0.322, highlighting the efficacy of action-aligned supervision and terrain-aware optimization for robust off-road driving.",
      "short_abstract": "While Vision-Language Models (VLMs) enable high-level semantic reasoning for end-to-end autonomous driving, particularly in unstructured environments, existing off-road datasets suffer from language annotations that are weakly aligned with vehicle actions and...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 18,
        "evidence": [
          "title: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.21249",
      "pdf_url": "https://arxiv.org/pdf/2604.21249.pdf"
    },
    {
      "id": "2604.21124",
      "anchor": "paper-2604-21124",
      "title": "Enabling Mixed criticality applications for the Versal AI-Engines",
      "authors": [
        "Vincent Sprave",
        "Martin Wilhelm",
        "Daniele Passaretti",
        "Alberto Garcia-Ortiz",
        "Thilo Pionteck"
      ],
      "abstract": "Adaptive Systems-on-Chips (SoCs) are increasingly being used in mixed criticality systems (MCSs), such as in autonomous driving, aviation and medical systems. In this context, AMD has proposed the Versal SoC, which has a heterogeneous architecture including, among other components, an Artificial Intelligence Engine (AIE), which is a 2D array of processors and memory tiles designed for AI and signal processing workloads. While this AIE offers significant potential for accelerating real-time data processing tasks, this has not yet been explored in the context of MCSs since individual tasks with different criticality levels cannot be dynamically assigned to tiles due to the static mapping of dataflow graphs and tasks. In this work, we propose a dynamic task dispatching infrastructure that enables task switching on the AIE at runtime. Based on this infrastructure, we present an MCS design that dynamically assigns tasks of different criticality to a pool of AIE tiles, depending on the criticality mode of the system. Our approach overcomes the limitations of static dataflow graph mappings and, for the first time, exploits the parallel processing capabilities of the AIE for MCSs. We also present a comprehensive timing analysis of the overhead introduced by the task dispatcher infrastructure, focusing on control logic, context switching and data copy operations. This shows that these operations have low variance and are negligible compared to the overall execution time, demonstrating that our infrastructure is suitable for MCSs. Finally, we evaluate the proposed infrastructure using an autonomous driving workload with tasks that have variable execution times and different criticality levels. In this case study, we maximized AIE utilization, reducing idle time by 65.5 %, while measuring an execution time overhead of less than 0.002 %, and doubling the throughput of low-criticality tasks.",
      "short_abstract": "Adaptive Systems-on-Chips (SoCs) are increasingly being used in mixed criticality systems (MCSs), such as in autonomous driving, aviation and medical systems. In this context, AMD has proposed the Versal SoC, which has a heterogeneous architecture including,...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Mapping & Localization: abstract: mapping",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.21124",
      "pdf_url": "https://arxiv.org/pdf/2604.21124.pdf"
    },
    {
      "id": "2604.20423",
      "anchor": "paper-2604-20423",
      "title": "OVPD: A Virtual-Physical Fusion Testing Dataset of OnSite Auton-omous Driving Challenge",
      "authors": [
        "Yuhang Zhang",
        "Jiarui Zhang",
        "Bowen Jian",
        "Xin Zhou",
        "Zhichao Lv",
        "Peng Hang",
        "Rongjie Yu",
        "Ye Tian",
        "Jian Sun"
      ],
      "abstract": "The rapid iteration of autonomous driving algorithms has created a growing demand for high-fidelity, replayable, and diagnosable testing data. However, many public datasets lack real vehicle dynamics feedback and closed-loop interaction with surrounding traffic and road infrastructure, limiting their ability to reflect deployment readiness. To address this gap, we present OVPD (OnSite Virtual-Physical Dataset), a virtual-physical fusion testing dataset released from the 2025 OnSite Autonomous Driving Challenge. Centered on real-vehicle-in-the-loop testing, OVPD integrates virtual background traffic with vehicle-infrastructure perception to build controllable and interactive closed-loop test environments on a proving ground. The dataset contains 20 testing clips from 20 teams over a scenario chain of 15 atomic scenarios, totaling nearly 3 hours of multi-modal data, including vehicle trajectories and states, control commands, and digital-twin-rendered surround-view observations. OVPD supports long-tail planning and decision-making validation, open-loop or platform-enabled closed-loop evaluation, and comprehensive assessment across safety, efficiency, comfort, rule compliance, and traffic impact, providing actionable evidence for failure diagnosis and iterative improvement. The dataset is available via: https://huggingface.co/datasets/Yuhang253820/Onsite_OPVD",
      "short_abstract": "The rapid iteration of autonomous driving algorithms has created a growing demand for high-fidelity, replayable, and diagnosable testing data. However, many public datasets lack real vehicle dynamics feedback and closed-loop interaction with surrounding...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 11,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.20423",
      "pdf_url": "https://arxiv.org/pdf/2604.20423.pdf"
    },
    {
      "id": "2604.20289",
      "anchor": "paper-2604-20289",
      "title": "X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference",
      "authors": [
        "Yixiao Zeng",
        "Jianlei Zheng",
        "Chaoda Zheng",
        "Shijia Chen",
        "Mingdian Liu",
        "Tongping Liu",
        "Tengwei Luo",
        "Yu Zhang",
        "Boyang Wang",
        "Linkun Xu",
        "Siyuan Lu",
        "Bo Tian",
        "Xianming Liu"
      ],
      "abstract": "Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Recent driving world models built on autoregressive video diffusion achieve high-fidelity, controllable multi-camera generation, but their inference cost remains a bottleneck for interactive deployment. However, existing diffusion caching methods are designed for offline video generation with multiple denoising steps, and do not transfer to this scenario. Few-step distilled models have no inter-step redundancy left for these methods to reuse, and sequence-level parallelization techniques require future conditioning that closed-loop interactive generation does not provide. We present X-Cache, a training-free acceleration method that caches along a different axis: across consecutive generation chunks rather than across denoising steps. X-Cache maintains per-block residual caches that persist across chunks, and applies a dual-metric gating mechanism over a structure- and action-aware block-input fingerprint to independently decide whether each block should recompute or reuse its cached residual. To prevent approximation errors from permanently contaminating the autoregressive KV cache, X-Cache identifies KV update chunks (the forward passes that write clean keys and values into the persistent cache) and unconditionally forces full computation on these chunks, cutting off error propagation. We implement X-Cache on X-world, a production multi-camera action-conditioned driving world model built on multi-block causal DiT with few-step denoising and rolling KV cache. X-Cache achieves 71% block skip rate with 2.6x wall-clock speedup while maintaining minimum degradation.",
      "short_abstract": "Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Recent driving world models built on autoregressive video diffusion achieve high-fidelity, controllable...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "World Model",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.20289",
      "pdf_url": "https://arxiv.org/pdf/2604.20289.pdf"
    },
    {
      "id": "2604.20278",
      "anchor": "paper-2604-20278",
      "title": "Lightweight Low-SNR-Robust Semantic Communication System for Autonomous Driving",
      "authors": [
        "Ruixing Ren",
        "Minjie Wei",
        "Junhui Zhao"
      ],
      "abstract": "Image transmission for vehicle-to-vehicle collaborative perception in autonomous driving faces challenges including limited on-board terminal resources, time-varying wireless channel fading, and poor robustness under low signal-to-noise (SNR) ratio. Traditional separate source-channel coding schemes suffer from the cliff effect, while existing semantic communication models are limited by large parameter sizes and weak digital compatibility. This paper proposes a lightweight, low-SNR-robust deep joint source-channel coding (JSCC) semantic communication system. First, structured pruning is implemented based on batch normalization layer scaling factors and L1 regularization, which significantly reduces model complexity while ensuring image reconstruction quality. Second, a uniform quantization and M-QAM modulation scheme adapted to JSCC features is designed, and a training-deployment separation strategy is adopted to address the non-differentiable quantization problem, enabling compatibility with existing digital communication systems. Simulation results on the Cityscapes dataset show that the pruned model maintains comparable performance and robustness to the original one, even with over half of its parameters removed. Notably, the proposed scheme exhibits significant advantages over conventional communication methods under low SNR conditions.",
      "short_abstract": "Image transmission for vehicle-to-vehicle collaborative perception in autonomous driving faces challenges including limited on-board terminal resources, time-varying wireless channel fading, and poor robustness under low signal-to-noise (SNR) ratio....",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 12,
        "evidence": [
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.20278",
      "pdf_url": "https://arxiv.org/pdf/2604.20278.pdf"
    },
    {
      "id": "2604.20231",
      "anchor": "paper-2604-20231",
      "title": "Toward Cooperative Driving in Mixed Traffic: An Adaptive Potential Game-Based Approach with Field Test Verification",
      "authors": [
        "Shiyu Fang",
        "Xiaocong Zhao",
        "Xuekai Liu",
        "Peng Hang",
        "Jianqiang Wang",
        "Yunpeng Wang",
        "Jian Sun"
      ],
      "abstract": "Connected autonomous vehicles (CAVs), which represent a significant advancement in autonomous driving technology, have the potential to greatly increase traffic safety and efficiency through cooperative decision-making. However, existing methods often overlook the individual needs and heterogeneity of cooperative participants, making it difficult to transfer them to environments where they coexist with human-driven vehicles (HDVs).To address this challenge, this paper proposes an adaptive potential game (APG) cooperative driving framework. First, the system utility function is established on the basis of a general form of individual utility and its monotonic relationship, allowing for the simultaneous optimization of both individual and system objectives. Second, the Shapley value is introduced to compute each vehicle's marginal utility within the system, allowing its varying impact to be quantified. Finally, the HDV preference estimation is dynamically refined by continuously comparing the observed HDV behavior with the APG's estimated actions, leading to improvements in overall system safety and efficiency. Ablation studies demonstrate that adaptively updating Shapley values and HDV preference estimation significantly improve cooperation success rates in mixed traffic. Comparative experiments further highlight the APG's advantages in terms of safety and efficiency over other cooperative methods. Moreover, the applicability of the approach to real-world scenarios was validated through field tests.",
      "short_abstract": "Connected autonomous vehicles (CAVs), which represent a significant advancement in autonomous driving technology, have the potential to greatly increase traffic safety and efficiency through cooperative decision-making. However, existing methods often...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 19,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "title: verification"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.20231",
      "pdf_url": "https://arxiv.org/pdf/2604.20231.pdf"
    },
    {
      "id": "2604.20191",
      "anchor": "paper-2604-20191",
      "title": "From Scene to Object: Text-Guided Dual-Gaze Prediction",
      "authors": [
        "Zehong Ke",
        "Yanbo Jiang",
        "Jinhao Li",
        "Zhiyuan Liu",
        "Yiqian Tu",
        "Qingwen Meng",
        "Heye Huang",
        "Jianqiang Wang"
      ],
      "abstract": "Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaze rather than fine-grained object-level annotations, inherently failing to support text-grounded cognitive modeling. Consequently, while Vision-Language Models (VLMs) hold great potential for semantic reasoning, this critical data limitations leads to severe text-vision decoupling and visual-bias hallucinations. To break this bottleneck and achieve precise object-level attention prediction, this paper proposes a novel dual-branch gaze prediction framework, establishing a complete paradigm from data construction to model architecture. First, we construct G-W3DA, a object-level driver attention dataset. By integrating a multimodal large language model with the Segment Anything Model 3 (SAM3), we decouple macroscopic heatmaps into object-level masks under rigorous cross-validation, fundamentally eliminating annotation hallucinations. Building upon this high-quality data foundation, we propose the DualGaze-VLM architecture. This architecture extracts the hidden states of semantic queries and dynamically modulates visual features via a Condition-Aware SE-Gate, achieving intent-driven precise spatial anchoring. Extensive experiments on the W3DA benchmark demonstrate that DualGaze-VLM consistently surpasses existing state-of-the-art (SOTA) models in spatial alignment metrics, notably achieving up to a 17.8% improvement in Similarity (SIM) under safety-critical scenarios. Furthermore, a visual Turing test reveals that the attention heatmaps generated by DualGaze-VLM are perceived as authentic by 88.22% of human evaluators, proving its capability to generate rational cognitive priors.",
      "short_abstract": "Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaze rather than fine-grained object-level annotations, inherently failing to support text-grounded cognitive...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Dataset",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 1,
        "evidence": [
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.20191",
      "pdf_url": "https://arxiv.org/pdf/2604.20191.pdf"
    },
    {
      "id": "2604.20621",
      "anchor": "paper-2604-20621",
      "title": "SoK: The Next Frontier in AV Security: Systematizing Perception Attacks and the Emerging Threat of Multi-Sensor Fusion",
      "authors": [
        "Shahriar Rahman Khan",
        "Tariqul Islam",
        "Raiful Hasan"
      ],
      "abstract": "Autonomous vehicles (AVs) increasingly rely on multi-sensor perception pipelines that combine data from cameras, lidar, radar, and other modalities to interpret the environment. This SoK systematizes 48 peer-reviewed studies on perception-layer attacks against AVs, tracking the field's evolution from single-sensor exploits to complex cross-modal threats that compromise multi-sensor fusion (MSF). We develop a unified taxonomy of 20 attack vectors organized by sensor type, attack stage, medium, and perception module, revealing patterns that expose underexplored vulnerabilities in fusion logic and cross-sensor dependencies. Our analysis identifies key research gaps, including limited real-world testing, short-term evaluation bias, and the absence of defenses that account for inter-sensor consistency. To illustrate one such gap, we validate a fusion-level vulnerability through a proof-of-concept simulation combining infrared and lidar spoofing. The findings highlight a fundamental shift in AV security: as systems fuse more sensors for robustness, attackers exploit the very redundancy meant to ensure safety. We conclude with directions for fusion-aware defense design and a research agenda for trustworthy perception in autonomous systems.",
      "short_abstract": "Autonomous vehicles (AVs) increasingly rely on multi-sensor perception pipelines that combine data from cameras, lidar, radar, and other modalities to interpret the environment. This SoK systematizes 48 peer-reviewed studies on perception-layer attacks...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 45,
        "margin": 9,
        "evidence": [
          "abstract: safety",
          "title: security",
          "abstract: security",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.20621",
      "pdf_url": "https://arxiv.org/pdf/2604.20621.pdf"
    },
    {
      "id": "2604.20428",
      "anchor": "paper-2604-20428",
      "title": "Lexicographic Minimum-Violation Motion Planning using Signal Temporal Logic",
      "authors": [
        "Patrick Halder",
        "Lothar Kiltz",
        "Hannes Homburger",
        "Johannes Reuter",
        "Matthias Althoff"
      ],
      "abstract": "Motion planning for autonomous vehicles often requires satisfying multiple conditionally conflicting specifications. In situations where not all specifications can be met simultaneously, minimum-violation motion planning maintains system operation by minimizing violations of specifications in accordance with their priorities. Signal temporal logic (STL) provides a formal language for rigorously defining these specifications and enables the quantitative evaluation of their violations. However, a total ordering of specifications yields a lexicographic optimization problem, which is typically computationally expensive to solve using standard methods. We address this problem by transforming the multi-objective lexicographic optimization problem into a single-objective scalar optimization problem using non-uniform quantization and bit-shifting. Specifically, we extend a deterministic model predictive path integral (MPPI) solver to efficiently solve optimization problems without quadratic input cost. Additionally, a novel predicate-robustness measure that combines spatial and temporal violations is introduced. Our results show that the proposed method offers an interpretable and scalable solution for lexicographic STL minimum-violation motion planning within a single-objective solver framework.",
      "short_abstract": "Motion planning for autonomous vehicles often requires satisfying multiple conditionally conflicting specifications. In situations where not all specifications can be met simultaneously, minimum-violation motion planning maintains system operation by...",
      "published": "2026-04-22",
      "updated": "2026-04-22",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 30,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 119,
      "arxiv_url": "https://arxiv.org/abs/2604.20428",
      "pdf_url": "https://arxiv.org/pdf/2604.20428.pdf"
    },
    {
      "id": "2604.20895",
      "anchor": "paper-2604-20895",
      "title": "Towards a Systematic Risk Assessment of Deep Neural Network Limitations in Autonomous Driving Perception",
      "authors": [
        "Svetlana Pavlitska",
        "Christopher Gerking",
        "J. Marius Zöllner"
      ],
      "abstract": "Safety and security are essential for the admission and acceptance of automated and autonomous vehicles. Deep neural networks (DNNs) are widely used for perception and further components of the autonomous driving (AD) stack. However, they possess several limitations, including lack of generalization, efficiency, explainability, plausibility, and robustness. These insufficiencies can pose significant risks to autonomous driving systems. However, hazards, threats, and risks associated with DNN limitations in this domain have not been systematically studied so far. In this work, we propose a joint workflow for risk assessment combining the hazard analysis and risk assessment (HARA) following ISO 26262 and threat analysis and risk assessment (TARA) following the ISO/SAE 21434 to identify and analyze risks arising from inherent DNN limitations in AD perception.",
      "short_abstract": "Safety and security are essential for the admission and acceptance of automated and autonomous vehicles. Deep neural networks (DNNs) are widely used for perception and further components of the autonomous driving (AD) stack. However, they possess several...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 19,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: attack or threat",
          "title: risk assessment",
          "abstract: risk assessment",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.20895",
      "pdf_url": "https://arxiv.org/pdf/2604.20895.pdf"
    },
    {
      "id": "2604.19741",
      "anchor": "paper-2604-19741",
      "title": "CityRAG: Stepping Into a City via Spatially-Grounded Video Generation",
      "authors": [
        "Gene Chou",
        "Charles Herrmann",
        "Kyle Genova",
        "Boyang Deng",
        "Songyou Peng",
        "Bharath Hariharan",
        "Jason Y. Zhang",
        "Noah Snavely",
        "Philipp Henzler"
      ],
      "abstract": "We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V) or image (I2V) prompt. However, the capability to reconstruct the real world under arbitrary weather conditions and dynamic object configurations is essential for downstream applications including autonomous driving and robotics simulation. To this end, we present CityRAG, a video generative model that leverages large corpora of geo-registered data as context to ground generation to the physical scene, while maintaining learned priors for complex motion and appearance changes. CityRAG relies on temporally unaligned training data, which teaches the model to semantically disentangle the underlying scene from its transient attributes. Our experiments demonstrate that CityRAG can generate coherent minutes-long, physically grounded video sequences, maintain weather and lighting conditions over thousands of frames, achieve loop closure, and navigate complex trajectories to reconstruct real-world geography.",
      "short_abstract": "We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V) or image (I2V)...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19741",
      "pdf_url": "https://arxiv.org/pdf/2604.19741.pdf"
    },
    {
      "id": "2604.19710",
      "anchor": "paper-2604-19710",
      "title": "SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model",
      "authors": [
        "Zewei Zhou",
        "Ruining Yang",
        "Xuewei",
        "Qi",
        "Yiluan Guo",
        "Sherry X. Chen",
        "Tao Feng",
        "Kateryna Pistunova",
        "Yishan Shen",
        "Lili Su",
        "Jiaqi Ma"
      ],
      "abstract": "Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high latency in action generation using an autoregressive generation framework and exhibit limited robustness. In this paper, we propose SpanVLA, a novel end-to-end autonomous driving framework, integrating an autoregressive reasoning and a flow-matching action expert. First, SpanVLA introduces an efficient bridge to leverage the vision and reasoning guidance of VLM to efficiently plan future trajectories using a flow-matching policy conditioned on historical trajectory initialization, which significantly reduces inference time. Second, to further improve the performance and robustness of the SpanVLA model, we propose a GRPO-based post-training method to enable the VLA model not only to learn from positive driving samples but also to learn how to avoid the typical negative behaviors and learn recovery behaviors. We further introduce mReasoning, a new real-world driving reasoning dataset, focusing on complex, reasoning-demanding scenarios and negative-recovery samples. Extensive experiments on the NAVSIM (v1 and v2) demonstrate the competitive performance of the SpanVLA model. Additionally, the qualitative results across diverse scenarios highlight the planning performance and robustness of our model.",
      "short_abstract": "Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high latency in action...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "VLA",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 29,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19710",
      "pdf_url": "https://arxiv.org/pdf/2604.19710.pdf"
    },
    {
      "id": "2604.19596",
      "anchor": "paper-2604-19596",
      "title": "PC2Model: ISPRS benchmark on 3D point cloud to model registration",
      "authors": [
        "Mehdi Maboudi",
        "Said Harb",
        "Jackson Ferrao",
        "Kourosh Khoshelham",
        "Yelda Turkan",
        "Karam Mawas"
      ],
      "abstract": "Point cloud registration involves aligning one point cloud with another or with a three-dimensional (3D) model, enabling the integration of multimodal data into a unified representation. This is essential in applications such as construction monitoring, autonomous driving, robotics, and virtual or augmented reality (VR/AR). With the increasing accessibility of point cloud acquisition technologies, such as Light Detection and Ranging (LiDAR) and structured light scanning, along with recent advances in deep learning, the research focus has increasingly shifted towards downstream tasks, particularly point cloud-to-model (PC2Model) registration. While data-driven methods aim to automate this process, they struggle with sparsity, noise, clutter, and occlusions in real-world scans, which limit their performance. To address these challenges, this paper introduces the PC2Model benchmark, a publicly available dataset designed to support the training and evaluation of both classical and data-driven methods. Developed under the leadership of ICWG II/Ib, the PC2Model benchmark adopts a hybrid design that combines simulated point clouds with, in some cases, real-world scans and their corresponding 3D models. Simulated data provide precise ground truth and controlled conditions, while real-world data introduce sensor and environmental artefacts. This design supports robust training and evaluation across domains and enables the systematic analysis of model transferability from simulated to real-world scenarios. The dataset is publicly accessible at: https://zenodo.org/records/17581812",
      "short_abstract": "Point cloud registration involves aligning one point cloud with another or with a three-dimensional (3D) model, enabling the integration of multimodal data into a unified representation. This is essential in applications such as construction monitoring,...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 13,
        "evidence": [
          "title: point cloud",
          "abstract: point cloud",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19596",
      "pdf_url": "https://arxiv.org/pdf/2604.19596.pdf"
    },
    {
      "id": "2604.19412",
      "anchor": "paper-2604-19412",
      "title": "VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing",
      "authors": [
        "Yanbin Huang",
        "Yisen Li",
        "Guiyao Tie",
        "Xiaoye Qu",
        "Pan Zhou",
        "Hongfei Wang",
        "Zhaofan Zou",
        "Hao Sun",
        "Xuelong Li"
      ],
      "abstract": "Large vision-language models (LVLMs) frequently suffer from Object Hallucination (OH), wherein they generate descriptions containing objects that are not actually present in the input image. This phenomenon is particularly problematic in real-world applications such as medical imaging and autonomous driving, where accuracy is critical. Recent studies suggest that the hallucination problem may stem from language priors: biases learned during pretraining that cause LVLMs to generate words based on their statistical co-occurrence. To mitigate this problem, we propose Visual Contrastive Editing (VCE), a novel post-hoc method that identifies and suppresses hallucinatory tendencies by analyzing the model's response to contrastive visual perturbations. Using Singular Value Decomposition (SVD), we decompose the model's activation patterns to isolate hallucination subspaces and apply targeted parameter edits to attenuate its influence. Unlike existing approaches that require fine-tuning or labeled data, VCE operates as a label-free intervention, making it both scalable and practical for deployment in resource-constrained settings. Experimental results demonstrate that VCE effectively reduces object hallucination across multiple benchmarks while maintaining the model's original computational efficiency.",
      "short_abstract": "Large vision-language models (LVLMs) frequently suffer from Object Hallucination (OH), wherein they generate descriptions containing objects that are not actually present in the input image. This phenomenon is particularly problematic in real-world...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19412",
      "pdf_url": "https://arxiv.org/pdf/2604.19412.pdf"
    },
    {
      "id": "2604.19379",
      "anchor": "paper-2604-19379",
      "title": "PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving",
      "authors": [
        "Yining Pan",
        "Shijie Li",
        "Yuchen Wu",
        "Xulei Yang",
        "Na Zhao"
      ],
      "abstract": "This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts commonly encountered in real-world autonomous driving. A straightforward solution is to employ a pseudo-labeling strategy, which is widely used in UDA to generate supervision for unlabeled target data, combined with an mm-3DPS backbone. However, existing supervised mm-3DPS methods rely heavily on strong cross-modal complementarity between LiDAR and RGB inputs, making them fragile under domain shifts where one modality degrades (e.g., poor lighting or adverse weather). Moreover, conventional pseudo-labeling typically retains only high-confidence regions, leading to fragmented masks and incomplete object supervision, which are issues particularly detrimental to panoptic segmentation. To address these challenges, we propose PanDA, the first UDA framework specifically designed for multimodal 3D panoptic segmentation. To improve robustness against single-sensor degradation, we introduce an asymmetric multimodal augmentation that selectively drops regions to simulate domain shifts and improve robust representation learning. To enhance pseudo-label completeness and reliability, we further develop a dual-expert pseudo-label refinement module that extracts domain-invariant priors from both 2D and 3D modalities. Extensive experiments across diverse domain shifts, spanning time, weather, location, and sensor variations, significantly surpass state-of-the-art UDA baselines for 3D semantic segmentation.",
      "short_abstract": "This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts commonly encountered in real-world autonomous driving. A straightforward...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 17,
        "evidence": [
          "title: scene segmentation",
          "abstract: scene segmentation",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19379",
      "pdf_url": "https://arxiv.org/pdf/2604.19379.pdf"
    },
    {
      "id": "2604.19257",
      "anchor": "paper-2604-19257",
      "title": "Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images",
      "authors": [
        "Hongyuan Liu",
        "Bochao Zou",
        "Qiankun Liu",
        "Haochen Yu",
        "Qi Mei",
        "Jianfei Jiang",
        "Chen Liu",
        "Cheng Bi",
        "Zhao Wang",
        "Xueyang Zhang",
        "Yifei Zhan",
        "Jiansheng Chen",
        "Huimin Ma"
      ],
      "abstract": "Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D vehicle generation methods are often trained on synthetic data with significant domain gaps from real-world distributions. The generated models often exhibit arbitrary poses and undefined scales, resulting in poor visual consistency when integrated into driving scenes. In this paper, we present Unposed-to-3D, a novel framework that learns to reconstruct 3D vehicles from real-world driving images using image-only supervision. Our approach consists of two stages. In the first stage, we train an image-to-3D reconstruction network using posed images with known camera parameters. In the second stage, we remove camera supervision and use a camera prediction head that directly estimates the camera parameters from unposed images. The predicted pose is then used for differentiable rendering to provide self-supervised photometric feedback, enabling the model to learn 3D geometry purely from unposed images. To ensure simulation readiness, we further introduce a scale-aware module to predict real-world size information, and a harmonization module that adapts the generated vehicles to the target driving scene with consistent lighting and appearance. Extensive experiments demonstrate that Unposed-to-3D effectively reconstructs realistic, pose-consistent, and harmonized 3D vehicle models from real-world images, providing a scalable path toward creating high-quality assets for driving scene simulation and digital twin environments.",
      "short_abstract": "Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D vehicle generation methods are often trained on synthetic data with significant domain gaps from real-world...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: camera",
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19257",
      "pdf_url": "https://arxiv.org/pdf/2604.19257.pdf"
    },
    {
      "id": "2604.19206",
      "anchor": "paper-2604-19206",
      "title": "When Can We Trust Deep Neural Networks? Towards Reliable Industrial Deployment with an Interpretability Guide",
      "authors": [
        "Hang-Cheng Dong",
        "Yuhao Jiang",
        "Yibo Jiao",
        "Lu Zou",
        "Kai Zheng",
        "Bingguo Liu",
        "Dong Ye",
        "Guodong Liu"
      ],
      "abstract": "The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely hampered by their lack of reliability. A single undetected erroneous prediction can lead to catastrophic outcomes. Unfortunately, there is often no alternative but to place trust in the outputs of a trained AI system, which operates without an internal safeguard to flag unreliable predictions, even in cases of high accuracy. We propose a post-hoc explanation-based indicator to detect false negatives in binary defect detection networks. To our knowledge, this is the first method to proactively identify potentially erroneous network outputs. Our core idea leverages the difference between class-specific discriminative heatmaps and class-agnostic ones. We compute the difference in their intersection over union (IoU) as a reliability score. An adversarial enhancement method is further introduced to amplify this disparity. Evaluations on two industrial defect detection benchmarks show our method effectively identifies false negatives. With adversarial enhancement, it achieves 100\\% recall, albeit with a trade-off for true negatives. Our work thus advocates for a new and trustworthy deployment paradigm: data-model-explanation-output, moving beyond conventional end-to-end systems to provide critical support for reliable AI in real-world applications.",
      "short_abstract": "The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely hampered by their lack of reliability. A single undetected erroneous prediction can lead to catastrophic...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Benchmark",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 12,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19206",
      "pdf_url": "https://arxiv.org/pdf/2604.19206.pdf"
    },
    {
      "id": "2604.19145",
      "anchor": "paper-2604-19145",
      "title": "ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving",
      "authors": [
        "Lin Sha",
        "Haiyun Guo",
        "Tao Wang",
        "Cong Zhang",
        "Min Huang",
        "Jinqiao Wang",
        "Qinghai Miao"
      ],
      "abstract": "Vision-Language Models (VLMs) have become central to autonomous driving systems, yet their deployment is severely bottlenecked by the massive computational overhead of multi-view camera and multi-frame video input. Existing token pruning methods, primarily designed for single-image inputs, treat each frame or view in isolation and thus fail to exploit the inherent spatio-temporal redundancies in driving scenarios. To bridge this gap, we propose ST-Prune, a training-free, plug-and-play framework comprising two complementary modules: Motion-aware Temporal Pruning (MTP) and Ring-view Spatial Pruning (RSP). MTP addresses temporal redundancy by encoding motion volatility and temporal recency as soft constraints within the diversity selection objective, prioritizing dynamic trajectories and current-frame content over static historical background. RSP further resolves spatial redundancy by exploiting the ring-view camera geometry to penalize bilateral cross-view similarity, eliminating duplicate projections and residual background that temporal pruning alone cannot suppress. These two modules together constitute a complete spatio-temporal pruning process, preserving key scene information under strict compression. Validated across four benchmarks spanning perception, prediction, and planning, ST-Prune establishes new state-of-the-art for training-free token pruning. Notably, even at 90\\% token reduction, ST-Prune achieves near-lossless performance with certain metrics surpassing the full-model baseline, while maintaining inference speeds comparable to existing pruning approaches.",
      "short_abstract": "Vision-Language Models (VLMs) have become central to autonomous driving systems, yet their deployment is severely bottlenecked by the massive computational overhead of multi-view camera and multi-frame video input. Existing token pruning methods, primarily...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: camera",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19145",
      "pdf_url": "https://arxiv.org/pdf/2604.19145.pdf"
    },
    {
      "id": "2604.19914",
      "anchor": "paper-2604-19914",
      "title": "AI Incident Monitoring through a Public Health Lens",
      "authors": [
        "Sophia Abraham",
        "Taiye Chen",
        "Cyril Chhun",
        "Giovanna Jaramillo-Gutierrez",
        "Simon Mylius",
        "Sayash Raaj",
        "Peter Slattery",
        "Sean McGregor"
      ],
      "abstract": "Artificial intelligence systems are now deployed at scale across sectors, accompanied by a growing number of real-world incidents ranging from misinformation and cybercrime to autonomous-system failures. Databases of AI incidents index these events, but they cannot measure ``risk'' (i.e., a joint measure of likelihood and severity) without additional data regarding the prevalence of risk-associated systems and their incident reporting rates. As a result, policymakers, companies, and the general public lack a means to weigh the benefits of AI against their in-context risks. Inspired by public-health processes, which presume noisy and incomplete disease surveillance, we identify six phases of incident emergence. We demonstrate the framework through a detailed case study of autonomous vehicles, whose mandatory reporting requirements produces reliable incident-rate ground truth expressed in distance traveled. The case study shows that an informed panel of domain experts (e.g., self-driving experts) can combine their domain expertise, incident data, and a collection of statistical and visualization tools to arrive at incident phase determinations serving public needs. We further demonstrate the approach with a deepfake incident case study and chart a path for future research in incident phase determination.",
      "short_abstract": "Artificial intelligence systems are now deployed at scale across sectors, accompanied by a growing number of real-world incidents ranging from misinformation and cybercrime to autonomous-system failures. Databases of AI incidents index these events, but they...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19914",
      "pdf_url": "https://arxiv.org/pdf/2604.19914.pdf"
    },
    {
      "id": "2604.19838",
      "anchor": "paper-2604-19838",
      "title": "Resolving space-sharing conflicts in road user interactions through uncertainty reduction: An active inference-based computational model",
      "authors": [
        "Julian F. Schumann",
        "Johan Engström",
        "Ran Wei",
        "Shu-Yuan Liu",
        "Jens Kober",
        "Arkady Zgonnikov"
      ],
      "abstract": "Understanding how road users resolve space-sharing conflicts is important both for traffic safety and the safe deployment of autonomous vehicles. While existing models have captured specific aspects of such interactions (e.g., explicit communication), a theoretically-grounded computational framework has been lacking. In this paper, we extend a previously developed active inference-based driver behavior model to simulate interactive behavior of two agents. Our model captures three complementary mechanisms for uncertainty reduction in interaction: (i) implicit communication via direct behavioral coupling, (ii) reliance on normative expectations (stop signs, priority rules, etc.), and (iii) explicit communication. In a simplified intersection scenario, we show that normative and explicit communication cues can increase the likelihood of a successful conflict resolution. However, this relies on agents acting as expected. In situations where another agent (intentionally or unintentionally) violates normative expectations or communicates misleading information, reliance on these cues may induce collisions. These findings illustrate how active inference can provide a novel framework for modeling road user interactions which is also applicable in other fields.",
      "short_abstract": "Understanding how road users resolve space-sharing conflicts is important both for traffic safety and the safe deployment of autonomous vehicles. While existing models have captured specific aspects of such interactions (e.g., explicit communication), a...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "title: uncertainty",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19838",
      "pdf_url": "https://arxiv.org/pdf/2604.19838.pdf"
    },
    {
      "id": "2604.19452",
      "anchor": "paper-2604-19452",
      "title": "Robust Nonlinear Trajectory Tracking Control for Autonomous Racing on Three-Dimensional Tracks",
      "authors": [
        "Joscha F. Bongard",
        "Georg Jank",
        "Simon Sagmeister",
        "Boris Lohmann"
      ],
      "abstract": "We propose a robust nonlinear model predictive control (MPC) scheme for trajectory-tracking control of autonomous vehicles at the limits of handling on non-planar road surfaces. We derive the dynamics from first principles and selectively omit terms with negligible dynamic influence to maintain real-time capability. The resulting MPC with a three-dimensional (3D) dynamic single-track model integrates relevant dynamic effects directly into the prediction model and leverages them to improve prediction accuracy and therefore control performance. Even if the influence of terrain-induced vertical loads on the total acceleration potential is modeled, tire-road interactions are subject to uncertainty and disturbance. The uncertainty-aware constraint tightening scheme introduces a margin to constraint bounds to keep the vehicle controllable and stable in this environment. To validate our proposed approach, we perform high-fidelity dynamic double-track vehicle dynamics simulations on a model of a real circuit. We find that our algorithm can improve trajectory-tracking accuracy while maintaining low computation times.",
      "short_abstract": "We propose a robust nonlinear model predictive control (MPC) scheme for trajectory-tracking control of autonomous vehicles at the limits of handling on non-planar road surfaces. We derive the dynamics from first principles and selectively omit terms with...",
      "published": "2026-04-21",
      "updated": "2026-04-21",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 22,
        "evidence": [
          "abstract: model predictive control",
          "abstract: vehicle dynamics",
          "title: control",
          "abstract: control",
          "abstract: steering or braking",
          "title: trajectory tracking"
        ]
      },
      "recency": "archive",
      "age_days": 120,
      "arxiv_url": "https://arxiv.org/abs/2604.19452",
      "pdf_url": "https://arxiv.org/pdf/2604.19452.pdf"
    },
    {
      "id": "2604.18993",
      "anchor": "paper-2604-18993",
      "title": "AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos",
      "authors": [
        "Jiagao Hu",
        "Daiguo Zhou",
        "Danzhen Fu",
        "Fuhao Li",
        "Zepeng Wang",
        "Fei Wang",
        "Wenhua Liao",
        "Jiayi Xie",
        "Haiyang Sun"
      ],
      "abstract": "Perception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-world video data in adverse weather. Existing weather generation approaches struggle to balance visual quality and annotation reusability. We present AutoAWG, a controllable Adverse Weather video Generation framework for Autonomous driving. Our method employs a semantics-guided adaptive fusion of multiple controls to balance strong weather stylization with high-fidelity preservation of safety-critical targets; leverages a vanishing point-anchored temporal synthesis strategy to construct training sequences from static images, thereby reducing reliance on synthetic data; and adopts masked training to enhance long-horizon generation stability. On the nuScenes validation set, AutoAWG significantly outperforms prior state-of-the-art methods: without first-frame conditioning, FID and FVD are relatively reduced by 50.0% and 16.1%; with first-frame conditioning, they are further reduced by 8.7% and 7.2%, respectively. Extensive qualitative and quantitative results demonstrate advantages in style fidelity, temporal consistency, and semantic--structural integrity, underscoring the practical value of AutoAWG for improving downstream perception in autonomous driving. Our code is available at: https://github.com/higherhu/AutoAWG",
      "short_abstract": "Perception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-world video data in adverse weather. Existing weather generation approaches struggle to balance visual quality...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 6,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.18993",
      "pdf_url": "https://arxiv.org/pdf/2604.18993.pdf"
    },
    {
      "id": "2604.18940",
      "anchor": "paper-2604-18940",
      "title": "Localization-Guided Foreground Augmentation in Autonomous Driving",
      "authors": [
        "Jiawei Yong",
        "Deyuan Qu",
        "Qi Chen",
        "Kentaro Oguchi",
        "Shintaro Fukushima"
      ],
      "abstract": "Autonomous driving systems often degrade under adverse visibility conditions-such as rain, nighttime, or snow-where online scene geometry (e.g., lane dividers, road boundaries, and pedestrian crossings) becomes sparse or fragmented. While high-definition (HD) maps can provide missing structural context, they are costly to construct and maintain at scale. We propose Localization-Guided Foreground Augmentation (LG-FA), a lightweight and plug-and-play inference module that enhances foreground perception by enriching geometric context online. LG-FA: (i) incrementally constructs a sparse global vector layer from per-frame Bird's-Eye View (BEV) predictions; (ii) estimates ego pose via class-constrained geometric alignment, jointly improving localization and completing missing local topology; and (iii) reprojects the augmented foreground into a unified global frame to improve per-frame predictions. Experiments on challenging nuScenes sequences demonstrate that LG-FA improves the geometric completeness and temporal stability of BEV representations, reduces localization error, and produces globally consistent lane and topology reconstructions. The module can be seamlessly integrated into existing BEV-based perception systems without backbone modification. By providing a reliable geometric context prior, LG-FA enhances temporal consistency and supplies stable structural support for downstream modules such as tracking and decision-making.",
      "short_abstract": "Autonomous driving systems often degrade under adverse visibility conditions-such as rain, nighttime, or snow-where online scene geometry (e.g., lane dividers, road boundaries, and pedestrian crossings) becomes sparse or fragmented. While high-definition (HD)...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 15,
        "evidence": [
          "title: localization",
          "abstract: localization"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.18940",
      "pdf_url": "https://arxiv.org/pdf/2604.18940.pdf"
    },
    {
      "id": "2604.18918",
      "anchor": "paper-2604-18918",
      "title": "From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing",
      "authors": [
        "Linfeng Liang",
        "Xiao Cheng",
        "Tsong Yueh Chen",
        "Xi Zheng"
      ],
      "abstract": "Simulation-based testing of autonomous driving systems (ADS) must uncover realistic and diverse failures in dense, heterogeneous traffic. However, existing search-based seeding methods (e.g., genetic algorithms) struggle in high-dimensional spaces, often collapsing to limited modes and missing many failure scenarios. We present PtoP, a framework that combines adaptive random seed generation with Stein Variational Gradient Descent (SVGD) to produce diverse, failure-inducing initial conditions. SVGD balances attraction toward high-risk regions and repulsion among particles, yielding risk-seeking yet well-distributed seeds across multiple failure modes. PtoP is plug-and-play and enhances existing online testing methods (e.g., reinforcement learning--based testers) by providing principled seeds. Evaluation in CARLA on two industry-grade ADS (Apollo, Autoware) and a native end-to-end system shows that PtoP improves safety violation rate (up to 27.68%), scenario diversity (9.6%), and map coverage (16.78%) over baselines.",
      "short_abstract": "Simulation-based testing of autonomous driving systems (ADS) must uncover realistic and diverse failures in dense, heterogeneous traffic. However, existing search-based seeding methods (e.g., genetic algorithms) struggle in high-dimensional spaces, often...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 14,
        "evidence": [
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.18918",
      "pdf_url": "https://arxiv.org/pdf/2604.18918.pdf"
    },
    {
      "id": "2604.18831",
      "anchor": "paper-2604-18831",
      "title": "Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model",
      "authors": [
        "Haiyang Wu",
        "Juan J. Gonzales Torres",
        "George Vosselman",
        "Ville Lehtola"
      ],
      "abstract": "Frame-wise semantic segmentation of indoor lidar scans is a fundamental step toward higher-level 3D scene understanding and mapping applications. However, acquiring frame-wise ground truth for training deep learning models is costly and time-consuming. This challenge is largely addressed, for imagery, by Visual Foundation Models (VFMs) which segment image frames. The same VFMs may be used to train a lidar scan frame segmentation model via a 2D-to-3D distillation pipeline. The success of such distillation has been shown for autonomous driving scenes, but not yet for indoor scenes. Here, we study the feasibility of repeating this success for indoor scenes, in a frame-wise distillation manner by coupling each lidar scan with a VFM-processed camera image. The evaluation is done using indoor SLAM datasets, where pseudo-labels are used for downstream evaluation. Also, a small manually annotated lidar dataset is provided for validation, as there are no other lidar frame-wise indoor datasets with semantics. Results show that the distilled model achieves up to 56% mIoU under pseudo-label evaluation and around 36% mIoU with real-label, demonstrating the feasibility of cross-modal distillation for indoor lidar semantic segmentation without manual annotations.",
      "short_abstract": "Frame-wise semantic segmentation of indoor lidar scans is a fundamental step toward higher-level 3D scene understanding and mapping applications. However, acquiring frame-wise ground truth for training deep learning models is costly and time-consuming. This...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 23,
        "evidence": [
          "title: scene segmentation",
          "abstract: scene segmentation",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: camera",
          "abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.18831",
      "pdf_url": "https://arxiv.org/pdf/2604.18831.pdf"
    },
    {
      "id": "2604.18486",
      "anchor": "paper-2604-18486",
      "title": "Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation",
      "authors": [
        "Jinghui Lu",
        "Jiayi Guan",
        "Zhijian Huang",
        "Jinlong Li",
        "Guang Li",
        "Lingdong Kong",
        "Yingyan Li",
        "Han Wang",
        "Shaoqing Xu",
        "Yuechen Luo",
        "Fang Li",
        "Chenxu Dang",
        "Junli Wang",
        "Tao Xu",
        "Jing Wu",
        "Jianhua Wu",
        "Xiaoshuai Hao",
        "Wen Zhang",
        "Tianyi Jiang",
        "Lingfeng Zhang",
        "Lei Zhou",
        "Yingbo Tang",
        "Jie Wang",
        "Yinfeng Gao",
        "Xizhou Bu"
      ],
      "abstract": "Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is prohibitive for real-time deployment. Latent CoT methods attempt to close this gap by compressing reasoning into continuous hidden states, but consistently fall short of their explicit counterparts. We suggest that this is due to purely linguistic latent representations compressing a symbolic abstraction of the world, rather than the causal dynamics that actually govern driving. Thus, we present OneVL (One-step latent reasoning and planning with Vision-Language explanations), a unified VLA and World Model framework that routes reasoning through compact latent tokens supervised by dual auxiliary decoders. Alongside a language decoder that reconstructs text CoT, we introduce a visual world model decoder that predicts future-frame tokens, forcing the latent space to internalize the causal dynamics of road geometry, agent motion, and environmental change. A three-stage training pipeline progressively aligns these latents with trajectory, language, and visual objectives, ensuring stable joint optimization. In inference, the auxiliary decoders are discarded, and all latent tokens are prefilled in a single parallel pass, matching the speed of answer-only prediction. Across four benchmarks, OneVL becomes the first latent CoT method to surpass explicit CoT, delivering superior accuracy at answer-only latency. These results show that with world model supervision, latent CoT produces more generalizable representations than verbose token-by-token reasoning. Code has been open-sourced to the community. Project Page: https://xiaomi-embodied-intelligence.github.io/OneVL",
      "short_abstract": "Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is prohibitive for real-time deployment. Latent CoT methods attempt to close this...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "VLA",
        "World Model",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 3,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: world model",
          "abstract: future-frame prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.18486",
      "pdf_url": "https://arxiv.org/pdf/2604.18486.pdf"
    },
    {
      "id": "2604.18476",
      "anchor": "paper-2604-18476",
      "title": "SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection",
      "authors": [
        "Hao Vo",
        "Khoa Vo",
        "Thinh Phan",
        "Ngo Xuan Cuong",
        "Gianfranco Doretto",
        "Hien Nguyen",
        "Anh Nguyen",
        "Ngan Le"
      ],
      "abstract": "Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performance while overlooking the severe long-tail imbalance inherent in real-world datasets. In practice, many rare but safety-critical categories such as children, strollers, or emergency vehicles are heavily underrepresented, leading to biased learning and degraded performance. This challenge is further exacerbated by pronounced inter-class ambiguity (e.g., visually similar subclasses) and substantial intra-class diversity (e.g., objects varying widely in appearance, scale, pose, or context), which together hinder reliable long-tail recognition. In this work, we introduce SemLT3D, a Semantic-Guided Expert Distillation framework designed to enrich the representation space for underrepresented classes through semantic priors. SemLT3D consists of: (1) a language-guided mixture-of-experts module that routes 3D queries to specialized experts according to their semantic affinity, enabling the model to better disentangle confusing classes and specialize on tail distributions; and (2) a semantic projection distillation pipeline that aligns 3D queries with CLIP-informed 2D semantics, producing more coherent and discriminative features across diverse visual manifestations. Although motivated by long-tail imbalance, the semantically structured learning in SemLT3D also improves robustness under broader appearance variations and challenging corner cases, offering a principled step toward more reliable camera-only 3D perception.",
      "short_abstract": "Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performance while overlooking the severe long-tail imbalance inherent in real-world...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 24,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection",
          "abstract: LiDAR",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.18476",
      "pdf_url": "https://arxiv.org/pdf/2604.18476.pdf"
    },
    {
      "id": "2604.18468",
      "anchor": "paper-2604-18468",
      "title": "Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation",
      "authors": [
        "Tianshi Cao",
        "Jiawei Ren",
        "Yuxuan Zhang",
        "Jaewoo Seo",
        "Jiahui Huang",
        "Shikhar Solanki",
        "Haotian Zhang",
        "Mingfei Guo",
        "Haithem Turki",
        "Muxingzi Li",
        "Yue Zhu",
        "Sipeng Zhang",
        "Zan Gojcic",
        "Sanja Fidler",
        "Kangxue Yin"
      ],
      "abstract": "Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural scene reconstruction converts driving logs into interactive 3D environments for simulation, but it does not produce complete 3D object assets required for agent manipulation and large-viewpoint novel-view synthesis. To address this challenge, we present Asset Harvester, an image-to-3D model and end-to-end pipeline that converts sparse, in-the-wild object observations from real driving logs into complete, simulation-ready assets. Rather than relying on a single model component, we developed a system-level design for real-world AV data that combines large-scale curation of object-centric training tuples, geometry-aware preprocessing across heterogeneous sensors, and a robust training recipe that couples sparse-view-conditioned multiview generation with 3D Gaussian lifting. Within this system, SparseViewDiT is explicitly designed to address limited-angle views and other real-world data challenges. Together with hybrid data curation, augmentation, and self-distillation, this system enables scalable conversion of sparse AV object observations into reusable 3D assets.",
      "short_abstract": "Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural scene reconstruction converts driving logs into interactive 3D environments for...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 6,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.18468",
      "pdf_url": "https://arxiv.org/pdf/2604.18468.pdf"
    },
    {
      "id": "2604.17915",
      "anchor": "paper-2604-17915",
      "title": "OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models",
      "authors": [
        "Yiwei Zhang",
        "Xuesong Chen",
        "Jin Gao",
        "Hanshi Wang",
        "Fudong Ge",
        "Weiming Hu",
        "Shaoshuai Shi",
        "Zhipeng Zhang"
      ],
      "abstract": "Vision-Language Models(VLMs) excel at autoregressive text generation, yet end-to-end autonomous driving requires multi-task learning with structured outputs and heterogeneous decoding behaviors, such as autoregressive language generation, parallel object detection and trajectory regression. To accommodate these differences, existing systems typically introduce separate or cascaded decoders, resulting in architectural fragmentation and limited backbone reuse. In this work, we present a unified autonomous driving framework built upon a pretrained VLM, where heterogeneous decoding behaviors are reconciled within a single transformer decoder. We demonstrate that pretrained VLM attention exhibits strong transferability beyond pure language modeling. By organizing visual and structured query tokens within a single causal decoder, structured queries can naturally condition on visual context through the original attention mechanism. Textual and structured outputs share a common attention backbone, enabling stable joint optimization across heterogeneous tasks. Trajectory planning is realized within the same causal LLM decoder by introducing structured trajectory queries. This unified formulation enables planning to share the pretrained attention backbone with images and perception tokens. Extensive experiments on end-to-end autonomous driving benchmarks demonstrate state-of-the-art performance, including 0.28 L2 and 0.18 collision rate on nuScenes open-loop evaluation and competitive results (86.8 PDMS) on NAVSIM closed-loop evaluation. The full model preserves multi-modal generation capability, while an efficient inference mode achieves approximately 40% lower latency. Code and models are available at https://github.com/Z1zyw/OneDrive",
      "short_abstract": "Vision-Language Models(VLMs) excel at autoregressive text generation, yet end-to-end autonomous driving requires multi-task learning with structured outputs and heterogeneous decoding behaviors, such as autoregressive language generation, parallel object...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 19,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.17915",
      "pdf_url": "https://arxiv.org/pdf/2604.17915.pdf"
    },
    {
      "id": "2604.17862",
      "anchor": "paper-2604-17862",
      "title": "M100: An Orchestrated Dataflow Architecture Powering General AI Computing",
      "authors": [
        "Yan Xie",
        "Changkui Mao",
        "Changsong Wu",
        "Chao Lu",
        "Chao Suo",
        "Cheng Qian",
        "Chun Yang",
        "Danyang Zhu",
        "Hengchang Xiong",
        "Hongzhan Lu",
        "Hongzhen Liu",
        "Jiafu Liu",
        "Jie Chen",
        "Jie Dai",
        "Junfeng Tang",
        "Kai Liu",
        "Kun Li",
        "Lipeng Ge",
        "Meng Sun",
        "Min Luo",
        "Peng Chen",
        "Peng Wang",
        "Shaodong Yang",
        "Shibin Tang",
        "Shibo Chen"
      ],
      "abstract": "As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility for diverse AI workloads, they often fall short in efficiency and cost-effectiveness. Various Domain-Specific Architectures (DSAs) excel at particular AI tasks but struggle to extend across broader applications or adapt to the rapidly evolving AI landscape. M100 is Li Auto's response: a performant, cost-effective architecture for AI inference in Autonomous Driving (AD), Large Language Models (LLMs), and intelligent human interactions, domains crucial to today's most competitive automobile platforms. M100 employs a dataflow parallel architecture, where compiler-architecture co-design orchestrates not only computation but, more critically, data movement across time and space. Leveraging dataflow computing efficiency, our hardware-software co-design improves system performance while reducing hardware complexity and cost. M100 largely eliminates caching: tensor computations are driven by compiler- and runtime-managed data streams flowing between computing elements and on/off-chip memories, yielding greater efficiency and scalability than cache-based systems. Another key principle was selecting the right operational granularity for scheduling, issuing, and execution across compiler, firmware, and hardware. Recognizing commonalities in AI workloads, we chose the tensor as the fundamental data element. M100 demonstrates general AI computing capability across diverse inference applications, including UniAD (for AD) and LLaMA (for LLMs). Benchmarks show M100 outperforms GPGPU architectures in AD applications with higher utilization, representing a promising direction for future general AI computing.",
      "short_abstract": "As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility for diverse AI workloads, they often fall short in efficiency and...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: hardware",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.17862",
      "pdf_url": "https://arxiv.org/pdf/2604.17862.pdf"
    },
    {
      "id": "2604.17841",
      "anchor": "paper-2604-17841",
      "title": "Driving risk emerges from the required two-dimensional joint evasive acceleration",
      "authors": [
        "Hao Cheng",
        "Yanbo Jiang",
        "Wenhao Yu",
        "Rui Zhou",
        "Jiang Bian",
        "Keyu Chen",
        "Zhiyuan Liu",
        "Heye Huang",
        "Hailun Zhang",
        "Fang Zhang",
        "Jianqiang Wang",
        "Sifa Zheng"
      ],
      "abstract": "Most autonomous driving safety benchmarks use time-to-collision (TTC) to assess risk and guide safe behaviour. However, TTC-based methods treat risk as a one-dimensional closing problem, despite the inherently two-dimensional nature of collision avoidance, and therefore cannot faithfully capture risk or its evolution over time. Here, we report evasive acceleration (EA), a hyperparameter-free and physically interpretable two-dimensional paradigm for risk quantification. By evaluating all possible directions of collision avoidance, EA defines risk as the minimum magnitude of a constant relative acceleration vector required to alter the relative motion and make the interaction collision-free. Using interaction data from five open datasets and more than 600 real crashes, we derive percentile-based warning thresholds and show that EA provides the earliest statistically significant warning across all thresholds. Moreover, EA provides the best discrimination of eventual collision outcomes and improves information retention by 54.2-241.4% over all compared baselines. Adding EA to existing methods yields 17.5-95.5 times more information gain than adding existing methods to EA, indicating that EA captures much of the outcome-relevant information in existing methods while contributing substantial additional nonredundant information. Overall, EA better captures the structure of collision risk and provides a foundation for next-generation autonomous driving systems.",
      "short_abstract": "Most autonomous driving safety benchmarks use time-to-collision (TTC) to assess risk and guide safe behaviour. However, TTC-based methods treat risk as a one-dimensional closing problem, despite the inherently two-dimensional nature of collision avoidance,...",
      "published": "2026-04-20",
      "updated": "2026-04-20",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: steering or braking",
          "abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 121,
      "arxiv_url": "https://arxiv.org/abs/2604.17841",
      "pdf_url": "https://arxiv.org/pdf/2604.17841.pdf"
    },
    {
      "id": "2604.22824",
      "anchor": "paper-2604-22824",
      "title": "WeatherSeg: Weather-Robust Image Segmentation using Teacher-Student Dual Learning and Classifier-Updating Attention",
      "authors": [
        "Zhang Zhang",
        "Yifeng Zeng",
        "Houshi Jiang",
        "Yinghui Pan"
      ],
      "abstract": "WeatherSeg, an advanced semi-supervised segmentation framework, addresses autonomous driving's environmental perception challenges in adverse weather while reducing annotation costs. This framework integrates a Dual Teacher-Student Weight-Sharing Model (DTSWSM) that enables knowledge distillation from weather-affected images, and a Classifier Weight Updating Attention Mechanism (CWUAM) that dynamically adjusts classifier weights based on environmental attributes. Comprehensive evaluations demonstrate that WeatherSeg significantly outperforms baseline models in both accuracy and robustness across various weather conditions, including clear, rainy, cloudy, and foggy scenarios, establishing it as an effective solution for all-weather semantic segmentation in autonomous driving and related applications.",
      "short_abstract": "WeatherSeg, an advanced semi-supervised segmentation framework, addresses autonomous driving's environmental perception challenges in adverse weather while reducing annotation costs. This framework integrates a Dual Teacher-Student Weight-Sharing Model...",
      "published": "2026-04-19",
      "updated": "2026-04-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 1,
        "evidence": [
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 122,
      "arxiv_url": "https://arxiv.org/abs/2604.22824",
      "pdf_url": "https://arxiv.org/pdf/2604.22824.pdf"
    },
    {
      "id": "2604.17651",
      "anchor": "paper-2604-17651",
      "title": "Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception",
      "authors": [
        "Siyuan Meng",
        "Chengbo Ai"
      ],
      "abstract": "World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego-vehicle perspective, leaving the infrastructure viewpoint unexplored. We argue that infrastructure-centric world models offer a fundamentally complementary capability: the bird's-eye, multi-sensor, persistent viewpoint that roadside systems uniquely possess. Central to our thesis is a spatio-temporal complementarity: fixed roadside sensors excel at temporal depth, accumulating long-term behavioral distributions including rare safety-critical events, while vehicle-borne sensors excel at spatial breadth, sampling diverse scenes across large road networks. This paper presents a vision for Infrastructure-centric World Models (I-WM) in three phases: (I) generative scene understanding with quality-aware uncertainty propagation, (II) physics-informed predictive dynamics with multi-agent counterfactual reasoning, and (III) collaborative world models for V2X communication via latent space alignment. We propose a dual-layer architecture, annotation-free perception as a multi-modal data engine feeding end-to-end generative world models, with a phased sensor strategy from LiDAR through 4D radar and signal phase data to event cameras. We establish a taxonomy of driving world model paradigms, position I-WM relative to LeCun's JEPA, Li Fei-Fei's spatial intelligence, and VLA architectures, and introduce Infrastructure VLA (I-VLA) as a novel unification of roadside perception, language commands, and traffic control actions. Our vision builds upon existing multi-LiDAR pipelines and identifies open-source foundations for each phase, providing a path toward infrastructure that understands and anticipates traffic.",
      "short_abstract": "World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego-vehicle perspective, leaving the infrastructure viewpoint unexplored. We argue that infrastructure-centric...",
      "published": "2026-04-19",
      "updated": "2026-04-19",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X",
        "VLA",
        "World Model"
      ],
      "classification": {
        "confidence": "low",
        "score": 23,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: title: perception",
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Perception & Sensor Fusion: abstract: radar",
          "Perception & Sensor Fusion: abstract: camera",
          "Perception & Sensor Fusion: abstract: scene understanding",
          "Systems, Deployment & Connectivity: abstract: V2X",
          "Systems, Deployment & Connectivity: abstract: cooperative systems"
        ]
      },
      "recency": "archive",
      "age_days": 122,
      "arxiv_url": "https://arxiv.org/abs/2604.17651",
      "pdf_url": "https://arxiv.org/pdf/2604.17651.pdf"
    },
    {
      "id": "2604.17391",
      "anchor": "paper-2604-17391",
      "title": "RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted Certification",
      "authors": [
        "Nick Andreasyan",
        "Mikhail Struve",
        "Alexey Popov",
        "Maksim Nikolaev",
        "Vadim Vashkelis"
      ],
      "abstract": "RISC-V is emerging as a viable platform for automotive-grade embedded computing, with recent ISO 26262 ASIL-D certifications demonstrating readiness for safety-critical deployment in autonomous driving systems. However, functional safety in automotive systems is fundamentally a certification problem rather than a processor problem. The dominant costs arise from diagnostic coverage analysis, toolchain qualification, fault injection campaigns, safety-case generation, and compliance with ISO 26262, ISO 21448 (SOTIF), and ISO/SAE 21434. This paper analyzes the role of RISC-V in automotive functional safety, focusing on ISA openness, formal verifiability, custom extension control, debug transparency, and vendor-independent qualification. We examine autonomous driving safety requirements and map them to RISC-V architectural challenges such as lockstep execution, safety islands, mixed-criticality isolation, and secure debug. Rather than proposing a single algorithmic breakthrough, we present an analytical framework and research roadmap centered on certification economics as the primary optimization objective. We also discuss how selected ML methods, including LLM-assisted FMEDA generation, knowledge-graph-based safety case automation, reinforcement learning for fault injection, and graph neural networks for diagnostic coverage, can support certification workflows. We argue that the strongest outcome is not a faster core, but an ASIL-D-ready certifiable RISC-V platform.",
      "short_abstract": "RISC-V is emerging as a viable platform for automotive-grade embedded computing, with recent ISO 26262 ASIL-D certifications demonstrating readiness for safety-critical deployment in autonomous driving systems. However, functional safety in automotive systems...",
      "published": "2026-04-19",
      "updated": "2026-04-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 17,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: SOTIF"
        ]
      },
      "recency": "archive",
      "age_days": 122,
      "arxiv_url": "https://arxiv.org/abs/2604.17391",
      "pdf_url": "https://arxiv.org/pdf/2604.17391.pdf"
    },
    {
      "id": "2604.17281",
      "anchor": "paper-2604-17281",
      "title": "Safety-Aware AoI Scheduling for LEO Satellite-Assisted Autonomous Driving",
      "authors": [
        "Kangkang Sun",
        "Junyi He",
        "Juntong Liu",
        "Xiuzhen Chen",
        "Jianhua Li",
        "Minyi Guo"
      ],
      "abstract": "Autonomous platoons traversing infrastructure gaps increasingly depend on LEO satellite backhaul for safety-critical updates, yet no existing framework jointly addresses compound Doppler from simultaneous satellite and vehicle motion, sub-slot handover outages that exceed collision-alert deadlines, and heterogeneous freshness requirements across three vehicular priority classes. The core challenge is a \\emph{timescale mismatch}: coarse control slots hide sub-slot outages, which makes both AoI spike analysis and safety verification ill-posed. Ping-pong handover oscillations further compound AoI cost in a way that purely reactive schedulers cannot mitigate. We address these challenges through a unified framework that couples a two-timescale AoI model with tiered time-average safety constraints enforced by virtual queues. A closed-form ping-pong AoI envelope reveals that cumulative penalty grows quadratically in oscillation length, analytically justifying oscillation suppression as the highest-leverage safety mechanism. The resulting drift-plus-penalty template is instantiated as SafeScale-MATD3 with proactive handover timing and multi-task dual-critic MARL. A key finding is that suppressing brief but repeated ping-pong oscillations yields larger safety returns than shortening any single outage, and that tick-level AoI accounting is a necessary condition for verifiable collision-alert guarantees under LEO handovers. Simulations show that SafeScale-MATD3 is the only method satisfying the strict 1 % collision-alert violation budget, reducing violation rate by 4 to 5.5 times versus baselines, while achieving 35 % lower collision-alert AoI and strict Pareto dominance on the energy and freshness tradeoff.",
      "short_abstract": "Autonomous platoons traversing infrastructure gaps increasingly depend on LEO satellite backhaul for safety-critical updates, yet no existing framework jointly addresses compound Doppler from simultaneous satellite and vehicle motion, sub-slot handover...",
      "published": "2026-04-19",
      "updated": "2026-04-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 15,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: verification"
        ]
      },
      "recency": "archive",
      "age_days": 122,
      "arxiv_url": "https://arxiv.org/abs/2604.17281",
      "pdf_url": "https://arxiv.org/pdf/2604.17281.pdf"
    },
    {
      "id": "2604.17135",
      "anchor": "paper-2604-17135",
      "title": "OptiMVMap: Offline Vectorized Map Construction via Optimal Multi-vehicle Perspectives",
      "authors": [
        "Zedong Dan",
        "Zijie Wang",
        "Wei Zhang",
        "Xiangru Lin",
        "Weiming Zhang",
        "Xiao Tan",
        "Jingdong Wang",
        "Liang Lin",
        "Guanbin Li"
      ],
      "abstract": "Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego-vehicle trajectories, which fundamentally suffer from viewpoint insufficiency: while memory-based methods extend observation time by aggregating ego-trajectory frames, they lack the spatial diversity needed to reveal occluded regions. Incorporating views from surrounding vehicles offers complementary perspectives, yet naive fusion introduces three key challenges: computational cost from large candidate pools, redundancy from near-collinear viewpoints, and noise from pose errors and occlusion artifacts. We present OptiMVMap, which reformulates multi-vehicle mapping as a select-then-fuse problem to address these challenges systematically. An Optimal Vehicle Selection (OVS) module strategically identifies a compact subset of helpers that maximally reduce ego-centric uncertainty in occluded regions, addressing computation and redundancy challenges. Cross-Vehicle Attention (CVA) and Semantic-aware Noise Filter (SNF) then perform pose-tolerant alignment and artifact suppression before BEV-level fusion, addressing the noise challenge. This targeted pipeline yields more complete and topologically faithful maps with substantially fewer views than indiscriminate aggregation. On nuScenes and Argoverse2, OptiMVMap improves MapTRv2 by +10.5 mAP and +9.3 mAP, respectively, and surpasses memory-augmented baselines MVMap and HRMapNet by +6.2 mAP and +3.8 mAP on nuScenes. These results demonstrate that uncertainty-guided selection of helper vehicles is essential for efficient and accurate multi-vehicle vectorized mapping. The code is released at https://github.com/DanZeDong/OptiMVMap.",
      "short_abstract": "Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego-vehicle trajectories, which fundamentally suffer from viewpoint insufficiency: while...",
      "published": "2026-04-18",
      "updated": "2026-04-18",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 14,
        "evidence": [
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 123,
      "arxiv_url": "https://arxiv.org/abs/2604.17135",
      "pdf_url": "https://arxiv.org/pdf/2604.17135.pdf"
    },
    {
      "id": "2604.16582",
      "anchor": "paper-2604-16582",
      "title": "Camo-M3FD: A New Benchmark Dataset for Cross-Spectral Camouflaged Pedestrian Detection",
      "authors": [
        "Henry O. Velesaca",
        "Andrea Mero",
        "Guillermo A. Castillo",
        "Angel D. Sappa"
      ],
      "abstract": "Pedestrian detection is fundamental to autonomous driving, robotics, and surveillance. Despite progress in deep learning, reliable identification remains challenging due to occlusions, cluttered backgrounds, and degraded visibility. While multispectral detection-combining visible and thermal sensors-mitigates poor visibility, the challenge of camouflaged pedestrians remains largely unexplored. Existing Camouflaged Object Detection (COD) benchmarks focus on biological species, leaving a gap in safety-critical human detection where targets blend into their surroundings. To address this, we introduce Camo-M3FD (derived from the M3FD dataset), a novel benchmark for cross-spectral camouflaged pedestrian detection, consisting of registered visible-thermal image pairs. The dataset is curated using quantitative metrics to ensure high foreground-background similarity. We provide high-quality pixel-level masks and establish a standardized evaluation framework using state-of-the-art COD models. Our results demonstrate that while thermal signals provide indispensable localization cues, multispectral fusion is essential for refining structural details. Camo-M3FD serves as a foundational resource for developing robust and safety-critical detection systems. The dataset is available on GitHub: https://cod-espol.github.io/Camo-M3FD/",
      "short_abstract": "Pedestrian detection is fundamental to autonomous driving, robotics, and surveillance. Despite progress in deep learning, reliable identification remains challenging due to occlusions, cluttered backgrounds, and degraded visibility. While multispectral...",
      "published": "2026-04-17",
      "updated": "2026-04-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 124,
      "arxiv_url": "https://arxiv.org/abs/2604.16582",
      "pdf_url": "https://arxiv.org/pdf/2604.16582.pdf"
    },
    {
      "id": "2604.16184",
      "anchor": "paper-2604-16184",
      "title": "Real-Time Solution-Seeking for Game-Theoretic Autonomous Driving via Time-Distributed Iterations",
      "authors": [
        "Shaoqing Liu",
        "Mushuang Liu"
      ],
      "abstract": "Computational complexity has been a major challenge in game-theoretic model predictive control (GT-MPC), as real-time solutions to a game (e.g., Nash equilibria (NEs)) have to be computed at each sampling instant of an MPC. This challenge is especially critical in autonomous driving, where interactions may involve many agents, and decisions must be made at fast sampling rates. We show that this challenge can be addressed through time-distributed solution-seeking iterations designed based on, e.g., Newton and Newton--Kantorovich methods. Specifically, the autonomous vehicle decision-making problem is first formulated as a GT-MPC problem. To ensure solution attainability, a potential game framework is adopted. Within this framework, both potential-function optimization and best-response dynamics are used to seek the NE. To enable real-time implementation, Newton and Newton--Kantorovich methods are employed to solve the optimization problems arising in the NE-seeking algorithms, with their iterations distributed over time. Numerical experiments on an intersection-crossing scenario demonstrate that the proposed methods achieve effective real-time performance.",
      "short_abstract": "Computational complexity has been a major challenge in game-theoretic model predictive control (GT-MPC), as real-time solutions to a game (e.g., Nash equilibria (NEs)) have to be computed at each sampling instant of an MPC. This challenge is especially...",
      "published": "2026-04-17",
      "updated": "2026-04-17",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 5,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "archive",
      "age_days": 124,
      "arxiv_url": "https://arxiv.org/abs/2604.16184",
      "pdf_url": "https://arxiv.org/pdf/2604.16184.pdf"
    },
    {
      "id": "2604.16086",
      "anchor": "paper-2604-16086",
      "title": "Stylistic-STORM (ST-STORM) : Perceiving the Semantic Nature of Appearance",
      "authors": [
        "Hamed Ouattara",
        "Pierre Duthon",
        "Pascal Houssam Salmane",
        "Frédéric Bernardin",
        "Omar Ait Aider"
      ],
      "abstract": "One of the dominant paradigms in self-supervised learning (SSL), illustrated by MoCo or DINO, aims to produce robust representations by capturing features that are insensitive to certain image transformations such as illumination, or geometric changes. This strategy is appropriate when the objective is to recognize objects independently of their appearance. However, it becomes counterproductive as soon as appearance itself constitutes the discriminative signal. In weather analysis, for example, rain streaks, snow granularity, atmospheric scattering, as well as reflections and halos, are not noise: they carry the essential information. In critical applications such as autonomous driving, ignoring these cues is risky, since grip and visibility depend directly on ground conditions and atmospheric conditions. We introduce ST-STORM, a hybrid SSL framework that treats appearance (style) as a semantic modality to be disentangled from content. Our architecture explicitly separates two latent streams, regulated by gating mechanisms. The Content branch aims at a stable semantic representation through a JEPA scheme coupled with a contrastive objective, promoting invariance to appearance variations. In parallel, the Style branch is constrained to capture appearance signatures (textures, contrasts, scattering) through feature prediction and reconstruction under an adversarial constraint. We evaluate ST-STORM on several tasks, including object classification (ImageNet-1K), fine-grained weather characterization, and melanoma detection (ISIC 2024 Challenge). The results show that the Style branch effectively isolates complex appearance phenomena (F1=97% on Multi-Weather and F1=94% on ISIC 2024 with 10% labeled data), without degrading the semantic performance (F1=80% on ImageNet-1K) of the Content branch, and improves the preservation of critical appearance",
      "short_abstract": "One of the dominant paradigms in self-supervised learning (SSL), illustrated by MoCo or DINO, aims to produce robust representations by capturing features that are insensitive to certain image transformations such as illumination, or geometric changes. This...",
      "published": "2026-04-17",
      "updated": "2026-04-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 4,
        "evidence": [
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 124,
      "arxiv_url": "https://arxiv.org/abs/2604.16086",
      "pdf_url": "https://arxiv.org/pdf/2604.16086.pdf"
    },
    {
      "id": "2604.15795",
      "anchor": "paper-2604-15795",
      "title": "Fed3D: Federated 3D Object Detection",
      "authors": [
        "Suyan Dai",
        "Chenxi Liu",
        "Fazeng Li",
        "Peican Lin"
      ],
      "abstract": "3D object detection models trained in one server plays an important role in autonomous driving, robotics manipulation, and augmented reality scenarios. However, most existing methods face severe privacy concern when deployed on a multi-robot perception network to explore large-scale 3D scene. Meanwhile, it is highly challenging to employ conventional federated learning methods on 3D object detection scenes, due to the 3D data heterogeneity and limited communication bandwidth. In this paper, we take the first attempt to propose a novel Federated 3D object detection framework (i.e., Fed3D), to enable distributed learning for 3D object detection with privacy preservation. Specifically, considering the irregular input 3D object in local robot and various category distribution between robots could cause local heterogeneity and global heterogeneity, respectively. We then propose a local-global class-aware loss for the 3D data heterogeneity issue, which could balance gradient back-propagation rate of different 3D categories from local and global aspects. To reduce communication cost on each round, we develop a federated 3D prompt module, which could only learn and communicate the prompts with few learnable parameters. To the end, several extensive experiments on federated 3D object detection show that our Fed3D model significantly outperforms state-of-the-art algorithms with lower communication cost when providing the limited local training data.",
      "short_abstract": "3D object detection models trained in one server plays an important role in autonomous driving, robotics manipulation, and augmented reality scenarios. However, most existing methods face severe privacy concern when deployed on a multi-robot perception...",
      "published": "2026-04-17",
      "updated": "2026-04-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 15,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 124,
      "arxiv_url": "https://arxiv.org/abs/2604.15795",
      "pdf_url": "https://arxiv.org/pdf/2604.15795.pdf"
    },
    {
      "id": "2604.15705",
      "anchor": "paper-2604-15705",
      "title": "Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning",
      "authors": [
        "Xiaoyu Yang",
        "En Yu",
        "Wei Duan",
        "Jie Lu"
      ],
      "abstract": "Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with complex human values and domain-specific requirements. Nevertheless, current research primarily focuses on mitigating exogenous distribution shifts arising from data-centric factors, the non-stationarity inherent in the endogenous reasoning remains largely unexplored. In this work, a critical vulnerability is revealed within MLLMs: they are highly susceptible to endogenous reasoning drift, across both thinking and perception perspectives. It manifests as unpredictable distribution changes that emerge spontaneously during the autoregressive generation process, independent of external environmental perturbations. To adapt it, we first theoretically define endogenous reasoning drift within the RFT of MLLMs as the multi-modal concept drift. In this context, this paper proposes Counterfactual Preference Optimization ++ (CPO++), a comprehensive and autonomous framework adapted to the multi-modal concept drift. It integrates counterfactual reasoning with domain knowledge to execute controlled perturbations across thinking and perception, employing preference optimization to disentangle spurious correlations. Extensive empirical evaluations across two highly dynamic and safety-critical domains: medical diagnosis and autonomous driving. They demonstrate that the proposed framework achieves superior performance in reasoning coherence, decision-making precision, and inherent robustness against extreme interference. The methodology also exhibits exceptional zero-shot cross-domain generalization, providing a principled foundation for reliable multi-modal reasoning in safety-critical applications.",
      "short_abstract": "Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with complex human values and domain-specific requirements. Nevertheless, current research primarily focuses on...",
      "published": "2026-04-17",
      "updated": "2026-04-17",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "abstract: safety",
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 124,
      "arxiv_url": "https://arxiv.org/abs/2604.15705",
      "pdf_url": "https://arxiv.org/pdf/2604.15705.pdf"
    },
    {
      "id": "2604.16702",
      "anchor": "paper-2604-16702",
      "title": "Autonomous Vehicle Collision Avoidance With Racing Parameterized Deep Reinforcement Learning",
      "authors": [
        "Shathushan Sivashangaran",
        "Vihaan Dutta",
        "Apoorva Khairnar",
        "Sepideh Gohari",
        "Azim Eskandarian"
      ],
      "abstract": "Road traffic accidents are a leading cause of fatalities worldwide. In the US, human error causes 94% of crashes, resulting in excess of 7,000 pedestrian fatalities and $500 billion in costs annually. Autonomous Vehicles (AVs) with emergency collision avoidance systems that operate at the limits of vehicle dynamics at a high frequency, a dual constraint of nonlinear kinodynamic accuracy and computational efficiency, further enhance safety benefits during adverse weather and cybersecurity breaches, and to evade dangerous human driving when AVs and human drivers share roads. This paper parameterizes a Deep Reinforcement Learning (DRL) collision avoidance policy Out-Of-Distribution (OOD) utilizing race car overtaking, without explicit geometric mimicry reference trajectory guidance, in simulation, with a physics-informed, simulator exploit-aware reward to encode nonlinear vehicle kinodynamics. Two policies are evaluated, a default uni-direction and a reversed heading variant that navigates in the opposite direction to other cars, which both consistently outperform a Model Predictive Control and Artificial Potential Function (MPC-APF) baseline, with zero-shot transfer to proportionally scaled hardware, across three intersection collision scenarios, at 31x fewer Floating Point Operations (FLOPS) and 64x lower inference latency. The reversed heading policy outperforms the default racing overtaking policy in head-to-head collisions by 30% and the baseline by 50%, and matches the former in side collisions, where both DRL policies evade 10% greater than numerical optimal control.",
      "short_abstract": "Road traffic accidents are a leading cause of fatalities worldwide. In the US, human error causes 94% of crashes, resulting in excess of 7,000 pedestrian fatalities and $500 billion in costs annually. Autonomous Vehicles (AVs) with emergency collision...",
      "published": "2026-04-17",
      "updated": "2026-04-17",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 3,
        "evidence": [
          "abstract: model predictive control",
          "abstract: vehicle dynamics",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 124,
      "arxiv_url": "https://arxiv.org/abs/2604.16702",
      "pdf_url": "https://arxiv.org/pdf/2604.16702.pdf"
    },
    {
      "id": "2604.15308",
      "anchor": "paper-2604-15308",
      "title": "RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework",
      "authors": [
        "Hao Gao",
        "Shaoyu Chen",
        "Yifan Zhu",
        "Yuehao Song",
        "Wenyu Liu",
        "Qian Zhang",
        "Xinggang Wang"
      ],
      "abstract": "High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interactions. Although diffusion-based planners are effective at modeling complex trajectory distributions, they often suffer from stochastic instabilities and the lack of corrective negative feedback when trained purely with imitation learning. To address these issues, we propose RAD-2, a unified generator-discriminator framework for closed-loop planning. Specifically, a diffusion-based generator is used to produce diverse trajectory candidates, while an RL-optimized discriminator reranks these candidates according to their long-term driving quality. This decoupled design avoids directly applying sparse scalar rewards to the full high-dimensional trajectory space, thereby improving optimization stability. To further enhance reinforcement learning, we introduce Temporally Consistent Group Relative Policy Optimization, which exploits temporal coherence to alleviate the credit assignment problem. In addition, we propose On-policy Generator Optimization, which converts closed-loop feedback into structured longitudinal optimization signals and progressively shifts the generator toward high-reward trajectory manifolds. To support efficient large-scale training, we introduce BEV-Warp, a high-throughput simulation environment that performs closed-loop evaluation directly in Bird's-Eye View feature space via spatial warping. RAD-2 reduces the collision rate by 56% compared with strong diffusion-based planners. Real-world deployment further demonstrates improved perceived safety and driving smoothness in complex urban traffic.",
      "short_abstract": "High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interactions. Although diffusion-based planners are effective at modeling complex trajectory distributions, they...",
      "published": "2026-04-16",
      "updated": "2026-04-16",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: robustness",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 125,
      "arxiv_url": "https://arxiv.org/abs/2604.15308",
      "pdf_url": "https://arxiv.org/pdf/2604.15308.pdf"
    },
    {
      "id": "2604.15291",
      "anchor": "paper-2604-15291",
      "title": "AD4AD: Benchmarking Visual Anomaly Detection Models for Safer Autonomous Driving",
      "authors": [
        "Fabrizio Genilotti",
        "Arianna Stropeni",
        "Gionata Grotto",
        "Francesco Borsatti",
        "Manuel Barusco",
        "Davide Dalle Pezze",
        "Gian Antonio Susto"
      ],
      "abstract": "The reliability of a machine vision system for autonomous driving depends heavily on its training data distribution. When a vehicle encounters significantly different conditions, such as atypical obstacles, its perceptual capabilities can degrade substantially. Unlike many domains where errors carry limited consequences, failures in autonomous driving translate directly into physical risk for passengers, pedestrians, and other road users. To address this challenge, we explore Visual Anomaly Detection (VAD) as a solution. VAD enables the identification of anomalous objects not present during training, allowing the system to alert the driver when an unfamiliar situation is detected. Crucially, VAD models produce pixel-level anomaly maps that can guide driver attention to specific regions of concern without requiring any prior assumptions about the nature or form of the hazard. We benchmark eight state-of-the-art VAD methods on AnoVox, the largest synthetic dataset for anomaly detection in autonomous driving. In particular, we evaluate performance across four backbone architectures spanning from large networks to lightweight ones such as MobileNet and DeiT-Tiny. Our results demonstrate that VAD transfers effectively to road scenes. Notably, Tiny-Dinomaly achieves the best accuracy-efficiency trade-off for edge deployment, matching full-scale localization performance at a fraction of the memory cost. This study represents a concrete step toward safer, more responsible deployment of autonomous vehicles, ultimately improving protection for passengers, pedestrians, and all road users.",
      "short_abstract": "The reliability of a machine vision system for autonomous driving depends heavily on its training data distribution. When a vehicle encounters significantly different conditions, such as atypical obstacles, its perceptual capabilities can degrade...",
      "published": "2026-04-16",
      "updated": "2026-04-16",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 0,
        "evidence": [
          "Mapping & Localization: abstract: localization",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 125,
      "arxiv_url": "https://arxiv.org/abs/2604.15291",
      "pdf_url": "https://arxiv.org/pdf/2604.15291.pdf"
    },
    {
      "id": "2604.13853",
      "anchor": "paper-2604-13853",
      "title": "Mosaic: An Extensible Framework for Composing Rule-Based and Learned Motion Planners",
      "authors": [
        "Nick Le Large",
        "Marlon Steiner",
        "Lingguang Wang",
        "Willi Poh",
        "Jan-Hendrik Pauls",
        "Ömer Şahin Taş",
        "Christoph Stiller"
      ],
      "abstract": "Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable behavior, they often fail to grasp the complexity and uncertainty of real-world traffic. Conversely, learned planners exhibit strong adaptability but suffer from reduced transparency and occasional safety violations. We introduce Mosaic, a framework for structured decision-making that integrates both paradigms through arbitration graphs. By decoupling trajectory verification and selection from the generation of trajectories by individual planners, every decision becomes transparent and traceable. This separation lets verification and trajectory selection contribute independently: centralized verification acts as a safety floor, reducing at-fault collisions from 25 for each standalone planner to 16. In contrast, per-step trajectory selection acts as a performance ceiling, combining the complementary strengths of a rule-based and a learned planner. In experimental evaluation on nuPlan, Mosaic achieves 95.56 CLS-NR and 94.18 CLS-R on the Val14 closed-loop benchmark, setting a new state of the art. On the interPlan benchmark, focused on highly interactive and out-of-distribution scenarios, Mosaic scores 54.10 CLS-R, outperforming its best constituent planner by 22.8% -- all without retraining or requiring additional data. The code is available at github.com/KIT-MRT/mosaic.",
      "short_abstract": "Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable behavior, they often fail to grasp the complexity and uncertainty of real-world traffic. Conversely, learned...",
      "published": "2026-04-15",
      "updated": "2026-04-15",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 1,
        "evidence": [
          "abstract: motion planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 126,
      "arxiv_url": "https://arxiv.org/abs/2604.13853",
      "pdf_url": "https://arxiv.org/pdf/2604.13853.pdf"
    },
    {
      "id": "2604.13691",
      "anchor": "paper-2604-13691",
      "title": "Towards Autonomous Driving with Short-Packet Rate Splitting: Age of Information Analysis and Optimization",
      "authors": [
        "Zirui Zheng",
        "Yingyang Chen",
        "Xinyue Pei",
        "Xingwei Wang",
        "Zhiquan Liu",
        "Theodoros A. Tsiftsis",
        "Miaowen Wen",
        "Pingzhi Fan"
      ],
      "abstract": "To address the high mobility impacts and the ultra-reliable and low-latency communication (URLLC) requirements in autonomous driving scenarios, rate-splitting multiple access (RSMA) combined with short-packet communication (SPC) emerges as a promising solution.Autonomous vehicles rely on real-time information exchange to ensure safety and coordination, making information freshness essential.By jointly capturing transmission delays and packet errors, age of information (AoI) serves as a comprehensive metric for freshness.In this paper, we investigate short-packet rate splitting to enhance information freshness measured by the AoI.By splitting the unicast messages into common and private parts, encoding all common parts together with the multicast message into a common stream, and encoding each private part into a private stream, RSMA effectively manages interference and enables achieving lower AoI.By considering critical factors such as transmit power, vehicle velocity, blocklength, and the number of transmit antennas, we derive closed-form expressions for the average AoI (AAoI) of the common stream under partial decoding and the overall AAoI under complete decoding.To enhance the AAoI performance, we propose the multi-start two-step successive convex approximation (SCA) algorithm.This algorithm first optimizes the power allocation and subsequently optimizes the rate splitting under the quality of service (QoS) trade-off constraint.Simulation results demonstrate that our short-packet rate-splitting scheme significantly improves the AAoI performance while ensuring system fairness and enabling ultra-low AAoI through the common stream, meeting the requirements of autonomous driving applications.Moreover, the trade-off between the common and overall performance is revealed, indicating that the overall performance can be further enhanced while maintaining the advantages of the common stream.",
      "short_abstract": "To address the high mobility impacts and the ultra-reliable and low-latency communication (URLLC) requirements in autonomous driving scenarios, rate-splitting multiple access (RSMA) combined with short-packet communication (SPC) emerges as a promising...",
      "published": "2026-04-15",
      "updated": "2026-04-15",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 3,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 126,
      "arxiv_url": "https://arxiv.org/abs/2604.13691",
      "pdf_url": "https://arxiv.org/pdf/2604.13691.pdf"
    },
    {
      "id": "2604.13587",
      "anchor": "paper-2604-13587",
      "title": "Hybrid Architecture Gets Fluid: A New Paradigm for Direction-of-arrival Estimation in 6G Networks",
      "authors": [
        "Ye Tian",
        "Jiaji Ren",
        "Tuo Wu",
        "Wei Liu",
        "Maged Elkashlan",
        "Matthew C. Valenti",
        "Naofal Al-Dhahir",
        "Hing Cheung So"
      ],
      "abstract": "High-precision direction-of-arrival (DOA) estimation, as a key sensing capability for 6G-enabled applications such as autonomous driving and extended reality, is increasingly dependent on the effective exploitation of spatial degrees of freedom (DOFs). This paper integrates two frontier DOFs-oriented paradigms and proposes a fluid antenna-enabled hybrid analog-digital (FA-HAD) architecture, which features an extremely lightweight front-end configuration mechanism and efficient spatial DOFs exploitation. Within this architecture, a collaborative spatial-phase sampling strategy is first developed to enable real-time 2-D DOA estimation under compressive observations, and a single-source CRLB analysis is provided to quantify the achievable performance limit, offering quantitative guidance for accuracy-overhead trade-offs. Furthermore, an efficient virtual-array spatial covariance matrix reconstruction method is proposed to recover a physically meaningful covariance representation, thereby providing a covariance-domain interface that is directly reusable by a broad class of existing covariance-based array processing and array design techniques, which strengthens the scalability and transferability of the proposed architecture. Building upon the reconstructed SCM, a Jacobi-Anger expansion based dimension-reduced MUSIC estimator is further derived for arbitrary planar arrays with a favorable computational cost. Simulation results demonstrate that the proposed FA-HAD framework attains DOA accuracy close to fully digital systems while substantially reducing RF hardware complexity and training overhead.",
      "short_abstract": "High-precision direction-of-arrival (DOA) estimation, as a key sensing capability for 6G-enabled applications such as autonomous driving and extended reality, is increasingly dependent on the effective exploitation of spatial degrees of freedom (DOFs). This...",
      "published": "2026-04-15",
      "updated": "2026-04-15",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 15,
        "evidence": [
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: hardware",
          "title: communications"
        ]
      },
      "recency": "archive",
      "age_days": 126,
      "arxiv_url": "https://arxiv.org/abs/2604.13587",
      "pdf_url": "https://arxiv.org/pdf/2604.13587.pdf"
    },
    {
      "id": "2604.14454",
      "anchor": "paper-2604-14454",
      "title": "CooperDrive: Enhancing Driving Decisions Through Cooperative Perception",
      "authors": [
        "Deyuan Qu",
        "Qi Chen",
        "Takayuki Shimizu",
        "Onur Altintas"
      ],
      "abstract": "Autonomous vehicles equipped with robust onboard perception, localization, and planning still face limitations in occlusion and non-line-of-sight (NLOS) scenarios, where delayed reactions can increase collision risk. We propose CooperDrive, a cooperative perception framework that augments situational awareness and enables earlier, safer driving decisions. CooperDrive offers two key advantages: (i) each vehicle retains its native perception, localization, and planning stack, and (ii) a lightweight object-level sharing and fusion strategy bridges perception and planning. Specifically, CooperDrive reuses detector Bird's-Eye View (BEV) features to estimate accurate vehicle poses without additional heavy encoders, thereby reconstructing BEV representations and feeding the planner with low latency. On the planning side, CooperDrive leverages the expanded object set to anticipate potential conflicts earlier and adjust speed and trajectory proactively, thereby transforming reactive behaviors into predictive and safer driving decisions. Real-world closed-loop tests at occlusion-heavy NLOS intersections demonstrate that CooperDrive increases reaction lead time, minimum time-to-collision (TTC), and stopping margin, while requiring only 90 kbps bandwidth and maintaining an average end-to-end latency of 89 ms.",
      "short_abstract": "Autonomous vehicles equipped with robust onboard perception, localization, and planning still face limitations in occlusion and non-line-of-sight (NLOS) scenarios, where delayed reactions can increase collision risk. We propose CooperDrive, a cooperative...",
      "published": "2026-04-15",
      "updated": "2026-04-15",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 19,
        "margin": 5,
        "evidence": [
          "abstract: planning",
          "title: decision-making",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 126,
      "arxiv_url": "https://arxiv.org/abs/2604.14454",
      "pdf_url": "https://arxiv.org/pdf/2604.14454.pdf"
    },
    {
      "id": "2604.13891",
      "anchor": "paper-2604-13891",
      "title": "Beyond Conservative Automated Driving in Multi-Agent Scenarios via Coupled Model Predictive Control and Deep Reinforcement Learning",
      "authors": [
        "Saeed Rahmani",
        "Gözde Körpe",
        "Zhenlin",
        "Xu",
        "Bruno Brito",
        "Simeon Craig Calvert",
        "Bart van Arem"
      ],
      "abstract": "Automated driving at unsignalized intersections is challenging due to complex multi-vehicle interactions and the need to balance safety and efficiency. Model Predictive Control (MPC) offers structured constraint handling through optimization but relies on hand-crafted rules that often produce overly conservative behavior. Deep Reinforcement Learning (RL) learns adaptive behaviors from experience but often struggles with safety assurance and generalization to unseen environments. In this study, we present an integrated MPC-RL framework to improve navigation performance in multi-agent scenarios. Experiments show that MPC-RL outperforms standalone MPC and end-to-end RL across three traffic-density levels. Collectively, MPC-RL reduces the collision rate by 21% and improves the success rate by 6.5% compared to pure MPC. We further evaluate zero-shot transfer to a highway merging scenario without retraining. Both MPC-based methods transfer substantially better than end-to-end PPO, which highlights the role of the MPC backbone in cross-scenario robustness. The framework also shows faster loss stabilization than end-to-end RL during training, which indicates a reduced learning burden. These results suggest that the integrated approach can improve the balance between safety performance and efficiency in multi-agent intersection scenarios, while the MPC component provides a strong foundation for generalization across driving environments. The implementation code is available open-source.",
      "short_abstract": "Automated driving at unsignalized intersections is challenging due to complex multi-vehicle interactions and the need to balance safety and efficiency. Model Predictive Control (MPC) offers structured constraint handling through optimization but relies on...",
      "published": "2026-04-15",
      "updated": "2026-04-15",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 22,
        "evidence": [
          "title: model predictive control",
          "abstract: model predictive control",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 126,
      "arxiv_url": "https://arxiv.org/abs/2604.13891",
      "pdf_url": "https://arxiv.org/pdf/2604.13891.pdf"
    },
    {
      "id": "2604.12918",
      "anchor": "paper-2604-12918",
      "title": "Radar-Camera BEV Multi-Task Learning with Cross-Task Attention Bridge for Joint 3D Detection and Segmentation",
      "authors": [
        "Ahmet İnanç",
        "Özgür Erkent"
      ],
      "abstract": "Bird's-eye-view (BEV) representations are the dominant paradigm for 3D perception in autonomous driving, providing a unified spatial canvas where detection and segmentation features are geometrically registered to the same physical coordinate system. However, existing radar-camera fusion methods treat these tasks in isolation, missing the opportunity for cross-task feature sharing: object-level geometric cues from detection can sharpen segmentation, while dense road-layout context from segmentation can anchor detection. We propose \\textbf{CTAB} (Cross-Task Attention Bridge), a bidirectional module that exchanges features between detection and segmentation branches via multi-scale deformable attention in shared BEV space. CTAB is integrated into a multi-task framework with an Instance Normalization-based segmentation decoder and learnable BEV upsampling to provide a more detailed BEV representation. On nuScenes, CTAB improves segmentation on 7 classes over the joint multi-task baseline at essentially neutral detection. On a 4-class subset (drivable area, pedestrian crossing, walkway, vehicle), our joint multi-task model achieves 51.0 mIoU-4 while simultaneously providing competitive 3D detection.",
      "short_abstract": "Bird's-eye-view (BEV) representations are the dominant paradigm for 3D perception in autonomous driving, providing a unified spatial canvas where detection and segmentation features are geometrically registered to the same physical coordinate system. However,...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 31,
        "evidence": [
          "abstract: perception",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12918",
      "pdf_url": "https://arxiv.org/pdf/2604.12918.pdf"
    },
    {
      "id": "2604.12656",
      "anchor": "paper-2604-12656",
      "title": "FeaXDrive: Feasibility-aware Trajectory-Centric Diffusion Planning for End-to-End Autonomous Driving",
      "authors": [
        "Baoyun Wang",
        "Zhuoren Li",
        "Ran Yu",
        "Yu Che",
        "Xinrui Zhang",
        "Ming Liu",
        "Jia Hu",
        "Chen Lv",
        "Bo Leng"
      ],
      "abstract": "End-to-end diffusion planning has shown strong potential for autonomous driving, but the physical feasibility of generated trajectories remains insufficiently addressed. In particular, generated trajectories may exhibit local geometric irregularities, violate trajectory-level kinematic constraints, or deviate from the drivable area, indicating that the commonly used noise-centric formulation in diffusion planning is not yet well aligned with the trajectory space where feasibility is more naturally characterized. To address this issue, we propose FeaXDrive, a feasibility-aware trajectory-centric diffusion planning method for end-to-end autonomous driving. The core idea is to treat the clean trajectory as the unified object for feasibility-aware modeling throughout the diffusion process. Built on this trajectory-centric formulation, FeaXDrive integrates adaptive curvature-constrained training to improve intrinsic geometric and kinematic feasibility, drivable-area guidance within reverse diffusion sampling to enhance consistency with the drivable area, and feasibility-aware GRPO post-training to further improve planning performance while balancing trajectory-space feasibility. Experiments on the NAVSIM benchmark show that FeaXDrive achieves strong closed-loop planning performance while substantially improving trajectory-space feasibility. These findings highlight the importance of explicitly modeling trajectory-space feasibility in end-to-end diffusion planning and provide a step toward more reliable and physically grounded autonomous driving planners.",
      "short_abstract": "End-to-end diffusion planning has shown strong potential for autonomous driving, but the physical feasibility of generated trajectories remains insufficiently addressed. In particular, generated trajectories may exhibit local geometric irregularities, violate...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12656",
      "pdf_url": "https://arxiv.org/pdf/2604.12656.pdf"
    },
    {
      "id": "2604.12418",
      "anchor": "paper-2604-12418",
      "title": "RACF: A Resilient Autonomous Car Framework with Object Distance Correction",
      "authors": [
        "Chieh Tsai",
        "Hossein Rastgoftar",
        "Salim Hariri"
      ],
      "abstract": "Autonomous vehicles are increasingly deployed in safety-critical applications, where sensing failures or cyberphysical attacks can lead to unsafe operations resulting in human loss and/or severe physical damages. Reliable real-time perception is therefore critically important for their safe operations and acceptability. For example, vision-based distance estimation is vulnerable to environmental degradation and adversarial perturbations, and existing defenses are often reactive and too slow to promptly mitigate their impacts on safe operations. We present a Resilient Autonomous Car Framework (RACF) that incorporates an Object Distance Correction Algorithm (ODCA) to improve perception-layer robustness through redundancy and diversity across a depth camera, LiDAR, and physics-based kinematics. Within this framework, when obstacle distance estimation produced by depth camera is inconsistent, a cross-sensor gate activates the correction algorithm to fix the detected inconsistency. We have experiment with the proposed resilient car framework and evaluate its performance on a testbed implemented using the Quanser QCar 2 platform. The presented framework achieved up to 35% RMSE reduction under strong corruption and improves stop compliance and braking latency, while operating in real time. These results demonstrate a practical and lightweight approach to resilient perception for safety-critical autonomous driving",
      "short_abstract": "Autonomous vehicles are increasingly deployed in safety-critical applications, where sensing failures or cyberphysical attacks can lead to unsafe operations resulting in human loss and/or severe physical damages. Reliable real-time perception is therefore...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: attack or threat",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12418",
      "pdf_url": "https://arxiv.org/pdf/2604.12418.pdf"
    },
    {
      "id": "2604.12331",
      "anchor": "paper-2604-12331",
      "title": "HyperLiDAR: Adaptive Post-Deployment LiDAR Segmentation via Hyperdimensional Computing",
      "authors": [
        "Ivannia Gomez Moreno",
        "Yi Yao",
        "Ye Tian",
        "Xiaofan Yu",
        "Flavio Ponzina",
        "Michael Sullivan",
        "Jingyi Zhang",
        "Mingyu Yang",
        "Hun Seok Kim",
        "Tajana Rosing"
      ],
      "abstract": "LiDAR semantic segmentation plays a pivotal role in 3D scene understanding for edge applications such as autonomous driving. However, significant challenges remain for real-world deployments, particularly for on-device post-deployment adaptation. Real-world environments can shift as the system navigates through different locations, leading to substantial performance degradation without effective and timely model adaptation. Furthermore, edge systems operate under strict computational and energy constraints, making it infeasible to adapt conventional segmentation models (based on large neural networks) directly on-device. To address the above challenges, we introduce HyperLiDAR, the first lightweight, post-deployment LiDAR segmentation framework based on Hyperdimensional Computing (HDC). The design of HyperLiDAR fully leverages the fast learning and high efficiency of HDC, inspired by how the human brain processes information. To further improve the adaptation efficiency, we identify the high data volume per scan as a key bottleneck and introduce a buffer selection strategy that focuses learning on the most informative points. We conduct extensive evaluations on two state-of-the-art LiDAR segmentation benchmarks and two representative devices. Our results show that HyperLiDAR outperforms or achieves comparable adaptation performance to state-of-the-art segmentation methods, while achieving up to a 13.8x speedup in retraining.",
      "short_abstract": "LiDAR semantic segmentation plays a pivotal role in 3D scene understanding for edge applications such as autonomous driving. However, significant challenges remain for real-world deployments, particularly for on-device post-deployment adaptation. Real-world...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 19,
        "margin": 5,
        "evidence": [
          "abstract: scene segmentation",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12331",
      "pdf_url": "https://arxiv.org/pdf/2604.12331.pdf"
    },
    {
      "id": "2604.12857",
      "anchor": "paper-2604-12857",
      "title": "Artificial Intelligence for Modeling and Simulation of Mixed Automated and Human Traffic",
      "authors": [
        "Saeed Rahmani",
        "Shiva Rasouli",
        "Daphne Cornelisse",
        "Eugene Vinitsky",
        "Bart van Arem",
        "Simeon C. Calvert"
      ],
      "abstract": "Autonomous vehicles (AVs) are now operating on public roads, which makes their testing and validation more critical than ever. Simulation offers a safe and controlled environment for evaluating AV performance in varied conditions. However, existing simulation tools mainly focus on graphical realism and rely on simple rule-based models and therefore fail to accurately represent the complexity of driving behaviors and interactions. Artificial intelligence (AI) has shown strong potential to address these limitations; however, despite the rapid progress across AI methodologies, a comprehensive survey of their application to mixed autonomy traffic simulation remains lacking. Existing surveys either focus on simulation tools without examining the AI methods behind them, or cover ego-centric decision-making without addressing the broader challenge of modeling surrounding traffic. Moreover, they do not offer a unified taxonomy of AI methods covering individual behavior modeling to full scene simulation. To address these gaps, this survey provides a structured review and synthesis of AI methods for modeling AV and human driving behavior in mixed autonomy traffic simulation. We introduce a taxonomy that organizes methods into three families: agent-level behavior models, environment-level simulation methods, and cognitive and physics-informed methods. The survey analyzes how existing simulation platforms fall short of the needs of mixed autonomy research and outlines directions to narrow this gap. It also provides a chronological overview of AI methods and reviews evaluation protocols and metrics, simulation tools, and datasets. By covering both traffic engineering and computer science perspectives, we aim to bridge the gap between these two communities.",
      "short_abstract": "Autonomous vehicles (AVs) are now operating on public roads, which makes their testing and validation more critical than ever. Simulation offers a safe and controlled environment for evaluating AV performance in varied conditions. However, existing simulation...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "Survey"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Planning & Decision-Making: abstract: decision-making",
          "Safety, Security & Verification: abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12857",
      "pdf_url": "https://arxiv.org/pdf/2604.12857.pdf"
    },
    {
      "id": "2604.12628",
      "anchor": "paper-2604-12628",
      "title": "A Comparison of Reinforcement Learning and Optimal Control Methods for Path Planning",
      "authors": [
        "Qiang Le",
        "Yaguang Yang",
        "Isaac E. Weintraub"
      ],
      "abstract": "Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find ideal paths, the computational time is often too slow for real-time decision-making. To solve this challenge, we propose a method based on Deep Deterministic Policy Gradient (DDPG) and model the threat as a simple, circular `no-go' zone. A mission failure is claimed if the vehicle enters this `no-go' zone at any time or does not reach a neighborhood of the destination. The DDPG agent is trained to learn a direct mapping from its current state (position and velocity) to a series of feasible actions that guide the agent to safely reach its goal. A reward function and two neural networks, critic and actor, are used to describe the environment and guide the control efforts. The DDPG trains the agent to find the largest possible set of starting points (``feasible set'') wherein a safe path to the goal is guaranteed. This provides critical information for mission planning, showing beforehand whether a task is achievable from a given starting point, assisting pre-mission planning activities. The approach is validated in simulation. A comparison between the DDPG method and a traditional optimal control (pseudo-spectral) method is carried out. The results show that the learning-based agent may produce effective paths while being significantly faster, making it a better fit for real-time applications. However, there are areas (``infeasible set'') where the DDPG agent cannot find paths to the destination, and the paths in the feasible set may not be optimal. These preliminary results guide our future research: (1) improve the reward function to enlarge the DDPG feasible set, (2) examine the feasible set obtained by the pseudo-spectral method, and (3) investigate the arc-search IPM method for the path planning problem.",
      "short_abstract": "Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find ideal paths, the computational time is often too slow for real-time decision-making. To solve this challenge, we...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 28,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12628",
      "pdf_url": "https://arxiv.org/pdf/2604.12628.pdf"
    },
    {
      "id": "2604.12590",
      "anchor": "paper-2604-12590",
      "title": "Situation-Aware Feedback-Predictive Control Framework for Lane-Less Dense Traffic",
      "authors": [
        "Parthib Khound"
      ],
      "abstract": "Navigating dense, lane-less traffic remains one of the most challenging scenarios for autonomous vehicles, especially in emerging regions where road structure and driver behavior are highly unpredictable. This paper presents a hybrid control framework tailored for such environments, integrating a $360^\\circ$ zone-based perception module with a dual-layer control strategy that combines classical feedback and predictive optimization. The longitudinal feedback controller computes reference speed based on braking distance and steering dynamics, while the lateral controller tracks a virtual optimal lane derived from the spatial distribution of neighboring vehicles. The predictive planner samples control inputs over a time horizon and selects the most feasible trajectory using a multi-term cost function. Simulation results across diverse one-way traffic scenarios demonstrate the framework's robustness, responsiveness, and suitability for chaotic, unstructured traffic.",
      "short_abstract": "Navigating dense, lane-less traffic remains one of the most challenging scenarios for autonomous vehicles, especially in emerging regions where road structure and driver behavior are highly unpredictable. This paper presents a hybrid control framework...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 8,
        "evidence": [
          "abstract: controller",
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12590",
      "pdf_url": "https://arxiv.org/pdf/2604.12590.pdf"
    },
    {
      "id": "2604.12408",
      "anchor": "paper-2604-12408",
      "title": "Security and Resilience in Autonomous Vehicles: A Proactive Design Approach",
      "authors": [
        "Chieh Tsai",
        "Murad Mehrab Abrar",
        "Salim Hariri"
      ],
      "abstract": "Autonomous vehicles (AVs) promise efficient, clean and cost-effective transportation systems, but their reliance on sensors, wireless communications, and decision-making systems makes them vulnerable to cyberattacks and physical threats. This chapter presents novel design techniques to strengthen the security and resilience of AVs. We first provide a taxonomy of potential attacks across different architectural layers, from perception and control manipulation to Vehicle-to-Any (V2X) communication exploits and software supply chain compromises. Building on this analysis, we present an AV Resilient architecture that integrates redundancy, diversity, and adaptive reconfiguration strategies, supported by anomaly- and hash-based intrusion detection techniques. Experimental validation on the Quanser QCar platform demonstrates the effectiveness of these methods in detecting depth camera blinding attacks and software tampering of perception modules. The results highlight how fast anomaly detection combined with fallback and backup mechanisms ensures operational continuity, even under adversarial conditions. By linking layered threat modeling with practical defense implementations, this work advances AV resilience strategies for safer and more trustworthy autonomous vehicles.",
      "short_abstract": "Autonomous vehicles (AVs) promise efficient, clean and cost-effective transportation systems, but their reliance on sensors, wireless communications, and decision-making systems makes them vulnerable to cyberattacks and physical threats. This chapter presents...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 19,
        "evidence": [
          "title: security",
          "abstract: security",
          "abstract: validation or testing",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12408",
      "pdf_url": "https://arxiv.org/pdf/2604.12408.pdf"
    },
    {
      "id": "2604.12293",
      "anchor": "paper-2604-12293",
      "title": "Defining an Evaluation Method for External Human-Machine Interfaces",
      "authors": [
        "Jose Gonzalez-Belmonte",
        "Jaerock Kwon"
      ],
      "abstract": "As the number of fatalities involving Autonomous Vehicles increase, the need for a universal method of communicating between vehicles and other agents on the road has also increased. Over the past decade, numerous proposals of external Human-Machine Interfaces (eHMIs) have been brought forward with the purpose of bridging this communication gap, with none yet to be determined as the ideal one. This work proposes a universal evaluation method conformed of 223 questions to objectively evaluate and compare different proposals and arrive at a conclusion. The questionnaire is divided into 7 categories that evaluate different aspects of any given proposal that uses eHMIs: ease of standardization, cost effectiveness, accessibility, ease of understanding, multifacetedness in communication, positioning, and readability. In order to test the method it was used on four existing proposals, plus a baseline using only kinematic motions, in order to both exemplify the application of the evaluation method and offer a baseline score for future comparison. The result of this testing suggests that the ideal method of machine-human communication is a combination of intentionally-designed vehicle kinematics and distributed well-placed text-based displays, but it also reveals knowledge gaps in the readability of eHMIs and the speed at which different observers may learn their meaning. This paper proposes future work related to these uncertainties, along with future testing with the proposed method.",
      "short_abstract": "As the number of fatalities involving Autonomous Vehicles increase, the need for a universal method of communicating between vehicles and other agents on the road has also increased. Over the past decade, numerous proposals of external Human-Machine...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 15,
        "evidence": [
          "title: human-machine interaction",
          "abstract: human-machine interaction"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12293",
      "pdf_url": "https://arxiv.org/pdf/2604.12293.pdf"
    },
    {
      "id": "2604.12425",
      "anchor": "paper-2604-12425",
      "title": "Forecasting the Past: Gradient-Based Distribution Shift Detection in Trajectory Prediction",
      "authors": [
        "Michele De Vita",
        "Julian Wiederer",
        "Vasileios Belagiannis"
      ],
      "abstract": "Trajectory prediction models often fail in real-world automated driving due to distributional shifts between training and test conditions. Such distributional shifts, whether behavioural or environmental, pose a critical risk by causing the model to make incorrect forecasts in unfamiliar situations. We propose a self-supervised method that trains a decoder in a post-hoc fashion on the self-supervised task of forecasting the second half of observed trajectories from the first half. The L2 norm of the gradient of this forecasting loss with respect to the decoder's final layer defines a score to identify distribution shifts. Our approach, first, does not affect the trajectory prediction model, ensuring no interference with original prediction performance and second, demonstrates substantial improvements on distribution shift detection for trajectory prediction on the Shifts and Argoverse datasets. Moreover, we show that this method can also be used to early detect collisions of a deep Q-Network motion planner in the Highway simulator. Source code is available at https://github.com/Michedev/forecasting-the-past.",
      "short_abstract": "Trajectory prediction models often fail in real-world automated driving due to distributional shifts between training and test conditions. Such distributional shifts, whether behavioural or environmental, pose a critical risk by causing the model to make...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 40,
        "margin": 38,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: forecasting",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.12425",
      "pdf_url": "https://arxiv.org/pdf/2604.12425.pdf"
    },
    {
      "id": "2604.13424",
      "anchor": "paper-2604-13424",
      "title": "Integrated Routing and Intersection Control for Mixed Traffic",
      "authors": [
        "Filippos N. Tzortzoglou",
        "Pengbo Zhu",
        "Andreas A. Malikopoulos"
      ],
      "abstract": "The rapid development of cyber-physical systems is driving a transition toward mixed traffic environments comprising both human-driven and connected and automated vehicles (CAVs). This shift presents a unique opportunity to leverage the efficient operation of CAVs to improve overall network throughput. This paper introduces a hierarchical framework designed to bridge macroscopic routing optimization at the network level with microscopic vicinity control at signalized intersections. The upper layer utilizes aggregated traffic information to provide proactive routing guidance for CAVs, aiming to minimize total travel time. The lower layer leverages local vehicle states to jointly optimize traffic light phases and individual CAV trajectories, aiming to reduce intersection crossing delays and optimize energy consumption, respectively. The effectiveness of the proposed framework is validated through SUMO on the Sioux Falls benchmark network. Results demonstrate that the integration of these macroscopic and microscopic layers yields significantly better performance compared to applying either layer in isolation, significantly improving network throughput and reducing congestion.",
      "short_abstract": "The rapid development of cyber-physical systems is driving a transition toward mixed traffic environments comprising both human-driven and connected and automated vehicles (CAVs). This shift presents a unique opportunity to leverage the efficient operation of...",
      "published": "2026-04-14",
      "updated": "2026-04-14",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 2,
        "evidence": [
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 127,
      "arxiv_url": "https://arxiv.org/abs/2604.13424",
      "pdf_url": "https://arxiv.org/pdf/2604.13424.pdf"
    },
    {
      "id": "2604.12239",
      "anchor": "paper-2604-12239",
      "title": "Physics-Grounded Monocular Vehicle Distance Estimation Using Standardized License Plate Typography",
      "authors": [
        "Manognya Lokesh Reddy",
        "Zheng Liu"
      ],
      "abstract": "Accurate inter-vehicle distance estimation is a cornerstone of Advanced Driver Assistance Systems (ADAS) and autonomous driving. While LiDAR and radar provide high precision, their high cost prohibits widespread adoption in mass-market vehicles. Monocular camera-based estimation offers a low-cost alternative but suffers from fundamental scale ambiguity. Recent deep learning methods for monocular depth achieve impressive results yet require expensive supervised training, suffer from domain shift, and produce predictions that are difficult to certify for safety-critical deployment. This paper presents a framework that exploits the standardized typography of United States license plates as passive fiducial markers for metric ranging, resolving scale ambiguity through explicit geometric priors without any training data or active illumination. First, a four-method parallel plate detector achieves robust plate reading across the full automotive lighting range. Second, a three-stage state identification engine fusing optical character recognition text matching, multi-design color scoring, and a lightweight neural network classifier provides robust identification across all ambient conditions. Third, hybrid depth fusion with inverse-variance weighting and online scale alignment, combined with a one-dimensional constant-velocity Kalman filter, delivers smoothed distance, relative velocity, and time-to-collision for collision warning. Baseline validation on a controlled static dataset reproduces a 2.3% coefficient of variation in character height measurements and a 36% reduction in distance-estimate variance compared with plate-width methods from prior work.",
      "short_abstract": "Accurate inter-vehicle distance estimation is a cornerstone of Advanced Driver Assistance Systems (ADAS) and autonomous driving. While LiDAR and radar provide high precision, their high cost prohibits widespread adoption in mass-market vehicles. Monocular...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.12239",
      "pdf_url": "https://arxiv.org/pdf/2604.12239.pdf"
    },
    {
      "id": "2604.12208",
      "anchor": "paper-2604-12208",
      "title": "Unveiling the Surprising Efficacy of Navigation Understanding in End-to-End Autonomous Driving",
      "authors": [
        "Zhihua Hua",
        "Junli Wang",
        "Pengfei LI",
        "Qihao Jin",
        "Bo Zhang",
        "Kehua Sheng",
        "Yilun Chen",
        "Zhongxue Gan",
        "Wenchao Ding"
      ],
      "abstract": "Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-end autonomous driving systems tend to over-rely on local scene understanding while failing to utilize global navigation information. These systems exhibit weak correlation between their planning capabilities and navigation input, and struggle to perform navigation-following in complex scenarios. To overcome this limitation, we propose the Sequential Navigation Guidance (SNG) framework, an efficient representation of global navigation information based on real-world navigation patterns. The SNG encompasses both navigation paths for constraining long-term trajectories and turn-by-turn (TBT) information for real-time decision-making logic. We constructed the SNG-QA dataset, a visual question answering (VQA) dataset based on SNG that aligns global and local planning. Additionally, we introduce an efficient model SNG-VLA that fuses local planning with global planning. The SNG-VLA achieves state-of-the-art performance through precise navigation information modeling without requiring auxiliary loss functions from perception tasks. Project page: SNG-VLA",
      "short_abstract": "Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-end autonomous driving systems tend to over-rely on local scene understanding...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "VLA"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 27,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.12208",
      "pdf_url": "https://arxiv.org/pdf/2604.12208.pdf"
    },
    {
      "id": "2604.11854",
      "anchor": "paper-2604-11854",
      "title": "MVAdapt: Zero-Shot Multi-Vehicle Adaptation for End-to-End Autonomous Driving",
      "authors": [
        "Haesung Oh",
        "Jaeheung Park"
      ],
      "abstract": "End-to-End (E2E) autonomous driving models are usually trained and evaluated with a fixed ego-vehicle, even though their driving policy is implicitly tied to vehicle dynamics. When such a model is deployed on a vehicle with different size, mass, or drivetrain characteristics, its performance can degrade substantially; we refer to this problem as the vehicle-domain gap. To address it, we propose MVAdapt, a physics-conditioned adaptation framework for multi-vehicle E2E driving. MVAdapt combines a frozen TransFuser++ scene encoder with a lightweight physics encoder and a cross-attention module that conditions scene features on vehicle properties before waypoint decoding. In the CARLA Leaderboard 1.0 benchmark, MVAdapt improves over naive transfer and multi-embodiment adaptation baselines on both in-distribution and unseen vehicles. We further show two complementary behaviors: strong zero-shot transfer on many unseen vehicles, and data-efficient few-shot calibration for severe physical outliers. These results suggest that explicitly conditioning E2E driving policies on vehicle physics is an effective step toward more transferable autonomous driving models. All codes are available at https://github.com/hae-sung-oh/MVAdapt",
      "short_abstract": "End-to-End (E2E) autonomous driving models are usually trained and evaluated with a fixed ego-vehicle, even though their driving policy is implicitly tied to vehicle dynamics. When such a model is deployed on a vehicle with different size, mass, or drivetrain...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 27,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end",
          "abstract: driving policy"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.11854",
      "pdf_url": "https://arxiv.org/pdf/2604.11854.pdf"
    },
    {
      "id": "2604.11707",
      "anchor": "paper-2604-11707",
      "title": "Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction",
      "authors": [
        "Efstathios Karypidis",
        "Spyros Gidaris",
        "Nikos Komodakis"
      ],
      "abstract": "Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as autonomous driving. We present Re2Pix, a hierarchical video prediction framework that decomposes forecasting into two stages: semantic representation prediction and representation-guided visual synthesis. Instead of directly predicting future RGB frames, our approach first forecasts future scene structure in the feature space of a frozen vision foundation model, and then conditions a latent diffusion model on these predicted representations to render photorealistic frames. This decomposition enables the model to focus first on scene dynamics and then on appearance generation. A key challenge arises from the train-test mismatch between ground-truth representations available during training and predicted ones used at inference. To address this, we introduce two conditioning strategies, nested dropout and mixed supervision, that improve robustness to imperfect autoregressive predictions. Experiments on challenging driving benchmarks demonstrate that the proposed semantics-first design significantly improves temporal semantic consistency, perceptual quality, and training efficiency compared to strong diffusion baselines. We provide the implementation code at https://github.com/Sta8is/Re2Pix",
      "short_abstract": "Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as autonomous driving. We present Re2Pix, a hierarchical video prediction framework that decomposes...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.11707",
      "pdf_url": "https://arxiv.org/pdf/2604.11707.pdf"
    },
    {
      "id": "2604.11081",
      "anchor": "paper-2604-11081",
      "title": "MapATM: Enhancing HD Map Construction through Actor Trajectory Modeling",
      "authors": [
        "Mingyang Li",
        "Brian Lee",
        "Rui Zuo",
        "Brent Bacchus",
        "Priyantha Mudalige",
        "Qinru Qiu"
      ],
      "abstract": "High-definition (HD) mapping tasks, which perform lane detections and predictions, are extremely challenging due to non-ideal conditions such as view occlusions, distant lane visibility, and adverse weather conditions. Those conditions often result in compromised lane detection accuracy and reduced reliability within autonomous driving systems. To address these challenges, we introduce MapATM, a novel deep neural network that effectively leverages historical actor trajectory information to improve lane detection accuracy, where actors refer to moving vehicles. By utilizing actor trajectories as structural priors for road geometry, MapATM achieves substantial performance enhancements, notably increasing AP by 4.6 for lane dividers and mAP by 2.6 on the challenging NuScenes dataset, representing relative improvements of 10.1% and 6.1%, respectively, compared to strong baseline methods. Extensive qualitative evaluations further demonstrate MapATM's capability to consistently maintain stable and robust map reconstruction across diverse and complex driving scenarios, underscoring its practical value for autonomous driving applications.",
      "short_abstract": "High-definition (HD) mapping tasks, which perform lane detections and predictions, are extremely challenging due to non-ideal conditions such as view occlusions, distant lane visibility, and adverse weather conditions. Those conditions often result in...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 28,
        "evidence": [
          "title: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.11081",
      "pdf_url": "https://arxiv.org/pdf/2604.11081.pdf"
    },
    {
      "id": "2604.11406",
      "anchor": "paper-2604-11406",
      "title": "Using Unwrapped Full Color Space Recording to Measure the Exposedness of Vehicle Exterior Parts for External Human Machine Interfaces",
      "authors": [
        "Jose Gonzalez-Belmonte",
        "Jaerock Kwon"
      ],
      "abstract": "One of the concerns with autonomous vehicles is their ability to communicate their intent to other road users, specially pedestrians, in order to prevent accidents. External Human-Machine Interfaces (eHMIs) are the proposed solution to this issue, through the introduction of electronic devices on the exterior of a vehicle that communicate when the vehicle is planning on slowing down or yielding. This paper uses the technique of unwrapping the faces of a mesh onto a texture where every pixel is a unique color, as well as a series of animated simulations made and ran in the Unity game engine, to measure how many times is each point on a 2015 Ford F-150 King Ranch is unobstructed to a pedestrian attempting to cross the road at a four-way intersection. By cross-referencing the results with a color-coded map of the labeled parts on the exterior of the vehicle, it was concluded that while the bumper, grill, and hood were the parts of the vehicle visible to the crossing pedestrian most often, the existence of other vehicles on the same lane that might obstruct the view of these makes them insufficient. The study recommends instead a distributive approach to eHMIs by using both the windshield and frontal fenders as simultaneous placements for these devices.",
      "short_abstract": "One of the concerns with autonomous vehicles is their ability to communicate their intent to other road users, specially pedestrians, in order to prevent accidents. External Human-Machine Interfaces (eHMIs) are the proposed solution to this issue, through the...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 17,
        "evidence": [
          "title: human-machine interaction",
          "abstract: human-machine interaction"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.11406",
      "pdf_url": "https://arxiv.org/pdf/2604.11406.pdf"
    },
    {
      "id": "2604.10993",
      "anchor": "paper-2604-10993",
      "title": "On Switched Event-triggered Full State-constrained Formation Control for Multi-vehicle Systems",
      "authors": [
        "Zihan Li",
        "Ziming Wang",
        "Xin Wang"
      ],
      "abstract": "Vehicular formation control is an important component of intelligent transportation systems (ITSs). In practical implementations, the controller design needs to satisfy multiple state constraints, including inter-vehicle spacing and vehicle speed. When system states approach the constraint boundaries, control singularity and excessive control effort may arise, which limits the practical applicability of existing methods. To address this problem, this paper investigates a class of nonlinear vehicular formation systems for autonomous vehicles (AVs) with uncertain dynamics and develops a switched event-triggered control framework. A smooth nonlinear mapping is first introduced to transform the constrained state space into an unconstrained one, thereby avoiding singularity near the constraint boundaries. A radial basis function neural network (RBFNN) is then employed to approximate the unknown nonlinear dynamics online, based on which an adaptive controller is constructed via the backstepping technique. In addition, a switched event-triggered mechanism (SETM) is designed to increase the control update frequency during the transient stage and reduce the communication burden during the steady-state stage. Lyapunov-based analysis proves that all signals in the closed-loop system remain uniformly bounded and that Zeno behavior is excluded. Simulation results verify that the proposed method achieves stable platoon formation under prescribed state constraints while significantly reducing communication updates.",
      "short_abstract": "Vehicular formation control is an important component of intelligent transportation systems (ITSs). In practical implementations, the controller design needs to satisfy multiple state constraints, including inter-vehicle spacing and vehicle speed. When system...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: controller",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.10993",
      "pdf_url": "https://arxiv.org/pdf/2604.10993.pdf"
    },
    {
      "id": "2604.11549",
      "anchor": "paper-2604-11549",
      "title": "Human Centered Non Intrusive Driver State Modeling Using Personalized Physiological Signals in Real World Automated Driving",
      "authors": [
        "David Puertas-Ramirez",
        "Raul Fernandez-Matellan",
        "David Martin Gomez",
        "Jesus G. Boticario"
      ],
      "abstract": "In vehicles with partial or conditional driving automation (SAE Levels 2-3), the driver remains responsible for supervising the system and responding to take-over requests. Therefore, reliable driver monitoring is essential for safe human-automation collaboration. However, most existing Driver Monitoring Systems rely on generalized models that ignore individual physiological variability. In this study, we examine the feasibility of personalized driver state modeling using non-intrusive physiological sensing during real-world automated driving. We conducted experiments in an SAE Level 2 vehicle using an Empatica E4 wearable sensor to capture multimodal physiological signals, including electrodermal activity, heart rate, temperature, and motion data. To leverage deep learning architectures designed for images, we transformed the physiological signals into two-dimensional representations and processed them using a multimodal architecture based on pre-trained ResNet50 feature extractors. Experiments across four drivers demonstrate substantial interindividual variability in physiological patterns related to driver awareness. Personalized models achieved an average accuracy of 92.68%, whereas generalized models trained on multiple users dropped to an accuracy of 54%, revealing substantial limitations in cross-user generalization. These results underscore the necessity of adaptive, personalized driver monitoring systems for future automated vehicles and imply that autonomous systems should adapt to each driver's unique physiological profile.",
      "short_abstract": "In vehicles with partial or conditional driving automation (SAE Levels 2-3), the driver remains responsible for supervising the system and responding to take-over requests. Therefore, reliable driver monitoring is essential for safe human-automation...",
      "published": "2026-04-13",
      "updated": "2026-04-13",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 25,
        "evidence": [
          "title: driver monitoring",
          "abstract: driver monitoring",
          "abstract: takeover"
        ]
      },
      "recency": "archive",
      "age_days": 128,
      "arxiv_url": "https://arxiv.org/abs/2604.11549",
      "pdf_url": "https://arxiv.org/pdf/2604.11549.pdf"
    },
    {
      "id": "2604.10856",
      "anchor": "paper-2604-10856",
      "title": "BridgeSim: Unveiling the OL-CL Gap in End-to-End Autonomous Driving",
      "authors": [
        "Seth Z. Zhao",
        "Luobin Wang",
        "Hongwei Ruan",
        "Yuxin Bao",
        "Yilan Chen",
        "Ziyang Leng",
        "Abhijit Ravichandran",
        "Honglin He",
        "Zewei Zhou",
        "Xu Han",
        "Abhishek Peri",
        "Zhiyu Huang",
        "Pranav Desai",
        "Henrik Christensen",
        "Jiaqi Ma",
        "Bolei Zhou"
      ],
      "abstract": "Open-loop (OL) to closed-loop (CL) gap (OL-CL gap) exists when OL-pretrained policies scoring high in OL evaluations fail to transfer effectively in closed-loop (CL) deployment. In this paper, we unveil the root causes of this systemic failure and propose a practical remedy. Specifically, we demonstrate that OL policies suffer from Observational Domain Shift and Objective Mismatch. We show that while the former is largely recoverable with adaptation techniques, the latter creates a structural inability to model complex reactive behaviors, which forms the primary OL-CL gap. We find that a wide range of OL policies learn a biased Q-value estimator that neglects both the reactive nature of CL simulations and the temporal awareness needed to reduce compounding errors. To this end, we propose a Test-Time Adaptation (TTA) framework that calibrates observational shift, reduces state-action biases, and enforces temporal consistency. Extensive experiments show that TTA effectively mitigates planning biases and yields superior scaling dynamics than its baseline counterparts. Furthermore, our analysis highlights the existence of blind spots in standard OL evaluation protocols that fail to capture the realities of closed-loop deployment.",
      "short_abstract": "Open-loop (OL) to closed-loop (CL) gap (OL-CL gap) exists when OL-pretrained policies scoring high in OL evaluations fail to transfer effectively in closed-loop (CL) deployment. In this paper, we unveil the root causes of this systemic failure and propose a...",
      "published": "2026-04-12",
      "updated": "2026-04-12",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 129,
      "arxiv_url": "https://arxiv.org/abs/2604.10856",
      "pdf_url": "https://arxiv.org/pdf/2604.10856.pdf"
    },
    {
      "id": "2604.10780",
      "anchor": "paper-2604-10780",
      "title": "LIDARLearn: A Unified Deep Learning Library for 3D Point Cloud Classification, Segmentation, and Self-Supervised Representation Learning",
      "authors": [
        "Said Ohamouddou",
        "Hanaa El Afia",
        "Abdellatif El Afia",
        "Raddouane Chiheb"
      ],
      "abstract": "Three-dimensional (3D) point cloud analysis has become central to applications ranging from autonomous driving and robotics to forestry and ecological monitoring. Although numerous deep learning methods have been proposed for point cloud understanding, including supervised backbones, self-supervised pre-training (SSL), and parameter-efficient fine-tuning (PEFT), their implementations are scattered across incompatible codebases with differing data pipelines, evaluation protocols, and configuration formats, making fair comparisons difficult. We introduce \\lib{}, a unified, extensible PyTorch library that integrates over 55 model configurations covering 29 supervised architectures, seven SSL pre-training methods, and five PEFT strategies, all within a single registry-based framework supporting classification, semantic segmentation, part segmentation, and few-shot learning. \\lib{} provides standardised training runners, cross-validation with stratified $K$-fold splitting, automated LaTeX/CSV table generation, built-in Friedman/Nemenyi statistical testing with critical-difference diagrams for rigorous multi-model comparison, and a comprehensive test suite with 2\\,200+ automated tests validating every configuration end-to-end. The code is available at https://github.com/said-ohamouddou/LIDARLearn under the MIT licence.",
      "short_abstract": "Three-dimensional (3D) point cloud analysis has become central to applications ranging from autonomous driving and robotics to forestry and ecological monitoring. Although numerous deep learning methods have been proposed for point cloud understanding,...",
      "published": "2026-04-12",
      "updated": "2026-04-12",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "abstract: scene segmentation",
          "title: point cloud",
          "abstract: point cloud"
        ]
      },
      "recency": "archive",
      "age_days": 129,
      "arxiv_url": "https://arxiv.org/abs/2604.10780",
      "pdf_url": "https://arxiv.org/pdf/2604.10780.pdf"
    },
    {
      "id": "2604.10662",
      "anchor": "paper-2604-10662",
      "title": "Energy-Efficient Federated Edge Learning For Small-Scale Datasets in Large IoT Networks",
      "authors": [
        "Haihui Xie",
        "Wenkun Wen",
        "Shuwu Chen",
        "Zhaogang Shu",
        "Minghua Xia"
      ],
      "abstract": "Large-scale Internet of Things (IoT) networks enable intelligent services such as smart cities and autonomous driving, but often face resource constraints. Collecting heterogeneous sensory data, especially in small-scale datasets, is challenging, and independent edge nodes can lead to inefficient resource utilization and reduced learning performance. To address these issues, this paper proposes a collaborative optimization framework for energy-efficient federated edge learning with small-scale datasets. We first derive an expected learning loss to quantify the relationship between the number of training samples and learning objectives. A stochastic online learning algorithm is then designed to adapt to data variations, and a resource optimization problem with a convergence bound is formulated. Finally, an online distributed algorithm efficiently solves large-scale optimization problems with high scalability. Extensive simulations and autonomous navigation case studies with collision avoidance demonstrate that the proposed approach significantly improves learning performance and resource efficiency compared to state-of-the-art benchmarks.",
      "short_abstract": "Large-scale Internet of Things (IoT) networks enable intelligent services such as smart cities and autonomous driving, but often face resource constraints. Collecting heterogeneous sensory data, especially in small-scale datasets, is challenging, and...",
      "published": "2026-04-12",
      "updated": "2026-04-12",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 9,
        "evidence": [
          "abstract: cooperative systems",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 129,
      "arxiv_url": "https://arxiv.org/abs/2604.10662",
      "pdf_url": "https://arxiv.org/pdf/2604.10662.pdf"
    },
    {
      "id": "2604.21941",
      "anchor": "paper-2604-21941",
      "title": "When Altruism Meets Autonomy: Managing Bottleneck Congestion with Strategic Autonomous Vehicles",
      "authors": [
        "Kexin Wang",
        "Haohui He",
        "Ruolin Li"
      ],
      "abstract": "Weaving ramps are critical bottlenecks in highway networks due to conflicting traffic flows and complex interactions among heterogeneous vehicle types. In mixed-autonomy settings, the presence of controllable autonomous vehicles (AVs) introduces new opportunities to influence system-level outcomes, yet the structural impact of such control remains poorly understood. This paper develops a unified equilibrium framework to capture, predict, and optimize aggregate lane-choice behavior in weaving ramps with heterogeneous vehicle populations. We first formulate a Wardrop-based model capturing the selfish behavior of human-driven vehicles (HDVs) and establish existence, uniqueness, and validity of the resulting equilibrium. We then introduce a Stackelberg--Wardrop formulation in which AVs act as strategic leaders optimizing system performance, while HDVs respond through equilibrium adaptation. The framework is further generalized to incorporate heterogeneous behavioral preferences of HDVs and AVs via a Social Value Orientation (SVO) model. Our analysis reveals a fundamental structural property of mixed-autonomy traffic systems: under selfish HDV behavior, the impact of AV penetration is inherently non-increasing, exhibiting plateau regions where performance remains unchanged and improves only at critical thresholds. These results provide principled guidance for the design of AV control and incentive mechanisms in the presence of selfish human behavior, and demonstrate how strategically controlled autonomous agents can be deployed to induce system-level efficiency gains in mixed-autonomy transportation networks.",
      "short_abstract": "Weaving ramps are critical bottlenecks in highway networks due to conflicting traffic flows and complex interactions among heterogeneous vehicle types. In mixed-autonomy settings, the presence of controllable autonomous vehicles (AVs) introduces new...",
      "published": "2026-04-12",
      "updated": "2026-04-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 129,
      "arxiv_url": "https://arxiv.org/abs/2604.21941",
      "pdf_url": "https://arxiv.org/pdf/2604.21941.pdf"
    },
    {
      "id": "2605.23915",
      "anchor": "paper-2605-23915",
      "title": "SEIDM: A Safe and Efficient Intelligent Driver Model for Autonomous Driving Behavior",
      "authors": [
        "Yuyang Yao",
        "Shaocheng Luo"
      ],
      "abstract": "The Intelligent Driver Model (IDM) is a cornerstone of Adaptive Cruise Control (ACC), valued for its interpretable parameters and effectiveness in car-following behavior modeling. However, its inherent conservatism leads to prolonged stabilization and reduced traffic efficiency, which have received limited attention. In this paper, we propose SEIDM (Safe and Efficient Intelligent Driver Model), an enhanced IDM extension designed to improve traffic flow efficiency without sacrificing safety. SEIDM introduces an adaptive safety factor to dynamically modulate the impact of the safe deceleration term in acceleration decisions. This allows vehicles to follow more assertively under safe conditions while behaving more cautiously in potential hazards. Extensive urban traffic simulations show that SEIDM achieves significantly shorter stabilization spacing and faster convergence to traffic flow equilibrium, outperforming the original IDM and its variants in traffic stability and efficiency.",
      "short_abstract": "The Intelligent Driver Model (IDM) is a cornerstone of Adaptive Cruise Control (ACC), valued for its interpretable parameters and effectiveness in car-following behavior modeling. However, its inherent conservatism leads to prolonged stabilization and reduced...",
      "published": "2026-04-11",
      "updated": "2026-04-11",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control",
          "Control & Vehicle Dynamics: abstract: steering or braking",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 130,
      "arxiv_url": "https://arxiv.org/abs/2605.23915",
      "pdf_url": "https://arxiv.org/pdf/2605.23915.pdf"
    },
    {
      "id": "2604.10436",
      "anchor": "paper-2604-10436",
      "title": "SignReasoner: Compositional Reasoning for Complex Traffic Sign Understanding via Functional Structure Units",
      "authors": [
        "Ruibin Wang",
        "Zhenyu Lin",
        "Xinhai Zhao"
      ],
      "abstract": "Accurate semantic understanding of complex traffic signs-including those with intricate layouts, multi-lingual text, and composite symbols-is critical for autonomous driving safety. Current models, both specialized small ones and large Vision Language Models (VLMs), suffer from a significant bottleneck: a lack of compositional generalization, leading to failure when encountering novel sign configurations. To overcome this, we propose SignReasoner, a novel paradigm that transforms general VLMs into expert traffic sign reasoners. Our core innovation is Functional Structure Unit (FSU), which shifts from common instance-based modeling to flexible function-based decomposition. By breaking down complex signs into minimal, core functional blocks (e.g., Direction, Notice, Lane), our model learns the underlying structural grammar, enabling robust generalization to unseen compositions. We define this decomposition as the FSU-Reasoning task and introduce a two-stage VLM post-training pipeline to maximize performance: Iterative Caption-FSU Distillation that enhances the model's accuracy in both FSU-reasoning and caption generation; FSU-GRPO that uses Tree Edit Distance (TED) to compute FSU differences as the rewards in GRPO algorithm, boosting reasoning abilities. Experiments on the newly proposed FSU-Reasoning benchmark, TrafficSignEval, show that SignReasoner achieves new SOTA with remarkable data efficiency and no architectural modification, significantly improving the traffic sign understanding in various VLMs.",
      "short_abstract": "Accurate semantic understanding of complex traffic signs-including those with intricate layouts, multi-lingual text, and composite symbols-is critical for autonomous driving safety. Current models, both specialized small ones and large Vision Language Models...",
      "published": "2026-04-11",
      "updated": "2026-04-11",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "abstract: safety",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 130,
      "arxiv_url": "https://arxiv.org/abs/2604.10436",
      "pdf_url": "https://arxiv.org/pdf/2604.10436.pdf"
    },
    {
      "id": "2604.10169",
      "anchor": "paper-2604-10169",
      "title": "MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction",
      "authors": [
        "Wenchang Duan",
        "Zhenguo Gao",
        "Jinguo Xian",
        "Yi Shi"
      ],
      "abstract": "Trajectory prediction is a key component of autonomous driving systems because future motions directly affect collision checking, behavior planning, and control. The task remains challenging under dense interactions, heterogeneous behaviors, multimodal futures, and limited on-board computation. Existing graph, attention, and generative predictors improve interaction reasoning or uncertainty modeling, but their high-capacity designs are often costly for real-time deployment. Lightweight predictors and conventional distillation reduce inference cost, yet usually rely on static imitation and do not explicitly correct safety-relevant teacher bias. This paper proposes \\textbf{MAVEN-T}, a reinforced heterogeneous distillation framework for real-time multi-agent trajectory prediction. A high-capacity teacher models directed local interactions with a surround-aware graph encoder, combines efficient temporal filtering with shifted-window spatial attention, and decodes maneuver-specific futures through a sparse Mixture-of-Experts head. A compact GRU--Squeeze-and-Excitation student with a Low-Rank Adapted policy head is trained by feature-, attention-, and semantic-level distillation. To align prediction with downstream behavior, the student is further refined by Proximal Policy Optimization rewards for collision avoidance, comfort, and progress, while a complexity-aware curriculum and Elastic Weight Consolidation stabilize stage-wise training. Experiments on NGSIM, HighD, MoCAD, Argoverse~2, and the Waymo Open Motion Dataset evaluate accuracy, efficiency, generalization, robustness, and closed-loop safety. The student achieves 6.2$\\times$ parameter compression, 3.7$\\times$ inference acceleration, and 14.6,ms latency on an NVIDIA Jetson AGX Orin while maintaining competitive accuracy.",
      "short_abstract": "Trajectory prediction is a key component of autonomous driving systems because future motions directly affect collision checking, behavior planning, and control. The task remains challenging under dense interactions, heterogeneous behaviors, multimodal...",
      "published": "2026-04-11",
      "updated": "2026-04-11",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 14,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 130,
      "arxiv_url": "https://arxiv.org/abs/2604.10169",
      "pdf_url": "https://arxiv.org/pdf/2604.10169.pdf"
    },
    {
      "id": "2604.10391",
      "anchor": "paper-2604-10391",
      "title": "FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception",
      "authors": [
        "Rahul Ahuja",
        "Mudit Jain",
        "Bala Murali Manoghar Sai Sudhakar",
        "Venkatraman Narayanan",
        "Pratik Likhar",
        "Varun Ravi Kumar",
        "Senthil Yogamani"
      ],
      "abstract": "Vision foundation models (VFMs) and Bird's Eye View (BEV) representation have advanced visual perception substantially, yet their internal spatial representations assume the rectilinear geometry of pinhole cameras. Fisheye cameras, widely deployed on production autonomous vehicles for their surround-view coverage, exhibit severe radial distortion that renders these representations geometrically inconsistent. At the same time, the scarcity of large-scale fisheye annotations makes retraining foundation models from scratch impractical. We present \\ours, a lightweight framework that adapts frozen VFMs to fisheye geometry through two components: a frozen DINOv2 backbone with Low-Rank Adaptation (LoRA) that transfers rich self-supervised features to fisheye without task-specific pretraining, and Fisheye Rotary Position Embedding (FishRoPE), which reparameterizes the attention mechanism in the spherical coordinates of the fisheye projection so that both self-attention and cross-attention operate on angular separation rather than pixel distance. FishRoPE is architecture-agnostic, introduces negligible computational overhead, and naturally reduces to the standard formulation under pinhole geometry. We evaluate \\ours on WoodScape 2D detection (54.3 mAP) and SynWoodScapes BEV segmentation (65.1 mIoU), where it achieves state-of-the-art results on both benchmarks.",
      "short_abstract": "Vision foundation models (VFMs) and Bird's Eye View (BEV) representation have advanced visual perception substantially, yet their internal spatial representations assume the rectilinear geometry of pinhole cameras. Fisheye cameras, widely deployed on...",
      "published": "2026-04-11",
      "updated": "2026-04-11",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 16,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 130,
      "arxiv_url": "https://arxiv.org/abs/2604.10391",
      "pdf_url": "https://arxiv.org/pdf/2604.10391.pdf"
    },
    {
      "id": "2604.09305",
      "anchor": "paper-2604-09305",
      "title": "VAGNet: Vision-based Accident Anticipation with Global Features",
      "authors": [
        "Vipooshan Vipulananthan",
        "Charith D. Chitraranjan"
      ],
      "abstract": "Traffic accidents are a leading cause of fatalities and injuries across the globe. Therefore, the ability to anticipate hazardous situations in advance is essential. Automated accident anticipation enables timely intervention through driver alerts and collision avoidance maneuvers, forming a key component of advanced driver assistance systems. In autonomous driving, such predictive capabilities support proactive safety behaviors, such as initiating defensive driving and human takeover when required. Using dashcam video as input offers a cost-effective solution, but it is challenging due to the complexity of real-world driving scenes. Accident anticipation systems need to operate in real-time. However, current methods involve extracting features from each detected object, which is computationally intensive. We propose VAGNet, a deep neural network that learns to predict accidents from dash-cam video using global features of traffic scenes without requiring explicit object-level features. The network consists of transformer and graph modules, and we use the vision foundation model VideoMAE-V2 for global feature extraction. Experiments on four benchmark datasets (DAD, DoTA, DADA, and Nexar) show that our method anticipates accidents with higher average precision and mean time-to-accident while being computationally more efficient compared to existing methods.",
      "short_abstract": "Traffic accidents are a leading cause of fatalities and injuries across the globe. Therefore, the ability to anticipate hazardous situations in advance is essential. Automated accident anticipation enables timely intervention through driver alerts and...",
      "published": "2026-04-10",
      "updated": "2026-04-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: abstract: takeover",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 131,
      "arxiv_url": "https://arxiv.org/abs/2604.09305",
      "pdf_url": "https://arxiv.org/pdf/2604.09305.pdf"
    },
    {
      "id": "2604.09232",
      "anchor": "paper-2604-09232",
      "title": "Neural Distribution Prior for LiDAR Out-of-Distribution Detection",
      "authors": [
        "Zizhao Li",
        "Zhengkang Xiang",
        "Jiayang Ao",
        "Feng Liu",
        "Joseph West",
        "Kourosh Khoshelham"
      ],
      "abstract": "LiDAR-based perception is critical for autonomous driving due to its robustness to poor lighting and visibility conditions. Yet, current models operate under the closed-set assumption and often fail to recognize unexpected out-of-distribution (OOD) objects in the open world. Existing OOD scoring functions exhibit limited performance because they ignore the pronounced class imbalance inherent in LiDAR OOD detection and assume a uniform class distribution. To address this limitation, we propose the Neural Distribution Prior (NDP), a framework that models the distributional structure of network predictions and adaptively reweights OOD scores based on alignment with a learned distribution prior. NDP dynamically captures the logit distribution patterns of training data and corrects class-dependent confidence bias through an attention-based module. We further introduce a Perlin noise-based OOD synthesis strategy that generates diverse auxiliary OOD samples from input scans, enabling robust OOD training without external datasets. Extensive experiments on the SemanticKITTI and STU benchmarks demonstrate that NDP substantially improves OOD detection performance, achieving a point-level AP of 61.31% on the STU test set, which is more than 10$\\times$ higher than the previous best result. Our framework is compatible with various existing OOD scoring formulations, providing an effective solution for open-world LiDAR perception.",
      "short_abstract": "LiDAR-based perception is critical for autonomous driving due to its robustness to poor lighting and visibility conditions. Yet, current models operate under the closed-set assumption and often fail to recognize unexpected out-of-distribution (OOD) objects in...",
      "published": "2026-04-10",
      "updated": "2026-04-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 13,
        "evidence": [
          "abstract: perception",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 131,
      "arxiv_url": "https://arxiv.org/abs/2604.09232",
      "pdf_url": "https://arxiv.org/pdf/2604.09232.pdf"
    },
    {
      "id": "2604.09206",
      "anchor": "paper-2604-09206",
      "title": "Long-SCOPE: Fully Sparse Long-Range Cooperative 3D Perception",
      "authors": [
        "Jiahao Wang",
        "Zikun Xu",
        "Yuner Zhang",
        "Zhongwei Jiang",
        "Chenyang Lu",
        "Shuocheng Yang",
        "Yuxuan Wang",
        "Jiaru Zhong",
        "Chuang Zhang",
        "Shaobing Xu",
        "Jianqiang Wang"
      ],
      "abstract": "Cooperative 3D perception via Vehicle-to-Everything communication is a promising paradigm for enhancing autonomous driving, offering extended sensing horizons and occlusion resolution. However, the practical deployment of existing methods is hindered at long distances by two critical bottlenecks: the quadratic computational scaling of dense BEV representations and the fragility of feature association mechanisms under significant observation and alignment errors. To overcome these limitations, we introduce Long-SCOPE, a fully sparse framework designed for robust long-distance cooperative 3D perception. Our method features two novel components: a Geometry-guided Query Generation module to accurately detect small, distant objects, and a learnable Context-Aware Association module that robustly matches cooperative queries despite severe positional noise. Experiments on the V2X-Seq and Griffin datasets validate that Long-SCOPE achieves state-of-the-art performance, particularly in challenging 100-150 m long-range settings, while maintaining highly competitive computation and communication costs.",
      "short_abstract": "Cooperative 3D perception via Vehicle-to-Everything communication is a promising paradigm for enhancing autonomous driving, offering extended sensing horizons and occlusion resolution. However, the practical deployment of existing methods is hindered at long...",
      "published": "2026-04-10",
      "updated": "2026-04-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 9,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 131,
      "arxiv_url": "https://arxiv.org/abs/2604.09206",
      "pdf_url": "https://arxiv.org/pdf/2604.09206.pdf"
    },
    {
      "id": "2604.09059",
      "anchor": "paper-2604-09059",
      "title": "Learning Vision-Language-Action World Models for Autonomous Driving",
      "authors": [
        "Guoqing Wang",
        "Pin Tang",
        "Xiangxuan Ren",
        "Guodongfang Zhao",
        "Bailan Feng",
        "Chao Ma"
      ],
      "abstract": "Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and control within a unified multimodal framework. However, they often lack explicit modeling of temporal dynamics and global world consistency, which limits their foresight and safety. In contrast, world models can simulate plausible future scenes but generally struggle to reason about or evaluate the imagined future they generate. In this work, we present VLA-World, a simple yet effective VLA world model that unifies predictive imagination with reflective reasoning to improve driving foresight. VLA-World first uses an action-derived feasible trajectory to guide the generation of the next-frame image, capturing rich spatial and temporal cues that describe how the surrounding environment evolves. The model then reasons over this self-generated future imagined frame to refine the predicted trajectory, achieving higher performance and better interpretability. To support this pipeline, we curate nuScenes-GR-20K, a generative reasoning dataset derived from nuScenes, and employ a three-stage training strategy that includes pretraining, supervised fine-tuning, and reinforcement learning. Extensive experiments demonstrate that VLA-World consistently surpasses state-of-the-art VLA and world-model baselines on both planning and future-generation benchmarks. Project page: https://vlaworld.github.io",
      "short_abstract": "Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and control within a unified multimodal framework. However, they often lack explicit modeling of temporal...",
      "published": "2026-04-10",
      "updated": "2026-04-10",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "VLA",
        "World Model",
        "Reinforcement Learning",
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 17,
        "evidence": [
          "abstract: end-to-end driving",
          "title: vision-language-action",
          "abstract: vision-language-action",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 131,
      "arxiv_url": "https://arxiv.org/abs/2604.09059",
      "pdf_url": "https://arxiv.org/pdf/2604.09059.pdf"
    },
    {
      "id": "2604.13098",
      "anchor": "paper-2604-13098",
      "title": "C$^2$T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination",
      "authors": [
        "Yuyang Chen",
        "Kaiyan Zhao",
        "Yiming Wang",
        "Ming Yang",
        "Bin Rao",
        "Zhenning Li"
      ],
      "abstract": "State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (TLCs) and Connected Autonomous Vehicles (CAVs). However, the performance of these systems is fundamentally capped by their hand-crafted, myopic rewards (e.g., intersection pressure), which fail to capture high-level, human-centric goals like safety, flow stability, and comfort. To overcome this limitation, we introduce C2T, a novel framework that learns a common-sense coordination model from traffic-vehicle dynamics. C2T distills \"common-sense\" knowledge from a Large Language Model (LLM) into a learned intrinsic reward function. This new reward is then used to guide the coordination policy of a cooperative multi-intersection TLC MARL system on CityFlow-based multi-intersection benchmarks. Our framework significantly outperforms strong MARL baselines in traffic efficiency, safety, and an energy-related proxy. We further highlight C2T's flexibility in principle, allowing distinct \"efficiency-focused\" versus \"safety-focused\" policies by modifying the LLM prompt.",
      "short_abstract": "State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (TLCs) and Connected Autonomous Vehicles (CAVs). However, the performance of these systems is fundamentally...",
      "published": "2026-04-10",
      "updated": "2026-04-10",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 3,
        "evidence": [
          "abstract: vehicle dynamics",
          "abstract: controller",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 131,
      "arxiv_url": "https://arxiv.org/abs/2604.13098",
      "pdf_url": "https://arxiv.org/pdf/2604.13098.pdf"
    },
    {
      "id": "2604.09351",
      "anchor": "paper-2604-09351",
      "title": "Decentralized Opinion-Integrated Decision making at Unsignalized Intersections via Signed Networks",
      "authors": [
        "Bhaskar Varma",
        "Ying Shuai Quan",
        "Karl D. von Ellenrieder",
        "Paolo Falcone"
      ],
      "abstract": "In this letter, we consider the problem of decentralized decision making among connected autonomous vehicles at unsignalized intersections, where existing centralized approaches do not scale gracefully under mixed maneuver intentions and coordinator failure. We propose a closed-loop opinion-dynamic decision model for intersection coordination, where vehicles exchange intent through dual signed networks: a conflict topology based communication network and a commitment-driven belief network that enable cooperation without a centralized coordinator. Continuous opinion states modulate velocity optimizer weights prior to commitment; a closed-form predictive feasibility gate then freezes each vehicle's decision into a GO or YIELD commitment, which propagates back through the belief network to pre-condition neighbor behavior ahead of physical conflicts. Crossing order emerges from geometric feasibility and arrival priority without the use of joint optimization or a solver. The approach is validated across three scenarios spanning fully competitive, merge, and mixed conflict topologies. The results demonstrate collision-free coordination and lower last-vehicle exit times compared to first come first served (FCFS) in all conflict non-trivial configurations.",
      "short_abstract": "In this letter, we consider the problem of decentralized decision making among connected autonomous vehicles at unsignalized intersections, where existing centralized approaches do not scale gracefully under mixed maneuver intentions and coordinator failure....",
      "published": "2026-04-10",
      "updated": "2026-04-10",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: decision-making",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 131,
      "arxiv_url": "https://arxiv.org/abs/2604.09351",
      "pdf_url": "https://arxiv.org/pdf/2604.09351.pdf"
    },
    {
      "id": "2604.08827",
      "anchor": "paper-2604-08827",
      "title": "Quantum Patches: Enhancing Robustness of Quantum Machine Learning Models",
      "authors": [
        "Ban Q. Tran",
        "Chuong K. Luong",
        "Viet Q. Nguyen",
        "Duong M. Chu",
        "Susan Mengel"
      ],
      "abstract": "Machine learning models and their applications, such as autonomous driving systems, are becoming increasingly common and are essential components of human daily life. However, due to their sensitivity to perturbed noise, these models are easily susceptible to adversarial attacks. Not only are classical machine learning models affected, but quantum machine learning (QML) models have also been proven to be vulnerable to adversarial attacks, which degrade their performance. To defend against these types of attacks, several classical methods have been proposed. Among these, a prominent approach uses various types of pseudo-noise during training to enhance the model's robustness against real-world attacks. One of the recently emerging solutions is to leverage the unique properties of quantum circuits to create quantum-based pseudo-noise similar to real perturbed noise to counter adversarial attacks. This paper proposes a solution that utilizes random quantum circuits (RQCs) as adversarial data to help QML models overcome these adversarial attacks. The results reported in this paper show that the data generated by RQC actually provides a similar effect to models trained with adversarial data on high-feature datasets. This quantum-based pseudo-noise resulted in a significant reduction in the attack rate in the CIFAR-10 data set, from \\textbf{89. 8\\%} to \\textbf{68.45\\%}. For the CINIC-10 dataset, the successful attack rate decreased from \\textbf{94.23\\%} to \\textbf{78.68\\%}. This research opens up avenues for applying unique quantum properties, such as superposition, entanglement, and even decoherence, to enhance the quality of machine learning models.",
      "short_abstract": "Machine learning models and their applications, such as autonomous driving systems, are becoming increasingly common and are essential components of human daily life. However, due to their sensitivity to perturbed noise, these models are easily susceptible to...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "abstract: attack or threat",
          "title: robustness",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08827",
      "pdf_url": "https://arxiv.org/pdf/2604.08827.pdf"
    },
    {
      "id": "2604.08719",
      "anchor": "paper-2604-08719",
      "title": "LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving",
      "authors": [
        "Hao Shao",
        "Letian Wang",
        "Yang Zhou",
        "Yuxuan Hu",
        "Zhuofan Zong",
        "Steven L. Waslander",
        "Wei Zhan",
        "Hongsheng Li"
      ],
      "abstract": "Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck for large-scale deployment. To address this challenge, some works use LLMs and VLMs for vision-language understanding and reasoning, enabling vehicles to interpret rare and safety-critical situations when generating actions. Others study generative world models to capture the spatio-temporal evolution of driving scenes, allowing agents to imagine possible futures before acting. Inspired by human intelligence, which unifies understanding and imagination, we explore a unified model for autonomous driving. We present LMGenDrive, the first framework that combines LLM-based multimodal understanding with generative world models for end-to-end closed-loop driving. Given multi-view camera inputs and natural-language instructions, LMGenDrive generates both future driving videos and control signals. This design provides complementary benefits: video prediction improves spatio-temporal scene modeling, while the LLM contributes strong semantic priors and instruction grounding from large-scale pretraining. We further propose a progressive three-stage training strategy, from vision pretraining to multi-step long-horizon driving, to improve stability and performance. LMGenDrive supports both low-latency online planning and autoregressive offline video generation. Experiments show that it significantly outperforms prior methods on challenging closed-loop benchmarks, with clear gains in instruction following, spatio-temporal understanding, and robustness to rare scenarios. These results suggest that unifying multimodal understanding and generation is a promising direction for more generalizable and robust embodied decision-making systems.",
      "short_abstract": "Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck for large-scale deployment. To address this challenge, some works use LLMs and VLMs for vision-language...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 23,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08719",
      "pdf_url": "https://arxiv.org/pdf/2604.08719.pdf"
    },
    {
      "id": "2604.08535",
      "anchor": "paper-2604-08535",
      "title": "Fail2Drive: Benchmarking Closed-Loop Driving Generalization",
      "authors": [
        "Simon Gerstenecker",
        "Andreas Geiger",
        "Katrin Renz"
      ],
      "abstract": "Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe and scalable testing, existing benchmarks rarely measure true generalization: they typically reuse training scenarios at test time. Success can therefore reflect memorization rather than robust driving behavior. We introduce Fail2Drive, the first paired-route benchmark for closed-loop generalization in CARLA, with 200 routes and 17 new scenario classes spanning appearance, layout, behavioral, and robustness shifts. Each shifted route is matched with an in-distribution counterpart, isolating the effect of the shift and turning qualitative failures into quantitative diagnostics. Evaluating multiple state-of-the-art models reveals consistent degradation, with an average success-rate drop of 22.8\\%. Our analysis uncovers unexpected failure modes, such as ignoring objects clearly visible in the LiDAR and failing to learn the fundamental concepts of free and occupied space. To accelerate follow-up work, Fail2Drive includes an open-source toolbox for creating new scenarios and validating solvability via a privileged expert policy. Together, these components establish a reproducible foundation for benchmarking and improving closed-loop driving generalization. We open-source all code, data, and tools at https://github.com/autonomousvision/fail2drive .",
      "short_abstract": "Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe and scalable testing, existing benchmarks rarely measure true generalization: they typically reuse training...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Benchmark",
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 3,
        "evidence": [
          "abstract: validation or testing",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08535",
      "pdf_url": "https://arxiv.org/pdf/2604.08535.pdf"
    },
    {
      "id": "2604.08457",
      "anchor": "paper-2604-08457",
      "title": "CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning",
      "authors": [
        "Rui Gan",
        "Junyi Ma",
        "Pei Li",
        "Xingyou Yang",
        "Kai Chen",
        "Sikai Chen",
        "Bin Ran"
      ],
      "abstract": "Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical traffic scenarios remains insufficiently evaluated due to the ego-vehicle focus of existing benchmarks. To bridge this gap, we present \\textbf{CrashSight}, a large-scale vision-language benchmark for roadway crash understanding using real-world roadside camera data. The dataset comprises 250 crash videos, annotated with 13K multiple-choice question-answer pairs organized under a two-tier taxonomy. Tier 1 evaluates the visual grounding of scene context and involved parties, while Tier 2 probes higher-level reasoning, including crash mechanics, causal attribution, temporal progression, and post-crash outcomes. We benchmark 8 state-of-the-art VLMs and show that, despite strong scene description capabilities, current models struggle with temporal and causal reasoning in safety-critical scenarios. We provide a detailed analysis of failure scenarios and discuss directions for improving VLM crash understanding. The benchmark provides a standardized evaluation framework for infrastructure-assisted perception in cooperative autonomous driving. The CrashSight benchmark, including the full dataset and code, is accessible at https://mcgrche.github.io/crashsight.",
      "short_abstract": "Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical traffic scenarios...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 11,
        "evidence": [
          "abstract: perception",
          "abstract: camera",
          "title: scene understanding",
          "abstract: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08457",
      "pdf_url": "https://arxiv.org/pdf/2604.08457.pdf"
    },
    {
      "id": "2604.08366",
      "anchor": "paper-2604-08366",
      "title": "Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems",
      "authors": [
        "Tolga Dimlioglu",
        "Nadine Chang",
        "Maying Shen",
        "Rafid Mahmood",
        "Jose M. Alvarez"
      ],
      "abstract": "Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, the training data, must address different evaluation criteria necessary for the models to be deployable in real-world environments. Data selection policies can guide the development of the training set, but current frameworks do not account for the ambiguity in how data points affect different metrics. In this work, we propose Mixture Optimization via Scaling-Aware Iterative Collection (MOSAIC), a general data selection framework that operates by: (i) partitioning the dataset into domains; (ii) fitting neural scaling laws from each data domain to the evaluation metrics; and (iii) optimizing a data mixture by iteratively adding data from domains that maximize the change in metrics. We apply MOSAIC to autonomous driving (AD), where an End-to-End (E2E) planner model is evaluated on the Extended Predictive Driver Model Score (EPDMS), an aggregate of driving rule compliance metrics. Here, MOSAIC outperforms a diverse set of baselines on EPDMS with up to 80\\% less data.",
      "short_abstract": "Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, the training data, must address different evaluation criteria necessary for the models to be deployable in...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 30,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08366",
      "pdf_url": "https://arxiv.org/pdf/2604.08366.pdf"
    },
    {
      "id": "2604.08266",
      "anchor": "paper-2604-08266",
      "title": "Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models",
      "authors": [
        "Jing Gu",
        "Niccolò Cavagnero",
        "Gijs Dubbelman"
      ],
      "abstract": "Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems to handle rare and complex scenarios. While integrating LLMs into Vision-Language-Action (VLA) models has yielded state-of-the-art performance, their massive parameter counts pose severe challenges for latency-sensitive and energy-efficient deployment. Distilling LLM knowledge into a compact driving model offers a compelling solution to retain these reasoning capabilities while maintaining a manageable computational footprint. Although previous works have demonstrated the efficacy of distillation, these efforts have primarily focused on relatively simple scenarios and open-loop evaluations. Therefore, in this work, we investigate LLM distillation in more complex, interactive scenarios under closed-loop evaluation. We demonstrate that through a combination of latent feature distillation and ground-truth trajectory supervision, an efficient vision-only student model \\textbf{Orion-Lite} can even surpass the performance of its massive VLA teacher, ORION. Setting a new state-of-the-art on the rigorous Bench2Drive benchmark, with a Driving Score of 80.6. Ultimately, this reveals that vision-only architectures still possess significant, untapped potential for high-performance reactive planning.",
      "short_abstract": "Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems to handle rare and complex scenarios. While integrating LLMs into Vision-Language-Action (VLA) models has...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: vision-language-action"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08266",
      "pdf_url": "https://arxiv.org/pdf/2604.08266.pdf"
    },
    {
      "id": "2604.08074",
      "anchor": "paper-2604-08074",
      "title": "DinoRADE: Full Spectral Radar-Camera Fusion with Vision Foundation Model Features for Multi-class Object Detection in Adverse Weather",
      "authors": [
        "Christof Leitgeb",
        "Thomas Puchleitner",
        "Max Peter Ronecker",
        "Daniel Watzenig"
      ],
      "abstract": "Reliable and weather-robust perception systems are essential for safe autonomous driving and typically employ multi-modal sensor configurations to achieve comprehensive environmental awareness. While recent automotive FMCW Radar-based approaches achieved remarkable performance on detection tasks in adverse weather conditions, they exhibited limitations in resolving fine-grained spatial details particularly critical for detecting smaller and vulnerable road users (VRUs). Furthermore, existing research has not adequately addressed VRU detection in adverse weather datasets such as K-Radar. We present DinoRADE, a Radar-centered detection pipeline that processes dense Radar tensors and aggregates vision features around transformed reference points in the camera perspective via deformable cross-attention. Vision features are provided by a DINOv3 Vision Foundation Model. We present a comprehensive performance evaluation on the K-Radar dataset in all weather conditions and are among the first to report detection performance individually for five object classes. Additionally, we compare our method with existing single-class detection approaches and outperform recent Radar-camera approaches by 12.1%. The code is available under https://github.com/chr-is-tof/RADE-Net.",
      "short_abstract": "Reliable and weather-robust perception systems are essential for safe autonomous driving and typically employ multi-modal sensor configurations to achieve comprehensive environmental awareness. While recent automotive FMCW Radar-based approaches achieved...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 33,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08074",
      "pdf_url": "https://arxiv.org/pdf/2604.08074.pdf"
    },
    {
      "id": "2604.08031",
      "anchor": "paper-2604-08031",
      "title": "Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles",
      "authors": [
        "Jiawei Liu",
        "Xun Gong",
        "Fen Fang",
        "Muli Yang",
        "Bohao Qu",
        "Yunfeng Hu",
        "Hong Chen",
        "Xulei Yang",
        "Qing Guo"
      ],
      "abstract": "Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals, without sacrificing interpretability and traceability, remains a challenge. This study proposes an instruction-realization framework that leverages a large language model (LLM) to interpret instructions, generates executable scripts that schedule multiple model predictive control (MPC)-based motion planners based on real-time feedback, and converts planned trajectories into control signals. This scheduling-centric design decouples semantic reasoning from vehicle control at different timescales, establishing a transparent, traceable decision-making chain from high-level instructions to low-level actions. Due to the absence of high-fidelity evaluation tools, this study introduces a benchmark for open-ended instruction realization in a closed-loop setting. Comprehensive experiments reveal that the framework significantly improves task-completion rates over instruction-realization baselines, reduces LLM query costs, achieves safety and compliance on par with specialized AD approaches, and exhibits considerable tolerance to LLM inference latency. For more qualitative illustrations and a clearer understanding.",
      "short_abstract": "Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals, without sacrificing...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark",
        "Hardware / Real-Time",
        "Explainability",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 11,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: model predictive control",
          "Control & Vehicle Dynamics: abstract: vehicle control",
          "Control & Vehicle Dynamics: abstract: control",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: title: systems constraints",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08031",
      "pdf_url": "https://arxiv.org/pdf/2604.08031.pdf"
    },
    {
      "id": "2604.08008",
      "anchor": "paper-2604-08008",
      "title": "SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving",
      "authors": [
        "Felix Embacher",
        "Jonas Uhrig",
        "Marius Cordts",
        "Markus Enzweiler"
      ],
      "abstract": "Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dataset sizes continue to grow, the key challenge shifts from collecting more data to efficiently identifying the most relevant samples. We introduce SearchAD, a large-scale rare image retrieval dataset for AD containing over 423k frames drawn from 11 established datasets. SearchAD provides high-quality manual annotations of more than 513k bounding boxes covering 90 rare categories. It specifically targets the needle-in-a-haystack problem of locating extremely rare classes, with some appearing fewer than 50 times across the entire dataset. Unlike existing benchmarks, which focused on instance-level retrieval, SearchAD emphasizes semantic image retrieval with a well-defined data split, enabling text-to-image and image-to-image retrieval, few-shot learning, and fine-tuning of multi-modal retrieval models. Comprehensive evaluations show that text-based methods outperform image-based ones due to stronger inherent semantic grounding. While models directly aligning spatial visual features with language achieve the best zero-shot results, and our fine-tuning baseline significantly improves performance, absolute retrieval capabilities remain unsatisfactory. With a held-out test set on a public benchmark server, SearchAD establishes the first large-scale dataset for retrieval-driven data curation and long-tail perception research in AD: https://iis-esslingen.github.io/searchad/",
      "short_abstract": "Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dataset sizes continue to grow, the key challenge shifts from collecting more data to efficiently identifying...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08008",
      "pdf_url": "https://arxiv.org/pdf/2604.08008.pdf"
    },
    {
      "id": "2604.07991",
      "anchor": "paper-2604-07991",
      "title": "MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models",
      "authors": [
        "Zile Guo",
        "Zhan Chen",
        "Enze Zhu",
        "Kan Wei",
        "Yongkang Zou",
        "Xiaoxuan Liu",
        "Lei Wang"
      ],
      "abstract": "Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex 3D dynamics is essential for autonomous navigation and robust decision-making in unconstrained environments. However, under the highly dynamic camera trajectories typical of UAV views, existing world models often struggle to maintain spatiotemporal physical consistency. A key reason lies in the distribution bias of current training data: most existing datasets exhibit restricted 2.5D motion patterns, such as ground-constrained autonomous driving scenes or relatively smooth human-centric egocentric videos, and therefore lack realistic high-dynamic 6-DoF UAV motion priors. To address this gap, we present MotionScape, a large-scale real-world UAV-view video dataset with highly dynamic motion for world modeling. MotionScape contains over 30 hours of 4K UAV-view videos, totaling more than 4.5M frames. This novel dataset features semantically and geometrically aligned training samples, where diverse real-world UAV videos are tightly coupled with accurate 6-DoF camera trajectories and fine-grained natural language descriptions. To build the dataset, we develop an automated multi-stage processing pipeline that integrates CLIP-based relevance filtering, temporal segmentation, robust visual SLAM for trajectory recovery, and large-language-model-driven semantic annotation. Extensive experiments show that incorporating such semantically and geometrically aligned annotations effectively improves the ability of existing world models to simulate complex 3D dynamics and handle large viewpoint shifts, thereby benefiting decision-making and planning for UAV agents in complex environments. The dataset is publicly available at https://github.com/Thelegendzz/MotionScape",
      "short_abstract": "Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex 3D dynamics is...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Dataset",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 9,
        "evidence": [
          "title: world model",
          "abstract: world model",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.07991",
      "pdf_url": "https://arxiv.org/pdf/2604.07991.pdf"
    },
    {
      "id": "2604.07980",
      "anchor": "paper-2604-07980",
      "title": "Object-Centric Stereo Ranging for Autonomous Driving: From Dense Disparity to Census-Based Template Matching",
      "authors": [
        "Qihao Huang"
      ],
      "abstract": "Accurate depth estimation is critical for autonomous driving perception systems, particularly for long range vehicle detection on highways. Traditional dense stereo matching methods such as Block Matching (BM) and Semi Global Matching (SGM) produce per pixel disparity maps but suffer from high computational cost, sensitivity to radiometric differences between stereo cameras, and poor accuracy at long range where disparity values are small. In this report, we present a comprehensive stereo ranging system that integrates three complementary depth estimation approaches: dense BM/SGM disparity, object centric Census based template matching, and monocular geometric priors, within a unified detection ranging tracking pipeline. Our key contribution is a novel object centric Census based template matching algorithm that performs GPU accelerated sparse stereo matching directly within detected bounding boxes, employing a far close divide and conquer strategy, forward backward verification, occlusion aware sampling, and robust multi block aggregation. We further describe an online calibration refinement framework that combines auto rectification offset search, radar stereo voting based disparity correction, and object level radar stereo association for continuous extrinsic drift compensation. The complete system achieves real time performance through asynchronous GPU pipeline design and delivers robust ranging across diverse driving conditions including nighttime, rain, and varying illumination.",
      "short_abstract": "Accurate depth estimation is critical for autonomous driving perception systems, particularly for long range vehicle detection on highways. Traditional dense stereo matching methods such as Block Matching (BM) and Semi Global Matching (SGM) produce per pixel...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 4,
        "evidence": [
          "abstract: perception",
          "abstract: radar",
          "abstract: camera",
          "abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.07980",
      "pdf_url": "https://arxiv.org/pdf/2604.07980.pdf"
    },
    {
      "id": "2604.07944",
      "anchor": "paper-2604-07944",
      "title": "On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning",
      "authors": [
        "Amirhossein Afsharrad",
        "Amirhesam Abedsoltan",
        "Ahmadreza Moradipari",
        "Sanjay Lall"
      ],
      "abstract": "Large language models (LLMs) have recently demonstrated strong potential for autonomous vehicle motion planning by reformulating trajectory prediction as a language generation problem. However, deploying capable LLMs in resource-constrained onboard systems remains a fundamental challenge. In this paper, we study how to effectively transfer motion planning knowledge from a large teacher LLM to a smaller, more deployable student model. We build on the GPT-Driver framework, which represents driving scenes as language prompts and generates waypoint trajectories with chain-of-thought reasoning, and investigate two student training paradigms: (i) on-policy generalized knowledge distillation (GKD), which trains the student on its own self-generated outputs using dense token-level feedback from the teacher, and (ii) a dense-feedback reinforcement learning (RL) baseline that uses the teacher's log-probabilities as per-token reward signals in a policy gradient framework. Experiments on the nuScenes benchmark show that GKD substantially outperforms the RL baseline and closely approaches teacher-level performance despite a 5$\\times$ reduction in model size. These results highlight the practical value of on-policy distillation as a principled and effective approach to deploying LLM-based planners in autonomous driving systems.",
      "short_abstract": "Large language models (LLMs) have recently demonstrated strong potential for autonomous vehicle motion planning by reformulating trajectory prediction as a language generation problem. However, deploying capable LLMs in resource-constrained onboard systems...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 16,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.07944",
      "pdf_url": "https://arxiv.org/pdf/2604.07944.pdf"
    },
    {
      "id": "2604.07912",
      "anchor": "paper-2604-07912",
      "title": "ParkSense: Where Should a Delivery Driver Park? Leveraging Idle AV Compute and Vision-Language Models",
      "authors": [
        "Die Hu",
        "Henan Li"
      ],
      "abstract": "Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant entrances. We propose ParkSense, a framework that repurposes idle compute during low-risk AV states -- queuing at red lights, traffic congestion, parking-lot crawl -- to run a Vision-Language Model (VLM) on pre-cached satellite and street view imagery, identifying entrances and legal parking zones. We formalize the Delivery-Aware Precision Parking (DAPP) problem, show that a quantized 7B VLM completes inference in 4-8 seconds on HW4-class hardware, and estimate annual per-driver income gains of 3,000-8,000 USD in the U.S. Five open research directions are identified at this unexplored intersection of autonomous driving, computer vision, and last-mile logistics.",
      "short_abstract": "Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant entrances. We propose ParkSense, a framework that repurposes idle compute during low-risk AV states -- queuing...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy",
          "Systems, Deployment & Connectivity: abstract: hardware"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.07912",
      "pdf_url": "https://arxiv.org/pdf/2604.07912.pdf"
    },
    {
      "id": "2604.08689",
      "anchor": "paper-2604-08689",
      "title": "An Energy-Efficient Lyapunov-Based Cooperative Adaptive Cruise Controller for Electric Vehicles",
      "authors": [
        "Hamed Faghihian",
        "Parisa Ansari Bonab",
        "Arman Sargolzaei"
      ],
      "abstract": "As electric vehicles (EVs) are increasingly adopted as platforms for connected and automated vehicles (CAVs), enhancing their energy efficiency becomes critical. With the emergence of vehicle-to-vehicle (V2V) communication, cooperative adaptive cruise control (CACC) offers improved traffic flow, safety, and energy efficiency by enabling real-time coordination among EVs. However, conventional CACC algorithms neglected acceleration and regenerative braking dynamics in their implementation. To address this gap, this paper proposes a third-order dynamic model for EVs which has been derived from real-world experimental data. We also propose a novel, practical, and energy-efficient Lyapunov-based CACC controller explicitly designed for EV platoons. The proposed controller is requiring lower control gains while ensuring string stability and energy efficiency. To validate its effectiveness, we conduct both simulation and experimental environments, demonstrating that our approach reduces velocity fluctuations, maintains string stability at lower headway times, and improves energy efficiency of the CACC platoon by up to 38.5% compared to a baseline CACC.",
      "short_abstract": "As electric vehicles (EVs) are increasingly adopted as platforms for connected and automated vehicles (CAVs), enhancing their energy efficiency becomes critical. With the emergence of vehicle-to-vehicle (V2V) communication, cooperative adaptive cruise control...",
      "published": "2026-04-09",
      "updated": "2026-04-09",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 11,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 132,
      "arxiv_url": "https://arxiv.org/abs/2604.08689",
      "pdf_url": "https://arxiv.org/pdf/2604.08689.pdf"
    },
    {
      "id": "2604.07250",
      "anchor": "paper-2604-07250",
      "title": "Geo-EVS: Geometry-Conditioned Extrapolative View Synthesis for Autonomous Driving",
      "authors": [
        "Yatong Lan",
        "Rongkui Tang",
        "Lei He"
      ],
      "abstract": "Extrapolative novel view synthesis can reduce camera-rig dependency in autonomous driving by generating standardized virtual views from heterogeneous sensors. Existing methods degrade outside recorded trajectories because extrapolated poses provide weak geometric support and no dense target-view supervision. The key is to explicitly expose the model to out-of-trajectory condition defects during training. We propose Geo-EVS, a geometry-conditioned framework under sparse supervision. Geo-EVS has two components. Geometry-Aware Reprojection (GAR) uses fine-tuned VGGT to reconstruct colored point clouds and reproject them to observed and virtual target poses, producing geometric condition maps. This design unifies the reprojection path between training and inference. Artifact-Guided Latent Diffusion (AGLD) injects reprojection-derived artifact masks during training so the model learns to recover structure under missing support. For evaluation, we use a LiDAR-Projected Sparse-Reference (LPSR) protocol when dense extrapolated-view ground truth is unavailable. On Waymo, Geo-EVS improves sparse-view synthesis quality and geometric accuracy, especially in high-angle and low-coverage settings. It also improves downstream 3D detection.",
      "short_abstract": "Extrapolative novel view synthesis can reduce camera-rig dependency in autonomous driving by generating standardized virtual views from heterogeneous sensors. Existing methods degrade outside recorded trajectories because extrapolated poses provide weak...",
      "published": "2026-04-08",
      "updated": "2026-04-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 8,
        "evidence": [
          "abstract: point cloud",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 133,
      "arxiv_url": "https://arxiv.org/abs/2604.07250",
      "pdf_url": "https://arxiv.org/pdf/2604.07250.pdf"
    },
    {
      "id": "2604.07126",
      "anchor": "paper-2604-07126",
      "title": "RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Prediction",
      "authors": [
        "Diyi Liu",
        "Zihan Niu",
        "Tu Xu",
        "Xingchen Zhang",
        "Lishan Sun"
      ],
      "abstract": "Predicting traffic agent trajectories plays an important role in autonomous driving, traffic operations, transportation safety analysis, etc. Although many deep learning algorithms are devised to predict future agent trajectories, the trajectory prediction problem is still challenging due to the complexity of decision-making process, interactions with surrounding vehicles, and the existence of multiple possible intentions for the traveling agents even under similar scenarios. Most existing methods are limited by the requirement of graph structures (e.g., Graph Neural Network) or the requirement of manually labeled intentions. In this study, we propose a pure Transformer-based deep learning model for multi-modal trajectory prediction considering temporal dependencies and agent-agent spatial interactions. After encoding the historical trajectories, two parallel decoders are employed to generate trajectories and probabilities on separate decoder tracks. The model is evaluated on two real-world datasets, one highway dataset and the other pedestrian dataset with solid performance. One important insight is that following a ``propose-then-select'' strategy, the agent-agent spatial interactions are only considered for probability estimation instead of trajectory generation. In summary, the proposed model provides a potential direction to design more robust and effective multi-modal trajectory prediction models.",
      "short_abstract": "Predicting traffic agent trajectories plays an important role in autonomous driving, traffic operations, transportation safety analysis, etc. Although many deep learning algorithms are devised to predict future agent trajectories, the trajectory prediction...",
      "published": "2026-04-08",
      "updated": "2026-04-08",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 21,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 133,
      "arxiv_url": "https://arxiv.org/abs/2604.07126",
      "pdf_url": "https://arxiv.org/pdf/2604.07126.pdf"
    },
    {
      "id": "2604.06832",
      "anchor": "paper-2604-06832",
      "title": "Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM",
      "authors": [
        "Chengyue Wu",
        "Shiyi Lan",
        "Yonggan Fu",
        "Sensen Gao",
        "Jin Wang",
        "Jincheng Yu",
        "Jose M. Alvarez",
        "Pavlo Molchanov",
        "Ping Luo",
        "Song Han",
        "Ligeng Zhu",
        "Enze Xie"
      ],
      "abstract": "Vision-language models (VLMs) predominantly rely on autoregressive decoding, which generates tokens one at a time and fundamentally limits inference throughput. This limitation is especially acute in physical AI scenarios such as robotics and autonomous driving, where VLMs are deployed on edge devices at batch size one, making AR decoding memory-bandwidth-bound and leaving hardware parallelism underutilized. While block-wise discrete diffusion has shown promise for parallel text generation, extending it to VLMs remains challenging due to the need to jointly handle continuous visual representations and discrete text tokens while preserving pretrained multimodal capabilities. We present Fast-dVLM, a block-diffusion-based VLM that enables KV-cache-compatible parallel decoding and speculative block decoding for inference acceleration. We systematically compare two AR-to-diffusion conversion strategies: a two-stage approach that first adapts the LLM backbone with text-only diffusion fine-tuning before multimodal training, and a direct approach that converts the full AR VLM in one stage. Under comparable training budgets, direct conversion proves substantially more efficient by leveraging the already multimodally aligned VLM; we therefore adopt it as our recommended recipe. We introduce a suite of multimodal diffusion adaptations, block size annealing, causal context attention, auto-truncation masking, and vision efficient concatenation, that collectively enable effective block diffusion in the VLM setting. Extensive experiments across 11 multimodal benchmarks show Fast-dVLM matches its autoregressive counterpart in generation quality. With SGLang integration and FP8 quantization, Fast-dVLM achieves over 6x end-to-end inference speedup over the AR baseline.",
      "short_abstract": "Vision-language models (VLMs) predominantly rely on autoregressive decoding, which generates tokens one at a time and fundamentally limits inference throughput. This limitation is especially acute in physical AI scenarios such as robotics and autonomous...",
      "published": "2026-04-08",
      "updated": "2026-04-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: hardware",
          "Systems, Deployment & Connectivity: abstract: systems constraints",
          "End-to-End & VLA: abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 133,
      "arxiv_url": "https://arxiv.org/abs/2604.06832",
      "pdf_url": "https://arxiv.org/pdf/2604.06832.pdf"
    },
    {
      "id": "2604.06750",
      "anchor": "paper-2604-06750",
      "title": "How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study",
      "authors": [
        "Roberto Brusnicki",
        "Mattia Piccinini",
        "Johannes Betz"
      ],
      "abstract": "Vision-Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poorly characterized, particularly regarding how input configurations affect their capabilities. We introduce VENUSS (VLM Evaluation oN Understanding Sequential Scenes), a framework for systematic sensitivity analysis of VLM performance on sequential driving scenes, establishing baselines for future research. Building upon existing datasets, VENUSS extracts temporal sequences from driving videos, and generates structured evaluations across custom categories. By comparing 25+ existing VLMs across 2,600+ scenarios, we reveal how even top models achieve only 57% accuracy, not matching human performance under similar constraints (65%) and exposing significant capability gaps. Our analysis shows that VLMs excel with static object detection but struggle with understanding vehicle dynamics and temporal relations. VENUSS offers the first systematic sensitivity analysis of VLMs focused on how input image configurations - resolution, frame count, temporal intervals, spatial layouts, and presentation modes - affect performance on sequential driving scenes. Supplementary material available at https://TUM-AVS.github.io/VENUSS/.",
      "short_abstract": "Vision-Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poorly characterized, particularly regarding how input configurations affect their capabilities. We introduce...",
      "published": "2026-04-08",
      "updated": "2026-04-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: vehicle dynamics",
          "Perception & Sensor Fusion: abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 133,
      "arxiv_url": "https://arxiv.org/abs/2604.06750",
      "pdf_url": "https://arxiv.org/pdf/2604.06750.pdf"
    },
    {
      "id": "2604.06665",
      "anchor": "paper-2604-06665",
      "title": "VDPP: Video Depth Post-Processing for Speed and Scalability",
      "authors": [
        "Daewon Yoon",
        "Injun Baek",
        "Sangyu Han",
        "Yearim Kim",
        "Nojun Kwak"
      ],
      "abstract": "Video depth estimation is essential for providing 3D scene structure in applications ranging from autonomous driving to mixed reality. Current end-to-end video depth models have established state-of-the-art performance. Although current end-to-end (E2E) models have achieved state-of-the-art performance, they function as tightly coupled systems that suffer from a significant adaptation lag whenever superior single-image depth estimators are released. To mitigate this issue, post-processing methods such as NVDS offer a modular plug-and-play alternative to incorporate any evolving image depth model without retraining. However, existing post-processing methods still struggle to match the efficiency and practicality of E2E systems due to limited speed, accuracy, and RGB reliance. In this work, we revitalize the role of post-processing by proposing VDPP (Video Depth Post-Processing), a framework that improves the speed and accuracy of post-processing methods for video depth estimation. By shifting the paradigm from computationally expensive scene reconstruction to targeted geometric refinement, VDPP operates purely on geometric refinements in low-resolution space. This design achieves exceptional speed (>43.5 FPS on NVIDIA Jetson Orin Nano) while matching the temporal coherence of E2E systems, with dense residual learning driving geometric representations rather than full reconstructions. Furthermore, our VDPP's RGB-free architecture ensures true scalability, enabling immediate integration with any evolving image depth model. Our results demonstrate that VDPP provides a superior balance of speed, accuracy, and memory efficiency, making it the most practical solution for real-time edge deployment. Our project page is at https://github.com/injun-baek/VDPP",
      "short_abstract": "Video depth estimation is essential for providing 3D scene structure in applications ranging from autonomous driving to mixed reality. Current end-to-end video depth models have established state-of-the-art performance. Although current end-to-end (E2E)...",
      "published": "2026-04-08",
      "updated": "2026-04-08",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 0,
        "evidence": [
          "End-to-End & VLA: abstract: end-to-end",
          "Perception & Sensor Fusion: abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 133,
      "arxiv_url": "https://arxiv.org/abs/2604.06665",
      "pdf_url": "https://arxiv.org/pdf/2604.06665.pdf"
    },
    {
      "id": "2604.07286",
      "anchor": "paper-2604-07286",
      "title": "CADENCE: Context-Adaptive Depth Estimation for Navigation and Computational Efficiency",
      "authors": [
        "Timothy K Johnsen",
        "Marco Levorato"
      ],
      "abstract": "Autonomous vehicles deployed in remote environments typically rely on embedded processors, compact batteries, and lightweight sensors. These hardware limitations conflict with the need to derive robust representations of the environment, which often requires executing computationally intensive deep neural networks for perception. To address this challenge, we present CADENCE, an adaptive system that dynamically scales the computational complexity of a slimmable monocular depth estimation network in response to navigation needs and environmental context. By closing the loop between perception fidelity and actuation requirements, CADENCE ensures high-precision computing is only used when mission-critical. We conduct evaluations on our released open-source testbed that integrates Microsoft AirSim with an NVIDIA Jetson Orin Nano. As compared to a state-of-the-art static approach, CADENCE decreases sensor acquisitions, power consumption, and inference latency by 9.67%, 16.1%, and 74.8%, respectively. The results demonstrate an overall reduction in energy expenditure by 75.0%, along with an increase in navigation accuracy by 7.43%.",
      "short_abstract": "Autonomous vehicles deployed in remote environments typically rely on embedded processors, compact batteries, and lightweight sensors. These hardware limitations conflict with the need to derive robust representations of the environment, which often requires...",
      "published": "2026-04-08",
      "updated": "2026-04-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 7,
        "evidence": [
          "abstract: perception",
          "title: depth estimation",
          "abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 133,
      "arxiv_url": "https://arxiv.org/abs/2604.07286",
      "pdf_url": "https://arxiv.org/pdf/2604.07286.pdf"
    },
    {
      "id": "2604.16452",
      "anchor": "paper-2604-16452",
      "title": "Compiling OpenSCENARIO 2.1 for Scenario-Based Testing in CARLA",
      "authors": [
        "Thoshitha Gamage",
        "Lasanthi Gamage"
      ],
      "abstract": "While the ASAM OpenSCENARIO 2.1 Domain-Specific Language (DSL) enables declarative, intent-driven authoring for Scenario-Based Testing (SBT), its integration into open-source simulators like CARLA remains limited by legacy parsers. We propose a multi-pass modern compiler architecture that translates the OpenSCENARIO 2.1 DSL directly into executable CARLA behaviors. The pipeline features an ANTLR4 frontend for Abstract Syntax Tree (AST) generation, a semantic middle-end, and a runtime backend that synthesizes deterministic py_trees behavior trees. Mapping the standardized domain ontology directly to CARLA's procedural API via a custom method registry eliminates the need for external logic solvers. A demonstrative multi-actor cut-in and evasive maneuver, selected from a wider suite of validated scenarios, confirms the compiler's ability to process concurrent actions, dynamic mathematical expressions, and asynchronous signaling. This framework establishes a functional baseline for reproducible, large-scale SBT, paving the way for future C++ optimizations to mitigate current Python-based computational overhead.",
      "short_abstract": "While the ASAM OpenSCENARIO 2.1 Domain-Specific Language (DSL) enables declarative, intent-driven authoring for Scenario-Based Testing (SBT), its integration into open-source simulators like CARLA remains limited by legacy parsers. We propose a multi-pass...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.16452",
      "pdf_url": "https://arxiv.org/pdf/2604.16452.pdf"
    },
    {
      "id": "2604.06339",
      "anchor": "paper-2604-06339",
      "title": "Evolution of Video Generative Foundations",
      "authors": [
        "Teng Hu",
        "Jiangning Zhang",
        "Hongrui Huang",
        "Ran Yi",
        "Zihan Su",
        "Jieyu Weng",
        "Zhucun Xue",
        "Lizhuang Ma",
        "Ming-Hsuan Yang",
        "Dacheng Tao"
      ],
      "abstract": "The rapid advancement of Artificial Intelligence Generated Content (AIGC) has revolutionized video generation, enabling systems ranging from proprietary pioneers like OpenAI's Sora, Google's Veo3, and Bytedance's Seedance to powerful open-source contenders like Wan and HunyuanVideo to synthesize temporally coherent and semantically rich videos. These advancements pave the way for building \"world models\" that simulate real-world dynamics, with applications spanning entertainment, education, and virtual reality. However, existing reviews on video generation often focus on narrow technical fields, e.g., Generative Adversarial Networks (GAN) and diffusion models, or specific tasks (e. g., video editing), lacking a comprehensive perspective on the field's evolution, especially regarding Auto-Regressive (AR) models and integration of multimodal information. To address these gaps, this survey firstly provides a systematic review of the development of video generation technology, tracing its evolution from early GANs to dominant diffusion models, and further to emerging AR-based and multimodal techniques. We conduct an in-depth analysis of the foundational principles, key advancements, and comparative strengths/limitations. Then, we explore emerging trends in multimodal video generation, emphasizing the integration of diverse data types to enhance contextual awareness. Finally, by bridging historical developments and contemporary innovations, this survey offers insights to guide future research in video generation and its applications, including virtual/augmented reality, personalized education, autonomous driving simulations, digital entertainment, and advanced world models, in this rapidly evolving field. For more details, please refer to the project at https://github.com/sjtuplayer/Awesome-Video-Foundations.",
      "short_abstract": "The rapid advancement of Artificial Intelligence Generated Content (AIGC) has revolutionized video generation, enabling systems ranging from proprietary pioneers like OpenAI's Sora, Google's Veo3, and Bytedance's Seedance to powerful open-source contenders...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "World Model",
        "Survey"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Prediction & World Models: abstract: world model",
          "Safety, Security & Verification: abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.06339",
      "pdf_url": "https://arxiv.org/pdf/2604.06339.pdf"
    },
    {
      "id": "2604.06332",
      "anchor": "paper-2604-06332",
      "title": "Telescope: Learnable Hyperbolic Foveation for Ultra-Long-Range Object Detection",
      "authors": [
        "Parker Ewen",
        "Dmitriy Rivkin",
        "Mario Bijelic",
        "Felix Heide"
      ],
      "abstract": "Autonomous highway driving, especially for long-haul heavy trucks, requires detecting objects at long ranges beyond 500 meters to satisfy braking distance requirements at high speeds. At long distances, vehicles and other critical objects occupy only a few pixels in high-resolution images, causing state-of-the-art object detectors to fail. This challenge is compounded by the limited effective range of commercially available LiDAR sensors, which fall short of ultra-long range thresholds because of quadratic loss of resolution with distance, making image-based detection the most practically scalable solution given commercially available sensor constraints. We introduce Telescope, a two-stage detection model designed for ultra-long range autonomous driving. Alongside a powerful detection backbone, this model contains a novel re-sampling layer and image transformation to address the fundamental challenges of detecting small, distant objects. Telescope achieves $76\\%$ relative improvement in mAP in ultra-long range detection compared to state-of-the-art methods (improving from an absolute mAP of 0.185 to 0.326 at distances beyond 250 meters), requires minimal computational overhead, and maintains strong performance across all detection ranges.",
      "short_abstract": "Autonomous highway driving, especially for long-haul heavy trucks, requires detecting objects at long ranges beyond 500 meters to satisfy braking distance requirements at high speeds. At long distances, vehicles and other critical objects occupy only a few...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 15,
        "margin": 13,
        "evidence": [
          "title: object detection",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.06332",
      "pdf_url": "https://arxiv.org/pdf/2604.06332.pdf"
    },
    {
      "id": "2604.05908",
      "anchor": "paper-2604-05908",
      "title": "Appearance Decomposition Gaussian Splatting for Multi-Traversal Reconstruction",
      "authors": [
        "Yangyi Xiao",
        "Siting Zhu",
        "Baoquan Yang",
        "Tianchen Deng",
        "Yongbo Chen",
        "Hesheng Wang"
      ],
      "abstract": "Multi-traversal scene reconstruction is important for high-fidelity autonomous driving simulation and digital twin construction. This task involves integrating multiple sequences captured from the same geographical area at different times. In this context, a primary challenge is the significant appearance inconsistency across traversals caused by varying illumination and environmental conditions, despite the shared underlying geometry. This paper presents ADM-GS (Appearance Decomposition Gaussian Splatting for Multi-Traversal Reconstruction), a framework that applies an explicit appearance decomposition to the static background to alleviate appearance entanglement across traversals. For the static background, we decompose the appearance into traversal-invariant material, representing intrinsic material properties, and traversal-dependent illumination, capturing lighting variations. Specifically, we propose a neural light field that utilizes a frequency-separated hybrid encoding strategy. By incorporating surface normals and explicit reflection vectors, this design separately captures low-frequency diffuse illumination and high-frequency specular reflections. Quantitative evaluations on the Argoverse 2 and Waymo Open datasets demonstrate the effectiveness of ADM-GS. In multi-traversal experiments, our method achieves a +0.98 dB PSNR improvement over existing latent-based baselines while producing more consistent appearance across traversals. Code will be available at https://github.com/IRMVLab/ADM-GS.",
      "short_abstract": "Multi-traversal scene reconstruction is important for high-fidelity autonomous driving simulation and digital twin construction. This task involves integrating multiple sequences captured from the same geographical area at different times. In this context, a...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.05908",
      "pdf_url": "https://arxiv.org/pdf/2604.05908.pdf"
    },
    {
      "id": "2604.05780",
      "anchor": "paper-2604-05780",
      "title": "Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion",
      "authors": [
        "Yu Xue",
        "Longjun Gao",
        "Yuanqi Su",
        "HaoAng Lu",
        "Xiaoning Zhang"
      ],
      "abstract": "Monocular Semantic Scene Completion (SSC) aims to reconstruct complete 3D semantic scenes from a single RGB image, offering a cost-effective solution for autonomous driving and robotics. However, the inherently imbalanced nature of voxel distributions, where over 93% of voxels are empty and foreground classes are rare, poses significant challenges. Existing methods often suffer from redundant emphasis on uninformative voxels and poor generalization to long-tailed categories. To address these issues, we propose VoxSAMNet (Voxel Sparsity-Aware Modulation Network), a unified framework that explicitly models voxel sparsity and semantic imbalance. Our approach introduces: (1) a Dummy Shortcut for Feature Refinement (DSFR) module that bypasses empty voxels via a shared dummy node while refining occupied ones with deformable attention; and (2) a Foreground Modulation Strategy combining Foreground Dropout (FD) and Text-Guided Image Filter (TGIF) to alleviate overfitting and enhance class-relevant features. Extensive experiments on the public benchmarks SemanticKITTI and SSCBench-KITTI-360 demonstrate that VoxSAMNet achieves state-of-the-art performance, surpassing prior monocular and stereo baselines with mIoU scores of 18.2% and 20.2%, respectively. Our results highlight the importance of sparsity-aware and semantics-guided design for efficient and accurate 3D scene completion, offering a promising direction for future research.",
      "short_abstract": "Monocular Semantic Scene Completion (SSC) aims to reconstruct complete 3D semantic scenes from a single RGB image, offering a cost-effective solution for autonomous driving and robotics. However, the inherently imbalanced nature of voxel distributions, where...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.05780",
      "pdf_url": "https://arxiv.org/pdf/2604.05780.pdf"
    },
    {
      "id": "2604.05449",
      "anchor": "paper-2604-05449",
      "title": "Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning",
      "authors": [
        "Kang Ding",
        "Hongsong Wang",
        "Jie Gui",
        "Lei He"
      ],
      "abstract": "End-to-end autonomous driving resides not in the integration of perception and planning, but rather in the dynamic multi-agent game within a unified representation space. Most existing end-to-end models treat all agents equally, hindering the decoupling of real collision threats from complex backgrounds. To address this issue, We introduce the concept of Risk-Prioritized Game Planning, and propose GameAD, a novel framework that models end-to-end autonomous driving as a risk-aware game problem. The GameAD integrates Risk-Aware Topology Anchoring, Strategic Payload Adapter, Minimax Risk-Aware Sparse Attention, and Risk Consistent Equilibrium Stabilization to enable game theoretic decision making with risk prioritized interactions. We also present the Planning Risk Exposure metric, which quantifies the cumulative risk intensity of planned trajectories over a long horizon for safe autonomous driving. Extensive experiments on the nuScenes and Bench2Drive datasets show that our approach significantly outperforms state-of-the-art methods, especially in terms of trajectory safety.",
      "short_abstract": "End-to-end autonomous driving resides not in the integration of perception and planning, but rather in the dynamic multi-agent game within a unified representation space. Most existing end-to-end models treat all agents equally, hindering the decoupling of...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.05449",
      "pdf_url": "https://arxiv.org/pdf/2604.05449.pdf"
    },
    {
      "id": "2604.07378",
      "anchor": "paper-2604-07378",
      "title": "Evaluation as Evolution: Transforming Adversarial Diffusion into Closed-Loop Curricula for Autonomous Vehicles",
      "authors": [
        "Yicheng Guo",
        "Jiaqi Liu",
        "Chengkai Xu",
        "Peng Hang",
        "Jian Sun"
      ],
      "abstract": "Autonomous vehicles in interactive traffic environments are often limited by the scarcity of safety-critical tail events in static datasets, which biases learned policies toward average-case behaviors and reduces robustness. Existing evaluation methods attempt to address this through adversarial stress testing, but are predominantly open-loop and post-hoc, making it difficult to incorporate discovered failures back into the training process. We introduce Evaluation as Evolution ($E^2$), a closed-loop framework that transforms adversarial generation from a static validation step into an adaptive evolutionary curriculum. Specifically, $E^2$ formulates adversarial scenario synthesis as transport-regularized sparse control over a learned reverse-time SDE prior. To make this high-dimensional generation tractable, we utilize topology-driven support selection to identify critical interacting agents, and introduce Topological Anchoring to stabilize the process. This approach enables the targeted discovery of failure cases while strictly constraining deviations from realistic data distributions. Empirically, $E^2$ improves collision failure discovery by 9.01% on the nuScenes dataset and up to 21.43% on the nuPlan dataset over the strongest baselines, while maintaining low invalidity and high realism. It further yields substantial robustness gains when the resulting boundary cases are recycled for closed-loop policy fine-tuning.",
      "short_abstract": "Autonomous vehicles in interactive traffic environments are often limited by the scarcity of safety-critical tail events in static datasets, which biases learned policies toward average-case behaviors and reduces robustness. Existing evaluation methods...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 23,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.07378",
      "pdf_url": "https://arxiv.org/pdf/2604.07378.pdf"
    },
    {
      "id": "2604.06387",
      "anchor": "paper-2604-06387",
      "title": "Uncertainty Estimation for Deep Reconstruction in Actuatic Disaster Scenarios with Autonomous Vehicles",
      "authors": [
        "Samuel Yanes Luis",
        "Alejandro Casado Pérez",
        "Alejandro Mendoza Barrionuevo",
        "Dame Seck Diop",
        "Sergio Toral Marín",
        "Daniel Gutiérrez Reina"
      ],
      "abstract": "Accurate reconstruction of environmental scalar fields from sparse onboard observations is essential for autonomous vehicles engaged in aquatic monitoring. Beyond point estimates, principled uncertainty quantification is critical for active sensing strategies such as Informative Path Planning, where epistemic uncertainty drives data collection decisions. This paper compares Gaussian Processes, Monte Carlo Dropout, Deep Ensembles, and Evidential Deep Learning for simultaneous scalar field reconstruction and uncertainty decomposition under three perceptual models representative of real sensor modalities. Results show that Evidential Deep Learning achieves the best reconstruction accuracy and uncertainty calibration across all sensor configurations at the lowest inference cost, while Gaussian Processes are fundamentally limited by their stationary kernel assumption and become intractable as observation density grows. These findings support Evidential Deep Learning as the preferred method for uncertainty-aware field reconstruction in real-time autonomous vehicle deployments.",
      "short_abstract": "Accurate reconstruction of environmental scalar fields from sparse onboard observations is essential for autonomous vehicles engaged in aquatic monitoring. Beyond point estimates, principled uncertainty quantification is critical for active sensing strategies...",
      "published": "2026-04-07",
      "updated": "2026-04-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: abstract: motion planning",
          "Planning & Decision-Making: abstract: planning",
          "Safety, Security & Verification: title: uncertainty",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 134,
      "arxiv_url": "https://arxiv.org/abs/2604.06387",
      "pdf_url": "https://arxiv.org/pdf/2604.06387.pdf"
    },
    {
      "id": "2604.16436",
      "anchor": "paper-2604-16436",
      "title": "Fuzzy Encoding-Decoding to Improve Spiking Q-Learning Performance in Autonomous Driving",
      "authors": [
        "Aref Ghoreishee",
        "Abhishek Mishra",
        "Lifeng Zhou",
        "John Walsh",
        "Anup Das",
        "Nagarajan Kandasamy"
      ],
      "abstract": "This paper develops an end-to-end fuzzy encoder-decoder architecture for enhancing vision-based multi-modal deep spiking Q-networks in autonomous driving. The method addresses two core limitations of spiking reinforcement learning: information loss stemming from the conversion of dense visual inputs into sparse spike trains, and the limited representational capacity of spike-based value functions, which often yields weakly discriminative Q-value estimates. The encoder introduces trainable fuzzy membership functions to generate expressive, population-based spike representations, and the decoder uses a lightweight neural decoder to reconstruct continuous Q-values from spiking outputs. Experiments on the HighwayEnv benchmark show that the proposed architecture substantially improves decision-making accuracy and closes the performance gap between spiking and non-spiking multi-modal Q-networks. The results highlight the potential of this framework for efficient and real-time autonomous driving with spiking neural networks.",
      "short_abstract": "This paper develops an end-to-end fuzzy encoder-decoder architecture for enhancing vision-based multi-modal deep spiking Q-networks in autonomous driving. The method addresses two core limitations of spiking reinforcement learning: information loss stemming...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: communications",
          "Planning & Decision-Making: abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.16436",
      "pdf_url": "https://arxiv.org/pdf/2604.16436.pdf"
    },
    {
      "id": "2604.05378",
      "anchor": "paper-2604-05378",
      "title": "ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving",
      "authors": [
        "Kaiser Hamid",
        "Can Cui",
        "Nade Liang"
      ],
      "abstract": "Recent progress in vision-language-action (VLA) models has enabled language-conditioned driving agents to execute natural-language navigation commands in closed-loop simulation, yet standard evaluations largely assume instructions are precise and well-formed. In deployment, instructions vary in phrasing and specificity, may omit critical qualifiers, and can occasionally include misleading, authority-framed text, leaving instruction-level robustness under-measured. We introduce ICR-Drive, a diagnostic framework for instruction counterfactual robustness in end-to-end language-conditioned autonomous driving. ICR-Drive generates controlled instruction variants spanning four perturbation families: Paraphrase, Ambiguity, Noise, and Misleading, where Misleading variants conflict with the navigation goal and attempt to override intent. We replay identical CARLA routes under matched simulator configurations and seeds to isolate performance changes attributable to instruction language. Robustness is quantified using standard CARLA Leaderboard metrics and per-family performance degradation relative to the baseline instruction. Experiments on LMDrive and BEVDriver show that minor instruction changes can induce substantial performance drops and distinct failure modes, revealing a reliability gap for deploying embodied foundation models in safety-critical driving.",
      "short_abstract": "Recent progress in vision-language-action (VLA) models has enabled language-conditioned driving agents to execute natural-language navigation commands in closed-loop simulation, yet standard evaluations largely assume instructions are precise and well-formed....",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Simulation",
        "VLA"
      ],
      "classification": {
        "confidence": "medium",
        "score": 18,
        "margin": 4,
        "evidence": [
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.05378",
      "pdf_url": "https://arxiv.org/pdf/2604.05378.pdf"
    },
    {
      "id": "2604.05115",
      "anchor": "paper-2604-05115",
      "title": "Probabilistic Tree Inference Enabled by FDSOI Ferroelectric FETs",
      "authors": [
        "Pengyu Ren",
        "Xingtian Wang",
        "Boyang Cheng",
        "Jiahui Duan",
        "Giuk Kim",
        "Xuezhong Niu",
        "Halid Mulaosmanovic",
        "Stefan Duenkel",
        "Sven Beyer",
        "X. Sharon Hu",
        "Ningyuan Cao",
        "Kai Ni"
      ],
      "abstract": "Artificial intelligence applications in autonomous driving, medical diagnostics, and financial systems increasingly demand machine learning models that can provide robust uncertainty quantification, interpretability, and noise resilience. Bayesian decision trees (BDTs) are attractive for these tasks because they combine probabilistic reasoning, interpretable decision-making, and robustness to noise. However, existing hardware implementations of BDTs based on CPUs and GPUs are limited by memory bottlenecks and irregular processing patterns, while multi-platform solutions exploiting analog content-addressable memory (ACAM) and Gaussian random number generators (GRNGs) introduce integration complexity and energy overheads. Here we report a monolithic FDSOI-FeFET hardware platform that natively supports both ACAM and GRNG functionalities. The ferroelectric polarization of FeFETs enables compact, energy-efficient multi-bit storage for ACAM, and band-to-band tunneling in the gate-to-drain overlap region and subsequent hole storage in the floating body provides a high-quality entropy source for GRNG. System-level evaluations demonstrate that the proposed architecture provides robust uncertainty estimation, interpretability, and noise tolerance with high energy efficiency. Under both dataset noise and device variations, it achieves over 40% higher classification accuracy on MNIST compared to conventional decision trees. Moreover, it delivers more than two orders of magnitude speedup over CPU and GPU baselines and over four orders of magnitude improvement in energy efficiency, making it a scalable solution for deploying BDTs in resource-constrained and safety-critical environments.",
      "short_abstract": "Artificial intelligence applications in autonomous driving, medical diagnostics, and financial systems increasingly demand machine learning models that can provide robust uncertainty quantification, interpretability, and noise resilience. Bayesian decision...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: robustness",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.05115",
      "pdf_url": "https://arxiv.org/pdf/2604.05115.pdf"
    },
    {
      "id": "2604.05070",
      "anchor": "paper-2604-05070",
      "title": "Part-Level 3D Gaussian Vehicle Generation with Joint and Hinge Axis Estimation",
      "authors": [
        "Shiyao Qian",
        "Yuan Ren",
        "Dongfeng Bai",
        "Bingbing Liu"
      ],
      "abstract": "Simulation is essential for autonomous driving, yet current frameworks often model vehicles as rigid assets and fail to capture part-level articulation. With perception algorithms increasingly leveraging dynamics such as wheel steering or door opening, realistic simulation requires animatable vehicle representations. Existing CAD-based pipelines are limited by library coverage and fixed templates, preventing faithful reconstruction of in-the-wild instances. We propose a generative framework that, from a single image or sparse multi-view input, synthesizes an animatable 3D Gaussian vehicle. Our method addresses two challenges: (i) large 3D asset generators are optimized for static quality but not articulation, leading to distortions at part boundaries when animated; and (ii) segmentation alone cannot provide the kinematic parameters required for motion. To overcome this, we introduce a part-edge refinement module that enforces exclusive Gaussian ownership and a kinematic reasoning head that predicts joint positions and hinge axes of movable parts. Together, these components enable faithful part-aware simulation, bridging the gap between static generation and animatable vehicle models.",
      "short_abstract": "Simulation is essential for autonomous driving, yet current frameworks often model vehicles as rigid assets and fail to capture part-level articulation. With perception algorithms increasingly leveraging dynamics such as wheel steering or door opening,...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Control & Vehicle Dynamics: abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.05070",
      "pdf_url": "https://arxiv.org/pdf/2604.05070.pdf"
    },
    {
      "id": "2604.04887",
      "anchor": "paper-2604-04887",
      "title": "HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes",
      "authors": [
        "Mauricio Soroco",
        "Francesco Pittaluga",
        "Zaid Tasneem",
        "Abhishek Aich",
        "Bingbing Zhuang",
        "Wuyang Chen",
        "Manmohan Chandraker",
        "Ziyu Jiang"
      ],
      "abstract": "Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real-world testing provides. Yet existing instruction guided image editors, trained on object-centric or artistic data, struggle with dense, safety-critical driving layouts. We propose HorizonWeaver, which tackles three fundamental challenges in driving scene editing: (1) multi-level granularity, requiring coherent object- and scene-level edits in dense environments; (2) rich high-level semantics, preserving diverse objects while following detailed instructions; and (3) ubiquitous domain shifts, handling changes in climate, layout, and traffic across unseen environments. The core of HorizonWeaver is a set of complementary contributions across data, model, and training: (1) Data: Large-scale dataset generation, where we build a paired real/synthetic dataset from Boreas, nuScenes, and Argoverse2 to improve generalization; (2) Model: Language-Guided Masks for fine-grained editing, where semantics-enriched masks and prompts enable precise, language-guided edits; and (3) Training: Content preservation and instruction alignment, where joint losses enforce scene consistency and instruction fidelity. Together, HorizonWeaver provides a scalable framework for photorealistic, instruction-driven editing of complex driving scenes, collecting 255K images across 13 editing categories and outperforming prior methods in L1, CLIP, and DINO metrics, achieving +46.4% user preference and improving BEV segmentation IoU by +33%. Project page: https://msoroco.github.io/horizonweaver/",
      "short_abstract": "Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real-world testing provides. Yet existing instruction guided image editors, trained on object-centric or artistic data, struggle with...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 5,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.04887",
      "pdf_url": "https://arxiv.org/pdf/2604.04887.pdf"
    },
    {
      "id": "2604.04857",
      "anchor": "paper-2604-04857",
      "title": "The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models",
      "authors": [
        "Runhao Mao",
        "Hanshi Wang",
        "Yixiang Yang",
        "Qianli Ma",
        "Jingmeng Zhou",
        "Zhipeng Zhang"
      ],
      "abstract": "The integration of Vision-Language Models (VLMs) into autonomous driving promises to solve long-tail scenarios, but this paradigm faces the critical and unaddressed challenge of catastrophic forgetting. The very fine-tuning process used to adapt these models to driving-specific data simultaneously erodes their invaluable pre-trained world knowledge, creating a self-defeating paradox that undermines the core reason for their use. This paper provides the first systematic investigation into this phenomenon. We introduce a new large-scale dataset of 180K scenes, which enables the first-ever benchmark specifically designed to quantify catastrophic forgetting in autonomous driving. Our analysis reveals that existing methods suffer from significant knowledge degradation. To address this, we propose the Drive Expert Adapter (DEA), a novel framework that circumvents this trade-off by shifting adaptation from the weight space to the prompt space. DEA dynamically routes inference through different knowledge experts based on scene-specific cues, enhancing driving-task performance without corrupting the model's foundational parameters. Extensive experiments demonstrate that our approach not only achieves state-of-the-art results on driving tasks but also effectively mitigates catastrophic forgetting, preserving the essential generalization capabilities that make VLMs a transformative force for autonomous systems. Data and model are released at FidelityDrivingBench.",
      "short_abstract": "The integration of Vision-Language Models (VLMs) into autonomous driving promises to solve long-tail scenarios, but this paradigm faces the critical and unaddressed challenge of catastrophic forgetting. The very fine-tuning process used to adapt these models...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset",
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.04857",
      "pdf_url": "https://arxiv.org/pdf/2604.04857.pdf"
    },
    {
      "id": "2604.04797",
      "anchor": "paper-2604-04797",
      "title": "Multi-Modal Sensor Fusion using Hybrid Attention for Autonomous Driving",
      "authors": [
        "Mayank Mayank",
        "Bharanidhar Duraisamy",
        "Florian Geiß",
        "Abhinav Valada"
      ],
      "abstract": "Accurate 3D object detection for autonomous driving requires complementary sensors. Cameras provide dense semantics but unreliable depth, while millimeter-wave radar offers precise range and velocity measurements with sparse geometry. We propose MMF-BEV, a radar-camera BEV fusion framework that leverages deformable attention for cross-modal feature alignment on the View-of-Delft (VoD) 4D radar dataset [1]. MMF-BEV builds a BEVDepth [2] camera branch and a RadarBEVNet [3] radar branch, each enhanced with Deformable Self-Attention, and fuses them via a Deformable Cross-Attention module. We evaluate three configurations: camera-only, radar-only, and hybrid fusion. A sensor contribution analysis quantifies per-distance modality weighting, providing interpretable evidence of sensor complementarity. A two-stage training strategy - pre-training the camera branch with depth supervision, then jointly training radar and fusion modules stabilizes learning. Experiments on VoD show that MMF-BEV consistently outperforms unimodal baselines and achieves competitive results against prior fusion methods across all object classes in both the full annotated area and near-range Region of Interest.",
      "short_abstract": "Accurate 3D object detection for autonomous driving requires complementary sensors. Cameras provide dense semantics but unreliable depth, while millimeter-wave radar offers precise range and velocity measurements with sparse geometry. We propose MMF-BEV, a...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 23,
        "evidence": [
          "abstract: object detection",
          "title: sensor fusion",
          "abstract: radar",
          "abstract: camera",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.04797",
      "pdf_url": "https://arxiv.org/pdf/2604.04797.pdf"
    },
    {
      "id": "2604.04630",
      "anchor": "paper-2604-04630",
      "title": "Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers",
      "authors": [
        "Jiancheng Wang",
        "Lidan Liang",
        "Yong Wang",
        "Zengzhen Su",
        "Haifeng Xia",
        "Yuanting Yan",
        "Wei Wang"
      ],
      "abstract": "Visual language model (VLM) is rapidly being integrated into safety-critical systems such as autonomous driving, making it an important attack surface for potential backdoor attacks. Existing backdoor attacks mainly rely on unimodal, explicit, and easily detectable triggers, making it difficult to construct both covert and stable attack channels in autonomous driving scenarios. GLA introduces two naturalistic triggers: graffiti-based visual patterns generated via stable diffusion inpainting, which seamlessly blend into urban scenes, and cross-language text triggers, which introduce distributional shifts while maintaining semantic consistency to build robust language-side trigger signals. Experiments on DriveVLM show that GLA requires only a 10\\% poisoning ratio to achieve a 90\\% Attack Success Rate (ASR) and a 0\\% False Positive Rate (FPR). More insidiously, the backdoor does not weaken the model on clean tasks, but instead improves metrics such as BLEU-1, making it difficult for traditional performance-degradation-based detection methods to identify the attack. This study reveals underestimated security threats in self-driving VLMs and provides a new attack paradigm for backdoor evaluation in safety-critical multimodal systems.",
      "short_abstract": "Visual language model (VLM) is rapidly being integrated into safety-critical systems such as autonomous driving, making it an important attack surface for potential backdoor attacks. Existing backdoor attacks mainly rely on unimodal, explicit, and easily...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 27,
        "margin": 27,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.04630",
      "pdf_url": "https://arxiv.org/pdf/2604.04630.pdf"
    },
    {
      "id": "2604.04518",
      "anchor": "paper-2604-04518",
      "title": "Reproducibility study on how to find Spurious Correlations, Shortcut Learning, Clever Hans or Group-Distributional non-robustness and how to fix them",
      "authors": [
        "Ole Delzer",
        "Sidney Bender"
      ],
      "abstract": "Deep Neural Networks (DNNs) are increasingly utilized in high-stakes domains like medical diagnostics and autonomous driving where model reliability is critical. However, the research landscape for ensuring this reliability is terminologically fractured across communities that pursue the same goal of ensuring models rely on causally relevant features rather than confounding signals. While frameworks such as distributionally robust optimization (DRO), invariant risk minimization (IRM), shortcut learning, simplicity bias, and the Clever Hans effect all address model failure due to spurious correlations, researchers typically only reference work within their own domains. This reproducibility study unifies these perspectives through a comparative analysis of correction methods under challenging constraints like limited data availability and severe subgroup imbalance. We evaluate recently proposed correction methods based on explainable artificial intelligence (XAI) techniques alongside popular non-XAI baselines using both synthetic and real-world datasets. Findings show that XAI-based methods generally outperform non-XAI approaches, with Counterfactual Knowledge Distillation (CFKD) proving most consistently effective at improving generalization. Our experiments also reveal that the practical application of many methods is hindered by a dependency on group labels, as manual annotation is often infeasible and automated tools like Spectral Relevance Analysis (SpRAy) struggle with complex features and severe imbalance. Furthermore, the scarcity of minority group samples in validation sets renders model selection and hyperparameter tuning unreliable, posing a significant obstacle to the deployment of robust and trustworthy models in safety-critical areas.",
      "short_abstract": "Deep Neural Networks (DNNs) are increasingly utilized in high-stakes domains like medical diagnostics and autonomous driving where model reliability is critical. However, the research landscape for ensuring this reliability is terminologically fractured...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "title: robustness",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.04518",
      "pdf_url": "https://arxiv.org/pdf/2604.04518.pdf"
    },
    {
      "id": "2604.04775",
      "anchor": "paper-2604-04775",
      "title": "Community Driving-Safety Deterioration as a Push Factor for Public Endorsement of AI Driving Capability",
      "authors": [
        "Amir Rafe",
        "Subasish Das"
      ],
      "abstract": "Road traffic crashes claim approximately 1.19 million lives annually worldwide, and human error accounts for the vast majority, yet the autonomous vehicle acceptance literature models adoption almost exclusively through technology-centered pull factors such as perceived usefulness and trust. This study examines a moderated mediation model in which perceived community driving-safety concern (PCSC) predicts evaluations of AI versus human driving capability, mediated by Generalized AI Orientation and moderated by personal driving frequency. Weighted structural equation modeling is applied to a nationally representative U.S. probability sample from Pew Research Center's American Trends Panel Wave 152, using Weighted Least Squares Mean and Variance Adjusted (WLSMV)-estimated confirmatory factor analysis on ordinal indicators, bias-corrected bootstrap inference, and seven robustness checks including Imai sensitivity analysis, E-value confounding thresholds, and propensity score matching. Results reveal a dual-pathway mechanism constituting an inconsistent mediation: PCSC exerts a small positive direct effect on AI driving evaluation, consistent with a domain-specific push interpretation, while simultaneously suppressing Generalized AI Orientation, which is itself a strong positive predictor of AI driving evaluation. Conditional indirect effects are negative and statistically significant at low, mean, and high levels of driving frequency. These findings establish a risk-spillover mechanism whereby community driving-safety concern promotes domain-specific AI endorsement yet suppresses domain-general AI enthusiasm, yielding a near-zero net total effect.",
      "short_abstract": "Road traffic crashes claim approximately 1.19 million lives annually worldwide, and human error accounts for the vast majority, yet the autonomous vehicle acceptance literature models adoption almost exclusively through technology-centered pull factors such...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 18,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.04775",
      "pdf_url": "https://arxiv.org/pdf/2604.04775.pdf"
    },
    {
      "id": "2604.04573",
      "anchor": "paper-2604-04573",
      "title": "SAIL: Scene-aware Adaptive Iterative Learning for Long-Tail Trajectory Prediction in Autonomous Vehicles",
      "authors": [
        "Bin Rao",
        "Haicheng Liao",
        "Chengyue Wang",
        "Keqiang Li",
        "Zhenning Li",
        "Hai Yang"
      ],
      "abstract": "Autonomous vehicles (AVs) rely on accurate trajectory prediction for safe navigation in diverse traffic environments, yet existing models struggle with long-tail scenarios-rare but safety-critical events characterized by abrupt maneuvers, high collision risks, and complex interactions. These challenges stem from data imbalance, inadequate definitions of long-tail trajectories, and suboptimal learning strategies that prioritize common behaviors over infrequent ones. To address this, we propose SAIL, a novel framework that systematically tackles the long-tail problem by first defining and modeling trajectories across three key attribute dimensions: prediction error, collision risk, and state complexity. Our approach then synergizes an attribute-guided augmentation and feature extraction process with a highly adaptive contrastive learning strategy. This strategy employs a continuous cosine momentum schedule, similarity-weighted hard-negative mining, and a dynamic pseudo-labeling mechanism based on evolving feature clustering. Furthermore, it incorporates a focusing mechanism to intensify learning on hard-positive samples within each identified class. This comprehensive design enables SAIL to excel at identifying and forecasting diverse and challenging long-tail events. Extensive evaluations on the nuScenes and ETH/UCY datasets demonstrate SAIL's superior performance, achieving up to 28.8% reduction in prediction error on the hardest 1% of long-tail samples compared to state-of-the-art baselines, while maintaining competitive accuracy across all scenarios. This framework advances reliable AV trajectory prediction in real-world, mixed-autonomy settings.",
      "short_abstract": "Autonomous vehicles (AVs) rely on accurate trajectory prediction for safe navigation in diverse traffic environments, yet existing models struggle with long-tail scenarios-rare but safety-critical events characterized by abrupt maneuvers, high collision...",
      "published": "2026-04-06",
      "updated": "2026-04-06",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 27,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 135,
      "arxiv_url": "https://arxiv.org/abs/2604.04573",
      "pdf_url": "https://arxiv.org/pdf/2604.04573.pdf"
    },
    {
      "id": "2604.04349",
      "anchor": "paper-2604-04349",
      "title": "Adversarial Robustness Analysis of Cloud-Assisted Autonomous Driving Systems",
      "authors": [
        "Maher Al Islam",
        "Amr S. El-Wakeel"
      ],
      "abstract": "Autonomous vehicles increasingly rely on deep learning-based perception and control, which impose substantial computational demands. Cloud-assisted architectures offload these functions to remote servers, enabling enhanced perception and coordinated decision-making through the Internet of Vehicles (IoV). However, this paradigm introduces cross-layer vulnerabilities, where adversarial manipulation of perception models and network impairments in the vehicle-cloud link can jointly undermine safety-critical autonomy. This paper presents a hardware-in-the-loop IoV testbed that integrates real-time perception, control, and communication to evaluate such vulnerabilities in cloud-assisted autonomous driving. A YOLOv8-based object detector deployed on the cloud is subjected to whitebox adversarial attacks using the Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD), while network adversaries induce delay and packet loss in the vehicle-cloud loop. Results show that adversarial perturbations significantly degrade perception performance, with PGD reducing detection precision and recall from 0.73 and 0.68 in the clean baseline to 0.22 and 0.15 at epsilon= 0.04. Network delays of 150-250 ms, corresponding to transient losses of approximately 3-4 frames, and packet loss rates of 0.5-5 % further destabilize closed-loop control, leading to delayed actuation and rule violations. These findings highlight the need for cross-layer resilience in cloud-assisted autonomous driving systems.",
      "short_abstract": "Autonomous vehicles increasingly rely on deep learning-based perception and control, which impose substantial computational demands. Cloud-assisted architectures offload these functions to remote servers, enabling enhanced perception and coordinated...",
      "published": "2026-04-05",
      "updated": "2026-04-05",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 26,
        "margin": 2,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "title: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 136,
      "arxiv_url": "https://arxiv.org/abs/2604.04349",
      "pdf_url": "https://arxiv.org/pdf/2604.04349.pdf"
    },
    {
      "id": "2604.04331",
      "anchor": "paper-2604-04331",
      "title": "GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction",
      "authors": [
        "Yedong Shen",
        "Shiqi Zhang",
        "Sha Zhang",
        "Yifan Duan",
        "Xinran Zhang",
        "Wenhao Yu",
        "Lu Zhang",
        "Jiajun Deng",
        "Yanyong Zhang"
      ],
      "abstract": "Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches typically rely on background for static scene reconstruction, limiting the ability to recover regions occluded by dynamic objects. In this paper, we propose GA-GS, a Generation-Assisted Gaussian Splatting method for Static Scene Reconstruction. The key innovation of our work lies in leveraging generation to assist in reconstructing occluded regions. We employ a motion-aware module to segment and remove dynamic regions, and thenuse a diffusion model to inpaint the occluded areas, providing pseudo-ground-truth supervision. To balance contributions from real background and generated region, we introduce a learnable authenticity scalar for each Gaussian primitive, which dynamically modulates opacity during splatting for authenticity-aware rendering and supervision. Since no existing dataset provides ground-truth static scene of video with dynamic objects, we construct a dataset named Trajectory-Match, using a fixed-path robot to record each scene with/without dynamic objects, enabling quantitative evaluation in reconstruction of occluded regions. Extensive experiments on both the DAVIS and our dataset show that GA-GS achieves state-of-the-art performance in static scene reconstruction, especially in challenging scenarios with large-scale, persistent occlusions.",
      "short_abstract": "Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches typically rely on background for static scene reconstruction, limiting the...",
      "published": "2026-04-05",
      "updated": "2026-04-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 136,
      "arxiv_url": "https://arxiv.org/abs/2604.04331",
      "pdf_url": "https://arxiv.org/pdf/2604.04331.pdf"
    },
    {
      "id": "2604.04198",
      "anchor": "paper-2604-04198",
      "title": "DriveVA: Video Action Models are Zero-Shot Drivers",
      "authors": [
        "Mengmeng Liu",
        "Diankun Zhang",
        "Jiuming Liu",
        "Jianfeng Cui",
        "Hongwei Xie",
        "Guang Chen",
        "Hangjun Ye",
        "Michael Ying Yang",
        "Francesco Nex",
        "Hao Cheng"
      ],
      "abstract": "Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world-model-based planning methods have shown strong capabilities in scene understanding and multi-modal future prediction, yet their generalization across datasets and sensor configurations remains limited. In addition, their loosely coupled planning paradigm often leads to poor video-trajectory consistency during visual imagination. To overcome these limitations, we propose DriveVA, a novel autonomous driving world model that jointly decodes future visual forecasts and action sequences in a shared latent generative process. DriveVA inherits rich priors on motion dynamics and physical plausibility from well-pretrained large-scale video generation models to capture continuous spatiotemporal evolution and causal interaction patterns. To this end, DriveVA employs a DiT-based decoder to jointly predict future action sequences (trajectories) and videos, enabling tighter alignment between planning and scene evolution. We also introduce a video continuation strategy to strengthen long-duration rollout consistency. DriveVA achieves an impressive PDM-based planning performance of 90.9 PDM score on the NAVSIM benchmark. Extensive experiments also demonstrate the zero-shot capability and cross-domain generalization of DriveVA, which reduces average L2 error and collision rate by 78.9% and 83.3% on nuScenes and 52.5% and 52.4% on the Bench2Drive built on CARLA v2 compared with the state-of-the-art world-model-based planner.",
      "short_abstract": "Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world-model-based planning methods have shown strong capabilities in...",
      "published": "2026-04-05",
      "updated": "2026-04-05",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 6,
        "evidence": [
          "abstract: world model",
          "abstract: future prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 136,
      "arxiv_url": "https://arxiv.org/abs/2604.04198",
      "pdf_url": "https://arxiv.org/pdf/2604.04198.pdf"
    },
    {
      "id": "2604.04015",
      "anchor": "paper-2604-04015",
      "title": "Enabling Deterministic User-Level Interrupts in Real-Time Processors via Hardware Extension",
      "authors": [
        "Hongbin Yang",
        "Huanle Zhang",
        "Runyu Pan"
      ],
      "abstract": "The growing complexity of real-time embedded systems demands strong isolation of software components into separate protection domains to reduce attack surfaces and limit fault propagation. However, application-supplied device interrupt handlers -- even untrusted -- have to remain in the kernel to minimize interrupt latency, undermining security and burdening manual certifications. Current hardware extensions accelerate interrupts only when the target protection domain is scheduled by the kernel; consequently, they are limited to improving average-case performance but not worst-case latency, and do not meet the requirements of critical real-time applications such as autonomous vehicles or robots. To overcome this limitation, we propose a novel hardware extension that enables direct, deterministic switching to the appropriate protection domain upon user-level interrupt arrival -- without kernel intervention -- even when that domain is dormant. Our hardware extension reduces worst-case latency by more than 50x with a 19% increase in core area (2% of total die area) and 4.1% increase in dynamic power. To the best of our knowledge, this is the first integrated mechanism to guarantee user-level interrupt delivery with a nanosecond-scale yet bounded worst-case latency.",
      "short_abstract": "The growing complexity of real-time embedded systems demands strong isolation of software components into separate protection domains to reduce attack surfaces and limit fault propagation. However, application-supplied device interrupt handlers -- even...",
      "published": "2026-04-05",
      "updated": "2026-04-05",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 17,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "title: hardware",
          "abstract: hardware",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 136,
      "arxiv_url": "https://arxiv.org/abs/2604.04015",
      "pdf_url": "https://arxiv.org/pdf/2604.04015.pdf"
    },
    {
      "id": "2604.14209",
      "anchor": "paper-2604-14209",
      "title": "Towards Verified and Targeted Explanations through Formal Methods",
      "authors": [
        "Hanchen David Wang",
        "Diego Manzanas Lopez",
        "Preston K. Robinette",
        "Ipek Oguz",
        "Taylor T. Johnson",
        "Meiyi Ma"
      ],
      "abstract": "As deep neural networks are deployed in safety-critical domains such as autonomous driving and medical diagnosis, stakeholders need explanations that are interpretable but also trustworthy with formal guarantees. Existing XAI methods fall short: heuristic attribution techniques (e.g., LIME, Integrated Gradients) highlight influential features but offer no mathematical guarantees about decision boundaries, while formal methods verify robustness yet remain untargeted, analyzing the nearest boundary regardless of whether it represents a critical risk. In safety-critical systems, not all misclassifications carry equal consequences; confusing a \"Stop\" sign for a \"60 kph\" sign is far more dangerous than confusing it with a \"No Passing\" sign. We introduce ViTaX (Verified and Targeted Explanations), a formal XAI framework that generates targeted semifactual explanations with mathematical guarantees. For a given input (class y) and a user-specified critical alternative (class t), ViTaX: (1) identifies the minimal feature subset most sensitive to the y->t transition, and (2) applies formal reachability analysis to guarantee that perturbing these features by epsilon cannot flip the classification to t. We formalize this through Targeted epsilon-Robustness, certifying whether a feature subset remains robust under perturbation toward a specific target class. ViTaX is the first method to provide formally guaranteed explanations of a model's resilience against user-identified alternatives. Evaluations on MNIST, GTSRB, EMNIST, and TaxiNet demonstrate over 30% fidelity improvement with minimal explanation cardinality.",
      "short_abstract": "As deep neural networks are deployed in safety-critical domains such as autonomous driving and medical diagnosis, stakeholders need explanations that are interpretable but also trustworthy with formal guarantees. Existing XAI methods fall short: heuristic...",
      "published": "2026-04-04",
      "updated": "2026-04-04",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 4,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 137,
      "arxiv_url": "https://arxiv.org/abs/2604.14209",
      "pdf_url": "https://arxiv.org/pdf/2604.14209.pdf"
    },
    {
      "id": "2604.03814",
      "anchor": "paper-2604-03814",
      "title": "InCaRPose: In-Cabin Relative Camera Pose Estimation Model and Dataset",
      "authors": [
        "Felix Stillger",
        "Lukas Hahn",
        "Frederik Hasecke",
        "Tobias Meisen"
      ],
      "abstract": "Camera extrinsic calibration is a fundamental task in computer vision. However, precise relative pose estimation in constrained, highly distorted environments, such as in-cabin automotive monitoring (ICAM), remains challenging. We present InCaRPose, a Transformer-based architecture designed for robust relative pose prediction between image pairs, which can be used for camera extrinsic calibration. By leveraging frozen backbone features such as DINOv3 and a Transformer-based decoder, our model effectively captures the geometric relationship between a reference and a target view. Unlike traditional methods, our approach achieves absolute metric-scale translation within the physically plausible adjustment range of in-cabin camera mounts in a single inference step, which is critical for ICAM, where accurate real-world distances are required for safety-relevant perception. We specifically address the challenges of highly distorted fisheye cameras in automotive interiors by training exclusively on synthetic data. Our model is capable of generalization to real-world cabin environments without relying on the exact same camera intrinsics and additionally achieves competitive performance on the public 7-Scenes dataset. Despite having limited training data, InCaRPose maintains high precision in both rotation and translation, even with a ViT-Small backbone. This enables real-time performance for time-critical inference, such as driver monitoring in supervised autonomous driving. We release our real-world In-Cabin-Pose test dataset consisting of highly distorted vehicle-interior images and our code at https://github.com/felixstillger/InCaRPose.",
      "short_abstract": "Camera extrinsic calibration is a fundamental task in computer vision. However, precise relative pose estimation in constrained, highly distorted environments, such as in-cabin automotive monitoring (ICAM), remains challenging. We present InCaRPose, a...",
      "published": "2026-04-04",
      "updated": "2026-04-04",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset",
        "Synthetic Data",
        "Hardware / Real-Time",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 5,
        "evidence": [
          "title: pose estimation",
          "abstract: pose estimation"
        ]
      },
      "recency": "archive",
      "age_days": 137,
      "arxiv_url": "https://arxiv.org/abs/2604.03814",
      "pdf_url": "https://arxiv.org/pdf/2604.03814.pdf"
    },
    {
      "id": "2604.03581",
      "anchor": "paper-2604-03581",
      "title": "HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving",
      "authors": [
        "Wenhao Yao",
        "Xinglong Sun",
        "Zhenxin Li",
        "Shiyi Lan",
        "Zi Wang",
        "Jose M. Alvarez",
        "Zuxuan Wu"
      ],
      "abstract": "End-to-end planning has emerged as a dominant paradigm for autonomous driving, where recent models often adopt a scoring-selection framework to choose trajectories from a large set of candidates, with diffusion-based decoding showing strong promise. However, directly selecting from the entire candidate space remains difficult to optimize, and Gaussian perturbations used in diffusion often introduce unrealistic trajectories that complicate the denoising process. In addition, for training these models, reinforcement learning (RL) has shown promise, but existing end-to-end RL approaches typically rely on a single coupled reward without structured signals, limiting optimization effectiveness. To address these challenges, we propose HAD, an end-to-end planning framework with a Hierarchical Diffusion Policy that decomposes planning into a coarse-to-fine process. To improve trajectory generation, we introduce Structure-Preserved Trajectory Expansion, which produces realistic candidates while maintaining kinematic structure. For policy learning, we develop Metric-Decoupled Policy Optimization (MDPO) to enable structured RL optimization across multiple driving objectives. Extensive experiments show that HAD achieves new state-of-the-art performance on both NAVSIM and HUGSIM, outperforming prior arts by a huge margin: +2.3 EPDMS on NAVSIM and +4.9 Route Completion on HUGSIM.",
      "short_abstract": "End-to-end planning has emerged as a dominant paradigm for autonomous driving, where recent models often adopt a scoring-selection framework to choose trajectories from a large set of candidates, with diffusion-based decoding showing strong promise. However,...",
      "published": "2026-04-04",
      "updated": "2026-04-04",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 24,
        "evidence": [
          "title: end-to-end driving",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 137,
      "arxiv_url": "https://arxiv.org/abs/2604.03581",
      "pdf_url": "https://arxiv.org/pdf/2604.03581.pdf"
    },
    {
      "id": "2604.03827",
      "anchor": "paper-2604-03827",
      "title": "Confidence Intervals for Rate Estimation with Importance Sampling in Autonomous Vehicle Evaluation",
      "authors": [
        "Aiyou Chen",
        "Ruixuan Rachel Zhou",
        "Joseph J. Lee",
        "Nicholas Chamandy",
        "Henning Hohnhold"
      ],
      "abstract": "Accounting for both rare events and complex sampling presents challenges when quantifying uncertainty for rate estimation in autonomous vehicle performance evaluation. In this paper, we introduce a statistical formulation of this problem and develop a unified compound Poisson model framework for unbiased rate estimation through the Horvitz Thompson estimator. Though asymptotic theory for the model is available, the inference of confidence intervals (CIs) in the presence of rare events requires new investigation. We also advocate for a new monotonicity criterion for rate CIs--summing the rates of disjoint types of events should produce not only a higher point estimate but also higher confidence bounds than for the individual rates--that facilitates interpretability in real applications. We propose a novel exponential bootstrap (EB) method for CI construction based on a fiducial argument; it satisfies the monotonicity property, while novel extensions of some existing methods do not. Comprehensive numerical studies show that EB performs well for a wide range of settings relevant to our applications. Fast implementation of EB based on saddlepoint approximation is also developed, which may be of independent interest.",
      "short_abstract": "Accounting for both rare events and complex sampling presents challenges when quantifying uncertainty for rate estimation in autonomous vehicle performance evaluation. In this paper, we introduce a statistical formulation of this problem and develop a unified...",
      "published": "2026-04-04",
      "updated": "2026-04-04",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 137,
      "arxiv_url": "https://arxiv.org/abs/2604.03827",
      "pdf_url": "https://arxiv.org/pdf/2604.03827.pdf"
    },
    {
      "id": "2604.03497",
      "anchor": "paper-2604-03497",
      "title": "Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving",
      "authors": [
        "Zilin Huang",
        "Zhengyang Wan",
        "Zihao Sheng",
        "Boyue Wang",
        "Junwei You",
        "Sikai Chen"
      ],
      "abstract": "Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-crafted rewards with semantically grounded signals; however, deploying such simulation-trained policies on real vehicles remains a fundamental challenge, because they rely on simulator-native observations and simulator-coupled action semantics with no counterpart on physical hardware. We identify a general principle: the simulation-to-reality gap decomposes into two largely orthogonal axes, a sensing-and-dynamics domain gap and a task-and-geometry gap, the former closable without real-world policy training by re-projecting real perception and control onto the policy's training manifold. We formalize this as a transfer guarantee that bounds the deployment gap by three independently controllable error terms, and instantiate it as Sim2Real-AD, which combines a Geometric Observation Bridge, a Physics-Aware Action Mapping, a Two-Phase Progressive Training curriculum, and a Real-time Deployment Pipeline. As a proof of concept, a CARLA-trained VLM-guided RL policy is transferred zero-shot to a full-scale battery-electric Ford E-Transit van in Madison, WI, USA, and drives across car-following, obstacle-avoidance, and stop-sign scenarios using no real-world training data. To our knowledge, this is among the first zero-shot closed-loop deployments of a CARLA-trained VLM-guided RL policy on a full-scale real vehicle, and the decomposition offers a principled, broadly applicable route for moving simulation-trained, foundation-model-guided policies into the physical world, supporting energy-efficient intelligent driving on electrified transportation platforms. The demo video, code, and model checkpoint are available at: https://zilin-huang.github.io/Sim2Real-AD-website/.",
      "short_abstract": "Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-crafted rewards with semantically grounded signals; however, deploying such simulation-trained policies on real vehicles...",
      "published": "2026-04-03",
      "updated": "2026-04-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 2,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware"
        ]
      },
      "recency": "archive",
      "age_days": 138,
      "arxiv_url": "https://arxiv.org/abs/2604.03497",
      "pdf_url": "https://arxiv.org/pdf/2604.03497.pdf"
    },
    {
      "id": "2604.03462",
      "anchor": "paper-2604-03462",
      "title": "SpectralSplat: Appearance-Disentangled Feed-Forward Gaussian Splatting for Driving Scenes",
      "authors": [
        "Quentin Herau",
        "Tianshuo Xu",
        "Depu Meng",
        "Jiezhi Yang",
        "Chensheng Peng",
        "Spencer Sherk",
        "Yihan Hu",
        "Wei Zhan"
      ],
      "abstract": "Feed-forward 3D Gaussian Splatting methods have achieved impressive reconstruction quality for autonomous driving scenes, yet they entangle scene geometry with transient appearance properties such as lighting, weather, and time of day. This coupling prevents relighting, appearance transfer, and consistent rendering across multi-traversal data captured under varying environmental conditions. We present SpectralSplat, a method that disentangles appearance from geometry within a feed-forward Gaussian Splatting framework. Our key insight is to factor color prediction into an appearance-agnostic base stream and and appearance-conditioned adapted stream, both produced by a shared MLP conditioned on a global appearance embedding derived from DINOv2 features. To enforce disentanglement, we train with paired observations generated by a hybrid relighting pipeline that combines physics-based intrinsic decomposition with diffusion based generative refinement, and supervise with complementary consistency, reconstruction, cross-appearance, and base color losses. We further introduce an appearance-adaptable temporal history that stores appearance-agnostic features, enabling accumulated Gaussians to be re-rendered under arbitrary target appearances. Experiments demonstrate that SpectralSplat preserves the reconstruction quality of the underlying backbone while enabling controllable appearance transfer and temporally consistent relighting across driving sequences.",
      "short_abstract": "Feed-forward 3D Gaussian Splatting methods have achieved impressive reconstruction quality for autonomous driving scenes, yet they entangle scene geometry with transient appearance properties such as lighting, weather, and time of day. This coupling prevents...",
      "published": "2026-04-03",
      "updated": "2026-04-03",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Prediction & World Models: abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 138,
      "arxiv_url": "https://arxiv.org/abs/2604.03462",
      "pdf_url": "https://arxiv.org/pdf/2604.03462.pdf"
    },
    {
      "id": "2604.03349",
      "anchor": "paper-2604-03349",
      "title": "YOLOv11 Demystified: A Practical Guide to High-Performance Object Detection",
      "authors": [
        "Nikhileswara Rao Sulake"
      ],
      "abstract": "YOLOv11 is the latest iteration in the You Only Look Once (YOLO) series of real-time object detectors, introducing novel architectural modules to improve feature extraction and small-object detection. In this paper, we present a detailed analysis of YOLOv11, including its backbone, neck, and head components. The model key innovations, the C3K2 blocks, Spatial Pyramid Pooling - Fast (SPPF), and C2PSA (Cross Stage Partial with Spatial Attention) modules enhance spatial feature processing while preserving speed. We compare YOLOv11 performance to prior YOLO versions on standard benchmarks, highlighting improvements in mean Average Precision (mAP) and inference speed. Our results demonstrate that YOLOv11 achieves superior accuracy without sacrificing real-time capabilities, making it well-suited for applications in autonomous driving, surveillance, and video analytics.This work formalizes YOLOv11 in a research context, providing a clear reference for future studies.",
      "short_abstract": "YOLOv11 is the latest iteration in the You Only Look Once (YOLO) series of real-time object detectors, introducing novel architectural modules to improve feature extraction and small-object detection. In this paper, we present a detailed analysis of YOLOv11,...",
      "published": "2026-04-03",
      "updated": "2026-04-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 138,
      "arxiv_url": "https://arxiv.org/abs/2604.03349",
      "pdf_url": "https://arxiv.org/pdf/2604.03349.pdf"
    },
    {
      "id": "2604.02930",
      "anchor": "paper-2604-02930",
      "title": "BEVPredFormer: Spatio-temporal Attention for BEV Instance Prediction in Autonomous Driving",
      "authors": [
        "Miguel Antunes-García",
        "Santiago Montiel-Marín",
        "Fabio Sánchez-García",
        "Rodrigo Gutiérrez-Moreno",
        "Rafael Barea",
        "Luis M. Bergasa"
      ],
      "abstract": "A robust awareness of how dynamic scenes evolve is essential for Autonomous Driving systems, as they must accurately detect, track, and predict the behaviour of surrounding obstacles. Traditional perception pipelines that rely on modular architectures tend to suffer from cumulative errors and latency. Instance Prediction models provide a unified solution, performing Bird's-Eye-View segmentation and motion estimation across current and future frames using information directly obtained from different sensors. However, a key challenge in these models lies in the effective processing of the dense spatial and temporal information inherent in dynamic driving environments. This level of complexity demands architectures capable of capturing fine-grained motion patterns and long-range dependencies without compromising real-time performance. We introduce BEVPredFormer, a novel camera-only architecture for BEV instance prediction that uses attention-based temporal processing to improve temporal and spatial comprehension of the scene and relies on an attention-based 3D projection of the camera information. BEVPredFormer employs a recurrent-free design that incorporates gated transformer layers, divided spatio-temporal attention mechanisms, and multi-scale head tasks. Additionally, we incorporate a difference-guided feature extraction module that enhances temporal representations. Extensive ablation studies validate the effectiveness of each architectural component. When evaluated on the nuScenes dataset, BEVPredFormer was on par or surpassed State-Of-The-Art methods, highlighting its potential for robust and efficient Autonomous Driving perception.",
      "short_abstract": "A robust awareness of how dynamic scenes evolve is essential for Autonomous Driving systems, as they must accurately detect, track, and predict the behaviour of surrounding obstacles. Traditional perception pipelines that rely on modular architectures tend to...",
      "published": "2026-04-03",
      "updated": "2026-04-03",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 5,
        "evidence": [
          "abstract: perception",
          "abstract: camera",
          "title: bird's-eye view",
          "abstract: bird's-eye view"
        ]
      },
      "recency": "archive",
      "age_days": 138,
      "arxiv_url": "https://arxiv.org/abs/2604.02930",
      "pdf_url": "https://arxiv.org/pdf/2604.02930.pdf"
    },
    {
      "id": "2604.02714",
      "anchor": "paper-2604-02714",
      "title": "ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving",
      "authors": [
        "Zihao Sheng",
        "Xin Ye",
        "Jingru Luo",
        "Sikai Chen",
        "Liu Ren"
      ],
      "abstract": "End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on expert demonstrations. However, imitation learning inherently limits the model to replicating observed behaviors without exploring diverse driving strategies, leaving it brittle in novel or out-of-distribution scenarios. Reinforcement learning (RL) offers a natural remedy by enabling policy exploration beyond the expert distribution. Yet VLA models, typically trained on offline datasets, lack directly observable state transitions, necessitating a learned world model to anticipate action consequences. In this work, we propose a unified understanding-and-generation framework that leverages world modeling to simultaneously enable meaningful exploration and provide dense supervision. Specifically, we augment trajectory prediction with future RGB and depth image generation as dense world modeling objectives, requiring the model to learn fine-grained visual and geometric representations that substantially enrich the planning backbone. Beyond serving as a supervisory signal, the world model further acts as a source of intrinsic reward for policy exploration: its image prediction uncertainty naturally measures a trajectory's novelty relative to the training distribution, where high uncertainty indicates out-of-distribution scenarios that, if safe, represent valuable learning opportunities. We incorporate this exploration signal into a safety-gated reward and optimize the policy via Group Relative Policy Optimization (GRPO). Experiments on the NAVSIM and nuScenes benchmarks demonstrate the effectiveness of our approach, achieving a state-of-the-art PDMS score of 93.7 and an EPDMS of 88.8 on NAVSIM. The code is available at https://zihaosheng.github.io/ExploreVLA/.",
      "short_abstract": "End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on expert demonstrations. However, imitation learning inherently limits the model to...",
      "published": "2026-04-03",
      "updated": "2026-04-03",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "VLA",
        "World Model",
        "Reinforcement Learning",
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 42,
        "margin": 31,
        "evidence": [
          "title: end-to-end driving",
          "abstract: end-to-end driving",
          "abstract: vision-language-action",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 138,
      "arxiv_url": "https://arxiv.org/abs/2604.02714",
      "pdf_url": "https://arxiv.org/pdf/2604.02714.pdf"
    },
    {
      "id": "2604.02710",
      "anchor": "paper-2604-02710",
      "title": "V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views",
      "authors": [
        "Junwei You",
        "Pei Li",
        "Zhuoyu Jiang",
        "Weizhe Tang",
        "Zilin Huang",
        "Rui Gan",
        "Jiaxi Liu",
        "Yan Zhao",
        "Sikai Chen",
        "Bin Ran"
      ],
      "abstract": "Multimodal large language models (MLLMs) have shown strong potential for autonomous driving, yet existing benchmarks remain largely ego-centric and therefore cannot systematically assess model performance in infrastructure-centric and cooperative driving conditions. In this work, we introduce V2X-QA, a real-world dataset and benchmark for evaluating MLLMs across vehicle-side, infrastructure-side, and cooperative viewpoints. V2X-QA is built around a view-decoupled evaluation protocol that enables controlled comparison under vehicle-only, infrastructure-only, and cooperative driving conditions within a unified multiple-choice question answering (MCQA) framework. The benchmark is organized into a twelve-task taxonomy spanning perception, prediction, and reasoning and planning, and is constructed through expert-verified MCQA annotation to enable fine-grained diagnosis of viewpoint-dependent capabilities. Benchmark results across ten representative state-of-the-art proprietary and open-source models show that viewpoint accessibility substantially affects performance, and infrastructure-side reasoning supports meaningful macroscopic traffic understanding. Results also indicate that cooperative reasoning remains challenging since it requires cross-view alignment and evidence integration rather than simply additional visual input. To address these challenges, we introduce V2X-MoE, a benchmark-aligned baseline with explicit view routing and viewpoint-specific LoRA experts. The strong performance of V2X-MoE further suggests that explicit viewpoint specialization is a promising direction for multi-view reasoning in autonomous driving. Overall, V2X-QA provides a foundation for studying multi-perspective reasoning, reliability, and cooperative physical intelligence in connected autonomous driving. The dataset and V2X-MoE resources are publicly available at: https://github.com/junwei0001/V2X-QA.",
      "short_abstract": "Multimodal large language models (MLLMs) have shown strong potential for autonomous driving, yet existing benchmarks remain largely ego-centric and therefore cannot systematically assess model performance in infrastructure-centric and cooperative driving...",
      "published": "2026-04-03",
      "updated": "2026-04-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Dataset",
        "Benchmark",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 33,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: cooperative systems"
        ]
      },
      "recency": "archive",
      "age_days": 138,
      "arxiv_url": "https://arxiv.org/abs/2604.02710",
      "pdf_url": "https://arxiv.org/pdf/2604.02710.pdf"
    },
    {
      "id": "2604.02639",
      "anchor": "paper-2604-02639",
      "title": "Cross-Vehicle 3D Geometric Consistency for Self-Supervised Surround Depth Estimation on Articulated Vehicles",
      "authors": [
        "Weimin Liu",
        "Jiyuan Qiu",
        "Wenjun Wang",
        "Joshua H. Meng"
      ],
      "abstract": "Surround depth estimation provides a cost-effective alternative to LiDAR for 3D perception in autonomous driving. While recent self-supervised methods explore multi-camera settings to improve scale awareness and scene coverage, they are primarily designed for passenger vehicles and rarely consider articulated vehicles or robotics platforms. The articulated structure introduces complex cross-segment geometry and motion coupling, making consistent depth reasoning across views more challenging. In this work, we propose \\textbf{ArticuSurDepth}, a self-supervised framework for surround-view depth estimation on articulated vehicles that enhances depth learning through cross-view and cross-vehicle geometric consistency guided by structural priors from vision foundation model. Specifically, we introduce multi-view spatial context enrichment strategy and a cross-view surface normal constraint to improve structural coherence across spatial and temporal contexts. We further incorporate camera height regularization with ground plane-awareness to encourage metric depth estimation, together with cross-vehicle pose consistency that bridges motion estimation between articulated segments. To validate our proposed method, an articulated vehicle experiment platform was established with a dataset collected over it. Experiment results demonstrate state-of-the-art (SoTA) performance of depth estimation on our self-collected dataset as well as on DDAD, nuScenes, and KITTI benchmarks.",
      "short_abstract": "Surround depth estimation provides a cost-effective alternative to LiDAR for 3D perception in autonomous driving. While recent self-supervised methods explore multi-camera settings to improve scale awareness and scene coverage, they are primarily designed for...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 20,
        "evidence": [
          "abstract: perception",
          "abstract: LiDAR",
          "abstract: camera",
          "title: depth estimation",
          "abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.02639",
      "pdf_url": "https://arxiv.org/pdf/2604.02639.pdf"
    },
    {
      "id": "2604.02603",
      "anchor": "paper-2604-02603",
      "title": "Rascene: High-Fidelity 3D Scene Imaging with mmWave Communication Signals",
      "authors": [
        "Kunzhe Song",
        "Geo Jie Zhou",
        "Xiaoming Liu",
        "Huacheng Zeng"
      ],
      "abstract": "Robust 3D environmental perception is critical for applications such as autonomous driving and robot navigation. However, optical sensors such as cameras and LiDAR often fail under adverse conditions, including smoke, fog, and non-ideal lighting. Although specialized radar systems can operate in these environments, their reliance on bespoke hardware and licensed spectrum limits scalability and cost-effectiveness. This paper introduces Rascene, an integrated sensing and communication (ISAC) framework that leverages ubiquitous mmWave OFDM communication signals for 3D scene imaging. To overcome the sparse and multipath-ambiguous nature of individual radio frames, Rascene performs multi-frame, spatially adaptive fusion with confidence-weighted forward projection, enabling the recovery of geometric consensus across arbitrary poses. Experimental results demonstrate that our method reconstructs 3D scenes with high precision, offering a new pathway toward low-cost, scalable, and robust 3D perception.",
      "short_abstract": "Robust 3D environmental perception is critical for applications such as autonomous driving and robot navigation. However, optical sensors such as cameras and LiDAR often fail under adverse conditions, including smoke, fog, and non-ideal lighting. Although...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 11,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception",
          "Perception & Sensor Fusion: abstract: LiDAR",
          "Perception & Sensor Fusion: abstract: radar",
          "Perception & Sensor Fusion: abstract: camera",
          "Systems, Deployment & Connectivity: abstract: hardware",
          "Systems, Deployment & Connectivity: title: communications",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.02603",
      "pdf_url": "https://arxiv.org/pdf/2604.02603.pdf"
    },
    {
      "id": "2604.02573",
      "anchor": "paper-2604-02573",
      "title": "Dynamic Risk Generation for Autonomous Driving: Naturalistic Reconstruction of Vehicle-E-Scooter Interactions",
      "authors": [
        "Abin Mathew",
        "Zhitong He",
        "Lingxi Li",
        "Yaobin Chen"
      ],
      "abstract": "The increasing, high-risk interactions between vehicles and vulnerable micromobility users, such as e-scooter riders, challenge vehicular safety functions and Automated Driving (AD) techniques, often resulting in severe consequences due to the dynamic uncertainty of e-scooter motion. Despite advances in data-driven AD methods, traffic data addressing the e-scooter interaction problem, particularly for safety-critical moments, remains underdeveloped. This paper proposes a pipeline that utilizes collected on-road traffic data and creates configurable synthetic interactions for validating vehicle motion planning algorithms. A Social Force Model (SFM) is applied to offer more dynamic and potentially risky movements for the e-scooter, thereby testing the functionality and reliability of the vehicle collision avoidance systems. A case study based on a real-world interaction scenario was conducted to verify the practicality and effectiveness of the established simulator. Simulation experiments successfully demonstrate the capability of extending the target scenario to more critical interactions that may result in a potential collision.",
      "short_abstract": "The increasing, high-risk interactions between vehicles and vulnerable micromobility users, such as e-scooter riders, challenge vehicular safety functions and Automated Driving (AD) techniques, often resulting in severe consequences due to the dynamic...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 9,
        "margin": 1,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.02573",
      "pdf_url": "https://arxiv.org/pdf/2604.02573.pdf"
    },
    {
      "id": "2604.02279",
      "anchor": "paper-2604-02279",
      "title": "The Self Driving Portfolio: Agentic Architecture for Institutional Asset Management",
      "authors": [
        "Andrew Ang",
        "Nazym Azimbayev",
        "Andrey Kim"
      ],
      "abstract": "Agentic AI shifts the investor's role from analytical execution to oversight. We present an agentic strategic asset allocation pipeline in which approximately 50 specialized agents produce capital market assumptions, construct portfolios using over 20 competing methods, and critique and vote on each other's output. A researcher agent proposes new portfolio construction methods not yet represented, and a meta-agent compares past forecasts against realized returns and rewrites agent code and prompts to improve future performance. The entire pipeline is governed by the Investment Policy Statement--the same document that guides human portfolio managers can now constrain and direct autonomous agents.",
      "short_abstract": "Agentic AI shifts the investor's role from analytical execution to oversight. We present an agentic strategic asset allocation pipeline in which approximately 50 specialized agents produce capital market assumptions, construct portfolios using over 20...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Human Factors & Policy: abstract: law or policy"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.02279",
      "pdf_url": "https://arxiv.org/pdf/2604.02279.pdf"
    },
    {
      "id": "2604.03325",
      "anchor": "paper-2604-03325",
      "title": "Safety-Aligned 3D Object Detection: Single-Vehicle, Cooperative, and End-to-End Perspectives",
      "authors": [
        "Brian Hsuan-Cheng Liao",
        "Chih-Hong Cheng",
        "Hasan Esen",
        "Alois Knoll"
      ],
      "abstract": "Perception plays a central role in connected and autonomous vehicles (CAVs), underpinning not only conventional modular driving stacks, but also cooperative perception systems and recent end-to-end driving models. While deep learning has greatly improved perception performance, its statistical nature makes perfect predictions difficult to attain. Meanwhile, standard training objectives and evaluation benchmarks treat all perception errors equally, even though only a subset is safety-critical. In this paper, we investigate safety-aligned evaluation and optimization for 3D object detection that explicitly characterize high-impact errors. Building on our previously proposed safety-oriented metric, NDS-USC, and safety-aware loss function, EC-IoU, we make three contributions. First, we present an expanded study of single-vehicle 3D object detection models across diverse neural network architectures and sensing modalities, showing that gains under standard metrics such as mAP and NDS may not translate to safety-oriented criteria represented by NDS-USC. With EC-IoU, we reaffirm the benefit of safety-aware fine-tuning for improving safety-critical detection performance. Second, we conduct an ego-centric, safety-oriented evaluation of AV-infrastructure cooperative object detection models, underscoring its superiority over vehicle-only models and demonstrating a safety impact analysis that illustrates the potential contribution of cooperative models to \"Vision Zero.\" Third, we integrate EC-IoU into SparseDrive and show that safety-aware perception hardening can reduce collision rate by nearly 30% and improve system-level safety directly in an end-to-end perception-to-planning framework. Overall, our results indicate that safety-aligned perception evaluation and optimization offer a practical path toward enhancing CAV safety across single-vehicle, cooperative, and end-to-end autonomy settings.",
      "short_abstract": "Perception plays a central role in connected and autonomous vehicles (CAVs), underpinning not only conventional modular driving stacks, but also cooperative perception systems and recent end-to-end driving models. While deep learning has greatly improved...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 19,
        "margin": 1,
        "evidence": [
          "abstract: perception",
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.03325",
      "pdf_url": "https://arxiv.org/pdf/2604.03325.pdf"
    },
    {
      "id": "2604.02282",
      "anchor": "paper-2604-02282",
      "title": "Deep Neural Network Based Roadwork Detection for Autonomous Driving",
      "authors": [
        "Sebastian Wullrich",
        "Nicolai Steinke",
        "Daniel Goehring"
      ],
      "abstract": "Road construction sites create major challenges for both autonomous vehicles and human drivers due to their highly dynamic and heterogeneous nature. This paper presents a real-time system that detects and localizes roadworks by combining a YOLO neural network with LiDAR data. The system identifies individual roadwork objects while driving, merges them into coherent construction sites and records their outlines in world coordinates. The model training was based on an adapted US dataset and a new dataset collected from test drives with a prototype vehicle in Berlin, Germany. Evaluations on real-world road construction sites showed a localization accuracy below 0.5 m. The system can support traffic authorities with up-to-date roadwork data and could enable autonomous vehicles to navigate construction sites more safely in the future.",
      "short_abstract": "Road construction sites create major challenges for both autonomous vehicles and human drivers due to their highly dynamic and heterogeneous nature. This paper presents a real-time system that detects and localizes roadworks by combining a YOLO neural network...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 6,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.02282",
      "pdf_url": "https://arxiv.org/pdf/2604.02282.pdf"
    },
    {
      "id": "2604.02128",
      "anchor": "paper-2604-02128",
      "title": "SEAL: An Open, Auditable, and Fair Data Generation Framework for AI-Native 6G Networks",
      "authors": [
        "Sunder Ali Khowaja",
        "Kapal Dev",
        "Engin Zeydan",
        "Madhusanka Liyanage"
      ],
      "abstract": "AI-native 6G networks promise to transform the telecom industry by enabling dynamic resource allocation, predictive maintenance, and ultra-reliable low-latency communications across all layers, which are essential for applications such as smart cities, autonomous vehicles, and immersive XR. However, the deployment of 6G systems results in severe data scarcity, hindering the training of efficient AI models. Synthetic data generation is extensively used to fill this gap; however, it introduces challenges related to dataset bias, auditability, and compliance with regulatory frameworks. In this regard, we propose the Synthetic Data Generation with Ethics Audit Loop (SEAL) framework, which extends baseline modular pipelines with an Ethical and Regulatory Compliance by Design (ERCD) module and a Federated Learning (FL) feedback system. The ERCD integrates fairness, bias detection, and standardized audit trails for regulatory mapping, while the FL enables privacy-preserving calibration using aggregated insights from real testbeds to close the reality-simulation gap. Results show that the SEAL framework outperforms existing methods in terms of Frechet Inception Distance, equalized odds, and accuracy. These results validate the framework's ability to generate auditable and bias-mitigated synthetic data for responsible AI-native 6G development.",
      "short_abstract": "AI-native 6G networks promise to transform the telecom industry by enabling dynamic resource allocation, predictive maintenance, and ultra-reliable low-latency communications across all layers, which are essential for applications such as smart cities,...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Dataset",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "medium",
        "score": 13,
        "margin": 9,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.02128",
      "pdf_url": "https://arxiv.org/pdf/2604.02128.pdf"
    },
    {
      "id": "2604.01791",
      "anchor": "paper-2604-01791",
      "title": "PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency",
      "authors": [
        "Leezy Han",
        "Seunggyu Kim",
        "Dongseok Shim",
        "Hyeonbeom Lee"
      ],
      "abstract": "Monocular depth estimation (MDE) has been widely adopted in the perception systems of autonomous vehicles and mobile robots. However, existing approaches often struggle to maintain temporal consistency in depth estimation across consecutive frames. This inconsistency not only causes jitter but can also lead to estimation failures when the depth range changes abruptly. To address these challenges, this paper proposes a consistency-aware monocular depth estimation framework that leverages wheel odometry from a mobile robot to achieve stable and coherent depth predictions over time. Specifically, we estimate camera pose and sparse depth from triangulation using optical flow between consecutive frames. The sparse depth estimates are used to update a recursive Bayesian estimate of the metric scale, which is then applied to rescale the relative depth predicted by a pre-trained depth estimation foundation model. The proposed method is evaluated on the KITTI, TartanAir, MS2, and our own dataset, demonstrating robust and accurate depth estimation performance.",
      "short_abstract": "Monocular depth estimation (MDE) has been widely adopted in the perception systems of autonomous vehicles and mobile robots. However, existing approaches often struggle to maintain temporal consistency in depth estimation across consecutive frames. This...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 13,
        "evidence": [
          "abstract: perception",
          "abstract: camera",
          "title: depth estimation",
          "abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.01791",
      "pdf_url": "https://arxiv.org/pdf/2604.01791.pdf"
    },
    {
      "id": "2604.01790",
      "anchor": "paper-2604-01790",
      "title": "Set-Theoretic Receding Horizon Control for Obstacle Avoidance and Overtaking in Autonomous Highway Driving",
      "authors": [
        "Gianni Cario",
        "Valentino Carriuolo",
        "Alessandro Casavola",
        "Gianfranco Gagliardi",
        "Marco Lupia",
        "Franco Angelo Torchiaro"
      ],
      "abstract": "This article addresses obstacle avoidance motion planning for autonomous vehicles, specifically focusing on highway overtaking maneuvers. The control design challenge is handled by considering a mathematical vehicle model that captures both lateral and longitudinal dynamics. Unlike existing numerical optimization methods that suffer from significant online computational overhead, this work extends the state-of-the-art by leveraging a fast set-theoretic ellipsoidal Model Predictive Control (Fast-MPC) technique. While originally restricted to stabilization tasks, the proposed framework is successfully adapted to handle motion planning for vehicles modeled as uncertain polytopic discrete-time linear systems. The control action is computed online via a set-membership evaluation against a structured sequence of nested inner ellipsoidal approximations of the exact one-step ahead controllable set within a receding horizon framework. A six-degrees-of-freedom (6-DOF) nonlinear model characterizes the vehicle dynamics, while a polytopic embedding approximates the nonlinearities within a linear framework with parameter uncertainties. Finally, to assess performance and real-time feasibility, comparative co-simulations against a baseline Non-Linear MPC (NLMPC) were conducted. Using the high-fidelity CARLA 3D simulator, results demonstrate that the proposed approach seamlessly rejects dynamic traffic disturbances while reducing online computational time by over 90% compared to standard optimization-based approaches.",
      "short_abstract": "This article addresses obstacle avoidance motion planning for autonomous vehicles, specifically focusing on highway overtaking maneuvers. The control design challenge is handled by considering a mathematical vehicle model that captures both lateral and...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 10,
        "evidence": [
          "abstract: model predictive control",
          "abstract: vehicle dynamics",
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.01790",
      "pdf_url": "https://arxiv.org/pdf/2604.01790.pdf"
    },
    {
      "id": "2604.01696",
      "anchor": "paper-2604-01696",
      "title": "A Graph Neural Network Approach for Solving the Ranked Assignment Problem in Multi-Object Tracking",
      "authors": [
        "Robin Dehler",
        "Martin Herrmann",
        "Jan Strohbeck",
        "Michael Buchholz"
      ],
      "abstract": "Associating measurements with tracks is a crucial step in Multi-Object Tracking (MOT) to guarantee the safety of autonomous vehicles. To manage the exponentially growing number of track hypotheses, truncation becomes necessary. In the $δ$-Generalized Labeled Multi-Bernoulli ($δ$-GLMB) filter application, this truncation typically involves the ranked assignment problem, solved by Murty's algorithm or the Gibbs sampling approach, both with limitations in terms of complexity or accuracy, respectively. With the motivation to improve these limitations, this paper addresses the ranked assignment problem arising from data association tasks with an approach that employs Graph Neural Networks (GNNs). The proposed Ranked Assignment Prediction Graph Neural Network (RAPNet) uses bipartite graphs to model the problem, harnessing the computational capabilities of deep learning. The conclusive evaluation compares the RAPNet with Murty's algorithm and the Gibbs sampler, showing accuracy improvements compared to the Gibbs sampler.",
      "short_abstract": "Associating measurements with tracks is a crucial step in Multi-Object Tracking (MOT) to guarantee the safety of autonomous vehicles. To manage the exponentially growing number of track hypotheses, truncation becomes necessary. In the $δ$-Generalized Labeled...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.01696",
      "pdf_url": "https://arxiv.org/pdf/2604.01696.pdf"
    },
    {
      "id": "2604.02206",
      "anchor": "paper-2604-02206",
      "title": "LEO: Graph Attention Network based Hybrid Multi Sensor Extended Object Fusion and Tracking for Autonomous Driving Applications",
      "authors": [
        "Mayank Mayank",
        "Bharanidhar Duraisamy",
        "Florian Geiss"
      ],
      "abstract": "Accurate shape and trajectory estimation of dynamic objects is essential for reliable automated driving. Classical Bayesian extended-object models offer theoretical robustness and efficiency but depend on completeness of a-priori and update-likelihood functions, while deep learning methods bring adaptability at the cost of dense annotations and high compute. We bridge these strengths with LEO (Learned Extension of Objects), a spatio-temporal Graph Attention Network that fuses multi-modal production-grade sensor tracks to learn adaptive fusion weights, ensure temporal consistency, and represent multi-scale shapes. Using a task-specific parallelogram ground-truth formulation, LEO models complex geometries (e.g. articulated trucks and trailers) and generalizes across sensor types, configurations, object classes, and regions, remaining robust for challenging and long-range targets. Evaluations on the Mercedes-Benz DRIVE PILOT SAE L3 dataset demonstrate real-time computational efficiency suitable for production systems; additional validation on public datasets such as View of Delft (VoD) further confirms cross-dataset generalization.",
      "short_abstract": "Accurate shape and trajectory estimation of dynamic objects is essential for reliable automated driving. Classical Bayesian extended-object models offer theoretical robustness and efficiency but depend on completeness of a-priori and update-likelihood...",
      "published": "2026-04-02",
      "updated": "2026-04-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 6,
        "evidence": [
          "abstract: deployment or real-time",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 139,
      "arxiv_url": "https://arxiv.org/abs/2604.02206",
      "pdf_url": "https://arxiv.org/pdf/2604.02206.pdf"
    },
    {
      "id": "2604.01466",
      "anchor": "paper-2604-01466",
      "title": "Efficient Equivariant Transformer for Self-Driving Agent Modeling",
      "authors": [
        "Scott Xu",
        "Dian Chen",
        "Kelvin Wong",
        "Chris Zhang",
        "Kion Fallah",
        "Raquel Urtasun"
      ],
      "abstract": "Accurately modeling agent behaviors is an important task in self-driving. It is also a task with many symmetries, such as equivariance to the order of agents and objects in the scene or equivariance to arbitrary roto-translations of the entire scene as a whole; i.e., SE(2)-equivariance. The transformer architecture is a ubiquitous tool for modeling these symmetries. While standard self-attention is inherently permutation equivariant, explicit pairwise relative positional encodings have been the standard for introducing SE(2)-equivariance. However, this approach introduces an additional cost that is quadratic in the number of agents, limiting its scalability to larger scenes and batch sizes. In this work, we propose DriveGATr, a novel transformer-based architecture for agent modeling that achieves SE(2)-equivariance without the computational cost of existing methods. Inspired by recent advances in geometric deep learning, DriveGATr encodes scene elements as multivectors in the 2D projective geometric algebra $\\mathbb{R}^*_{2,0,1}$ and processes them with a stack of equivariant transformer blocks. Crucially, DriveGATr models geometric relationships using standard attention between multivectors, eliminating the need for costly explicit pairwise relative positional encodings. Experiments on the Waymo Open Motion Dataset demonstrate that DriveGATr is comparable to the state-of-the-art in traffic simulation and establishes a superior Pareto front for performance vs computational cost.",
      "short_abstract": "Accurately modeling agent behaviors is an important task in self-driving. It is also a task with many symmetries, such as equivariance to the order of agents and objects in the scene or equivariance to arbitrary roto-translations of the entire scene as a...",
      "published": "2026-04-01",
      "updated": "2026-04-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 140,
      "arxiv_url": "https://arxiv.org/abs/2604.01466",
      "pdf_url": "https://arxiv.org/pdf/2604.01466.pdf"
    },
    {
      "id": "2604.01346",
      "anchor": "paper-2604-01346",
      "title": "Safety, Security, and Cognitive Risks in World Models",
      "authors": [
        "Manoj Parmar"
      ],
      "abstract": "World models - learned internal simulators of environment dynamics - are rapidly becoming foundational to autonomous decision-making in robotics, autonomous vehicles, and agentic AI. By predicting future states in compressed latent spaces, they enable sample-efficient planning and long-horizon imagination without direct environment interaction. Yet this predictive power introduces a distinctive set of safety, security, and cognitive risks. Adversaries can corrupt training data, poison latent representations, and exploit compounding rollout errors to cause significant degradation in safety-critical deployments. At the alignment layer, world model-equipped agents are more capable of goal misgeneralisation, deceptive alignment, and reward hacking. At the human layer, authoritative world model predictions foster automation bias, miscalibrated trust, and planning hallucination. This paper surveys the world model landscape; introduces formal definitions of trajectory persistence and representational risk; presents a five-profile attacker taxonomy; and develops a unified threat model drawing on MITRE ATLAS and the OWASP LLM Top 10. We provide an empirical proof-of-concept demonstrating trajectory-persistent adversarial attacks on a GRU-based RSSM ($\\mathcal{A}_1 = 2.26\\times$ amplification, $-59.5\\%$ reward reduction under adversarial fine-tuning), validate architecture-dependence via a stochastic RSSM proxy ($\\mathcal{A}_1 = 0.65\\times$), and probe a real DreamerV3 checkpoint (non-zero action drift confirmed). We propose interdisciplinary mitigations spanning adversarial hardening, alignment engineering, NIST AI RMF and EU AI Act governance, and human-factors design, arguing that world models require the same rigour as flight-control software or medical devices.",
      "short_abstract": "World models - learned internal simulators of environment dynamics - are rapidly becoming foundational to autonomous decision-making in robotics, autonomous vehicles, and agentic AI. By predicting future states in compressed latent spaces, they enable...",
      "published": "2026-04-01",
      "updated": "2026-04-01",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 40,
        "margin": 24,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: security",
          "abstract: security",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 140,
      "arxiv_url": "https://arxiv.org/abs/2604.01346",
      "pdf_url": "https://arxiv.org/pdf/2604.01346.pdf"
    },
    {
      "id": "2604.01254",
      "anchor": "paper-2604-01254",
      "title": "Simulating Realistic LiDAR Data Under Adverse Weather for Autonomous Vehicles: A Physics-Informed Learning Approach",
      "authors": [
        "Vivek Anand",
        "Bharat Lohani",
        "Rakesh Mishra",
        "Gaurav Pandey"
      ],
      "abstract": "Accurate LiDAR simulation is crucial for autonomous driving, especially under adverse weather conditions. Existing methods struggle to capture the complex interactions between LiDAR signals and atmospheric phenomena, leading to unrealistic representations. This paper presents a physics-informed learning framework (PICWGAN) for generating realistic LiDAR data under adverse weather conditions. By integrating physicsdriven constraints for modeling signal attenuation and geometryconsistent degradations into a physics-informed learning pipeline, the proposed method reduces the sim-to-real gap. Evaluations on real-world datasets (CADC for snow, Boreas for rain) and the VoxelScape dataset show that our approach closely mimics realworld intensity patterns. Quantitative metrics, including MSE, SSIM, KL divergence, and Wasserstein distance, demonstrate statistically consistent intensity distributions. Additionally, models trained on data enhanced by our framework outperform baselines in downstream 3D object detection, achieving performance comparable to models trained on real-world data. These results highlight the effectiveness of the proposed approach in improving the realism of LiDAR data and enabling robust perception under adverse weather conditions.",
      "short_abstract": "Accurate LiDAR simulation is crucial for autonomous driving, especially under adverse weather conditions. Existing methods struggle to capture the complex interactions between LiDAR signals and atmospheric phenomena, leading to unrealistic representations....",
      "published": "2026-04-01",
      "updated": "2026-04-01",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 17,
        "evidence": [
          "abstract: perception",
          "abstract: object detection",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 140,
      "arxiv_url": "https://arxiv.org/abs/2604.01254",
      "pdf_url": "https://arxiv.org/pdf/2604.01254.pdf"
    },
    {
      "id": "2604.00832",
      "anchor": "paper-2604-00832",
      "title": "Steering through Time: Blending Longitudinal Data with Simulation to Rethink Human-Autonomous Vehicle Interaction",
      "authors": [
        "Yasaman Hakiminejad",
        "Shiva Azimi",
        "Luis Gomero",
        "Elizabeth Pantesco",
        "Irene P. Kan",
        "Meltem Izzetoglu",
        "Arash Tavakoli"
      ],
      "abstract": "As semi-automated vehicles (SAVs) become more common, ensuring effective human-vehicle interaction during control handovers remains a critical safety challenge. Existing studies often rely on single-session simulator experiments or naturalistic driving datasets, which often lack temporal context on drivers' cognitive and physiological states before takeover events. This study introduces a hybrid framework combining longitudinal mobile sensing with high-fidelity driving simulation to examine driver readiness in semi-automated contexts. In a pilot study with 38 participants, we collected 7 days of wearable physiological data and daily surveys on stress, arousal, valence, and sleep quality, followed by an in-lab simulation with scripted takeover events under varying secondary task conditions. Multimodal sensing, including eye tracking, fNIRS, and physiological measures, captured real-time responses. Preliminary analysis shows the framework's feasibility and individual variability in baseline and in-task measures; for example, fixation duration and takeover control time differed by task type, and RMSSD showed high inter-individual stability. This proof-of-concept supports the development of personalized, context-aware driver monitoring by linking temporally layered data with real-time performance.",
      "short_abstract": "As semi-automated vehicles (SAVs) become more common, ensuring effective human-vehicle interaction during control handovers remains a critical safety challenge. Existing studies often rely on single-session simulator experiments or naturalistic driving...",
      "published": "2026-04-01",
      "updated": "2026-04-01",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Hardware / Real-Time",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 10,
        "margin": 2,
        "evidence": [
          "abstract: driver monitoring",
          "abstract: takeover"
        ]
      },
      "recency": "archive",
      "age_days": 140,
      "arxiv_url": "https://arxiv.org/abs/2604.00832",
      "pdf_url": "https://arxiv.org/pdf/2604.00832.pdf"
    },
    {
      "id": "2604.00621",
      "anchor": "paper-2604-00621",
      "title": "Heterogeneous Mean Field Game Framework for LEO Satellite-Assisted V2X Networks",
      "authors": [
        "Kangkang Sun",
        "Jianhua Li",
        "Xiuzhen Chen",
        "Mingzhe Chen",
        "Minyi Guo"
      ],
      "abstract": "Coordinating mixed fleets of massive vehicles under stringent delay constraints is a central scalability bottleneck in next-generation mobile computing networks, especially when passenger cars, freight trucks, and autonomous vehicles share the same radio and multi-access edge computing (MEC) infrastructure. Heterogeneous mean field games (HMFG) are a principled framework for this setting, but a fundamental design question remains open: how many agent types should be used for a fleet of size $N$? The difficulty is a two-sided trade-off that existing theory does not resolve: using more types improves heterogeneity representation, but it reduces per-class sample size and weakens the mean-field approximation accuracy. This paper resolves that trade-off through an explicit $\\varepsilon$-Nash error decomposition, a closed-form type-selection law, a heterogeneity-aware equilibrium solver, and a robust extension to time-varying LEO backhaul dynamics. For the 1D queue state space, the optimal type count satisfies $K^*(N)=Θ(N^{1/3})$; for the joint queue-channel model ($d=2$), the scaling becomes $K^*(N)=Θ(N^{1/5})$ with logarithmic correction. The unified formula $K^*(N)=Θ(N^{α/(α+β)})$ provides dimension-dependent design guidance, reducing type granularity to a principled, set-once system parameter rather than a per-deployment tuning burden. Experiments validate the 1D scaling law with empirical slope $0.334 \\pm 0.004$, achieve $2.3\\times$ faster PDHG convergence at $K=5$, and deliver up to $29.5\\%$ lower delay and $60\\%$ higher throughput than homogeneous baselines. Unlike model-free DRL methods whose training complexity scales with the state-action space, the proposed HMFG solver has per-iteration complexity $O(K^2 N_q N_t)$ independent of fleet size $N$, making it suitable for large-scale mobile edge computing deployment.",
      "short_abstract": "Coordinating mixed fleets of massive vehicles under stringent delay constraints is a central scalability bottleneck in next-generation mobile computing networks, especially when passenger cars, freight trucks, and autonomous vehicles share the same radio and...",
      "published": "2026-04-01",
      "updated": "2026-04-01",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 31,
        "evidence": [
          "title: V2X",
          "abstract: deployment or real-time",
          "abstract: edge or cloud",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 140,
      "arxiv_url": "https://arxiv.org/abs/2604.00621",
      "pdf_url": "https://arxiv.org/pdf/2604.00621.pdf"
    },
    {
      "id": "2604.08588",
      "anchor": "paper-2604-08588",
      "title": "Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models",
      "authors": [
        "Matthew DosSantos DiSorbo",
        "Harang Ju"
      ],
      "abstract": "Effective automation hinges on deciding when to act and when to escalate. We model this as a decision under uncertainty: an LLM forms a prediction, estimates its probability of being correct, and compares the expected costs of acting and escalating. Using this framework across five domains of recorded human decisions-demand forecasting, content recommendation, content moderation, loan approval, and autonomous driving-and across multiple model families, we find marked differences in the implicit thresholds models use to trade off these costs. These thresholds vary substantially and are not predicted by architecture or scale, while self-estimates are miscalibrated in model-specific ways. We then test interventions that target this decision process by varying cost ratios, providing accuracy signals, and training models to follow the desired escalation rule. Prompting helps mainly for reasoning models. SFT on chain-of-thought targets yields the most robust policies, which generalize across datasets, cost ratios, prompt framings, and held-out domains. These results suggest that escalation behavior is a model-specific property that should be characterized before deployment, and that robust alignment benefits from training models to reason explicitly about uncertainty and decision costs.",
      "short_abstract": "Effective automation hinges on deciding when to act and when to escalate. We model this as a decision under uncertainty: an LLM forms a prediction, estimates its probability of being correct, and compares the expected costs of acting and escalating. Using...",
      "published": "2026-03-31",
      "updated": "2026-03-31",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Prediction & World Models: abstract: forecasting",
          "Prediction & World Models: abstract: prediction",
          "Safety, Security & Verification: abstract: robustness",
          "Safety, Security & Verification: abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 141,
      "arxiv_url": "https://arxiv.org/abs/2604.08588",
      "pdf_url": "https://arxiv.org/pdf/2604.08588.pdf"
    },
    {
      "id": "2604.00069",
      "anchor": "paper-2604-00069",
      "title": "Perspective: Towards sustainable exploration of chemical spaces with machine learning",
      "authors": [
        "Leonardo Medrano Sandonas",
        "David Balcells",
        "Anton Bochkarev",
        "Jacqueline M. Cole",
        "Volker L. Deringer",
        "Werner Dobrautz",
        "Adrian Ehrenhofer",
        "Thorben Frank",
        "Pascal Friederich",
        "Rico Friedrich",
        "Janine George",
        "Luca Ghiringhelli",
        "Alejandra Hinostroza Caldas",
        "Veronika Juraskova",
        "Hannes Kneiding",
        "Yury Lysogorskiy",
        "Johannes T. Margraf",
        "Hanna Türk",
        "Anatole von Lilienfeld",
        "Milica Todorović",
        "Alexandre Tkatchenko",
        "Mariana Rossi",
        "Gianaurelio Cuniberti"
      ],
      "abstract": "Artificial intelligence is transforming molecular and materials science, but its growing computational and data demands raise critical sustainability challenges. In this Perspective, we examine resource considerations across the AI-driven discovery pipeline--from quantum-mechanical (QM) data generation and model training to automated, self-driving research workflows--building on discussions from the ``SusML workshop: Towards sustainable exploration of chemical spaces with machine learning'' held in Dresden, Germany. In this context, the availability of large quantum datasets has enabled rigorous benchmarking and rapid methodological progress, while also incurring substantial energy and infrastructure costs. We highlight emerging strategies to enhance efficiency, including general-purpose machine learning (ML) models, multi-fidelity approaches, model distillation, and active learning. Moreover, incorporating physics-based constraints within hierarchical workflows, where fast ML surrogates are applied broadly and high-accuracy QM methods are used selectively, can further optimize resource use without compromising reliability. Equally important is bridging the gap between idealized computational predictions and real-world conditions by accounting for synthesizability and multi-objective design criteria, which is essential for practical impact. Finally, we argue that sustainable progress will rely on open data and models, reusable workflows, and domain-specific AI systems that maximize scientific value per unit of computation, enabling efficient and responsible discovery of technological materials and therapeutics.",
      "short_abstract": "Artificial intelligence is transforming molecular and materials science, but its growing computational and data demands raise critical sustainability challenges. In this Perspective, we examine resource considerations across the AI-driven discovery...",
      "published": "2026-03-31",
      "updated": "2026-03-31",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 141,
      "arxiv_url": "https://arxiv.org/abs/2604.00069",
      "pdf_url": "https://arxiv.org/pdf/2604.00069.pdf"
    },
    {
      "id": "2604.16406",
      "anchor": "paper-2604-16406",
      "title": "Heterogeneous Self-Play for Realistic Highway Traffic Simulation",
      "authors": [
        "Jinkai Qiu",
        "Alessandro Saviolo",
        "Chaojie Wang",
        "Mingke Wang",
        "Xiaoyu Huang"
      ],
      "abstract": "Realistic highway simulation is critical for scalable safety evaluation of autonomous vehicles, particularly for interactions that are too rare to study from logged data alone. Yet highway traffic generation remains challenging because it requires broad coverage across speeds and maneuvers, controllable generation of rare safety-critical scenarios, and behavioral credibility in multi-agent interactions. We present PHASE, Policy for Heterogeneous Agent Self-play on Expressway, a context-aware self-play framework that addresses these three requirements through explicit per-agent conditioning for controllability, synthetic scenario generation for broad highway coverage, and closed-loop multi-agent training for realistic interaction dynamics. PHASE further supports different vehicle profiles, for example, passenger cars and articulated trailer trucks, within a single policy via vehicle-aware dynamics and context-conditioned actions, and stabilizes self-play with early termination of unrecoverable states, at-fault collision attribution, highway-aware reward shaping, coupled curricula, and robust policy optimization. Despite being trained only on synthetic data, PHASE transfers zero-shot to 512 unseen high-interaction real scenarios in exiD, achieving a 96.3% success rate and reducing ADE/FDE from 6.57/12.07 m to 2.44/5.25 m relative to a prior self-play baseline. In a learned trajectory embedding space, it also improves behavioral realism over IDM, reducing Frechet trajectory distance by 13.1% and energy distance by 20.2%. These results show that synthetic self-play can provide a scalable route to controllable and realistic highway scenario generation without direct imitation of expert logs.",
      "short_abstract": "Realistic highway simulation is critical for scalable safety evaluation of autonomous vehicles, particularly for interactions that are too rare to study from logged data alone. Yet highway traffic generation remains challenging because it requires broad...",
      "published": "2026-03-31",
      "updated": "2026-03-31",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Simulation",
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 2,
        "evidence": [
          "abstract: safety",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 141,
      "arxiv_url": "https://arxiv.org/abs/2604.16406",
      "pdf_url": "https://arxiv.org/pdf/2604.16406.pdf"
    },
    {
      "id": "2604.00150",
      "anchor": "paper-2604-00150",
      "title": "Data-Driven Reachability of Nonlinear Lipschitz Systems via Koopman Operator Embeddings",
      "authors": [
        "Alireza Naderi",
        "Ahmad Hafez",
        "Abdulla Fawzy",
        "Amr Alanwar"
      ],
      "abstract": "Data-driven safety verification of robotic systems often relies on zonotopic reachability analysis due to its scalability and computational efficiency. However, for nonlinear systems, these methods can become overly conservative, especially over long prediction horizons and under measurement noise. We propose a data-driven reachability framework based on the Koopman operator and zonotopic set representations that lifts the nonlinear system into a finite-dimensional, linear, state-input-dependent model. Reachable sets are then computed in the lifted space and projected back to the original state space to obtain guaranteed over-approximations of the true dynamics. The proposed method reduces conservatism while preserving formal safety guarantees, and we prove that the resulting reachable sets over-approximate the true reachable sets. Numerical simulations and real-world experiments on an autonomous vehicle show that the proposed approach yields substantially tighter reachable set over-approximations than both model-based and linear data-driven methods, particularly over long horizons.",
      "short_abstract": "Data-driven safety verification of robotic systems often relies on zonotopic reachability analysis due to its scalability and computational efficiency. However, for nonlinear systems, these methods can become overly conservative, especially over long...",
      "published": "2026-03-31",
      "updated": "2026-03-31",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "abstract: verification"
        ]
      },
      "recency": "archive",
      "age_days": 141,
      "arxiv_url": "https://arxiv.org/abs/2604.00150",
      "pdf_url": "https://arxiv.org/pdf/2604.00150.pdf"
    },
    {
      "id": "2603.29135",
      "anchor": "paper-2603-29135",
      "title": "Quality-Controlled Active Learning via Gaussian Processes for Robust Structure-Property Learning in Autonomous Microscopy",
      "authors": [
        "Jawad Chowdhury",
        "Ganesh Narasimha",
        "Jan-Chi Yang",
        "Yongtao Liu",
        "Rama Vasudevan"
      ],
      "abstract": "Autonomous experimental systems are increasingly used in materials research to accelerate scientific discovery, but their performance is often limited by low-quality, noisy data. This issue is especially problematic in data-intensive structure-property learning tasks such as Image-to-Spectrum (Im2Spec) and Spectrum-to-Image (Spec2Im) translations, where standard active learning strategies can mistakenly prioritize poor-quality measurements. We introduce a gated active learning framework that combines curiosity-driven sampling with a physics-informed quality control filter based on the Simple Harmonic Oscillator model fits, allowing the system to automatically exclude low-fidelity data during acquisition. Evaluations on a pre-acquired dataset of band-excitation piezoresponse spectroscopy (BEPS) data from PbTiO3 thin films with spatially localized noise show that the proposed method outperforms random sampling, standard active learning, and multitask learning strategies. The gated approach enhances both Im2Spec and Spec2Im by handling noise during training and acquisition, leading to more reliable forward and inverse predictions. In contrast, standard active learners often misinterpret noise as uncertainty and end up acquiring bad samples that hurt performance. Given its promising applicability, we further deployed the framework in real-time experiments on BiFeO3 thin films, demonstrating its effectiveness in real autonomous microscopy experiments. Overall, this work supports a shift toward hybrid autonomy in self-driving labs, where physics-informed quality assessment and active decision-making work hand-in-hand for more reliable discovery.",
      "short_abstract": "Autonomous experimental systems are increasingly used in materials research to accelerate scientific discovery, but their performance is often limited by low-quality, noisy data. This issue is especially problematic in data-intensive structure-property...",
      "published": "2026-03-30",
      "updated": "2026-03-30",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: robustness",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 142,
      "arxiv_url": "https://arxiv.org/abs/2603.29135",
      "pdf_url": "https://arxiv.org/pdf/2603.29135.pdf"
    },
    {
      "id": "2603.28889",
      "anchor": "paper-2603-28889",
      "title": "DeepEye: A Steerable Self-driving Data Agent System",
      "authors": [
        "Boyan Li",
        "Yiran Peng",
        "Yupeng Xie",
        "Sirong Lu",
        "Yizhang Zhu",
        "Xing Mu",
        "Xinyu Liu",
        "Yuyu Luo"
      ],
      "abstract": "Large Language Models (LLMs) have revolutionized natural language interaction with data. The \"holy grail\" of data analytics is to build autonomous Data Agents that can self-drive complex data analysis workflows. However, current implementations are still limited to linear \"ChatBI\" systems. These systems struggle with joint analysis across heterogeneous data sources (e.g., databases, documents, and data files) and often encounter \"context explosion\" in complex and iterative data analysis workflows. To address these challenges, we present DeepEye, a production-ready data agent system that adopts a workflow-centric architecture to ensure scalability and trustworthiness. DeepEye introduces a Unified Multimodal Orchestration protocol, enabling seamless integration of structured and unstructured data sources. To mitigate hallucinations, it employs Hierarchical Reasoning with context isolation, decomposing complex intents into autonomous AgentNodes and deterministic ToolNodes. Furthermore, DeepEye incorporates a database-inspired Workflow Engine (comprising a Compiler, Validator, Optimizer, and Executor) that guarantees structural correctness and accelerates execution via runtime topological optimization. In this demonstration, we showcase DeepEye's ability to orchestrate complex workflows to generate diverse multimodal outputs -- including Data Videos, Dashboards, and Analytical Reports -- highlighting its advantages in transparent execution, automated optimization, and human-in-the-loop reliability.",
      "short_abstract": "Large Language Models (LLMs) have revolutionized natural language interaction with data. The \"holy grail\" of data analytics is to build autonomous Data Agents that can self-drive complex data analysis workflows. However, current implementations are still...",
      "published": "2026-03-30",
      "updated": "2026-03-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Human Factors & Policy: abstract: human-machine interaction"
        ]
      },
      "recency": "archive",
      "age_days": 142,
      "arxiv_url": "https://arxiv.org/abs/2603.28889",
      "pdf_url": "https://arxiv.org/pdf/2603.28889.pdf"
    },
    {
      "id": "2603.28965",
      "anchor": "paper-2603-28965",
      "title": "Koopman Operator Framework for Modeling and Control of Off-Road Vehicle on Deformable Terrain",
      "authors": [
        "Kartik Loya",
        "Phanindra Tallapragada"
      ],
      "abstract": "This work presents a hybrid physics-informed and data-driven modeling framework for predictive control of autonomous off-road vehicles operating on deformable terrain. Traditional high-fidelity terramechanics models are often too computationally demanding to be directly used in control design. Modern Koopman operator methods can be used to represent the complex terramechanics and vehicle dynamics in a linear form. We develop a framework whereby a Koopman linear system can be constructed using data from simulations of a vehicle moving on deformable terrain. For vehicle simulations, the deformable-terrain terramechanics are modeled using Bekker-Wong theory, and the vehicle is represented as a simplified five-degree-of-freedom (5-DOF) system. The Koopman operators are identified from large simulation datasets for sandy loam and clay using a recursive subspace identification method, where Grassmannian distance is used to prioritize informative data segments during training. The advantage of this approach is that the Koopman operator learned from simulations can be updated with data from the physical system in a seamless manner, making this a hybrid physics-informed and data-driven approach. Prediction results demonstrate stable short-horizon accuracy and robustness under mild terrain-height variations. When embedded in a constrained MPC, the learned predictor enables stable closed-loop tracking of aggressive maneuvers while satisfying steering and torque limits.",
      "short_abstract": "This work presents a hybrid physics-informed and data-driven modeling framework for predictive control of autonomous off-road vehicles operating on deformable terrain. Traditional high-fidelity terramechanics models are often too computationally demanding to...",
      "published": "2026-03-30",
      "updated": "2026-03-30",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 18,
        "evidence": [
          "abstract: model predictive control",
          "abstract: vehicle dynamics",
          "title: control",
          "abstract: control",
          "abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 142,
      "arxiv_url": "https://arxiv.org/abs/2603.28965",
      "pdf_url": "https://arxiv.org/pdf/2603.28965.pdf"
    },
    {
      "id": "2603.28888",
      "anchor": "paper-2603-28888",
      "title": "A Semantic Observer Layer for Autonomous Vehicles: Pre-Deployment Feasibility Study of VLMs for Low-Latency Anomaly Detection",
      "authors": [
        "Kunal Runwal",
        "Swaraj Gajare",
        "Daniel Adejumo",
        "Omkar Ankalkope",
        "Siddhant Baroth",
        "Aliasghar Arab"
      ],
      "abstract": "Semantic anomalies-context-dependent hazards that pixel-level detectors cannot reason about-pose a critical safety risk in autonomous driving. We propose a \\emph{semantic observer layer}: a quantized vision-language model (VLM) running at 1--2\\,Hz alongside the primary AV control loop, monitoring for semantic edge cases, and triggering fail-safe handoffs when detected. Using Nvidia Cosmos-Reason1-7B with NVFP4 quantization and FlashAttention2, we achieve ~500 ms inference a ~50x speedup over the unoptimized FP16 baseline (no quantization, standard PyTorch attention) on the same hardware--satisfying the observer timing budget. We benchmark accuracy, latency, and quantization behavior in static and video conditions, identify NF4 recall collapse (10.6%) as a hard deployment constraint, and a hazard analysis mapping performance metrics to safety goals. The results establish a pre-deployment feasibility case for the semantic observer architecture on embodied-AI AV platforms.",
      "short_abstract": "Semantic anomalies-context-dependent hazards that pixel-level detectors cannot reason about-pose a critical safety risk in autonomous driving. We propose a \\emph{semantic observer layer}: a quantized vision-language model (VLM) running at 1--2\\,Hz alongside...",
      "published": "2026-03-30",
      "updated": "2026-03-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 19,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time",
          "abstract: hardware",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 142,
      "arxiv_url": "https://arxiv.org/abs/2603.28888",
      "pdf_url": "https://arxiv.org/pdf/2603.28888.pdf"
    },
    {
      "id": "2603.28625",
      "anchor": "paper-2603-28625",
      "title": "Dynamic Lookahead Distance via Reinforcement Learning-Based Pure Pursuit for Autonomous Racing",
      "authors": [
        "Mohamed Elgouhary",
        "Amr S. El-Wakeel"
      ],
      "abstract": "Pure Pursuit (PP) is a widely used path-tracking algorithm in autonomous vehicles due to its simplicity and real-time performance. However, its effectiveness is sensitive to the choice of lookahead distance: shorter values improve cornering but can cause instability on straights, while longer values improve smoothness but reduce accuracy in curves. We propose a hybrid control framework that integrates Proximal Policy Optimization (PPO) with the classical Pure Pursuit controller to adjust the lookahead distance dynamically during racing. The PPO agent maps vehicle speed and multi-horizon curvature features to an online lookahead command. It is trained using Stable-Baselines3 in the F1TENTH Gym simulator with a KL penalty and learning-rate decay for stability, then deployed in a ROS2 environment to guide the controller. Experiments in simulation compare the proposed method against both fixed-lookahead Pure Pursuit and an adaptive Pure Pursuit baseline. Additional real-car experiments compare the learned controller against a fixed-lookahead Pure Pursuit controller. Results show that the learned policy improves lap-time performance and repeated lap completion on unseen tracks, while also transferring zero-shot to hardware. The learned controller adapts the lookahead by increasing it on straights and reducing it in curves, demonstrating effectiveness in augmenting a classical controller by online adaptation of a single interpretable parameter. On unseen tracks, the proposed method achieved 33.16 s on Montreal and 46.05 s on Yas Marina, while tolerating more aggressive speed-profile scaling than the baselines and achieving the best lap times among the tested settings. Initial real-car experiments further support sim-to-real transfer on a 1:10-scale autonomous racing platform",
      "short_abstract": "Pure Pursuit (PP) is a widely used path-tracking algorithm in autonomous vehicles due to its simplicity and real-time performance. However, its effectiveness is sensitive to the choice of lookahead distance: shorter values improve cornering but can cause...",
      "published": "2026-03-30",
      "updated": "2026-03-30",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Simulation",
        "Reinforcement Learning",
        "Hardware / Real-Time",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 1,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware"
        ]
      },
      "recency": "archive",
      "age_days": 142,
      "arxiv_url": "https://arxiv.org/abs/2603.28625",
      "pdf_url": "https://arxiv.org/pdf/2603.28625.pdf"
    },
    {
      "id": "2603.28333",
      "anchor": "paper-2603-28333",
      "title": "Integrating Multimodal Large Language Model Knowledge into Amodal Completion",
      "authors": [
        "Heecheol Yun",
        "Eunho Yang"
      ],
      "abstract": "With the widespread adoption of autonomous vehicles and robotics, amodal completion, which reconstructs the occluded parts of people and objects in an image, has become increasingly crucial. Just as humans infer hidden regions based on prior experience and common sense, this task inherently requires physical knowledge about real-world entities. However, existing approaches either depend solely on the image generation ability of visual generative models, which lack such knowledge, or leverage it only during the segmentation stage, preventing it from explicitly guiding the completion process. To address this, we propose AmodalCG, a novel framework that harnesses the real-world knowledge of Multimodal Large Language Models (MLLMs) to guide amodal completion. Our framework first assesses the extent of occlusion to selectively invoke MLLM guidance only when the target object is heavily occluded. If guidance is required, the framework further incorporates MLLMs to reason about both the (1) extent and (2) content of the missing regions. Finally, a visual generative model integrates these guidance and iteratively refines imperfect completions that may arise from inaccurate MLLM guidance. Experimental results on various real-world images show impressive improvements compared to all existing works, suggesting MLLMs as a promising direction for addressing challenging amodal completion.",
      "short_abstract": "With the widespread adoption of autonomous vehicles and robotics, amodal completion, which reconstructs the occluded parts of people and objects in an image, has become increasingly crucial. Just as humans infer hidden regions based on prior experience and...",
      "published": "2026-03-30",
      "updated": "2026-03-30",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 142,
      "arxiv_url": "https://arxiv.org/abs/2603.28333",
      "pdf_url": "https://arxiv.org/pdf/2603.28333.pdf"
    },
    {
      "id": "2603.27912",
      "anchor": "paper-2603-27912",
      "title": "Safety Guardrails in the Sky: Realizing Control Barrier Functions on the VISTA F-16 Jet",
      "authors": [
        "Andrew W. Singletary",
        "Max H. Cohen",
        "Tamas G. Molnar",
        "Aaron D. Ames"
      ],
      "abstract": "The advancement of autonomous systems -- from legged robots to self-driving vehicles and aircraft -- necessitates executing increasingly high-performance and dynamic motions without ever putting the system or its environment in harm's way. In this paper, we introduce Guardrails -- a novel runtime assurance mechanism that guarantees dynamic safety for autonomous systems, allowing them to safely evolve on the edge of their operational domains. Rooted in the theory of control barrier functions, Guardrails offers a control strategy that carefully blends commands from a human or AI operator with safe control actions to guarantee safe behavior. To demonstrate its capabilities, we implemented Guardrails on an F-16 fighter jet and conducted flight tests where Guardrails supervised a human pilot to enforce g-limits, altitude bounds, geofence constraints, and combinations thereof. Throughout extensive flight testing, Guardrails successfully ensured safety, keeping the pilot in control when safe to do so and minimally modifying unsafe pilot inputs otherwise.",
      "short_abstract": "The advancement of autonomous systems -- from legged robots to self-driving vehicles and aircraft -- necessitates executing increasingly high-performance and dynamic motions without ever putting the system or its environment in harm's way. In this paper, we...",
      "published": "2026-03-29",
      "updated": "2026-03-29",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 11,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 143,
      "arxiv_url": "https://arxiv.org/abs/2603.27912",
      "pdf_url": "https://arxiv.org/pdf/2603.27912.pdf"
    },
    {
      "id": "2603.27933",
      "anchor": "paper-2603-27933",
      "title": "From Passersby to Placemaking: Designing Autonomous Vehicle-Pedestrian Encounters for an Urban Shared Space",
      "authors": [
        "Yiyuan Wang",
        "Martin Tomitsch",
        "Marius Hoggenmüller",
        "Senuri Wijenayake",
        "Wai Yan",
        "Luke Hespanhol"
      ],
      "abstract": "Autonomous vehicles (AVs) tend to disrupt the atmosphere and pedestrian experience in urban shared spaces, undermining the focus of these spaces on people and placemaking. We investigate how external human-machine interfaces (eHMIs) supporting AV-pedestrian interaction can be extended to consider the characteristics of an urban shared space. Inspired by urban HCI, we devised three place-based eHMI designs that (i) enhance a conventional intent eHMI and (ii) exhibit content and physical integration with the space. In an evaluation study, 25 participants experienced the eHMIs in an immersive simulation of the space via virtual reality and shared their impressions through think-aloud, interviews, and questionnaires. Results showed that the place-based eHMIs had a notable effect on influencing the perception of AV interaction, including aspects like visual aesthetics and sense of reassurance, and on fostering a sense of place, such as social interactivity and the intentionality to coexist. In measuring qualities of pedestrian experience, we found that perceived safety significantly correlated with user experience and affect, including the attractiveness of eHMIs and feelings of pleasantness. The paper opens the avenue for exploring how eHMIs may contribute to the placemaking goals of pedestrian-centric spaces and improve the experience of people encountering AVs within these environments.",
      "short_abstract": "Autonomous vehicles (AVs) tend to disrupt the atmosphere and pedestrian experience in urban shared spaces, undermining the focus of these spaces on people and placemaking. We investigate how external human-machine interfaces (eHMIs) supporting AV-pedestrian...",
      "published": "2026-03-29",
      "updated": "2026-03-29",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: human-machine interaction",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 143,
      "arxiv_url": "https://arxiv.org/abs/2603.27933",
      "pdf_url": "https://arxiv.org/pdf/2603.27933.pdf"
    },
    {
      "id": "2603.27418",
      "anchor": "paper-2603-27418",
      "title": "Buzz Buzz: Haptic Cuing of Road Conditions in Autonomous Cars for Drivers Engaged in Secondary Tasks",
      "authors": [
        "Shivam Pandey"
      ],
      "abstract": "Can drivers' situation awareness during automated driving be maintained using haptic cues that provide information about road and traffic scenarios while the drivers are engaged in a secondary task? And can this be done without disengaging them from the secondary task? Multiple Resource Theory predicts that using different sensory channels can improve multiple-task performance. Using haptics to provide information avoids the audio-visual channels likely occupied by the secondary task. An experiment was conducted to assess whether drivers' situation awareness could be maintained using haptic cues. Drivers played Fruit Ninja as the secondary task while seated in a driving simulator with a Level 4 autonomous system driving. A mixed design was used for the experiment with the presence of haptic cues and the presentation time of situation awareness questions as the between-subjects conditions. Five road and traffic scenarios comprised the within-subjects part of the design. Subjects who received haptic cues had a higher number of correct responses to the situation awareness questions and looked up at the simulator screen fewer times than those who were not provided cues. Subjects did not find the cues to be disruptive and gave good satisfaction scores to the haptic device. Additionally, subjects across all conditions seemed to have performed equally well in playing Fruit Ninja. It appears that haptic cuing can maintain drivers' situation awareness during automated driving while drivers are engaged in a secondary task. Practical implications of these findings for implementing haptic cues in autonomous vehicles are also discussed.",
      "short_abstract": "Can drivers' situation awareness during automated driving be maintained using haptic cues that provide information about road and traffic scenarios while the drivers are engaged in a secondary task? And can this be done without disengaging them from the...",
      "published": "2026-03-28",
      "updated": "2026-03-28",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 144,
      "arxiv_url": "https://arxiv.org/abs/2603.27418",
      "pdf_url": "https://arxiv.org/pdf/2603.27418.pdf"
    },
    {
      "id": "2603.27207",
      "anchor": "paper-2603-27207",
      "title": "Autonomous overtaking trajectory optimization using reinforcement learning and opponent pose estimation",
      "authors": [
        "Matej Rene Cihlar",
        "Luka Šiktar",
        "Branimir Ćaran",
        "Marko Švaco"
      ],
      "abstract": "Vehicle overtaking is one of the most complex driving maneuvers for autonomous vehicles. To achieve optimal autonomous overtaking, driving systems rely on multiple sensors that enable safe trajectory optimization and overtaking efficiency. This paper presents a reinforcement learning mechanism for multi-agent autonomous racing environments, enabling overtaking trajectory optimization, based on LiDAR and depth image data. The developed reinforcement learning agent uses pre-generated raceline data and sensor inputs to compute the steering angle and linear velocity for optimal overtaking. The system uses LiDAR with a 2D detection algorithm and a depth camera with YOLO-based object detection to identify the vehicle to be overtaken and its pose. The LiDAR and the depth camera detection data are fused using a UKF for improved opponent pose estimation and trajectory optimization for overtaking in racing scenarios. The results show that the proposed algorithm successfully performs overtaking maneuvers in both simulation and real-world experiments, with pose estimation RMSE of (0.0816, 0.0531) m in (x, y).",
      "short_abstract": "Vehicle overtaking is one of the most complex driving maneuvers for autonomous vehicles. To achieve optimal autonomous overtaking, driving systems rely on multiple sensors that enable safe trajectory optimization and overtaking efficiency. This paper presents...",
      "published": "2026-03-28",
      "updated": "2026-03-28",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 7,
        "evidence": [
          "title: pose estimation",
          "abstract: pose estimation"
        ]
      },
      "recency": "archive",
      "age_days": 144,
      "arxiv_url": "https://arxiv.org/abs/2603.27207",
      "pdf_url": "https://arxiv.org/pdf/2603.27207.pdf"
    },
    {
      "id": "2603.27108",
      "anchor": "paper-2603-27108",
      "title": "MotiMem: Motion-Aware Approximate Memory for Energy-Efficient Neural Perception in Autonomous Vehicles",
      "authors": [
        "Haohua Que",
        "Mingkai Liu",
        "Jiayue Xie",
        "Haojia Gao",
        "Jiajun Sun",
        "Hongyi Xu",
        "Handong Yao",
        "Fei Qiao"
      ],
      "abstract": "High-resolution sensors are critical for robust autonomous perception but impose a severe memory wall on battery-constrained electric vehicles. In these systems, data movement energy often outweighs computation. Traditional image compression is ill-suited as it is semantically blind and optimizes for storage rather than bus switching activity. We propose MotiMem, a hardware-software co-designed interface. Exploiting temporal coherence,MotiMem uses lightweight 2D Motion Propagation to dynamically identify Regions of Interest (RoI). Complementing this, a Hybrid Sparsity-Aware Coding scheme leverages adaptive inversion and truncation to induce bitlevel sparsity. Extensive experiments across nuScenes, Waymo, and KITTI with 16 detection models demonstrate that MotiMem reduces memory-interface dynamic energy by approximately 43 percent while retaining approximately 93 percent of the object detection accuracy, establishing a new Pareto frontier significantly superior to standard codecs like JPEG and WebP.",
      "short_abstract": "High-resolution sensors are critical for robust autonomous perception but impose a severe memory wall on battery-constrained electric vehicles. In these systems, data movement energy often outweighs computation. Traditional image compression is ill-suited as...",
      "published": "2026-03-27",
      "updated": "2026-03-27",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 13,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 145,
      "arxiv_url": "https://arxiv.org/abs/2603.27108",
      "pdf_url": "https://arxiv.org/pdf/2603.27108.pdf"
    },
    {
      "id": "2603.26462",
      "anchor": "paper-2603-26462",
      "title": "DTP-Attack: A decision-based black-box adversarial attack on trajectory prediction",
      "authors": [
        "Jiaxiang Li",
        "Jun Yan",
        "Daniel Watzenig",
        "Huilin Yin"
      ],
      "abstract": "Trajectory prediction systems are critical for autonomous vehicle safety, yet remain vulnerable to adversarial attacks that can cause catastrophic traffic behavior misinterpretations. Existing attack methods require white-box access with gradient information and rely on rigid physical constraints, limiting real-world applicability. We propose DTP-Attack, a decision-based black-box adversarial attack framework tailored for trajectory prediction systems. Our method operates exclusively on binary decision outputs without requiring model internals or gradients, making it practical for real-world scenarios. DTP-Attack employs a novel boundary walking algorithm that navigates adversarial regions without fixed constraints, naturally maintaining trajectory realism through proximity preservation. Unlike existing approaches, our method supports both intention misclassification attacks and prediction accuracy degradation. Extensive evaluation on nuScenes and Apolloscape datasets across state-of-the-art models including Trajectron++ and Grip++ demonstrates superior performance. DTP-Attack achieves 41 - 81% attack success rates for intention misclassification attacks that manipulate perceived driving maneuvers with perturbations below 0.45 m, and increases prediction errors by 1.9 - 4.2 for accuracy degradation. Our method consistently outperforms existing black-box approaches while maintaining high controllability and reliability across diverse scenarios. These results reveal fundamental vulnerabilities in current trajectory prediction systems, highlighting urgent needs for robust defenses in safety-critical autonomous driving applications.",
      "short_abstract": "Trajectory prediction systems are critical for autonomous vehicle safety, yet remain vulnerable to adversarial attacks that can cause catastrophic traffic behavior misinterpretations. Existing attack methods require white-box access with gradient information...",
      "published": "2026-03-27",
      "updated": "2026-03-27",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 28,
        "margin": 6,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 145,
      "arxiv_url": "https://arxiv.org/abs/2603.26462",
      "pdf_url": "https://arxiv.org/pdf/2603.26462.pdf"
    },
    {
      "id": "2603.26343",
      "anchor": "paper-2603-26343",
      "title": "Hermes Seal: Zero-Knowledge Assurance for Autonomous Vehicle Communications",
      "authors": [
        "Munawar Hasan",
        "Apostol Vassilev",
        "Edward Griffor",
        "Thoshitha Gamage"
      ],
      "abstract": "The application of zero-knowledge proofs (ZKPs) in autonomous systems is an emerging area of research, motivated by the growing need for regulatory compliance, transparent auditing, and trustworthy operation in decentralized environments. zk-SNARK is a powerful cryptographic tool that allows a party (the prover) to prove to another party (the verifier) that a statement about its own internal state is true, without revealing sensitive or proprietary data about that state. This paper proposes Hermes Seal: a zk-SNARK-based ZKP framework for enabling privacy-preserving, verifiable communication in vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) networks. The framework allows autonomous systems to generate cryptographic proofs of perception and decision-related computations without revealing proprietary models, sensor data, or internal system states, thereby supporting interoperability across heterogeneous autonomous systems. We present two real-world case studies implemented and empirically evaluated within our framework, demonstrating a step toward verifiable autonomous system information exchanges. The first demonstrates real-time proof generation and verification, achieving 8 ms proof generation and 1 ms verification on a GPU, while the second evaluates the performance of an autonomous vehicle perception stack, enabling proof of computation without exposing proprietary or confidential data. Furthermore, the framework can be integrated into AV perception stacks to facilitate verifiable interoperability and privacy-preserving cooperative perception. The demonstration code for this project is open source, available on Github.",
      "short_abstract": "The application of zero-knowledge proofs (ZKPs) in autonomous systems is an emerging area of research, motivated by the growing need for regulatory compliance, transparent auditing, and trustworthy operation in decentralized environments. zk-SNARK is a...",
      "published": "2026-03-27",
      "updated": "2026-03-27",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 9,
        "evidence": [
          "abstract: V2X",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 145,
      "arxiv_url": "https://arxiv.org/abs/2603.26343",
      "pdf_url": "https://arxiv.org/pdf/2603.26343.pdf"
    },
    {
      "id": "2603.25968",
      "anchor": "paper-2603-25968",
      "title": "Neuro-Cognitive Reward Modeling for Human-Centered Autonomous Vehicle Control",
      "authors": [
        "Zhuoli Zhuang",
        "Yu-Cheng Chang",
        "Yu-Kai Wang",
        "Thomas Do",
        "Chin-Teng Lin"
      ],
      "abstract": "Recent advancements in computer vision have accelerated the development of autonomous driving. Despite these advancements, training machines to drive in a way that aligns with human expectations remains a significant challenge. Human factors are still essential, as humans possess a sophisticated cognitive system capable of rapidly interpreting scene information and making accurate decisions. Aligning machine with human intent has been explored with Reinforcement Learning with Human Feedback (RLHF). Conventional RLHF methods rely on collecting human preference data by manually ranking generated outputs, which is time-consuming and indirect. In this work, we propose an electroencephalography (EEG)-guided decision-making framework to incorporate human cognitive insights without behaviour response interruption into reinforcement learning (RL) for autonomous driving. We collected EEG signals from 20 participants in a realistic driving simulator and analyzed event-related potentials (ERP) in response to sudden environmental changes. Our proposed framework employs a neural network to predict the strength of ERP based on the cognitive information from visual scene information. Moreover, we explore the integration of such cognitive information into the reward signal of the RL algorithm. Experimental results show that our framework can improve the collision avoidance ability of the RL algorithm, highlighting the potential of neuro-cognitive feedback in enhancing autonomous driving systems. Our project page is: https://alex95gogo.github.io/Cognitive-Reward/.",
      "short_abstract": "Recent advancements in computer vision have accelerated the development of autonomous driving. Despite these advancements, training machines to drive in a way that aligns with human expectations remains a significant challenge. Human factors are still...",
      "published": "2026-03-26",
      "updated": "2026-03-26",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Simulation",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 18,
        "margin": 12,
        "evidence": [
          "title: vehicle control",
          "title: control"
        ]
      },
      "recency": "archive",
      "age_days": 146,
      "arxiv_url": "https://arxiv.org/abs/2603.25968",
      "pdf_url": "https://arxiv.org/pdf/2603.25968.pdf"
    },
    {
      "id": "2603.25963",
      "anchor": "paper-2603-25963",
      "title": "BEVMAPMATCH: Multimodal BEV Neural Map Matching for Robust Re-Localization of Autonomous Vehicles",
      "authors": [
        "Shounak Sural",
        "Ragunathan Rajkumar"
      ],
      "abstract": "Localization in GNSS-denied and GNSS-degraded environments is a challenge for the safe widespread deployment of autonomous vehicles. Such GNSS-challenged environments require alternative methods for robust localization. In this work, we propose BEVMapMatch, a framework for robust vehicle re-localization on a known map without the need for GNSS priors. BEVMapMatch uses a context-aware lidar+camera fusion method to generate multimodal Bird's Eye View (BEV) segmentations around the ego vehicle in both good and adverse weather conditions. Leveraging a search mechanism based on cross-attention, the generated BEV segmentation maps are then used for the retrieval of candidate map patches for map-matching purposes. Finally, BEVMapMatch uses the top retrieved candidate for finer alignment against the generated BEV segmentation, achieving accurate global localization without the need for GNSS. Multiple frames of generated BEV segmentation further improve localization accuracy. Extensive evaluations show that BEVMapMatch outperforms existing methods for re-localization in GNSS-denied and adverse environments, with a Recall@1m of 39.8%, being nearly twice as much as the best performing re-localization baseline. Our code and data will be made available at https://github.com/ssuralcmu/BEVMapMatch.git.",
      "short_abstract": "Localization in GNSS-denied and GNSS-degraded environments is a challenge for the safe widespread deployment of autonomous vehicles. Such GNSS-challenged environments require alternative methods for robust localization. In this work, we propose BEVMapMatch, a...",
      "published": "2026-03-26",
      "updated": "2026-03-26",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 9,
        "evidence": [
          "title: localization",
          "abstract: localization",
          "abstract: GNSS or GPS"
        ]
      },
      "recency": "archive",
      "age_days": 146,
      "arxiv_url": "https://arxiv.org/abs/2603.25963",
      "pdf_url": "https://arxiv.org/pdf/2603.25963.pdf"
    },
    {
      "id": "2603.25672",
      "anchor": "paper-2603-25672",
      "title": "Can Users Specify Driving Speed? Bench2Drive-Speed: Benchmark and Baselines for Desired-Speed Conditioned Autonomous Driving",
      "authors": [
        "Yuqian Shao",
        "Xiaosong Jia",
        "Langechuan Liu",
        "Junchi Yan"
      ],
      "abstract": "End-to-end autonomous driving (E2E-AD) has achieved remarkable progress. However, one practical and useful function has been long overlooked: users may wish to customize the desired speed of the policy or specify whether to allow the autonomous vehicle to overtake. To bridge this gap, we present Bench2Drive-Speed, a benchmark with metrics, dataset, and baselines for desired-speed conditioned autonomous driving. We introduce explicit inputs of users' desired target-speed and overtake/follow instructions to driving policy models. We design quantitative metrics, including Speed-Adherence Score and Overtake Score, to measure how faithfully policies follow user specifications, while remaining compatible with standard autonomous driving metrics. To enable training of speed-conditioned policies, one approach is to collect expert demonstrations that strictly follow speed requirements, an expensive and unscalable process in the real world. An alternative is to adapt existing regular driving data by treating the speed observed in future frames as the target speed for training. To investigate this, we construct CustomizedSpeedDataset, composed of 2,100 clips annotated with experts demonstrations, enabling systematic investigation of supervision strategies. Our experiments show that, under proper re-annotation, models trained on regular driving data perform comparably to on expert demonstrations, suggesting that speed supervision can be introduced without additional complex real-world data collection. Furthermore, we find that while target-speed following can be achieved without degrading regular driving performance, executing overtaking commands remains challenging due to the inherent difficulty of interactive behaviors. All code, datasets and baselines are available at https://github.com/Thinklab-SJTU/Bench2Drive-Speed",
      "short_abstract": "End-to-end autonomous driving (E2E-AD) has achieved remarkable progress. However, one practical and useful function has been long overlooked: users may wish to customize the desired speed of the policy or specify whether to allow the autonomous vehicle to...",
      "published": "2026-03-26",
      "updated": "2026-03-26",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset",
        "Benchmark",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end",
          "abstract: driving policy"
        ]
      },
      "recency": "archive",
      "age_days": 146,
      "arxiv_url": "https://arxiv.org/abs/2603.25672",
      "pdf_url": "https://arxiv.org/pdf/2603.25672.pdf"
    },
    {
      "id": "2603.25275",
      "anchor": "paper-2603-25275",
      "title": "V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative Perception",
      "authors": [
        "Weijia Li",
        "Haoen Xiang",
        "Tianxu Wang",
        "Shuaibing Wu",
        "Qiming Xia",
        "Cheng Wang",
        "Chenglu Wen"
      ],
      "abstract": "Modern autonomous vehicle perception systems are often constrained by occlusions, blind spots, and limited sensing range. While existing cooperative perception paradigms, such as Vehicle-to-Vehicle (V2V) and Vehicle-to-Infrastructure (V2I), have demonstrated their effectiveness in mitigating these challenges, they remain limited to ground-level collaboration and cannot fully address large-scale occlusions or long-range perception in complex environments. To advance research in cross-view cooperative perception, we present V2U4Real, the first large-scale real-world multi-modal dataset for Vehicle-to-UAV (V2U) cooperative object perception. V2U4Real is collected by a ground vehicle and a UAV equipped with multi-view LiDARs and RGB cameras. The dataset covers urban streets, university campuses, and rural roads under diverse traffic scenarios, comprising over 56K LiDAR frames, 56K multi-view camera images, and 700K annotated 3D bounding boxes across four classes. To support a wide range of research tasks, we establish benchmarks for single-agent 3D object detection, cooperative 3D object detection, and object tracking. Comprehensive evaluations of several state-of-the-art models demonstrate the effectiveness of V2U cooperation in enhancing perception robustness and long-range awareness. The V2U4Real dataset and codebase is available at https://github.com/VjiaLi/V2U4Real.",
      "short_abstract": "Modern autonomous vehicle perception systems are often constrained by occlusions, blind spots, and limited sensing range. While existing cooperative perception paradigms, such as Vehicle-to-Vehicle (V2V) and Vehicle-to-Infrastructure (V2I), have demonstrated...",
      "published": "2026-03-26",
      "updated": "2026-03-26",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset",
        "Benchmark",
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 21,
        "margin": 3,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: object detection",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 146,
      "arxiv_url": "https://arxiv.org/abs/2603.25275",
      "pdf_url": "https://arxiv.org/pdf/2603.25275.pdf"
    },
    {
      "id": "2603.24931",
      "anchor": "paper-2603-24931",
      "title": "COIN: Collaborative Interaction-Aware Multi-Agent Reinforcement Learning for Self-Driving Systems",
      "authors": [
        "Yifeng Zhang",
        "Jieming Chen",
        "Tingguang Zhou",
        "Tanishq Duhan",
        "Jianghong Dong",
        "Yuhong Cao",
        "Guillaume Sartoretti"
      ],
      "abstract": "Multi-Agent Self-Driving (MASD) systems provide an effective solution for coordinating autonomous vehicles to reduce congestion and enhance both safety and operational efficiency in future intelligent transportation systems. Multi-Agent Reinforcement Learning (MARL) has emerged as a promising approach for developing advanced end-to-end MASD systems. However, achieving efficient and safe collaboration in dynamic MASD systems remains a significant challenge in dense scenarios with complex agent interactions. To address this challenge, we propose a novel collaborative(CO-) interaction-aware(-IN) MARL framework, named COIN. Specifically, we develop a new counterfactual individual-global twin delayed deep deterministic policy gradient (CIG-TD3) algorithm, crafted in a \"centralized training, decentralized execution\" (CTDE) manner, which aims to jointly optimize the individual objectives (navigation) and the global objectives (collaboration) of agents. We further introduce a dual-level interaction-aware centralized critic architecture that captures both local pairwise interactions and global system-level dependencies, enabling more accurate global value estimation and improved credit assignment for collaborative policy learning. We conduct extensive simulation experiments in dense urban traffic environments, which demonstrate that COIN consistently outperforms other advanced baseline methods in both safety and efficiency across various system sizes. These results highlight its superiority in complex and dynamic MASD scenarios, as further validated through real-world robot demonstrations. Supplementary videos are available at https://marmotlab.github.io/COIN/",
      "short_abstract": "Multi-Agent Self-Driving (MASD) systems provide an effective solution for coordinating autonomous vehicles to reduce congestion and enhance both safety and operational efficiency in future intelligent transportation systems. Multi-Agent Reinforcement Learning...",
      "published": "2026-03-25",
      "updated": "2026-03-25",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems"
        ]
      },
      "recency": "archive",
      "age_days": 147,
      "arxiv_url": "https://arxiv.org/abs/2603.24931",
      "pdf_url": "https://arxiv.org/pdf/2603.24931.pdf"
    },
    {
      "id": "2603.24138",
      "anchor": "paper-2603-24138",
      "title": "Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models",
      "authors": [
        "Lukas Theiner",
        "Maik Pfefferkorn",
        "Yongpeng Zhao",
        "Sebastian Hirt",
        "Rolf Findeisen"
      ],
      "abstract": "Tuning control policies manually to meet high-level objectives is often time-consuming. Bayesian optimization provides a data-efficient framework for automating this process using numerical evaluations of an objective function. However, many systems, particularly those involving humans, require optimization based on subjective criteria. Preferential Bayesian optimization addresses this by learning from pairwise comparisons instead of quantitative measurements, but relying solely on preference data can be inefficient. We propose a multi-fidelity, multi-modal Bayesian optimization framework that integrates low-fidelity numerical data with high-fidelity human preferences. Our approach employs Gaussian process surrogate models with both hierarchical, autoregressive and non-hierarchical, coregionalization-based structures, enabling efficient learning from mixed-modality data. We illustrate the framework by tuning an autonomous vehicle's trajectory planner, showing that combining numerical and preference data significantly reduces the need for experiments involving the human decision maker while effectively adapting driving style to individual preferences.",
      "short_abstract": "Tuning control policies manually to meet high-level objectives is often time-consuming. Bayesian optimization provides a data-efficient framework for automating this process using numerical evaluations of an objective function. However, many systems,...",
      "published": "2026-03-25",
      "updated": "2026-03-25",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 11,
        "evidence": [
          "title: controller",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 147,
      "arxiv_url": "https://arxiv.org/abs/2603.24138",
      "pdf_url": "https://arxiv.org/pdf/2603.24138.pdf"
    },
    {
      "id": "2603.24971",
      "anchor": "paper-2603-24971",
      "title": "Quantum Inspired Vehicular Network Optimization for Intelligent Decision Making in Smart Cities",
      "authors": [
        "Kamran Ahmad Awan",
        "Sonia Khan",
        "Eman Abdullah Aldakheel",
        "Saif Al-Kuwari",
        "Ahmed Farouk"
      ],
      "abstract": "Connected and automated vehicles require city-scale coordination under strict latency and reliability constraints. However, many existing approaches optimize communication and mobility separately, which can degrade performance during network outages and under compute contention. This paper presents QIVNOM, a quantum-inspired framework that jointly optimizes vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication together with urban traffic control on classical edge--cloud hardware, without requiring a quantum processor. QIVNOM encodes candidate routing--signal plans as probabilistic superpositions and updates them using sphere-projected gradients with annealed sampling to minimize a regularized objective. An entanglement-style regularizer couples networking and mobility decisions, while Tchebycheff multi-objective scalarization with feasibility projection enforces constraints on latency and reliability. The proposed framework is evaluated in METR-LA--calibrated SUMO--OMNeT++/Veins simulations over a $5\\times5$~km urban map with IEEE 802.11p and 5G NR sidelink. Results show that QIVNOM reduces mean end-to-end latency to 57.3~ms, approximately $20\\%$ lower than the best baseline. Under incident conditions, latency decreases from 79~ms to 62~ms ($-21.5\\%$), while under roadside unit (RSU) outages, it decreases from 86~ms to 67~ms ($-22.1\\%$). Packet delivery reaches $96.7\\%$ (an improvement of $+2.3$ percentage points), and reliability remains $96.7\\%$ overall, including $96.8\\%$ under RSU outages versus $94.1\\%$ for the baseline. In corridor-closure scenarios, travel performance also improves, with average travel time reduced to 12.8~min and congestion lowered to $33\\%$, compared with 14.5~min and $37\\%$ for the baseline.",
      "short_abstract": "Connected and automated vehicles require city-scale coordination under strict latency and reliability constraints. However, many existing approaches optimize communication and mobility separately, which can degrade performance during network outages and under...",
      "published": "2026-03-25",
      "updated": "2026-03-25",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 26,
        "margin": 14,
        "evidence": [
          "abstract: V2X",
          "abstract: connected vehicles",
          "abstract: hardware",
          "title: communications",
          "abstract: communications",
          "abstract: systems constraints",
          "abstract: roadside infrastructure"
        ]
      },
      "recency": "archive",
      "age_days": 147,
      "arxiv_url": "https://arxiv.org/abs/2603.24971",
      "pdf_url": "https://arxiv.org/pdf/2603.24971.pdf"
    },
    {
      "id": "2603.23607",
      "anchor": "paper-2603-23607",
      "title": "LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset",
      "authors": [
        "Royden Wagner",
        "Omer Sahin Tas",
        "Jaime Villa",
        "Felix Hauser",
        "Yinzhe Shen",
        "Marlon Steiner",
        "Dominik Strutz",
        "Carlos Fernandez",
        "Christian Kinzig",
        "Guillermo S. Guitierrez-Cabello",
        "Hendrik Königshof",
        "Fabian Immel",
        "Richard Schwarzkopf",
        "Nils Alexander Rack",
        "Kevin Rösch",
        "Kaiwen Wang",
        "Jan-Hendrik Pauls",
        "Martin Lauer",
        "Igor Gilitschenski",
        "Holger Caesar",
        "Christoph Stiller"
      ],
      "abstract": "In real-world domains such as self-driving, generalization to rare scenarios remains a fundamental challenge. To address this, we introduce a new dataset designed for end-to-end driving that focuses on long-tail driving events. We provide multi-view video data, trajectories, high-level instructions, and detailed reasoning traces, facilitating in-context learning and few-shot generalization. The resulting benchmark for multimodal models, such as VLMs and VLAs, goes beyond safety and comfort metrics by evaluating instruction following and semantic coherence between model outputs. The multilingual reasoning traces in English, Spanish, and Chinese are from domain experts with diverse cultural backgrounds. Thus, our dataset is a unique resource for studying how different forms of reasoning affect driving competence. Our dataset is available at: https://hf.co/datasets/kit-mrt/kitscenes-longtail",
      "short_abstract": "In real-world domains such as self-driving, generalization to rare scenarios remains a fundamental challenge. To address this, we introduce a new dataset designed for end-to-end driving that focuses on long-tail driving events. We provide multi-view video...",
      "published": "2026-03-24",
      "updated": "2026-03-24",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 5,
        "evidence": [
          "abstract: end-to-end driving",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 148,
      "arxiv_url": "https://arxiv.org/abs/2603.23607",
      "pdf_url": "https://arxiv.org/pdf/2603.23607.pdf"
    },
    {
      "id": "2603.23310",
      "anchor": "paper-2603-23310",
      "title": "Modeling Edge-to-Cloud Offloading Workloads for Autonomous Vehicles",
      "authors": [
        "Longkun Li",
        "Evangelos Pournaras"
      ],
      "abstract": "Autonomous vehicles generate large volumes of data for applications such as fleet monitoring, model retraining, and high-definition map updates. Existing studies often rely on generic traffic traces, which do not capture the characteristics of autonomous driving workloads. This paper proposes a system-level workload modeling framework for vehicle-to-cloud data. We classify offloaded data into three types: telemetry, event-driven fleet learning, and high-definition map updates, while we model their generation using a parameterized formulation based on empirical data. Using a real-world mobility trace from Munich, we analyze the resulting workloads over time and space. The results show that workload scales with vehicle penetration, exhibits temporal structure and spatial imbalance across access points, and is distinguished from baseline traffic models.",
      "short_abstract": "Autonomous vehicles generate large volumes of data for applications such as fleet monitoring, model retraining, and high-definition map updates. Existing studies often rely on generic traffic traces, which do not capture the characteristics of autonomous...",
      "published": "2026-03-24",
      "updated": "2026-03-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Mapping & Localization: abstract: HD map"
        ]
      },
      "recency": "archive",
      "age_days": 148,
      "arxiv_url": "https://arxiv.org/abs/2603.23310",
      "pdf_url": "https://arxiv.org/pdf/2603.23310.pdf"
    },
    {
      "id": "2603.23162",
      "anchor": "paper-2603-23162",
      "title": "LiZIP: An Auto-Regressive Compression Framework for LiDAR Point Clouds",
      "authors": [
        "Aditya Shibu",
        "Kayvan Karim",
        "Claudio Zito"
      ],
      "abstract": "The massive volume of data generated by LiDAR sensors in autonomous vehicles creates a bottleneck for real-time processing and vehicle-to-everything (V2X) transmission. Existing lossless compression methods often force a trade-off: industry standard algorithms (e.g., LASzip) lack adaptability, while deep learning approaches suffer from prohibitive computational costs. This paper proposes LiZIP, a lightweight, near-lossless zero-drift compression framework based on neural predictive coding. By utilizing a compact Multi-Layer Perceptron (MLP) to predict point coordinates from local context, LiZIP efficiently encodes only the sparse residuals. We evaluate LiZIP on the NuScenes and Argoverse datasets, benchmarking against GZip, LASzip, and Google Draco (configured with 24-bit quantization to serve as a high-precision geometric baseline). Results demonstrate that LiZIP consistently achieves superior compression ratios across varying environments. The proposed system achieves a 7.5%-14.8% reduction in file size compared to the industry-standard LASzip and outperforms Google Draco by 8.8%-11.3% across diverse datasets. Furthermore, the system demonstrates generalization capabilities on the unseen Argoverse dataset without retraining. Against the general purpose GZip algorithm, LiZIP achieves a reduction of 38%-48%. This efficiency offers a distinct advantage for bandwidth constrained V2X applications and large scale cloud archival.",
      "short_abstract": "The massive volume of data generated by LiDAR sensors in autonomous vehicles creates a bottleneck for real-time processing and vehicle-to-everything (V2X) transmission. Existing lossless compression methods often force a trade-off: industry standard...",
      "published": "2026-03-24",
      "updated": "2026-03-24",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 10,
        "evidence": [
          "title: point cloud",
          "title: LiDAR",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 148,
      "arxiv_url": "https://arxiv.org/abs/2603.23162",
      "pdf_url": "https://arxiv.org/pdf/2603.23162.pdf"
    },
    {
      "id": "2603.23037",
      "anchor": "paper-2603-23037",
      "title": "YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception",
      "authors": [
        "Marios Impraimakis",
        "Daniel Vazquez",
        "Feiyu Zhou"
      ],
      "abstract": "The interpretable object detection capabilities of a novel Kolmogorov-Arnold network framework are examined here. The approach refers to a key limitation in computer vision for autonomous vehicles perception, and beyond. These systems offer limited transparency regarding the reliability of their confidence scores in visually degraded or ambiguous scenes. To address this limitation, a Kolmogorov-Arnold network is employed as an interpretable post-hoc surrogate to model the trustworthiness of the You Only Look Once (Yolov10) detections using seven geometric and semantic features. The additive spline-based structure of the Kolmogorov-Arnold network enables direct visualisation of each feature's influence. This produces smooth and transparent functional mappings that reveal when the model's confidence is well supported and when it is unreliable. Experiments on both Common Objects in Context (COCO), and images from the University of Bath campus demonstrate that the framework accurately identifies low-trust predictions under blur, occlusion, or low texture. This provides actionable insights for filtering, review, or downstream risk mitigation. Furthermore, a bootstrapped language-image (BLIP) foundation model generates descriptive captions of each scene. This tool enables a lightweight multimodal interface without affecting the interpretability layer. The resulting system delivers interpretable object detection with trustworthy confidence estimates. It offers a powerful tool for transparent and practical perception component for autonomous and multimodal artificial intelligence applications.",
      "short_abstract": "The interpretable object detection capabilities of a novel Kolmogorov-Arnold network framework are examined here. The approach refers to a key limitation in computer vision for autonomous vehicles perception, and beyond. These systems offer limited...",
      "published": "2026-03-24",
      "updated": "2026-03-24",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 20,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "title: object detection",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 148,
      "arxiv_url": "https://arxiv.org/abs/2603.23037",
      "pdf_url": "https://arxiv.org/pdf/2603.23037.pdf"
    },
    {
      "id": "2603.22035",
      "anchor": "paper-2603-22035",
      "title": "Future-Interactions-Aware Trajectory Prediction via Braid Theory",
      "authors": [
        "Caio Azevedo",
        "Stefano Sabatini",
        "Sascha Hornauer",
        "Fabien Moutarde"
      ],
      "abstract": "To safely operate, an autonomous vehicle must know the future behavior of a potentially high number of interacting agents around it, a task often posed as multi-agent trajectory prediction. Many previous attempts to model social interactions and solve the joint prediction task either add extensive computational requirements or rely on heuristics to label multi-agent behavior types. Braid theory, in contrast, provides a powerful exact descriptor of multi-agent behavior by projecting future trajectories into braids that express how trajectories cross with each other over time; a braid then corresponds to a specific mode of coordination between the multiple agents in the future. In past work, braids have been used lightly to reason about interacting agents and restrict the attention window of predicted agents. We show that leveraging more fully the expressivity of the braid representation and using it to condition the trajectories themselves leads to even further gains in joint prediction performance, with negligible added complexity either in training or at inference time. We do so by proposing a novel auxiliary task, braid prediction, done in parallel with the trajectory prediction task. By classifying edges between agents into their correct crossing types in the braid representation, the braid prediction task is able to imbue the model with improved social awareness, which is reflected in joint predictions that more closely adhere to the actual multi-agent behavior. This simple auxiliary task allowed us to obtain significant improvements in joint metrics on three separate datasets. We show how the braid prediction task infuses the model with future intention awareness, leading to more accurate joint predictions. Code is available at github.com/caiocj1/traj-pred-braid-theory.",
      "short_abstract": "To safely operate, an autonomous vehicle must know the future behavior of a potentially high number of interacting agents around it, a task often posed as multi-agent trajectory prediction. Many previous attempts to model social interactions and solve the...",
      "published": "2026-03-23",
      "updated": "2026-03-23",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 28,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 149,
      "arxiv_url": "https://arxiv.org/abs/2603.22035",
      "pdf_url": "https://arxiv.org/pdf/2603.22035.pdf"
    },
    {
      "id": "2603.21987",
      "anchor": "paper-2603-21987",
      "title": "LRC-WeatherNet: LiDAR, RADAR, and Camera Fusion Network for Real-time Weather-type Classification in Autonomous Driving",
      "authors": [
        "Nour Alhuda Albashir",
        "Lars Pernickel",
        "Danial Hamoud",
        "Idriss Gouigah",
        "Eren Erdal Aksoy"
      ],
      "abstract": "Autonomous vehicles face major perception and navigation challenges in adverse weather such as rain, fog, and snow, which degrade the performance of LiDAR, RADAR, and RGB camera sensors. While each sensor type offers unique strengths, such as RADAR robustness in poor visibility and LiDAR precision in clear conditions, they also suffer distinct limitations when exposed to environmental obstructions. This study proposes LRC-WeatherNet, a novel multi-sensor fusion framework that integrates LiDAR, RADAR, and camera data for real-time classification of weather conditions. By employing both early fusion using a unified Bird's Eye View representation and mid-level gated fusion of modality-specific feature maps, our approach adapts to the varying reliability of each sensor under changing weather. Evaluated on the extensive MSU-4S dataset covering nine weather types, LRC-WeatherNet achieves superior classification performance and computational efficiency, significantly outperforming unimodal baselines in adverse conditions. This work is the first to combine all three modalities for robust, real-time weather classification in autonomous driving. We release our trained models and source code in https://github.com/nouralhudaalbashir/LRC-WeatherNet.",
      "short_abstract": "Autonomous vehicles face major perception and navigation challenges in adverse weather such as rain, fog, and snow, which degrade the performance of LiDAR, RADAR, and RGB camera sensors. While each sensor type offers unique strengths, such as RADAR robustness...",
      "published": "2026-03-23",
      "updated": "2026-03-23",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 41,
        "margin": 23,
        "evidence": [
          "abstract: perception",
          "abstract: sensor fusion",
          "title: LiDAR",
          "abstract: LiDAR",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 149,
      "arxiv_url": "https://arxiv.org/abs/2603.21987",
      "pdf_url": "https://arxiv.org/pdf/2603.21987.pdf"
    },
    {
      "id": "2603.21958",
      "anchor": "paper-2603-21958",
      "title": "Interaction-Aware Predictive Environmental Control Barrier Function for Emergency Lane Change",
      "authors": [
        "Ying Shuai Quan",
        "Paolo Falcone",
        "Jonas Sjöberg"
      ],
      "abstract": "Safety-critical motion planning in mixed traffic remains challenging for autonomous vehicles, especially when it involves interactions between the ego vehicle (EV) and surrounding vehicles (SVs). In dense traffic, the feasibility of a lane change depends strongly on how SVs respond to the EV motion. This paper presents an interaction-aware safety framework that incorporates such interactions into a control barrier function (CBF)-based safety assessment. The proposed method predicts near-future vehicle positions over a finite horizon, thereby capturing reactive SV behavior and embedding it into the CBF-based safety constraint. To address uncertainty in the SV response model, a robust extension is developed by treating the model mismatch as a bounded disturbance and incorporating an online uncertainty estimate into the barrier condition. Compared with classical environmental CBF methods that neglect SV reactions, the proposed approach provides a less conservative and more informative safety representation for interactive traffic scenarios, while improving robustness to uncertainty in the modeled SV behavior.",
      "short_abstract": "Safety-critical motion planning in mixed traffic remains challenging for autonomous vehicles, especially when it involves interactions between the ego vehicle (EV) and surrounding vehicles (SVs). In dense traffic, the feasibility of a lane change depends...",
      "published": "2026-03-23",
      "updated": "2026-03-23",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: title: control",
          "Control & Vehicle Dynamics: abstract: control",
          "Planning & Decision-Making: abstract: motion planning",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 149,
      "arxiv_url": "https://arxiv.org/abs/2603.21958",
      "pdf_url": "https://arxiv.org/pdf/2603.21958.pdf"
    },
    {
      "id": "2603.21810",
      "anchor": "paper-2603-21810",
      "title": "Partial Attention in Deep Reinforcement Learning for Safe Multi-Agent Control",
      "authors": [
        "Turki Bin Mohaya",
        "Peter Seiler"
      ],
      "abstract": "Attention mechanisms excel at learning sequential patterns by discriminating data based on relevance and importance. This provides state-of-the-art performance in advanced generative artificial intelligence models. This paper applies this concept of an attention mechanism for multi-agent safe control. We specifically consider the design of a neural network to control autonomous vehicles in a highway merging scenario. The environment is modeled as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). Within a QMIX framework, we include partial attention for each autonomous vehicle, thus allowing each ego vehicle to focus on the most relevant neighboring vehicles. Moreover, we propose a comprehensive reward signal that considers the global objectives of the environment (e.g., safety and vehicle flow) and the individual interests of each agent. Simulations are conducted in the Simulation of Urban Mobility (SUMO). The results show better performance compared to other driving algorithms in terms of safety, driving speed, and reward.",
      "short_abstract": "Attention mechanisms excel at learning sequential patterns by discriminating data based on relevance and importance. This provides state-of-the-art performance in advanced generative artificial intelligence models. This paper applies this concept of an...",
      "published": "2026-03-23",
      "updated": "2026-03-23",
      "category": "Control & Vehicle Dynamics",
      "primary_category": "Control & Vehicle Dynamics",
      "tags": [
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 8,
        "margin": 4,
        "evidence": [
          "title: control",
          "abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 149,
      "arxiv_url": "https://arxiv.org/abs/2603.21810",
      "pdf_url": "https://arxiv.org/pdf/2603.21810.pdf"
    },
    {
      "id": "2603.21066",
      "anchor": "paper-2603-21066",
      "title": "The Role of Road Features and Vehicle Dynamics in Cost-Effective Autonomous Vehicles Safety Testing: Insights from Instance Space Analysis",
      "authors": [
        "Victor Crespo-Rodriguez",
        "Christian Birchler",
        "Neelofar",
        "Aldeida Aleti",
        "Sebastiano Panichella"
      ],
      "abstract": "Context: Simulation-based testing is a cost-efficient alternative to field testing for Autonomous Vehicles (AVs), but generating safety-critical test cases is challenging due to the vast search space. Prior work has studied static (road features) and dynamic (AV behavior) features of test scenarios separately, but their inter-dependencies are underexplored. Objective: In this paper, we describe an empirical to analyze how static and dynamic featuresof test scenarios, and their inter-dependencies, influence AV test scenario outcomes. Method: This study proposes an integrated approach using Instance Space Analysis (ISA) toevaluate both types of features, identify key influences on AV safety, and predict test outcomeswithout execution. Results: Our study identifies critical features affecting test outcomes (effective/ineffective, depending on whether it leads to a safety-critical condition). Results show that combining static and dynamic features improves prediction accuracy, confirmed by models trained on both feature types outperforming models trained with only one type of feature. Conclusion: The interplay of static and dynamic features enhances fault detection in AV testing. This research underscores the importance of integrating both types of features to create more effective testing frameworks for autonomous systems. Key contributions include: (1) a unified framework for AV safety assessment, (2) identification of influential features using ISA, and (3) efficient test outcome prediction for optimized regression testing.",
      "short_abstract": "Context: Simulation-based testing is a cost-efficient alternative to field testing for Autonomous Vehicles (AVs), but generating safety-critical test cases is challenging due to the vast search space. Prior work has studied static (road features) and dynamic...",
      "published": "2026-03-22",
      "updated": "2026-03-22",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 28,
        "margin": 13,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 150,
      "arxiv_url": "https://arxiv.org/abs/2603.21066",
      "pdf_url": "https://arxiv.org/pdf/2603.21066.pdf"
    },
    {
      "id": "2603.20525",
      "anchor": "paper-2603-20525",
      "title": "High-Speed, All-Terrain Autonomy: Ensuring Safety at the Limits of Mobility",
      "authors": [
        "James R. Baxter",
        "Bogdan I. Epureanu",
        "Paramsothy Jayakumar",
        "Tulga Ersal"
      ],
      "abstract": "A novel local trajectory planner, capable of controlling an autonomous off-road vehicle on rugged terrain at high-speed is presented. Autonomous vehicles are currently unable to safely operate off-road at high-speed, as current approaches either fail to predict and mitigate rollovers induced by rough terrain or are not real-time feasible. To address this challenge, a novel model predictive control (MPC) formulation is developed for local trajectory planning. A new dynamics model for off-road vehicles on rough, non-planar terrain is derived and used for prediction. Extreme mobility, including tire liftoff without rollover, is safely enabled through a new energy-based constraint. The formulation is analytically shown to mitigate rollover types ignored by many state-of-the-art methods, and real-time feasibility is achieved through parallelized GPGPU computation. The planner's ability to provide safe, extreme trajectories is studied through both simulated trials and full-scale physical experiments. The results demonstrate fewer rollovers and more successes compared to a state-of-the-art baseline across several challenging scenarios that push the vehicle to its mobility limits.",
      "short_abstract": "A novel local trajectory planner, capable of controlling an autonomous off-road vehicle on rugged terrain at high-speed is presented. Autonomous vehicles are currently unable to safely operate off-road at high-speed, as current approaches either fail to...",
      "published": "2026-03-20",
      "updated": "2026-03-20",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "title: safety"
        ]
      },
      "recency": "archive",
      "age_days": 152,
      "arxiv_url": "https://arxiv.org/abs/2603.20525",
      "pdf_url": "https://arxiv.org/pdf/2603.20525.pdf"
    },
    {
      "id": "2604.09631",
      "anchor": "paper-2604-09631",
      "title": "Hardware Utilization and Inference Performance of Edge Object Detection Under Fault Injection",
      "authors": [
        "Faezeh Pasandideh",
        "Mehdi Azarafza",
        "Achim Rettberg"
      ],
      "abstract": "As deep learning models are deployed on resource constrained edge platforms in autonomous driving systems, reli able knowledge of hardware behavior under resource degradation becomes an essential requirement. Therefore, we introduce a systematic characterization of CPU load, GPU utilization, RAM consumption, power draw, throughput, and thermal behaviour of TensorRT-optimized YOLOv10s, YOLOv11s and YOLO2026n pipelines running on NVIDIA Jetson Nano under a large-scale fault injection campaign targeting both lane-following and ob ject detection tasks. Faults are synthesized using a decoupled framework that leverages large language models (LLMs) and latent diffusion models (LDMs), based on original data from our JetBot platform data collection. Results show that across both tasks and both models the inference engines keep GPU occupancy stable, temperature rise under control, and power consumption within safe limits, while memory usage settles into a consistent release pattern after the initial warm-up phase. Object detection tends to show somewhat more variability in memory and thermal behavior, yet both tasks point to the same conclusion: the TensorRT pipelines hold up well even when the input data is heavily degraded. These findings offer a hardware-level view of model reliability that sits alongside, rather than against, the broader body of work focused on inference performance at the edge.",
      "short_abstract": "As deep learning models are deployed on resource constrained edge platforms in autonomous driving systems, reli able knowledge of hardware behavior under resource degradation becomes an essential requirement. Therefore, we introduce a systematic...",
      "published": "2026-03-19",
      "updated": "2026-03-19",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 18,
        "margin": 6,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: occupancy"
        ]
      },
      "recency": "archive",
      "age_days": 153,
      "arxiv_url": "https://arxiv.org/abs/2604.09631",
      "pdf_url": "https://arxiv.org/pdf/2604.09631.pdf"
    },
    {
      "id": "2603.19556",
      "anchor": "paper-2603-19556",
      "title": "Planning Autonomous Vehicle Maneuvering in Work Zones Through Game-Theoretic Trajectory Generation",
      "authors": [
        "Mayar Nour",
        "Atrisha Sarkar",
        "Mohamed H. Zaki"
      ],
      "abstract": "Work zone navigation remains one of the most challenging manoeuvres for autonomous vehicles (AVs), where constrained geometries and unpredictable traffic patterns create a high-risk environment. Despite extensive research on AV trajectory planning, few studies address the decision-making required to navigate work zones safely. This paper proposes a novel game-theoretic framework for trajectory generation and control to enhance the safety of lane changes in a work zone environment. By modelling the lane change manoeuvre as a non-cooperative game between vehicles, we use a game-theoretic planner to generate trajectories that balance safety, progress, and traffic stability. The simulation results show that the proposed game-theoretic model reduces the frequency of conflicts by 35 percent and decreases the probability of high risk safety events compared to traditional vehicle behaviour planning models in safety-critical highway work-zone scenarios.",
      "short_abstract": "Work zone navigation remains one of the most challenging manoeuvres for autonomous vehicles (AVs), where constrained geometries and unpredictable traffic patterns create a high-risk environment. Despite extensive research on AV trajectory planning, few...",
      "published": "2026-03-19",
      "updated": "2026-03-19",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 35,
        "margin": 31,
        "evidence": [
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: decision-making",
          "title: trajectory generation",
          "abstract: trajectory generation",
          "abstract: navigation"
        ]
      },
      "recency": "archive",
      "age_days": 153,
      "arxiv_url": "https://arxiv.org/abs/2603.19556",
      "pdf_url": "https://arxiv.org/pdf/2603.19556.pdf"
    },
    {
      "id": "2603.19533",
      "anchor": "paper-2603-19533",
      "title": "Pedestrian Crossing Intent Prediction via Psychological Features and Transformer Fusion",
      "authors": [
        "Sima Ashayer",
        "Hoang H. Nguyen",
        "Yu Liang",
        "Mina Sartipi"
      ],
      "abstract": "Pedestrian intention prediction needs to be accurate for autonomous vehicles to navigate safely in urban environments. We present a lightweight, socially informed architecture for pedestrian intention prediction. It fuses four behavioral streams (attention, position, situation, and interaction) using highway encoders, a compact 4-token Transformer, and global self-attention pooling. To quantify uncertainty, we incorporate two complementary heads: a variational bottleneck whose KL divergence captures epistemic uncertainty, and a Mahalanobis distance detector that identifies distributional shift. Together, these components yield calibrated probabilities and actionable risk scores without compromising efficiency. On the PSI 1.0 benchmark, our model outperforms recent vision language models by achieving 0.9 F1, 0.94 AUC-ROC, and 0.78 MCC by using only structured, interpretable features. On the more diverse PSI 2.0 dataset, where, to the best of our knowledge, no prior results exist, we establish a strong initial baseline of 0.78 F1 and 0.79 AUC-ROC. Selective prediction based on Mahalanobis scores increases test accuracy by up to 0.4 percentage points at 80% coverage. Qualitative attention heatmaps further show how the model shifts its cross-stream focus under ambiguity. The proposed approach is modality-agnostic, easy to integrate with vision language pipelines, and suitable for risk-aware intent prediction on resource-constrained platforms.",
      "short_abstract": "Pedestrian intention prediction needs to be accurate for autonomous vehicles to navigate safely in urban environments. We present a lightweight, socially informed architecture for pedestrian intention prediction. It fuses four behavioral streams (attention,...",
      "published": "2026-03-19",
      "updated": "2026-03-19",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 18,
        "evidence": [
          "title: behavior prediction",
          "abstract: behavior prediction",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 153,
      "arxiv_url": "https://arxiv.org/abs/2603.19533",
      "pdf_url": "https://arxiv.org/pdf/2603.19533.pdf"
    },
    {
      "id": "2603.19188",
      "anchor": "paper-2603-19188",
      "title": "Markov Potential Game and Multi-Agent Reinforcement Learning for Autonomous Driving",
      "authors": [
        "Huiwen Yan",
        "Mushuang Liu"
      ],
      "abstract": "Autonomous driving (AD) requires safe and reliable decision-making among interacting agents, e.g., vehicles, bicycles, and pedestrians. Multi-agent reinforcement learning (MARL) modeled by Markov games (MGs) provides a suitable framework to characterize such agents' interactions during decision-making. Nash equilibria (NEs) are often the desired solution in an MG. However, it is typically challenging to compute an NE in general-sum games, unless the game is a Markov potential game (MPG), which ensures the NE attainability under a few learning algorithms such as gradient play. However, it has been an open question how to construct an MPG and whether these construction rules are suitable for AD applications. In this paper, we provide sufficient conditions under which an MG is an MPG and show that these conditions can accommodate general driving objectives for autonomous vehicles (AVs) using highway forced merge scenarios as illustrative examples. A parameter-sharing neural network (NN) structure is designed to enable decentralized policy execution. The trained driving policy from MPGs is evaluated in both simulated and naturalistic traffic datasets. Comparative studies with single-agent RL and with human drivers whose behaviors are recorded in the traffic datasets are reported, respectively.",
      "short_abstract": "Autonomous driving (AD) requires safe and reliable decision-making among interacting agents, e.g., vehicles, bicycles, and pedestrians. Multi-agent reinforcement learning (MARL) modeled by Markov games (MGs) provides a suitable framework to characterize such...",
      "published": "2026-03-19",
      "updated": "2026-03-19",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 3,
        "evidence": [
          "abstract: law or policy",
          "abstract: human driver"
        ]
      },
      "recency": "archive",
      "age_days": 153,
      "arxiv_url": "https://arxiv.org/abs/2603.19188",
      "pdf_url": "https://arxiv.org/pdf/2603.19188.pdf"
    },
    {
      "id": "2603.19119",
      "anchor": "paper-2603-19119",
      "title": "Exact-Time Safety Recovery using Time-Varying Control Barrier Functions with Optimal Barrier Tracking",
      "authors": [
        "Yingqing Chen",
        "Christos G. Cassandras",
        "Wei Xiao",
        "Anni Li"
      ],
      "abstract": "This paper is motivated by controllers developed for autonomous vehicles which occasionally result into conditions where safety is no longer guaranteed. We develop an exact-time safety recovery framework for any control-affine nonlinear system when its state is outside a safe region using time-varying Control Barrier Functions (CBFs) with optimal barrier tracking. Unlike conventional formulations that provide only conservative upper bounds on recovery time convergence, the proposed approach guarantees recovery to the safe set at a prescribed time. The key mechanism is an active barrier tracking condition that forces the barrier function to follow exactly a designer-specified recovery trajectory. This transforms safety recovery into a trajectory design problem. The recovery trajectory is parameterized and optimized to achieve optimal performance while preserving feasibility under input constraints, avoiding the aggressive corrective actions typically induced by conventional finite-time formulations. The safety recovery framework is applied to the roundabout traffic coordination problem for Connected and Automated Vehicles (CAVs), where any initially violated safe merging constraint is replaced by an exact-time recovery barrier constraint to ensure safety guarantee restoration before CAV conflict points are reached. Simulation results demonstrate improved feasibility and performance.",
      "short_abstract": "This paper is motivated by controllers developed for autonomous vehicles which occasionally result into conditions where safety is no longer guaranteed. We develop an exact-time safety recovery framework for any control-affine nonlinear system when its state...",
      "published": "2026-03-19",
      "updated": "2026-03-19",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 5,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 153,
      "arxiv_url": "https://arxiv.org/abs/2603.19119",
      "pdf_url": "https://arxiv.org/pdf/2603.19119.pdf"
    },
    {
      "id": "2603.18595",
      "anchor": "paper-2603-18595",
      "title": "RUBICONe: Wireless RAFT-Unified Behaviors for Intervehicular Cooperative Operations and Negotiations",
      "authors": [
        "Zhenghua Hu",
        "Tairan Dan",
        "Zeyu Tao",
        "Jiacheng Qian",
        "Amedeo Morat",
        "Lorenzo Romano",
        "Alessandro Massafra",
        "Hao Xu"
      ],
      "abstract": "Just as Caesar declared \"alea iacta est\" (the die is cast) upon crossing the Rubicone river, lane change decisions in autonomous vehicles also represent critical points of no return. RUBICONe addresses this challenge by recognizing that lane change decision-making relying solely on a single vehicle's perception would be as precarious as crossing an unknown river alone. By implementing a distributed consensus framework that extends the RAFT algorithm with wireless connectivity, RUBICONe enables multiple vehicles to collectively process and aggregate their perceptions. Using multiple software-defined radio (SDR) devices as the experimental platform, this study demonstrates how consensus-based decision-making significantly reduces the impact of environmental interference and mitigates the risk of misjudgments by individual vehicles. Just as crossing the Rubicone marked a point of irrevocable action backed by collective intelligence, RUBICONe ensures that lane change decisions are made with comprehensive situational awareness and distributed consensus, showcasing the reliability gain of consensus in wireless communications.",
      "short_abstract": "Just as Caesar declared \"alea iacta est\" (the die is cast) upon crossing the Rubicone river, lane change decisions in autonomous vehicles also represent critical points of no return. RUBICONe addresses this challenge by recognizing that lane change...",
      "published": "2026-03-19",
      "updated": "2026-03-19",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 11,
        "margin": 7,
        "evidence": [
          "title: cooperative systems",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 153,
      "arxiv_url": "https://arxiv.org/abs/2603.18595",
      "pdf_url": "https://arxiv.org/pdf/2603.18595.pdf"
    },
    {
      "id": "2603.20285",
      "anchor": "paper-2603-20285",
      "title": "AgentComm-Bench: Stress-Testing Cooperative Embodied AI Under Latency, Packet Loss, and Bandwidth Collapse",
      "authors": [
        "Aayam Bansal",
        "Ishaan Gangwani"
      ],
      "abstract": "Cooperative multi-agent methods for embodied AI are almost universally evaluated under idealized communication: zero latency, no packet loss, and unlimited bandwidth. Real-world deployment on robots with wireless links, autonomous vehicles on congested networks, or drone swarms in contested spectrum offers no such guarantees. We introduce AgentComm-Bench, a benchmark suite and evaluation protocol that systematically stress-tests cooperative embodied AI under six communication impairment dimensions: latency, packet loss, bandwidth collapse, asynchronous updates, stale memory, and conflicting sensor evidence. AgentComm-Bench spans three task families: cooperative perception, multi-agent waypoint navigation, and cooperative zone search, and evaluates five communication strategies, including a lightweight method we propose based on redundant message coding with staleness-aware fusion. Our experiments reveal that communication-dependent tasks degrade catastrophically: stale memory and bandwidth collapse cause over 96% performance drops in navigation, while content corruption (stale or conflicting data) reduces perception F1 by over 85%. Vulnerability depends on the interaction between impairment type and task design; perception fusion is robust to packet loss but amplifies corrupted data. Redundant message coding more than doubles navigation performance under 80% packet loss. We release AgentComm-Bench as a practical evaluation protocol and recommend that cooperative embodied AI work report performance under multiple impairment conditions.",
      "short_abstract": "Cooperative multi-agent methods for embodied AI are almost universally evaluated under idealized communication: zero latency, no packet loss, and unlimited bandwidth. Real-world deployment on robots with wireless links, autonomous vehicles on congested...",
      "published": "2026-03-18",
      "updated": "2026-03-18",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Benchmark",
        "Cooperative / V2X",
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 14,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: deployment or real-time",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 154,
      "arxiv_url": "https://arxiv.org/abs/2603.20285",
      "pdf_url": "https://arxiv.org/pdf/2603.20285.pdf"
    },
    {
      "id": "2603.18201",
      "anchor": "paper-2603-18201",
      "title": "A Computationally Efficient Learning of Artificial Intelligence System Reliability Considering Error Propagation",
      "authors": [
        "Fenglian Pan",
        "Yinwei Zhang",
        "Yili Hong",
        "Larry Head",
        "Jian Liu"
      ],
      "abstract": "Artificial Intelligence (AI) systems are increasingly prominent in emerging smart cities, yet their reliability remains a critical concern. These systems typically operate through a sequence of interconnected functional stages, where upstream errors may propagate to downstream stages, ultimately affecting overall system reliability. Quantifying such error propagation is essential for accurate modeling of AI system reliability. However, this task is challenging due to: i) data availability: real-world AI system reliability data are often scarce and constrained by privacy concerns; ii) model validity: recurring error events across sequential stages are interdependent, violating the independence assumptions of statistical inference; and iii) computational complexity: AI systems process large volumes of high-speed data, resulting in frequent and complex recurrent error events that are difficult to track and analyze. To address these challenges, this paper leverages a physics-based autonomous vehicle simulation platform with a justifiable error injector to generate high-quality data for AI system reliability analysis. Building on this data, a new reliability modeling framework is developed to explicitly characterize error propagation across stages. Model parameters are estimated using a computationally efficient, theoretically guaranteed composite likelihood expectation - maximization algorithm. Its application to the reliability modeling for autonomous vehicle perception systems demonstrates its predictive accuracy and computational efficiency.",
      "short_abstract": "Artificial Intelligence (AI) systems are increasingly prominent in emerging smart cities, yet their reliability remains a critical concern. These systems typically operate through a sequence of interconnected functional stages, where upstream errors may...",
      "published": "2026-03-18",
      "updated": "2026-03-18",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 3,
        "evidence": [
          "Perception & Sensor Fusion: abstract: perception"
        ]
      },
      "recency": "archive",
      "age_days": 154,
      "arxiv_url": "https://arxiv.org/abs/2603.18201",
      "pdf_url": "https://arxiv.org/pdf/2603.18201.pdf"
    },
    {
      "id": "2603.17751",
      "anchor": "paper-2603-17751",
      "title": "Multisource Human-in-the-Loop Digital Twin Testbed for Connected and Autonomous Vehicles in Mixed Traffic Flow",
      "authors": [
        "Jianghong Dong",
        "Chunying Yang",
        "Mengchi Cai",
        "Chaoyi Chen",
        "Qing Xu",
        "Jianqiang Wang",
        "Jiawei Wang",
        "Keqiang Li"
      ],
      "abstract": "In the emerging mixed traffic environments, Connected and Autonomous Vehicles (CAVs) have to interact with surrounding human-driven vehicles (HDVs). This paper introduces MSH-MCCT (Multi-Source Human-in-the-Loop Mixed Cloud Control Testbed), a novel CAV testbed that captures complex interactions between various CAVs and HDVs. Utilizing the Mixed Digital Twin concept, which combines Mixed Reality with Digital Twin, MSH-MCCT integrates physical, virtual, and mixed platforms, along with multi-source control inputs. Bridged by the mixed platform, MSH-MCCT allows human drivers and CAV algorithms to operate both physical and virtual vehicles within multiple fields of view. Particularly, this testbed facilitates the coexistence and real-time interaction of physical and virtual CAVs \\& HDVs, significantly enhancing the experimental flexibility and scalability. Experiments on vehicle platooning in mixed traffic showcase the potential of MSH-MCCT to conduct CAV testing with multi-source real human drivers in the loop through driving simulators of diverse fidelity. The videos for the experiments are available at our project website: https://dongjh20.github.io/MSH-MCCT.",
      "short_abstract": "In the emerging mixed traffic environments, Connected and Autonomous Vehicles (CAVs) have to interact with surrounding human-driven vehicles (HDVs). This paper introduces MSH-MCCT (Multi-Source Human-in-the-Loop Mixed Cloud Control Testbed), a novel CAV...",
      "published": "2026-03-18",
      "updated": "2026-03-18",
      "category": "Human Factors & Policy",
      "primary_category": "Human Factors & Policy",
      "tags": [
        "Simulation",
        "Human Interaction"
      ],
      "classification": {
        "confidence": "medium",
        "score": 23,
        "margin": 4,
        "evidence": [
          "title: human-machine interaction",
          "abstract: human-machine interaction",
          "abstract: human driver"
        ]
      },
      "recency": "archive",
      "age_days": 154,
      "arxiv_url": "https://arxiv.org/abs/2603.17751",
      "pdf_url": "https://arxiv.org/pdf/2603.17751.pdf"
    },
    {
      "id": "2603.16742",
      "anchor": "paper-2603-16742",
      "title": "When the City Teaches the Car: Label-Free 3D Perception from Infrastructure",
      "authors": [
        "Zhen Xu",
        "Jinsu Yoo",
        "Cristian Bautista",
        "Zanming Huang",
        "Tai-Yu Pan",
        "Zhenzhen Liu",
        "Katie Z Luo",
        "Mark Campbell",
        "Bharath Hariharan",
        "Wei-Lun Chao"
      ],
      "abstract": "Building robust 3D perception for self-driving still relies heavily on large-scale data collection and manual annotation, yet this paradigm becomes impractical as deployment expands across diverse cities and regions. Meanwhile, modern cities are increasingly instrumented with roadside units (RSUs), static sensors deployed along roads and at intersections to monitor traffic. This raises a natural question: can the city itself help train the vehicle? We propose infrastructure-taught, label-free 3D perception, a paradigm in which RSUs act as stationary, unsupervised teachers for ego vehicles. Leveraging their fixed viewpoints and repeated observations, RSUs learn local 3D detectors from unlabeled data and broadcast predictions to passing vehicles, which are aggregated as pseudo-label supervision for training a standalone ego detector. The resulting model requires no infrastructure or communication at test time. We instantiate this idea as a fully label-free three-stage pipeline and conduct a concept-and-feasibility study in a CARLA-based multi-agent environment. With CenterPoint, our pipeline achieves 82.3% AP for detecting vehicles, compared to a fully supervised ego upper bound of 94.4%. We further systematically analyze each stage, evaluate its scalability, and demonstrate complementarity with existing ego-centric label-free methods. Together, these results suggest that city infrastructure itself can potentially provide a scalable supervisory signal for autonomous vehicles, positioning infrastructure-taught learning as a promising orthogonal paradigm for reducing annotation cost in 3D perception.",
      "short_abstract": "Building robust 3D perception for self-driving still relies heavily on large-scale data collection and manual annotation, yet this paradigm becomes impractical as deployment expands across diverse cities and regions. Meanwhile, modern cities are increasingly...",
      "published": "2026-03-17",
      "updated": "2026-03-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "title: perception",
          "abstract: perception"
        ]
      },
      "recency": "archive",
      "age_days": 155,
      "arxiv_url": "https://arxiv.org/abs/2603.16742",
      "pdf_url": "https://arxiv.org/pdf/2603.16742.pdf"
    },
    {
      "id": "2603.17236",
      "anchor": "paper-2603-17236",
      "title": "Neural Radiance Maps for Extraterrestrial Navigation and Path Planning",
      "authors": [
        "Adam Dai",
        "Shubh Gupta",
        "Grace Gao"
      ],
      "abstract": "Autonomous vehicles such as the Mars rovers currently lead the vanguard of surface exploration on extraterrestrial planets and moons. In order to accelerate the pace of exploration and science objectives, it is critical to plan safe and efficient paths for these vehicles. However, current rover autonomy is limited by a lack of global maps which can be easily constructed and stored for onboard re-planning. Recently, Neural Radiance Fields (NeRFs) have been introduced as a detailed 3D scene representation which can be trained from sparse 2D images and efficiently stored. We propose to use NeRFs to construct maps for online use in autonomous navigation, and present a planning framework which leverages the NeRF map to integrate local and global information. Our approach interpolates local cost observations across global regions using kernel ridge regression over terrain features extracted from the NeRF map, allowing the rover to re-route itself around untraversable areas discovered during online operation. We validate our approach in high-fidelity simulation and demonstrate lower cost and higher percentage success rate path planning compared to various baselines.",
      "short_abstract": "Autonomous vehicles such as the Mars rovers currently lead the vanguard of surface exploration on extraterrestrial planets and moons. In order to accelerate the pace of exploration and science objectives, it is critical to plan safe and efficient paths for...",
      "published": "2026-03-17",
      "updated": "2026-03-17",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 40,
        "margin": 40,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "title: navigation",
          "abstract: navigation"
        ]
      },
      "recency": "archive",
      "age_days": 155,
      "arxiv_url": "https://arxiv.org/abs/2603.17236",
      "pdf_url": "https://arxiv.org/pdf/2603.17236.pdf"
    },
    {
      "id": "2603.17108",
      "anchor": "paper-2603-17108",
      "title": "LLM-Powered Flood Depth Estimation from Social Media Imagery: A Vision-Language Model Framework with Mechanistic Interpretability for Transportation Resilience",
      "authors": [
        "Nafis Fuad",
        "Xiaodong Qian"
      ],
      "abstract": "Urban flooding poses an escalating threat to transportation network continuity, yet no operational system currently provides real-time, street-level flood depth information at the centimeter resolution required for dynamic routing, electric vehicle (EV) safety, and autonomous vehicle (AV) operations. This study presents FloodLlama, a fine-tuned open-source vision-language model (VLM) for continuous flood depth estimation from single street-level images, supported by a multimodal sensing pipeline using TikTok data. A synthetic dataset of approximately 190000 images was generated, covering seven vehicle types, four weather conditions, and 41 depth levels (0-40 cm at 1 cm resolution). Progressive curriculum training enabled coarse-to-fine learning, while LLaMA 3.2-11B Vision was fine-tuned using QLoRA. Evaluation across 34797 trials reveals a depth-dependent prompt effect: simple prompts perform better for shallow flooding, whereas chain-of-thought (CoT) reasoning improves performance at greater depths. FloodLlama achieves a mean absolute error (MAE) below 0.97 cm and Acc@5cm above 93.7% for deep flooding, exceeding 96.8% for shallow depths. A five-phase mechanistic interpretability framework identifies layer L23 as the critical depth-encoding transition and enables selective fine-tuning that reduces trainable parameters by 76-80% while maintaining accuracy. The Tier 3 configuration achieves 98.62% accuracy on real-world data and shows strong robustness under visual occlusion. A TikTok-based data pipeline, validated on 676 annotated flood frames from Detroit, demonstrates the feasibility of real-time, crowd-sourced flood sensing. The proposed framework provides a scalable, infrastructure-free solution with direct implications for EV safety, AV deployment, and resilient transportation management.",
      "short_abstract": "Urban flooding poses an escalating threat to transportation network continuity, yet no operational system currently provides real-time, street-level flood depth information at the centimeter resolution required for dynamic routing, electric vehicle (EV)...",
      "published": "2026-03-17",
      "updated": "2026-03-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 2,
        "evidence": [
          "title: depth estimation",
          "abstract: depth estimation"
        ]
      },
      "recency": "archive",
      "age_days": 155,
      "arxiv_url": "https://arxiv.org/abs/2603.17108",
      "pdf_url": "https://arxiv.org/pdf/2603.17108.pdf"
    },
    {
      "id": "2603.17056",
      "anchor": "paper-2603-17056",
      "title": "DesertFormer: Transformer-Based Semantic Segmentation for Off-Road Desert Terrain Classification in Autonomous Navigation Systems",
      "authors": [
        "Yasaswini Chebolu"
      ],
      "abstract": "Reliable terrain perception is a fundamental requirement for autonomous navigation in unstructured, off-road environments. Desert landscapes present unique challenges due to low chromatic contrast between terrain categories, extreme lighting variability, and sparse vegetation that defy the assumptions of standard road-scene segmentation models. We present DesertFormer, a semantic segmentation pipeline for off-road desert terrain analysis based on SegFormer B2 with a hierarchical Mix Transformer (MiT-B2) backbone. The system classifies terrain into ten ecologically meaningful categories -- Trees, Lush Bushes, Dry Grass, Dry Bushes, Ground Clutter, Flowers, Logs, Rocks, Landscape, and Sky -- enabling safety-aware path planning for ground robots and autonomous vehicles. Trained on a purpose-built dataset of 4,176 annotated off-road images at 512x512 resolution, DesertFormer achieves a mean Intersection-over-Union (mIoU) of 64.4% and pixel accuracy of 86.1%, representing a +24.2% absolute improvement over a DeepLabV3 MobileNetV2 baseline (41.0% mIoU). We further contribute a systematic failure analysis identifying the primary confusion patterns -- Ground Clutter to Landscape and Dry Grass to Landscape -- and propose class-weighted training and copy-paste augmentation for rare terrain categories. Code, checkpoints, and an interactive inference dashboard are released at https://github.com/Yasaswini-ch/Vision-based-Desert-Terrain-Segmentation-using-SegFormer.",
      "short_abstract": "Reliable terrain perception is a fundamental requirement for autonomous navigation in unstructured, off-road environments. Desert landscapes present unique challenges due to low chromatic contrast between terrain categories, extreme lighting variability, and...",
      "published": "2026-03-17",
      "updated": "2026-03-17",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 19,
        "margin": 3,
        "evidence": [
          "abstract: perception",
          "title: scene segmentation",
          "abstract: scene segmentation"
        ]
      },
      "recency": "archive",
      "age_days": 155,
      "arxiv_url": "https://arxiv.org/abs/2603.17056",
      "pdf_url": "https://arxiv.org/pdf/2603.17056.pdf"
    },
    {
      "id": "2603.16688",
      "anchor": "paper-2603-16688",
      "title": "Assessment of Latent Pedestrian-Vehicle Interaction Risk Profiles at Midblock Crossing in VR",
      "authors": [
        "Rulla Al-Haideri",
        "Bilal Farooq",
        "Elisabetta Cherchi"
      ],
      "abstract": "Pedestrian safety at midblock crossings is a critical concern in mixed traffic environments where autonomous vehicles (AVs) and human-driven vehicles (HDVs) share the road. Pedestrians often infer intent from vehicle motion in AV encounters, making them vulnerable to small shifts in conflict margins. This study investigates whether virtual reality (VR) crossing sessions separate into distinct interaction risk profiles and whether AV-only sessions shift profile prevalence compared to HDV-only sessions. Using large-scale immersive VR experiments from Toronto, Canada, and Newcastle, England, we compute surrogate safety measures (SSMs) and apply latent profile analysis (LPA) to identify distinct pedestrian crossing stances, ranging from risk-accepting to highly cautious. Key findings show that Newcastle exhibits a higher prevalence of high-urgency risk profiles in AV-only sessions, indicating that AVs contribute to higher-risk encounters. In contrast, Toronto shows no significant difference between AV-only and HDV-only sessions, suggesting that contextual factors influence the impact of AVs on pedestrian safety.",
      "short_abstract": "Pedestrian safety at midblock crossings is a critical concern in mixed traffic environments where autonomous vehicles (AVs) and human-driven vehicles (HDVs) share the road. Pedestrians often infer intent from vehicle motion in AV encounters, making them...",
      "published": "2026-03-17",
      "updated": "2026-03-17",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 4,
        "evidence": [
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 155,
      "arxiv_url": "https://arxiv.org/abs/2603.16688",
      "pdf_url": "https://arxiv.org/pdf/2603.16688.pdf"
    },
    {
      "id": "2604.16345",
      "anchor": "paper-2604-16345",
      "title": "Bridging the Experimental Last Mile: Digitizing Laboratory Know-How for Safe AI-Assisted Support",
      "authors": [
        "Akira Miura",
        "Yuki Sasahara",
        "Momoka Demura",
        "Yuji Masubuchi",
        "Tetsuya Asai",
        "Chikahiko Mitsui"
      ],
      "abstract": "While advances in materials informatics have accelerated the development of Self-Driving Laboratories (SDLs), human-led experiments remain standard in many educational and exploratory research laboratories. In specific lab settings, formal documentation alone is often insufficient for safe and reliable operation. We refer to the gap between formal documentation and reliable execution in such settings as the experimental last mile; this gap mainly involves site-specific operational know-how, including local rules, routine checks, procedural details, and safety-conscious actions that are can be verbalizable but are often under-documented in standard manuals. In this proof-of-concept study, we developed a human-in-the-loop AI assistant that combines first-person experimental video, multimodal AI, and retrieval-augmented generation (RAG). Using powder X-ray diffraction experiments and student-recorded video data as inputs, the system extracts site-specific laboratory knowledge from recorded procedures, including physical techniques and audible confirmation that conventional manuals could omit. It then provides grounded responses based on the resulting manual. To reduce the risk of unsupported outputs, the system employs a two-layer safety design: source restriction through RAG and strict system-prompt constraints. Instructor-based evaluation showed alignment with expected guidance for questions covered by the manual. For out-of-scope queries, the system appropriately refused to answer, indicating a reduced risk of hallucination. Expert evaluation further indicated that the generated advisory reports were useful and safe (utility: 3.25/4.00; safety: 4.00/4.00). These results suggest the feasibility of a framework for bridging the experimental last mile in which AI supports laboratory practice under explicit human supervision rather than replacing human judgment.",
      "short_abstract": "While advances in materials informatics have accelerated the development of Self-Driving Laboratories (SDLs), human-led experiments remain standard in many educational and exploratory research laboratories. In specific lab settings, formal documentation alone...",
      "published": "2026-03-16",
      "updated": "2026-03-16",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: human-machine interaction",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 156,
      "arxiv_url": "https://arxiv.org/abs/2604.16345",
      "pdf_url": "https://arxiv.org/pdf/2604.16345.pdf"
    },
    {
      "id": "2603.16050",
      "anchor": "paper-2603-16050",
      "title": "The Era of End-to-End Autonomy: Transitioning from Rule-Based Driving to Large Driving Models",
      "authors": [
        "Eduardo Nebot",
        "Julie Stephany Berrio Perez"
      ],
      "abstract": "Autonomous driving is undergoing a shift from modular rule based pipelines toward end to end (E2E) learning systems. This paper examines this transition by tracing the evolution from classical sense perceive plan control architectures to large driving models (LDMs) capable of mapping raw sensor input directly to driving actions. We analyze recent developments including Tesla's Full Self Driving (FSD) V12 V14, Rivian's Unified Intelligence platform, NVIDIA Cosmos, and emerging commercial robotaxi deployments, focusing on architectural design, deployment strategies, safety considerations and industry implications. A key emerging product category is supervised E2E driving, often referred to as FSD (Supervised) or L2 plus plus, which several manufacturers plan to deploy from 2026 onwards. These systems can perform most of the Dynamic Driving Task (DDT) in complex environments while requiring human supervision, shifting the driver's role to safety oversight. Early operational evidence suggests E2E learning handles the long tail distribution of real world driving scenarios and is becoming a dominant commercial strategy. We also discuss how similar architectural advances may extend beyond autonomous vehicles (AV) to other embodied AI systems, including humanoid robotics.",
      "short_abstract": "Autonomous driving is undergoing a shift from modular rule based pipelines toward end to end (E2E) learning systems. This paper examines this transition by tracing the evolution from classical sense perceive plan control architectures to large driving models...",
      "published": "2026-03-16",
      "updated": "2026-03-16",
      "category": "End-to-End & VLA",
      "primary_category": "End-to-End & VLA",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 28,
        "evidence": [
          "title: driving foundation model",
          "abstract: driving foundation model",
          "title: end-to-end",
          "abstract: end-to-end"
        ]
      },
      "recency": "archive",
      "age_days": 156,
      "arxiv_url": "https://arxiv.org/abs/2603.16050",
      "pdf_url": "https://arxiv.org/pdf/2603.16050.pdf"
    },
    {
      "id": "2603.16083",
      "anchor": "paper-2603-16083",
      "title": "Structured prototype regularization for synthetic-to-real driving scene parsing",
      "authors": [
        "Jiahe Fan",
        "Xiao Ma",
        "Sergey Vityazev",
        "George Giakos",
        "Shaolong Shu",
        "Rui Fan"
      ],
      "abstract": "Driving scene parsing is critical for autonomous vehicles to operate reliably in complex real-world traffic environments. To reduce the reliance on costly pixel-level annotations, synthetic datasets with automatically generated labels have become a popular alternative. However, models trained on synthetic data often perform poorly when applied to real-world scenes due to the synthetic-to-real domain gap. Despite the success of unsupervised domain adaptation in narrowing this gap, most existing methods mainly focus on global feature alignment while overlooking the semantic structure of the feature space. As a result, semantic relations among classes are insufficiently modeled, limiting the model's ability to generalize. To address these challenges, this study introduces a novel unsupervised domain adaptation framework that explicitly regularizes semantic feature structures to significantly enhance driving scene parsing performance in real-world scenarios. Specifically, the proposed method enforces inter-class separation and intra-class compactness by leveraging class-specific prototypes, thereby enhancing the discriminability and structural coherence of feature clusters. An entropy-based noise filtering strategy improves the reliability of pseudo labels, while a pixel-level attention mechanism further refines feature alignment. Extensive experiments on representative benchmarks demonstrate that the proposed method consistently outperforms recent state-of-the-art methods. These results underscore the importance of preserving semantic structure for robust synthetic-to-real adaptation in driving scene parsing tasks.",
      "short_abstract": "Driving scene parsing is critical for autonomous vehicles to operate reliably in complex real-world traffic environments. To reduce the reliance on costly pixel-level annotations, synthetic datasets with automatically generated labels have become a popular...",
      "published": "2026-03-16",
      "updated": "2026-03-16",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Synthetic Data"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 156,
      "arxiv_url": "https://arxiv.org/abs/2603.16083",
      "pdf_url": "https://arxiv.org/pdf/2603.16083.pdf"
    },
    {
      "id": "2603.14841",
      "anchor": "paper-2603-14841",
      "title": "Real-Time Driver Safety Scoring Through Inverse Crash Probability Modeling",
      "authors": [
        "Joyjit Roy",
        "Samaresh Kumar Singh",
        "Sushanta Das"
      ],
      "abstract": "Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offer limited actionable insights for real-time driver feedback. These approaches often lack continuous risk quantification, interpretability, and explicit consideration of vulnerable road users (VRUs), such as pedestrians and cyclists. This research introduces SafeDriver-IQ, a framework that transforms binary crash classifiers into continuous 0-100 safety scores by combining national crash statistics with naturalistic driving data from autonomous vehicles. The framework fuses National Highway Traffic Safety Administration (NHTSA) crash records with Waymo Open Motion Dataset scenarios, engineers domain-informed features, and incorporates a calibration layer grounded in transportation safety literature. Evaluation across 15 complementary analyses indicates that the framework reliably differentiates high-risk from low-risk driving conditions with strong discriminative performance. Findings further reveal that 87% of crashes involve multiple co-occurring risk factors, with non-linear compounding effects that increase the risk to 4.5x baseline. SafeDriver-IQ delivers proactive, explainable safety intelligence relevant to advanced driver-assistance systems (ADAS), fleet management, and urban infrastructure planning. Beyond the specific application, the inverse modeling paradigm is domain-agnostic. Any binary risk classifier can be converted into a continuous, explainable safety-scoring system using the same pipeline without retraining. This framework shifts the focus from reactive crash counting to real-time risk prevention.",
      "short_abstract": "Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offer limited actionable insights for real-time driver feedback. These approaches often lack continuous risk quantification,...",
      "published": "2026-03-16",
      "updated": "2026-03-16",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Hardware / Real-Time",
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 156,
      "arxiv_url": "https://arxiv.org/abs/2603.14841",
      "pdf_url": "https://arxiv.org/pdf/2603.14841.pdf"
    },
    {
      "id": "2603.14616",
      "anchor": "paper-2603-14616",
      "title": "Functional Safety Analysis for Infrastructure-Enabled Depot Autonomy System",
      "authors": [
        "Gaurav Pandey",
        "Gregory Stevens",
        "Henry Liu"
      ],
      "abstract": "This paper presents the functional safety analysis for an Infrastructure-Enabled Depot Autonomy (IX-DA) system. The IX-DA system automates the marshalling of delivery vehicles within a controlled depot environment, navigating connected autonomous vehicles (CAVs) between drop-off zones, service stations (washing, calibration, charging, loading), and pick-up zones without human intervention. We describe the system architecture comprising three principal subsystems -- the connected autonomous vehicle, the infrastructure sensing and compute layer, and the human operator interface -- and derive their functional requirements. Using ISO 26262-compliant Hazard Analysis and Risk Assessment (HARA) methodology, we identify eight hazardous events, evaluate them across different operating scenarios, and assign Automotive Safety Integrity Levels~(ASILs) ranging from Quality Management (QM) to ASIL C. Six safety goals are derived and allocated to vehicle and infrastructure subsystems. The analysis demonstrates that high-speed uncontrolled operation imposes the most demanding safety requirements (ASIL C), while controlled low-speed operation reduces most goals to QM, offering a practical pathway for phased deployment.",
      "short_abstract": "This paper presents the functional safety analysis for an Infrastructure-Enabled Depot Autonomy (IX-DA) system. The IX-DA system automates the marshalling of delivery vehicles within a controlled depot environment, navigating connected autonomous vehicles...",
      "published": "2026-03-15",
      "updated": "2026-03-15",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 9,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: risk assessment"
        ]
      },
      "recency": "archive",
      "age_days": 157,
      "arxiv_url": "https://arxiv.org/abs/2603.14616",
      "pdf_url": "https://arxiv.org/pdf/2603.14616.pdf"
    },
    {
      "id": "2603.14124",
      "anchor": "paper-2603-14124",
      "title": "Experimental Evaluation of Security Attacks on Self-Driving Car Platforms",
      "authors": [
        "Viet K. Nguyen",
        "Nathan Lee",
        "Mohammad Husain"
      ],
      "abstract": "Deep learning-based perception pipelines in autonomous ground vehicles are vulnerable to both adversarial manipulation and network-layer disruption. We present a systematic, on-hardware experimental evaluation of five attack classes: FGSM, PGD, man-in-the-middle (MitM), denial-of-service (DoS), and phantom attacks on low-cost autonomous vehicle platforms (JetRacer and Yahboom). Using a standardized 13-second experimental protocol and comprehensive automated logging, we systematically characterize three dimensions of attack behavior:(i) control deviation, (ii) computational cost, and (iii) runtime responsiveness. Our analysis reveals that distinct attack classes produce consistent and separable \"fingerprints\" across these dimensions: perception attacks (MitM output manipulation and phantom projection) generate high steering deviation signatures with nominal computational overhead, PGD produces combined steering perturbation and computational load signatures across multiple dimensions, and DoS exhibits frame rate and latency degradation signatures with minimal control-plane perturbation. We demonstrate that our fingerprinting framework generalizes across both digital attacks (adversarial perturbations, network manipulation) and environmental attacks (projected false features), providing a foundation for attack-aware monitoring systems and targeted, signature-based defense mechanisms.",
      "short_abstract": "Deep learning-based perception pipelines in autonomous ground vehicles are vulnerable to both adversarial manipulation and network-layer disruption. We present a systematic, on-hardware experimental evaluation of five attack classes: FGSM, PGD,...",
      "published": "2026-03-14",
      "updated": "2026-03-14",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 24,
        "evidence": [
          "title: security",
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 158,
      "arxiv_url": "https://arxiv.org/abs/2603.14124",
      "pdf_url": "https://arxiv.org/pdf/2603.14124.pdf"
    },
    {
      "id": "2603.12647",
      "anchor": "paper-2603-12647",
      "title": "LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction",
      "authors": [
        "ZY Chen",
        "F Zhu",
        "H Zhu",
        "DY Kong",
        "XK Kuang",
        "YJ Zhang",
        "CM Jiang"
      ],
      "abstract": "Recent 3D Gaussian Splatting (3DGS) methods have demonstrated the feasibility of self-driving scene reconstruction and novel view synthesis. However, most existing methods either rely solely on cameras or use LiDAR only for Gaussian initialization or depth supervision, while the rich scene information contained in point clouds, such as reflectance, and the complementarity between LiDAR and RGB have not been fully exploited, leading to degradation in challenging self-driving scenes, such as those with high ego-motion and complex lighting. To address these issues, we propose a robust and efficient LiDAR-reflectance-guided Salient Gaussian Splatting method (LR-SGS) for self-driving scenes, which introduces a structure-aware Salient Gaussian representation, initialized from geometric and reflectance feature points extracted from LiDAR and refined through a salient transform and improved density control to capture edge and planar structures. Furthermore, we calibrate LiDAR intensity into reflectance and attach it to each Gaussian as a lighting-invariant material channel, jointly aligned with RGB to enforce boundary consistency. Extensive experiments on the Waymo Open Dataset demonstrate that LR-SGS achieves superior reconstruction performance with fewer Gaussians and shorter training time. In particular, on Complex Lighting scenes, our method surpasses OmniRe by 1.18 dB PSNR.",
      "short_abstract": "Recent 3D Gaussian Splatting (3DGS) methods have demonstrated the feasibility of self-driving scene reconstruction and novel view synthesis. However, most existing methods either rely solely on cameras or use LiDAR only for Gaussian initialization or depth...",
      "published": "2026-03-13",
      "updated": "2026-03-13",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 9,
        "evidence": [
          "abstract: point cloud",
          "title: LiDAR",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 159,
      "arxiv_url": "https://arxiv.org/abs/2603.12647",
      "pdf_url": "https://arxiv.org/pdf/2603.12647.pdf"
    },
    {
      "id": "2603.13675",
      "anchor": "paper-2603-13675",
      "title": "Hidden Risks of Unmonitored GPUs in Intelligent Transportation Systems",
      "authors": [
        "Sefatun-Noor Puspa",
        "Mashrur Chowdhury"
      ],
      "abstract": "Graphics processing units (GPUs) power many intelligent transportation systems (ITS) and automated driving applications, but remain largely unmonitored for safety and security. This article highlights GPU misuse as a critical blind spot, showing how unmanaged GPU workloads silently degrade real-time performance, demonstrating the need for stronger security measures in ITS.",
      "short_abstract": "Graphics processing units (GPUs) power many intelligent transportation systems (ITS) and automated driving applications, but remain largely unmonitored for safety and security. This article highlights GPU misuse as a critical blind spot, showing how unmanaged...",
      "published": "2026-03-13",
      "updated": "2026-03-13",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 6,
        "evidence": [
          "abstract: safety",
          "abstract: security"
        ]
      },
      "recency": "archive",
      "age_days": 159,
      "arxiv_url": "https://arxiv.org/abs/2603.13675",
      "pdf_url": "https://arxiv.org/pdf/2603.13675.pdf"
    },
    {
      "id": "2603.13136",
      "anchor": "paper-2603-13136",
      "title": "Unifying Decision Making and Trajectory Planning in Automated Driving through Time-Varying Potential Fields",
      "authors": [
        "David Costa",
        "Francesco Cerrito",
        "Massimo Canale",
        "Carlo Novara"
      ],
      "abstract": "This paper proposes a unified decision making and local trajectory planning framework based on Time-Varying Artificial Potential Fields (TVAPFs). The TVAPF explicitly models the predicted motion via bounded uncertainty of dynamic obstacles over the planning horizon, using information from perception and V2X sources when available. TVAPFs are embedded into a finite horizon optimal control problem that jointly selects the driving maneuver and computes a feasible, collision free trajectory. The effectiveness and real-time suitability of the approach are demonstrated through a simulation test in a multi-actor scenario with real road topology, highlighting the advantages of the unified TVAPF-based formulation.",
      "short_abstract": "This paper proposes a unified decision making and local trajectory planning framework based on Time-Varying Artificial Potential Fields (TVAPFs). The TVAPF explicitly models the predicted motion via bounded uncertainty of dynamic obstacles over the planning...",
      "published": "2026-03-13",
      "updated": "2026-03-13",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 48,
        "margin": 39,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "title: decision-making",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 159,
      "arxiv_url": "https://arxiv.org/abs/2603.13136",
      "pdf_url": "https://arxiv.org/pdf/2603.13136.pdf"
    },
    {
      "id": "2603.11987",
      "anchor": "paper-2603-11987",
      "title": "LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories",
      "authors": [
        "Qianpu Sun",
        "Xiaowei Chi",
        "Yuhan Rui",
        "Ying Li",
        "Kuangzhi Ge",
        "Jiajun Li",
        "Sirui Han",
        "Shanghang Zhang"
      ],
      "abstract": "Artificial intelligence is increasingly catalyzing scientific automation, with multimodal large language model (MLLM) agents evolving from lab assistants into self-driving lab operators. This transition imposes stringent safety requirements on laboratory environments, where fragile glassware, hazardous substances, and high-precision laboratory equipment render planning errors or misinterpreted risks potentially irreversible. However, the safety awareness and decision-making reliability of embodied agents in such high-stakes settings remain insufficiently defined and evaluated. To bridge this gap, we introduce LABSHIELD, a realistic multi-view benchmark designed to assess MLLMs in hazard identification and safety-critical reasoning. Grounded in U.S. Occupational Safety and Health Administration (OSHA) standards and the Globally Harmonized System (GHS), LABSHIELD establishes a rigorous safety taxonomy spanning 164 operational tasks with diverse manipulation complexities and risk profiles. We evaluate 20 proprietary models, 9 open-source models, and 3 embodied models under a dual-track evaluation framework. Our results reveal a systematic gap between general-domain MCQ accuracy and Semi-open QA safety performance, with models exhibiting an average drop of 32.0% in professional laboratory scenarios, particularly in hazard interpretation and safety-aware planning. These findings underscore the urgent necessity for safety-centric reasoning frameworks to ensure reliable autonomous scientific experimentation in embodied laboratory contexts. The full dataset will be released soon.",
      "short_abstract": "Artificial intelligence is increasingly catalyzing scientific automation, with multimodal large language model (MLLM) agents evolving from lab assistants into self-driving lab operators. This transition imposes stringent safety requirements on laboratory...",
      "published": "2026-03-12",
      "updated": "2026-03-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: title: planning",
          "Planning & Decision-Making: abstract: planning",
          "Planning & Decision-Making: abstract: decision-making",
          "Safety, Security & Verification: title: safety",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 160,
      "arxiv_url": "https://arxiv.org/abs/2603.11987",
      "pdf_url": "https://arxiv.org/pdf/2603.11987.pdf"
    },
    {
      "id": "2603.12607",
      "anchor": "paper-2603-12607",
      "title": "CarPLAN: Context-Adaptive and Robust Planning with Dynamic Scene Awareness for Autonomous Driving",
      "authors": [
        "Junyong Yun",
        "Jungho Kim",
        "ByungHyun Lee",
        "Dongyoung Lee",
        "Sehwan Choi",
        "Seunghyeop Nam",
        "Kichun Jo",
        "Jun Won Choi"
      ],
      "abstract": "Imitation learning (IL) is widely used for motion planning in autonomous driving due to its data efficiency and access to real-world driving data. For safe and robust real-world driving, IL-based planning requires capturing the complex driving contexts inherent in real-world data and enabling context-adaptive decision-making, rather than relying solely on expert trajectory imitation. In this paper, we propose CarPLAN, a novel IL-based motion planning framework that explicitly enhances driving context understanding and enables adaptive planning across diverse traffic scenarios. Our contributions are twofold: We introduce Displacement-Aware Predictive Encoding (DPE) to improve the model's spatial awareness by predicting future displacement vectors between the Autonomous Vehicle (AV) and surrounding scene elements. This allows the planner to account for relational spacing when generating trajectories. In addition to the standard imitation loss, we incorporate an augmented loss term that captures displacement prediction errors, ensuring planning decisions consider relative distances from other agents. To improve the model's ability to handle diverse driving contexts, we propose Context-Adaptive Multi-Expert Decoder (CMD), which leverages the Mixture of Experts (MoE) framework. CMD dynamically selects the most suitable expert decoders based on scene structure at each Transformer layer, enabling adaptive and context-aware planning in dynamic environments. We evaluate CarPLAN on the nuPlan benchmark and demonstrate state-of-the-art performance across all closed-loop simulation metrics. In particular, CarPLAN exhibits robust performance on challenging scenarios such as Test14-Hard, validating its effectiveness in complex driving conditions. Additional experiments on the Waymax benchmark further demonstrate its generalization capability across different benchmark settings.",
      "short_abstract": "Imitation learning (IL) is widely used for motion planning in autonomous driving due to its data efficiency and access to real-world driving data. For safe and robust real-world driving, IL-based planning requires capturing the complex driving contexts...",
      "published": "2026-03-12",
      "updated": "2026-03-12",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [
        "Imitation Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 21,
        "margin": 13,
        "evidence": [
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 160,
      "arxiv_url": "https://arxiv.org/abs/2603.12607",
      "pdf_url": "https://arxiv.org/pdf/2603.12607.pdf"
    },
    {
      "id": "2603.12036",
      "anchor": "paper-2603-12036",
      "title": "Single Pixel Image Classification using an Ultrafast Digital Light Projector",
      "authors": [
        "Aisha Kanwal",
        "Graeme E. Johnstone",
        "Fahimeh Dehkhoda",
        "Johannes H. Herrnsdorf",
        "Robert K. Henderson",
        "Martin D. Dawson",
        "Xavier Porte",
        "Michael J. Strain"
      ],
      "abstract": "Pattern recognition and image classification are essential tasks in machine vision. Autonomous vehicles, for example, require being able to collect the complex information contained in a changing environment and classify it in real time. Here, we experimentally demonstrate image classification at multi-kHz frame rates combining the technique of single pixel imaging (SPI) with a low complexity machine learning model. The use of a microLED-on-CMOS digital light projector for SPI enables ultrafast pattern generation for sub-ms image encoding. We investigate the classification accuracy of our experimental system against the broadly accepted benchmarking task of the MNIST digits classification. We compare the classification performance of two machine learning models: An extreme learning machine (ELM) and a backpropagation trained deep neural network. The complexity of both models is kept low so the overhead added to the inference time is comparable to the image generation time. Crucially, our single pixel image classification approach is based on a spatiotemporal transformation of the information, entirely bypassing the need for image reconstruction. By exploring the performance of our SPI based ELM as binary classifier we demonstrate its potential for efficient anomaly detection in ultrafast imaging scenarios.",
      "short_abstract": "Pattern recognition and image classification are essential tasks in machine vision. Autonomous vehicles, for example, require being able to collect the complex information contained in a changing environment and classify it in real time. Here, we...",
      "published": "2026-03-12",
      "updated": "2026-03-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 5,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 160,
      "arxiv_url": "https://arxiv.org/abs/2603.12036",
      "pdf_url": "https://arxiv.org/pdf/2603.12036.pdf"
    },
    {
      "id": "2603.11572",
      "anchor": "paper-2603-11572",
      "title": "Quantum computing for transport research: an introduction, systematic review, and perspective",
      "authors": [
        "Lachlan Oberg",
        "Paul Corry",
        "Moji Ghadimi",
        "Ashish Bhaskar"
      ],
      "abstract": "Transport engineering has significant potential to benefit from quantum computing. The rise of intelligent transport systems, autonomous vehicles, and the Internet of Things has created an unprecedented demand for efficient information processing and computational optimisation. Accordingly, transport engineers and scientists have explored the ever-improving capabilities of quantum computers in an effort to meet this demand. Motivated by this growing interest, this paper sets out four aims: (1) to introduce the fundamental aspects of quantum computing relevant to the transport domain, (2) to identify transport-related problems which are suitable for quantum acceleration, (3) to develop a pipeline for solving these problems, and (4) to provide a systematic review of the existing literature. For the latter, a systematic search of the Scopus database (and supplemented by additional citation sources) identified 103 studies for inclusion following PRISMA 2020 guidelines. While a diverse set of use cases have been proposed, we conclude that future research should prioritise problems where quantum computation offers a clear practical benefit. To this end, we suggest promising directions to guide further work in this burgeoning subfield.",
      "short_abstract": "Transport engineering has significant potential to benefit from quantum computing. The rise of intelligent transport systems, autonomous vehicles, and the Internet of Things has created an unprecedented demand for efficient information processing and...",
      "published": "2026-03-12",
      "updated": "2026-03-12",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Survey"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: steering or braking"
        ]
      },
      "recency": "archive",
      "age_days": 160,
      "arxiv_url": "https://arxiv.org/abs/2603.11572",
      "pdf_url": "https://arxiv.org/pdf/2603.11572.pdf"
    },
    {
      "id": "2603.11380",
      "anchor": "paper-2603-11380",
      "title": "DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding",
      "authors": [
        "Mingzhe Tao",
        "Ruiping Liu",
        "Junwei Zheng",
        "Yufan Chen",
        "Kedi Ying",
        "M. Saquib Sarfraz",
        "Kailun Yang",
        "Jiaming Zhang",
        "Rainer Stiefelhagen"
      ],
      "abstract": "Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leveraging multi-sensor information to understand adverse driving scenarios in autonomous vehicles. To address this gap, we propose the DriveXQA, a multimodal dataset for autonomous driving VQA. In addition to four visual modalities, five sensor failure cases, and five weather conditions, it includes $102,505$ QA pairs categorized into three types: global scene level, allocentric level, and ego-vehicle centric level. Since no existing MLLM framework adopts multiple complementary visual modalities as input, we design MVX-LLM, a token-efficient architecture with a Dual Cross-Attention (DCA) projector that fuses the modalities to alleviate information redundancy. Experiments demonstrate that our DCA achieves improved performance under challenging conditions such as foggy (GPTScore: $53.5$ vs. $25.1$ for the baseline).",
      "short_abstract": "Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leveraging multi-sensor information to...",
      "published": "2026-03-11",
      "updated": "2026-03-11",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 7,
        "evidence": [
          "title: scene understanding"
        ]
      },
      "recency": "archive",
      "age_days": 161,
      "arxiv_url": "https://arxiv.org/abs/2603.11380",
      "pdf_url": "https://arxiv.org/pdf/2603.11380.pdf"
    },
    {
      "id": "2603.10688",
      "anchor": "paper-2603-10688",
      "title": "MapGCLR: Geospatial Contrastive Learning of Representations for Online Vectorized HD Map Construction",
      "authors": [
        "Jonas Merkert",
        "Alexander Blumberg",
        "Jan-Hendrik Pauls",
        "Christoph Stiller"
      ],
      "abstract": "Autonomous vehicles rely on map information to understand the world around them. However, the creation and maintenance of offline high-definition (HD) maps remains costly. A more scalable alternative lies in online HD map construction, which only requires map annotations at training time. To further reduce the need for annotating vast training labels, self-supervised training provides an alternative. This work focuses on improving the latent birds-eye-view (BEV) feature grid representation within a vectorized online HD map construction model by enforcing geospatial consistency between overlapping BEV feature grids as part of a contrastive loss function. To ensure geospatial overlap for contrastive pairs, we introduce an approach to analyze the overlap between traversals within a given dataset and generate subsidiary dataset splits following adjustable multi-traversal requirements. We train the same model supervised using a reduced set of single-traversal labeled data and self-supervised on a broader unlabeled set of data following our multi-traversal requirements, effectively implementing a semi-supervised approach. Our approach outperforms the supervised baseline across the board, both quantitatively in terms of the downstream tasks vectorized map perception performance and qualitatively in terms of segmentation in the principal component analysis (PCA) visualization of the BEV feature space.",
      "short_abstract": "Autonomous vehicles rely on map information to understand the world around them. However, the creation and maintenance of offline high-definition (HD) maps remains costly. A more scalable alternative lies in online HD map construction, which only requires map...",
      "published": "2026-03-11",
      "updated": "2026-03-11",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 31,
        "evidence": [
          "title: HD map",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 161,
      "arxiv_url": "https://arxiv.org/abs/2603.10688",
      "pdf_url": "https://arxiv.org/pdf/2603.10688.pdf"
    },
    {
      "id": "2603.09255",
      "anchor": "paper-2603-09255",
      "title": "Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning",
      "authors": [
        "Kanishkha Jaisankar",
        "Pranav M. Pawar",
        "Diana Susan Joseph",
        "Raja Muthalagu",
        "Mithun Mukherjee",
        "Dnyaneshawar Mantri",
        "Ramjee Prasad"
      ],
      "abstract": "Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars. These techniques play a crucial role in enabling self-driving cars to perceive and understand their surroundings, allowing them to safely navigate and make decisions in real-time. Using Neural Networks self-driving cars can accurately identify and classify objects such as pedestrians, other vehicles, and traffic signals. Using deep learning and analyzing data from sensors such as cameras and radar, self-driving cars can predict the likely movement of other objects and plan their own actions accordingly. In this study, a novel approach to enhance the performance of self-driving cars by using pre-trained and custom-made neural networks for key tasks, including traffic sign classification, vehicle detection, lane detection, and behavioral cloning is provided. The methodology integrates several innovative techniques, such as geometric and color transformations for data augmentation, image normalization, and transfer learning for feature extraction. These techniques are applied to diverse datasets, including the German Traffic Sign Recognition Benchmark (GTSRB), road and lane segmentation datasets, vehicle detection datasets, and data collected using the Udacity self-driving car simulator to evaluate the model efficacy. The primary objective of the work is to review the state-of-the-art in deep learning and computer vision for self-driving cars. The findings of the work are effective in solving various challenges related to self-driving cars like traffic sign classification, lane prediction, vehicle detection, and behavioral cloning, and provide valuable insights into improving the robustness and reliability of autonomous systems, paving the way for future research and deployment of safer and more efficient self-driving technologies.",
      "short_abstract": "Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars. These techniques play a crucial role in enabling self-driving cars to perceive and understand their surroundings, allowing them to safely...",
      "published": "2026-03-10",
      "updated": "2026-03-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "high",
        "score": 20,
        "margin": 15,
        "evidence": [
          "abstract: radar",
          "abstract: camera",
          "title: lane detection",
          "abstract: lane detection",
          "abstract: traffic sign recognition"
        ]
      },
      "recency": "archive",
      "age_days": 162,
      "arxiv_url": "https://arxiv.org/abs/2603.09255",
      "pdf_url": "https://arxiv.org/pdf/2603.09255.pdf"
    },
    {
      "id": "2603.09455",
      "anchor": "paper-2603-09455",
      "title": "Declarative Scenario-based Testing with RoadLogic",
      "authors": [
        "Ezio Bartocci",
        "Alessio Gambi",
        "Felix Gigler",
        "Cristinel Mateis",
        "Dejan Ničković"
      ],
      "abstract": "Scenario-based testing is a key method for cost-effective and safe validation of autonomous vehicles (AVs). Existing approaches rely on imperative scenario definitions, requiring developers to manually enumerate numerous variants to achieve coverage. Declarative languages, such as ASAM OpenSCENARIO DSL (OS2), raise the abstraction level but lack systematic methods for instantiating concrete and specification-compliant scenarios. To our knowledge, currently, no open-source solution provides this capability. We present RoadLogic that bridges declarative OS2 specifications and executable simulations. It uses Answer Set Programming to generate abstract plans satisfying scenario constraints, motion planning to refine the plans into feasible trajectories, and specification-based monitoring to verify correctness. We evaluate RoadLogic on instantiating representative OS2 scenarios executed in the CommonRoad framework. Results show that RoadLogic consistently produces realistic, specification-satisfying simulations within minutes and captures diverse behavioral variants through parameter sampling, thus opening the door to systematic scenario-based testing for autonomous driving systems.",
      "short_abstract": "Scenario-based testing is a key method for cost-effective and safe validation of autonomous vehicles (AVs). Existing approaches rely on imperative scenario definitions, requiring developers to manually enumerate numerous variants to achieve coverage....",
      "published": "2026-03-10",
      "updated": "2026-03-10",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 4,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 162,
      "arxiv_url": "https://arxiv.org/abs/2603.09455",
      "pdf_url": "https://arxiv.org/pdf/2603.09455.pdf"
    },
    {
      "id": "2603.09420",
      "anchor": "paper-2603-09420",
      "title": "Class-Incremental Motion Forecasting",
      "authors": [
        "Nicolas Schischka",
        "Nikhil Gosala",
        "B Ravi Kiran",
        "Senthil Yogamani",
        "Abhinav Valada"
      ],
      "abstract": "Motion forecasting enables autonomous vehicles to anticipate scene evolution by predicting the future trajectories of dynamic agents. However, existing approaches typically assume a closed-world setting with a fixed object taxonomy and access to high-quality perception, limiting their applicability in the real world where perception is imperfect, and new object classes may emerge over time. In this work, we introduce class-incremental motion forecasting, a novel setting in which new object classes are sequentially introduced over time and future object trajectories are predicted directly from camera images. We propose the first end-to-end framework for this setting, which adapts to newly introduced classes while mitigating catastrophic forgetting of previously learned ones. Our method generates motion forecasting pseudo-labels for known classes and matches them with 2D instance masks from an open-vocabulary segmentation model. This 3D-to-2D keypoint voting mechanism filters inconsistent and overconfident predictions, while a query feature variance-based replay strategy samples informative past sequences to preserve prior knowledge. Extensive evaluations on nuScenes and Argoverse 2 show that our approach successfully preserves performance on known classes while effectively adapting to novel ones. We further demonstrate zero-shot transfer to real-world driving and show that the framework extends naturally to open- and closed-loop end-to-end class-incremental planning on nuScenes and NeuroNCAP. Code and models will be made publicly available at https://omen.cs.uni-freiburg.de.",
      "short_abstract": "Motion forecasting enables autonomous vehicles to anticipate scene evolution by predicting the future trajectories of dynamic agents. However, existing approaches typically assume a closed-world setting with a fixed object taxonomy and access to high-quality...",
      "published": "2026-03-10",
      "updated": "2026-03-10",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 27,
        "evidence": [
          "title: motion forecasting",
          "abstract: motion forecasting",
          "title: forecasting",
          "abstract: forecasting"
        ]
      },
      "recency": "archive",
      "age_days": 162,
      "arxiv_url": "https://arxiv.org/abs/2603.09420",
      "pdf_url": "https://arxiv.org/pdf/2603.09420.pdf"
    },
    {
      "id": "2603.09399",
      "anchor": "paper-2603-09399",
      "title": "Vision-Augmented On-Track System Identification for Autonomous Racing via Attention-Based Priors and Iterative Neural Correction",
      "authors": [
        "Zhiping Wu",
        "Cheng Hu",
        "Yiqin Wang",
        "Lei Xie",
        "Hongye Su"
      ],
      "abstract": "Operating autonomous vehicles at the absolute limits of handling requires precise, real-time identification of highly non-linear tire dynamics. However, traditional online optimization methods suffer from \"cold-start\" initialization failures and struggle to model high-frequency transient dynamics. To address these bottlenecks, this paper proposes a novel vision-augmented, iterative system identification framework. First, a lightweight CNN (MobileNetV3) translates visual road textures into a continuous heuristic friction prior, providing a robust \"warm-start\" for parameter optimization. Next, a S4 model captures complex temporal dynamic residuals, circumventing the memory and latency limitations of traditional MLPs and RNNs. Finally, a derivative-free Nelder-Mead algorithm iteratively extracts physically interpretable Pacejka tire parameters via a hybrid virtual simulation. Co-simulation in CarSim demonstrates that the lightweight vision backbone reduces friction estimation error by 76.1 using 85 fewer FLOPs, accelerating cold-start convergence by 71.4. Furthermore, the S4-augmented framework improves parameter extraction accuracy and decreases lateral force RMSE by over 60 by effectively capturing complex vehicle dynamics, demonstrating superior performance compared to conventional neural architectures.",
      "short_abstract": "Operating autonomous vehicles at the absolute limits of handling requires precise, real-time identification of highly non-linear tire dynamics. However, traditional online optimization methods suffer from \"cold-start\" initialization failures and struggle to...",
      "published": "2026-03-10",
      "updated": "2026-03-10",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 0,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: vehicle dynamics",
          "Systems, Deployment & Connectivity: abstract: deployment or real-time",
          "Systems, Deployment & Connectivity: abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 162,
      "arxiv_url": "https://arxiv.org/abs/2603.09399",
      "pdf_url": "https://arxiv.org/pdf/2603.09399.pdf"
    },
    {
      "id": "2603.09256",
      "anchor": "paper-2603-09256",
      "title": "Platooning as a Service (PlaaS): A Sustainable Transportation Framework for Connected and Autonomous Vehicles",
      "authors": [
        "Bhosale Akshay Tanaji",
        "Sayak Roychowdhury",
        "Anand Abraham"
      ],
      "abstract": "Vehicular platooning is a well-known transportation technique that helps reduce fuel consumption, carbon emissions, and road congestion. When integrated with Connected and Autonomous Vehicle (CAV) technologies, platooning enhances the overall safety and efficiency of the transportation system. This article presents Platooning as a Service (PlaaS) platform, a decision-support framework to promote sustainable transportation through platooning. We have formulated this problem as a Stackelberg game, with the platoon service provider (PSP) as the leader and the service users as the followers. The PSP sets the pricing policy, and the follower responds by choosing the distance to be travelled with the platoon. The optimal service contract between the PSP and the follower vehicle is established with the Stackelberg equilibrium, which is derived using Karush-Kuhn-Tucker optimality conditions. Additionally, we have examined the impact of government subsidies on the PlaaS platform in reducing carbon emissions. Our model has been applied to a specific example problem to illustrate the benefits for both players. We have also derived managerial insights through sensitivity analysis, exploring the effect of different velocity levels, vehicle dimensions, government subsidies, and operational urgency on the utilities of the players and CO2 emissions. Our analysis shows that PSP gains a higher profit from high delay-cost vehicles performing time-critical operations with higher platoon velocity. However, the benefits related to fuel consumption are only realized at moderate platoon velocities.",
      "short_abstract": "Vehicular platooning is a well-known transportation technique that helps reduce fuel consumption, carbon emissions, and road congestion. When integrated with Connected and Autonomous Vehicle (CAV) technologies, platooning enhances the overall safety and...",
      "published": "2026-03-10",
      "updated": "2026-03-10",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 12,
        "evidence": [
          "title: connected vehicles",
          "abstract: connected vehicles"
        ]
      },
      "recency": "archive",
      "age_days": 162,
      "arxiv_url": "https://arxiv.org/abs/2603.09256",
      "pdf_url": "https://arxiv.org/pdf/2603.09256.pdf"
    },
    {
      "id": "2603.09695",
      "anchor": "paper-2603-09695",
      "title": "DRIFT: Dual-Representation Inter-Fusion Transformer for Automated Driving Perception with 4D Radar Point Clouds",
      "authors": [
        "Siqi Pei",
        "Andras Palffy",
        "Dariu M. Gavrila"
      ],
      "abstract": "4D radars, which provide 3D point cloud data along with Doppler velocity, are attractive components of modern automated driving systems due to their low cost and robustness under adverse weather conditions. However, they provide a significantly lower point cloud density than LiDAR sensors. This makes it important to exploit not only local but also global contextual scene information. This paper proposes DRIFT, a model that effectively captures and fuses both local and global contexts through a dual-path architecture. The model incorporates a point path to aggregate fine-grained local features and a pillar path to encode coarse-grained global features. These two parallel paths are intertwined via novel feature-sharing layers at multiple stages, enabling full utilization of both representations. DRIFT is evaluated on the widely used View-of-Delft (VoD) dataset and a proprietary internal dataset. It outperforms the baselines on the tasks of object detection and/or free road estimation. For example, DRIFT achieves a mean average precision (mAP) of 52.6% (compared to, say, 45.4% of CenterPoint) on the VoD dataset.",
      "short_abstract": "4D radars, which provide 3D point cloud data along with Doppler velocity, are attractive components of modern automated driving systems due to their low cost and robustness under adverse weather conditions. However, they provide a significantly lower point...",
      "published": "2026-03-10",
      "updated": "2026-03-10",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 37,
        "margin": 35,
        "evidence": [
          "title: perception",
          "abstract: object detection",
          "title: point cloud",
          "abstract: point cloud",
          "abstract: LiDAR",
          "title: radar"
        ]
      },
      "recency": "archive",
      "age_days": 162,
      "arxiv_url": "https://arxiv.org/abs/2603.09695",
      "pdf_url": "https://arxiv.org/pdf/2603.09695.pdf"
    },
    {
      "id": "2603.08645",
      "anchor": "paper-2603-08645",
      "title": "Retrieval-Augmented Gaussian Avatars: Improving Expression Generalization",
      "authors": [
        "Matan Levy",
        "Gavriel Habib",
        "Issar Tzachor",
        "Dvir Samuel",
        "Rami Ben-Ari",
        "Nir Darshan",
        "Or Litany",
        "Dani Lischinski"
      ],
      "abstract": "Template-free animatable head avatars can achieve high visual fidelity by learning expression-dependent facial deformation directly from a subject's capture, avoiding parametric face templates and hand-designed blendshape spaces. However, since learned deformation is supervised only by the expressions observed for a single identity, these models suffer from limited expression coverage and often struggle when driven by motions that deviate from the training distribution. We introduce RAF (Retrieval-Augmented Faces), a simple training-time augmentation designed for template-free head avatars that learn deformation from data. RAF constructs a large unlabeled expression bank and, during training, replaces a subset of the subject's expression features with nearest-neighbor expressions retrieved from this bank while still reconstructing the subject's original frames. This exposes the deformation field to a broader range of expression conditions, encouraging stronger identity-expression decoupling and improving robustness to expression distribution shift without requiring paired cross-identity data, additional annotations, or architectural changes. We further analyze how retrieval augmentation increases expression diversity and validate retrieval quality with a user study showing that retrieved neighbors are perceptually closer in expression and pose. Experiments on the NeRSemble benchmark demonstrate that RAF consistently improves expression fidelity over the baseline, in both self-driving and cross-driving scenarios.",
      "short_abstract": "Template-free animatable head avatars can achieve high visual fidelity by learning expression-dependent facial deformation directly from a subject's capture, avoiding parametric face templates and hand-designed blendshape spaces. However, since learned...",
      "published": "2026-03-09",
      "updated": "2026-03-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 163,
      "arxiv_url": "https://arxiv.org/abs/2603.08645",
      "pdf_url": "https://arxiv.org/pdf/2603.08645.pdf"
    },
    {
      "id": "2603.08611",
      "anchor": "paper-2603-08611",
      "title": "FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection",
      "authors": [
        "Anqi Joyce Yang",
        "James Tu",
        "Nikita Dvornik",
        "Enxu Li",
        "Raquel Urtasun"
      ],
      "abstract": "In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or traffic control devices. However, many safety-critical objects (e.g., construction worker) appear infrequently in nominal traffic conditions, leading to a severe shortage of training examples from driving data alone. Recent vision foundation models, which are trained on a large corpus of data, can serve as a good source of external prior knowledge to improve generalization. We propose FOMO-3D, the first multi-modal 3D detector to leverage vision foundation models for long-tailed 3D detection. Specifically, FOMO-3D exploits rich semantic and depth priors from OWLv2 and Metric3Dv2 within a two-stage detection paradigm that first generates proposals with a LiDAR-based branch and a novel camera-based branch, and refines them with attention especially to image features from OWL. Evaluations on real-world driving data show that using rich priors from vision foundation models with careful multi-modal fusion designs leads to large gains for long-tailed 3D detection. Project website is at https://waabi.ai/fomo3d/.",
      "short_abstract": "In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or traffic control devices. However, many safety-critical objects (e.g., construction worker) appear infrequently...",
      "published": "2026-03-09",
      "updated": "2026-03-09",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 17,
        "margin": 13,
        "evidence": [
          "title: object detection",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 163,
      "arxiv_url": "https://arxiv.org/abs/2603.08611",
      "pdf_url": "https://arxiv.org/pdf/2603.08611.pdf"
    },
    {
      "id": "2603.08438",
      "anchor": "paper-2603-08438",
      "title": "Graph Based Semantic Encoder Decoder Framework for Task Oriented Communications in Connected Autonomous Vehicles",
      "authors": [
        "Soheyb Ribouh",
        "Phil Polo Ditsia Di Ngoma"
      ],
      "abstract": "Connected autonomous vehicles (CAVs) require reliable and efficient communication frameworks to support safety critical and task-oriented applications such as collision avoidance, cooperative perception, and traffic risk assessment. Traditional communication paradigms, which focus on transmitting raw bits, often incur excessive bandwidth consumption and fail to preserve the semantic relevance of transmitted information. To bridge this gap, we propose a Graph-Based Semantic Encoder-Decoder (GBSED) architecture tailored for task-oriented communications in CAV networks. The encoder leverages scene graphs to capture spatial and semantic relationships among road entities, combined with a semantic compression algorithm that reduces the size of the extracted graph based representations by up to 99% compared to raw images, while the decoder reconstructs task relevant representations rather than raw data. This design enables a significant reduction in communication overhead while maintaining high semantic fidelity, exceeding 0.9 at SNR levels above 10dB, for downstream vehicular tasks. We evaluate the proposed framework through simulations in autonomous driving scenarios, where the semantic encoder and decoder are integrated into a MIMO OFDM physical layer system. The results demonstrate high prediction success rates for risk assessment, improved robustness under the 3GPP CDL channel, and significant compression gains, confirming that the proposed semantic communication framework is a promising solution for future 6G systems.",
      "short_abstract": "Connected autonomous vehicles (CAVs) require reliable and efficient communication frameworks to support safety critical and task-oriented applications such as collision avoidance, cooperative perception, and traffic risk assessment. Traditional communication...",
      "published": "2026-03-09",
      "updated": "2026-03-09",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 13,
        "evidence": [
          "abstract: cooperative systems",
          "title: connected vehicles",
          "abstract: connected vehicles",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 163,
      "arxiv_url": "https://arxiv.org/abs/2603.08438",
      "pdf_url": "https://arxiv.org/pdf/2603.08438.pdf"
    },
    {
      "id": "2603.09086",
      "anchor": "paper-2603-09086",
      "title": "Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges",
      "authors": [
        "Rongxiang Zeng",
        "Yongqi Dong"
      ],
      "abstract": "Emerging generative world models and vision-language-action (VLA) systems are rapidly reshaping automated driving by enabling scalable simulation, long-horizon forecasting, and capability-rich decision making. Across these directions, latent representations serve as the central computational substrate: they compress high-dimensional multi-sensor observations, enable temporally coherent rollouts, and provide interfaces for planning, reasoning, and controllable generation. This paper proposes a unifying latent-space framework that synthesizes recent progress in world models for automated driving. The framework organizes the design space by the target and form of latent representations (latent worlds, latent actions, latent generators; continuous states, discrete tokens, and hybrids) and by structural priors for geometry, topology, and semantics. Building on this taxonomy, the paper articulates five cross-cutting internal mechanics (i.e, structural isomorphism, long-horizon temporal stability, semantic and reasoning alignment, value-aligned objectives and post-training, as well as adaptive computation and deliberation) and connects these design choices to robustness, generalization, and deployability. The work also proposes concrete evaluation prescriptions, including a closed-loop metric suite and a resource-aware deliberation cost, designed to reduce the open-loop / closed-loop mismatch. Finally, the paper identifies actionable research directions toward advancing latent world model for decision-ready, verifiable, and resource-efficient automated driving.",
      "short_abstract": "Emerging generative world models and vision-language-action (VLA) systems are rapidly reshaping automated driving by enabling scalable simulation, long-horizon forecasting, and capability-rich decision making. Across these directions, latent representations...",
      "published": "2026-03-09",
      "updated": "2026-03-09",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "VLA",
        "World Model"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 12,
        "evidence": [
          "abstract: forecasting",
          "title: world model",
          "abstract: world model"
        ]
      },
      "recency": "archive",
      "age_days": 163,
      "arxiv_url": "https://arxiv.org/abs/2603.09086",
      "pdf_url": "https://arxiv.org/pdf/2603.09086.pdf"
    },
    {
      "id": "2603.08138",
      "anchor": "paper-2603-08138",
      "title": "Toward Governing Perception in Safety-Critical Mediated Reality on the Move",
      "authors": [
        "Pascal Jansen"
      ],
      "abstract": "Wearable Augmented Reality (AR) is increasingly deployed in on-the-move contexts such as automated driving, cycling, and pedestrian navigation. To date, most systems rely on additive overlays that highlight hazards, intentions, or predictions without altering the scene itself. However, advances in head-mounted displays and computer vision now enable Diminished and Modified Reality techniques that suppress, transform, or substitute scene elements. These capabilities conceptually extend AR into Mediated Reality (MR), shifting the design space from \"what to add\" to \"what is perceptually available.\" Because such mediation reshapes the evidential basis for situation awareness and trust calibration, it raises novel interaction challenges. This position paper argues that MR on the move must become governable, as users need mechanisms to configure, inspect, and understand mediation without compromising safety. Additionally, this position paper outlines design challenges related to governance granularity, epistemic signaling, and accountability, and frames MR on the move as a research agenda for governable perceptual mediation in dynamic, safety-critical environments.",
      "short_abstract": "Wearable Augmented Reality (AR) is increasingly deployed in on-the-move contexts such as automated driving, cycling, and pedestrian navigation. To date, most systems rely on additive overlays that highlight hazards, intentions, or predictions without altering...",
      "published": "2026-03-09",
      "updated": "2026-03-09",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 7,
        "evidence": [
          "title: safety",
          "abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 163,
      "arxiv_url": "https://arxiv.org/abs/2603.08138",
      "pdf_url": "https://arxiv.org/pdf/2603.08138.pdf"
    },
    {
      "id": "2603.08080",
      "anchor": "paper-2603-08080",
      "title": "MRDrive: An Open Source Mixed Reality Driving Simulator for Automotive User Research",
      "authors": [
        "Patrick Ebel",
        "Michał Patryk Miazga",
        "Martin Lorenz",
        "Timur Getselev",
        "Pavlo Bazilinskyy",
        "Celine Conzen"
      ],
      "abstract": "Designing and evaluating in-vehicle interfaces requires experimental platforms that combine ecological validity with experimental control. Driving simulators are widely used for this purpose. However, they face a fundamental trade-off: high-fidelity physical simulators are costly and difficult to adapt, while virtual reality simulators provide flexibility at the expense of physical interaction with the vehicle. In this work, we present MRDrive, an open mixed-reality driving simulator designed to support HCI research on in-vehicle interaction, attention, and explainability in manual and automated driving contexts. MRDrive enables drivers and passengers to interact with a real vehicle cabin while being fully immersed in a virtual driving environment. We demonstrate the capabilities of MRDrive through a small pilot study that illustrates how the simulator can be used to collect and analyze eye-tracking and touch interaction data in an automated driving scenario. MRDRive is available at: https://github.com/ciao-group/mrdrive",
      "short_abstract": "Designing and evaluating in-vehicle interfaces requires experimental platforms that combine ecological validity with experimental control. Driving simulators are widely used for this purpose. However, they face a fundamental trade-off: high-fidelity physical...",
      "published": "2026-03-09",
      "updated": "2026-03-09",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation",
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 3,
        "margin": 1,
        "evidence": [
          "Human Factors & Policy: abstract: explainability",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 163,
      "arxiv_url": "https://arxiv.org/abs/2603.08080",
      "pdf_url": "https://arxiv.org/pdf/2603.08080.pdf"
    },
    {
      "id": "2603.07794",
      "anchor": "paper-2603-07794",
      "title": "4DRC-OCC: Robust Semantic Occupancy Prediction Through Fusion of 4D Radar and Camera",
      "authors": [
        "David Ninfa",
        "Andras Palffy",
        "Holger Caesar"
      ],
      "abstract": "Autonomous driving requires robust perception across diverse environmental conditions, yet 3D semantic occupancy prediction remains challenging under adverse weather and lighting. In this work, we present the first study combining 4D radar and camera data for 3D semantic occupancy prediction. Our fusion leverages the complementary strengths of both modalities: 4D radar provides reliable range, velocity, and angle measurements in challenging conditions, while cameras contribute rich semantic and texture information. We further show that integrating depth cues from camera pixels enables lifting 2D images to 3D, improving scene reconstruction accuracy. Additionally, we introduce a fully automatically labeled dataset for training semantic occupancy models, substantially reducing reliance on costly manual annotation. Experiments demonstrate the robustness of 4D radar across diverse scenarios, highlighting its potential to advance autonomous vehicle perception.",
      "short_abstract": "Autonomous driving requires robust perception across diverse environmental conditions, yet 3D semantic occupancy prediction remains challenging under adverse weather and lighting. In this work, we present the first study combining 4D radar and camera data for...",
      "published": "2026-03-08",
      "updated": "2026-03-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 23,
        "evidence": [
          "abstract: perception",
          "title: occupancy",
          "abstract: occupancy",
          "title: radar",
          "abstract: radar",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 164,
      "arxiv_url": "https://arxiv.org/abs/2603.07794",
      "pdf_url": "https://arxiv.org/pdf/2603.07794.pdf"
    },
    {
      "id": "2603.07464",
      "anchor": "paper-2603-07464",
      "title": "Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection",
      "authors": [
        "Rui Ding",
        "Meng Yang",
        "Nanning Zheng"
      ],
      "abstract": "Monocular 3D object detection is a promising yet ill-posed task for autonomous vehicles due to the lack of accurate depth information. Cross-modality knowledge distillation could effectively transfer depth information from LiDAR to image-based network. However, modality gap between image and LiDAR seriously limits its accuracy. In this paper, we systematically investigate the negative transfer problem induced by modality gap in cross-modality distillation for the first time, including not only the architecture inconsistency issue but more importantly the feature overfitting issue. We propose a selective learning approach named MonoSTL to overcome these issues, which encourages positive transfer of depth information from LiDAR while alleviates the negative transfer on image-based network. On the one hand, we utilize similar architectures to ensure spatial alignment of features between image-based and LiDAR-based networks. On the other hand, we develop two novel distillation modules, namely Depth-Aware Selective Feature Distillation (DASFD) and Depth-Aware Selective Relation Distillation (DASRD), which selectively learn positive features and relationships of objects by integrating depth uncertainty into feature and relation distillations, respectively. Our approach can be seamlessly integrated into various CNN-based and DETR-based models, where we take three recent models on KITTI and a recent model on NuScenes for validation. Extensive experiments show that our approach considerably improves the accuracy of the base models and thereby achieves the best accuracy compared with all recently released SOTA models.",
      "short_abstract": "Monocular 3D object detection is a promising yet ill-posed task for autonomous vehicles due to the lack of accurate depth information. Cross-modality knowledge distillation could effectively transfer depth information from LiDAR to image-based network....",
      "published": "2026-03-08",
      "updated": "2026-03-08",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 14,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "abstract: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 164,
      "arxiv_url": "https://arxiv.org/abs/2603.07464",
      "pdf_url": "https://arxiv.org/pdf/2603.07464.pdf"
    },
    {
      "id": "2603.07314",
      "anchor": "paper-2603-07314",
      "title": "Faster-HEAL: An Efficient and Privacy-Preserving Collaborative Perception Framework for Heterogeneous Autonomous Vehicles",
      "authors": [
        "Armin Maleki",
        "Hayder Radha"
      ],
      "abstract": "Collaborative perception (CP) is a promising paradigm for improving situational awareness in autonomous vehicles by overcoming the limitations of single-agent perception. However, most existing approaches assume homogeneous agents, which restricts their applicability in real-world scenarios where vehicles use diverse sensors and perception models. This heterogeneity introduces a feature domain gap that degrades detection performance. Prior works address this issue by retraining entire models/major components, or using feature interpreters for each new agent type, which is computationally expensive, compromises privacy, and may reduce single-agent accuracy. We propose Faster-HEAL, a lightweight and privacy-preserving CP framework that fine-tunes a low-rank visual prompt to align heterogeneous features with a unified feature space while leveraging pyramid fusion for robust feature aggregation. This approach reduces the trainable parameters by 94%, enabling efficient adaptation to new agents without retraining large models. Experiments on the OPV2V-H dataset show that Faster-HEAL improves detection performance by 2% over state-of-the-art methods with significantly lower computational overhead, offering a practical solution for scalable heterogeneous CP.",
      "short_abstract": "Collaborative perception (CP) is a promising paradigm for improving situational awareness in autonomous vehicles by overcoming the limitations of single-agent perception. However, most existing approaches assume homogeneous agents, which restricts their...",
      "published": "2026-03-07",
      "updated": "2026-03-07",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 0,
        "evidence": [
          "Perception & Sensor Fusion: title: perception",
          "Perception & Sensor Fusion: abstract: perception",
          "Systems, Deployment & Connectivity: title: cooperative systems",
          "Systems, Deployment & Connectivity: abstract: cooperative systems"
        ]
      },
      "recency": "archive",
      "age_days": 165,
      "arxiv_url": "https://arxiv.org/abs/2603.07314",
      "pdf_url": "https://arxiv.org/pdf/2603.07314.pdf"
    },
    {
      "id": "2603.06850",
      "anchor": "paper-2603-06850",
      "title": "Nonlinear Performance Degradation of Vision-Based Teleoperation under Network Latency",
      "authors": [
        "Aws Khalil",
        "Jaerock Kwon"
      ],
      "abstract": "Teleoperation is increasingly being adopted as a critical fallback for autonomous vehicles. However, the impact of network latency on vision-based, perception-driven control remains insufficiently studied. The present work investigates the nonlinear degradation of closed-loop stability in camera-based lane keeping under varying network delays. To conduct this study, we developed the Latency-Aware Vision Teleoperation testbed (LAVT), a research-oriented ROS 2 framework that enables precise, distributed one-way latency measurement and reproducible delay injection. Using LAVT, we performed 180 closed-loop experiments in simulation across diverse road geometries. Our findings reveal a sharp collapse in stability between 150 ms and 225 ms of one-way perception latency, where route completion rates drop from 100% to below 50% as oscillatory instability and phase-lag effects emerge. We further demonstrate that additional control-channel delay compounds these effects, significantly accelerating system failure even under constant visual latency. By combining this systematic empirical characterization with the LAVT testbed, this work provides quantitative insights into perception-driven instability and establishes a reproducible baseline for future latency-compensation and predictive control strategies. Project page, supplementary video, and code are available at https://bimilab.github.io/paper-LAVT",
      "short_abstract": "Teleoperation is increasingly being adopted as a critical fallback for autonomous vehicles. However, the impact of network latency on vision-based, perception-driven control remains insufficiently studied. The present work investigates the nonlinear...",
      "published": "2026-03-06",
      "updated": "2026-03-06",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 31,
        "evidence": [
          "title: teleoperation",
          "abstract: teleoperation",
          "title: communications",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 166,
      "arxiv_url": "https://arxiv.org/abs/2603.06850",
      "pdf_url": "https://arxiv.org/pdf/2603.06850.pdf"
    },
    {
      "id": "2603.06544",
      "anchor": "paper-2603-06544",
      "title": "Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving",
      "authors": [
        "Yuhan Zhou",
        "Mehri Sattari",
        "Haihua Chen",
        "Kewei Sha"
      ],
      "abstract": "Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-making. In practice, data quality (DQ) varies across sources and modalities due to environmental conditions and sensor limitations, yet AV research has largely prioritized algorithm design over DQ analysis. This work focuses on redundancy as a fundamental but underexplored DQ issue in AV datasets. Using the nuScenes and Argoverse 2 (AV2) datasets, we model and measure redundancy in multisource camera data and multimodal image-LiDAR data, and evaluate how removing redundant labels affects the YOLOv8 object detection task. Experimental results show that selectively removing redundant multisource image object labels from cameras with shared fields of view improves detection. In nuScenes, mAP${50}$ gains from $0.66$ to $0.70$, $0.64$ to $0.67$, and from $0.53$ to $0.55$, on three representative overlap regions, while detection on other overlapping camera pairs remains at the baseline even under stronger pruning. In AV2, $4.1$-$8.6\\%$ of labels are removed, and mAP${50}$ stays near the $0.64$ baseline. Multimodal analysis also reveals substantial redundancy between image and LiDAR data. These findings demonstrate that redundancy is a measurable and actionable DQ factor with direct implications for AV performance. This work highlights the role of redundancy as a data quality factor in AV perception and motivates a data-centric perspective for evaluating and improving AV datasets. Code, data, and implementation details are publicly available at: https://github.com/yhZHOU515/RedundancyAD",
      "short_abstract": "Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-making. In practice, data quality (DQ) varies across sources and modalities due to environmental conditions and sensor...",
      "published": "2026-03-06",
      "updated": "2026-03-06",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "abstract: perception",
          "abstract: object detection",
          "abstract: LiDAR",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 166,
      "arxiv_url": "https://arxiv.org/abs/2603.06544",
      "pdf_url": "https://arxiv.org/pdf/2603.06544.pdf"
    },
    {
      "id": "2603.06054",
      "anchor": "paper-2603-06054",
      "title": "Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving",
      "authors": [
        "Nikos Theodoridis",
        "Reenu Mohandas",
        "Ganesh Sistu",
        "Anthony Scanlan",
        "Ciarán Eising",
        "Tim Brophy"
      ],
      "abstract": "The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios. However, these models often fail on simple visual questions that are highly relevant to automated driving, and the reasons behind these failures remain poorly understood. In this work, we examine the intermediate activations of VLMs and assess the extent to which specific visual concepts are linearly encoded, with the goal of identifying bottlenecks in the flow of visual information. Specifically, we create counterfactual image sets that differ only in a targeted visual concept and then train linear probes to distinguish between them using the activations of five state-of-the-art (SOTA) VLMs, including two training variants for one model. Our results show that concepts such as the presence of an object or agent in a scene are explicitly and linearly encoded, whereas other spatial visual concepts, such as the orientation of an object or agent, are only implicitly encoded by the spatial structure retained by the vision encoder, or are not linearly encoded at all. In parallel, we observe that in certain cases, even when a concept is linearly encoded in the model's activations, the model still fails to answer correctly. This leads us to identify two failure modes. The first is perceptual failure, where the visual information required to answer a question is not linearly encoded in the model's activations. The second is cognitive failure, where the visual information is present but the model fails to align it correctly with language semantics. Finally, we show that increasing the distance of the object in question quickly degrades the linear separability of the corresponding visual concept...",
      "short_abstract": "The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios. However, these models often fail on simple...",
      "published": "2026-03-06",
      "updated": "2026-03-06",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 166,
      "arxiv_url": "https://arxiv.org/abs/2603.06054",
      "pdf_url": "https://arxiv.org/pdf/2603.06054.pdf"
    },
    {
      "id": "2603.05350",
      "anchor": "paper-2603-05350",
      "title": "A Shift-Invariant Deep Learning Framework for Automated Analysis of XPS Spectra",
      "authors": [
        "Issa Saddiq",
        "Yuxin Fan",
        "Robert G. Palgrave",
        "Mark A. Isaacs",
        "David Morgan",
        "Keith T. Butler"
      ],
      "abstract": "X-ray Photoelectron Spectroscopy (XPS) is a crucial technique for material surface analysis, yet interpreting its spectra is often challenging for both human analysts and automated methods due to the prevalence of variable spectral shifts and overlapping peaks. This project introduces a machine learning solution using a Spatial Transformer Network (STN), a type of neural network that implicitly learns to align spectra. An STN model was designed to classify the chemical environments present in an input spectrum, using functional groups as a proxy. The model was trained and tested on a large synthetic dataset of 100,000 spectra, created by linearly combining real experimental data from a library of 104 polymers. \\cite{RN22} To simulate experimental variability, random uniform shifts and broadening were applied to the data. The STN was found to effectively correct for random electrostatic shifts (up to 3.0 eV) and achieved relatively high accuracy ($\\sim$ 82\\%) in identifying functional groups, despite utilizing a much simpler architecture than previous work. These findings demonstrate that neural networks can effectively learn the underlying relationships between spectral features and chemical composition when they are able to intrinsically account for variable shifts. This work advances the development of more reliable automated XPS analysis, offering potential as an assistive tool for researchers and as a core component in future autonomous systems like self-driving laboratories.",
      "short_abstract": "X-ray Photoelectron Spectroscopy (XPS) is a crucial technique for material surface analysis, yet interpreting its spectra is often challenging for both human analysts and automated methods due to the prevalence of variable spectral shifts and overlapping...",
      "published": "2026-03-05",
      "updated": "2026-03-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 167,
      "arxiv_url": "https://arxiv.org/abs/2603.05350",
      "pdf_url": "https://arxiv.org/pdf/2603.05350.pdf"
    },
    {
      "id": "2603.05716",
      "anchor": "paper-2603-05716",
      "title": "Introducing the transitional autonomous vehicle lane-changing dataset: Empirical Experiments",
      "authors": [
        "Abhinav Sharma",
        "Zijun He",
        "Danjue Chen"
      ],
      "abstract": "Transitional autonomous vehicles (tAVs), which operate beyond SAE Level 1-2 automation but short of full autonomy, are increasingly sharing the road with human-driven vehicles (HDVs). As these systems interact during complex maneuvers such as lane changes, new patterns may emerge with implications for traffic stability and safety. Assessing these dynamics, particularly during mandatory lane changes, requires high-resolution trajectory data, yet datasets capturing tAV lane-changing behavior are scarce. This study introduces the North Carolina Transitional Autonomous Vehicle Lane-Changing (NC-tALC) Dataset, a high-fidelity trajectory dataset designed to characterize tAV interactions during lane-changing maneuvers. The dataset includes two controlled experimental series. In the first, tAV lane-changing experiments, a tAV executes lane changes in the presence of adaptive cruise control (ACC) equipped target vehicles, enabling analysis of lane-changing execution. In the second, tAV responding experiments, two tAVs act as followers and respond to cut-in maneuvers initiated by another tAV, enabling analysis of follower response dynamics. The dataset contains 152 trials (72 lane-changing and 80 responding trials) sampled at 20 Hz with centimeter-level RTK-GPS accuracy. The NC-tALC dataset provides a rigorous empirical foundation for evaluating tAV decision-making and interaction dynamics in controlled mandatory lane-changing scenarios.",
      "short_abstract": "Transitional autonomous vehicles (tAVs), which operate beyond SAE Level 1-2 automation but short of full autonomy, are increasingly sharing the road with human-driven vehicles (HDVs). As these systems interact during complex maneuvers such as lane changes,...",
      "published": "2026-03-05",
      "updated": "2026-03-05",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 0,
        "evidence": [
          "Planning & Decision-Making: abstract: decision-making",
          "Safety, Security & Verification: abstract: safety"
        ]
      },
      "recency": "archive",
      "age_days": 167,
      "arxiv_url": "https://arxiv.org/abs/2603.05716",
      "pdf_url": "https://arxiv.org/pdf/2603.05716.pdf"
    },
    {
      "id": "2603.05165",
      "anchor": "paper-2603-05165",
      "title": "V2N-Based Algorithm and Communication Protocol for Autonomous Non-Stop Intersections",
      "authors": [
        "Lorenzo Farina",
        "Lorenzo Mario Amorosa",
        "Marco Rapelli",
        "Barbara Maví Masini",
        "Claudio Casetti",
        "Alessandro Bazzi"
      ],
      "abstract": "Intersections are critical areas for road safety and traffic efficiency, accounting for a significant portion of vehicle crashes and fatalities. While connected and autonomous vehicle (CAV) technologies offer a promising solution for autonomous intersection management, many existing proposals either rely on computationally heavy centralized controllers or overlook the practical impairments of real-world communication networks. This paper introduces seamless mobility of vehicles over intersections (Moveover), a novel algorithm comprising a vehicle-to-network (V2N) communication protocol designed to let vehicles cross autonomous intersections without stopping. Moveover delegates trajectory and speed profile selection to individual vehicles, allowing each CAV to optimize them according to its unique kinematic characteristics. Simultaneously, a local intersection controller prevents collisions through deterministic conflict zone reservations. The algorithm is rigorously evaluated under both ideal and non-ideal networking conditions, specifically modeling 4G and 5G communication delays, across multiple layouts including single-lane, multi-lane, and roundabouts. Furthermore, we test Moveover on a real urban map with multiple intersections. Simulation results demonstrate that Moveover significantly outperforms baseline strategies, offering substantial improvements in travel times and reduced pollutant emissions.",
      "short_abstract": "Intersections are critical areas for road safety and traffic efficiency, accounting for a significant portion of vehicle crashes and fatalities. While connected and autonomous vehicle (CAV) technologies offer a promising solution for autonomous intersection...",
      "published": "2026-03-05",
      "updated": "2026-03-05",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 167,
      "arxiv_url": "https://arxiv.org/abs/2603.05165",
      "pdf_url": "https://arxiv.org/pdf/2603.05165.pdf"
    },
    {
      "id": "2603.05042",
      "anchor": "paper-2603-05042",
      "title": "CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection",
      "authors": [
        "Zhaonian Kuang",
        "Rui Ding",
        "Haotian Wang",
        "Xinhu Zheng",
        "Meng Yang",
        "Gang Hua"
      ],
      "abstract": "Multi-camera 3D object detection (MC3D) has attracted increasing attention with the growing deployment of multi-sensor physical agents, such as robots and autonomous vehicles. However, MC3D models still struggle to generalize to unseen platforms with new multi-camera configurations. Current solutions simply employ a meta-camera for unified representation but lack comprehensive consideration. In this paper, we revisit this issue and identify that the devil lies in spatial prior discrepancies across source and target configurations, including different intrinsics, extrinsics, and array layouts. To address this, we propose CoIn3D, a generalizable MC3D framework that enables strong transferability from source configurations to unseen target ones. CoIn3D explicitly incorporates all identified spatial priors into both feature embedding and image observation through spatial-aware feature modulation (SFM) and camera-aware data augmentation (CDA), respectively. SFM enriches feature space by integrating four spatial representations, such as focal length, ground depth, ground gradient, and Plücker coordinate. CDA improves observation diversity under various configurations via a training-free dynamic novel-view image synthesis scheme. Extensive experiments demonstrate that CoIn3D achieves strong cross-configuration performance on landmark datasets such as NuScenes, Waymo, and Lyft, under three dominant MC3D paradigms represented by BEVDepth, BEVFormer, and PETR.",
      "short_abstract": "Multi-camera 3D object detection (MC3D) has attracted increasing attention with the growing deployment of multi-sensor physical agents, such as robots and autonomous vehicles. However, MC3D models still struggle to generalize to unseen platforms with new...",
      "published": "2026-03-05",
      "updated": "2026-03-05",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 21,
        "evidence": [
          "title: object detection",
          "abstract: object detection",
          "title: camera",
          "abstract: camera"
        ]
      },
      "recency": "archive",
      "age_days": 167,
      "arxiv_url": "https://arxiv.org/abs/2603.05042",
      "pdf_url": "https://arxiv.org/pdf/2603.05042.pdf"
    },
    {
      "id": "2603.04695",
      "anchor": "paper-2603-04695",
      "title": "Selecting Spots by Explicitly Predicting Intention from Motion History Improves Performance in Autonomous Parking",
      "authors": [
        "Long Kiu Chung",
        "David Isele",
        "Faizan M. Tariq",
        "Sangjae Bae",
        "Shreyas Kousik",
        "Jovin D'sa"
      ],
      "abstract": "In many applications of social navigation, existing works have shown that predicting and reasoning about human intentions can help robotic agents make safer and more socially acceptable decisions. In this work, we study this problem for autonomous valet parking (AVP), where an autonomous vehicle ego agent must drop off its passengers, explore the parking lot, find a parking spot, negotiate for the spot with other vehicles, and park in the spot without human supervision. Specifically, we propose an AVP pipeline that selects parking spots by explicitly predicting where other agents are going to park from their motion history using learned models and probabilistic belief maps. To test this pipeline, we build a simulation environment with reactive agents and realistic modeling assumptions on the ego agent, such as occlusion-aware observations, and imperfect trajectory prediction. Simulation experiments show that our proposed method outperforms existing works that infer intentions from future predicted motion or embed them implicitly in end-to-end models, yielding better results in prediction accuracy, social acceptance, and task completion. Our key insight is that, in parking, where driving regulations are more lax, explicit intention prediction is crucial for reasoning about diverse and ambiguous long-term goals, which cannot be reliably inferred from short-term motion prediction alone, but can be effectively learned from motion history.",
      "short_abstract": "In many applications of social navigation, existing works have shown that predicting and reasoning about human intentions can help robotic agents make safer and more socially acceptable decisions. In this work, we study this problem for autonomous valet...",
      "published": "2026-03-04",
      "updated": "2026-03-04",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "abstract: trajectory prediction",
          "abstract: motion forecasting",
          "abstract: behavior prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 168,
      "arxiv_url": "https://arxiv.org/abs/2603.04695",
      "pdf_url": "https://arxiv.org/pdf/2603.04695.pdf"
    },
    {
      "id": "2603.04208",
      "anchor": "paper-2603-04208",
      "title": "GSeg3D: A High-Precision Grid-Based Algorithm for Safety-Critical Ground Segmentation in LiDAR Point Clouds",
      "authors": [
        "Muhammad Haider Khan Lodhi",
        "Christoph Hertzberg"
      ],
      "abstract": "Ground segmentation in point cloud data is the process of separating ground points from non-ground points. This task is fundamental for perception in autonomous driving and robotics, where safety and reliable operation depend on the precise detection of obstacles and navigable surfaces. Existing methods often fall short of the high precision required in safety-critical environments, leading to false detections that can compromise decision-making. In this work, we present a ground segmentation approach designed to deliver consistently high precision, supporting the stringent requirements of autonomous vehicles and robotic systems operating in real-world, safety-critical scenarios.",
      "short_abstract": "Ground segmentation in point cloud data is the process of separating ground points from non-ground points. This task is fundamental for perception in autonomous driving and robotics, where safety and reliable operation depend on the precise detection of...",
      "published": "2026-03-04",
      "updated": "2026-03-04",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 8,
        "evidence": [
          "abstract: perception",
          "title: point cloud",
          "abstract: point cloud",
          "title: LiDAR"
        ]
      },
      "recency": "archive",
      "age_days": 168,
      "arxiv_url": "https://arxiv.org/abs/2603.04208",
      "pdf_url": "https://arxiv.org/pdf/2603.04208.pdf"
    },
    {
      "id": "2603.03978",
      "anchor": "paper-2603-03978",
      "title": "Map-Agnostic And Interactive Safety-Critical Scenario Generation via Multi-Objective Tree Search",
      "authors": [
        "Wenyun Li",
        "Zejian Deng",
        "Chen Sun"
      ],
      "abstract": "Generating safety-critical scenarios is essential for validating the robustness of autonomous driving systems, yet existing methods often struggle to produce collisions that are both realistic and diverse while ensuring explicit interaction logic among traffic participants. This paper presents a novel framework for traffic-flow level safety-critical scenario generation via multi-objective Monte Carlo Tree Search (MCTS). We reframe trajectory feasibility and naturalistic behavior as optimization objectives within a unified evaluation function, enabling the discovery of diverse collision events without compromising realism. A hybrid Upper Confidence Bound (UCB) and Lower Confidence Bound (LCB) search strategy is introduced to balance exploratory efficiency with risk-averse decision-making. Furthermore, our method is map-agnostic and supports interactive scenario generation with each vehicle individually powered by SUMO's microscopic traffic models, enabling realistic agent behaviors in arbitrary geographic locations imported from OpenStreetMap. We validate our approach across four high-risk accident zones in Hong Kong's complex urban environments. Experimental results demonstrate that our framework achieves an 85\\% collision failure rate while generating trajectories with superior feasibility and comfort metrics. The resulting scenarios exhibit greater complexity, as evidenced by increased vehicle mileage and CO\\(_2\\) emissions. Our work provides a principled solution for stress testing autonomous vehicles through the generation of realistic yet infrequent corner cases at traffic-flow level.",
      "short_abstract": "Generating safety-critical scenarios is essential for validating the robustness of autonomous driving systems, yet existing methods often struggle to produce collisions that are both realistic and diverse while ensuring explicit interaction logic among...",
      "published": "2026-03-04",
      "updated": "2026-03-04",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 23,
        "margin": 19,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 168,
      "arxiv_url": "https://arxiv.org/abs/2603.03978",
      "pdf_url": "https://arxiv.org/pdf/2603.03978.pdf"
    },
    {
      "id": "2603.03848",
      "anchor": "paper-2603-03848",
      "title": "Dual-Interaction-Aware Cooperative Control Strategy for Alleviating Mixed Traffic Congestion",
      "authors": [
        "Zhengxuan Liu",
        "Yuxin Cai",
        "Yijing Wang",
        "Xiangkun He",
        "Chen Lv",
        "Zhiqiang Zuo"
      ],
      "abstract": "As Intelligent Transportation System (ITS) develops, Connected and Automated Vehicles (CAVs) are expected to significantly reduce traffic congestion through cooperative strategies, such as in bottleneck areas. However, the uncertainty and diversity in the behaviors of Human-Driven Vehicles (HDVs) in mixed traffic environments present major challenges for CAV cooperation. This paper proposes a Dual-Interaction-Aware Cooperative Control (DIACC) strategy that enhances both local and global interaction perception within the Multi-Agent Reinforcement Learning (MARL) framework for Connected and Automated Vehicles (CAVs) in mixed traffic bottleneck scenarios. The DIACC strategy consists of three key innovations: 1) A Decentralized Interaction-Adaptive Decision-Making (D-IADM) module that enhances actor's local interaction perception by distinguishing CAV-CAV cooperative interactions from CAV-HDV observational interactions. 2) A Centralized Interaction-Enhanced Critic (C-IEC) that improves critic's global traffic understanding through interaction-aware value estimation, providing more accurate guidance for policy updates. 3) A reward design that employs softmin aggregation with temperature annealing to prioritize interaction-intensive scenarios in mixed traffic. Additionally, a lightweight Proactive Safety-based Action Refinement (PSAR) module applies rule-based corrections to accelerate training convergence. Experimental results demonstrate that DIACC significantly improves traffic efficiency and adaptability compared to rule-based and benchmark MARL models.",
      "short_abstract": "As Intelligent Transportation System (ITS) develops, Connected and Automated Vehicles (CAVs) are expected to significantly reduce traffic congestion through cooperative strategies, such as in bottleneck areas. However, the uncertainty and diversity in the...",
      "published": "2026-03-04",
      "updated": "2026-03-04",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X",
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 8,
        "evidence": [
          "title: cooperative systems",
          "abstract: cooperative systems",
          "abstract: connected vehicles"
        ]
      },
      "recency": "archive",
      "age_days": 168,
      "arxiv_url": "https://arxiv.org/abs/2603.03848",
      "pdf_url": "https://arxiv.org/pdf/2603.03848.pdf"
    },
    {
      "id": "2603.03574",
      "anchor": "paper-2603-03574",
      "title": "Safety-Centered Scenario Generation for Autonomous Vehicles",
      "authors": [
        "Kiruthiga Chandra Shekar",
        "Aliasghar Moj Arab"
      ],
      "abstract": "This paper presents a scenario generation framework that creates diverse, parametrized, and safety-critical driving situations to validate the safety features of autonomous vehicles in simulation [15]. By modeling factors such as road geometry, traffic participants, environmental conditions, and perception uncertainties, the framework enables repeatable and scalable testing of safety mechanisms, including emergency braking, evasive maneuvers, and vulnerable road user protection. The framework supports both regulatory and edge case scenarios, mapped to hazards and safety goals derived from Hazard Analysis and Risk Assessment (HARA), ensuring traceability to ISO 26262 functional safety requirements and performance limitations. The output from these simulations provides quantitative safety metrics such as time-to-collision, minimum distance, braking and steering performance, and residual collision severity. These metrics enable the systematic evaluation of evasive maneuvering as a safety feature, while highlighting system limitations and edge-case vulnerabilities. Integration of scenario-based simulation with safety engineering principles offers accelerated validation cycles, improved test coverage at reduced cost, and stronger evidence for regulatory and stakeholder confidence.",
      "short_abstract": "This paper presents a scenario generation framework that creates diverse, parametrized, and safety-critical driving situations to validate the safety features of autonomous vehicles in simulation [15]. By modeling factors such as road geometry, traffic...",
      "published": "2026-03-03",
      "updated": "2026-03-03",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 25,
        "margin": 22,
        "evidence": [
          "title: safety",
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: risk assessment",
          "abstract: uncertainty"
        ]
      },
      "recency": "archive",
      "age_days": 169,
      "arxiv_url": "https://arxiv.org/abs/2603.03574",
      "pdf_url": "https://arxiv.org/pdf/2603.03574.pdf"
    },
    {
      "id": "2603.03452",
      "anchor": "paper-2603-03452",
      "title": "Impact of Localization Errors on Label Quality for Online HD Map Construction",
      "authors": [
        "Alexander Blumberg",
        "Jonas Merkert",
        "Richard Fehler",
        "Fabian Immel",
        "Frank Bieder",
        "Jan-Hendrik Pauls",
        "Christoph Stiller"
      ],
      "abstract": "High-definition (HD) maps are crucial for autonomous vehicles, but their creation and maintenance is very costly. This motivates the idea of online HD map construction. To provide a continuous large-scale stream of training data, existing HD maps can be used as labels for onboard sensor data from consumer vehicle fleets. However, compared to current, well curated HD map perception datasets, this fleet data suffers from localization errors, resulting in distorted map labels. We introduce three kinds of localization errors, Ramp, Gaussian, and Perlin noise, to examine their influence on generated map labels. We train a variant of MapTRv2, a state-of-the-art online HD map construction model, on the Argoverse 2 dataset with various levels of localization errors and assess the degradation of model performance. Since localization errors affect distant labels more severely, but are also less significant to driving performance, we introduce a distance-based map construction metric. Our experiments reveal that localization noise affects the model performance significantly. We demonstrate that errors in heading angle exert a more substantial influence than position errors, as angle errors result in a greater distortion of labels as distance to the vehicle increases. Furthermore, we can demonstrate that the model benefits from non-distorted ground truth (GT) data and that the performance decreases more than linearly with the increase in noisy data. Our study additionally provides a qualitative evaluation of the extent to which localization errors influence the construction of HD maps.",
      "short_abstract": "High-definition (HD) maps are crucial for autonomous vehicles, but their creation and maintenance is very costly. This motivates the idea of online HD map construction. To provide a continuous large-scale stream of training data, existing HD maps can be used...",
      "published": "2026-03-03",
      "updated": "2026-03-03",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "high",
        "score": 56,
        "margin": 53,
        "evidence": [
          "title: localization",
          "abstract: localization",
          "title: HD map",
          "abstract: HD map",
          "title: mapping",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 169,
      "arxiv_url": "https://arxiv.org/abs/2603.03452",
      "pdf_url": "https://arxiv.org/pdf/2603.03452.pdf"
    },
    {
      "id": "2603.03039",
      "anchor": "paper-2603-03039",
      "title": "Exploiting Repetitions and Interference Cancellation for the 6G-V2X Sidelink Autonomous Mode",
      "authors": [
        "Alessandro Bazzi",
        "Vittorio Todisco",
        "Antonella Molinaro",
        "Antoine O. Berthet",
        "Richard A. Stirling-Gallacher",
        "Claudia Campolo"
      ],
      "abstract": "In recent years, the Third Generation Partnership Project (3GPP) has developed the new radio-vehicle-to-everything (NR-V2X) sidelink standard, to enable direct communication between connected and autonomous vehicles (CAVs). Users can autonomously select radio resources for their transmissions with the Mode 2 channel access scheme, which can also operate under out-of-coverage conditions. However, Mode 2 performance is hindered by interference and packet collisions arising from dynamic mobile environments and limitations in assessing radio resource availability. The 3GPP specifications allow transmitting multiple copies of the same packet to improve reliability, though at the cost of increased channel congestion. This paper proposes to leverage receivers equipped with successive interference cancellation (SIC) capabilities, to exploit packet repetitions. Specifically, once a packet is successfully decoded the interfering contribution carried by repetitions can be cancelled from future or past received signals, enabling the decoding of new packets. Extensive highway scenario simulations demonstrate that the proposed solution significantly outperforms the legacy Mode 2 scheme, especially under high interference conditions, achieving improvements exceeding 100% in some cases.",
      "short_abstract": "In recent years, the Third Generation Partnership Project (3GPP) has developed the new radio-vehicle-to-everything (NR-V2X) sidelink standard, to enable direct communication between connected and autonomous vehicles (CAVs). Users can autonomously select radio...",
      "published": "2026-03-03",
      "updated": "2026-03-03",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 30,
        "margin": 30,
        "evidence": [
          "title: V2X",
          "abstract: V2X",
          "abstract: connected vehicles",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 169,
      "arxiv_url": "https://arxiv.org/abs/2603.03039",
      "pdf_url": "https://arxiv.org/pdf/2603.03039.pdf"
    },
    {
      "id": "2603.02200",
      "anchor": "paper-2603-02200",
      "title": "Adaptive Confidence Regularization for Multimodal Failure Detection",
      "authors": [
        "Moru Liu",
        "Hao Dong",
        "Olga Fink",
        "Mario Trapp"
      ],
      "abstract": "The deployment of multimodal models in high-stakes domains, such as self-driving vehicles and medical diagnostics, demands not only strong predictive performance but also reliable mechanisms for detecting failures. In this work, we address the largely unexplored problem of failure detection in multimodal contexts. We propose Adaptive Confidence Regularization (ACR), a novel framework specifically designed to detect multimodal failures. Our approach is driven by a key observation: in most failure cases, the confidence of the multimodal prediction is significantly lower than that of at least one unimodal branch, a phenomenon we term confidence degradation. To mitigate this, we introduce an Adaptive Confidence Loss that penalizes such degradations during training. In addition, we propose Multimodal Feature Swapping, a novel outlier synthesis technique that generates challenging, failure-aware training examples. By training with these synthetic failures, ACR learns to more effectively recognize and reject uncertain predictions, thereby improving overall reliability. Extensive experiments across four datasets, three modalities, and multiple evaluation settings demonstrate that ACR achieves consistent and robust gains. The source code will be available at https://github.com/mona4399/ACR.",
      "short_abstract": "The deployment of multimodal models in high-stakes domains, such as self-driving vehicles and medical diagnostics, demands not only strong predictive performance but also reliable mechanisms for detecting failures. In this work, we address the largely...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 10,
        "margin": 7,
        "evidence": [
          "abstract: robustness",
          "title: failure",
          "abstract: failure"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.02200",
      "pdf_url": "https://arxiv.org/pdf/2603.02200.pdf"
    },
    {
      "id": "2603.02538",
      "anchor": "paper-2603-02538",
      "title": "PathSpace: Rapid continuous map approximation for efficient SLAM using B-Splines in constrained environments",
      "authors": [
        "Aduen Benjumea",
        "Andrew Bradley",
        "Alexander Rast",
        "Matthias Rolf"
      ],
      "abstract": "Simultaneous Localization and Mapping (SLAM) plays a crucial role in enabling autonomous vehicles to navigate previously unknown environments. Semantic SLAM mostly extends visual SLAM, leveraging the higher density information available to reason about the environment in a more human-like manner. This allows for better decision making by exploiting prior structural knowledge of the environment, usually in the form of labels. Current semantic SLAM techniques still mostly rely on a dense geometric representation of the environment, limiting their ability to apply constraints based on context. We propose PathSpace, a novel semantic SLAM framework that uses continuous B-splines to represent the environment in a compact manner, while also maintaining and reasoning through the continuous probability density functions required for probabilistic reasoning. This system applies the multiple strengths of B-splines in the context of SLAM to interpolate and fit otherwise discrete sparse environments. We test this framework in the context of autonomous racing, where we exploit pre-specified track characteristics to produce significantly reduced representations at comparable levels of accuracy to traditional landmark based methods and demonstrate its potential in limiting the resources used by a system with minimal accuracy loss.",
      "short_abstract": "Simultaneous Localization and Mapping (SLAM) plays a crucial role in enabling autonomous vehicles to navigate previously unknown environments. Semantic SLAM mostly extends visual SLAM, leveraging the higher density information available to reason about the...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Mapping & Localization",
      "primary_category": "Mapping & Localization",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 33,
        "margin": 29,
        "evidence": [
          "abstract: localization",
          "title: SLAM",
          "abstract: SLAM",
          "abstract: mapping"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.02538",
      "pdf_url": "https://arxiv.org/pdf/2603.02538.pdf"
    },
    {
      "id": "2603.02528",
      "anchor": "paper-2603-02528",
      "title": "LLM-MLFFN: Multi-Level Autonomous Driving Behavior Feature Fusion via Large Language Model",
      "authors": [
        "Xiangyu Li",
        "Tianyi Wang",
        "Xi Cheng",
        "Rakesh Chowdary Machineni",
        "Zhaomiao Guo",
        "Sikai Chen",
        "Junfeng Jiao",
        "Christian Claudel"
      ],
      "abstract": "Accurate classification of autonomous vehicle (AV) driving behaviors is critical for safety validation, performance diagnosis, and traffic integration analysis. However, existing approaches primarily rely on numerical time-series modeling and often lack semantic abstraction, limiting interpretability and robustness in complex traffic environments. This paper presents LLM-MLFFN, a novel large language model (LLM)-enhanced multi-level feature fusion network designed to address the complexities of multi-dimensional driving data. The proposed LLM-MLFFN framework integrates priors from largescale pre-trained models and employs a multi-level approach to enhance classification accuracy. LLM-MLFFN comprises three core components: (1) a multi-level feature extraction module that extracts statistical, behavioral, and dynamic features to capture the quantitative aspects of driving behaviors; (2) a semantic description module that leverages LLMs to transform raw data into high-level semantic features; and (3) a dual-channel multi-level feature fusion network that combines numerical and semantic features using weighted attention mechanisms to improve robustness and prediction accuracy. Evaluation on the Waymo open trajectory dataset demonstrates the superior performance of the proposed LLM-MLFFN, achieving a classification accuracy of over 94%, surpassing existing machine learning models. Ablation studies further validate the critical contributions of multi-level fusion, feature extraction strategies, and LLM-derived semantic reasoning. These results suggest that integrating structured feature modeling with language-driven semantic abstraction provides a principled and interpretable pathway for robust autonomous driving behavior classification.",
      "short_abstract": "Accurate classification of autonomous vehicle (AV) driving behaviors is critical for safety validation, performance diagnosis, and traffic integration analysis. However, existing approaches primarily rely on numerical time-series modeling and often lack...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 9,
        "margin": 7,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.02528",
      "pdf_url": "https://arxiv.org/pdf/2603.02528.pdf"
    },
    {
      "id": "2603.02194",
      "anchor": "paper-2603-02194",
      "title": "From Leaderboard to Deployment: Code Quality Challenges in AV Perception Repositories",
      "authors": [
        "Mateus Karvat",
        "Bram Adams",
        "Sidney Givigi"
      ],
      "abstract": "Autonomous vehicle (AV) perception models are typically evaluated solely on benchmark performance metrics, with limited attention to code quality, production readiness and long-term maintainability. This creates a significant gap between research excellence and real-world deployment in safety-critical systems subject to international safety standards. To address this gap, we present the first large-scale empirical study of software quality in AV perception repositories, systematically analyzing 178 unique models from the KITTI and NuScenes 3D Object Detection leaderboards. Using static analysis tools (Pylint, Bandit, and Radon), we evaluated code errors, security vulnerabilities, maintainability, and development practices. Our findings revealed that only 7.3% of the studied repositories meet basic production-readiness criteria, defined as having zero critical errors and no high-severity security vulnerabilities. Security issues are highly concentrated, with the top five issues responsible for almost 80% of occurrences, which prompted us to develop a set of actionable guidelines to prevent them. Additionally, the adoption of Continuous Integration/Continuous Deployment pipelines was correlated with better code maintainability. Our findings highlight that leaderboard performance does not reflect production readiness and that targeted interventions could substantially improve the quality and safety of AV perception code.",
      "short_abstract": "Autonomous vehicle (AV) perception models are typically evaluated solely on benchmark performance metrics, with limited attention to code quality, production readiness and long-term maintainability. This creates a significant gap between research excellence...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Perception & Sensor Fusion",
      "primary_category": "Perception & Sensor Fusion",
      "tags": [
        "Benchmark"
      ],
      "classification": {
        "confidence": "medium",
        "score": 16,
        "margin": 4,
        "evidence": [
          "title: perception",
          "abstract: perception",
          "abstract: object detection"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.02194",
      "pdf_url": "https://arxiv.org/pdf/2603.02194.pdf"
    },
    {
      "id": "2603.01864",
      "anchor": "paper-2603-01864",
      "title": "Streaming Real-Time Trajectory Prediction Using Endpoint-Aware Modeling",
      "authors": [
        "Alexander Prutsch",
        "David Schinagl",
        "Horst Possegger"
      ],
      "abstract": "Future trajectories of neighboring traffic agents have a significant influence on the path planning and decision-making of autonomous vehicles. While trajectory forecasting is a well-studied field, research mainly focuses on snapshot-based prediction, where each scenario is treated independently of its global temporal context. However, real-world autonomous driving systems need to operate in a continuous setting, requiring real-time processing of data streams with low latency and consistent predictions over successive timesteps. We leverage this continuous setting to propose a lightweight yet highly accurate streaming-based trajectory forecasting approach. We integrate valuable information from previous predictions with a novel endpoint-aware modeling scheme. Our temporal context propagation uses the trajectory endpoints of the previous forecasts as anchors to extract targeted scenario context encodings. Our approach efficiently guides its scene encoder to extract highly relevant context information without needing refinement iterations or segment-wise decoding. Our experiments highlight that our approach effectively relays information across consecutive timesteps. Unlike methods using multi-stage refinement processing, our approach significantly reduces inference latency, making it well-suited for real-world deployment. We achieve state-of-the-art streaming trajectory prediction results on the Argoverse~2 multi-agent and single-agent benchmarks, while requiring substantially fewer resources.",
      "short_abstract": "Future trajectories of neighboring traffic agents have a significant influence on the path planning and decision-making of autonomous vehicles. While trajectory forecasting is a well-studied field, research mainly focuses on snapshot-based prediction, where...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 17,
        "evidence": [
          "title: trajectory prediction",
          "abstract: trajectory prediction",
          "abstract: forecasting",
          "title: prediction",
          "abstract: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.01864",
      "pdf_url": "https://arxiv.org/pdf/2603.01864.pdf"
    },
    {
      "id": "2603.01560",
      "anchor": "paper-2603-01560",
      "title": "(hu)Man vs. Machine: In the Future of Motorsport, can Autonomous Vehicles Compete?",
      "authors": [
        "Armand Amaritei",
        "Amber-Lily Blackadder",
        "Sebastian Donnelly",
        "Lora Hernandez",
        "James Vine",
        "Alexander Rast",
        "Matthias Rolf",
        "Andrew Bradley"
      ],
      "abstract": "Motorsport has historically driven technological innovation in the automotive industry. Autonomous racing provides a proving ground to push the limits of performance of autonomous vehicle (AV) systems. In principle, AVs could be at least as fast, if not faster, than humans. However, human driven racing provides broader audience appeal thus far, and is more strategically challenging. Both provide opportunities to push each other even further technologically, yet competitions remain separate. This paper evaluates whether the future of motorsport could encompass joint competition between humans and AVs. Analysis of the current state of the art, as well as recent competition outcomes, shows that while technical performance has reached comparable levels, there are substantial challenges in racecraft, strategy and safety that need to be overcome. Outstanding issues involved in mixed human-AI racing, ranging from an initial assessment of critical factors such as system-level latencies, to effective planning and risk guarantees are explored. The crucial non-technical aspect of audience engagement and appeal regarding the changing character of motorsport is addressed. In the wider context of motorsport and AVs, this work outlines a proposed agenda for future research to 'keep pushing the possible', in the true spirit of motorsport.",
      "short_abstract": "Motorsport has historically driven technological innovation in the automotive industry. Autonomous racing provides a proving ground to push the limits of performance of autonomous vehicle (AV) systems. In principle, AVs could be at least as fast, if not...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 1,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.01560",
      "pdf_url": "https://arxiv.org/pdf/2603.01560.pdf"
    },
    {
      "id": "2603.02526",
      "anchor": "paper-2603-02526",
      "title": "Event-Driven Safe and Resilient Control of Automated and Human-Driven Vehicles under EU-FDI Attacks",
      "authors": [
        "Yi Zhang",
        "Yichao Wang",
        "Wei Xiao",
        "Mohamadamin Rajabinezhad",
        "Shan Zuo"
      ],
      "abstract": "This paper studies the safe and resilient control of Connected and Automated Vehicles (CAVs) operating in mixed traffic environments where they must interact with Human-Driven Vehicles (HDVs) under uncertain dynamics and exponentially unbounded false data injection (EU-FDI) attacks. These attacks pose serious threats to safety-critical applications. While resilient control strategies can mitigate adversarial effects, they often overlook collision avoidance requirements. Conversely, safety-critical approaches tend to assume nominal operating conditions and lack resilience to adversarial inputs. To address these challenges, we propose an event-driven safe and resilient (EDSR) control framework that integrates event-driven Control Barrier Functions (CBFs) and Control Lyapunov Functions (CLFs) with adaptive attack-resilient control. The framework further incorporates data-driven estimation of HDV behaviors to ensure safety and resilience against EU-FDI attacks. Specifically, we focus on the lane-changing maneuver of CAVs in the presence of unpredictable HDVs and EU-FDI attacks on acceleration inputs. The event-driven approach reduces computational load while maintaining real-time safety guarantees. Simulation results, including comparisons with conventional safety-critical control methods that lack resilience, validate the effectiveness and robustness of the proposed EDSR framework in achieving collision-free maneuvers, stable velocity regulation, and resilient operation under adversarial conditions.",
      "short_abstract": "This paper studies the safe and resilient control of Connected and Automated Vehicles (CAVs) operating in mixed traffic environments where they must interact with Human-Driven Vehicles (HDVs) under uncertain dynamics and exponentially unbounded false data...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 22,
        "margin": 12,
        "evidence": [
          "abstract: safety",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.02526",
      "pdf_url": "https://arxiv.org/pdf/2603.02526.pdf"
    },
    {
      "id": "2603.01611",
      "anchor": "paper-2603-01611",
      "title": "Predictive Lane-Change and Routing Coordination in Bus-Priority Mixed Traffic Corridors",
      "authors": [
        "Tanlu Liang",
        "Ting Bai",
        "Andreas A. Malikopoulos"
      ],
      "abstract": "In this paper, we investigate the coordination of vehicle maneuvers in mixed-traffic corridors where connected and automated vehicles, human-driven vehicles, and buses interact under dedicated bus lane operations. We develop a segment-based network coordination framework that jointly optimizes lane-change and routing decisions of connected and automated vehicles to improve dedicated lane utilization while preserving bus priority. The proposed framework incorporates a predictive bus-protection mechanism that restricts vehicle access to protected lane segments within a monitoring horizon, together with a utility-driven lane-change strategy that accounts for anticipated travel time gains, downstream routing feasibility, and lane-change stability. By explicitly coupling network-level routing decisions with lane-level interaction control, the method proactively mitigates conflicts on dedicated lanes before congestion effects materialize. The proposed approach is evaluated through microscopic traffic simulations in SUMO using a realistic urban corridor. Simulation results demonstrate that the framework enhances bus schedule adherence and reduces average travel times for both automated and human-driven vehicles, while maintaining stable lane-change behavior without increasing maneuver frequency.",
      "short_abstract": "In this paper, we investigate the coordination of vehicle maneuvers in mixed-traffic corridors where connected and automated vehicles, human-driven vehicles, and buses interact under dedicated bus lane operations. We develop a segment-based network...",
      "published": "2026-03-02",
      "updated": "2026-03-02",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 4,
        "evidence": [
          "abstract: connected vehicles",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 170,
      "arxiv_url": "https://arxiv.org/abs/2603.01611",
      "pdf_url": "https://arxiv.org/pdf/2603.01611.pdf"
    },
    {
      "id": "2603.01202",
      "anchor": "paper-2603-01202",
      "title": "Opportunities and Challenges of Operating Semi-Autonomous Vehicles: A Layered Vulnerability Perspective",
      "authors": [
        "Soumita Mukherjee",
        "Priya Kumar",
        "Laura Cabrera"
      ],
      "abstract": "This study examines how vulnerability is produced for human operators of Tesla's Full Self-Driving (FSD), a Level 2 semi-autonomous vehicle (SAV) system, by applying Florencia Luna's layered vulnerability framework. While existing road safety models conceptualize vulnerability as a fixed attribute of external road users, emerging evidence suggests that semi-autonomous vehicle operators themselves experience dynamic and situational vulnerability as they supervise automated systems that they do not fully control. To investigate this phenomenon, we conducted semi-structured interviews with 17 active FSD users, analyzing their accounts through a combined deductive-inductive coding process aligned with Luna's framework. Findings reveal three interacting layers of operator vulnerability, namely psychological, operational, and social. Vulnerability emerged not from any single layer but from how these layers converged in specific situations, creating fluctuating supervisory demands and uneven capacity to recognize and manage risk. The findings extend debates on contextual trust calibration, automation complacency, and meaningful human control by demonstrating how factors commonly treated as liabilities such as trust or informal learning, can both increase and mitigate vulnerability depending on context. This analysis determines the need for design and regulatory interventions that address psychological, operational, and social conditions together rather than in isolation, and highlights how responsibility is implicitly shifted onto individual operators within inadequately supported supervisory regimes.",
      "short_abstract": "This study examines how vulnerability is produced for human operators of Tesla's Full Self-Driving (FSD), a Level 2 semi-autonomous vehicle (SAV) system, by applying Florencia Luna's layered vulnerability framework. While existing road safety models...",
      "published": "2026-03-01",
      "updated": "2026-03-01",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 4,
        "margin": 2,
        "evidence": [
          "Safety, Security & Verification: abstract: safety",
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 171,
      "arxiv_url": "https://arxiv.org/abs/2603.01202",
      "pdf_url": "https://arxiv.org/pdf/2603.01202.pdf"
    },
    {
      "id": "2603.01325",
      "anchor": "paper-2603-01325",
      "title": "From GEV to ResLogit: Spatially Correlated Discrete Choice Models for Pedestrian Movement Prediction",
      "authors": [
        "Rulla Al-Haideri",
        "Bilal Farooq"
      ],
      "abstract": "High frequency pedestrian motion forecasting when interacting with autonomous vehicles (AVs) can be enhanced through the use of behavioural frameworks, such as discrete choice models, that can explicitly account for correlation among similar movement alternatives. We formulate the pedestrian next step choice as a spatial discrete choice defined by a grid of speed adjustment and heading change. Using naturalistic pedestrian-AV encounters from nuScenes and Argoverse 2 (1 sec decision interval), we estimate a multinomial logit baseline and four spatial generalized extreme value (GEV) specifications (SCL, GSCL, SCNL, and GSCNL). We then compare them to a residual neural network logit (ResLogit) model that learns cross alternative effects while retaining an interpretable linear utility component. Across the evaluated data, spatial GEV structures yield only marginal improvements over multinomial logit, whereas ResLogit achieves a substantially better fit and produces behaviourally coherent errors concentrated among neighbouring grid cells. The results suggest that in dense, high frequency spatial choice sets, learning based residual corrections can capture proximity induced correlation more effectively than analyst specified GEV nesting structures, while maintaining interpretability.",
      "short_abstract": "High frequency pedestrian motion forecasting when interacting with autonomous vehicles (AVs) can be enhanced through the use of behavioural frameworks, such as discrete choice models, that can explicitly account for correlation among similar movement...",
      "published": "2026-03-01",
      "updated": "2026-03-01",
      "category": "Prediction & World Models",
      "primary_category": "Prediction & World Models",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "medium",
        "score": 14,
        "margin": 12,
        "evidence": [
          "abstract: motion forecasting",
          "abstract: forecasting",
          "title: prediction"
        ]
      },
      "recency": "archive",
      "age_days": 171,
      "arxiv_url": "https://arxiv.org/abs/2603.01325",
      "pdf_url": "https://arxiv.org/pdf/2603.01325.pdf"
    },
    {
      "id": "2603.13296",
      "anchor": "paper-2603-13296",
      "title": "Rationale Behind Human-Led Autonomous Truck Platooning",
      "authors": [
        "Yukun Lu",
        "Chenzhao Li",
        "Xintong Jiang",
        "Qiaoxuan Zhang"
      ],
      "abstract": "Autonomous trucking has progressed rapidly in recent years, transitioning from early demonstrations to OEM-integrated commercial deployments. However, fully driverless freight operations across heterogeneous climates, infrastructure conditions, and regulatory environments remain technically and socially challenging. This paper presents a systematic rationale for human-led autonomous truck platooning as a pragmatic intermediate pathway. First, we analyze 53 major truck accidents across North America (2021-2026) and show that human-related factors remain the dominant contributors to severe crashes, highlighting both the need for advanced assistance/automated driving systems and the complexity of real-world driving environments. Second, we review recent industry developments and identify persistent limitations in long-tail edge cases, winter operations, remote-region logistics, and large-scale safety validation. Based on these findings, we argue that a human-in-the-loop (HiL) platooning architecture offers layered redundancy, adaptive judgment in uncertain conditions, and a scalable validation framework. Furthermore, the dual-use capability of follower vehicles enables an evolutionary transition from coordinated platooning to independent autonomous operation. Rather than representing a compromise, human-led platooning provides a technically grounded and societally aligned bridge toward large-scale autonomous freight deployment.",
      "short_abstract": "Autonomous trucking has progressed rapidly in recent years, transitioning from early demonstrations to OEM-integrated commercial deployments. However, fully driverless freight operations across heterogeneous climates, infrastructure conditions, and regulatory...",
      "published": "2026-03-01",
      "updated": "2026-03-01",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 7,
        "margin": 2,
        "evidence": [
          "abstract: safety",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 171,
      "arxiv_url": "https://arxiv.org/abs/2603.13296",
      "pdf_url": "https://arxiv.org/pdf/2603.13296.pdf"
    },
    {
      "id": "2602.24096",
      "anchor": "paper-2602-24096",
      "title": "DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer",
      "authors": [
        "Yuxuan Zhang",
        "Katarína Tóthová",
        "Zian Wang",
        "Kangxue Yin",
        "Haithem Turki",
        "Riccardo de Lutio",
        "Yen-Yu Chang",
        "Or Litany",
        "Sanja Fidler",
        "Zan Gojcic"
      ],
      "abstract": "Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a promising solution as it enables simulating a wide variety of scenarios from real-world data alone in an automated and scalable way. However, while methods such as NeRF and 3D Gaussian Splatting can produce visually compelling results, they often exhibit artifacts particularly when rendering novel views, and fail to realistically integrate inserted dynamic objects, especially when they were captured from different scenes. To overcome these limitations, we introduce DiffusionHarmonizer, an online generative enhancement framework that transforms renderings from such imperfect scenes into temporally consistent outputs while improving their realism. At its core is a single-step temporally-conditioned enhancer that is converted from a pretrained multi-step image diffusion model, capable of running in online simulators on a single GPU. The key to training it effectively is a custom data curation pipeline that constructs synthetic-real pairs emphasizing appearance harmonization, artifact correction, and lighting realism. The result is a scalable system that significantly elevates simulation fidelity in both research and production environments.",
      "short_abstract": "Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a promising solution as it enables simulating a wide variety of scenarios from real-world data alone in an...",
      "published": "2026-02-27",
      "updated": "2026-02-27",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Simulation"
      ],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 173,
      "arxiv_url": "https://arxiv.org/abs/2602.24096",
      "pdf_url": "https://arxiv.org/pdf/2602.24096.pdf"
    },
    {
      "id": "2603.00217",
      "anchor": "paper-2603-00217",
      "title": "Physical Evaluation of Naturalistic Adversarial Patches for Camera-Based Traffic-Sign Detection",
      "authors": [
        "Brianna D'Urso",
        "Tahmid Hasan Sakib",
        "Syed Rafay Hasan",
        "Terry N. Guo"
      ],
      "abstract": "This paper studies how well Naturalistic Adversarial Patches (NAPs) transfer to a physical traffic sign setting when the detector is trained on a customized dataset for an autonomous vehicle (AV) environment. We construct a composite dataset, CompGTSRB (which is customized dataset for AV environment), by pasting traffic sign instances from the German Traffic Sign Recognition Benchmark (GTSRB) onto undistorted backgrounds captured from the target platform. CompGTSRB is used to train a YOLOv5 model and generate patches using a Generative Adversarial Network (GAN) with latent space optimization, following existing NAP methods. We carried out a series of experiments on our Quanser QCar testbed utilizing the front CSI camera provided in QCar. Across configurations, NAPs reduce the detector's STOP class confidence. Different configurations include distance, patch sizes, and patch placement. These results along with a detailed step-by-step methodology indicate the utility of CompGTSRB dataset and the proposed systematic physical protocols for credible patch evaluation. The research further motivate researching the defenses that address localized patch corruption in embedded perception pipelines.",
      "short_abstract": "This paper studies how well Naturalistic Adversarial Patches (NAPs) transfer to a physical traffic sign setting when the detector is trained on a customized dataset for an autonomous vehicle (AV) environment. We construct a composite dataset, CompGTSRB (which...",
      "published": "2026-02-27",
      "updated": "2026-02-27",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Dataset"
      ],
      "classification": {
        "confidence": "low",
        "score": 16,
        "margin": 2,
        "evidence": [
          "title: attack or threat",
          "abstract: attack or threat"
        ]
      },
      "recency": "archive",
      "age_days": 173,
      "arxiv_url": "https://arxiv.org/abs/2603.00217",
      "pdf_url": "https://arxiv.org/pdf/2603.00217.pdf"
    },
    {
      "id": "2602.23871",
      "anchor": "paper-2602-23871",
      "title": "Bandwidth-adaptive Cloud-Assisted 360-Degree 3D Perception for Autonomous Vehicles",
      "authors": [
        "Faisal Hawladera",
        "Rui Meireles",
        "Gamal Elghazaly",
        "Ana Aguiar",
        "Raphaël Frank"
      ],
      "abstract": "A key challenge for autonomous driving lies in maintaining real-time situational awareness regarding surrounding obstacles under strict latency constraints. The high processing requirements coupled with limited onboard computational resources can cause delay issues, particularly in complex urban settings. To address this, we propose leveraging Vehicle-to-Everything (V2X) communication to partially offload processing to the cloud, where compute resources are abundant, thus reducing overall latency. Our approach utilizes transformer-based models to fuse multi-camera sensor data into a comprehensive Bird's-Eye View (BEV) representation, enabling accurate 360-degree 3D object detection. The computation is dynamically split between the vehicle and the cloud based on the number of layers processed locally and the quantization level of the features. To further reduce network load, we apply feature vector clipping and compression prior to transmission. In a real-world experimental evaluation, our hybrid strategy achieved a 72 \\% reduction in end-to-end latency compared to a traditional onboard solution. To adapt to fluctuating network conditions, we introduce a dynamic optimization algorithm that selects the split point and quantization level to maximize detection accuracy while satisfying real-time latency constraints. Trace-based evaluation under realistic bandwidth variability shows that this adaptive approach improves accuracy by up to 20 \\% over static parameterization with the same latency performance.",
      "short_abstract": "A key challenge for autonomous driving lies in maintaining real-time situational awareness regarding surrounding obstacles under strict latency constraints. The high processing requirements coupled with limited onboard computational resources can cause delay...",
      "published": "2026-02-27",
      "updated": "2026-02-27",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 31,
        "margin": 14,
        "evidence": [
          "abstract: V2X",
          "abstract: deployment or real-time",
          "title: edge or cloud",
          "abstract: communications",
          "title: systems constraints",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 173,
      "arxiv_url": "https://arxiv.org/abs/2602.23871",
      "pdf_url": "https://arxiv.org/pdf/2602.23871.pdf"
    },
    {
      "id": "2602.22940",
      "anchor": "paper-2602-22940",
      "title": "Considering Perspectives for Automated Driving Ethics: Collective Risk in Vehicular Motion Planning",
      "authors": [
        "Leon Tolksdorf",
        "Arturo Tejada",
        "Christian Birkner",
        "Nathan van de Wouw"
      ],
      "abstract": "Recent automated vehicle (AV) motion planning strategies evolve around minimizing risk in road traffic. However, they exclusively consider risk from the AV's perspective and, as such, do not address the ethicality of its decisions for other road users. We argue that this does not reduce the risk of each road user, as risk may be different from the perspective of each road user. Indeed, minimizing the risk from the AV's perspective may not imply that the risk from the perspective of other road users is also being minimized; in fact, it may even increase. To test this hypothesis, we propose an AV motion planning strategy that supports switching risk minimization strategies between all road user perspectives. We find that the risk from the perspective of other road users can generally be considered different to the risk from the AV's perspective. Taking a collective risk perspective, i.e., balancing the risks of all road users, we observe an AV that minimizes overall traffic risk the best, while putting itself at slightly higher risk for the benefit of others, which is consistent with human driving behavior. In addition, adopting a collective risk minimization strategy can also be beneficial to the AV's travel efficiency by acting assertively when other road users maintain a low risk estimate of the AV. Yet, the AV drives conservatively when its planned actions are less predictable to other road users, i.e., associated with high risk. We argue that such behavior is a form of self-reflection and a natural prerequisite for socially acceptable AV behavior. We conclude that to facilitate ethicality in road traffic that includes AVs, the risk-perspective of each road user must be considered in the decision-making of AVs.",
      "short_abstract": "Recent automated vehicle (AV) motion planning strategies evolve around minimizing risk in road traffic. However, they exclusively consider risk from the AV's perspective and, as such, do not address the ethicality of its decisions for other road users. We...",
      "published": "2026-02-26",
      "updated": "2026-02-26",
      "category": "Planning & Decision-Making",
      "primary_category": "Planning & Decision-Making",
      "tags": [],
      "classification": {
        "confidence": "high",
        "score": 36,
        "margin": 24,
        "evidence": [
          "title: motion planning",
          "abstract: motion planning",
          "title: planning",
          "abstract: planning",
          "abstract: decision-making"
        ]
      },
      "recency": "archive",
      "age_days": 174,
      "arxiv_url": "https://arxiv.org/abs/2602.22940",
      "pdf_url": "https://arxiv.org/pdf/2602.22940.pdf"
    },
    {
      "id": "2602.23051",
      "anchor": "paper-2602-23051",
      "title": "An Empirical Analysis of Cooperative Perception for Occlusion Risk Mitigation",
      "authors": [
        "Aihong Wang",
        "Tenghui Xie",
        "Fuxi Wen",
        "Jun Li"
      ],
      "abstract": "Occlusions present a significant challenge for connected and automated vehicles, as they can obscure critical road users from perception systems. Traditional risk metrics often fail to capture the cumulative nature of these threats over time adequately. In this paper, we propose a novel and universal risk assessment metric, the Risk of Tracking Loss (RTL), which aggregates instantaneous risk intensity throughout occluded periods. This provides a holistic risk profile that encompasses both high-intensity, short-term threats and prolonged exposure. Utilizing diverse and high-fidelity real-world datasets, a large-scale statistical analysis is conducted to characterize occlusion risk and validate the effectiveness of the proposed metric. The metric is applied to evaluate different vehicle-to-everything (V2X) deployment strategies. Our study shows that full V2X penetration theoretically eliminates this risk, the reduction is highly nonlinear; a substantial statistical benefit requires a high penetration threshold of 75-90%. To overcome this limitation, we propose a novel asymmetric communication framework that allows even non-connected vehicles to receive warnings. Experimental results demonstrate that this paradigm achieves better risk mitigation performance. We found that our approach at 25% penetration outperforms the traditional symmetric model at 75%, and benefits saturate at only 50% penetration. This work provides a crucial risk assessment metric and a cost-effective, strategic roadmap for accelerating the safety benefits of V2X deployment.",
      "short_abstract": "Occlusions present a significant challenge for connected and automated vehicles, as they can obscure critical road users from perception systems. Traditional risk metrics often fail to capture the cumulative nature of these threats over time adequately. In...",
      "published": "2026-02-26",
      "updated": "2026-02-26",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 12,
        "evidence": [
          "abstract: V2X",
          "title: cooperative systems",
          "abstract: connected vehicles",
          "abstract: deployment or real-time",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 174,
      "arxiv_url": "https://arxiv.org/abs/2602.23051",
      "pdf_url": "https://arxiv.org/pdf/2602.23051.pdf"
    },
    {
      "id": "2602.22531",
      "anchor": "paper-2602-22531",
      "title": "Interpretable self-driving sputter epitaxy: from black-box optimization to human-usable growth rules",
      "authors": [
        "Yuki K. Wakabayashi",
        "Yui Ogawa",
        "Franz Benedict Romero",
        "Takuma Otsuka",
        "Yoshitaka Taniyasu"
      ],
      "abstract": "Self-driving laboratories have emerged as powerful tools for navigating high-dimensional process spaces, yet systems remain black-box optimizers that yield limited transferable process understanding. Here, we demonstrate an interpretable self-driving laboratory framework that transforms autonomous optimization into human-usable growth rules. As a stringent benchmark, we apply this framework to RF magnetron sputtering, addressing a long-standing challenge of achieving high-quality beta-Ga2O3 heteroepitaxy and single-crystalline beta-Ga2O3 homoepitaxy via sputtering. By combining Bayesian optimization with automated optical evaluation of the Urbach energy as a metric of sub-bandgap disorder, the self-driving system efficiently identifies heteroepitaxial growth conditions yielding a minimum Urbach energy of 182 meV, the lowest value for sputtered beta-Ga2O3 films. Importantly, the optimized growth window is transferable, realizing single-crystalline beta-Ga2O3 homoepitaxy without further optimization, corroborated by scanning transmission electron microscopy. To convert the closed-loop dataset into interpretable growth rules, we train a random forest surrogate and distill it into response curves and quantified pairwise interactions across the four-dimensional growth-parameter space. This analysis identifies substrate temperature as the primary control knob, with RF power and gas flows acting largely additively and only a modest temperature-oxygen coupling delineating the narrow window for high-quality growth, establishing a general route from autonomous experimentation to transferable growth rules.",
      "short_abstract": "Self-driving laboratories have emerged as powerful tools for navigating high-dimensional process spaces, yet systems remain black-box optimizers that yield limited transferable process understanding. Here, we demonstrate an interpretable self-driving...",
      "published": "2026-02-25",
      "updated": "2026-02-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 2,
        "margin": 2,
        "evidence": [
          "Control & Vehicle Dynamics: abstract: control"
        ]
      },
      "recency": "archive",
      "age_days": 175,
      "arxiv_url": "https://arxiv.org/abs/2602.22531",
      "pdf_url": "https://arxiv.org/pdf/2602.22531.pdf"
    },
    {
      "id": "2602.22041",
      "anchor": "paper-2602-22041",
      "title": "Using Feasible Action-Space Reduction by Groups to fill Causal Responsibility Gaps in Spatial Interactions",
      "authors": [
        "Ashwin George",
        "Vassil Guenov",
        "Arkady Zgonnikov",
        "David A. Abbink",
        "Luciano Cavalcante Siebert"
      ],
      "abstract": "Heralding the advent of autonomous vehicles and mobile robots that interact with humans, responsibility in spatial interaction is burgeoning as a research topic. Even though metrics of responsibility tailored to spatial interactions have been proposed, they are mostly focused on the responsibility of individual agents. Metrics of causal responsibility focusing on individuals fail in cases of causal overdeterminism - when many actors simultaneously cause an outcome. To fill the gaps in causal responsibility left by individual-focused metrics, we formulate a metric for the causal responsibility of groups. To identify assertive agents that are causally responsible for the trajectory of an affected agent, we further formalise the types of assertive influences and propose a tiering algorithm for systematically identifying assertive agents. Finally, we use scenario-based simulations to illustrate the benefits of considering groups and how the emergence of group effects vary with interaction dynamics and the proximity of agents.",
      "short_abstract": "Heralding the advent of autonomous vehicles and mobile robots that interact with humans, responsibility in spatial interaction is burgeoning as a research topic. Even though metrics of responsibility tailored to spatial interactions have been proposed, they...",
      "published": "2026-02-25",
      "updated": "2026-02-25",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 0,
        "margin": 0,
        "evidence": [
          "no strong primary-topic signal"
        ]
      },
      "recency": "archive",
      "age_days": 175,
      "arxiv_url": "https://arxiv.org/abs/2602.22041",
      "pdf_url": "https://arxiv.org/pdf/2602.22041.pdf"
    },
    {
      "id": "2604.09578",
      "anchor": "paper-2604-09578",
      "title": "Explainable Planning for Hybrid Systems",
      "authors": [
        "Mir Md Sajid Sarwar"
      ],
      "abstract": "The recent advancement in artificial intelligence (AI) technologies facilitates a paradigm shift toward automation. Autonomous systems are fully or partially replacing manually crafted ones. At the core of these systems is automated planning. With the advent of powerful planners, automated planning is now applied to many complex and safety-critical domains, including smart energy grids, self-driving cars, warehouse automation, urban and air traffic control, search and rescue operations, surveillance, robotics, and healthcare. There is a growing need to generate explanations of AI-based systems, which is one of the major challenges the planning community faces today. The thesis presents a comprehensive study on explainable artificial intelligence planning (XAIP) for hybrid systems that capture a representation of real-world problems closely.",
      "short_abstract": "The recent advancement in artificial intelligence (AI) technologies facilitates a paradigm shift toward automation. Autonomous systems are fully or partially replacing manually crafted ones. At the core of these systems is automated planning. With the advent...",
      "published": "2026-02-24",
      "updated": "2026-02-24",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Explainability"
      ],
      "classification": {
        "confidence": "low",
        "score": 12,
        "margin": 0,
        "evidence": [
          "Human Factors & Policy: title: explainability",
          "Human Factors & Policy: abstract: explainability",
          "Planning & Decision-Making: title: planning",
          "Planning & Decision-Making: abstract: planning"
        ]
      },
      "recency": "archive",
      "age_days": 176,
      "arxiv_url": "https://arxiv.org/abs/2604.09578",
      "pdf_url": "https://arxiv.org/pdf/2604.09578.pdf"
    },
    {
      "id": "2602.20963",
      "anchor": "paper-2602-20963",
      "title": "A Robotic Testing Platform for Pipelined Discovery of Resilient Soft Actuators",
      "authors": [
        "Ang Li",
        "Alexander Yin",
        "Alexander White",
        "Sahib Sandhu",
        "Matthew Francoeur",
        "Victor Jimenez-Santiago",
        "Van Remenar",
        "Codrin Tugui",
        "Mihai Duduta"
      ],
      "abstract": "Short lifetime under high electrical fields hinders the widespread robotic application of linear dielectric elastomer actuators (DEAs). Systematic scanning is difficult due to time-consuming per-sample testing and the high-dimensional parameter space affecting performance. To address this, we propose an optimization pipeline enabled by a novel testing robot capable of scanning DEA lifetime. The robot integrates electro-mechanical property measurement, programmable voltage input, and multi-channel testing capacity. Using it, we scanned the lifetime of Elastosil-based linear actuators across parameters including input voltage magnitude, frequency, electrode material concentration, and electrical connection filler. The optimal parameter combinations improved operational lifetime under boundary operating conditions by up to 100% and were subsequently scaled up to achieve higher force and displacement output. The final product demonstrated resilience on a modular, scalable quadruped walking robot with payload carrying capacity (>100% of its untethered body weight, and >700% of combined actuator weight). This work is the first to introduce a self-driving lab approach into robotic actuator design.",
      "short_abstract": "Short lifetime under high electrical fields hinders the widespread robotic application of linear dielectric elastomer actuators (DEAs). Systematic scanning is difficult due to time-consuming per-sample testing and the high-dimensional parameter space...",
      "published": "2026-02-24",
      "updated": "2026-02-24",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 12,
        "evidence": [
          "title: validation or testing",
          "abstract: validation or testing"
        ]
      },
      "recency": "archive",
      "age_days": 176,
      "arxiv_url": "https://arxiv.org/abs/2602.20963",
      "pdf_url": "https://arxiv.org/pdf/2602.20963.pdf"
    },
    {
      "id": "2602.20669",
      "anchor": "paper-2602-20669",
      "title": "Integrating Domain-Specialized Language Models with AI Measurement Tools for Deterministic Atomic-Resolution Experimentation",
      "authors": [
        "Zhuo Diao",
        "Kouma Matsumoto",
        "Linfeng Hou",
        "Masahiro Ohara",
        "Hayato Yamashita",
        "Masayuki Abe"
      ],
      "abstract": "Self-driving laboratories based on large language models promise to transform scientific discovery through general experimental automation. However, realizing this vision on precision platforms remains challenging, requiring deterministic execution and effective domain adaptation under strict physical constraints. We address these requirements through a framework that specializes in small language models for autonomous control of scanning probe microscopy, coordinating task-specific models with AI-driven measurement tools. We demonstrate real-time, atomic-resolution SPM experiments at room temperature, achieving instruction-level control and multi-step experimental planning. Fine-tuning reduces perplexity from 1.44 to 1.20 and improves reliability, with the adapted model reaching 99.3% and 95.2% command accuracy, outperforming OpenAI o4-mini on domain-specific tasks. This architecture achieves lower computational cost while maintaining deterministic execution and enabling deployment on consumer-grade hardware. This work bridges probabilistic language models with deterministic experimental control through a modular, domain-specialized architecture, providing a generalizable pathway toward scalable and trustworthy self-driving laboratories across diverse scientific platforms.",
      "short_abstract": "Self-driving laboratories based on large language models promise to transform scientific discovery through general experimental automation. However, realizing this vision on precision platforms remains challenging, requiring deterministic execution and...",
      "published": "2026-02-24",
      "updated": "2026-02-24",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [],
      "classification": {
        "confidence": "low",
        "score": 6,
        "margin": 3,
        "evidence": [
          "abstract: deployment or real-time",
          "abstract: hardware"
        ]
      },
      "recency": "archive",
      "age_days": 176,
      "arxiv_url": "https://arxiv.org/abs/2602.20669",
      "pdf_url": "https://arxiv.org/pdf/2602.20669.pdf"
    },
    {
      "id": "2602.20432",
      "anchor": "paper-2602-20432",
      "title": "Autonomous epitaxial atomic-layer synthesis via real-time computer vision of electron diffraction",
      "authors": [
        "Haotong Liang",
        "Yunlong Sun",
        "Ryan Paxson",
        "Chih-Yu Lee",
        "Alex T. Hall",
        "Zoey Warecki",
        "John Cumings",
        "Hideomi Koinuma",
        "Aaron Gilad Kusne",
        "Mikk Lippmaa",
        "Ichiro Takeuchi"
      ],
      "abstract": "Autonomous science platforms which make decisions on the fly are fundamentally changing the outlook for materials development. AI-driven schemes can effectively reduce the total number of iterations needed to arrive at the best stoichiometry for desired properties or optimum synthesis parameters by significant margins. Here, we demonstrate real-time closed-loop autonomous navigation of a multi-dimensional synthesis parameter space for fabricating phase-pure epitaxial films of a metastable functional oxide phase using pulsed laser deposition. Sequential growth iterations in search of the optimized recipe to stabilize the desired crystal phase were performed using frame-by-frame quantitative computer vision of electron diffraction images at the unit-cell level. Our scheme regularly resulted in > 30-fold reduction in the number of experiments compared to comprehensive parameter-space mapping. The real-time workflow developed here can be readily extended to other thin film synthesis platforms opening the door for self-driving atomic-level materials design as well as autonomous semiconductor manufacturing.",
      "short_abstract": "Autonomous science platforms which make decisions on the fly are fundamentally changing the outlook for materials development. AI-driven schemes can effectively reduce the total number of iterations needed to arrive at the best stoichiometry for desired...",
      "published": "2026-02-23",
      "updated": "2026-02-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "medium",
        "score": 12,
        "margin": 8,
        "evidence": [
          "title: deployment or real-time",
          "abstract: deployment or real-time"
        ]
      },
      "recency": "archive",
      "age_days": 177,
      "arxiv_url": "https://arxiv.org/abs/2602.20432",
      "pdf_url": "https://arxiv.org/pdf/2602.20432.pdf"
    },
    {
      "id": "2602.20521",
      "anchor": "paper-2602-20521",
      "title": "Towards Secure and Efficient DNN Accelerators via Hardware-Software Co-Design",
      "authors": [
        "Wei Xuan",
        "Zihao Xuan",
        "Rongliang Fu",
        "Ning Lin",
        "Kwunhang Wong",
        "Zikang Yuan",
        "Lang Feng",
        "Zhongrui Wang",
        "Tsung-Yi Ho",
        "Yuzhong Jiao",
        "Luhong Liang"
      ],
      "abstract": "The rapid deployment of deep neural network (DNN) accelerators in safety-critical domains such as autonomous vehicles, healthcare systems, and financial infrastructure necessitates robust mechanisms to safeguard data confidentiality and computational integrity. Existing security solutions for DNN accelerators, however, suffer from excessive hardware resource demands and frequent off-chip memory access overheads, which degrade performance and scalability. To address these challenges, this paper presents a secure and efficient memory protection framework for DNN accelerators with minimal overhead. First, we propose a bandwidth-aware cryptographic scheme that adapts encryption granularity based on memory traffic patterns, striking a balance between security and resource efficiency. Second, we observe that both the overlapping regions in the intra-layer tiling's sliding window pattern and those resulting from inter-layer tiling strategy discrepancies introduce substantial redundant memory accesses and repeated computational overhead in cryptography. Third, we introduce a multi-level authentication mechanism that effectively eliminates unnecessary off-chip memory accesses, enhancing performance and energy efficiency. Experimental results show that this work decreases performance overhead by over 12% and achieves 87% energy efficiency improvement for both server and edge neural processing units (NPUs), while ensuring robust scalability.",
      "short_abstract": "The rapid deployment of deep neural network (DNN) accelerators in safety-critical domains such as autonomous vehicles, healthcare systems, and financial infrastructure necessitates robust mechanisms to safeguard data confidentiality and computational...",
      "published": "2026-02-23",
      "updated": "2026-02-23",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Hardware / Real-Time"
      ],
      "classification": {
        "confidence": "high",
        "score": 19,
        "margin": 8,
        "evidence": [
          "abstract: deployment or real-time",
          "title: hardware",
          "abstract: hardware",
          "abstract: communications",
          "abstract: systems constraints"
        ]
      },
      "recency": "archive",
      "age_days": 177,
      "arxiv_url": "https://arxiv.org/abs/2602.20521",
      "pdf_url": "https://arxiv.org/pdf/2602.20521.pdf"
    },
    {
      "id": "2602.19596",
      "anchor": "paper-2602-19596",
      "title": "Learning Mutual View Information Graph for Adaptive Adversarial Collaborative Perception",
      "authors": [
        "Yihang Tao",
        "Senkang Hu",
        "Haonan An",
        "Zhengru Fang",
        "Hangcheng Cao",
        "Yuguang Fang"
      ],
      "abstract": "Collaborative perception (CP) enables data sharing among connected and autonomous vehicles (CAVs) to enhance driving safety. However, CP systems are vulnerable to adversarial attacks where malicious agents forge false objects via feature-level perturbations. Current defensive systems use threshold-based consensus verification by comparing collaborative and ego detection results. Yet, these defenses remain vulnerable to more sophisticated attack strategies that could exploit two critical weaknesses: (i) lack of robustness against attacks with systematic timing and target region optimization, and (ii) inadvertent disclosure of vulnerability knowledge through implicit confidence information in shared collaboration data. In this paper, we propose MVIG attack, a novel adaptive adversarial CP framework learning to capture vulnerability knowledge disclosed by different defensive CP systems from a unified mutual view information graph (MVIG) representation. Our approach combines MVIG representation with temporal graph learning to generate evolving fabrication risk maps and employs entropy-aware vulnerability search to optimize attack location, timing and persistence, enabling adaptive attacks with generalizability across various defensive configurations. Extensive evaluations on OPV2V and Adv-OPV2V datasets demonstrate that MVIG attack reduces defense success rates by up to 62\\% against state-of-the-art defenses while achieving 47\\% lower detection for persistent attacks at 29.9 FPS, exposing critical security gaps in CP systems. Code will be released at https://github.com/yihangtao/MVIG.git",
      "short_abstract": "Collaborative perception (CP) enables data sharing among connected and autonomous vehicles (CAVs) to enhance driving safety. However, CP systems are vulnerable to adversarial attacks where malicious agents forge false objects via feature-level perturbations....",
      "published": "2026-02-23",
      "updated": "2026-02-23",
      "category": "Safety, Security & Verification",
      "primary_category": "Safety, Security & Verification",
      "tags": [
        "Cooperative / V2X"
      ],
      "classification": {
        "confidence": "high",
        "score": 32,
        "margin": 16,
        "evidence": [
          "abstract: safety",
          "abstract: security",
          "abstract: verification",
          "title: attack or threat",
          "abstract: attack or threat",
          "abstract: robustness"
        ]
      },
      "recency": "archive",
      "age_days": 177,
      "arxiv_url": "https://arxiv.org/abs/2602.19596",
      "pdf_url": "https://arxiv.org/pdf/2602.19596.pdf"
    },
    {
      "id": "2602.18798",
      "anchor": "paper-2602-18798",
      "title": "See What I See: An Attention-Guiding eHMI Approach for Autonomous Vehicles",
      "authors": [
        "Jialong Li",
        "Zhenyu Mao",
        "Zhiyao Wang",
        "Yijun Lu",
        "Shogo Morita",
        "Nianyu Li",
        "Kenji Tei"
      ],
      "abstract": "As autonomous vehicles are gradually being deployed in the real world, external Human-Machine Interfaces (eHMIs) are expected to serve as a critical solution for enhancing vehicle-pedestrian communication. However, existing eHMI designs typically focus solely on the ego vehicle's status, which can inadvertently capture pedestrians' attention or encourage misguided reliance on the AV's signals, leading them to neglect scanning for other surrounding hazards. To address this, we propose the Attention-Guiding eHMI (AGeHMI), a projection-based visualization that employs directional cues and risk-based color coding to actively guide pedestrians' attention toward potential environmental dangers. Evaluation through a virtual reality user study (N = 20) suggests that AGeHMI effectively influences participants' visual attention distribution and significantly reduces potential collision risks with surrounding vehicles, while simultaneously improving subjective confidence and reducing cognitive workload.",
      "short_abstract": "As autonomous vehicles are gradually being deployed in the real world, external Human-Machine Interfaces (eHMIs) are expected to serve as a critical solution for enhancing vehicle-pedestrian communication. However, existing eHMI designs typically focus solely...",
      "published": "2026-02-21",
      "updated": "2026-02-21",
      "category": "Cross-cutting / Other",
      "primary_category": "Cross-cutting / Other",
      "tags": [
        "Human Interaction"
      ],
      "classification": {
        "confidence": "low",
        "score": 5,
        "margin": 3,
        "evidence": [
          "Human Factors & Policy: abstract: human-machine interaction",
          "Systems, Deployment & Connectivity: abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 179,
      "arxiv_url": "https://arxiv.org/abs/2602.18798",
      "pdf_url": "https://arxiv.org/pdf/2602.18798.pdf"
    },
    {
      "id": "2602.18740",
      "anchor": "paper-2602-18740",
      "title": "HONEST-CAV: Hierarchical Optimization of Network Signals and Trajectories for Connected and Automated Vehicles with Multi-Agent Reinforcement Learning",
      "authors": [
        "Ziyan Zhang",
        "Changxin Wan",
        "Peng Hao",
        "Kanok Boriboonsomsin",
        "Matthew J. Barth",
        "Yongkang Liu",
        "Seyhan Ucar",
        "Guoyuan Wu"
      ],
      "abstract": "This study presents a hierarchical, network-level traffic flow control framework for mixed traffic consisting of Human-driven Vehicles (HVs), Connected and Automated Vehicles (CAVs). The framework jointly optimizes vehicle-level eco-driving behaviors and intersection-level traffic signal control to enhance overall network efficiency and decrease energy consumption. A decentralized Multi-Agent Reinforcement Learning (MARL) approach by Value Decomposition Network (VDN) manages cycle-based traffic signal control (TSC) at intersections, while an innovative Signal Phase and Timing (SPaT) prediction method integrates a Machine Learning-based Trajectory Planning Algorithm (MLTPA) to guide CAVs in executing Eco-Approach and Departure (EAD) maneuvers. The framework is evaluated across varying CAV proportions and powertrain types to assess its effects on mobility and energy performance. Experimental results conducted in a 4*4 real-world network demonstrate that the MARL-based TSC method outperforms the baseline model (i.e., Webster method) in speed, fuel consumption, and idling time. In addition, with MLTPA, HONEST-CAV benefits the traffic system further in energy consumption and idling time. With a 60% CAV proportion, vehicle average speed, fuel consumption, and idling time can be improved/saved by 7.67%, 10.23%, and 45.83% compared with the baseline. Furthermore, discussions on CAV proportions and powertrain types are conducted to quantify the performance of the proposed method with the impact of automation and electrification.",
      "short_abstract": "This study presents a hierarchical, network-level traffic flow control framework for mixed traffic consisting of Human-driven Vehicles (HVs), Connected and Automated Vehicles (CAVs). The framework jointly optimizes vehicle-level eco-driving behaviors and...",
      "published": "2026-02-21",
      "updated": "2026-02-21",
      "category": "Systems, Deployment & Connectivity",
      "primary_category": "Systems, Deployment & Connectivity",
      "tags": [
        "Reinforcement Learning"
      ],
      "classification": {
        "confidence": "high",
        "score": 24,
        "margin": 16,
        "evidence": [
          "title: connected vehicles",
          "abstract: connected vehicles",
          "title: communications",
          "abstract: communications"
        ]
      },
      "recency": "archive",
      "age_days": 179,
      "arxiv_url": "https://arxiv.org/abs/2602.18740",
      "pdf_url": "https://arxiv.org/pdf/2602.18740.pdf"
    }
  ]
}
