SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation

Published in ECCV 2026, 2026

# SEM-ROVER ### Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation **ECCV 2026** 🌐 Project Page 📄 Paper

Replace sem-rover-preview.jpg with a screenshot from your project website. It will serve as the cover image for the publication.

Abstract

Generating realistic and geometrically consistent driving environments at scale remains a challenging problem. SEM-ROVER introduces a semantic voxel-guided diffusion framework built upon Σ-Voxfield, a structured 3D representation that enables efficient generation of large outdoor scenes while preserving multi-view consistency.

The method progressively expands scenes through semantic-aware outpainting and combines voxel-based generation with deferred rendering to synthesize realistic novel views without scene-specific optimization.


Key Contributions

  • 🚗 Large-scale 3D driving scene generation
  • 🧊 Novel Σ-Voxfield scene representation
  • 🌍 Semantic-guided diffusion model
  • 🔄 Progressive spatial outpainting
  • 📷 Multi-view consistent rendering
  • ⚡ Efficient deferred rendering pipeline

Resources

ResourceLink
🌐 Projecthttps://dahmanihiba.github.io/SEM-ROVER/
📄 Paperhttps://arxiv.org/abs/2604.06113

If you use this work in your research, please consider citing the paper.