SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation
Published in ECCV 2026, 2026
Replace
sem-rover-preview.jpgwith a screenshot from your project website. It will serve as the cover image for the publication.
Abstract
Generating realistic and geometrically consistent driving environments at scale remains a challenging problem. SEM-ROVER introduces a semantic voxel-guided diffusion framework built upon Σ-Voxfield, a structured 3D representation that enables efficient generation of large outdoor scenes while preserving multi-view consistency.
The method progressively expands scenes through semantic-aware outpainting and combines voxel-based generation with deferred rendering to synthesize realistic novel views without scene-specific optimization.
Key Contributions
- 🚗 Large-scale 3D driving scene generation
- 🧊 Novel Σ-Voxfield scene representation
- 🌍 Semantic-guided diffusion model
- 🔄 Progressive spatial outpainting
- 📷 Multi-view consistent rendering
- ⚡ Efficient deferred rendering pipeline
Resources
| Resource | Link |
|---|---|
| 🌐 Project | https://dahmanihiba.github.io/SEM-ROVER/ |
| 📄 Paper | https://arxiv.org/abs/2604.06113 |
If you use this work in your research, please consider citing the paper.
