SWAG: Splatting in the Wild images with Appearance-conditioned Gaussians

ECCV 2024
1Noah’s Ark, Huawei Paris Research Center
Paper image

Abstract

Implicit neural representation methods have shown impressive advancements in learning 3D scenes from unstructured in-the-wild photo collections but are still limited by the large computational cost of volumetric rendering. More recently, 3D Gaussian Splatting emerged as a much faster alternative with superior rendering quality and training efficiency, especially for small-scale and object-centric scenarios. Nevertheless, this technique suffers from poor performance on unstructured in-the-wild data. To tackle this, we extend over 3D Gaussian Splatting to handle unstructured image collections. We achieve this by modeling appearance to seize photometric variations in the rendered images. Additionally, we introduce a new mechanism to train transient Gaussians to handle the presence of scene occluders in an unsupervised manner. Experiments on diverse photo collection scenes and multi-pass acquisition of outdoor landmarks show the effectiveness of our method over prior works achieving state-of-the-art results with improved efficiency.

Novel View synthesis

Using only unstructured image collections, SWAG is able to generate high-quality renderings of new views.

Interpolations in Appearance space

The model is capable of rendering new views of the scene by adjusting the camera's viewpoint and modifying the appearance embedding, which allows us to alter the scene’s visual style.

Transient objects removal

Results on NeRF-OSR dataset

BibTeX

@misc{dahmani2024swagsplattingwildimages,
      title={SWAG: Splatting in the Wild images with Appearance-conditioned Gaussians},
      author={Hiba Dahmani and Moussab Bennehar and Nathan Piasco and Luis Roldao and Dzmitry Tsishkou},
      year={2024},
      eprint={2403.10427},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2403.10427},
}