Generative Semantic Scene Completion

arXiv preprint · under review

Shi Chen1,2  ·  Weifeng Ge1,*

1College of Computer Science and Artificial Intelligence, Fudan University  ·  2Robotics Institute, Carnegie Mellon University

*Corresponding author

The semantic scene completion challenge

Semantic scene completion: a sparse LiDAR sweep occupying about 1% of the volume, and the dense semantically labelled scene that must be predicted from it, with zoomed insets contrasting sparse geometry and lost semantics against dense structure and rich semantics.
One LiDAR sweep returns about one voxel in a hundred. Semantic scene completion has to recover the other ninety-nine — geometry and semantics together.

Abstract

Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000×. We recast SSC as generative semantic scene completion (GSSC): a single discrete-diffusion formulation in three roles. First, paired sparse–dense scene synthesis (PS3) generates matched sparse LiDAR observations with their dense semantic completions, addressing the long tail at its source and yielding the PS3-SemanticKITTI corpus we train on alongside SemanticKITTI. Second, semantic-guided generative scene completion (SGSC) generates the scene from noise with multinomial discrete diffusion, conditioned on the sparse scan through a bird’s-eye-view semantic map and a sparse 3D feature stream. Third, the same framework instead refines an existing completion in one flow-matching step: structured source discrete diffusion (S2D2). S2D2 improves the mIoU of SGSC’s own output and every external SSC base tested, without base retraining or test-time adaptation. On the strongest base, one step without test-time augmentation reaches 38.8% mIoU on the SemanticKITTI hidden test. To our knowledge that is the best causal, single-sweep, single-sample result on that leaderboard, +2.1 pp over the previous best published score under the same restriction. Four correction steps with eight-view test-time augmentation reach 39.2%, outside that restriction.

Results

Qualitative comparison on SemanticKITTI validation sequence 08: three frozen sources (JS3C-Net, SCPNet, and TALoS's adaptation of SCPNet), our refinement on the SCPNet base, and ground truth; a bicyclist scene above and a motorcyclist scene below.
One correction step, on a frozen base. Three frozen sources, our refinement of the SCPNet one, ground truth. Chips come from the N=4, +D4-TTA run, not the N=1 row the headline is indexed on.
Horizontal bar chart of SemanticKITTI hidden-test mIoU, nine bars. Longest first: double-dagger Ours + D4 ensemble (N=4, +D4 TTA) 39.2; SCPNet + S2D2 (N=1) 38.8; double-dagger TALoS (test-time adaptation) 37.9; SCPNet (published) 36.7; S3CNet 29.5; DiffSSC 27.4; JS3C-Net 23.8; SSA-SC 23.5; LMSCNet 17.6. The two bars that are ours, 38.8 and 39.2, are red; the other seven are grey published baselines. A leading double-dagger marks the two rows outside the headline predicate, our own ensemble row among them, and is explained under the axis as test-time adaptation or ensembling. The best bar inside the predicate is our 38.8, above SCPNet’s published 36.7.
38.8% mIoU at one step, no test-time augmentation: to our knowledge the best causal, single-sweep, single-sample result on this leaderboard, +2.1 pp over the previous best published score under that restriction. 39.2% — four steps with an eight-view D4 ensemble — is the entry Codabench displays, since the platform lists each team’s best score; it falls outside that restriction, as ‡ marks.

Interactive comparison

Switch scene and view; the camera holds, so the four views line up. The unlabelled input is painted one colour.

Scene
View

: frozen base · ours · Chips: N=4, +D4 TTA — outside the headline predicate

How it works

PS3 — paired sparse–dense synthesis

The PS3 offline data-augmentation pipeline: three cascaded coarse-to-fine multinomial diffusions screened by a Jensen-Shannon divergence filter, then a rare-class object bank and the HALO ray-tracer turning each synthetic scene into a sparse-dense pair.
Generate complete labelled scenes, then ray-trace them back into sparse sweeps — matched pairs the recorder never drove past.
Grouped log-scale bar chart of per-class voxel frequency, real only against real + synth-32K, with the synthetic pool's own frequency against real-only printed above the rarest classes: bicyclist 6239 times, motorcycle 1626 times, bicycle 15.4 times.
Per-class voxel frequency, real only against real + synth-32K. The printed factors are the synthetic pool’s own frequency against real-only — bicyclist 6,239× — while the pooled bars gain 3.2×–3,907×.

SGSC — completion from noise

The SGSC denoiser in its generation role: a multinomial discrete diffusion model conditioned on a bird's-eye-view semantic map and a sparse 3D feature stream, resolving categorical noise into a labelled scene.
The same denoiser run from pure categorical noise instead of from a base, conditioned on a bird’s-eye-view map and a sparse feature stream. No completion network underneath.

S2D2 — one-step refinement

S2D2 refinement: the velocity field carrying a frozen base prediction to ground truth in one Euler step on the per-voxel simplex, with an identity and error bound ruling out amplification, and a motorcyclist scene before and after.
Carry a frozen model’s output to ground truth in one step on the per-voxel simplex. No base retraining, no test-time adaptation.

Acknowledgements

This work was supported by the National Natural Science Foundation of China under Grant 624B1006 and by the Shanghai Science and Technology Committee under Grant 24511103900.

Evaluation uses the SemanticKITTI benchmark and its hidden test server. The comparison renders here — qualitative figure, gallery, 3D viewer — come from validation seq. 08. Nothing in the PS3 figure does: it shows a training-split frame, scenes synthesised after training on that split, and rare-class crops from it.

Those renders, and the point clouds the viewer loads, are voxelised and class-recoloured exports of SemanticKITTI ground-truth annotations — modified material, redistributed here. SemanticKITTI is © its authors under CC BY-NC-SA 4.0: credit the creators, non-commercial use only, share-alike. Those terms travel with these exports, which are offered under that licence and not this page’s own. semantic-kitti.org asks that both the SemanticKITTI paper (Behley et al.) and the original KITTI Vision Benchmark (Geiger et al.) be cited; BibTeX for both is below.

BibTeX

@misc{chen2026gssc,
  title         = {Generative Semantic Scene Completion},
  author        = {Chen, Shi and Ge, Weifeng},
  year          = {2026},
  eprint        = {2608.26737},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2608.26737},
  url           = {https://arxiv.org/abs/2608.26737},
  note          = {Under review}
}

The SemanticKITTI material shown on this page carries its own citation requirement. Both of these must accompany any reuse of it:

@inproceedings{behley2019semantickitti,
  title     = {{SemanticKITTI}: A Dataset for Semantic Scene
               Understanding of {LiDAR} Sequences},
  author    = {Behley, Jens and Garbade, Martin and Milioto, Andres
               and Quenzel, Jan and Behnke, Sven and Stachniss, Cyrill
               and Gall, Juergen},
  booktitle = {Proc. IEEE/CVF Int. Conf. on Computer Vision (ICCV)},
  pages     = {9297--9307},
  year      = {2019}
}

@inproceedings{geiger2012kitti,
  title     = {Are We Ready for Autonomous Driving? The {KITTI}
               Vision Benchmark Suite},
  author    = {Geiger, Andreas and Lenz, Philip and Urtasun, Raquel},
  booktitle = {Proc. IEEE Conf. on Computer Vision and Pattern
               Recognition (CVPR)},
  pages     = {3354--3361},
  year      = {2012}
}