DSP-SLAM

Wang (UCL) 2021 · Paper

One-line summary — Augments ORB-SLAM2 with category-level DeepSDF shape priors to reconstruct complete, dense object models online from monocular, stereo, or stereo+LiDAR input.

Problem

Object-level SLAM systems without priors (e.g. Fusion++) reconstruct each object only as well as the camera happened to observe it — partial viewpoints yield partial, low-quality models — while instance-database systems (SLAM++) require every object to be pre-scanned. Learned shape priors such as DeepSDF can complete unseen object parts from sparse observations, but integrating a deep implicit shape model into a real-time SLAM loop — with online pose tracking, sparse and partial data, and a joint map — was an open problem: prior-based reconstructors like FroDO were slow batch methods, and NodeSLAM needed dense depth. DSP-SLAM builds a joint map of dense object models for the foreground and sparse landmark points for the background, closing that gap.

Method & architecture

ORB-SLAM2 (monocular or stereo) provides camera tracking, keyframing, and the sparse 3D point cloud. At each keyframe, Mask R-CNN masks plus a 3D detector give each object instance I={B,M,D,Tco,0}I=\{\mathcal{B},\mathcal{M},\mathcal{D},\mathbf{T}_{co,0}\} — 2D box, mask, sparse 3D point observations (SLAM points, or as few as 50 LiDAR points), and an initial pose from a LiDAR/image 3D detector or PCA on the object points. Each object is a DeepSDF latent code zR64\mathbf{z}\in\mathbb{R}^{64} with decoder s=G(x,z)s=G(\mathbf{x},\mathbf{z}) and 7-DoF pose TcoSim(3)\mathbf{T}_{co}\in \mathbf{Sim}(3). Shape and pose are estimated by minimizing two energies. A surface-consistency term drives observed back-projected points onto the zero level set:

Esurf=1ΩsuΩsG2(Tocπ1 ⁣(u,D),z)E_{surf}=\frac{1}{\lvert\mathbf{\Omega}_{s}\rvert}\sum_{\mathbf{u}\in\mathbf{\Omega}_{s}}G^{2}\big(\mathbf{T}_{oc}\,\pi^{-1}\!\left(\mathbf{u},\mathcal{D}\right),\,\mathbf{z}\big)

Alone this lets shapes grow oversized under partial observation, so a differentiable SDF renderer adds silhouette-aware depth supervision: along each pixel ray, MM sampled depths get occupancy oio_i from the predicted SDF (piecewise-linear cutoff σ=0.01\sigma=0.01), ray-termination event probabilities ϕi=oij=1i1(1oj)\phi_{i}=o_{i}\prod_{j=1}^{i-1}(1-o_{j}), and an expected rendered depth d^u=i=1M+1ϕidi\hat{d}_{\mathbf{u}}=\sum_{i=1}^{M+1}\phi_{i}d_{i}, giving

Erend=1ΩruΩr(dud^u)2E_{rend}=\frac{1}{\lvert\mathbf{\Omega}_{r}\rvert}\sum_{\mathbf{u}\in\mathbf{\Omega}_{r}}(d_{\mathbf{u}}-\hat{d}_{\mathbf{u}})^{2}

where Ωr\mathbf{\Omega}_{r} adds pixels inside the box but outside the mask, assigned background depth 1.1dmax1.1\,d_{max} — penalizing shapes that leak outside the silhouette. The total energy E=λsEsurf+λrErend+λcz2E=\lambda_{s}E_{surf}+\lambda_{r}E_{rend}+\lambda_{c}\lVert\mathbf{z}\rVert^{2} (λs=100\lambda_s=100, λr=2.5\lambda_r=2.5, λc=0.25\lambda_c=0.25) is minimized by Gauss-Newton with analytical Jacobians through the network from z=0\mathbf{z}=\mathbf{0} — roughly an order of magnitude faster per iteration than first-order descent (20 ms vs 183 ms with both terms) and needing ~10 instead of 50 iterations. Reconstructed objects then enter a joint factor graph over camera poses CC, object poses OO, and points PP:

C,O,P=argmin{C,O,P}i,jeco(Twci,Twoj)Σi,j+i,kecp(Twci,wpk)Σi,kC^{*},O^{*},P^{*}=\mathop{\arg\min}_{\{C,O,P\}}\sum_{i,j}\big\lVert\mathbf{e}_{co}(\mathbf{T}_{wc_{i}},\mathbf{T}_{wo_{j}})\big\rVert_{\Sigma_{i,j}}+\sum_{i,k}\big\lVert\mathbf{e}_{cp}(\mathbf{T}_{wc_{i}},{}^{w}\mathbf{p}_{k})\big\rVert_{\Sigma_{i,k}}

with camera-object residual eco=log(Tco1Twc1Two)\mathbf{e}_{co}=\log(\mathbf{T}^{-1}_{co}\mathbf{T}^{-1}_{wc}\mathbf{T}_{wo}) and the standard ORB-SLAM2 reprojection residual ecp\mathbf{e}_{cp}, solved by Levenberg-Marquardt in g2o — objects act as additional landmarks. Data association matches detections to map objects by 3D-box distance (LiDAR) or shared feature matches (mono/stereo); re-observed objects get pose-only updates.

Results

On KITTI3D (7481 frames, single image + LiDAR, same DeepSDF prior and initialization as the baseline), DSP-SLAM beats auto-labelling on nearly all object pose metrics: BEV [email protected] 83.31 vs 80.70 (Easy) and 75.28 vs 63.36 (Moderate); nuScenes [email protected] 88.01 vs 86.52 (E) and 76.15 vs 64.44 (M) — with visibly better shapes (sedans no longer reconstructed “beetle”-shaped). On the KITTI odometry benchmark, stereo+LiDAR DSP-SLAM averages 0.70 % translation / 0.22 deg per 100 m — improving on its ORB-SLAM2 backbone (0.72/0.22) on object-rich sequences 03, 05, 06, 08, and matching SuMa++ (0.70/0.29) while using only a few hundred LiDAR points per frame; reducing to 50 points per object barely changes accuracy (0.72/0.22). Stereo-only runs at 0.75/0.25, and at 5 Hz with per-keyframe BA matches ORB-SLAM2 (0.72/0.22). The full system runs at ~10 fps and produces complete object reconstructions from monocular input on Freiburg Cars and Redwood-OS chairs.

Why it matters for SLAM

DSP-SLAM was the first SLAM system to integrate learned implicit shape priors for online object reconstruction, updating the SLAM++ vision — maps made of objects, not raw geometry — for the deep learning era: instead of a database of scanned CAD models, a latent shape space covers a whole category. Its second-order optimization through a neural SDF showed that deep shape fitting can live inside a real-time loop, and it is a key stepping stone between classical object-level SLAM (SLAM++, Fusion++, NodeSLAM) and later neural-field object mapping (vMAP) — a good template whenever you need semantically meaningful, complete object models rather than surfel soup.

Hands-on