Level 05 / 11

Applying Deep Learning

Learned frontends, differentiable backends, end-to-end systems, foundation-model & neural SLAM, scene understanding

Key Concepts

Level 5 is organized into five pillars: A. Frontend — learned perception components replacing hand-crafted modules B. Backend — learned/certifiable optimization replacing classical solvers C. Systems — end-to-end deep VO/SLAM pipelines D. Scene Understanding — semantic, language, and relational reasoning on SLAM maps E. Foundation-Model & Neural SLAM — pointmap transformers, NeRF- and 3DGS-based dense SLAM systems

A. Deep Frontend — Perception

Feature Detection & Matching

System Author/Year Key Concepts
NetVLAD Arandjelović 2016 VLAD, place recognition
SuperPoint DeTone 2017 Homographic Adaptation, Self-supervised, VGG encoder + detector/descriptor heads
HardNet Mishchuk 2017 Learned local descriptor
R2D2 Revaud 2019 Repeatable + Reliable detector/descriptor, explicit repeatability/reliability maps
KeyNet Barroso-Laguna 2019 Learned keypoint detector
HF-Net Sarlin 2019 Global feature, Local feature, Visual localization
SuperGlue Sarlin 2020 Self/Cross-attention GNN, Sinkhorn optimal assignment, dustbin for outliers
DISK Tyszkiewicz 2020 Policy gradient (RL) training, match success/failure as reward
Patch NetVLAD Hausler 2021 Multi-scale patch-level VLAD
LoFTR Sun 2021 Detector-free, Transformer coarse-to-fine dense matching
LightGlue Lindenberger 2023 Adaptive depth/width, 5-10× faster than SuperGlue
XFeat Potje 2024 0.3M params, 1400 FPS (RTX 4090), 64-dim descriptor, embedded-friendly
RoMa Edstedt 2024 DINOv2 foundation feature + coarse-to-fine dense matching
DeDoDe Edstedt 2024 Joint detect-and-describe in one stage
RoMa v2 Edstedt 2025 Harder-better-faster-denser dense feature matching

Depth Estimation

System Author/Year Key Concepts
MonoDepth Godard 2016 Left-Right photometric consistency, self-supervised
MiDaS Ranftl 2020 Multi-dataset mixing, scale-and-shift invariant loss, relative depth
DPT Ranftl 2021 Dense Prediction Transformer (ViT backbone), global context
ZoeDepth Bhat 2023 Zero-shot metric depth, Metric Bins Module
Metric3D Yin 2023 Camera intrinsic-conditioned metric depth, Canonical Camera Space
Depth Anything Yang 2024 62M images, foundation model for monocular depth
Depth Anything V2 Yang 2024 Improved with synthetic data, better edge preservation
Depth Anything 3 Lin 2025 Any-view geometry from arbitrary inputs, depth-ray prediction target, single plain transformer (DINOv2), teacher-student training
Marigold Ke 2024 Stable Diffusion for depth, fine detail, uncertainty via sampling
Align3R Lu 2025 Video temporal consistency, DUSt3R-based, CVPR 2025 Highlight
Masked Depth Modeling (LingBot-Depth) Tan 2026 Fixes RGB-D failures on glass/mirrors/metal

Optical Flow & Scene Flow

System Author/Year Key Concepts
FlowNet Dosovitskiy 2015 First end-to-end deep optical flow (SimpleNet / CorrNet)
FlowNet 2.0 Ilg 2017 Stacked networks, classical-level accuracy
PWC-Net Sun 2018 Pyramid-Warping-Cost volume, coarse-to-fine, 8.4M params
FlowNet3D Liu 2019 Point cloud scene flow, PointNet++ based
RAFT Teed 2020 All-Pairs Correlation + iterative ConvGRU update, ECCV Best Paper
RAFT-3D Teed 2021 Scene flow (3D motion) from RAFT
FlowFormer Huang 2022 Transformer on cost volume tokens, global context
SEA-RAFT Wang 2024 Efficient RAFT variant for real-time

Camera Pose Regression & Relocalization

System Author/Year Key Concepts
PoseNet Kendall 2015 CNN-based 6-DoF pose regression (APR), GoogLeNet backbone
DSAC Brachmann 2017 Differentiable RANSAC, Scene Coordinate Regression (SCR)
DSAC++ Brachmann 2018 Self-supervision, RGB-D support
CNN Pose Regression Limitations Sattler 2019 Pose regression ≈ image retrieval performance
LM-Reloc von Stumberg 2020 Deep direct relocalization
DSAC* Brachmann 2021 Visual relocalization from RGB/RGB-D, improved learning stability (TPAMI)
ACE Brachmann 2023 Accelerated Coordinate Encoding, 5-min training per scene
ACE Zero Brachmann 2024 Zero-shot SCR, no pre-built 3D map needed
ACE-G Bruns 2025 Generalizable SCR via query pretraining, new scenes without fine-tuning
ACE-SLAM Alzugaray 2025 Neural implicit real-time SLAM, network weights = map
hloc Sarlin 2019 Toolbox implementing HF-Net's hierarchical localization: coarse (NetVLAD) → fine (SuperGlue)

Object Detection & Segmentation for SLAM

System Author/Year Key Concepts
YOLO (v1→v11) Redmon 2016→2024 Real-time object detection, Ultralytics ecosystem
DETR Carion 2020 Transformer detection, anchor-free, no NMS
RT-DETR Zhao (Baidu) 2023 Real-time DETR, YOLO-speed + Transformer quality
RF-DETR Robinson 2025 Weight-sharing NAS over DETRs, accuracy-latency Pareto tuning, first real-time detector past 60 AP on COCO
SAM Kirillov 2023 Segment Anything, prompt-based, Foundation Model
SAM 2 Meta 2024 Video segmentation, Memory Attention, temporal consistency
SAM 3 Carion 2025 Promptable concept segmentation (noun-phrase / exemplar prompts), presence head, detector + memory-based video tracker
Grounding DINO Liu 2023 Text-prompted detection → SAM pipeline (Grounded SAM)
Open-YOLO 3D Boudjoghra 2024 2D open-vocab detection → 3D instance seg, 16× faster

B. Deep Backend — Optimization

Differentiable Bundle Adjustment

System Author/Year Key Concepts
BA-Net Tang 2019 FPN + differentiable LM layer, end-to-end SfM (ICLR)
DROID-SLAM Teed 2021 Dense optical flow + differentiable dense BA, all-pixels reprojection
DPVO Teed 2023 Patch-based DROID-SLAM, 30+ FPS real-time
Theseus Pineda (Meta) 2022 Differentiable nonlinear optimization library (PyTorch)
Lietorch Teed 2021 Lie group operations for PyTorch (SE(3)/SO(3))

C. End-to-End Deep VO / SLAM Systems

Self-supervised & Learned VO

System Author/Year Key Concepts
DeepVO Wang 2017 Supervised learning
SfM-Learner Zhou 2017 Unsupervised, deep depth + deep pose
DeMoN Ummenhofer 2017 Depth + Motion from two frames, encoder-decoder
UndeepVO Li 2018 Stereo self-supervised, absolute scale recovery
DeepTAM Zhou 2018 Deep tracking and mapping, cost volume based
DeepV2D Teed 2018 Iterative depth from video, differentiable geometry layers
Depth from Videos in the Wild Gordon 2019 Unconstrained video depth, learned camera intrinsics
Neural Ray Surfaces Vasiljevic 2020 Learned ray surface model, non-pinhole cameras
GradSLAM Murthy 2020 Differentiable SLAM framework (PyTorch, supports multiple SLAM backends)
DeepSLAM Li 2020 TrackingNet, MappingNet, LoopNet
MonoRec Wimbauer 2021 Self-supervised monocular 3D reconstruction, moving objects
TANDEM Koestler 2021 Real-time tracking + dense mapping via MVS depth, DSO-based

Learning-based SLAM Systems

System Author/Year Key Concepts
DROID-SLAM Teed 2021 Differentiable BA, dense optical flow, end-to-end learned
TartanVO Wang 2021 Generalizable visual odometry
DPV-SLAM Lipson 2024 DPVO + loop closure, full SLAM (ECCV 2024)
MAC-VO Qiu 2024 Learning-based VO, metric-aware
VoT Yugay 2025 Visual Odometry with Transformers (later retitled FVO)

Latent Representation SLAM

System Author/Year Key Concepts
CodeSLAM Bloesch 2018 Depth as 128-dim latent code, photometric BA on codes + poses
SceneCode Zhi 2019 Depth + semantic in single latent code, cross-modal constraints
DeepFactors Czarnowski 2020 Probabilistic depth codes + factor graph, GPU 30+ FPS
NodeSLAM Sucar 2020 Object-level DeepSDF codes, occupancy VAE per object
CodeMapping Matsuki 2021 Sparse SLAM + learned dense mapping, hybrid approach

Neural Rendering (reference)

NeRF/3DGS-based SLAM systems → see Pillar E below

System Author/Year Key Concepts
NeRF Mildenhall 2020 Neural Radiance Fields, novel view synthesis (foundational)
DIFIX3D+ Wu 2025 Single-step diffusion for 3D reconstruction artifact removal (post-processing)

D. Scene Understanding

Benchmarks & Foundations

System Author/Year Key Concepts
EFM3D Straub (Meta) 2024 Egocentric Foundation Model 3D benchmark, depth/surface/semantic from ego-video

3D Scene Graph

System Author/Year Key Concepts
Kimera / 3D Dynamic Scene Graph Rosinol 2020 Kimera-VIO, Kimera-Mesher, Kimera-PGMO, Kimera-Semantics, Kimera-DSG (stereo/mono visual-inertial pipeline)
Hydra Hughes (MIT SPARK) 2022 Real-time hierarchical Scene Graph (mesh→objects→places→rooms→buildings)
Hydra-Multi Chang 2023 Distributed multi-robot 3D Scene Graph
Clio Maggio (MIT SPARK) 2024 Open-set task-driven Scene Graph, CLIP embeddings per node
Khronos Schmid (MIT SPARK) 2024 Spatio-temporal Scene Graph, dynamic object history tracking
ConceptGraphs Gu 2023 Open-vocabulary 3D Scene Graph, SAM + CLIP + LLM relations

Semantic / Language-Grounded SLAM

System Author/Year Key Concepts
ConceptFusion Jatavallabhula (MIT) 2023 CLIP features fused into 3D map, open-vocabulary language queries
LERF Kerr 2023 Language Embedded Radiance Fields, DINO multi-scale, NeRF + CLIP
OpenScene Peng (ETH) 2023 Language features back-projected to 3D point clouds
SpatialLM Mao 2025 Point cloud → LLM, structured indoor modeling as Python scripts

Also see: LEGS, OpenGS-SLAM (Pillar E above); Open-YOLO 3D (Level 5 Object Detection)

E. Foundation-Model & Neural-Representation SLAM

Foundation-Model SLAM

System Author/Year Key Concepts
DUSt3R Wang 2024 Pointmap regression from image pairs, no calibration needed
MASt3R Leroy 2024 DUSt3R + local feature matching
MASt3R-SLAM Murai 2024 Real-time dense SLAM from MASt3R priors
VGGT Wang (Meta) 2025 Feed-forward inference of poses, depths, pointmaps, tracks from N views (CVPR 2025 Best Paper)
VGGT-SLAM Maggio 2025 Dense RGB SLAM optimized on the SL(4) manifold, VGGT frontend
VGGT-SLAM 2.0 Maggio 2026 Real-time dense feed-forward scene reconstruction
VGGT-Geo Qin 2026 Probabilistic geometric fusion of VGGT priors for dense indoor SLAM
IGGT Li 2025 Instance-grounded geometry transformer — unified 3D reconstruction + instance-level understanding
AMB3R Wang 2025 Accurate feed-forward metric-scale 3D reconstruction with backend, SfM/SLAM support
MASt3R-Fusion Zhou 2025 MASt3R feed-forward visual model + IMU + GNSS fusion

NeRF-based

System Author/Year Key Concepts
iMAP Sucar 2021 First NeRF-SLAM, single MLP, real-time tracking/mapping
BARF Lin 2021 Bundle-Adjusting NeRF, coarse-to-fine positional encoding, joint pose+NeRF opt (not full SLAM — pose+NeRF co-optimization)
NICE-SLAM Zhu & Peng 2022 Hierarchical feature grid (coarse/mid/fine), scalable
Co-SLAM Wang 2023 Hash grid (Instant-NGP) + coordinate encoding, 5-10× faster than NICE-SLAM
ESLAM Johari 2023 Tri-plane representation, O(N²) vs O(N³) memory
Point-SLAM Sandström 2023 Neural point cloud based
NeRF-SLAM Rosinol 2023 NeRF + classical SLAM pipeline
NICER-SLAM Zhu 2024 RGB-only NeRF-SLAM (no depth sensor), monocular depth integration
vMAP Kong 2023 Object-level NeRF-SLAM, per-object neural fields
GO-SLAM Zhang 2023 Global optimization + NeRF-SLAM, loop closure + global BA

3DGS-based

System Author/Year Key Concepts
SplaTAM Keetha 2024 Among the first 3DGS SLAM systems (concurrent with GS-SLAM, MonoGS), RGB-D, silhouette-guided densification
MonoGS Matsuki 2024 First monocular 3DGS SLAM (CVPR 2024 highlight), direct rasterization-based tracking, analytic camera Jacobians
GS-ICP SLAM Ha 2024 Gaussian-to-Gaussian ICP (Mahalanobis distance), geometric tracking
Photo-SLAM Huang 2024 Explicit geometry + implicit appearance (MLP color), anti-aliasing
RTG-SLAM Peng 2024 Real-time focus, adaptive Gaussian budget, Jetson Orin 25 FPS
EGG-Fusion Pan 2025 Geometry-aware Gaussian surfel fusion on the fly, information-filter-based, real-time
Online 3DGS Modeling Lee 2025 Online 3D Gaussian Splatting modeling with novel view selection
ActiveSplat Li 2025 Active mapping with 3DGS + Voronoi-based path planning
OpenGS-SLAM Yang 2025 Open-set dense semantic 3DGS SLAM, object-level scene understanding
LEGS Yu 2024 Language Embedded Gaussian Splats, real-time language-queryable 3D