Learned frontends, differentiable backends, end-to-end systems, foundation-model & neural SLAM, scene understanding
Key Concepts
- Learned vs hand-crafted — Replacing individual classical modules (features, depth, matching) with networks vs end-to-end learning
- Differentiability — Making classical optimization (RANSAC, BA) differentiable so it can be trained through
- Foundation models — Large pretrained models (CLIP, SAM, DUSt3R-family) as reusable perception backbones
Level 5 is organized into five pillars:
A. Frontend — learned perception components replacing hand-crafted modules
B. Backend — learned/certifiable optimization replacing classical solvers
C. Systems — end-to-end deep VO/SLAM pipelines
D. Scene Understanding — semantic, language, and relational reasoning on SLAM maps
E. Foundation-Model & Neural SLAM — pointmap transformers, NeRF- and 3DGS-based dense SLAM systems
A. Deep Frontend — Perception
Feature Detection & Matching
| System |
Author/Year |
Key Concepts |
| NetVLAD |
Arandjelović 2016 |
VLAD, place recognition |
| SuperPoint |
DeTone 2017 |
Homographic Adaptation, Self-supervised, VGG encoder + detector/descriptor heads |
| HardNet |
Mishchuk 2017 |
Learned local descriptor |
| R2D2 |
Revaud 2019 |
Repeatable + Reliable detector/descriptor, explicit repeatability/reliability maps |
| KeyNet |
Barroso-Laguna 2019 |
Learned keypoint detector |
| HF-Net |
Sarlin 2019 |
Global feature, Local feature, Visual localization |
| SuperGlue |
Sarlin 2020 |
Self/Cross-attention GNN, Sinkhorn optimal assignment, dustbin for outliers |
| DISK |
Tyszkiewicz 2020 |
Policy gradient (RL) training, match success/failure as reward |
| Patch NetVLAD |
Hausler 2021 |
Multi-scale patch-level VLAD |
| LoFTR |
Sun 2021 |
Detector-free, Transformer coarse-to-fine dense matching |
| LightGlue |
Lindenberger 2023 |
Adaptive depth/width, 5-10× faster than SuperGlue |
| XFeat |
Potje 2024 |
0.3M params, 1400 FPS (RTX 4090), 64-dim descriptor, embedded-friendly |
| RoMa |
Edstedt 2024 |
DINOv2 foundation feature + coarse-to-fine dense matching |
| DeDoDe |
Edstedt 2024 |
Joint detect-and-describe in one stage |
| RoMa v2 |
Edstedt 2025 |
Harder-better-faster-denser dense feature matching |
Depth Estimation
| System |
Author/Year |
Key Concepts |
| MonoDepth |
Godard 2016 |
Left-Right photometric consistency, self-supervised |
| MiDaS |
Ranftl 2020 |
Multi-dataset mixing, scale-and-shift invariant loss, relative depth |
| DPT |
Ranftl 2021 |
Dense Prediction Transformer (ViT backbone), global context |
| ZoeDepth |
Bhat 2023 |
Zero-shot metric depth, Metric Bins Module |
| Metric3D |
Yin 2023 |
Camera intrinsic-conditioned metric depth, Canonical Camera Space |
| Depth Anything |
Yang 2024 |
62M images, foundation model for monocular depth |
| Depth Anything V2 |
Yang 2024 |
Improved with synthetic data, better edge preservation |
| Depth Anything 3 |
Lin 2025 |
Any-view geometry from arbitrary inputs, depth-ray prediction target, single plain transformer (DINOv2), teacher-student training |
| Marigold |
Ke 2024 |
Stable Diffusion for depth, fine detail, uncertainty via sampling |
| Align3R |
Lu 2025 |
Video temporal consistency, DUSt3R-based, CVPR 2025 Highlight |
| Masked Depth Modeling (LingBot-Depth) |
Tan 2026 |
Fixes RGB-D failures on glass/mirrors/metal |
Optical Flow & Scene Flow
Camera Pose Regression & Relocalization
Object Detection & Segmentation for SLAM
| System |
Author/Year |
Key Concepts |
| YOLO (v1→v11) |
Redmon 2016→2024 |
Real-time object detection, Ultralytics ecosystem |
| DETR |
Carion 2020 |
Transformer detection, anchor-free, no NMS |
| RT-DETR |
Zhao (Baidu) 2023 |
Real-time DETR, YOLO-speed + Transformer quality |
| RF-DETR |
Robinson 2025 |
Weight-sharing NAS over DETRs, accuracy-latency Pareto tuning, first real-time detector past 60 AP on COCO |
| SAM |
Kirillov 2023 |
Segment Anything, prompt-based, Foundation Model |
| SAM 2 |
Meta 2024 |
Video segmentation, Memory Attention, temporal consistency |
| SAM 3 |
Carion 2025 |
Promptable concept segmentation (noun-phrase / exemplar prompts), presence head, detector + memory-based video tracker |
| Grounding DINO |
Liu 2023 |
Text-prompted detection → SAM pipeline (Grounded SAM) |
| Open-YOLO 3D |
Boudjoghra 2024 |
2D open-vocab detection → 3D instance seg, 16× faster |
B. Deep Backend — Optimization
Differentiable Bundle Adjustment
C. End-to-End Deep VO / SLAM Systems
Self-supervised & Learned VO
| System |
Author/Year |
Key Concepts |
| DeepVO |
Wang 2017 |
Supervised learning |
| SfM-Learner |
Zhou 2017 |
Unsupervised, deep depth + deep pose |
| DeMoN |
Ummenhofer 2017 |
Depth + Motion from two frames, encoder-decoder |
| UndeepVO |
Li 2018 |
Stereo self-supervised, absolute scale recovery |
| DeepTAM |
Zhou 2018 |
Deep tracking and mapping, cost volume based |
| DeepV2D |
Teed 2018 |
Iterative depth from video, differentiable geometry layers |
| Depth from Videos in the Wild |
Gordon 2019 |
Unconstrained video depth, learned camera intrinsics |
| Neural Ray Surfaces |
Vasiljevic 2020 |
Learned ray surface model, non-pinhole cameras |
| GradSLAM |
Murthy 2020 |
Differentiable SLAM framework (PyTorch, supports multiple SLAM backends) |
| DeepSLAM |
Li 2020 |
TrackingNet, MappingNet, LoopNet |
| MonoRec |
Wimbauer 2021 |
Self-supervised monocular 3D reconstruction, moving objects |
| TANDEM |
Koestler 2021 |
Real-time tracking + dense mapping via MVS depth, DSO-based |
Learning-based SLAM Systems
Latent Representation SLAM
Neural Rendering (reference)
NeRF/3DGS-based SLAM systems → see Pillar E below
| System |
Author/Year |
Key Concepts |
| NeRF |
Mildenhall 2020 |
Neural Radiance Fields, novel view synthesis (foundational) |
| DIFIX3D+ |
Wu 2025 |
Single-step diffusion for 3D reconstruction artifact removal (post-processing) |
D. Scene Understanding
Benchmarks & Foundations
| System |
Author/Year |
Key Concepts |
| EFM3D |
Straub (Meta) 2024 |
Egocentric Foundation Model 3D benchmark, depth/surface/semantic from ego-video |
3D Scene Graph
Semantic / Language-Grounded SLAM
Also see: LEGS, OpenGS-SLAM (Pillar E above); Open-YOLO 3D (Level 5 Object Detection)
E. Foundation-Model & Neural-Representation SLAM
Foundation-Model SLAM
NeRF-based
| System |
Author/Year |
Key Concepts |
| iMAP |
Sucar 2021 |
First NeRF-SLAM, single MLP, real-time tracking/mapping |
| BARF |
Lin 2021 |
Bundle-Adjusting NeRF, coarse-to-fine positional encoding, joint pose+NeRF opt (not full SLAM — pose+NeRF co-optimization) |
| NICE-SLAM |
Zhu & Peng 2022 |
Hierarchical feature grid (coarse/mid/fine), scalable |
| Co-SLAM |
Wang 2023 |
Hash grid (Instant-NGP) + coordinate encoding, 5-10× faster than NICE-SLAM |
| ESLAM |
Johari 2023 |
Tri-plane representation, O(N²) vs O(N³) memory |
| Point-SLAM |
Sandström 2023 |
Neural point cloud based |
| NeRF-SLAM |
Rosinol 2023 |
NeRF + classical SLAM pipeline |
| NICER-SLAM |
Zhu 2024 |
RGB-only NeRF-SLAM (no depth sensor), monocular depth integration |
| vMAP |
Kong 2023 |
Object-level NeRF-SLAM, per-object neural fields |
| GO-SLAM |
Zhang 2023 |
Global optimization + NeRF-SLAM, loop closure + global BA |
3DGS-based
| System |
Author/Year |
Key Concepts |
| SplaTAM |
Keetha 2024 |
Among the first 3DGS SLAM systems (concurrent with GS-SLAM, MonoGS), RGB-D, silhouette-guided densification |
| MonoGS |
Matsuki 2024 |
First monocular 3DGS SLAM (CVPR 2024 highlight), direct rasterization-based tracking, analytic camera Jacobians |
| GS-ICP SLAM |
Ha 2024 |
Gaussian-to-Gaussian ICP (Mahalanobis distance), geometric tracking |
| Photo-SLAM |
Huang 2024 |
Explicit geometry + implicit appearance (MLP color), anti-aliasing |
| RTG-SLAM |
Peng 2024 |
Real-time focus, adaptive Gaussian budget, Jetson Orin 25 FPS |
| EGG-Fusion |
Pan 2025 |
Geometry-aware Gaussian surfel fusion on the fly, information-filter-based, real-time |
| Online 3DGS Modeling |
Lee 2025 |
Online 3D Gaussian Splatting modeling with novel view selection |
| ActiveSplat |
Li 2025 |
Active mapping with 3DGS + Voronoi-based path planning |
| OpenGS-SLAM |
Yang 2025 |
Open-set dense semantic 3DGS SLAM, object-level scene understanding |
| LEGS |
Yu 2024 |
Language Embedded Gaussian Splats, real-time language-queryable 3D |