The Evolution of Monocular Depth Estimation:From Spatial Regression to GenerativeFoundations and the Reliability Gaps
DOI:
https://doi.org/10.53799/zsh31s95Keywords:
Monocular depth estimation, self-supervised learning, vision transformers, adverse weatherAbstract
Monocular Depth Estimation (MDE) is one of the most rigorously studied problems in modern computer vision, yet it is fundamentally ill-posed. Recovering absolute three-dimensional geometry from a single two-dimensional projection is mathematically impossible without strong inductive priors. This paper presents a structurally organized review of MDE's evolution spanning from 2005 to 2026. We trace the trajectory from handcrafted Markov random fields to brute-force pixel-wise regression with convolutional neural networks (CNNs), and then to photometric self-supervision, which liberated the field from expensive LiDAR sensor suites. We further analyzed the vision transformers (ViTs) and generative diffusion priors have achieved unprecedented zero-shot metric generalization - with models such as Metric3D v2 attaining an Absolute Relative Error (AbsRel) as low as 0.039 on the KITTI benchmark. We also show that the standard self supervised baseline Monodepth2 degrades catastrophically to an AbsRel of 1.185 under nighttime conditions in NuScenes-Night dataset, which is a massive performance collapse from its clear-weather baseline. Physical-prior models such as PhysDepth recover this to 0.118 AbsRel by embedding Rayleigh scattering theory directly into the network. At the efficiency frontier, architectures such as LEDepth achieve 0.101 AbsRel at 5.7 ms inference time with only 3.1M parameters. We conclude that the future of depth estimation lies in embedding rigid physical and spatiotemporal priors into foundational latent spaces to ensure unyielding reliability in the physical world.
References
[1] J. E. Cutting and P. M. Vishton, "Perceiving layout and knowing distances," in Perception of Space and Motion, Elsevier, 1995, pp. 69–117.
[2] M. Bijelic et al., "Seeing through fog without seeing fog," in Proc. IEEE/CVF CVPR, 2020, pp. 11682–11692.
[3] S. Shao et al., "Self-supervised monocular depth and ego-motion estimation in endoscopy," Medical Image Analysis, vol. 77, p. 102338, 2022.
[4] S. Nadeem and A. Kaufman, "Computer-aided detection of polyps in optical colonoscopy," Machine Vision and Applications, vol. 27, no. 7, pp. 1037–1053, 2016.
[5] A Saxena, S. H. Chung, and A. Y. Ng, "Learning depth from single monocular images," in Proc. NIPS, 2005, pp. 1161–1168.
[6] D. Eigen, C. Puhrsch, and R. Fergus, "Depth map prediction from a single image using a multi-scale deep network," arXiv:1406.2283, 2014.
[7] D. Hoiem, A. A. Efros, and M. Hebert, "Geometric context from a single image," in Proc. ICCV, 2005, pp. 654–661.
[8] K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," arXiv:1409.1556, 2015.
[9] R. Garg et al., "Unsupervised CNN for single view depth estimation," arXiv:1603.04992, 2016.
[10] C. Godard, O. Mac Aodha, and G. J. Brostow, "Unsupervised monocular depth estimation with left-right consistency," arXiv:1609.03677, 2017.
[11] Saxena, M. Sun, and A. Y. Ng, "Make3D: Learning 3D scene structure from a single still image," IEEE Trans. PAMI, vol. 31, no. 5, pp. 824–840, 2009.
[12] T. Lindeberg, "Scale-space theory: A basic tool for analysing structures at different scales," J. Applied Statistics, vol. 21, pp. 224–270, 1994.
[13] J. Diebel and S. Thrun, "An application of Markov random fields to range sensing," NIPS, pp. 291–298, 2005.
[14] N. Silberman et al., "Indoor segmentation and support inference from RGBD images," in Proc. ECCV, 2012, pp. 746–760.
[15] D. Eigen and R. Fergus, "Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture," arXiv:1411.4734, 2015.
[16] K. Karsch, C. Liu, and S. B. Kang, "Depth transfer: Depth extraction from video using non-parametric sampling," IEEE Trans. PAMI, vol. 36, no. 11, pp. 2144–2158, 2014.
[17] Laina et al., "Deeper depth prediction with fully convolutional residual networks," arXiv:1606.00373, 2016.
[18] He et al., "Deep residual learning for image recognition," arXiv:1512.03385, 2015.
[19] Long, E. Shelhamer, and T. Darrell, "Fully convolutional networks for semantic segmentation," arXiv:1411.4038, 2015.
[20] B. Li et al., "Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFs," in Proc. IEEE CVPR, 2015, pp. 1119–1127.
[21] P. Wang et al., "Towards unified depth and semantic prediction from a single image," in Proc. IEEE CVPR, 2015, pp. 2800–2809.
[22] D. Xu et al., "Multi-scale continuous CRFs as sequential deep networks for monocular depth estimation," arXiv:1704.02157, 2017.
[23] S. Xie and Z. Tu, "Holistically-nested edge detection," arXiv:1504.06375, 2015.
[24] Geiger, P. Lenz, and R. Urtasun, "Are we ready for autonomous driving? The KITTI vision benchmark suite," in Proc. IEEE CVPR, 2012, pp. 3354–3361.
[25] J. Xie, R. Girshick, and A. Farhadi, "Deep3D: Fully automatic 2D-to-3D video conversion with deep convolutional neural networks," arXiv:1604.03650, 2016.
[26] H. Fu et al., "Deep ordinal regression network for monocular depth estimation," arXiv:1806.02446, 2018.
[27] S. F. Bhat, I. Alhashim, and P. Wonka, "AdaBins: Depth estimation using adaptive bins," in Proc. IEEE/CVF CVPR, 2021, pp. 4008–4017.
[28] R. Ranftl et al., "Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer," arXiv:1907.01341, 2020.
[29] Z. Li et al., "BinsFormer: Revisiting adaptive bins for monocular depth estimation," arXiv:2204.00987, 2022.
[30] J. H. Lee et al., "From big to small: Multi-scale local planar guidance for monocular depth estimation," arXiv:1907.10326, 2021.
[31] G. Huang et al., "Densely connected convolutional networks," arXiv:1608.06993, 2018.
[32] S. Abdulwahab et al., "Deep monocular depth estimation based on content and contextual features," Sensors, vol. 23, no. 6, 2023.
[33] J. Hu et al., "Revisiting single image depth estimation: Toward higher resolution maps with accurate object boundaries," arXiv:1803.08673, 2018.
[34] L.-C. Chen et al., "DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs," arXiv:1606.00915, 2017.
[35] S. Paul, D. Mishra, and S. K. Marimuthu, "Nested DWT-based CNN architecture for monocular depth estimation," Sensors, vol. 23, no. 6, 2023.
[36] O. Ronneberger, P. Fischer, and T. Brox, "U-Net: Convolutional networks for biomedical image segmentation," arXiv:1505.04597, 2015.
[37] N. Rahaman et al., "On the spectral bias of neural networks," arXiv:1806.08734, 2019.
[38] T. Zhou et al., "Unsupervised learning of depth and ego-motion from video," arXiv:1704.07813, 2017.
[39] R. Mahjourian, M. Wicke, and A. Angelova, "Unsupervised learning of depth and ego-motion from monocular video using 3D geometric constraints," arXiv:1802.05522, 2018.
[40] J.-W. Bian et al., "Unsupervised scale-consistent depth and ego-motion learning from monocular video," arXiv:1908.10553, 2019.
[41] C. Godard et al., "Digging into self-supervised monocular depth estimation," arXiv:1806.01260, 2019.
[42] V. Guizilini et al., "3D packing for self-supervised monocular depth estimation," arXiv:1905.02693, 2020.
[43] Masoumian et al., "GCNDepth: Self-supervised monocular depth estimation based on graph convolutional network," arXiv:2112.06782, 2021.
[44] X. Lyu et al., "HR-Depth: High resolution self-supervised monocular depth estimation," arXiv:2012.07356, 2020.
[45] V. Casser et al., "Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos," arXiv:1811.06152, 2018.
[46] Zhou et al., "Manydepth2: Motion-aware self-supervised monocular depth estimation in dynamic scenes," arXiv:2312.15268, 2025.
[47] J. Watson et al., "The temporal opportunist: Self-supervised multi-frame monocular depth," arXiv:2104.14540, 2021.
[48] J. Kopf, X. Rong, and J.-B. Huang, "Robust consistent video depth estimation," arXiv:2012.05901, 2021.
[49] C. Wang et al., "Learning depth from monocular videos using direct methods," arXiv:1712.00175, 2017.
[50] Dosovitskiy et al., "An image is worth 16×16 words: Transformers for image recognition at scale," arXiv:2010.11929, 2021.
[51] R. Ranftl, A. Bochkovskiy, and V. Koltun, "Vision transformers for dense prediction," arXiv:2103.13413, 2021.
[52] Vaswani et al., "Attention is all you need," arXiv:1706.03762, 2023.
[53] Z. Liu et al., "Swin transformer: Hierarchical vision transformer using shifted windows," arXiv:2103.14030, 2021.
[54] Y. Li and X. Wei, "MobileDepth: Monocular depth estimation based on lightweight vision transformer," Applied Artificial Intelligence, vol. 38, 2024.
[55] Zhang et al., "LEDepth: A lightweight self-supervised monocular depth estimation network combining CNN and transformer," Signal, Image and Video Processing, vol. 19, 2025.
[56] X. Sui et al., "Lightweight monocular depth estimation using a fusion-improved transformer," Scientific Reports, vol. 14, 2024.
[57] Z. Wang et al., "Lightweight self-supervised monocular depth estimation through CNN and transformer integration," IEEE Access, vol. 12, pp. 167934–167943, 2024.
[58] S. F. Bhat et al., "ZoeDepth: Zero-shot transfer by combining relative and metric depth," arXiv:2302.12288, 2023.
[59] W. Yuan et al., "NeW CRFs: Neural window fully-connected CRFs for monocular depth estimation," arXiv:2203.01502, 2022.
[60] J. Bai et al., "GLPanoDepth: Global-to-local panoramic depth estimation," arXiv:2202.02796, 2022.
[61] K. Saunders, G. Vogiatzis, and L. Manso, "Self-supervised monocular depth estimation: Let's talk about the weather," arXiv:2307.08357, 2023.
[62] Y. Ding et al., "WaterMono: Teacher-guided anomaly masking and enhancement boosting for robust underwater self-supervised monocular depth estimation," arXiv:2406.13344, 2024.
[63] K. Peng et al., "PhysDepth: Plug-and-play physical refinement for monocular depth estimation in challenging environments," arXiv:2412.04666, 2026.
[64] K. Jiang et al., "Always clear depth: Robust monocular depth estimation under adverse weather," arXiv:2505.12199, 2025.
[65] K. Wang et al., "Regularizing nighttime weirdness: Efficient self-supervised monocular depth estimation in the dark," arXiv:2108.03830, 2021.
[66] Y. Huang et al., "DASP: Self-supervised nighttime monocular depth estimation with domain adaptation of spatiotemporal priors," IEEE Robotics and Automation Letters, vol. 11, no. 2, pp. 2074–2081, 2026.
[67] S. G. Narasimhan and S. K. Nayar, "Vision and the atmosphere," Int. J. Comput. Vision, vol. 48, no. 3, pp. 233–254, 2002.
[68] R. Rombach et al., "High-resolution image synthesis with latent diffusion models," arXiv:2112.10752, 2022.
[69] D. Akkaynak and T. Treibitz, "Sea-Thru: A method for removing water from underwater images," in Proc. IEEE CVPR, 2019, pp. 1682–1691.
[70] J. Wang and X. Ye, "Advancing monocular depth estimation by integrating underwater optical imaging priors into transformer-based network," Optics Express, vol. 33, pp. 32631–32648, 2025.
[71] J. Shen, Z. Huang, and L. Jiao, "Self-supervised monocular depth estimation on construction sites in low-light conditions and dynamic scenes," Automation in Construction, vol. 168, p. 105848, 2024.
[72] B. Ke et al., "Marigold: Affordable adaptation of diffusion-based image generators for image analysis," arXiv:2505.09358, 2025.
[73] J. Ho, A. Jain, and P. Abbeel, "Denoising diffusion probabilistic models," arXiv:2006.11239, 2020.
[74] L. Yang et al., "Depth Anything V2," arXiv:2406.09414, 2024.
[75] Roberts et al., "Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding," in Proc. IEEE/CVF ICCV, 2021, pp. 10912–10922.
[76] L. Piccinelli et al., "UniDepth: Universal monocular metric depth estimation," arXiv:2403.18913, 2024.
[77] M. Hu et al., "Metric3D v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation," IEEE Trans. PAMI, vol. 46, no. 12, pp. 10579–10596, 2024.
[78] G. Bae, I. Budvytis, and R. Cipolla, "Estimating and exploiting the aleatoric uncertainty in surface normal estimation," arXiv:2109.09881, 2021.
[79] R. Bhattarai and H. Rhodin, "Re-Depth Anything: Test-time depth refinement via self-supervised re-lighting," arXiv:2512.17908, 2025.
[80] R. Zhang and M. Shah, "Shape from intensity gradient," IEEE Trans. Systems, Man, Cybernetics, vol. 29, no. 3, pp. 318–325, 1999.
[81] Z. Cao et al., "PanDA: Towards panoramic depth anything with unlabeled panoramas and Möbius spatial augmentation," arXiv:2406.13378, 2025.
[82] B. Coors, A. P. Condurache, and A. Geiger, "SphereNet: Learning spherical representations for detection and classification in omnidirectional images," in Proc. ECCV, 2018.
[83] H. Caesar et al., "nuScenes: A multimodal dataset for autonomous driving," in Proc. IEEE/CVF CVPR, 2020, pp. 11621–11631.
[84] Y. Randall et al., "FLSea: Underwater visual-inertial and stereo-vision forward-looking datasets," Ph.D. dissertation, Univ. of Haifa, Israel, 2023.
[85] Zhang et al., "Lite-Mono: A lightweight CNN and transformer architecture for self-supervised monocular depth estimation," in Proc. IEEE/CVF CVPR, 2023, pp. 18537–18546.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 AIUB Journal of Science and Engineering

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
AJSE contents are under the terms of the Creative Commons Attribution License. This permits anyone to copy, distribute, transmit and adapt the work non-commercially provided the original work and source is appropriately cited.