Monocular Neural Depth Estimation With Metric Scale Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional depth estimation methods requiring multiple cameras or physical markers are computationally intensive, making them unsuitable for mobile applications, and existing single-image depth estimation techniques struggle with scale/shift ambiguity, hindering accurate placement of virtual objects in augmented reality.
Innovation Solution
A neural network-based system that generates an affine-invariant depth map from a single image, transformed to a metric scale using sparse depth estimates and additional signals like surface normals, gravity direction, and planar regions, resolving scale/shift ambiguity and enabling rapid virtual object placement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras and physical markers are used to reconstruct depth map, then measurement precision is improved, but device complexity and computation power requirements increase
Solution Approach 1:
The patent extracts and removes the unnecessary components (multiple cameras and physical markers) from the depth estimation system, retaining only a single image sensor while achieving accurate depth measurement through neural network processing of monocular images
Solution Approach 2:
The patent replaces the mechanical/optical system (multiple cameras and physical markers) with a computational system (neural network processing), substituting physical complexity with algorithmic intelligence to achieve depth estimation
2Measurement precision
If multiple cameras and physical markers are used to reconstruct depth map, then measurement precision is improved, but computation power requirements increase
Solution Approach 1:
The patent removes the computationally intensive multi-camera processing and physical marker tracking, extracting only the essential function of depth estimation which can be achieved through lightweight neural network processing of single images
Solution Approach 2:
The patent uses lightweight, portable image sensors and efficient neural network models that can be deployed on mobile devices with limited computational resources, replacing heavy-duty multi-camera systems
3Device complexity
If affine-invariant depth map is generated without scale information, then device complexity is reduced, but manufacturing precision (depth accuracy) deteriorates due to scale/shift ambiguity
Solution Approach 1:
The patent introduces an intermediary component (scale estimation module) that bridges the gap between affine-invariant depth maps and metric-scale depth maps, using detected planar regions and vanishing points to infer scale and shift parameters
Solution Approach 2:
The patent performs preliminary detection of planar regions and vanishing points in the image before generating the final depth map, using this information to pre-calculate scale and shift parameters that resolve the affine ambiguity
Data Source
AI summary
According to an aspect, a method for depth estimation includes receiving image data from a sensor system, generating, by a neural network, a first depth map based on the image data, where the first depth map has a first scale, obtaining depth estimates associated with the image data, and transforming the first depth map to a second depth map using the depth estimates, where the second depth map has a second scale.


