Monocular Neural Depth Estimation With Metric Scale Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional depth estimation methods requiring multiple cameras or physical markers are computationally intensive, making them unsuitable for mobile applications, and existing single-image depth estimation techniques struggle with scale/shift ambiguity, hindering accurate placement of virtual objects in augmented reality.

Innovation Solution

A neural network-based system that generates an affine-invariant depth map from a single image, transformed to a metric scale using sparse depth estimates and additional signals like surface normals, gravity direction, and planar regions, resolving scale/shift ambiguity and enabling rapid virtual object placement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple cameras and physical markers are used to reconstruct depth map, then measurement precision is improved, but device complexity and computation power requirements increase

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the unnecessary components (multiple cameras and physical markers) from the depth estimation system, retaining only a single image sensor while achieving accurate depth measurement through neural network processing of monocular images

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/optical system (multiple cameras and physical markers) with a computational system (neural network processing), substituting physical complexity with algorithmic intelligence to achieve depth estimation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If multiple cameras and physical markers are used to reconstruct depth map, then measurement precision is improved, but computation power requirements increase

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidcomputation power
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent removes the computationally intensive multi-camera processing and physical marker tracking, extracting only the essential function of depth estimation which can be achieved through lightweight neural network processing of single images

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses lightweight, portable image sensors and efficient neural network models that can be deployed on mobile devices with limited computational resources, replacing heavy-duty multi-camera systems

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Device complexity

If affine-invariant depth map is generated without scale information, then device complexity is reduced, but manufacturing precision (depth accuracy) deteriorates due to scale/shift ambiguity

Engineering Contradiction:
Improvesystem simplicityVSAvoiddepth map accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary component (scale estimation module) that bridges the gap between affine-invariant depth maps and metric-scale depth maps, using detected planar regions and vanishing points to infer scale and shift parameters

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary detection of planar regions and vanishing points in the image before generating the final depth map, using this information to pre-calculate scale and shift parameters that resolve the affine ambiguity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12430778B2Depth estimation using a neural network
Publication Date: 2025.09.30 GOOGLE LLC
  • US12430778B2 patent drawing
  • US12430778B2 patent drawing
  • US12430778B2 patent drawing

AI summary

According to an aspect, a method for depth estimation includes receiving image data from a sensor system, generating, by a neural network, a first depth map based on the image data, where the first depth map has a first scale, obtaining depth estimates associated with the image data, and transforming the first depth map to a second depth map using the depth estimates, where the second depth map has a second scale.