Monocular Visual-Inertial Depth Alignment for Metric Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing monocular depth estimation methods suffer from scale ambiguity and lack generalizability, leading to inaccurate and inconsistent depth measurements, which are critical for tasks like visual navigation and scene reconstruction.

Innovation Solution

A monocular visual-inertial system is used to integrate inertial data with monocular depth estimation, employing a global alignment phase to correct scale using least-squares estimation and a learning-based dense alignment phase to refine per-pixel scaling, leveraging a ScaleMapLearner network for improved metric accuracy and generalizability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If monocular depth estimation is used, then computational efficiency is improved, but metric accuracy deteriorates due to scale ambiguity

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidmetric accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary alignment process that bridges monocular depth estimation and metric accuracy. A alignment network is used as an intermediary component that takes monocular depth predictions and adjusts them using reference depth information from datasets like KITTI or ETH3D, thereby resolving the scale ambiguity issue while maintaining the efficiency of monocular processing

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter changes by transforming depth predictions through learnable scaling factors and offset parameters. The alignment network modifies the depth parameters (scale and offset) based on reference data, enabling the system to recover metric accuracy from affine-invariant monocular predictions by adjusting these parameters to match real-world measurements

Inventive Principle:
Principle #35Parameter changes

2Speed

If existing monocular depth estimation methods are used, then processing speed is maintained, but generalizability across environments deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidgeneralizability
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent implements universality by training the alignment network on multiple reference datasets (KITTI, ETH3D, nuScenes) with diverse environmental conditions. This multi-functional approach enables the same monocular depth estimation pipeline to generalize across different environments, from indoor scenes to outdoor urban areas, maintaining processing speed while improving adaptability through shared learning

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If scale ambiguity is present in monocular depth estimates, then computational simplicity is preserved, but measurement reliability deteriorates

Engineering Contradiction:
Improvecomputational simplicityVSAvoidmeasurement reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces feedback mechanisms where the alignment network continuously refines depth predictions by comparing them against reference depth maps from annotated datasets. This feedback loop adjusts the scale and offset parameters iteratively, ensuring reliable metric measurements while maintaining computational simplicity through efficient parameter optimization rather than complex multi-stage processing

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12400346B2Methods and apparatus for metric depth estimation using a monocular visual-inertial system
Publication Date: 2025.08.26 INTEL CORP
  • US12400346B2 patent drawing
  • US12400346B2 patent drawing
  • US12400346B2 patent drawing

AI summary

Methods, apparatus, systems, and articles of manufacture are disclosed for metric depth estimation using a monocular visual-inertial system. An example apparatus for metric depth estimation includes at least one memory, instructions in the apparatus, and processor circuitry to execute the instructions to access a globally-aligned depth prediction, the globally-aligned depth prediction generated based on a monocular depth estimator, access a dense scale map scaffolding, the dense scale map scaffolding generated based on visual-inertial odometry, regress a dense scale residual map determined using the globally-aligned depth prediction and the dense scale map scaffolding, and apply the dense scale residual map to the globally-aligned depth prediction.