Self-Supervised Scale-Aware Monocular Depth Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Monocular cameras face challenges in estimating depth due to limited resolution, image artifacts, and scale ambiguities, which hinder situational awareness and navigation in robotic devices, especially when relying on self-supervised training without additional ground-truth data.

Innovation Solution

A depth system that employs a training architecture using monocular video and incorporates a pose model with a velocity component in the loss function to learn scale-aware depth estimates, eliminating the need for secondary ground-truth data by computing velocity supervision loss from instantaneous camera velocity measurements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If monocular cameras are used for depth estimation, then cost is reduced, but depth accuracy and scale awareness deteriorate

Engineering Contradiction:
ImprovecostVSAvoiddepth accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The system uses the monocular camera itself to provide the velocity information needed for scale-aware depth estimation. The camera's own motion data (velocity) is fed back into the loss function to enable the network to learn metric depth without external sensors or ground truth data, making the system self-sufficient and cost-effective

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention changes the training parameter by incorporating velocity information into the loss function. This parameter modification enables the network to learn scale-aware depth estimates from monocular video sequences, transforming an underdetermined problem into a solvable one without adding hardware costs

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If self-supervised training is used without ground-truth data, then data availability is improved, but scale awareness deteriorates

Engineering Contradiction:
Improvedata availabilityVSAvoidscale awareness
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The system implements feedback by using the network's own predictions and the camera's velocity measurements to construct a supervisory signal. The velocity-supervised loss function provides continuous feedback during training, enabling the network to learn scale-aware depth estimates without requiring external ground truth data

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Velocity information acts as an intermediary that bridges the gap between monocular image data and metric depth estimation. This intermediate parameter (camera velocity) enables the training process to infer scale information without directly observing ground truth depth measurements

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If photometric loss is used for training, then training simplicity is improved, but scale ambiguity worsens

Engineering Contradiction:
Improvetraining simplicityVSAvoidscale ambiguity
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The invention merges photometric loss with velocity-supervised loss into a combined training objective. This combination maintains the simplicity of photometric loss while adding scale-awareness through velocity constraints, achieving both training ease and scale accuracy simultaneously

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11176709B2Systems and methods for self-supervised scale-aware training of a model for monocular depth estimation
Publication Date: 2021.11.16 TOYOTA JIDOSHA KK
  • US11176709B2 patent drawing
  • US11176709B2 patent drawing
  • US11176709B2 patent drawing

AI summary

System, methods, and other embodiments described herein relate to self-supervised training of a depth model for monocular depth estimation. In one embodiment, a method includes processing a first image of a pair according to the depth model to generate a depth map. The method includes processing the first image and a second image of the pair according to a pose model to generate a transformation that defines a relationship between the pair. The pair of images are separate frames depicting a scene of a monocular video. The method includes generating a monocular loss and a pose loss, the pose loss including at least a velocity component that accounts for motion of a camera between the training images. The method includes updating the pose model according to the pose loss and the depth model according to the monocular loss to improve scale awareness of the depth model in producing depth estimates.