Scale-Aware Monocular Depth Training with Sparse Radar Supervision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing depth estimation methods using monocular images suffer from scale ambiguities and require expensive or limited training data, leading to reduced situational awareness and navigation difficulties for robotic devices.

Innovation Solution

A two-stage training process is employed, combining self-supervised photometric loss from monocular images with semi-supervised learning using sparse radar data to recover a single scale factor, refining depth estimates without dense annotations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If monocular images are used for depth estimation, then cost is reduced and field-of-view is improved, but depth information is not explicitly available and scale ambiguities occur

Engineering Contradiction:
ImprovecostVSAvoiddepth information
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent uses an autoencoder as an intermediary component that learns to compress monocular images into latent representations containing depth information. The encoder extracts features and the decoder reconstructs images, with the latent space serving as a mediator that encodes depth cues without requiring explicit depth sensors

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces mechanical depth sensing systems (LiDAR, radar, stereo cameras) with a computational approach using monocular image processing and autoencoders. The system substitutes physical depth measurement mechanisms with learned representations from 2D images, achieving depth estimation without expensive or complex hardware

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If LiDAR sensors are used for depth perception, then depth measurement precision is improved, but cost increases and errors occur in certain weather conditions

Engineering Contradiction:
Improvedepth measurementVSAvoidcost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent replaces expensive LiDAR sensors with inexpensive monocular cameras. The system uses readily available, low-cost image sensors instead of sophisticated active ranging devices, achieving acceptable depth estimation performance through computational methods rather than expensive hardware

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Ease of manufacture

If radar sensors are used for depth perception, then cost is reduced and weather robustness is improved, but depth data sparsity increases

Engineering Contradiction:
ImprovecostVSAvoiddepth data sparsity
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent merges radar point cloud data with monocular image data in a unified training framework. The autoencoder processes both modalities together, learning to complement sparse radar measurements with rich image information, thereby reducing the impact of radar sparsity while maintaining cost advantages

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates a composite representation by combining radar point cloud data with monocular image features in the latent space of the autoencoder. This composite approach leverages the strengths of both sensors - the cost-effectiveness and weather robustness of radar with the rich spatial information from images

Inventive Principle:
Principle #40Composite materials

4Measurement precision

If stereo cameras are used for depth capture, then depth information quality is improved, but cost increases and field-of-view is limited

Engineering Contradiction:
Improvedepth informationVSAvoidfield-of-view
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The patent makes the monocular camera system universal by enabling it to perform both 2D image capture and 3D depth estimation functions. The autoencoder learns to extract depth information from standard monocular images, allowing a single sensor to serve multiple purposes without requiring specialized stereo configurations

Inventive Principle:
Principle #6Universality (Multi-functionality)

5Ease of manufacture

If self-supervised photometric loss is used for training, then training data requirements are reduced, but scale ambiguities persist

Engineering Contradiction:
Improvetraining data availabilityVSAvoidscale accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces feedback from radar point cloud data into the training process. The radar measurements provide ground truth depth information that feeds back to correct scale ambiguities in the monocular depth estimates, allowing the system to learn accurate scale factors without requiring densely annotated training data

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system uses sparse radar measurements (partial action) rather than requiring complete dense annotations for training. By leveraging even limited radar point cloud data, the system achieves scale-aware training without the excessive data requirements of fully supervised approaches

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250278847A1Scale-aware self-supervised monocular depth with sparse radar supervision
Publication Date: 2025.09.04 TOYOTA RESEARCH INSTITUTE INC
  • US20250278847A1 patent drawing
  • US20250278847A1 patent drawing
  • US20250278847A1 patent drawing

AI summary

Systems and methods are provided for training a depth model to recover scale factor for self-supervised depth estimation in monocular images. According to some embodiments, a method comprises receiving an image representing a scene of an environment; deriving a depth map for the image based on a depth model, the depth map comprising depth values for pixels of the image; estimating a first scale for the image based the depth values; receiving depth data captured by a range sensor, the depth data comprising a point cloud representing the scene of the environment, the point cloud comprising depth measures; determining a second scale for the point cloud based on the depth measures; determining a scale factor based the second scale and the first scale; and updating the depth model based on the scale factor, wherein the depth model generates metrically accurate depth estimates based on the scale factor.