Self-Supervised Depth Estimation Network Eliminates Labeled Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for training depth map generation networks require a large amount of labeled data, leading to significant time and resource consumption, and struggle to accurately estimate depth information for dynamic or occluded objects.

Innovation Solution

A learning device based on self-supervised learning that includes a first network for generating an estimated depth map and a second network for generating pose change information, using a composite image and pseudo depth map to calculate losses and update network parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a large amount of labeled data is used for training, then the depth map generation network can be trained, but significant time and resource consumption occurs

Engineering Contradiction:
Improvetraining capabilityVSAvoiddata preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-supervised learning by automatically generating supervision signals from the input images themselves through pose estimation and depth map generation, eliminating the need for external labeled data and manual annotation processes

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-processes images to generate depth maps and pose information that serve as supervision signals before the actual training process, preparing training data automatically without requiring external labeled datasets

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If traditional supervised learning is used, then labeled training data can be utilized, but the process consumes huge amounts of time and resources

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The network learns to generate accurate depth maps by receiving supervision signals from its own pose estimation and depth generation capabilities, creating a self-contained learning system that does not require external labeled data

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses the generated depth maps and pose information as feedback signals to continuously improve the depth map generation network through backpropagation, creating a closed-loop learning process

Inventive Principle:
Principle #23Feedback

3Reliability

If labeled data is manually annotated, then training data quality is ensured, but huge time and resource consumption occurs

Engineering Contradiction:
Improvedata qualityVSAvoiddata preparation ease
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The system automatically generates high-quality supervision signals through its own processing of input images, eliminating the need for manual annotation while maintaining data quality through self-supervised learning mechanisms

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system creates synthetic supervision signals by copying and transforming information from the input images through pose estimation and depth map generation, replacing the need for manual copied annotations

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250095174A1Learning Device, Learning Method And Test Device, Test Method Using The Same
Publication Date: 2025.03.20 HYUNDAI MOTOR CO LTD
  • US20250095174A1 patent drawing
  • US20250095174A1 patent drawing
  • US20250095174A1 patent drawing

AI summary

A learning device, a learning method thereof, a test device using the same, and a test method using the same are provided. The learning device may obtain a target image and a source image, generate an estimated depth map based on the target image via a first network, generate pose change information corresponding to a pose change between the target image and the source image, generate a composite image corresponding to the target image, determine a first loss based on the composite image and the target image, and determine a second loss, and back-propagate the first loss and the second loss and update a parameter of the first network and a parameter of the second network.