Self-Supervised Depth Estimation Network Eliminates Labeled Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for training depth map generation networks require a large amount of labeled data, leading to significant time and resource consumption, and struggle to accurately estimate depth information for dynamic or occluded objects.
Innovation Solution
A learning device based on self-supervised learning that includes a first network for generating an estimated depth map and a second network for generating pose change information, using a composite image and pseudo depth map to calculate losses and update network parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a large amount of labeled data is used for training, then the depth map generation network can be trained, but significant time and resource consumption occurs
Solution Approach 1:
The system performs self-supervised learning by automatically generating supervision signals from the input images themselves through pose estimation and depth map generation, eliminating the need for external labeled data and manual annotation processes
Solution Approach 2:
The system pre-processes images to generate depth maps and pose information that serve as supervision signals before the actual training process, preparing training data automatically without requiring external labeled datasets
2Measurement precision
If traditional supervised learning is used, then labeled training data can be utilized, but the process consumes huge amounts of time and resources
Solution Approach 1:
The network learns to generate accurate depth maps by receiving supervision signals from its own pose estimation and depth generation capabilities, creating a self-contained learning system that does not require external labeled data
Solution Approach 2:
The system uses the generated depth maps and pose information as feedback signals to continuously improve the depth map generation network through backpropagation, creating a closed-loop learning process
3Reliability
If labeled data is manually annotated, then training data quality is ensured, but huge time and resource consumption occurs
Solution Approach 1:
The system automatically generates high-quality supervision signals through its own processing of input images, eliminating the need for manual annotation while maintaining data quality through self-supervised learning mechanisms
Solution Approach 2:
The system creates synthetic supervision signals by copying and transforming information from the input images through pose estimation and depth map generation, replacing the need for manual copied annotations
Data Source
AI summary
A learning device, a learning method thereof, a test device using the same, and a test method using the same are provided. The learning device may obtain a target image and a source image, generate an estimated depth map based on the target image via a first network, generate pose change information corresponding to a pose change between the target image and the source image, generate a composite image corresponding to the target image, determine a first loss based on the composite image and the target image, and determine a second loss, and back-propagate the first loss and the second loss and update a parameter of the first network and a parameter of the second network.


