Stereo-Supervised Autoencoder Training for Monocular Depth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Stereo cameras used for depth perception in autonomous driving are costly and have poor applicability, necessitating a more efficient and cost-effective alternative for obtaining depth images.

Innovation Solution

Training an autoencoder using images from a stereo camera to generate depth images from monocular camera inputs, utilizing a neural network to learn efficient representations through unsupervised learning and generating disparity maps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a stereo camera is used to obtain depth images, then depth perception accuracy is improved, but system cost and complexity increase

Engineering Contradiction:
Improvedepth perception accuracyVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses a monocular camera to capture images and employs an autoencoder neural network to generate virtual depth information, effectively copying the depth perception capability of stereo cameras through computational methods rather than physical duplicate sensors

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical stereo camera system with a computational approach using a single camera combined with deep learning algorithms, substituting physical complexity with information processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If a stereo camera is used to obtain depth images, then depth perception accuracy is improved, but system cost increases

Engineering Contradiction:
Improvedepth perception accuracyVSAvoidsystem cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent employs a single monocular camera which is significantly cheaper than a stereo camera system, accepting that the computational model needs continuous training and updates to maintain performance

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If a stereo camera is used to obtain depth images, then depth perception capability is improved, but applicability decreases

Engineering Contradiction:
Improvedepth perception capabilityVSAvoidapplicability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent makes the monocular camera system universal by training the autoencoder on diverse datasets and implementing continuous learning mechanisms, enabling the same hardware to adapt to various driving scenarios and environmental conditions

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12462413B2Method for training autoencoder, electronic device, and storage medium
Publication Date: 2025.11.04 HON HAI PRECISION INDUSTRY CO LTD
  • US12462413B2 patent drawing
  • US12462413B2 patent drawing
  • US12462413B2 patent drawing

AI summary

A method for training an autoencoder implemented in an electronic device includes obtaining a stereoscopic image as the vehicle is in motion, the stereoscopic image includes a left image and a right image; generating a stereo disparity map according to the left image; generating a predicted right image according to the left image and the stereo disparity map; and calculating a first mean square error between the predicted right image and the right image.