Monocular Depth Estimation Using Binocular-Trained Autoencoders

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing depth estimation methods in autonomous driving rely on expensive and complex binocular cameras, which are prone to failure due to unpredicted pixels in critical areas, necessitating a more efficient and cost-effective solution for depth estimation.

Innovation Solution

A method using an autoencoder network trained with binocular images containing instance segmentation labels to generate depth maps from monocular images, leveraging binocular parallax for improved depth prediction without requiring large data sets or extensive labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If binocular cameras are used for depth estimation, then depth prediction accuracy is improved, but device cost and complexity increase

Engineering Contradiction:
Improvedepth prediction accuracyVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses a monocular camera to capture images and employs an autoencoder network trained on binocular image data to generate depth maps. This approach copies the depth estimation capability from binocular systems without requiring actual binocular hardware, thereby reducing device complexity while maintaining depth prediction accuracy through learned relationships from training data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical binocular camera system with a monocular camera combined with a deep learning-based autoencoder network. The network substitutes the physical binocular vision mechanism with a computational model that has been trained on binocular image pairs, achieving depth estimation through software/algorithms rather than hardware complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If binocular cameras are used for depth estimation, then depth prediction accuracy is improved, but cost increases

Engineering Contradiction:
Improvedepth prediction accuracyVSAvoidsystem cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent copies the depth estimation functionality from expensive binocular camera systems using a monocular camera paired with a trained autoencoder network. This allows achieving similar depth prediction accuracy without the high cost of binocular hardware, making the system more cost-effective while maintaining measurement precision.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces expensive binocular cameras with a cheaper monocular camera and uses a pre-trained autoencoder network model. The network, trained on large datasets of binocular images, serves as a disposable computational resource that provides accurate depth estimation without requiring expensive hardware, thereby reducing system cost while maintaining accuracy.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Device complexity

If traditional depth estimation methods are used, then implementation is straightforward, but reliability decreases due to unpredicted pixels in critical areas

Engineering Contradiction:
Improvemethod implementation simplicityVSAvoidobstacle detection reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent employs an autoencoder network that processes monocular images through encoding and decoding stages, with the decoder generating depth maps based on learned patterns from training data. This feedback mechanism allows the system to continuously refine depth predictions by comparing encoded features with decoded outputs, improving reliability in critical areas while maintaining implementation simplicity through automated processing.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary training of the autoencoder network on large datasets of binocular images and their corresponding depth maps before deployment. This preliminary action captures complex depth relationships and patterns, enabling the model to make reliable predictions on new monocular images without requiring complex real-time processing, thus improving reliability while keeping implementation straightforward.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The method achieves accurate depth estimation using monocular cameras, reducing costs and data requirements while enhancing the reliability of obstacle detection and 3D scene reconstruction in autonomous driving.

Implementation Method 1

leveraging binocular parallax for improved depth prediction

Methodology Applied
Scientific EffectBinocular parallax: Parallax

Data Source

PatentUS12423848B2Method for generating depth in images, electronic device, and non-transitory storage medium
Publication Date: 2025.09.23 HON HAI PRECISION INDUSTRY CO LTD
  • US12423848B2 patent drawing
  • US12423848B2 patent drawing
  • US12423848B2 patent drawing

AI summary

A method and system for generating depth in monocular images acquires multiple sets of binocular images to build a dataset containing instance segmentation labels as to content; training an work using the dataset with instance segmentation labels to obtain a trained autoencoder network; acquiring monocular image, the monocular image is input into the trained autoencoder network to obtain a first disparity map and the first disparity map is converted to obtain depth image corresponding to the monocular image. The method combines binocular images with instance segmentation images as training data for training an autoencoder network, monocular images can simply be input into the autoencoder network to output the disparity map. Depth estimation for monocular images is achieved by converting the disparity map to a depth image corresponding to the monocular image. An electronic device and a non-transitory storage are also disclosed.