Monocular Depth Estimation Using Binocular-Trained Autoencoders
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing depth estimation methods in autonomous driving rely on expensive and complex binocular cameras, which are prone to failure due to unpredicted pixels in critical areas, necessitating a more efficient and cost-effective solution for depth estimation.
Innovation Solution
A method using an autoencoder network trained with binocular images containing instance segmentation labels to generate depth maps from monocular images, leveraging binocular parallax for improved depth prediction without requiring large data sets or extensive labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If binocular cameras are used for depth estimation, then depth prediction accuracy is improved, but device cost and complexity increase
Solution Approach 1:
The patent uses a monocular camera to capture images and employs an autoencoder network trained on binocular image data to generate depth maps. This approach copies the depth estimation capability from binocular systems without requiring actual binocular hardware, thereby reducing device complexity while maintaining depth prediction accuracy through learned relationships from training data.
Solution Approach 2:
The patent replaces the mechanical binocular camera system with a monocular camera combined with a deep learning-based autoencoder network. The network substitutes the physical binocular vision mechanism with a computational model that has been trained on binocular image pairs, achieving depth estimation through software/algorithms rather than hardware complexity.
2Measurement precision
If binocular cameras are used for depth estimation, then depth prediction accuracy is improved, but cost increases
Solution Approach 1:
The patent copies the depth estimation functionality from expensive binocular camera systems using a monocular camera paired with a trained autoencoder network. This allows achieving similar depth prediction accuracy without the high cost of binocular hardware, making the system more cost-effective while maintaining measurement precision.
Solution Approach 2:
The patent replaces expensive binocular cameras with a cheaper monocular camera and uses a pre-trained autoencoder network model. The network, trained on large datasets of binocular images, serves as a disposable computational resource that provides accurate depth estimation without requiring expensive hardware, thereby reducing system cost while maintaining accuracy.
3Device complexity
If traditional depth estimation methods are used, then implementation is straightforward, but reliability decreases due to unpredicted pixels in critical areas
Solution Approach 1:
The patent employs an autoencoder network that processes monocular images through encoding and decoding stages, with the decoder generating depth maps based on learned patterns from training data. This feedback mechanism allows the system to continuously refine depth predictions by comparing encoded features with decoded outputs, improving reliability in critical areas while maintaining implementation simplicity through automated processing.
Solution Approach 2:
The patent performs preliminary training of the autoencoder network on large datasets of binocular images and their corresponding depth maps before deployment. This preliminary action captures complex depth relationships and patterns, enabling the model to make reliable predictions on new monocular images without requiring complex real-time processing, thus improving reliability while keeping implementation straightforward.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The method achieves accurate depth estimation using monocular cameras, reducing costs and data requirements while enhancing the reliability of obstacle detection and 3D scene reconstruction in autonomous driving.
Implementation Method 1
leveraging binocular parallax for improved depth prediction
Data Source
AI summary
A method and system for generating depth in monocular images acquires multiple sets of binocular images to build a dataset containing instance segmentation labels as to content; training an work using the dataset with instance segmentation labels to obtain a trained autoencoder network; acquiring monocular image, the monocular image is input into the trained autoencoder network to obtain a first disparity map and the first disparity map is converted to obtain depth image corresponding to the monocular image. The method combines binocular images with instance segmentation images as training data for training an autoencoder network, monocular images can simply be input into the autoencoder network to output the disparity map. Depth estimation for monocular images is achieved by converting the disparity map to a depth image corresponding to the monocular image. An electronic device and a non-transitory storage are also disclosed.


