Mono-Camera Depth Estimation Using Radar Doppler Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Monocular camera-based depth estimation in vehicular systems is challenging due to the loss of distance information in 2D images, leading to inaccurate and unreliable depth/distance estimations, especially in real-world scenarios where assumptions made for 2D-to-3D conversion may not be sufficiently correct.
Innovation Solution
A system and method that enhance depth estimation by using radar data for supervised training of a depth estimation model, where a processing unit receives 2D images from a mono-camera, generates estimated depth images, and corrects depth estimations using losses derived from differences between synthetic images, measured radar distances, and Doppler information, thereby improving the accuracy of depth estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If monocular camera is used for depth estimation, then device complexity is reduced, but measurement precision of depth/distance deteriorates
Solution Approach 1:
The patent combines monocular camera data with radar data (both distance and Doppler measurements) to create a fused dataset for training the depth estimation model. This merging of heterogeneous sensor data sources compensates for the inherent limitations of monocular cameras while maintaining relative system simplicity compared to multi-camera setups.
Solution Approach 2:
The patent introduces a supervised training model as an intermediary that learns to map 2D image features to 3D depth information using radar-annotated training data. This intermediary model bridges the gap between 2D image input and accurate depth output without requiring complex multi-camera hardware.
2Ease of operation
If assumptions are made for 2D-to-3D conversion, then depth estimation can be achieved, but reliability of depth estimation deteriorates
Solution Approach 1:
The patent performs preliminary action by collecting and preparing training data that pairs 2D camera images with corresponding radar distance and Doppler measurements before deploying the depth estimation model. This pre-training phase allows the model to learn accurate 2D-to-3D mappings without relying on unreliable assumptions during actual operation.
Solution Approach 2:
The patent uses radar measurements as feedback signals during the supervised training process to correct and refine the depth estimation model's predictions. The loss function incorporates differences between estimated depth and radar-measured distance, continuously improving reliability through feedback-driven optimization.
3Ease of manufacture
If manual data labeling is avoided, then ease of manufacture is improved, but measurement precision may deteriorate
Solution Approach 1:
The patent implements self-service by using radar sensors to automatically generate ground truth distance and velocity labels for training the depth estimation model. This eliminates the need for manual annotation while maintaining high precision, as radar provides accurate physical measurements that serve as reliable supervision signals.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution provides enhanced and accurate depth estimation for both stationary and moving objects, improving safety and functionality in autonomous and semi-autonomous vehicles by fusing heterogeneous data sources without requiring manual data labeling or Structure from Motion techniques, thus increasing spatial coverage and reliability.
Implementation Method 1
radar distance measurement
Implementation Method 2
radar doppler measurement
Data Source
AI summary
Systems and methods for depth estimation of images from a mono-camera by use of radar data by: receiving, a plurality of input 2-D images from the mono-camera; generating, by the processing unit, an estimated depth image by supervised training of an image estimation model; generating, by the processing unit, a synthetic image from a first input image and a second input image from the mono-camera by applying an estimated transform pose; comparing, by the processing unit, an estimated three-dimensional (3-D) point cloud to radar data by applying another estimated transform pose to a 3-D point cloud wherein the 3-D point cloud is estimated from a depth image by supervised training of the image estimation model to radar distance and radar doppler measurement; correcting a depth estimation of the estimated depth image by losses derived from differences: of the synthetic image and original images; of an estimated depth image and a measured radar distance; and of an estimated doppler information and measured radar doppler information.


