Camera-Radar Depth Estimation With Mask-Guided Sensor Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing depth estimation methods suffer from inaccuracies and high hardware costs, particularly in solutions using multi-line lidar, structured light, monocular depth estimation, and binocular cameras.
Innovation Solution
A method that integrates camera images with radar data to determine a reference line and radar line, creating a mask for accurate depth estimation using a depth estimation model trained on fused image and radar data, employing low-cost radar sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multi-line lidar or structured light is used for depth estimation, then measurement precision is improved, but device complexity and hardware costs increase
Solution Approach 1:
The patent combines radar data with camera images to create a fused input for depth estimation. The radar provides accurate distance measurements while the camera provides visual context, merging both data sources to achieve high precision depth estimation without requiring complex multi-line lidar or structured light systems.
Solution Approach 2:
The patent uses a depth estimation model trained on fused radar and image data to predict depth maps. This computational approach copies the functionality of expensive hardware systems through software-based depth prediction, achieving similar measurement precision with lower hardware complexity.
2Device complexity
If monocular depth estimation is used, then device complexity is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent merges radar data with monocular camera images to enhance depth estimation accuracy. The radar provides reliable distance information that compensates for the limitations of monocular depth estimation, achieving high precision while maintaining simple hardware configuration.
Solution Approach 2:
The patent introduces radar data as an intermediary element that mediates between the simple monocular camera input and the desired high-precision depth output. The radar acts as a bridge providing additional geometric constraints that improve depth estimation accuracy without increasing overall system complexity.
3Measurement precision
If binocular cameras are used for depth estimation, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent combines radar measurements with single-camera images to achieve depth estimation accuracy comparable to binocular systems. The radar provides direct distance measurements that replace the need for stereoscopic vision, achieving high precision with simpler hardware.
Solution Approach 2:
The patent replaces the mechanical stereoscopic system (binocular cameras requiring precise calibration and synchronization) with a hybrid radar-image system. This substitution maintains measurement precision while reducing hardware complexity and calibration requirements.
4Measurement precision
If radar data is integrated with camera images, then measurement precision is improved, but processing complexity increases
Solution Approach 1:
The patent performs preliminary registration of radar data with camera images before depth estimation, aligning both data sources in a common coordinate system. This preprocessing step, including ground area identification and mask creation, simplifies subsequent depth estimation by ensuring proper data alignment and reducing processing complexity during inference.
Data Source
AI summary
A method includes separately obtaining a to-be-detected image captured by a camera and radar data synchronously collected by a radar sensor; determining a reference line in the to-be-detected image; registering the radar data with the to-be-detected image, where a registered to-be-detected image includes the reference line and a radar line obtained based on the radar data, and the radar line includes pixels in the to-be-detected image that correspond to reflection points corresponding to the radar data; determining, based on the reference line and the radar line, a mask corresponding to the registered to-be-detected image; and inputting the radar data, the to-be-detected image, and the mask into a depth estimation model, to obtain a depth image corresponding to the to-be-detected image.


