Vehicle Distance Fusion Using LiDAR and Monocular Depth Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for measuring distance information, such as those using LiDAR and monocular cameras, face challenges in environments with no texture or periodic patterns, and struggle to maintain high accuracy across various scenes.
Innovation Solution
An information processing apparatus that combines distance information from a LiDAR sensor with distance information estimated from monocular camera images using a CNN, generating stable distance information by calibrating and transforming coordinates between the two sources, and using weighted averaging based on reliability determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If distance information is measured using images captured from a plurality of viewpoints by calculating degrees of similarity between local regions, then distance measurement can be performed, but it is difficult to calculate the correct distance if the subject has no texture or if the subject is a periodic pattern
Solution Approach 1:
The patent combines multiple distance information sources (stereo vision, LiDAR, and CNN-based monocular depth estimation) into a unified distance map. By merging the strength of each method, the system overcomes the limitation of stereo vision failing on textureless or periodic surfaces, as the CNN-based estimator can provide reliable distance information in these challenging scenes.
Solution Approach 2:
The patent creates a composite distance estimation system that integrates different sensing modalities (optical stereo imaging, laser ranging, and neural network-based monocular depth estimation). This composite approach allows the system to adapt to various scene conditions by relying on the most appropriate sensing modality for each specific situation.
2Adaptability or versatility
If a convolutional neural network (CNN) is used to estimate distance information from a monocular camera image, then distance estimation can be achieved, but the accuracy is insufficient compared to combined methods
Solution Approach 1:
The patent uses LiDAR-measured distance information as an intermediary to correct and refine the CNN-based distance estimates. The LiDAR data serves as a reliable reference that guides the CNN estimator, improving its accuracy while maintaining its advantage of working with monocular images.
Solution Approach 2:
The system implements a feedback mechanism where the LiDAR distance measurements are used to evaluate and correct the CNN-based distance estimates. This feedback loop allows the system to continuously improve the accuracy of monocular depth estimation by comparing it against the more reliable LiDAR measurements and adjusting accordingly.
3Measurement precision
If distance information from multiple sources is combined, then higher accuracy can be achieved, but the system complexity increases
Solution Approach 1:
The patent applies local quality by determining the reliability of each distance information source on a per-pixel or per-region basis. Instead of uniformly combining all distance sources, the system selectively weights and integrates them based on their local reliability, reducing unnecessary complexity while maintaining high accuracy where needed.
Solution Approach 2:
The system dynamically adjusts the weighting and integration strategy for combining distance information from different sources based on scene conditions, sensor reliability, and measurement quality. This dynamic approach allows the system to simplify processing in favorable conditions while automatically increasing robustness when challenges are detected.
Data Source
AI summary
This invention provides an information processing apparatus comprising a first acquiring unit configured to acquire first distance information from a distance sensor, a second acquiring unit configured to acquire an image from an image capturing device, a holding unit configured to hold a learning model for estimating distance information from images, an estimating unit configured to estimate second distance information corresponding to the image acquired by the second acquiring unit using the learning model, and a generating unit configured to generate third distance information based on the first distance information and the second distance information.


