LIDAR-Camera Depth Fusion for High-Resolution 3D Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current object detection methods for autonomous driving face challenges in recognizing and classifying objects in 3D spaces with multiple objects close by, particularly due to the limitations of camera-only, LIDAR-only, and simple fusion approaches, which result in inaccurate depth models and artifacts such as missing patches and dynamic range issues.
Innovation Solution
A method for training an artificial intelligence module using multiple sets of training data comprising camera images and depth images from flash LIDAR sensors, where depth measurement data from combined images of multiple flash LIDAR sensors serves as ground truth to enhance the resolution and accuracy of the depth model, overcoming limitations like rolling shutter effects and motion artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If camera-only or LIDAR-only or simple fusion approaches are used for object detection, then device complexity is reduced, but measurement precision and reliability of depth model deteriorate
Solution Approach 1:
The patent combines camera images and LIDAR depth images through a neural network-based fusion approach. The system merges the visual information from the camera with the depth information from the LIDAR sensor to generate a high-resolution depth model, achieving superior measurement precision compared to using either sensor alone or simple concatenation-based fusion.
Solution Approach 2:
The patent introduces a neural network as an intermediary component that processes and fuses the camera and LIDAR data. This intermediary learns to combine the features from both sensors optimally, transforming the input data into an accurate depth model without requiring complex manual fusion algorithms.
2Measurement precision
If high-resolution scanning LIDAR sensors are used for depth completion, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent uses a neural network to learn the mapping from low-resolution LIDAR depth images to high-resolution depth models. Instead of using expensive high-resolution scanning LIDAR sensors, the system copies the essential depth information patterns learned from training data and applies them to enhance the resolution of depth images from simpler, lower-resolution LIDAR sensors.
Solution Approach 2:
The patent transforms the resolution parameter of the depth image through learned transformations in the neural network. The system changes the spatial resolution parameter from low to high by leveraging correlations learned from training data, achieving high-resolution depth models without requiring high-resolution LIDAR hardware.
3Measurement precision
If ground truth generation by camera-based semi-global matching is used, then measurement precision is improved, but object-generated harmful factors increase due to motion artifacts
Solution Approach 1:
The patent performs preliminary fusion of camera and LIDAR data before depth completion, creating a preliminary depth estimate that incorporates both visual and depth information. This preliminary action helps establish a more robust foundation for the subsequent depth completion process, reducing the impact of motion artifacts that would otherwise propagate through the pipeline.
Solution Approach 2:
The patent implements a feedback mechanism where the fused features from camera and LIDAR are iteratively refined through the neural network. The system uses the combined information to correct and refine depth estimates, providing feedback that reduces motion artifacts and improves the overall accuracy of the depth model.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution enables the AI module to produce a high-resolution depth model with improved accuracy and reduced sensitivity to motion, allowing for better object detection and classification, even with lower resolution LIDAR sensors, and reduces artifacts like shadows and self-occlusions.
Implementation Method 1
a depth image of the scene taken by a flash LIDAR sensor of the vehicle
Data Source
AI summary
A method for training an artificial intelligence module for an advanced driver assistance system for obtaining a depth model of a scene uses data of a camera image and a depth image of a LIDAR sensor as input, and is trained for outputting the depth model of the scene. The depth model has a higher resolution than the depth image of the LIDAR sensor. The method of training the artificial intelligence module includes acquiring depth measurement data of the scene, which are obtained by combining images of a plurality of other LIDAR sensors, and which are used as a ground truth for the depth model of the scene.


