Flash LiDAR Camera Fusion for High-Resolution 3D Depth Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current object detection methods for autonomous driving face challenges in recognizing and classifying objects in 3D spaces with multiple objects close by, particularly due to the limitations of camera-only, LIDAR-only, and simple fusion approaches, which result in inaccurate depth models and artifacts.
Innovation Solution
A method for training an artificial intelligence module using multiple sets of training data comprising camera images and depth images from flash LIDAR sensors, where depth measurement data from combined flash LIDAR sensors is used as ground truth to produce a high-resolution depth model, overcoming limitations of low-resolution LIDAR and camera-based depth completion methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If camera-based depth completion methods are used, then depth model can be obtained from low-resolution depth image, but multiple artifacts including missing patches from motion of objects and dark or light areas on the edges of camera dynamic range occur
Solution Approach 1:
The patent combines data from multiple sensors (camera, flash LIDAR, and optionally other LIDAR sensors) to create a fused depth model. The fusion architecture integrates features from camera images and depth images at multiple levels, leveraging the strengths of each sensor modality to compensate for their individual weaknesses and produce a more reliable, artifact-free depth model.
Solution Approach 2:
The patent introduces an intermediate processing stage where features from camera and LIDAR sensors are separately extracted and then fused. This intermediary feature fusion layer acts as a mediator that combines the complementary information from both sensors before generating the final depth model, reducing artifacts from motion and dynamic range issues.
2Device complexity
If low-resolution LIDAR sensor is used, then device complexity is reduced, but LIDAR based object detection performs poorly
Solution Approach 1:
The patent merges data from multiple LIDAR sensors (including flash LIDAR and potentially other types) to effectively increase the resolution and quality of depth information. By combining measurements from multiple sensors with different characteristics, the system achieves high-resolution depth modeling without requiring a single high-resolution LIDAR sensor, thus maintaining lower device complexity.
Solution Approach 2:
The patent transitions from relying on a single LIDAR sensor's spatial resolution to utilizing multiple dimensions of data fusion - combining spatial information from multiple sensors, temporal information from sequential measurements, and complementary information from different sensor modalities (camera and LIDAR) to achieve high-resolution depth models.
3Device complexity
If simple concatenation based fusion approaches are used, then implementation is simplified, but performance is capped by information contained in individual sensors
Solution Approach 1:
The patent segments the fusion process into distinct stages: separate feature extraction from camera and LIDAR sensors, intermediate feature fusion at multiple levels, and final depth model generation. This segmented approach allows each component to be optimized independently while achieving superior overall performance compared to simple concatenation.
Solution Approach 2:
The patent moves beyond simple feature concatenation to multi-level feature fusion that operates at different hierarchical levels of the neural network. This dimensional expansion in the fusion process allows integration of features at various abstraction levels, capturing both low-level details and high-level semantic information for improved depth model quality.
4Measurement precision
If multiple flash LIDAR sensors are combined to obtain depth measurement data, then higher resolution depth model is achieved, but device complexity increases
Solution Approach 1:
The patent segments the processing of multiple LIDAR sensors into separate feature extraction streams that are then fused at intermediate layers. This segmentation allows the system to handle multiple sensor inputs efficiently without requiring complex point-to-point correspondence algorithms, reducing computational complexity while maintaining high resolution.
Solution Approach 2:
The patent uses the same flash LIDAR sensor multiple times in different spatial positions to effectively create multiple copies of the sensing capability. Rather than requiring fundamentally different sensor types or complex calibration systems, the approach replicates the proven flash LIDAR technology in multiple locations, achieving high resolution through geometric diversity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The method enhances the accuracy and resolution of depth models, reducing motion artifacts and self-occlusion, and allows for more robust object detection by using a time-synchronized multi-flash LIDAR setup, enabling improved performance in complex scenes.
Implementation Method 1
a depth image of the scene taken by a flash LIDAR sensor of the vehicle
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method for training an artificial intelligence module for an advanced driver assistance system for obtaining a depth model of a scene. The artificial intelligence module uses data of a camera image and a depth image of a flash LIDAR sensor as input, and is trained for outputting the depth model of the scene, wherein the depth model has a higher resolution than the depth image of the flash LIDAR sensor. The method of training the artificial intelligence module comprises acquiring depth measurement data of the scene, which are obtained by combining images of a plurality of flash LIDAR sensors, and which are used as a ground truth for the depth model of the scene.