Shared Vision Backbone for Dense LiDAR Depth Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems relying on LiDAR for 3D environment representation are costly and prone to errors in adverse weather conditions, with pseudo-LiDAR offering less accurate object detection due to aberrations and distortions when transforming 2D image data into 3D maps.
Innovation Solution
A method and apparatus that generate a dense LiDAR representation by fusing sparse depth estimates from multiple 2D representations, including RGB images, semantic maps, and radar images, using a shared backbone network to improve accuracy and robustness across different sensor configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If LiDAR sensors are used to generate accurate 3D representations, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent combines multiple 2D image representations from different sensors (RGB cameras, semantic maps, radar images) to generate a dense 3D LiDAR representation. By merging these multiple 2D inputs through a fusion network, the system achieves accurate 3D reconstruction without requiring a dedicated LiDAR sensor, thus resolving the contradiction between measurement precision and device complexity
Solution Approach 2:
The shared backbone network processes multiple types of 2D representations (RGB images, semantic maps, radar images) uniformly, enabling a single system to handle various sensor configurations and input types. This multi-functional approach allows the system to generate accurate 3D representations using existing multi-sensor setups without adding specialized LiDAR hardware
2Device complexity
If pseudo-LiDAR is used as an alternative to LiDAR, then device complexity is reduced, but measurement precision deteriorates due to aberrations and distortions
Solution Approach 1:
Instead of relying on a single pseudo-LiDAR transformation that introduces aberrations, the patent merges multiple 2D representations (RGB images, semantic maps, radar images) and processes them through a fusion network. This combination of multiple information sources compensates for the distortions inherent in pseudo-LiDAR transformations, maintaining measurement precision while avoiding complex LiDAR hardware
Solution Approach 2:
The depth fusion network acts as an intermediary that processes and reconciles depth estimates from multiple 2D representations before generating the final 3D LiDAR representation. This intermediary processing step corrects aberrations and distortions that would otherwise be present in direct pseudo-LiDAR transformations, improving measurement precision without requiring actual LiDAR sensors
3Measurement precision
If multiple 2D representations are fused to generate dense LiDAR, then measurement precision is improved, but computational overhead increases
Solution Approach 1:
The processing system is segmented into distinct components: a shared backbone network that extracts features from multiple 2D representations, and a depth fusion network that combines these features. This segmentation allows efficient reuse of the backbone network across different input types and reduces redundant computations, improving processing efficiency while maintaining the ability to fuse multiple representations for high precision
Solution Approach 2:
The shared backbone network performs preliminary feature extraction from all 2D representations (RGB images, semantic maps, radar images) before the depth fusion network combines them. By pre-processing and extracting essential features in advance, the system reduces the computational burden during the fusion stage, enabling accurate dense LiDAR generation with optimized computational overhead
Data Source
AI summary
A method for generating a dense light detection and ranging (LiDAR) representation by a vision system includes receiving, at a sparse depth network, one or more sparse representations of an environment. The method also includes generating a depth estimate of the environment depicted in an image captured by an image capturing sensor. The method further includes generating, via the sparse depth network, one or more sparse depth estimates based on receiving the one or more sparse representations. The method also includes fusing the depth estimate and the one or more sparse depth estimates to generate a dense depth estimate. The method further includes generating the dense LiDAR representation based on the dense depth estimate and controlling an action of the vehicle based on identifying a three-dimensional object in the dense LiDAR representation.


