Shared Vision Backbone for Dense LiDAR Depth Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems relying on LiDAR for 3D environment representation are costly and prone to errors in adverse weather conditions, with pseudo-LiDAR offering less accurate object detection due to aberrations and distortions when transforming 2D image data into 3D maps.

Innovation Solution

A method and apparatus that generate a dense LiDAR representation by fusing sparse depth estimates from multiple 2D representations, including RGB images, semantic maps, and radar images, using a shared backbone network to improve accuracy and robustness across different sensor configurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If LiDAR sensors are used to generate accurate 3D representations, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improve3D representation accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple 2D image representations from different sensors (RGB cameras, semantic maps, radar images) to generate a dense 3D LiDAR representation. By merging these multiple 2D inputs through a fusion network, the system achieves accurate 3D reconstruction without requiring a dedicated LiDAR sensor, thus resolving the contradiction between measurement precision and device complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared backbone network processes multiple types of 2D representations (RGB images, semantic maps, radar images) uniformly, enabling a single system to handle various sensor configurations and input types. This multi-functional approach allows the system to generate accurate 3D representations using existing multi-sensor setups without adding specialized LiDAR hardware

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If pseudo-LiDAR is used as an alternative to LiDAR, then device complexity is reduced, but measurement precision deteriorates due to aberrations and distortions

Engineering Contradiction:
Improvesensor system complexityVSAvoidobject detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

Instead of relying on a single pseudo-LiDAR transformation that introduces aberrations, the patent merges multiple 2D representations (RGB images, semantic maps, radar images) and processes them through a fusion network. This combination of multiple information sources compensates for the distortions inherent in pseudo-LiDAR transformations, maintaining measurement precision while avoiding complex LiDAR hardware

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The depth fusion network acts as an intermediary that processes and reconciles depth estimates from multiple 2D representations before generating the final 3D LiDAR representation. This intermediary processing step corrects aberrations and distortions that would otherwise be present in direct pseudo-LiDAR transformations, improving measurement precision without requiring actual LiDAR sensors

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If multiple 2D representations are fused to generate dense LiDAR, then measurement precision is improved, but computational overhead increases

Engineering Contradiction:
Improve3D representation accuracyVSAvoidcomputational overhead
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The processing system is segmented into distinct components: a shared backbone network that extracts features from multiple 2D representations, and a depth fusion network that combines these features. This segmentation allows efficient reuse of the backbone network across different input types and reduces redundant computations, improving processing efficiency while maintaining the ability to fuse multiple representations for high precision

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The shared backbone network performs preliminary feature extraction from all 2D representations (RGB images, semantic maps, radar images) before the depth fusion network combines them. By pre-processing and extracting essential features in advance, the system reduces the computational burden during the fusion stage, enabling accurate dense LiDAR generation with optimized computational overhead

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12148223B2Shared vision system backbone
Publication Date: 2024.11.19 TOYOTA JIDOSHA KK
  • US12148223B2 patent drawing
  • US12148223B2 patent drawing
  • US12148223B2 patent drawing

AI summary

A method for generating a dense light detection and ranging (LiDAR) representation by a vision system includes receiving, at a sparse depth network, one or more sparse representations of an environment. The method also includes generating a depth estimate of the environment depicted in an image captured by an image capturing sensor. The method further includes generating, via the sparse depth network, one or more sparse depth estimates based on receiving the one or more sparse representations. The method also includes fusing the depth estimate and the one or more sparse depth estimates to generate a dense depth estimate. The method further includes generating the dense LiDAR representation based on the dense depth estimate and controlling an action of the vehicle based on identifying a three-dimensional object in the dense LiDAR representation.