Multi-modal Sensor Fusion for Dense 3D Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sensors, such as LIDAR and cameras, provide sparse information at longer distances, limiting the creation of dense three-dimensional navigational maps for self-driving vehicles, which can lead to inadequate recognition of objects and increased risk of accidents, especially at higher speeds.

Innovation Solution

Integrating multi-modal and multi-dimensional sensor information from various sources, including 3D point clouds and 2D images, to generate a real-time, high-fidelity 3D navigational map through sensor fusion and deep learning, incorporating temporal and geo-referenced data to enhance semantic segmentation and depth estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the vehicle travels a given course several times to densely map the course by cameras, then the mapping density is improved, but the vehicle is constrained to only navigating known paths and cannot respond to ad hoc objects

Engineering Contradiction:
Improvemapping densityVSAvoidadaptability to ad hoc objects
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent combines multiple sensor modalities (camera and LIDAR) into a unified perception system. The camera provides dense visual information while LIDAR provides accurate depth and 3D spatial information, merging their strengths to achieve both dense mapping and real-time adaptability to unknown objects without requiring repeated traversals of the same path.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from 2D camera images to 3D point cloud representations by incorporating depth information from LIDAR. This dimensional transformation enables the system to perceive and map the environment in three-dimensional space, allowing real-time detection and response to ad hoc objects at various distances without constraining the vehicle to pre-mapped paths.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If sensors provide dense information close to the vehicle, then the recognition accuracy is improved, but the information becomes sparse at long range

Engineering Contradiction:
Improveobject recognition accuracyVSAvoiddetection range
Core Design Contradiction:
Measurement precisionVSLength of stationary object

Solution Approach 1:

The patent merges camera-based visual recognition (excellent for close-range dense information) with LIDAR-based depth sensing (excellent for long-range 3D mapping). This combination allows the system to maintain high recognition accuracy for nearby objects while extending detection capability to distant objects through LIDAR's point cloud coverage.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses LIDAR point clouds as an intermediary that bridges the gap between camera images and 3D spatial understanding. The point cloud representation serves as a mediator that provides depth information and 3D structure at long ranges, enabling the system to overcome the sparsity problem of camera-only systems at distance.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If the vehicle increases speed, then the productivity is improved, but the time to recognize objects and react is reduced

Engineering Contradiction:
Improvevehicle speedVSAvoidobject recognition and reaction time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements continuous multi-modal sensor fusion that operates continuously as the vehicle moves, rather than relying on discrete repeated passes. The system continuously integrates camera and LIDAR data to maintain an updated 3D environmental model, enabling real-time object detection and reaction even at high speeds without requiring the vehicle to slow down for mapping.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS10991156B2Multi-modal data fusion for enhanced 3D perception for platforms
Publication Date: 2021.04.27 SRI INTERNATIONAL
  • US10991156B2 patent drawing
  • US10991156B2 patent drawing
  • US10991156B2 patent drawing

AI summary

A method for providing a real time, three-dimensional (3D) navigational map for platforms includes integrating at least two sources of multi-modal and multi-dimensional platform sensor information to produce a more accurate 3D navigational map. The method receives both a 3D point cloud from a first sensor on a platform with a first modality and a 2D image from a second sensor on the platform with a second modality different from the first modality, generates a semantic label and a semantic label uncertainty associated with a first space point in the 3D point cloud, generates a semantic label and a semantic label uncertainty associated with a second space point in the 2D image, and fuses the first space semantic label and the first space semantic uncertainty with the second space semantic label and the second space semantic label uncertainty to create fused 3D spatial information to enhance the 3D navigational map.