Multi-Modal Sensor Fusion for Cooperative Perception Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional cooperative perception systems face challenges in real-time data transmission of raw sensing data between vehicles, leading to decreased detection accuracy due to limitations in detection systems, where objects may not be detected by either vehicle.

Innovation Solution

The system employs intermediate fusion by transmitting features extracted from raw sensing data between vehicles, allowing each vehicle to locally extract features and fuse them without requiring additional computation overhead or raw data transmission, using a scale-fusion network to combine features in a bird's eye view format.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If raw sensing data is transmitted between vehicles for cooperative perception, then detection accuracy can be improved, but data transmission bandwidth requirements increase significantly

Engineering Contradiction:
Improvedetection accuracyVSAvoiddata transmission bandwidth
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential feature representations from raw sensing data at each vehicle, rather than transmitting the complete raw data. This extraction process isolates the critical information needed for cooperative detection while discarding redundant data, thereby reducing transmission bandwidth requirements while preserving detection accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the cooperative perception system into independent feature extraction modules at each vehicle, which locally process raw data and transmit only the extracted features. This segmentation allows distributed processing and reduces the need for centralized raw data transmission, achieving both accuracy and bandwidth efficiency.

Inventive Principle:
Principle #1Segmentation

2Reliability

If prediction results from multiple vehicles are combined, then detection reliability can be improved, but detection accuracy decreases when objects are not detected by either vehicle

Engineering Contradiction:
Improvedetection reliabilityVSAvoiddetection accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary feature extraction and object detection at each vehicle before combining results. By pre-processing the data locally and extracting meaningful features in advance, the system ensures that only relevant detection information is transmitted and combined, improving both reliability and accuracy of the final cooperative detection results.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If feature-based intermediate fusion is used instead of raw data transmission, then data transmission bandwidth is reduced, but computation overhead at each vehicle increases

Engineering Contradiction:
Improvedata transmission bandwidthVSAvoidcomputation overhead
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent changes the parameter representation from raw sensing data to extracted feature representations. This parameter transformation reduces the dimensionality and complexity of transmitted data while maintaining the essential information needed for detection, thereby reducing bandwidth requirements without significantly increasing computation overhead at each vehicle.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250104409A1Methods and systems for fusing multi-modal sensor data
Publication Date: 2025.03.27 TOYOTA MOTOR ENG & MFG NORTH AMERICA INC
  • US20250104409A1 patent drawing
  • US20250104409A1 patent drawing
  • US20250104409A1 patent drawing

AI summary

A method of fusing multi-modal sensor data is provided. The method includes obtaining features for 3D data captured by a first sensor of an ego vehicle, obtaining features for images captured by second sensors of the ego vehicle, flattening the features for 3D data to first features in bird eye view, transforming the features for images into second features in bird eye view, concatenating the first features and the second features to obtain first concatenated multi-sensor features, and fusing the first concatenated multi-sensor features with second concatenated multi-sensor features received from another vehicle to obtain fused multi-sensor features.