Multi-Attention Sensor Fusion for Vehicle Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to effectively combine and process sensor data from different modalities, such as audio, image, and lidar, to improve object detection and classification in autonomous vehicles, as existing solutions do not leverage correlations between these modalities for enhanced detection and classification.

Innovation Solution

A machine-learning pipeline utilizing a multi-attention component correlates features across different sensor modalities, such as audio and image data, to enhance object detection and classification by drawing attention to relevant features in other modalities, thereby improving detection accuracy and confidence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If sensor data from different modalities is processed independently, then processing complexity is reduced, but detection accuracy and information completeness deteriorate

Engineering Contradiction:
Improveprocessing complexityVSAvoiddetection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges sensor data from multiple modalities (audio, image, lidar) into a unified processing framework using a machine-learned transformer model. The model processes all modalities simultaneously and enables cross-modal attention mechanisms, allowing the system to achieve high detection accuracy without independently processing each modality, thus resolving the contradiction between processing complexity and detection accuracy.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If sensor data from different modalities is combined for processing, then detection accuracy improves, but processing complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a machine-learned transformer model as an intermediary that handles the complex task of multi-modal data integration. This intermediary absorbs the processing complexity through its internal attention mechanisms and feature correlation capabilities, while presenting simplified, accurate detection results to the downstream system, thus resolving the contradiction between detection accuracy and processing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If features from one sensor modality are used to improve detections from another modality, then detection confidence improves, but computational requirements increase

Engineering Contradiction:
Improvedetection confidenceVSAvoidcomputational requirements
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The patent implements cross-modal attention mechanisms that selectively process only the most relevant features between modalities rather than exhaustively analyzing all possible feature combinations. The attention mechanism dynamically identifies and processes only the necessary cross-modal correlations needed for improved detection confidence, avoiding unnecessary computational overhead while maintaining high reliability.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250259454A1Multi-attention machine learning for object detection and classification
Publication Date: 2025.08.14 ZOOX INC
  • US20250259454A1 patent drawing
  • US20250259454A1 patent drawing
  • US20250259454A1 patent drawing

AI summary

Techniques for detecting, locating, and/or classifying objects based on multiple sensor data inputs received from different sensor modalities. The techniques may include receiving sensor data generated by different sensor modalities of a vehicle, the sensor data including at least first sensor data generated by a first sensor modality and second sensor data generated by a second sensor modality. In some examples, the sensor data may be input into a machine-learning pipeline. The machine-learning pipeline may be configured to determine locations of objects in an environment surrounding the vehicle based at least in part on a correlation, by the multi-attention component, of the first sensor data and the second sensor data. The techniques may also include receiving, from the machine-learning pipeline, an output indicating a location of an object in the environment.