Physics-Guided Deep Multimodal Embeddings for Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sensor fusion methods are limited to early-stage fusion of similar sensor types and fail to effectively combine complementary information from different types of sensors, leading to suboptimal robustness and accuracy in tasks like target detection and recognition.

Innovation Solution

A method and apparatus for sensor data fusion using a common embedding space, where features from multiple sensor modalities are projected and combined, constrained by physics properties, enabling late-stage fusion and attention-based mode fusion to enhance robustness and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If early stage fusion of similar sensor types is used, then processing simplicity is maintained, but fusion of complementary information from different sensor types cannot be achieved

Engineering Contradiction:
Improveprocessing simplicityVSAvoidfusion of complementary information
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces a common embedding space as an intermediary representation that bridges different sensor modalities. Each sensor type has its own neural network that projects features into this shared space, enabling fusion of complementary information while maintaining processing simplicity through a unified framework.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The common embedding space serves as a universal representation that can accommodate multiple sensor types (acoustic, electromagnetic, mechanical). This universal space allows different sensor modalities to be fused while maintaining a consistent processing approach across diverse sensor types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If hand-crafted features or deep-learned features from single data source are used, then task performance is achieved, but robustness and accuracy are suboptimal

Engineering Contradiction:
Improvetask performanceVSAvoidrobustness and accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent combines features from multiple sensor modalities (acoustic, electromagnetic, mechanical) into a composite representation in the common embedding space. This composite approach integrates complementary information from different physical domains, enhancing robustness and accuracy beyond what single sensor types can achieve.

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The patent merges feature representations from multiple independent sensor systems into a unified embedding space. By combining acoustic features, electromagnetic features, and mechanical features in a common space, the system achieves improved robustness and accuracy through synergistic integration.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If different sensor types with diverse physical characteristics are fused, then complementary information is combined, but data heterogeneity increases processing complexity

Engineering Contradiction:
Improvecomplementary information combinationVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transforms heterogeneous sensor data into a unified parameter space (the common embedding space). Each sensor type's diverse physical characteristics are mapped to standardized vector representations, changing the parameter domain from modality-specific to modality-agnostic, thereby reducing processing complexity while preserving complementary information.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12412091B2Physics-guided deep multimodal embeddings for task-specific data exploitation
Publication Date: 2025.09.09 SRI INTERNATIONAL
  • US12412091B2 patent drawing
  • US12412091B2 patent drawing
  • US12412091B2 patent drawing

AI summary

A method, apparatus and system for object detection in sensor data having at least two modalities using a common embedding space includes creating first modality vector representations of features of sensor data having a first modality and second modality vector representations of features of sensor data having a second modality, projecting the first and second modality vector representations into the common embedding space such that related embedded modality vectors are closer together in the common embedding space than unrelated modality vectors, combining the projected first and second modality vector representations, and determining a similarity between the combined modality vector representations and respective embedded vector representations of features of objects in the common embedding space to identify at least one object depicted by the captured sensor data. In some instances, data manipulation of the method, apparatus and system can be guided by physics properties of a sensor and/or sensor data.