Volumetric Descriptors for Multi-Modal Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in generating recognition descriptors that can effectively capture representations of an object across variations in a dimension of interest, such as time, frequency, wavelength, or other parameters, using a unified modal descriptor.

Innovation Solution

The development of a multi-modal sensitive recognition system that uses a modal recognition algorithm to derive multidimensional modal descriptors. This system configures a device to initiate actions based on these descriptors, improving automated diagnostics and reactions to changes over time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If separate descriptor sets are compiled for distinct images or representations of the same object, then each representation can be accurately characterized, but the system complexity increases and a unified representation across variations is not achieved

Engineering Contradiction:
Improveunified representation across variationsVSAvoiddescriptor set management
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple separate descriptor sets into a single unified volumetric descriptor that captures object representations across multiple variations simultaneously. This is achieved by organizing descriptors from different images or representations into a three-dimensional structure where the object's features are represented in a unified manner, reducing the number of separate descriptor sets that need to be managed while maintaining comprehensive object characterization.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from two-dimensional descriptor representations to three-dimensional volumetric descriptors. By adding a temporal or variation dimension to the descriptor structure, the system can represent object variations across multiple images or representations within a single unified framework, enabling comprehensive object characterization without increasing system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If traditional feature extraction algorithms are used for digital images, then feature extraction is straightforward, but the ability to capture object representations across multiple modalities and variations is limited

Engineering Contradiction:
Improvemulti-modality recognitionVSAvoiddescriptor generation complexity
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent creates a universal volumetric descriptor framework that can handle multiple modalities (visual, acoustic, tactile, etc.) and various object variations simultaneously. This multi-functional descriptor structure allows the same representation method to be applied across different sensing modalities and object states, enhancing adaptability while providing a systematic approach to managing the complexity of multi-modality recognition.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12205336B2Volumetric descriptors
Publication Date: 2025.01.21 NANT HOLDINGS IP LLC
  • US12205336B2 patent drawing
  • US12205336B2 patent drawing
  • US12205336B2 patent drawing

AI summary

Techniques are provided for multi-modal sensitive recognition. A digital data set for an object is obtained according to a modality, where the digital data set includes digital representations of the object at different values of a dimension of relevance of the modality. A reference location associated with the object is identified. A modal descriptor is derived for the modality according to an implementation of a multi-modal recognition algorithm by deriving a set of feature descriptors for the reference location and at the different values of the corresponding dimension of relevance, calculating a set of differences between the feature descriptors in the set of feature descriptors, and aggregating the set of differences into the modal descriptor. A device is then configured to initiate an action as a function of the modal descriptor.