Multi-Modal Object Recognition Using Audio Visual Keypoint Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current object recognition techniques in complex scenes, such as music performances, struggle to accurately identify and decompose multiple sound sources due to reliance on sequential recording and close microphone placement, limiting spatial effect control and user experience.

Innovation Solution

The integration of audio and visual information using keypoint selection and matching devices to recognize objects, employing multi-modal Bayesian estimation and acoustic feature analysis, allowing for simultaneous identification and localization of instruments in a scene.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sequential recording and close microphone placement are used, then sound source identification can be achieved, but spatial effect control is limited and device complexity increases

Engineering Contradiction:
Improvesound source identification accuracyVSAvoidmicrophone placement requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges audio and visual information processing to achieve object recognition in music performances. By combining audio feature extraction with visual keypoint detection and matching, the system can identify instruments and spatial effects without requiring complex microphone placements, thus reducing device complexity while maintaining identification accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces visual information as an intermediary to bridge audio processing and object recognition. Visual keypoints and descriptors serve as mediators that connect audio features with spatial location, enabling the system to infer instrument positions and spatial effects without direct acoustic measurement complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If visual information only is used for object recognition, then processing is simple, but recognition accuracy in complex scenes is insufficient

Engineering Contradiction:
Improveprocessing complexityVSAvoidobject recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges audio and visual information processing streams. Audio features are extracted and matched with visual keypoints and descriptors to achieve more accurate object recognition in complex scenes, while the integrated system manages processing complexity through coordinated multi-modal analysis

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite recognition system that integrates multiple data types (audio features, visual keypoints, visual descriptors). This composite approach leverages the complementary strengths of each modality to achieve higher recognition accuracy than single-modality systems while managing overall system complexity

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS9495591B2Object recognition using multi-modal matching scheme
Publication Date: 2016.11.15 QUALCOMM INC
  • US9495591B2 patent drawing
  • US9495591B2 patent drawing
  • US9495591B2 patent drawing

AI summary

Methods, systems and articles of manufacture for recognizing and locating one or more objects in a scene are disclosed. An image and/or video of the scene are captured. Using audio recorded at the scene, an object search of the captured scene is narrowed down. For example, the direction of arrival (DOA) of a sound can be determined and used to limit the search area in a captured image/video. In another example, keypoint signatures may be selected based on types of sounds identified in the recorded audio. A keypoint signature corresponds to a particular object that the system is configured to recognize. Objects in the scene may then be recognized using a shift invariant feature transform (SIFT) analysis comparing keypoints identified in the captured scene to the selected keypoint signatures.