Multimodal Speaker Identification via Early Audio-Video Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker identification systems often rely on combining audio and video inputs at the end of the detection process, which can limit accuracy and efficiency in identifying people or speakers, particularly in complex environments.

Innovation Solution

The system integrates multiple types of input, including audio and video, from the beginning of the detection process, using a classifier generated from a pool of features to efficiently identify regions where people or speakers might exist, incorporating techniques like sound source localization and feature selection to enhance detection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio and video inputs are combined at the end of the detection process (decision level fusion), then the system can process multiple input types, but the accuracy and efficiency of speaker identification is limited

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoiddetection efficiency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by integrating audio and video inputs at the beginning of the detection process rather than at the end. The system performs modality-specific processing separately for audio and video streams, generating intermediate detection results that are then fused. This early integration allows the system to leverage complementary information from both modalities throughout the detection pipeline, improving both accuracy and efficiency by avoiding redundant processing and enabling earlier elimination of non-speaker regions.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If multiple types of input (audio and video) are processed separately and then combined, then the system can maintain simplicity in processing each modality, but the overall detection accuracy in complex environments deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoiddetection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the detection process into separate modality-specific processing streams for audio and video inputs. Each stream processes its respective input type independently using optimized algorithms tailored to that modality's characteristics. The audio stream processes sound waveforms and spectral features, while the video stream processes visual frames and motion detection. These segmented streams are then fused at the decision level, allowing the system to maintain processing simplicity for each individual modality while achieving high detection accuracy through their coordinated integration.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8024189B2Identification of people using multiple types of input
Publication Date: 2011.09.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8024189B2 patent drawing
  • US8024189B2 patent drawing
  • US8024189B2 patent drawing

AI summary

Systems and methods for detecting people or speakers in an automated fashion are disclosed. A pool of features including more than one type of input (like audio input and video input) may be identified and used with a learning algorithm to generate a classifier that identifies people or speakers. The resulting classifier may be evaluated to detect people or speakers.