Multi-modal Speech Localization via Camera-Microphone Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

State-of-the-art speech recognizers fail to reliably associate speech with the correct speaker in environments with multiple speakers, as they struggle to differentiate audio sources accurately.

Innovation Solution

A multi-modal approach using image data from cameras and audio data from microphone arrays, where frequency domain representations of audio and positioning information of human faces are combined to identify the source of sound through a previously-trained audio source localization classifier.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If state-of-the-art speech recognizers are used in multi-speaker environments, then speech recognition is performed, but the ability to reliably associate speech with the correct speaker deteriorates

Engineering Contradiction:
Improvespeech-to-speaker association reliabilityVSAvoidmulti-speaker environment handling
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent combines audio data from microphone arrays with image data from camera arrays to create a multi-modal system. The audio source localization classifier processes both audio frequency domain representations and visual face positioning information together, merging multiple data types to improve speaker identification reliability in multi-speaker environments

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an audio source localization classifier as an intermediary component that bridges audio data and visual data. This classifier takes audio frequency representations and face positioning information as inputs and produces speaker identification outputs, mediating between different data modalities to enable reliable speech-to-speaker association

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If audio data from microphone arrays is processed alone, then audio processing is performed, but the precision of audio source localization deteriorates

Engineering Contradiction:
Improveaudio source localization precisionVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system merges audio data processing with visual data processing by feeding both audio frequency domain representations and image-derived face positioning data into the same audio source localization classifier. This combination of multiple data sources improves localization precision while the integrated architecture manages complexity

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If face positioning data from camera arrays is integrated with audio data, then speaker identification accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidmulti-modal processing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio source localization classifier serves multiple functions: it processes audio frequency domain representations, integrates face positioning information from camera arrays, and produces speaker identification outputs. This multi-functional component handles both audio and visual data streams, improving speaker identification accuracy while managing system complexity through a unified processing approach

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10847162B2Multi-modal speech localization
Publication Date: 2020.11.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10847162B2 patent drawing
  • US10847162B2 patent drawing
  • US10847162B2 patent drawing

AI summary

Multi-modal speech localization is achieved using image data captured by one or more cameras, and audio data captured by a microphone array. Audio data captured by each microphone of the array is transformed to obtain a frequency domain representation that is discretized in a plurality of frequency intervals. Image data captured by each camera is used to determine a positioning of each human face. Input data is provided to a previously-trained, audio source localization classifier, including: the frequency domain representation of the audio data captured by each microphone, and the positioning of each human face captured by each camera in which the positioning of each human face represents a candidate audio source. An identified audio source is indicated by the classifier based on the input data that is estimated to be the human face from which the audio data originated.