Multi-modal Speech Localization via Camera-Microphone Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
State-of-the-art speech recognizers fail to reliably associate speech with the correct speaker in environments with multiple speakers, as they struggle to differentiate audio sources accurately.
Innovation Solution
A multi-modal approach using image data from cameras and audio data from microphone arrays, where frequency domain representations of audio and positioning information of human faces are combined to identify the source of sound through a previously-trained audio source localization classifier.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If state-of-the-art speech recognizers are used in multi-speaker environments, then speech recognition is performed, but the ability to reliably associate speech with the correct speaker deteriorates
Solution Approach 1:
The patent combines audio data from microphone arrays with image data from camera arrays to create a multi-modal system. The audio source localization classifier processes both audio frequency domain representations and visual face positioning information together, merging multiple data types to improve speaker identification reliability in multi-speaker environments
Solution Approach 2:
The patent introduces an audio source localization classifier as an intermediary component that bridges audio data and visual data. This classifier takes audio frequency representations and face positioning information as inputs and produces speaker identification outputs, mediating between different data modalities to enable reliable speech-to-speaker association
2Measurement precision
If audio data from microphone arrays is processed alone, then audio processing is performed, but the precision of audio source localization deteriorates
Solution Approach 1:
The system merges audio data processing with visual data processing by feeding both audio frequency domain representations and image-derived face positioning data into the same audio source localization classifier. This combination of multiple data sources improves localization precision while the integrated architecture manages complexity
3Measurement precision
If face positioning data from camera arrays is integrated with audio data, then speaker identification accuracy is improved, but system complexity increases
Solution Approach 1:
The audio source localization classifier serves multiple functions: it processes audio frequency domain representations, integrates face positioning information from camera arrays, and produces speaker identification outputs. This multi-functional component handles both audio and visual data streams, improving speaker identification accuracy while managing system complexity through a unified processing approach
Data Source
AI summary
Multi-modal speech localization is achieved using image data captured by one or more cameras, and audio data captured by a microphone array. Audio data captured by each microphone of the array is transformed to obtain a frequency domain representation that is discretized in a plurality of frequency intervals. Image data captured by each camera is used to determine a positioning of each human face. Input data is provided to a previously-trained, audio source localization classifier, including: the frequency domain representation of the audio data captured by each microphone, and the positioning of each human face captured by each camera in which the positioning of each human face represents a candidate audio source. An identified audio source is indicated by the classifier based on the input data that is estimated to be the human face from which the audio data originated.


