Active Speaker Location Detection Using Multi-Array Audio and 3D Spatial Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing systems face challenges in accurately identifying the active speaker, particularly due to varying microphone array placements and environmental changes, which affect the precision of speaker location determination.
Innovation Solution
The system employs a video conferencing device with a first and second microphone array, along with image capture devices, to generate a three-dimensional model of the room and determine the location of the second microphone array relative to the video conferencing device. This allows for accurate calculation of the active speaker's location using sound source localization distributions from both arrays, combined with image data to distinguish between potential speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single microphone array is used in the video conferencing device, then the device structure is simple, but the accuracy of active speaker location determination deteriorates when microphone array placement varies
Solution Approach 1:
The system divides the audio detection function into multiple independent microphone arrays: a first microphone array in the video conferencing device and a second microphone array in a separate device. Each array independently captures audio data, and their results are combined to determine speaker location, improving accuracy without requiring a single complex array
Solution Approach 2:
The patent introduces a three-dimensional model of the room environment as an intermediary that mediates between the multiple microphone arrays and the final speaker location determination. The model provides a common spatial reference frame that allows data from differently placed arrays to be accurately integrated
2Measurement precision
If multiple microphone arrays are used to improve speaker location accuracy, then measurement precision improves, but device complexity and system configuration difficulty increase
Solution Approach 1:
The system uses a universal three-dimensional room model that serves multiple functions: it represents the physical environment, provides coordinate transformation references for different microphone arrays, and enables consistent speaker location determination across various array configurations. This universal model simplifies the integration of multiple arrays despite their different placements
Solution Approach 2:
The patent transforms the problem from directly comparing audio data from different arrays to first mapping all audio data into a common three-dimensional spatial reference frame defined by the room model. This parameter transformation (from array-specific coordinates to room-model coordinates) enables accurate integration despite varying array positions and orientations
3Ease of operation
If traditional audio-only methods are used to identify active speakers, then the system is simple to operate, but accuracy deteriorates in environments with multiple potential speakers
Solution Approach 1:
The system merges audio data from multiple microphone arrays with three-dimensional spatial information from the room model to create sound source localization distributions. This combination of audio signals with spatial context enables accurate identification of which participant is speaking, even when multiple people are present in the scene
Solution Approach 2:
The patent replaces traditional audio-only speaker identification with a hybrid system that uses acoustic principles (sound source localization) combined with visual-spatial information from the three-dimensional room model. This substitution of pure audio processing with a multi-modal approach improves accuracy while maintaining operational simplicity through automated processing
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution enables precise determination of the active speaker's location, even with changing microphone array positions and environmental configurations, ensuring accurate highlighting of the active speaker in video feeds.
Implementation Method 1
determine an estimated location in the three dimensional model of the active speaker using the audio data from the first microphone array, the audio data from the second microphone array, the location of the second microphone array, and an angular orientation of the second microphone array
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Examples related to determining a location of an active speaker are provided. In one example, image data (54) of a room from an image capture device (52) is received and a three dimensional model (64) is generated. First audio data (26) from a first microphone array (24) at the image capture device is received. Second audio data (34) from a second microphone array (30) laterally spaced from the image capture device is received. A location (82) of the second microphone array is determined. Using the audio data and the location and angular orientation (68) of the second microphone array, an estimated location (84) of the active speaker is determined. Using the estimated location, a setting (90) for the image capture device is determined and outputted to highlight the active speaker.