Active Speaker Location Detection Using Multi-Array Audio and 3D Spatial Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conferencing systems face challenges in accurately identifying the active speaker, particularly due to varying microphone array placements and environmental changes, which affect the precision of speaker location determination.

Innovation Solution

The system employs a video conferencing device with a first and second microphone array, along with image capture devices, to generate a three-dimensional model of the room and determine the location of the second microphone array relative to the video conferencing device. This allows for accurate calculation of the active speaker's location using sound source localization distributions from both arrays, combined with image data to distinguish between potential speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single microphone array is used in the video conferencing device, then the device structure is simple, but the accuracy of active speaker location determination deteriorates when microphone array placement varies

Engineering Contradiction:
Improvemicrophone array configurationVSAvoidspeaker location determination accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system divides the audio detection function into multiple independent microphone arrays: a first microphone array in the video conferencing device and a second microphone array in a separate device. Each array independently captures audio data, and their results are combined to determine speaker location, improving accuracy without requiring a single complex array

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a three-dimensional model of the room environment as an intermediary that mediates between the multiple microphone arrays and the final speaker location determination. The model provides a common spatial reference frame that allows data from differently placed arrays to be accurately integrated

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple microphone arrays are used to improve speaker location accuracy, then measurement precision improves, but device complexity and system configuration difficulty increase

Engineering Contradiction:
Improvespeaker location determination accuracyVSAvoidmicrophone array configuration
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses a universal three-dimensional room model that serves multiple functions: it represents the physical environment, provides coordinate transformation references for different microphone arrays, and enables consistent speaker location determination across various array configurations. This universal model simplifies the integration of multiple arrays despite their different placements

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms the problem from directly comparing audio data from different arrays to first mapping all audio data into a common three-dimensional spatial reference frame defined by the room model. This parameter transformation (from array-specific coordinates to room-model coordinates) enables accurate integration despite varying array positions and orientations

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If traditional audio-only methods are used to identify active speakers, then the system is simple to operate, but accuracy deteriorates in environments with multiple potential speakers

Engineering Contradiction:
Improvesystem operation simplicityVSAvoidactive speaker identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system merges audio data from multiple microphone arrays with three-dimensional spatial information from the room model to create sound source localization distributions. This combination of audio signals with spatial context enables accurate identification of which participant is speaking, even when multiple people are present in the scene

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent replaces traditional audio-only speaker identification with a hybrid system that uses acoustic principles (sound source localization) combined with visual-spatial information from the three-dimensional room model. This substitution of pure audio processing with a multi-modal approach improves accuracy while maintaining operational simplicity through automated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution enables precise determination of the active speaker's location, even with changing microphone array positions and environmental configurations, ensuring accurate highlighting of the active speaker in video feeds.

Implementation Method 1

determine an estimated location in the three dimensional model of the active speaker using the audio data from the first microphone array, the audio data from the second microphone array, the location of the second microphone array, and an angular orientation of the second microphone array

Methodology Applied
Scientific EffectSound source localization: Acoustics

Data Source

PatentEP3400705B1Active speaker location detection
Publication Date: 2021.03.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3400705B1 patent drawingFigure 1
  • EP3400705B1 patent drawingFigure 2
  • EP3400705B1 patent drawingFigure 3

AI summary

Examples related to determining a location of an active speaker are provided. In one example, image data (54) of a room from an image capture device (52) is received and a three dimensional model (64) is generated. First audio data (26) from a first microphone array (24) at the image capture device is received. Second audio data (34) from a second microphone array (30) laterally spaced from the image capture device is received. A location (82) of the second microphone array is determined. Using the audio data and the location and angular orientation (68) of the second microphone array, an estimated location (84) of the active speaker is determined. Using the estimated location, a setting (90) for the image capture device is determined and outputted to highlight the active speaker.