Dynamic Microphone Array Switching for Speaker Direction Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In videoconferencing settings, it is challenging for remote participants to determine which individual at a central venue is speaking due to the broad viewing angle required by cameras, which can result in displaying the entire conference room instead of focusing on the speaker.

Innovation Solution

The use of an array of directional microphones to automatically identify the direction of the speech source by switching between using audio streams from a subgroup of microphones and all microphones, allowing the display to show images centered on the speaker, thereby enhancing speaker identification and visualization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If a single camera with broad viewing angle is used to capture all participants, then all participants are visible, but it becomes difficult to identify which person is speaking

Engineering Contradiction:
Improveviewing angleVSAvoidspeaker identification
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The system segments the audio input by using multiple directional microphones to capture sound from different directions. Each microphone is associated with a specific spatial region, allowing the system to identify which participant is speaking by determining the direction of the sound source. This segmentation of audio capture by spatial direction resolves the contradiction by providing speaker identification while maintaining a broad viewing angle.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing system that receives audio signals from multiple directional microphones, determines the direction of the speech source, and uses this information to control camera selection or image processing. This intermediary layer bridges the gap between the broad viewing angle requirement and the need for speaker identification, allowing the system to maintain both characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If multiple cameras are used to focus on the speaker, then speaker identification is improved, but the system complexity increases

Engineering Contradiction:
Improvespeaker identificationVSAvoidnumber of cameras
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical solution of using multiple cameras with an electronic/audio-based solution. Instead of requiring multiple camera systems to identify and focus on speakers, the system uses an array of directional microphones to determine speaker location and then processes images from a single camera or selects from multiple cameras based on audio direction data. This substitution reduces device complexity while maintaining speaker identification capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system makes the audio system multi-functional by having the directional microphone array serve both audio capture and speaker identification functions. The same audio data used for meeting participation is also used to determine which participant is speaking, eliminating the need for separate identification hardware and reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If audio streams from all directional microphones are used, then the direction identification accuracy is improved, but the processing complexity increases

Engineering Contradiction:
Improvedirection identification accuracyVSAvoidaudio processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a dynamic approach to audio processing where the system adaptively selects which microphone data to use based on current acoustic conditions. Rather than always processing data from all microphones, the system dynamically adjusts the processing scope, using all microphones when needed for accuracy but reducing to subsets when conditions allow, thereby balancing precision with processing complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters dynamically based on acoustic environment assessment. When the acoustic conditions favor using all microphones for maximum accuracy, the system processes all audio streams. When conditions change or full processing is unnecessary, the system adjusts parameters to use only the necessary subset of microphone data, thus maintaining accuracy while controlling processing complexity through parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8130978B2Dynamic switching of microphone inputs for identification of a direction of a source of speech sounds
Publication Date: 2012.03.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8130978B2 patent drawing
  • US8130978B2 patent drawing
  • US8130978B2 patent drawing

AI summary

This disclosure describes techniques of automatically identifying a direction of a speech source relative to an array of directional microphones using audio streams from some or all of the directional microphones. Whether the direction of the speech source is identified using audio streams from some of the directional microphones or from all of the directional microphones depends on whether using audio streams from a subgroup of the directional microphones or using audio streams from all of the directional microphones is more likely to correctly identify the direction of the speech source. Switching between using audio streams from some of the directional microphones and using audio streams from all of the directional microphones may occur automatically to best identify the direction of the speech source. A display screen at a remote venue may then display images having angles of view that are centered generally in the direction of the speech source.