Dynamic Microphone Array Switching for Speaker Direction Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In videoconferencing settings, it is challenging for remote participants to determine which individual at a central venue is speaking due to the broad viewing angle required by cameras, which can result in displaying the entire conference room instead of focusing on the speaker.
Innovation Solution
The use of an array of directional microphones to automatically identify the direction of the speech source by switching between using audio streams from a subgroup of microphones and all microphones, allowing the display to show images centered on the speaker, thereby enhancing speaker identification and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If a single camera with broad viewing angle is used to capture all participants, then all participants are visible, but it becomes difficult to identify which person is speaking
Solution Approach 1:
The system segments the audio input by using multiple directional microphones to capture sound from different directions. Each microphone is associated with a specific spatial region, allowing the system to identify which participant is speaking by determining the direction of the sound source. This segmentation of audio capture by spatial direction resolves the contradiction by providing speaker identification while maintaining a broad viewing angle.
Solution Approach 2:
The patent introduces an intermediary processing system that receives audio signals from multiple directional microphones, determines the direction of the speech source, and uses this information to control camera selection or image processing. This intermediary layer bridges the gap between the broad viewing angle requirement and the need for speaker identification, allowing the system to maintain both characteristics.
2Loss of information
If multiple cameras are used to focus on the speaker, then speaker identification is improved, but the system complexity increases
Solution Approach 1:
The patent replaces the mechanical solution of using multiple cameras with an electronic/audio-based solution. Instead of requiring multiple camera systems to identify and focus on speakers, the system uses an array of directional microphones to determine speaker location and then processes images from a single camera or selects from multiple cameras based on audio direction data. This substitution reduces device complexity while maintaining speaker identification capability.
Solution Approach 2:
The system makes the audio system multi-functional by having the directional microphone array serve both audio capture and speaker identification functions. The same audio data used for meeting participation is also used to determine which participant is speaking, eliminating the need for separate identification hardware and reducing overall system complexity.
3Measurement precision
If audio streams from all directional microphones are used, then the direction identification accuracy is improved, but the processing complexity increases
Solution Approach 1:
The patent implements a dynamic approach to audio processing where the system adaptively selects which microphone data to use based on current acoustic conditions. Rather than always processing data from all microphones, the system dynamically adjusts the processing scope, using all microphones when needed for accuracy but reducing to subsets when conditions allow, thereby balancing precision with processing complexity.
Solution Approach 2:
The system changes processing parameters dynamically based on acoustic environment assessment. When the acoustic conditions favor using all microphones for maximum accuracy, the system processes all audio streams. When conditions change or full processing is unnecessary, the system adjusts parameters to use only the necessary subset of microphone data, thus maintaining accuracy while controlling processing complexity through parameter adjustment.
Data Source
AI summary
This disclosure describes techniques of automatically identifying a direction of a speech source relative to an array of directional microphones using audio streams from some or all of the directional microphones. Whether the direction of the speech source is identified using audio streams from some of the directional microphones or from all of the directional microphones depends on whether using audio streams from a subgroup of the directional microphones or using audio streams from all of the directional microphones is more likely to correctly identify the direction of the speech source. Switching between using audio streams from some of the directional microphones and using audio streams from all of the directional microphones may occur automatically to best identify the direction of the speech source. A display screen at a remote venue may then display images having angles of view that are centered generally in the direction of the speech source.


