Microphone Array Speaker Selection for Teleconferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In teleconferencing environments, existing systems fail to effectively separate and transmit speech signals from multiple active speakers, leading to low quality and intelligibility due to signal interference and noise, and lack the ability to selectively identify and transmit the speech of a desired speaker.

Innovation Solution

A method and apparatus using a microphone array module for signal separation, a speaker recognition system to identify active speakers, and a user interface for participants to select the desired speaker, enhancing speech quality and intelligibility by isolating the speech signal of the chosen speaker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single microphone is used to receive audio in a conference room, then the device complexity is low, but the speech quality and intelligibility deteriorate when multiple speakers are present due to signal superposition and interference

Engineering Contradiction:
Improvemicrophone configurationVSAvoidspeech quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent divides the single microphone function into multiple microphones arranged in an array. Each microphone captures audio from different spatial positions, enabling the system to separate and isolate speech signals from multiple speakers through beamforming techniques, thereby improving speech quality while maintaining reasonable device complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces beamforming algorithms as an intermediary processing layer between the microphones and the speech output. This intermediary technique processes the raw audio signals from multiple microphones to enhance desired speech signals while suppressing interference and noise, resolving the contradiction between simple hardware and high speech quality

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If microphone arrays with beamforming are used to separate speech signals, then the speech quality improves, but the device complexity and processing requirements increase

Engineering Contradiction:
Improvespeech qualityVSAvoidmicrophone array processing
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs self-identification of active speakers through automated speaker recognition algorithms that analyze the separated speech signals. This self-service capability eliminates the need for manual speaker selection and reduces the complexity of user interaction, offsetting the increased processing complexity with automated intelligence

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent dynamically adjusts beamforming parameters and signal processing settings based on the number and positions of active speakers detected in real-time. By changing processing parameters adaptively rather than using fixed complex configurations, the system maintains high speech quality while optimizing computational efficiency

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If the mixed signal from all speakers is transmitted to remote rooms, then the system operation is simple, but the intelligibility and quality of individual speaker speech deteriorate due to interference

Engineering Contradiction:
Improvesignal transmissionVSAvoidspeech intelligibility
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent extracts individual speaker speech signals from the mixed audio environment using beamforming and speaker recognition techniques. By separating and extracting the desired speaker's signal before transmission, the system maintains simple operation while preventing the loss of speech intelligibility that would occur with mixed signal transmission

Inventive Principle:
Principle #2Taking out (Extraction)

4Device complexity

If the system arbitrarily selects the strongest speaker's signal for transmission, then the processing complexity is reduced, but the ability to meet user needs deteriorates as the selected speaker may not be the desired one

Engineering Contradiction:
Improvespeaker selection processingVSAvoiduser preference accommodation
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements a feedback mechanism where speaker identities are transmitted to remote participants who can then select which speaker they wish to hear. This feedback loop allows the system to adapt to user preferences while maintaining reasonable processing complexity through automated speaker identification and selection interface

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables participants to selectively listen to the speech of a desired speaker, reducing interference and noise, and improving speech quality by separating and identifying individual speaker signals, allowing remote participants to choose which speaker to hear.

Implementation Method 1

With the use of microphone arrays, a desired signal may be advantageously extracted from the cacophony of audio sounds using beamforming, or more generally spatiotemporal filtering techniques

Methodology Applied
Scientific EffectBeamforming:

Data Source

PatentUS8503653B2Method and apparatus for active speaker selection using microphone arrays and speaker recognition
Publication Date: 2013.08.06 WSOU INVESTMENTS LLC
  • US8503653B2 patent drawing
  • US8503653B2 patent drawing
  • US8503653B2 patent drawing

AI summary

A method and apparatus for performing active speaker selection in teleconferencing applications illustratively comprises a microphone array module, a speaker recognition system, a user interface, and a speech signal selection module. The microphone array module separates the speech signal from each active speaker from those of other active speakers, providing a plurality of individual speaker's speech signals. The speaker recognition system identifies each currently active speaker using conventional speaker recognition/identification techniques. These identities are then transmitted to a remote teleconferencing location for display to remote participants via a user interface. The remote participants may then select one of the identified speakers, and the speech signal selection module then selects for transmission the speech signal associated with the selected identified speaker, thereby enabling the participants at the remote location to listen to the selected speaker and neglect the speech from other active speakers.