Speaker Position Identification Using Microphone Array Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems fail to reliably identify the active speaker in conference calls when multiple participants share a single audio input device, leading to difficulties for listeners in distinguishing who is speaking.

Innovation Solution

A system and method that uses a microphone array and position processing module to determine the physical position of the active speaker, transmitting this information for display to other sites, allowing listeners to identify the speaker through a graphical user interface that shows the speaker's position and potentially their identity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If multiple participants share a single audio input device, then device complexity is reduced and ease of operation is improved, but reliable identification of the active speaker becomes difficult

Engineering Contradiction:
Improveease of operationVSAvoidreliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The audio input device is segmented into multiple individual microphones arranged in a specific pattern. Each microphone captures sound from a particular spatial direction, allowing the system to distinguish between different speakers based on which microphone detects the sound first or most strongly. This segmentation enables reliable speaker identification while maintaining the simplicity of using a single shared device.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An intermediary processing system is introduced between the audio input device and the conference call system. This intermediary analyzes the temporal and spatial characteristics of the audio signals from multiple microphones to determine which participant is currently speaking, thereby enabling reliable speaker identification without requiring each participant to have their own dedicated audio device.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If each participant uses a distinct audio input device, then reliable identification of the active speaker is achieved, but device complexity increases and ease of operation deteriorates

Engineering Contradiction:
ImprovereliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Multiple individual audio input devices (microphones) are merged into a single integrated audio input device that shares a common housing and control system. This merged device maintains the ability to reliably identify speakers through its multiple sensing elements while eliminating the need for each participant to have their own separate device, thereby reducing overall system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The single shared audio input device is designed with multi-functionality to serve multiple participants simultaneously. It can distinguish between different speakers, detect their positions, and provide identification information to the conference call system, making it a universal solution that replaces the need for multiple dedicated devices.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If multiple participants share a single audio input device, then cost is reduced, but the ability to distinguish and identify speakers deteriorates

Engineering Contradiction:
Improvequantity of devicesVSAvoidloss of information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The system adds a spatial dimension to audio detection by using multiple microphones positioned at different locations and orientations. This dimensional expansion allows the system to distinguish speakers not just by their voice characteristics but by their spatial position and the temporal pattern of sound arrival at different microphones, thereby preventing information loss about speaker identity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system replaces reliance on mechanical or physical separation of audio devices with an electronic and computational approach. By using signal processing algorithms to analyze the temporal and spatial characteristics of audio from multiple microphones, the system achieves speaker identification without requiring physical separation of devices, thus maintaining cost-effectiveness while preventing information loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution provides improved identification of the active speaker by visually indicating their position and, if available, their identity, enhancing listener engagement and understanding during conference calls by clearly indicating who is speaking.

Implementation Method 1

The position processing module is coupled to receive acoustic signals from a microphone array

Methodology Applied
Scientific EffectAcoustic signals: Sound

Data Source

PatentUS9083822B1Speaker position identification and user interface for its representation
Publication Date: 2015.07.14 SHORETEL INC
  • US9083822B1 patent drawing
  • US9083822B1 patent drawing
  • US9083822B1 patent drawing

AI summary

A system, method and graphical user interface for determining a speaker's position and a generating a display showing the position of the speaker. In one embodiment, the system comprises a first speakerphone system and a second speakerphone system communicatively coupled to send and receive data. The speakerphone system comprises a display, an input device, a microphone array, a speaker, and a position processing module. The position processing module is coupled to receive acoustic signals from the microphone array. The position processing module uses these acoustic signals to determine a position of the speaker. The position information is then sent to other speakerphone system for presentation on the display.