Active Speaker Identification in Multi-Endpoint Conferencing Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conferencing systems struggle to accurately identify the active speaker in multi-location conferences, where participants are distributed across different locations, and provide real-time information about who is speaking.

Innovation Solution

A conferencing system that uses a processor to receive audio signals from multiple endpoints, generates voice identification information through registration audio signals, and identifies active speakers by comparing audio energy values and voice characteristics, transmitting this information to remote endpoints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio signals are received from multiple endpoints in a conferencing system, then the ability to communicate across locations is improved, but the difficulty of accurately identifying the active speaker increases

Engineering Contradiction:
Improvemulti-location communication capabilityVSAvoidactive speaker identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system segments the audio signal processing by endpoint, analyzing audio energy values from each location separately before comparing across endpoints. This allows accurate identification of which endpoint has the active speaker while maintaining multi-location communication capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary processing layer that collects audio signals from multiple endpoints, compares audio energy values, and determines active speaker identity before transmitting to participants. This intermediary function resolves the contradiction by enabling accurate identification despite multiple locations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If real-time audio processing is performed to identify active speakers, then communication effectiveness is improved, but system complexity increases

Engineering Contradiction:
Improvereal-time communication effectivenessVSAvoidsignal processing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system uses self-service processing where each endpoint independently measures its own audio energy values, and the comparison logic automatically determines the active speaker without requiring complex external processing. This reduces overall system complexity while maintaining real-time effectiveness.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameter being measured from raw audio signals to audio energy values, which are easier to process and compare in real-time. This parameter transformation simplifies the processing requirements while maintaining real-time communication effectiveness.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If audio energy values are compared across multiple endpoints, then active speaker identification is improved, but information loss about participant locations may occur

Engineering Contradiction:
Improveactive speaker detection precisionVSAvoidparticipant location information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system provides feedback to participants about both the active speaker identity and their own location information. This feedback mechanism ensures that location information is not lost but rather used to enhance the active speaker identification and communicate it back to participants.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system adds a new dimension to the information transmission by including both audio-based speaker identification and location-based participant information in the same communication stream. This dimensional expansion prevents information loss while improving detection precision.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9210269B2Active speaker indicator for conference participants
Publication Date: 2015.12.08 CISCO TECHNOLOGY INC
  • US9210269B2 patent drawing
  • US9210269B2 patent drawing
  • US9210269B2 patent drawing

AI summary

In one embodiment, a method includes receiving requests to join a conference from a plurality of user devices proximate a first endpoint. The requests include a username. The method also includes receiving an audio signal for the conference from the first endpoint. The first endpoint is operable to capture audio proximate the first endpoint. The method also includes transmitting the audio signal to a second endpoint, remote from the first endpoint. The method also includes identifying, by a processor, an active speaker proximate the first endpoint based on information received from the plurality of user devices.