Active Speaker Detection in Electronic Meetings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional videoconferencing systems struggle to detect active speakers when devices are muted or do not provide audio data, leading to difficulties in determining the active speaker and adjusting video displays accordingly.

Innovation Solution

An electronic meeting system that receives speaker detection signals from devices, independent of audio or video data, to determine which devices have detected an active speaker, allowing for the identification and display of the active participant even when devices are muted or do not provide audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If devices are muted or do not provide audio data, then privacy is preserved and bandwidth consumption is reduced, but the ability to detect active speakers is lost

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidactive speaker detection accuracy
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The system separates audio data transmission from speaker detection functionality. Devices are segmented into different operational modes where audio data can be muted while speaker detection signals continue to be transmitted independently, allowing the meeting system to detect active speakers without receiving full audio streams from all devices.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A dedicated speaker detection signal acts as an intermediary between the device and the meeting system. This separate detection channel allows the system to determine active speaker status without relying on audio data, enabling accurate speaker detection even when audio is muted or bandwidth is limited.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If audio data is transmitted from all devices, then active speaker detection accuracy is improved, but bandwidth consumption increases

Engineering Contradiction:
Improveactive speaker detection accuracyVSAvoidbandwidth consumption
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The speaker detection functionality is extracted from the audio data stream. Instead of analyzing audio data to detect speakers, the system uses a separate, dedicated speaker detection signal that is transmitted independently. This extraction allows accurate speaker detection without the need to process or transmit full audio data from all devices.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses a simplified copy or representation of speaker presence through detection signals rather than transmitting the actual audio data. These detection signals convey the essential information needed for speaker identification without the bandwidth requirements of full audio streams.

Inventive Principle:
Principle #26Copying

3Measurement precision

If speaker detection relies on audio data, then detection accuracy is improved, but the system cannot detect speakers when devices are muted

Engineering Contradiction:
Improvespeaker detection accuracyVSAvoiddetection capability under muted conditions
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts its detection mechanism based on audio availability. When audio data is available, the system can use traditional audio analysis. When devices are muted or audio is unavailable, the system automatically switches to using dedicated speaker detection signals, ensuring continuous speaker detection capability across different operational conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The speaker detection signal serves multiple functions: it works both when audio data is available and when devices are muted. This universal detection mechanism eliminates the limitation of traditional systems that fail to detect speakers in muted conditions, providing consistent speaker identification across all device states.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11282537B2Active speaker detection in electronic meetings for providing video from one device to plurality of other devices
Publication Date: 2022.03.22 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11282537B2 patent drawing
  • US11282537B2 patent drawing
  • US11282537B2 patent drawing

AI summary

Active speaker detection can include receiving speaker detection signals from a plurality of devices participating in an electronic meeting. Each speaker detection signal specifies a score indicating whether an active speaker is detected by a respective device of the plurality of devices that generates the speaker detection signal. Active speaker detection further can include determining, using a processor, a device of the plurality of devices that detects an active speaker based upon the speaker detection signals, wherein, in response to the determining, the method further comprises: providing video received from the determined device to the plurality of devices during the electronic meeting.