Virtual Meeting Speaker Separation Through Iterative Audio Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual conferencing systems struggle with accurately separating multiple speakers due to unknown numbers, variability in speaker environments, overlapping speech, unbalanced talk times, and gender variability, leading to low accuracy and reliability in speaker separation and identification.

Innovation Solution

An AI-based solution using machine learning algorithms, including divisive and agglomerative hierarchical clustering, to separate speakers by environment and gender, followed by re-segmentation to enhance accuracy and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speaker separation methods are used in virtual conferencing systems, then the system can handle multiple speakers, but the accuracy and reliability of speaker separation and identification deteriorates due to unknown numbers of speakers, variability in environments, overlapping speech, unbalanced talk times, and gender variability

Engineering Contradiction:
Improvespeaker separation accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the speaker separation task into multiple processing stages: initial clustering to group audio segments by speaker characteristics, iterative refinement to separate overlapping speech, and post-processing to handle unbalanced talk times. This multi-stage segmentation approach improves separation accuracy by breaking down the complex problem into manageable steps while systematically addressing each challenge (unknown speaker numbers, environmental variability, overlapping speech, and gender differences).

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs dynamic adaptive processing that adjusts separation parameters based on real-time analysis of speaker characteristics, environmental conditions, and speech overlap patterns. The iterative refinement process dynamically modifies clustering parameters and separation thresholds to optimize performance for each specific conferencing scenario, thereby improving reliability without requiring a fixed complex structure.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If advanced machine learning algorithms are implemented to improve speaker separation accuracy, then identification precision improves, but computational complexity and processing time increase

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by performing initial speaker clustering and environmental characterization before the main separation process. Audio segments are pre-grouped by speaker characteristics and environmental conditions, creating organized data structures that accelerate subsequent iterative refinement. This preliminary organization reduces the computational burden during real-time processing while maintaining high identification accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial action by focusing computational resources on the most challenging aspects of speaker separation - specifically targeting overlapping speech segments and boundary cases where speakers transition. Rather than uniformly processing all audio segments with maximum complexity, the system selectively applies advanced algorithms only where needed, reducing overall processing time while maintaining precision for difficult cases.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12380910B2Systems and methods for virtual meeting speaker separation
Publication Date: 2025.08.05 RINGCENTRAL INC
  • US12380910B2 patent drawing
  • US12380910B2 patent drawing
  • US12380910B2 patent drawing

AI summary

A computer-implemented machine learning method for improving speaker separation is provided. The method comprises processing audio data to generate prepared audio data and determining feature data and speaker data from the prepared audio data through a clustering iteration to generate an audio file. The method further comprises re-segmenting the audio file to generate a speaker segment and causing to display the speaker segment through a client device.