Iterative Speaker Embedding for Real-Time Speech Diarization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker diarization systems are limited by unsupervised clustering algorithms that fail to leverage labeled training data, struggle with speaker overlap, and are inefficient for real-time processing of long audio sequences, leading to inaccurate speech recognition.

Innovation Solution

An end-to-end neural diarization system (DIVE) that combines a temporal encoder, iterative speaker selector, and voice activity detector to iteratively select speaker embeddings and predict voice activity indicators, utilizing labeled training data and a collar-aware training process for improved accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If unsupervised clustering algorithms are used for speaker diarization, then the system can process multiple speakers, but the diarization accuracy deteriorates especially in the presence of overlap

Engineering Contradiction:
Improveability to process multiple speakersVSAvoiddiarization accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent replaces traditional unsupervised clustering algorithms with a neural network-based embedding system. The neural encoder extracts optimized speaker-discriminative embeddings that are specifically trained to handle overlapping speech, substituting the mechanical clustering process with a learned representation that achieves higher diarization accuracy in complex scenarios.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the speaker representation from clustered vectors to optimized embeddings in a higher-dimensional space. By changing the parameter space and using neural network encoders, the system creates speaker-discriminative embeddings that capture subtle acoustic variations, improving accuracy while maintaining the ability to handle multiple simultaneous speakers.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional speaker diarization systems are used, then they can handle basic speaker separation, but they struggle with real-time processing and long sequences due to memory constraints

Engineering Contradiction:
Improvespeaker separation capabilityVSAvoidreal-time processing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent processes audio sequences by dividing them into manageable segments that are encoded independently into embeddings. This segmentation approach allows the system to handle long sequences and real-time processing by processing audio in chunks, reducing memory constraints while maintaining reliable speaker separation across the entire sequence.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If optimized speaker-discriminative embeddings are extracted, then diarization accuracy improves, but computational complexity and processing time increase

Engineering Contradiction:
Improvediarization accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary encoding of the entire audio sequence into temporal embeddings before conducting speaker diarization. By pre-processing the audio to extract optimized speaker-discriminative embeddings upfront, the system reduces the computational complexity of subsequent diarization steps, achieving high accuracy without excessive processing time during actual speaker assignment.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12633305B2End-to-end speech diarization via iterative speaker embedding
Publication Date: 2026.05.19 GOOGLE LLC
  • US12633305B2 patent drawing
  • US12633305B2 patent drawing
  • US12633305B2 patent drawing

AI summary

A method includes receiving an input audio signal corresponding to utterances spoken by multiple speakers. The method also includes encoding the input audio signal into a sequence of T temporal embeddings. During each of a plurality of iterations each corresponding to a respective speaker of the multiple speakers, the method includes selecting a respective speaker embedding for the respective speaker by determining a probability that the corresponding temporal embedding includes a presence of voice activity by a single new speaker for which a speaker embedding was not previously selected during a previous iteration and selecting the respective speaker embedding for the respective speaker as the temporal embedding. The method also includes, at each time step, predicting a respective voice activity indicator for each respective speaker of the multiple speakers based on the respective speaker embeddings selected during the plurality of iterations and the temporal embedding.