Multi-Scale Speaker Diarization for Overlap and Boundary Errors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker diarization systems face challenges in accurately attributing short speech segments to the correct speaker due to insufficient speech information, while long segments are prone to errors from overlapping speakers or boundary issues.

Innovation Solution

A multi-scale approach using dynamically-weighted embeddings, involving a speaker embeddings model, clustering model, dynamic weights model, and speaker labeling model, processes speech segments of varying durations to generate accurate speaker labels by applying dynamic weights and cosine similarity, enabling high temporal resolution and efficient detection of overlapping speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If short speech segments are used for speaker diarization, then temporal resolution is improved, but attribution accuracy deteriorates due to insufficient speech information

Engineering Contradiction:
Improvetemporal resolutionVSAvoidspeaker attribution accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent segments speech into multiple time scales (short segments for temporal resolution, long segments for accuracy) and processes them through separate embedding models. The short segment embeddings capture precise timing information while long segment embeddings provide sufficient speech context for accurate speaker attribution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates composite speaker representations by combining embeddings from multiple time scales. The final speaker attribution is determined by aggregating evidence from both short-scale and long-scale embeddings, creating a more robust and accurate speaker identification system that overcomes the limitations of using单一 scale embeddings.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If long speech segments are used for speaker diarization, then speaker attribution accuracy is improved, but error rate increases due to overlapping speakers or boundary issues

Engineering Contradiction:
Improvespeaker attribution accuracyVSAvoiderror rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent divides long speech segments into multiple overlapping time scales, processing them through separate embedding models. This segmentation allows the system to identify and isolate overlapping speech regions, reducing errors from boundary issues while maintaining the sufficient speech context needed for accurate speaker attribution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses overlapping segments that extend beyond the minimal required boundaries, intentionally creating redundancy. This excessive action ensures that even if some segments contain errors from overlapping speakers, other overlapping segments provide sufficient correct information for accurate speaker attribution.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If multi-scale embeddings are processed, then speaker diarization accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvespeaker diarization accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the computational process into distinct embedding models for different time scales, allowing parallel processing of short and long segments. This segmentation enables efficient utilization of computational resources by processing different time scales independently and then aggregating results.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic weighting of embeddings from different time scales, where the system adaptively determines the relative importance of short-scale versus long-scale embeddings based on the specific speech context. This dynamic approach optimizes computational efficiency by emphasizing the most relevant time scale for each segment.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20260073937A1Multi-scale speaker diarization for conversational ai systems and applications
Publication Date: 2026.03.12 NVIDIA CORP
  • US20260073937A1 patent drawing
  • US20260073937A1 patent drawing
  • US20260073937A1 patent drawing

AI summary

Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing speaker diarization. The techniques include obtaining a speaker embedding for various reference times of a speech and for various differently-sized time intervals, identifying a plurality of clusters, each cluster associated with a different speaker of the speech. The techniques further include computing, using the speaker embeddings, a set of embedding weights for various differently-sized time intervals, and identifying, using the computed set of the embedding weights, one or more speakers speaking at a respective reference time.