Adaptive Diarization Model for Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker diarization systems in translation applications rely on explicit speaker indications, which are cumbersome, error-prone, and consume excessive power, leading to unnatural conversation and increased battery consumption.

Innovation Solution

A diarization model is trained to distinguish between speakers based on supervised data, allowing it to automatically identify speakers without explicit hints, reducing the need for button presses and silent periods, and adapt over time for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If explicit speaker indications (button presses, silent periods) are used for speaker diarization, then speaker identification can be achieved, but the system becomes cumbersome, error-prone, and consumes excessive power

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidconversation naturalness
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs self-service by automatically detecting speaker turns through audio signal analysis without requiring user intervention. The speaker diarization model autonomously identifies speaker boundaries and attributes by processing audio waveforms, eliminating the need for users to press buttons or create silent periods to indicate speaker changes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces mechanical user actions (button presses, intentional silent periods) with an automated acoustic analysis system. The speaker diarization model uses machine learning to detect speaker boundaries and attributes through audio signal processing, substituting the mechanical interaction paradigm with an intelligent automated system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If explicit speaker indications are required, then speaker identification can be performed, but power consumption increases and latency increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs self-service by automatically detecting speaker turns through audio signal analysis without requiring user intervention. The speaker diarization model autonomously identifies speaker boundaries and attributes by processing audio waveforms, eliminating the need for users to press buttons or create silent periods to indicate speaker changes.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If explicit speaker indications are required, then speaker identification can be performed, but human error increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoiderror rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system performs self-service by automatically detecting speaker turns through audio signal analysis without requiring user intervention. The speaker diarization model autonomously identifies speaker boundaries and attributes by processing audio waveforms, eliminating the need for users to press buttons or create silent periods to indicate speaker changes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces mechanical user actions (button presses, intentional silent periods) with an automated acoustic analysis system. The speaker diarization model uses machine learning to detect speaker boundaries and attributes through audio signal processing, substituting the mechanical interaction paradigm with an intelligent automated system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If supervised training data with hint/identity data is used, then the diarization model can be trained to distinguish speakers, but training data requirements increase

Engineering Contradiction:
Improvespeaker distinction capabilityVSAvoidtraining data volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system implements feedback by using the diarization model's own predictions to generate additional training data. The model processes audio waveforms, predicts speaker boundaries and attributes, and uses these predictions (along with any available hint data) to update and refine its parameters, creating a self-reinforcing learning loop that improves accuracy without requiring proportionally more external training data.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service by automatically detecting speaker turns through audio signal analysis without requiring user intervention. The speaker diarization model autonomously identifies speaker boundaries and attributes by processing audio waveforms, eliminating the need for users to press buttons or create silent periods to indicate speaker changes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11710496B2Adaptive diarization model and user interface
Publication Date: 2023.07.25 GOOGLE LLC
  • US11710496B2 patent drawing
  • US11710496B2 patent drawing
  • US11710496B2 patent drawing

AI summary

A computing device receives a first audio waveform representing a first utterance and a second utterance. The computing device receives identity data indicating that the first utterance corresponds to a first speaker and the second utterance corresponds to a second speaker. The computing device determines, based on the first utterance, the second utterance, and the identity data, a diarization model configured to distinguish between utterances by the first speaker and utterances by the second speaker. The computing device receives, exclusively of receiving further identity data indicating a source speaker of a third utterance, a second audio waveform representing the third utterance. The computing device determines, by way of the diarization model and independently of the further identity data of the first type, the source speaker of the third utterance. The computing device updates the diarization model based on the third utterance and the determined source speaker.