Hybrid Speaker Recognition System Using Dynamic TD-TI Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker recognition systems face challenges in efficiently verifying user identity, particularly in scenarios where computational resources are limited, and data security is a concern.

Innovation Solution

The implementation of a hybrid speaker recognition system that selectively uses both text-dependent (TD) and text-independent (TI) speaker recognition models, combining TD user measures and TI user measures to enhance verification accuracy and robustness, while conserving computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text-independent speaker recognition is used to verify user identity without constraining specific words, then speaker recognition accuracy is improved, but computational resources are consumed more heavily

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts the speaker recognition approach based on the characteristics of the audio input. It evaluates whether the input contains clear speaker characteristics or if it requires text-dependent constraints, and adapts the verification process accordingly to balance accuracy and computational efficiency

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters of the speaker recognition model based on the input characteristics. It adjusts the threshold for speaker verification and selects different processing paths depending on whether the audio input provides sufficient speaker-specific information, thereby optimizing resource usage while maintaining accuracy

Inventive Principle:
Principle #35Parameter changes

2Use of energy by moving object

If text-dependent speaker recognition with specific invocation phrases is used, then computational resources are conserved, but speaker recognition robustness deteriorates

Engineering Contradiction:
Improvecomputational resource consumptionVSAvoidspeaker recognition robustness
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The system dynamically selects between text-dependent and text-independent verification paths based on the quality and characteristics of the audio input. When the input contains clear invocation phrases, it uses the more efficient text-dependent path; otherwise, it transitions to the more robust text-independent path

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The speaker recognition process is segmented into multiple stages: initial text-dependent filtering using invocation phrases, followed by text-independent verification for robustness. This segmentation allows the system to benefit from both approaches without fully committing to the computationally expensive text-independent method for all inputs

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If hybrid speaker recognition system dynamically weights TD and TI user measures, then verification accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveverification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system changes the weighting parameters of TD and TI user measures dynamically based on the confidence levels and characteristics of each measurement. It adjusts these weights in real-time during verification to optimize accuracy while managing complexity through adaptive parameter tuning

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system uses feedback from the initial verification stage to adjust the weighting of subsequent measurements. The confidence levels from TD verification inform how much weight should be given to TI verification results, creating a feedback loop that improves accuracy while preventing unnecessary computational overhead

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250131916A1Text independent speaker recognition
Publication Date: 2025.04.24 GOOGLE LLC
  • US20250131916A1 patent drawing
  • US20250131916A1 patent drawing
  • US20250131916A1 patent drawing

AI summary

Text independent speaker recognition models can be utilized by an automated assistant to verify a particular user spoke a spoken utterance and/or to identify the user who spoke a spoken utterance. Implementations can include automatically updating a speaker embedding for a particular user based on previous utterances by the particular user. Additionally or alternatively, implementations can include verifying a particular user spoke a spoken utterance using output generated by both a text independent speaker recognition model as well as a text dependent speaker recognition model. Furthermore, implementations can additionally or alternatively include prefetching content for several users associated with a spoken utterance prior to determining which user spoke the spoken utterance.