Context-Aware Speaker Recognition Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text independent speaker recognition systems require lengthy enrollment procedures and are context-specific, failing to perform well in varying environments and conditions.

Innovation Solution

A context-aware training method for text independent speaker recognition models that adapts over time through continuous learning, using context data such as location, microphone properties, and user state analysis to improve recognition accuracy without dedicated training time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing text independent speaker recognition systems use traditional enrollment procedures, then speaker recognition models can be generated, but the enrollment time becomes excessively long (five or more minutes)

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidenrollment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously collecting and storing speech samples from users during their normal interaction with the device. These samples are accumulated in advance and used later for model training, eliminating the need for dedicated enrollment time while maintaining recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system serves itself by automatically collecting, processing, and training on speech data without requiring user intervention or dedicated enrollment procedures. The continuous learning process happens autonomously in the background, allowing the system to improve its speaker recognition capabilities without consuming user time.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If speaker recognition models are trained in specific contexts, then recognition accuracy improves for those contexts, but the models fail to perform well in other contexts or environments

Engineering Contradiction:
Improverecognition accuracy in specific contextVSAvoidperformance across varying contexts
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system creates universal speaker recognition models that function across multiple contexts and environments. By collecting speech samples from diverse situations and continuously adapting the models, the system achieves both context-specific accuracy and broad versatility, allowing reliable recognition regardless of where or how the user speaks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements dynamic model adaptation where speaker recognition models continuously evolve and adjust to new contexts and environments. This dynamic learning process allows the models to maintain high accuracy across varying conditions by adapting to changes in user behavior, environment, and speech patterns over time.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10339935B2Context-aware enrollment for text independent speaker recognition
Publication Date: 2019.07.02 INTEL CORP
  • US10339935B2 patent drawing
  • US10339935B2 patent drawing
  • US10339935B2 patent drawing

AI summary

Techniques are provided for training of a text independent (TI) speaker recognition (SR) model. A methodology implementing the techniques according to an embodiment includes measuring context data associated with collected TI speech utterances from a user and identifying the user based on received identity measurements. The method further includes performing a speech quality analysis and a speaker state analysis based on the utterances, and evaluating a training merit value of the utterances, based on the speech quality analysis and the speaker state analysis. If the training merit value exceeds a threshold value, the utterances are stored as training data in a training database. The database is indexed by the user identity and the context data. The method further includes determining whether the stored training data has achieved a sufficiency level for enrollment of a TI SR model, and training the TI SR model for the identified user and context.