Voice Model Retraining Using Usage Data to Reduce Recognition Errors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice model enrollment processes are restrictive due to noise conditions, leading to incomplete training and high false acceptance/rejection rates, which can deter users from continuing with hands-free voice command usage.

Innovation Solution

A method for retraining a voice model using additional training data collected during usage, dynamically improving the model's efficacy over time by relaxing initial enrollment restrictions and iteratively updating the voice model with successful voice command interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If initial training data is collected under restrictive noise conditions, then the voice model achieves higher initial accuracy, but the enrollment process completion rate decreases

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidenrollment completion rate
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs preliminary voice model training with initial training data collected under controlled noise conditions before actual usage. This preliminary action establishes a baseline accurate model that can later be iteratively improved with additional training data collected during normal usage, resolving the contradiction between initial accuracy requirements and enrollment completion rates.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The voice model is designed to be dynamic and updatable. After initial training with restricted data, the model continues to learn and adapt by incorporating additional training data collected during actual usage scenarios. This dynamic updating capability allows the system to start with high accuracy under controlled conditions while eventually adapting to real-world variability, thus resolving the enrollment completion issue.

Inventive Principle:
Principle #15Dynamics

2Loss of time

If the voice model is trained with limited initial training data, then the enrollment process is faster and simpler, but the model produces high false acceptance and false rejection rates

Engineering Contradiction:
Improveenrollment timeVSAvoidtrigger phrase recognition reliability
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system implements continuous learning by collecting additional training data during normal usage and iteratively retraining the voice model. This continuous action of data collection and model improvement maintains fast initial enrollment while progressively reducing false acceptance and false rejection rates through ongoing refinement with real-world usage data.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system incorporates feedback mechanisms where the performance of the voice model during actual usage (including false acceptances and rejections) informs subsequent training iterations. By analyzing real-world performance data and using it to retrain the model, the system continuously improves reliability while maintaining the benefit of fast initial enrollment.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the voice model is retrained iteratively with additional training data, then the model accuracy improves over time, but the computational resources and processing time increase

Engineering Contradiction:
Improvevoice command recognition accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system applies partial retraining by selecting and using only the most relevant additional training data collected during usage, rather than retraining with all available data. This partial action approach improves model accuracy through targeted learning from useful examples while minimizing unnecessary computational energy consumption associated with processing excessive or redundant training data.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10885899B2Retraining voice model for trigger phrase using training data collected during usage
Publication Date: 2021.01.05 MOTOROLA MOBILITY LLC
  • US10885899B2 patent drawing
  • US10885899B2 patent drawing

AI summary

A method includes receiving initial training data associated with a trigger phrase in a device and training a voice model in the device using the initial training data. The voice model is used to identify a plurality of voice commands in the device initiated using the trigger phrase. Collection of additional training data from the plurality of voice commands and retraining of the voice model in the device are iteratively performed using the additional training data. A device includes a microphone and a processor to receive initial training data associated with a trigger phrase using the microphone, train a voice model device using the initial training data, use the voice model to identify a plurality of voice commands initiated using the trigger phrase, and iteratively collect additional training data from the plurality of voice commands and retrain the voice model in the device using the additional training data.