Voice Template Distinctness in Speech Recognition Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice templates created in voice-directed workflow systems often lack distinctness, leading to similarity with other words in the application vocabulary, which erodes speech recognition performance, causing errors and reducing productivity.

Innovation Solution

A method for dynamic training analysis that compares voice templates to other templates during creation, prompting users to adjust spoken words to make them less similar, using instructions such as enunciation prompts or alternative words to ensure distinctness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If custom voice templates are created for each word, then speech recognition accuracy should improve, but template similarity between words increases leading to recognition errors

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtemplate distinctness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system performs preliminary comparison of voice templates against existing templates in the vocabulary before finalizing template creation. This pre-check prevents creation of similar templates that would cause recognition errors, ensuring distinctness is established beforehand rather than corrected afterward.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides feedback to the user when template similarity is detected, prompting them to create a new template with different pronunciation characteristics. This iterative feedback loop ensures the final template is sufficiently distinct from other vocabulary words while maintaining accurate representation of the intended word.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If workers create voice templates for all words, then speech recognition performance improves, but training time and productivity decrease

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system automatically identifies and creates templates only for words that are likely to cause confusion (such as phonetically similar words like 'five' and 'nine'), rather than requiring manual template creation for all vocabulary words. This localized approach maintains high recognition performance for critical words while significantly reducing overall training time.

Inventive Principle:
Principle #3Local quality

3Reliability

If voice templates are made more distinct through user adjustment, then template discrimination improves, but training complexity and user burden increase

Engineering Contradiction:
Improvetemplate discriminationVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically compares created templates against the vocabulary and provides guidance prompts to users when similarity is detected. This self-service mechanism handles the complexity of template comparison and distinction automatically, requiring users only to make minor adjustments rather than navigating complex training procedures.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10121466B2Methods for training a speech recognition system
Publication Date: 2018.11.06 HAND HELD PRODS INC
  • US10121466B2 patent drawing
  • US10121466B2 patent drawing
  • US10121466B2 patent drawing

AI summary

Speech recognition systems that use voice templates may create (or update) voice templates for a particular user by training (or re-training). If a training results in a vocabulary with similar voice templates, then the speech recognition system's performance may suffer. The present invention provides embraces methods for training a speech recognition system to prevent voice template similarity. In these methods, a trained word's voice template may be evaluated for similarity to other vocabulary templates prior to enrolling the voice template into the vocabulary. If template similarity is found, then a user may be prompted to retrain the system using an alternate word. Alternatively, the user may be prompted to retrain the system with the word spoken more clearly. This dynamic enrollment training analysis insures that all templates in the vocabulary are distinct.