Voice User Interface Training Data Optimization via Synthesis Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice user interfaces (VUIs) in electronic devices face challenges in accurately recognizing complex speech interactions, leading to misunderstandings, especially due to variations in user speech such as accents, volume, and background noise, which affects the user experience and the accuracy of speech-to-text conversions.

Innovation Solution

A method that synthesizes human speech for training phrases, captures the synthesized audio, and uses a speech-to-text framework to convert it into text, comparing the original and converted text to identify misinterpretations, generating a score to rank phrases that are most misunderstood, allowing for adjustments to improve VUI accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If VUI training data is used as-is without optimization, then the development process is simple and fast, but the speech recognition accuracy deteriorates due to misunderstandings of complex speech interactions

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtraining data optimization process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of training data before actual VUI deployment by synthesizing speech samples and evaluating them through speech-to-text frameworks to identify potential recognition issues in advance, allowing developers to fix problems before they affect real users

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback loop where synthesized speech is converted back to text and compared with original training phrases, generating scores that indicate which phrases are most misunderstood, allowing iterative improvement of training data quality

Inventive Principle:
Principle #23Feedback

2Measurement precision

If more comprehensive training data is collected to handle various speech variations, then speech recognition accuracy improves, but the time and resources required for data processing increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtraining data processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of collecting and processing extensive real-world speech data, the system creates synthetic copies of training phrases through text-to-speech synthesis, then evaluates these copies through speech-to-text frameworks to identify weaknesses without requiring large amounts of additional real data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system focuses evaluation efforts on the most critical training phrases by generating scores that rank which phrases are most misunderstood, allowing developers to prioritize optimization of high-impact phrases rather than uniformly processing all training data

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If the VUI is designed to handle complex speech interactions, then the user experience improves, but the recognition accuracy deteriorates due to increased complexity

Engineering Contradiction:
Improvespeech interaction capabilityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary evaluation of training data to identify which complex phrases are most vulnerable to misrecognition, allowing developers to strengthen training for these specific cases before deploying the VUI to handle complex interactions

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The feedback mechanism generates specific scores for different training phrases, highlighting which complex interactions are most problematic, enabling targeted improvements that maintain versatility while improving accuracy for complex speech patterns

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10565982B2Training data optimization in a service computing system for voice enablement of applications
Publication Date: 2020.02.18 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10565982B2 patent drawing
  • US10565982B2 patent drawing
  • US10565982B2 patent drawing

AI summary

Techniques for optimizing training data within voice user interface (VUI) of an application under development are disclosed. A VUI feedback module synthesizes human speech of a training phrase. This phrase is presented upon a speaker which is simultaneously captured upon a microphone. A speech to text framework converts the synthesized training phrase into text (textualized training phrase). The VUI feedback module compares the textualized training phrase to the actual training phrase and generates a speech training data structure that identifies similarities or dissimilarities between the textualized training phrase and the actual training phrase. This data structure may be utilized by an application developer computing system to identify training data that is most venerable to misinterpretation when a user interacts with the VUI. The VUI may subsequently be adjusted to account for the vulnerabilities to improve operations or user experience of the VUI.