Voice User Interface Training Data Optimization via Synthesis Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice user interfaces (VUIs) in electronic devices face challenges in accurately recognizing complex speech interactions, leading to misunderstandings, especially due to variations in user speech such as accents, volume, and background noise, which affects the user experience and the accuracy of speech-to-text conversions.
Innovation Solution
A method that synthesizes human speech for training phrases, captures the synthesized audio, and uses a speech-to-text framework to convert it into text, comparing the original and converted text to identify misinterpretations, generating a score to rank phrases that are most misunderstood, allowing for adjustments to improve VUI accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If VUI training data is used as-is without optimization, then the development process is simple and fast, but the speech recognition accuracy deteriorates due to misunderstandings of complex speech interactions
Solution Approach 1:
The system performs preliminary analysis of training data before actual VUI deployment by synthesizing speech samples and evaluating them through speech-to-text frameworks to identify potential recognition issues in advance, allowing developers to fix problems before they affect real users
Solution Approach 2:
The system implements a feedback loop where synthesized speech is converted back to text and compared with original training phrases, generating scores that indicate which phrases are most misunderstood, allowing iterative improvement of training data quality
2Measurement precision
If more comprehensive training data is collected to handle various speech variations, then speech recognition accuracy improves, but the time and resources required for data processing increase
Solution Approach 1:
Instead of collecting and processing extensive real-world speech data, the system creates synthetic copies of training phrases through text-to-speech synthesis, then evaluates these copies through speech-to-text frameworks to identify weaknesses without requiring large amounts of additional real data
Solution Approach 2:
The system focuses evaluation efforts on the most critical training phrases by generating scores that rank which phrases are most misunderstood, allowing developers to prioritize optimization of high-impact phrases rather than uniformly processing all training data
3Adaptability or versatility
If the VUI is designed to handle complex speech interactions, then the user experience improves, but the recognition accuracy deteriorates due to increased complexity
Solution Approach 1:
The system performs preliminary evaluation of training data to identify which complex phrases are most vulnerable to misrecognition, allowing developers to strengthen training for these specific cases before deploying the VUI to handle complex interactions
Solution Approach 2:
The feedback mechanism generates specific scores for different training phrases, highlighting which complex interactions are most problematic, enabling targeted improvements that maintain versatility while improving accuracy for complex speech patterns
Data Source
AI summary
Techniques for optimizing training data within voice user interface (VUI) of an application under development are disclosed. A VUI feedback module synthesizes human speech of a training phrase. This phrase is presented upon a speaker which is simultaneously captured upon a microphone. A speech to text framework converts the synthesized training phrase into text (textualized training phrase). The VUI feedback module compares the textualized training phrase to the actual training phrase and generates a speech training data structure that identifies similarities or dissimilarities between the textualized training phrase and the actual training phrase. This data structure may be utilized by an application developer computing system to identify training data that is most venerable to misinterpretation when a user interacts with the VUI. The VUI may subsequently be adjusted to account for the vulnerabilities to improve operations or user experience of the VUI.


