Semi-supervised ASR Training Data Generation via Targeted Word Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating ASR training data are inefficient in terms of time and transcription accuracy, and there is a disconnect between human and machine interfaces, leading to suboptimal training data that may contain redundant or less impactful information.
Innovation Solution
A semi-supervised method that selects optimal words or phrases based on the ASR engine's current corpus, presents them to users, receives speech utterances, and uses multiple transcription services for accurate transcription, allowing real-time editing and data augmentation to enhance training data quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional 3-step manual transcription process is used, then training data can be generated, but time consumption is excessive and transcription accuracy is low
Solution Approach 1:
The patent replaces the manual mechanical transcription process with an automated ASR system. The ASR engine automatically converts speech to text, eliminating the need for manual typing and verification steps. This substitution dramatically reduces time consumption while improving transcription accuracy through the use of advanced speech recognition algorithms and multiple verification mechanisms.
Solution Approach 2:
The ASR system performs self-verification through multiple transcription services and cross-validation mechanisms. The system automatically compares transcriptions from different services, identifies discrepancies, and resolves errors without human intervention. This self-service capability maintains high accuracy while minimizing time loss by eliminating manual verification steps.
2Reliability
If users provide training data without ASR knowledge, then data collection is simple, but training data quality is suboptimal with redundant or rare words
Solution Approach 1:
The system incorporates feedback loops where the ASR engine analyzes the training corpus and provides guidance to users about which words and phrases would be most beneficial to record. The system tracks transcription accuracy by word and actively seeks additional examples of frequently misrecognized words, creating a continuous improvement cycle that enhances training data quality without requiring users to have domain expertise.
Solution Approach 2:
The system performs preliminary analysis of the training corpus to identify gaps and weaknesses before users begin recording. It pre-calculates which words and phrases would have the highest impact on improving ASR performance and presents these as targeted suggestions to users, ensuring that collected data is highly valuable without requiring users to understand ASR technicalities.
3Reliability
If more training data is collected to improve ASR performance, then accuracy improves, but cost and time requirements increase significantly
Solution Approach 1:
The system dynamically adjusts the parameters of data collection based on the current state of the training corpus. Instead of uniformly collecting data across all vocabulary, it focuses resources on specific parameters such as frequently misrecognized words, underrepresented phonetic patterns, and high-impact phrases. This parameter-driven approach maximizes accuracy improvement per unit of data collected, reducing overall time and cost requirements.
Solution Approach 2:
The system applies different data collection strategies to different parts of the vocabulary based on their specific needs. High-frequency words receive different treatment compared to low-frequency words, and words with known transcription difficulties receive targeted attention. This localized approach ensures that data generation efficiency is optimized for each specific vocabulary segment rather than applying a one-size-fits-all method.
Data Source
AI summary
Techniques are disclosed for generating ASR training data. According to an embodiment, impactful ASR training corpora is generated efficiently, and the quality or relevance of ASR training corpora being generated is increased by leveraging knowledge of the ASR system being trained. An example methodology includes: selecting one of a word or phrase, based on knowledge and/or content of said ASR training corpora; presenting a textual representation of said word or phrase; receiving a speech utterance that includes said word or phrase; receiving a transcript for said speech utterance; presenting said transcript for review (to allow for editing, if needed); and storing said transcript and said audio file in an ASR system training database. The selecting may include, for instance, selecting a word or phrase that is under-represented in said database, and/or based upon an n-gram distribution on a language, and/or based upon known areas that tend to incur transcription mistakes.


