Automated Speech Recognition Training Data Generation System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition training data generation methods are inefficient and costly, as they rely on manual processes and lack effective automation for generating high-quality training data across multiple auto speech recognition (ASR) engines.
Innovation Solution
A system and method that utilizes a network environment with user terminals and a server to preprocess speech data, interface with multiple ASR engines, evaluate transcription data based on confidence scores, and manage training data generation, ensuring accurate pairing and standardization of speech and transcription data for training purposes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual processes are used for speech recognition training data generation, then data quality can be controlled, but time consumption and cost increase significantly
Solution Approach 1:
The system enables self-service automated generation of training data by having ASR engines process speech data independently, with automatic confidence score evaluation and quality assessment, eliminating the need for manual data generation while maintaining quality standards through automated verification mechanisms
Solution Approach 2:
The system implements feedback loops where ASR engines generate transcription data with confidence scores, which are then evaluated against quality thresholds. The feedback mechanism allows automatic rejection or reprocessing of low-quality data, ensuring maintained data quality while automating the process
2Reliability
If multiple ASR engines are used for transcription data generation, then data quality and reliability improve, but system complexity and processing time increase
Solution Approach 1:
The system merges the outputs of multiple ASR engines by collecting transcription data from each engine and evaluating their confidence scores collectively. This combination approach leverages the strengths of different engines while using a unified evaluation framework to manage complexity
Solution Approach 2:
The system creates a universal evaluation framework that works across multiple different ASR engines, using standardized confidence score assessment and quality thresholds. This multi-functional approach allows the same evaluation mechanism to handle diverse engine outputs, reducing overall system complexity
3Productivity
If automated processes are implemented for training data generation, then productivity increases, but data quality control becomes more difficult
Solution Approach 1:
The automated system incorporates feedback mechanisms where ASR engines provide confidence scores for their transcriptions, and the system automatically evaluates these scores against predefined quality thresholds. This feedback loop enables quality control without manual intervention, maintaining both productivity and data quality
Solution Approach 2:
The system replaces manual quality control mechanisms with automated electronic evaluation processes. Confidence score assessment, quality threshold comparison, and automatic acceptance/rejection decisions are handled by computational algorithms rather than human operators, maintaining quality control while enabling automation
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided is a system for generating speech recognition training data, the system including: a speech data processing module receiving speech data from a user terminal and performing data preprocessing on the received speech data; an auto speech recognition (ASR) interfacing module transmitting the preprocessed speech data to a plurality of ASR engines and acquiring a confidence score and transcription data of the speech data from the plurality of ASR engines; an ASR result evaluating module determining whether the speech data and the transcription data match each other; and a training data managing unit generating training data as a pair of the speech data and the transcription data determined to match each other.