ASR Model Training via Confidence-Filtered Transcription
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Automatic Speech Recognition (ASR) models struggle with accurate transcriptions due to poorly constructed training data sets and fail to keep pace with rapid changes in colloquial and idiomatic language, requiring improved training methods for enhanced performance.
Innovation Solution
A system and method for continuously refining the ASR model by analyzing speech activity, generating quality metrics, filtering transcriptions based on confidence scores, and allowing manual review and correction, which selects a subset of high-quality transcriptions to retrain the model, thereby improving transcription accuracy and speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional ASR models are trained with large data sets of thousands of hours of audio, then the model capacity increases, but the transcription accuracy still struggles to keep pace with rapid changes in colloquial and idiomatic language
Solution Approach 1:
The patent changes the quality parameters of training data by introducing confidence score thresholds and quality metrics to filter transcriptions. Instead of simply increasing data volume, the system selectively uses high-quality transcriptions (above threshold T) for training, thereby improving transcription accuracy while maintaining model reliability in adapting to language changes
Solution Approach 2:
The system implements feedback loops where transcriptions are continuously evaluated against quality metrics and confidence scores. High-quality transcriptions are fed back into the training set to retrain the ASR model, creating a self-improving system that adapts to colloquial and idiomatic language changes over time
2Measurement precision
If a filtering process is implemented to select transcriptions based on confidence scores and quality metrics, then transcription quality improves, but system complexity increases
Solution Approach 1:
The system performs preliminary filtering and quality assessment of transcriptions before they are used for training. By pre-evaluating transcriptions against confidence scores and quality metrics T, the system eliminates low-quality data early in the pipeline, improving transcription quality without requiring complex post-processing or retraining mechanisms
3Productivity
If the ASR model is continuously retrained with selected high-quality transcriptions, then transcription accuracy and speed improve, but computational resources and training time increase
Solution Approach 1:
Instead of retraining the model with all available transcriptions, the system applies partial action by selecting only a subset of high-quality transcriptions (those above threshold T) for retraining. This selective approach improves transcription speed and accuracy while minimizing computational resources and training time required
Data Source
AI summary
The disclosed system continuously refines a model used by an Automatic Speech Recognition (ASR) system to enable fast and accurate transcriptions of detected speech activity. The ASR system analyzes speech activity to generate text transcriptions and associated metrics (such as minimum Bayes risk and/or perplexity) that correspond to the quality of or confidence in each generated transcription. The system employs a filtering process to select certain text transcriptions based in part on one or more associated quality metrics. In addition, the system corrects for known systemic errors within the ASR system and provides a mechanism for manual review and correction of transcriptions. The system selects a subset of transcriptions based on factors including confidence score, and uses the selected subset of transcriptions to re-train the ASR model. By continuously retraining the ASR model, the system is able to provide ever faster and more accurate text transcriptions of detected speech activity.


