Unsupervised Speech Recognition Model Weighting via Confidence Levels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in improving performance due to variability in recognition depending on environments and speakers, requiring manual transcription and intervention, which is time-consuming and costly, and are limited by the use of only one learning method, leading to inefficiencies in real-world applications.
Innovation Solution
An unsupervised learning system that automatically generates and evaluates acoustic models by measuring confidence levels of data, classifying it into learning and adaptation data, and applying weights to improve speech recognition performance without human intervention, using a processor to generate and update models through unsupervised learning and generative adversarial networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual transcription is used to transcribe sampled speech data, then speech recognition performance can be improved through better training data, but time and cost are significantly increased
Solution Approach 1:
The system uses automatically transcribed speech data from the speech recognizer itself as training data, eliminating the need for manual transcription. The recognizer processes speech data and the resulting transcriptions are directly used for retraining, creating a self-improving system that reduces both time and cost while maintaining performance improvement.
Solution Approach 2:
The patent introduces an intermediary evaluation mechanism that automatically assesses the quality and suitability of automatically transcribed data before using it for training. This intermediary step filters and selects appropriate data without requiring manual transcription, bridging the gap between automatic transcription and high-quality training data requirements.
2Quantity of substance
If all sampled speech data is used for training, then more training data is available, but learning time increases significantly
Solution Approach 1:
The system extracts and selects only the most suitable speech data for training by evaluating transcription quality and data characteristics. Instead of using all sampled data, the patent identifies and extracts high-value training data that meets specific criteria, reducing the overall data volume while maintaining or improving training effectiveness and reducing learning time.
Solution Approach 2:
The patent applies different quality standards and selection criteria to different portions of the speech data based on their characteristics. High-confidence transcriptions are selected for training while low-confidence data is excluded or handled differently, ensuring that only quality data contributes to training, thereby optimizing the balance between data quantity and learning time.
3Device complexity
If only one learning method or adaptation method is used, then the system is simpler to implement, but speech recognition performance improvement is limited in real environments
Solution Approach 1:
The patent combines multiple learning methods and adaptation techniques into a unified system. It integrates unsupervised learning, supervised learning, and adaptation methods, allowing them to work together synergistically. This merging approach maintains relative system simplicity while significantly improving speech recognition performance in real-world environments through the complementary strengths of different methods.
Solution Approach 2:
The system is designed to perform multiple functions using a single integrated framework that can switch between or combine different learning and adaptation methods. This multi-functional approach allows the system to adapt to various real-world conditions without requiring separate specialized systems, balancing complexity with performance improvement capability.
4Measurement precision
If a person directly evaluates the performance of generated models through direct intervention, then model performance can be accurately assessed, but time and cost are significantly increased
Solution Approach 1:
The system implements self-evaluation capabilities where the speech recognizer automatically assesses its own performance on test data without requiring human intervention. This self-evaluation mechanism provides accurate performance metrics while eliminating the time and cost associated with manual evaluation, maintaining objectivity through automated testing and comparison.
Solution Approach 2:
The patent establishes an automated feedback loop where performance evaluation results are immediately fed back into the system to guide further training and improvement. This continuous feedback mechanism replaces manual evaluation with an automated process that not only saves time but also enables rapid iterative improvement by quickly identifying performance gaps and directing targeted training.
Data Source
AI summary
A learning system and method for updating recognition performance by assigning weights according to a confidence level of data are discussed. The unsupervised learning system includes a memory configured to store speech data received from a server that performs speech recognition; and a processor configured to measure confidence levels of pieces of learnable data stored in the memory and classify the pieces of learnable data into learning data and adaptation data, generate a learning model by performing unsupervised learning on the learning data, generate an adaption model using the adaptation data, and evaluate speech recognition performance for the learning model and the adaptation model, wherein the processor is configured to assign weights by applying the measured confidence levels to the learning model and the adaptation model and update recognition performance with the learning model and the adaptation model to which the weights are applied.


