Unsupervised Speech Recognition Model Weighting via Confidence Levels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in improving performance due to variability in recognition depending on environments and speakers, requiring manual transcription and intervention, which is time-consuming and costly, and are limited by the use of only one learning method, leading to inefficiencies in real-world applications.

Innovation Solution

An unsupervised learning system that automatically generates and evaluates acoustic models by measuring confidence levels of data, classifying it into learning and adaptation data, and applying weights to improve speech recognition performance without human intervention, using a processor to generate and update models through unsupervised learning and generative adversarial networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual transcription is used to transcribe sampled speech data, then speech recognition performance can be improved through better training data, but time and cost are significantly increased

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidtranscription time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses automatically transcribed speech data from the speech recognizer itself as training data, eliminating the need for manual transcription. The recognizer processes speech data and the resulting transcriptions are directly used for retraining, creating a self-improving system that reduces both time and cost while maintaining performance improvement.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary evaluation mechanism that automatically assesses the quality and suitability of automatically transcribed data before using it for training. This intermediary step filters and selects appropriate data without requiring manual transcription, bridging the gap between automatic transcription and high-quality training data requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If all sampled speech data is used for training, then more training data is available, but learning time increases significantly

Engineering Contradiction:
Improvetraining data quantityVSAvoidlearning time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system extracts and selects only the most suitable speech data for training by evaluating transcription quality and data characteristics. Instead of using all sampled data, the patent identifies and extracts high-value training data that meets specific criteria, reducing the overall data volume while maintaining or improving training effectiveness and reducing learning time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different quality standards and selection criteria to different portions of the speech data based on their characteristics. High-confidence transcriptions are selected for training while low-confidence data is excluded or handled differently, ensuring that only quality data contributes to training, thereby optimizing the balance between data quantity and learning time.

Inventive Principle:
Principle #3Local quality

3Device complexity

If only one learning method or adaptation method is used, then the system is simpler to implement, but speech recognition performance improvement is limited in real environments

Engineering Contradiction:
Improvelearning system complexityVSAvoidspeech recognition performance
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple learning methods and adaptation techniques into a unified system. It integrates unsupervised learning, supervised learning, and adaptation methods, allowing them to work together synergistically. This merging approach maintains relative system simplicity while significantly improving speech recognition performance in real-world environments through the complementary strengths of different methods.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to perform multiple functions using a single integrated framework that can switch between or combine different learning and adaptation methods. This multi-functional approach allows the system to adapt to various real-world conditions without requiring separate specialized systems, balancing complexity with performance improvement capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If a person directly evaluates the performance of generated models through direct intervention, then model performance can be accurately assessed, but time and cost are significantly increased

Engineering Contradiction:
Improvemodel performance evaluation accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system implements self-evaluation capabilities where the speech recognizer automatically assesses its own performance on test data without requiring human intervention. This self-evaluation mechanism provides accurate performance metrics while eliminating the time and cost associated with manual evaluation, maintaining objectivity through automated testing and comparison.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent establishes an automated feedback loop where performance evaluation results are immediately fed back into the system to guide further training and improvement. This continuous feedback mechanism replaces manual evaluation with an automated process that not only saves time but also enables rapid iterative improvement by quickly identifying performance gaps and directing targeted training.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11164565B2Unsupervised learning system and method for performing weighting for improvement in speech recognition performance and recording medium for performing the method
Publication Date: 2021.11.02 LG ELECTRONICS INC
  • US11164565B2 patent drawing
  • US11164565B2 patent drawing
  • US11164565B2 patent drawing

AI summary

A learning system and method for updating recognition performance by assigning weights according to a confidence level of data are discussed. The unsupervised learning system includes a memory configured to store speech data received from a server that performs speech recognition; and a processor configured to measure confidence levels of pieces of learnable data stored in the memory and classify the pieces of learnable data into learning data and adaptation data, generate a learning model by performing unsupervised learning on the learning data, generate an adaption model using the adaptation data, and evaluate speech recognition performance for the learning model and the adaptation model, wherein the processor is configured to assign weights by applying the measured confidence levels to the learning model and the adaptation model and update recognition performance with the learning model and the adaptation model to which the weights are applied.