Speech Recognition Model Update via Failure Cause Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face limitations in improving accuracy due to the inability to determine whether recognition failures are caused by acoustic or language models, leading to data imbalance and performance deterioration.
Innovation Solution
A method and device for analyzing speech recognition failure data to differentiate between acoustic and language model errors, updating and machine-learning the respective models based on the identified causes, thereby improving recognition performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If speech recognition failure data is updated without analyzing the cause, then the data volume increases, but the speech recognition accuracy deteriorates due to data imbalance
Solution Approach 1:
The patent segments speech recognition failures into two distinct categories: acoustic model failures and language model failures. By analyzing the cause of each failure and routing it to the appropriate model's learning database, the system prevents mixing unrelated data that would cause imbalance. This segmentation allows each model to learn from relevant failures only, maintaining data quality while increasing overall data volume.
Solution Approach 2:
The patent applies local quality by directing different types of failure data to different learning databases based on the failure cause. Acoustic model failures are added to the acoustic model's learning database, while language model failures are added to the language model's learning database. This ensures that each part of the system receives appropriately tailored data, improving the quality of learning for each specific model.
2Productivity
If acoustic model and language model learning are performed simultaneously without cause analysis, then both models are updated, but the performance deteriorates due to incorrect data assignment
Solution Approach 1:
The patent segments the model update process into two separate pathways based on failure cause analysis. When a failure is identified as acoustic model-related, only the acoustic model learning database is updated. When a failure is identified as language model-related, only the language model learning database is updated. This segmentation prevents incorrect data from being assigned to the wrong model, maintaining recognition performance while preserving update efficiency.
Solution Approach 2:
The patent implements feedback by analyzing the cause of each speech recognition failure and using this information to guide the model update process. The system continuously monitors recognition failures, determines whether they stem from acoustic or language model issues, and accordingly updates the appropriate model's learning database. This feedback loop ensures that model updates are based on relevant data, preventing performance deterioration.
3Reliability
If user pronunciation information is added to recognition dictionary without cause analysis, then recognition success rate improves for user-specific cases, but overall speech recognition accuracy is limited
Solution Approach 1:
The patent segments the approach to handling user pronunciation by first analyzing the cause of recognition failure. If the failure is due to acoustic model issues (including user-specific pronunciation), the system adds the data to the acoustic model learning database. This segmentation allows the system to address user-specific cases while maintaining overall accuracy through systematic cause-based routing.
Solution Approach 2:
The patent performs preliminary action by analyzing the cause of recognition failure before adding data to any learning database. This preliminary analysis ensures that data is only added to the appropriate database based on the actual cause, preventing premature or incorrect updates that would limit overall accuracy improvement.
Data Source
AI summary
Disclosed are a speech recognition method capable of communicating with other electronic devices and an external server in a 5G communication condition by performing speech recognition by executing an artificial intelligence (AI) algorithm and/or a machine learning algorithm. The speech recognition method may comprise performing speech recognition by using an acoustic model and a language model stored in a speech database, determining whether the speech recognition of the spoken sentence is successful, storing speech recognition failure data when the speech recognition of the spoken sentence fails, analyzing the speech recognition failure data of the spoken sentence and updating the acoustic model or the language model by adding the recognition failure data to a learning database of the acoustic model or the language model when the cause of the speech recognition failure is due to the acoustic model or the language model and machine-learning the acoustic model or the language model.


