Personal Language Model for Speech Recognition Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems lack accuracy in considering user-specific language and pronunciation variations, leading to suboptimal recognition results.
Innovation Solution
An electronic apparatus equipped with a personal language model and a personal acoustic model that learns from user input text and speech, allowing for improved speech recognition by adjusting probabilities based on user-specific data and retraining models for enhanced accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a general speech recognition module is used, then the system can recognize speech, but the recognition accuracy does not account for user-specific language and pronunciation variations
Solution Approach 1:
The speech recognition system is segmented into multiple components: a general speech recognition module for baseline recognition and a personal language model for user-specific adaptation. This segmentation allows each component to specialize in different aspects, with the personal language model focusing specifically on user characteristics to improve overall accuracy.
Solution Approach 2:
A personal language model acts as an intermediary between the general speech recognition module and the final recognition result. This intermediary layer processes the general module's output and refines it by applying user-specific language patterns and pronunciation characteristics, thereby improving accuracy without replacing the entire system.
2Measurement precision
If a personal language model is introduced to improve recognition accuracy, then user-specific characteristics are considered, but the system complexity increases
Solution Approach 1:
The personal language model is designed to be multi-functional, serving both as a refinement layer for the general speech recognition module and as a standalone adaptation mechanism. This universality allows the system to handle different user profiles and speech patterns without requiring separate dedicated systems for each function.
Solution Approach 2:
The personal language model is trained using the user's own speech data and language patterns, allowing the system to automatically adapt and improve recognition accuracy without requiring manual configuration or complex external intervention. The model serves itself by continuously learning from user interactions.
3Ease of operation
If speech recognition is performed without personalization, then the system operates simply, but it cannot provide accurate recognition for individual user characteristics
Solution Approach 1:
The personal language model is trained in advance using the user's speech data and language patterns before actual speech recognition tasks. This preliminary action of training and adaptation allows the system to automatically incorporate user characteristics without adding complexity to the real-time operation, maintaining ease of use while improving accuracy.
Data Source
AI summary
An electronic apparatus configured to acquire information on a plurality of candidate texts corresponding to input speech of a user through a general speech recognition module, determine text corresponding to the input speech from among the plurality of candidate texts using a trained personal language model, and output the text as a result of speech recognition of the input speech.


