Personalized Speech Recognition via Language Model Interpolation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in personalizing language models for individual users, leading to reduced accuracy when users speak infrequently used words, as they rely on a single general language model for all users, which is not adaptable to individual speech patterns and preferences.
Innovation Solution
A speech recognition apparatus and method that identifies a language model group based on user-specific characteristic data, including static and dynamic information, to generate a user-based language model by interpolating a general language model, enhancing speech recognition accuracy by tailoring the model to individual users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single general language model is used for all users, then the device complexity is reduced, but the speech recognition accuracy deteriorates for users who speak less frequently used words
Solution Approach 1:
The patent segments the single general language model into multiple user-specific language models, each tailored to individual speech patterns and vocabulary usage. This segmentation allows the system to maintain lower overall complexity by using interpolation techniques while significantly improving recognition accuracy for each user segment.
Solution Approach 2:
The patent changes the parameters of the language model by introducing user-specific characteristics such as frequently used words, speech patterns, and personal vocabulary. These parameter changes enable the system to adapt to individual users without requiring completely separate models for each user.
2Measurement precision
If user-specific language models are created for each user, then speech recognition accuracy is improved, but the device complexity increases
Solution Approach 1:
The patent implements dynamic language model generation where user-specific models are created on-demand based on available user data. The system dynamically adjusts the interpolation weights between general and user-specific models, allowing flexibility without permanently storing multiple complete models, thus managing complexity while maintaining accuracy.
Solution Approach 2:
The patent uses interpolation as an intermediary technique that combines the general language model with user-specific adaptations. This intermediary approach allows the system to benefit from user personalization without fully committing to storing and managing complete separate models for each user, thereby controlling device complexity.
3Quantity of substance
If a general language model is used for all users, then the model size is reduced, but the adaptability to individual speech patterns deteriorates
Solution Approach 1:
The patent applies partial action by generating only the necessary user-specific portions of the language model through interpolation, rather than creating complete separate models. This approach provides sufficient adaptability for individual speech patterns while keeping the overall model size manageable by using the general model as a base.
4Measurement precision
If user characteristic data is collected and processed, then personalization accuracy is improved, but the processing time increases
Solution Approach 1:
The patent performs preliminary action by collecting and processing user characteristic data in advance to build user profiles. This preprocessing allows the system to quickly generate personalized language models when needed, reducing the time penalty during actual speech recognition tasks while maintaining high personalization accuracy.
Data Source
AI summary
An apparatus includes a language model group identifier configured to identify a language model group based on determined characteristic data of a user, and a language model generator configured to generate a user-based language model by interpolating a general language model for speech recognition based on the identified language model group.


