Personalized Speech Recognition via Language Model Interpolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in personalizing language models for individual users, leading to reduced accuracy when users speak infrequently used words, as they rely on a single general language model for all users, which is not adaptable to individual speech patterns and preferences.

Innovation Solution

A speech recognition apparatus and method that identifies a language model group based on user-specific characteristic data, including static and dynamic information, to generate a user-based language model by interpolating a general language model, enhancing speech recognition accuracy by tailoring the model to individual users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single general language model is used for all users, then the device complexity is reduced, but the speech recognition accuracy deteriorates for users who speak less frequently used words

Engineering Contradiction:
Improvelanguage model structureVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the single general language model into multiple user-specific language models, each tailored to individual speech patterns and vocabulary usage. This segmentation allows the system to maintain lower overall complexity by using interpolation techniques while significantly improving recognition accuracy for each user segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of the language model by introducing user-specific characteristics such as frequently used words, speech patterns, and personal vocabulary. These parameter changes enable the system to adapt to individual users without requiring completely separate models for each user.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If user-specific language models are created for each user, then speech recognition accuracy is improved, but the device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidlanguage model structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements dynamic language model generation where user-specific models are created on-demand based on available user data. The system dynamically adjusts the interpolation weights between general and user-specific models, allowing flexibility without permanently storing multiple complete models, thus managing complexity while maintaining accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses interpolation as an intermediary technique that combines the general language model with user-specific adaptations. This intermediary approach allows the system to benefit from user personalization without fully committing to storing and managing complete separate models for each user, thereby controlling device complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If a general language model is used for all users, then the model size is reduced, but the adaptability to individual speech patterns deteriorates

Engineering Contradiction:
Improvelanguage model sizeVSAvoiduser-specific speech adaptation
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent applies partial action by generating only the necessary user-specific portions of the language model through interpolation, rather than creating complete separate models. This approach provides sufficient adaptability for individual speech patterns while keeping the overall model size manageable by using the general model as a base.

Inventive Principle:
Principle #16Partial or excessive action

4Measurement precision

If user characteristic data is collected and processed, then personalization accuracy is improved, but the processing time increases

Engineering Contradiction:
Improvepersonalization accuracyVSAvoidmodel generation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by collecting and processing user characteristic data in advance to build user profiles. This preprocessing allows the system to quickly generate personalized language models when needed, reducing the time penalty during actual speech recognition tasks while maintaining high personalization accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10242668B2Speech recognition apparatus and method
Publication Date: 2019.03.26 SAMSUNG ELECTRONICS CO LTD
  • US10242668B2 patent drawing
  • US10242668B2 patent drawing
  • US10242668B2 patent drawing

AI summary

An apparatus includes a language model group identifier configured to identify a language model group based on determined characteristic data of a user, and a language model generator configured to generate a user-based language model by interpolating a general language model for speech recognition based on the identified language model group.