Personal Language Model for Speech Recognition Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems lack accuracy in considering user-specific language and pronunciation variations, leading to suboptimal recognition results.

Innovation Solution

An electronic apparatus equipped with a personal language model and a personal acoustic model that learns from user input text and speech, allowing for improved speech recognition by adjusting probabilities based on user-specific data and retraining models for enhanced accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a general speech recognition module is used, then the system can recognize speech, but the recognition accuracy does not account for user-specific language and pronunciation variations

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiduser-specific adaptation capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The speech recognition system is segmented into multiple components: a general speech recognition module for baseline recognition and a personal language model for user-specific adaptation. This segmentation allows each component to specialize in different aspects, with the personal language model focusing specifically on user characteristics to improve overall accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A personal language model acts as an intermediary between the general speech recognition module and the final recognition result. This intermediary layer processes the general module's output and refines it by applying user-specific language patterns and pronunciation characteristics, thereby improving accuracy without replacing the entire system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a personal language model is introduced to improve recognition accuracy, then user-specific characteristics are considered, but the system complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidmodel structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The personal language model is designed to be multi-functional, serving both as a refinement layer for the general speech recognition module and as a standalone adaptation mechanism. This universality allows the system to handle different user profiles and speech patterns without requiring separate dedicated systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The personal language model is trained using the user's own speech data and language patterns, allowing the system to automatically adapt and improve recognition accuracy without requiring manual configuration or complex external intervention. The model serves itself by continuously learning from user interactions.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If speech recognition is performed without personalization, then the system operates simply, but it cannot provide accurate recognition for individual user characteristics

Engineering Contradiction:
Improvesystem operation simplicityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The personal language model is trained in advance using the user's speech data and language patterns before actual speech recognition tasks. This preliminary action of training and adaptation allows the system to automatically incorporate user characteristics without adding complexity to the real-time operation, maintaining ease of use while improving accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11631400B2Electronic apparatus and controlling method thereof
Publication Date: 2023.04.18 SAMSUNG ELECTRONICS CO LTD
  • US11631400B2 patent drawing
  • US11631400B2 patent drawing
  • US11631400B2 patent drawing

AI summary

An electronic apparatus configured to acquire information on a plurality of candidate texts corresponding to input speech of a user through a general speech recognition module, determine text corresponding to the input speech from among the plurality of candidate texts using a trained personal language model, and output the text as a result of speech recognition of the input speech.