Speech Recognition Adaptation via Automatic Text Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional electronic devices providing speech recognition services face challenges in accurately recognizing user-specific utterance characteristics, requiring users to invest time and effort in updating speech databases, which is inconvenient and may not properly account for individual differences in speech patterns.

Innovation Solution

An electronic device that acquires and processes speech data using automatic speech recognition (ASR) and natural language understanding (NLU) to identify and correct speech recognition results based on pre-stored text, allowing for the acquisition of training data without additional user effort, thereby adapting to user-specific utterance characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition services are provided using general speech databases, then the system can operate with existing resources, but the recognition accuracy for user-specific utterance characteristics deteriorates

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiduser effort for database updates
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically collects speech data during normal device operation and performs self-training without requiring user intervention. The processor identifies speech recognition failures, collects corresponding speech data, and retrains the recognition model autonomously, making the system serve itself rather than requiring user service.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system establishes a feedback loop where speech recognition results are continuously evaluated, and when recognition failures are detected, the corresponding speech data is collected and used to retrain the recognition model. This closed-loop feedback mechanism continuously improves recognition accuracy based on actual usage patterns.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If users manually update speech databases to improve recognition accuracy, then user-specific speech patterns can be captured, but the time and effort required from users increases

Engineering Contradiction:
Improveuser-specific speech recognition accuracyVSAvoidtime for manual database updates
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary data collection during normal operation by monitoring and storing speech data that results in recognition failures. This preliminary collection of relevant training data occurs automatically in the background, preparing the data needed for model retraining before it is actually needed, eliminating the need for users to manually gather speech samples.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system autonomously identifies when retraining is needed, collects the necessary speech data, and performs model retraining without user involvement. The entire process of improving user-specific recognition accuracy is automated, converting a manual user task into an autonomous system function.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If speech recognition systems use general databases without user-specific adaptation, then the system complexity remains low, but the recognition accuracy for individual users deteriorates

Engineering Contradiction:
Improveindividual user speech recognition accuracyVSAvoidspeech database structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech database is segmented into general speech data and user-specific speech data. The system maintains a base recognition model trained on general data and creates separate user-specific adaptation models. This segmentation allows the system to handle individual user characteristics without completely redesigning the entire speech recognition architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The speech recognition system transitions from a static general database to a dynamic structure that automatically adapts to individual users. The system dynamically collects user-specific speech data, retrains models based on individual patterns, and continuously evolves the recognition accuracy for each user without requiring manual configuration.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240135925A1Electronic device for performing speech recognition and operation method thereof
Publication Date: 2024.04.25 SAMSUNG ELECTRONICS CO LTD
  • US20240135925A1 patent drawing
  • US20240135925A1 patent drawing
  • US20240135925A1 patent drawing

AI summary

An electronic device according to an embodiment may include a microphone, a memory, and at least one processor(s). According to an embodiment, the at least one processor may be configured to acquire speech data corresponding to a user's speech via the microphone. The at least one processor according to an embodiment may be configured to acquire first text recognized on speech data by at least partially performing automatic speech recognition and/or natural language understanding. The at least one processor according to an embodiment may be configured to identify, based on the first text, second text stored in the memory. The at least one processor according to an embodiment may be configured to control to output the first text or the second text as a speech recognition result of the speech data, based on a difference between the first text and the second text. The at least one processor according to an embodiment may be configured to acquire training data for recognition of the user's speech, based on relevance between the first text and the second text with respect to the speech data.