Speech Recognition Adaptation via Automatic Text Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional electronic devices providing speech recognition services face challenges in accurately recognizing user-specific utterance characteristics, requiring users to invest time and effort in updating speech databases, which is inconvenient and may not properly account for individual differences in speech patterns.
Innovation Solution
An electronic device that acquires and processes speech data using automatic speech recognition (ASR) and natural language understanding (NLU) to identify and correct speech recognition results based on pre-stored text, allowing for the acquisition of training data without additional user effort, thereby adapting to user-specific utterance characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech recognition services are provided using general speech databases, then the system can operate with existing resources, but the recognition accuracy for user-specific utterance characteristics deteriorates
Solution Approach 1:
The system automatically collects speech data during normal device operation and performs self-training without requiring user intervention. The processor identifies speech recognition failures, collects corresponding speech data, and retrains the recognition model autonomously, making the system serve itself rather than requiring user service.
Solution Approach 2:
The system establishes a feedback loop where speech recognition results are continuously evaluated, and when recognition failures are detected, the corresponding speech data is collected and used to retrain the recognition model. This closed-loop feedback mechanism continuously improves recognition accuracy based on actual usage patterns.
2Measurement precision
If users manually update speech databases to improve recognition accuracy, then user-specific speech patterns can be captured, but the time and effort required from users increases
Solution Approach 1:
The system performs preliminary data collection during normal operation by monitoring and storing speech data that results in recognition failures. This preliminary collection of relevant training data occurs automatically in the background, preparing the data needed for model retraining before it is actually needed, eliminating the need for users to manually gather speech samples.
Solution Approach 2:
The system autonomously identifies when retraining is needed, collects the necessary speech data, and performs model retraining without user involvement. The entire process of improving user-specific recognition accuracy is automated, converting a manual user task into an autonomous system function.
3Measurement precision
If speech recognition systems use general databases without user-specific adaptation, then the system complexity remains low, but the recognition accuracy for individual users deteriorates
Solution Approach 1:
The speech database is segmented into general speech data and user-specific speech data. The system maintains a base recognition model trained on general data and creates separate user-specific adaptation models. This segmentation allows the system to handle individual user characteristics without completely redesigning the entire speech recognition architecture.
Solution Approach 2:
The speech recognition system transitions from a static general database to a dynamic structure that automatically adapts to individual users. The system dynamically collects user-specific speech data, retrains models based on individual patterns, and continuously evolves the recognition accuracy for each user without requiring manual configuration.
Data Source
AI summary
An electronic device according to an embodiment may include a microphone, a memory, and at least one processor(s). According to an embodiment, the at least one processor may be configured to acquire speech data corresponding to a user's speech via the microphone. The at least one processor according to an embodiment may be configured to acquire first text recognized on speech data by at least partially performing automatic speech recognition and/or natural language understanding. The at least one processor according to an embodiment may be configured to identify, based on the first text, second text stored in the memory. The at least one processor according to an embodiment may be configured to control to output the first text or the second text as a speech recognition result of the speech data, based on a difference between the first text and the second text. The at least one processor according to an embodiment may be configured to acquire training data for recognition of the user's speech, based on relevance between the first text and the second text with respect to the speech data.


