Speech Recognition Text Correction via Similar Pronunciation Patterns
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies often misinterpret user intent due to variations in pronunciation, leading to incorrect recognition and unintended results, causing user inconvenience and the need for re-utterance.
Innovation Solution
An electronic apparatus and method that corrects text obtained from speech recognition by identifying and correcting errors using pattern information to generate similar texts with similar pronunciations, allowing for repeated service operations without re-utterance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech recognition is used to process user input, then the operation speed and convenience are improved, but recognition accuracy deteriorates due to pronunciation variations
Solution Approach 1:
The system performs service operations with recognized text, then uses the operation result as feedback to determine whether to generate and test similar texts. This feedback loop enables automatic correction of recognition errors without requiring user re-utterance, thus maintaining high operation speed while improving recognition accuracy.
Solution Approach 2:
The system preliminarily generates similar texts based on pronunciation patterns before actual service execution fails. By preparing alternative text interpretations in advance, the system can quickly retry operations with corrected text, resolving recognition errors proactively rather than reactively.
2Ease of operation
If speech recognition is used for user input, then ease of operation is improved, but error rate increases due to misinterpretation of user intent
Solution Approach 1:
The system performs self-correction by automatically generating similar texts and retrying service operations when recognition errors are detected. This self-service mechanism eliminates the need for user intervention to correct errors, maintaining ease of operation while significantly reducing the effective error rate through automated verification and correction.
Solution Approach 2:
The system changes the text parameter by generating similar texts with alternative character compositions based on pronunciation patterns. When the original recognized text fails to produce the expected service result, the system modifies the text parameter to similar alternatives, thereby reducing errors caused by pronunciation variations.
3Measurement precision
If similar texts are generated and tested to correct recognition errors, then recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The system applies partial action by generating similar texts only for specific characters that are prone to misrecognition, rather than generating all possible text combinations. This selective approach improves recognition accuracy for error-prone cases while avoiding the exponential complexity increase that would result from exhaustive text generation.
Solution Approach 2:
The system segments the text correction process by identifying and targeting specific misrecognized characters individually. Instead of treating the entire text as a single correction unit, the system divides the problem into character-level segments, generating similar texts only for problematic characters, thereby reducing overall system complexity while maintaining accuracy improvement.
4Measurement precision
If similar texts are generated and sequentially input to correct errors, then user intent matching is improved, but time consumption increases
Solution Approach 1:
The system implements periodic action by attempting service operations with the original recognized text first, then periodically retrying with similar texts only when the initial attempt fails. This periodic retry mechanism improves intent matching accuracy by correcting errors when necessary, while minimizing time consumption by avoiding unnecessary retries when the original recognition is correct.
Data Source
AI summary
Disclosed is a method of controlling an electronic apparatus. The method of controlling an electronic apparatus includes: displaying a screen including an input area configured to receive a text, receiving a speech and obtaining a text corresponding to the speech, performing a service operation corresponding to the input area by inputting the obtained text to the input area, and based on a result of performing the service operation, obtaining a plurality of similar texts including a similar pronunciation with the obtained text, and repeatedly performing the service operation by sequentially inputting the plurality of obtained similar texts to the input area.


