Voice Recognition Ambiguity Resolution via User Selection Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems often misinterpret user commands due to multiple similar words being recognized, leading to incorrect execution commands and a lack of convenience in controlling display apparatuses through voice input.
Innovation Solution
A method and apparatus that enhance voice recognition by extracting similar words based on reliability values, setting a target word that meets predetermined conditions, and displaying a list of similar words, allowing users to select the correct command, while adjusting critical values to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If a critical value threshold method is used to set execution commands, then the voice recognition process is automated, but the recognition accuracy deteriorates when multiple similar words exist
Solution Approach 1:
The system displays recognized similar words to the user and receives feedback through selection input. This feedback loop allows the system to adjust and confirm the correct execution command based on user choice, thereby maintaining automation while improving recognition accuracy when similar words are present.
Solution Approach 2:
The display unit acts as an intermediary between the voice recognition process and the execution command setting. It presents similar words to the user and transmits the selected word back to the controller, serving as a mediator that resolves ambiguity without completely automating the decision process.
2Reliability
If a list of similar words is displayed for user selection, then recognition accuracy is improved, but the ease of operation deteriorates due to additional user steps required
Solution Approach 1:
The system performs partial action by displaying only when multiple similar words are detected. When a single clear match exists, no display is shown and the command is executed directly. This partial application of the display function maintains convenience for clear commands while improving accuracy for ambiguous ones.
Solution Approach 2:
The system changes the parameter of user interaction based on recognition confidence. When confidence is high (single match), the interaction parameter is minimized (no display). When confidence is low (multiple similar words), the interaction parameter is increased (display with selection). This dynamic parameter adjustment balances accuracy and convenience.
3Reliability
If internal voice recognition components are modified to improve recognition rate, then accuracy is improved, but the device complexity increases
Solution Approach 1:
The patent extracts the ambiguity resolution function from the internal voice recognition components and places it in the display and user interaction system. Instead of modifying acoustic models or pronunciation dictionaries, it separates the recognition function from the decision-making function, reducing complexity in the recognition components.
Solution Approach 2:
The display unit serves as an intermediary that handles the complex task of resolving word ambiguity without requiring modifications to the voice recognition components themselves. This mediator approach improves recognition accuracy while keeping the recognition system simple and unchanged.
Data Source
AI summary
A display apparatus which is capable of recognizing a voice and a method thereof are provided. The method includes receiving an uttered voice of a user, extracting a plurality of similar words which are similar to the uttered voice by extracting voice information from the uttered voice and measuring reliability of a plurality of words based on the extracted voice information, setting a word satisfying a predetermined condition from among the plurality of extracted similar words as a target word with respect to the uttered voice, and displaying at least one of the target word and a similar word list including similar words other than the target word. In this manner, a display apparatus may improve a recognition rate on an uttered voice of a user without changing an internal component related to voice recognition, such as, an acoustic model, a pronunciation dictionary, or the like.


