Voice Recognition Correction via Letter Spelling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition systems often misinterpret user inputs, leading to incorrect search results and requiring users to reinitiate the voice input process to correct errors, which is inefficient and time-consuming.
Innovation Solution
Implementing a method that allows users to select and correct misrecognized terms by spelling them out letter by letter, using parallel voice recognition processes to score and merge the corrected output with the original, thereby improving recognition accuracy and user experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice recognition processes are used, then the system can process voice inputs, but misrecognition errors occur leading to incorrect results
Solution Approach 1:
The system provides visual feedback by displaying the recognition output to the user, allowing them to identify and select misrecognized terms. The system then uses this user feedback to initiate correction processes, creating a closed-loop feedback mechanism that improves recognition accuracy through iterative correction.
Solution Approach 2:
The patent introduces an intermediary correction interface between the initial voice recognition and the final result. This intermediary layer allows users to review, select, and correct misrecognized terms before the final output is generated, acting as a mediator that resolves the conflict between automated recognition and accuracy.
2Measurement precision
If users correct misrecognized terms by restarting the voice input process, then correction can be achieved, but time is lost
Solution Approach 1:
The system performs preliminary action by displaying the recognition output to the user before final processing occurs. This allows users to identify and select terms needing correction in advance, rather than having to restart the entire process after receiving incorrect results.
Solution Approach 2:
The correction process is segmented into discrete steps: displaying recognition output, user selection of specific terms to correct, receiving correction input, and merging corrected portions. This segmentation allows partial correction of only the problematic terms rather than requiring complete re-input.
3Measurement precision
If parallel voice recognition processes are used, then correction accuracy improves, but system complexity increases
Solution Approach 1:
The patent merges multiple voice recognition processes (original recognition and correction recognition) into a unified output. The correction recognition output is merged with the original recognition output to produce the final corrected result, combining the strengths of both processes while maintaining a relatively simple overall architecture.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for natural language processing. One of the methods includes receiving a first voice input from a user device; generating a first recognition output; receiving a user selection of one or more terms in the first recognition output; receiving a second voice input spelling a correction of the user selection; determining a corrected recognition output for the selected portion; and providing a second recognition output that merges the first recognition output and the corrected recognition output.


