Voice Input Accuracy via Word Segment Probability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice input technologies face inaccuracies in recognizing homophones, leading to lower accuracy and user experience issues due to the inability to effectively handle ambiguous voice inputs.
Innovation Solution
A method and device system that determines an input character sequence based on a voice recognition model, calculates appearance probability information for word segments, and transmits this information to a user device, allowing for improved accuracy and flexibility in voice input by providing alternative options and context-based corrections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If voice recognition technology is used for input, then input efficiency is improved, but recognition accuracy deteriorates due to homophone ambiguity
Solution Approach 1:
The patent segments the voice recognition process into multiple stages: initial voice recognition to obtain candidate word sequences, appearance probability calculation to evaluate likelihoods, and correction mechanisms to refine results. This segmentation allows the system to handle homophone ambiguity systematically by breaking down the complex recognition task into manageable components that can be processed and corrected separately.
Solution Approach 2:
The patent implements feedback mechanisms where appearance probability information is calculated and fed back into the correction process. The system uses this feedback to identify and correct recognition errors, particularly for homophones, by comparing predicted probabilities against actual usage patterns and allowing user corrections to be learned and applied to future recognitions.
2Device complexity
If voice input is processed without accuracy information, then system complexity is reduced, but user experience deteriorates due to frequent errors
Solution Approach 1:
The patent performs preliminary calculations of appearance probability information during the voice recognition process itself, rather than adding it as a separate post-processing step. This preliminary action integrates accuracy assessment into the recognition flow, allowing the system to prepare correction data in advance without significantly increasing overall system complexity.
Solution Approach 2:
The patent introduces appearance probability information as an intermediary element between voice recognition and final output. This intermediary layer provides a bridge that enables accuracy correction without requiring complete redesign of the core recognition system, allowing the system to maintain simplicity while improving reliability through this intermediate correction layer.
3Speed
If homophone recognition is handled traditionally, then processing speed is maintained, but accuracy information is lost leading to incorrect word entries
Solution Approach 1:
The patent merges the voice recognition process with appearance probability calculation into a unified processing pipeline. By combining these functions, the system maintains processing speed while simultaneously generating accuracy information, eliminating the need for separate slow processing stages and preventing information loss through integrated handling of both recognition and probability assessment.
Data Source
AI summary
A network device for implementing voice input comprises an input-obtaining module for obtaining voice input information, a sequence-determining module for determining an input character sequence corresponding to the voice input information based on a voice recognition model, an accuracy-determining module for determining appearance-probability information corresponding to word segments in the input character sequence so as to obtain accuracy information of the word segments, and a transmitting module for transmitting, to a user device, the input character sequence and the accuracy information of the word segments corresponding to the voice input information.


