Speech Recognition Interface Correction Mechanism
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems on mobile devices lack an intuitive and efficient interface, leading to frustration due to interruptions in the dictation process caused by recognition mistakes.
Innovation Solution
A user interface that utilizes a presence-sensitive input device to switch between dictation and correction modes by moving an input object between designated icons on the screen, allowing for seamless correction of speech recognition errors without interrupting the dictation flow.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional speech recognition interface is used on mobile device, then speech to text conversion is achieved, but speech recognition mistakes interrupt continuity of dictation process
Solution Approach 1:
The interface is segmented into distinct functional zones: a first location for dictation control and a second location for correction control. This segmentation allows independent handling of dictation continuity and error correction, enabling users to maintain dictation flow while correcting mistakes without interrupting the overall process.
Solution Approach 2:
The patent introduces an intermediary correction mechanism that acts as a mediator between the speech recognition system and the user. By providing a dedicated correction location that operates independently, it mediates the conflict between maintaining dictation continuity and correcting recognition errors, allowing both functions to coexist without interruption.
2Reliability
If speech recognition mistakes are corrected, then accuracy is improved, but dictation flow is interrupted
Solution Approach 1:
The system performs preliminary error detection and highlights potential recognition mistakes during the dictation process. By identifying and marking errors in advance while maintaining dictation flow, users can correct them at their convenience without interrupting the overall dictation continuity, thus reducing time loss.
Solution Approach 2:
The patent maintains continuity of the dictation process by allowing error correction to occur without stopping the recording or breaking the dictation flow. The correction function operates in parallel with the dictation function, ensuring that the useful action of dictation continues uninterrupted while accuracy is improved through post-hoc corrections.
3Reliability
If traditional correction methods are used, then speech recognition errors are fixed, but user interface complexity increases
Solution Approach 1:
The correction function is extracted as a separate, dedicated location on the interface, distinct from the main dictation controls. This extraction simplifies the interface by providing a clear, intuitive separation between dictation and correction functions, reducing cognitive load and making the interface easier to use despite the added correction capability.
Solution Approach 2:
The second location on the interface serves multiple functions: it acts as a correction trigger, displays correction options, and manages the correction process. This multi-functionality consolidates what could be multiple separate controls into a single universal interface element, thereby reducing overall interface complexity while maintaining comprehensive error correction capability.
Data Source
AI summary
Certain implementations of the disclosed technology include systems and methods for an enhanced speech recognition interface. According to an example implementation, a method includes outputting a first icon and second icon for presentation on a display device; responsive to receiving an indication of an input object being maintained at a first location of an input device, causing a recording device to record an audio signal; responsive to receiving an indication that the input object has moved across the input device from the first location of the input device to a second location of the input device, causing the recording device to stop recording the audio signal; outputting text, based on the recorded audio signal, for presentation on the display device; and responsive to receiving an indication of the input object being maintained at the second location of the input device, causing a portion of the text to be removed from presentation on the display device.


