Speech Recognition Audio Input Editing Modes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face challenges with low accuracy and limited richness in audio input processing, leading to a poor user experience due to immature models and frequent errors in recognition results.
Innovation Solution
An audio input method and terminal device that incorporate two modes: an audio-input mode for initial recognition and an editing mode for processing corrections and additions, allowing users to switch between modes to enhance accuracy and richness of input content, with features like command execution, error correction, and content element addition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition technology is used to convert audio information to text, then audio input functionality is provided, but recognition accuracy is low and user experience is poor
Solution Approach 1:
The patent segments the audio input process into two distinct modes: audio-input mode for initial recognition and editing mode for correction. This segmentation allows different processing strategies to be applied at different stages, improving overall recognition accuracy while maintaining user experience through targeted error correction.
Solution Approach 2:
The patent implements a feedback mechanism where recognition results are displayed to users, who can then provide corrections in editing mode. This feedback loop enables continuous improvement of recognition accuracy by learning from user corrections, thereby enhancing both measurement precision and reliability.
2Productivity
If simple text display is used for audio recognition output, then processing speed is fast, but content richness is limited
Solution Approach 1:
The patent makes the audio input system multi-functional by enabling it to handle not only text display but also editing operations, command execution, and content enrichment. The same audio input interface serves multiple purposes, allowing fast processing while expanding content richness through integrated editing and command capabilities.
Solution Approach 2:
The patent introduces dynamic mode switching between audio-input mode and editing mode. This dynamic adaptability allows the system to adjust its functionality based on user needs, transitioning from simple text display to comprehensive content editing, thereby maintaining processing speed while enhancing content richness.
3Manufacturing precision
If manual selection of content for correction is required, then editing precision is high, but operation complexity increases
Solution Approach 1:
The patent enables self-service editing where the system automatically presents recognition results that users can directly review and correct. This approach maintains high editing precision by allowing user oversight while reducing operation complexity by eliminating the need for manual content selection - users simply review and correct what is already presented to them.
Data Source
AI summary
An audio input method includes: in an audio-input mode, receiving a first audio input by a user, recognizing the first audio to generate a first recognition result, and displaying corresponding verbal content to the user based on the first recognition result; and in an editing mode, receiving a second audio input by the user and recognizing and generating a second recognition result, converting the second recognition result to an editing instruction, and executing a corresponding operation based on the editing operation. The audio-input mode and the editing mode are switchable.


