Word-Level Candidate Generation for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in selecting the most suitable sentence from multiple candidate sentences provided by speech recognition systems, especially in limited display spaces of mobile terminals, leading to inconvenience in choosing the correct output.
Innovation Solution
A speech recognition system and method that includes a speech recognition result verifying unit and a word sequence displaying unit, where the system visually distinguishes candidate words within a word sequence, allowing users to replace words with their corresponding candidate words upon selection, thereby simplifying the selection process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all candidate sentences are displayed to the user, then the user can see all options, but the display space is insufficient and the user interface becomes complex
Solution Approach 1:
The patent segments the candidate sentences into individual word-level candidates. Instead of displaying complete candidate sentences, the system identifies and displays only the words that differ among candidates, allowing users to select alternative words to construct different sentence variations. This segmentation resolves the contradiction by preserving information about multiple candidates while using minimal display space.
Solution Approach 2:
The patent transitions from displaying sentences in one dimension (complete text) to displaying word candidates in another dimension (interactive selection). By presenting words that can be clicked or selected to generate different sentence interpretations, the system provides multiple sentence options without occupying additional display space, effectively moving the information representation to an interactive dimension.
2Measurement precision
If multiple candidate sentences are provided, then recognition accuracy improves, but user selection difficulty increases
Solution Approach 1:
The patent extracts only the critical differing words from multiple candidate sentences and presents them as selectable candidates. Instead of requiring users to compare and select from multiple complete sentences, the system extracts and displays only the words that create variations, significantly reducing the cognitive load and selection difficulty while maintaining the benefit of multiple recognition options.
Solution Approach 2:
The patent inverts the traditional approach by not asking users to select from complete sentences, but rather to select individual words that will automatically generate the corresponding sentence interpretations. This inversion simplifies the user's task from comparing multiple full sentences to selecting single words, greatly improving ease of operation while preserving recognition accuracy.
3Ease of operation
If candidate words are highlighted for selection, then user interaction is simplified, but the system complexity increases
Solution Approach 1:
The patent implements a self-service mechanism where the speech recognition system automatically identifies and highlights the differing words that need user attention. Instead of requiring complex user instructions or manual navigation through candidate sentences, the system autonomously determines which words to present as candidates and handles the generation of alternative sentences automatically when users select different options, simplifying user interaction despite increased processing complexity.
Data Source
AI summary
A speech recognition system and method based on word-level candidate generation are provided. The speech recognition system may include a speech recognition result verifying unit to verify a word sequence and a candidate word for at least one word included in the word sequence when the word sequence and the candidate word are provided as a result of speech recognition. A word sequence displaying unit may display the word sequence in which the at least one word is visually distinguishable from other words of the word sequence. The word sequence displaying unit may display the word sequence by replacing the at least one word with the candidate word when the at least one word is selected by a user.


