Voice Dialogue Error Correction Through Candidate Term Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice dialogue systems struggle with accurately determining the position of recognition errors in long sentences or sentences with similar pronunciations, leading to incorrect corrections and user frustration.
Innovation Solution
Implementing a multi-modality-based active error correction function that allows users to select and correct terms directly from candidate terms with high confidence using graphical interfaces, keyboards, or gestures, reducing the need for manual keyboard input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is used to correct errors in voice dialogue, then correction capability is improved, but accuracy deteriorates when speech is long or contains repeated content
Solution Approach 1:
The patent segments the correction process into multiple interaction rounds. Instead of requiring complete correction in one speech input, the system divides correction into iterative steps where each round addresses specific error types, breaking down complex recognition tasks into manageable segments that improve overall accuracy
Solution Approach 2:
The system performs preliminary analysis of the original speech to identify potential error positions and candidate corrections before presenting them to the user. This preliminary action allows the system to pre-process the correction candidates and present only relevant options, improving recognition accuracy in subsequent interaction rounds
2Measurement precision
If users clearly describe wrong characters and correction targets, then correction precision is improved, but operation complexity increases
Solution Approach 1:
The system automatically identifies potential error positions and generates candidate corrections without requiring users to explicitly describe what is wrong. The correction system serves itself by autonomously analyzing the speech, locating errors, and presenting correction options, thereby maintaining high precision while reducing operational complexity
Solution Approach 2:
The system provides feedback by presenting multiple candidate corrections with confidence scores to the user. This feedback mechanism allows users to select from pre-analyzed options rather than formulating complex correction descriptions, achieving both high precision and ease of operation through iterative interaction
Data Source
Figure 1~2
Figure 3~4
AI summary
Disclosed are method and apparatus for correcting voice dialogue, including: recognizing first text information of a dialogue speech input by a user, including a first semantic keyword determined from a plurality of candidate terms; feeding back a first result with the first semantic keyword to the user based on the first text information; feeding back the plurality of candidate terms to the user in response to the user's selection of the first semantic keyword from the first result; and receiving a second semantic keyword input by the user, correcting the first text information based on the second semantic keyword, determining corrected second text information, and feeding back a second result with the second semantic keyword to the user based on the second text information. The problem of true ambiguity can be solved, while improving the fault tolerance and processing capability of the dialogue apparatus for corresponding errors.