Speech Recognition Error Correction Using Whole or Partial Re-utterance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems lack flexibility in error correction, often requiring users to re-input entire utterances or specific parts, which can be inefficient and increase operational load, especially when dealing with burst errors or minor errors.
Innovation Solution
A speech recognition apparatus and method that determines whether a re-input speech is for whole or partial correction, adjusting the correction process accordingly, allowing for flexible selection between whole or part correction based on the type of error, using a unit configured to generate recognition candidates, store likelihoods, and determine the relation between initial and re-input utterances to correct recognition candidates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the user re-inputs the entire utterance for correction, then the recognition accuracy can be improved, but the operational load on the user increases
Solution Approach 1:
The patent segments the correction process into two distinct modes: whole utterance correction and partial utterance correction. The system automatically determines which segment (whole or part) to correct based on the error type detected, thereby reducing unnecessary user input while maintaining recognition accuracy.
Solution Approach 2:
The patent implements a dynamic correction mechanism where the system adapts its correction approach based on the detected error type. When burst errors are detected, whole utterance correction is applied; when isolated errors are detected, partial correction is applied. This dynamic adaptation optimizes both accuracy and user convenience.
2Measurement precision
If the system always performs whole utterance correction, then recognition accuracy is maintained, but the efficiency of correction decreases
Solution Approach 1:
The patent applies partial correction when only specific portions of the utterance contain errors. By correcting only the erroneous segments rather than the entire utterance, the system achieves the necessary recognition accuracy while significantly improving correction efficiency and reducing user burden.
Solution Approach 2:
The system dynamically selects between whole and partial correction based on error analysis. This dynamic approach ensures that correction efficiency is optimized by applying the minimal necessary correction action while maintaining recognition accuracy.
3Ease of operation
If the system always performs partial correction, then the operational load on the user is reduced, but the ability to handle burst errors decreases
Solution Approach 1:
The patent implements a dynamic error handling mechanism that automatically detects the error type (isolated error vs. burst error) and switches between partial and whole correction modes accordingly. This ensures reliable handling of all error types while maintaining ease of operation.
Solution Approach 2:
The system changes the correction parameter (scope of correction) based on the detected error characteristics. When burst errors are detected, the correction scope is expanded to cover the entire utterance; when isolated errors are detected, the correction scope is limited to the specific erroneous portion.
4Adaptability or versatility
If the system provides flexible correction modes, then user convenience is improved, but the system complexity increases
Solution Approach 1:
The patent implements a self-service correction system where the system automatically determines the appropriate correction mode (whole or partial) based on its own error detection capabilities. The user simply provides the re-input without needing to specify the correction mode, thereby achieving flexibility without increasing user-side complexity.
Solution Approach 2:
The system acts as an intermediary that automatically processes the decision between whole and partial correction. The error detection module analyzes the re-input and mediates the selection of correction scope, shielding the user from system complexity while providing adaptable correction services.
Data Source
AI summary
A speech recognition apparatus includes a generation unit generating a recognition candidate associated with a speech utterance and a likelihood; a storing unit storing the one recognition; a selecting unit selecting the recognition candidate as a recognition result of a first speech utterance; an utterance relation determining unit determining whether a second speech utterance which is input after the input of the first speech utterance is a speech re-utterance of a whole of the first speech utterance or a speech re-utterance of a part of the first speech utterance; a whole correcting unit correcting the recognition candidate of the whole of the first speech utterance when the second speech utterance is the whole of the first speech utterance; and a part correcting unit correcting the recognition candidate for the part of the first speech utterance when the second speech utterance is the part of the first speech utterance.


