Speech Recognition Error Correction via Acoustic-Text Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition error correction methods require manual selection and input of correction, making the process cumbersome and time-consuming.
Innovation Solution
An apparatus and method that utilize text and acoustic matching to automatically estimate and correct speech recognition errors by acquiring and converting spoken content, allowing for real-time correction of captions through text matching and acoustic analysis, enabling automated identification and replacement of incorrect character strings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual selection and input methods are used for correcting speech recognition errors, then correction accuracy can be maintained, but operational complexity and time consumption increase significantly
Solution Approach 1:
The system performs self-correction by automatically detecting recognition errors and generating correction suggestions without requiring manual selection or input. The correction suggestion generation unit creates candidate corrections based on the recognized text and context, allowing the system to service itself rather than relying on manual intervention.
Solution Approach 2:
The patent replaces the mechanical manual input process (keyboard typing, mouse selection) with an automated information processing system. The correction suggestion generation unit substitutes the mechanical interaction between user and system with an automated algorithmic process that analyzes recognition results and generates corrections programmatically.
2Measurement precision
If manual correction processes are used, then precise error identification is possible, but time consumption and operational cost increase
Solution Approach 1:
The system performs preliminary analysis of the recognized text to identify potential errors before final correction is applied. The correction suggestion generation unit proactively generates correction suggestions based on the recognition result, preparing corrections in advance rather than waiting for manual error identification after the fact.
Solution Approach 2:
The system uses feedback from the speech recognition process itself to identify and correct errors. The correction suggestion generation unit analyzes the recognition result and generates corrections based on this feedback, creating a closed-loop system where recognition outcomes inform subsequent correction actions automatically.
3Productivity
If automated correction is implemented, then operational efficiency improves, but system complexity increases
Solution Approach 1:
The system segments the correction process into distinct functional units: the correction suggestion generation unit that creates candidate corrections, and the correction suggestion output unit that presents these suggestions. This segmentation allows automated correction functionality to be added as modular components rather than requiring complete system redesign.
Solution Approach 2:
The correction suggestion generation unit acts as an intermediary between the speech recognition process and the final correction output. Rather than directly implementing complex automated correction, the system introduces this intermediate layer that generates suggestions, which then can be reviewed or automatically applied, mediating between automation and control.
Data Source
AI summary
An apparatus for correcting a character string in a text of an embodiment includes a first converter, a first output unit, a second converter, an estimation unit, and a second output unit. The first converter recognizes a first speech of a first speaker, and converts the first speech to a first text. The first output unit outputs a first caption image indicating the first text. The second converter recognizes a second speech of a second speaker for correcting a character string to be corrected in the first text, and converts the second speech to a second text. The estimation unit estimates the character string to be corrected, based on text matching between the first text and the second text. The second output unit outputs a second caption image indicating that the character string to be corrected is to be replaced with the second text.


