Speech Recognition Text Correction via Confirmed Rate Calculation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition devices face challenges in generating 100% reliable text from speech data, requiring manual correction by operators due to unconfirmed parts with low reliability, which is time-consuming and inefficient.
Innovation Solution
A support device with a confirmed rate calculator, candidate obtaining unit, and selector is introduced to calculate the confirmed utterance rate, obtain candidate character strings, and preferentially select the most likely candidate string based on utterance time, reducing the need for manual correction and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech recognition device generates text from speech data, then text generation capability is improved, but reliability of generated text deteriorates due to unconfirmed parts
Solution Approach 1:
The patent segments the generated text into confirmed parts and unconfirmed parts based on reliability thresholds. This segmentation allows the system to identify specific portions requiring manual correction while maintaining automatic processing for reliable portions, thus resolving the contradiction between productivity and reliability.
2Reliability
If operator manually corrects unconfirmed text parts, then reliability of text is improved, but working hours increase
Solution Approach 1:
The patent applies partial action by requiring manual correction only for unconfirmed parts that fall below a reliability threshold, rather than requiring correction of the entire text. This reduces the operator's workload and time investment while maintaining sufficient reliability for confirmed portions.
3Measurement precision
If speech recognition device creates multiple candidate character strings, then selection accuracy is improved, but operator workload increases due to enormous number of candidates
Solution Approach 1:
The patent applies local quality by differentiating between confirmed and unconfirmed parts of the text. For unconfirmed parts, the system presents only relevant candidate character strings rather than all possible candidates, reducing operator workload while maintaining selection accuracy for the specific problematic regions.
4Manufacturing precision
If operator sequentially corrects text from beginning, then correction completeness is improved, but efficiency deteriorates due to repeated listening to speech data
Solution Approach 1:
The patent performs preliminary action by automatically identifying and marking unconfirmed parts before the operator begins correction. This preliminary identification allows the operator to focus directly on problematic sections without sequentially listening to and evaluating all speech data from the beginning, thus improving correction efficiency while maintaining completeness.
Data Source
AI summary
A support device, program and support method for supporting generation of text from speech data. The support device includes a confirmed rate calculator, a candidate obtaining unit and a selector. The confirmed rate calculator calculates a confirmed utterance rate which is an utterance rate of a confirmed part having already-confirmed text in the speech data. The candidate obtaining unit obtains multiple candidate character strings resulting from a speech recognition of an unconfirmed part having unconfirmed text in the speech data. The selector preferentially selects, from among the plurality of candidate character strings, a candidate character string whose utterance time consumed in uttering the candidate character string at the confirmed utterance rate is closest to an utterance time of the unconfirmed part of the speech data.


