Speech Recognition Text Correction via Confirmed Rate Calculation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition devices face challenges in generating 100% reliable text from speech data, requiring manual correction by operators due to unconfirmed parts with low reliability, which is time-consuming and inefficient.

Innovation Solution

A support device with a confirmed rate calculator, candidate obtaining unit, and selector is introduced to calculate the confirmed utterance rate, obtain candidate character strings, and preferentially select the most likely candidate string based on utterance time, reducing the need for manual correction and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech recognition device generates text from speech data, then text generation capability is improved, but reliability of generated text deteriorates due to unconfirmed parts

Engineering Contradiction:
Improvetext generation capabilityVSAvoidreliability of generated text
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the generated text into confirmed parts and unconfirmed parts based on reliability thresholds. This segmentation allows the system to identify specific portions requiring manual correction while maintaining automatic processing for reliable portions, thus resolving the contradiction between productivity and reliability.

Inventive Principle:
Principle #1Segmentation

2Reliability

If operator manually corrects unconfirmed text parts, then reliability of text is improved, but working hours increase

Engineering Contradiction:
Improvereliability of corrected textVSAvoidworking hours for correction
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by requiring manual correction only for unconfirmed parts that fall below a reliability threshold, rather than requiring correction of the entire text. This reduces the operator's workload and time investment while maintaining sufficient reliability for confirmed portions.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If speech recognition device creates multiple candidate character strings, then selection accuracy is improved, but operator workload increases due to enormous number of candidates

Engineering Contradiction:
Improveselection accuracyVSAvoidoperator workload for candidate selection
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent applies local quality by differentiating between confirmed and unconfirmed parts of the text. For unconfirmed parts, the system presents only relevant candidate character strings rather than all possible candidates, reducing operator workload while maintaining selection accuracy for the specific problematic regions.

Inventive Principle:
Principle #3Local quality

4Manufacturing precision

If operator sequentially corrects text from beginning, then correction completeness is improved, but efficiency deteriorates due to repeated listening to speech data

Engineering Contradiction:
Improvecorrection completenessVSAvoidcorrection efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent performs preliminary action by automatically identifying and marking unconfirmed parts before the operator begins correction. This preliminary identification allows the operator to focus directly on problematic sections without sequentially listening to and evaluating all speech data from the beginning, thus improving correction efficiency while maintaining completeness.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8275614B2Support device, program and support method
Publication Date: 2012.09.25 CERENCE OPERATING CO
  • US8275614B2 patent drawing
  • US8275614B2 patent drawing
  • US8275614B2 patent drawing

AI summary

A support device, program and support method for supporting generation of text from speech data. The support device includes a confirmed rate calculator, a candidate obtaining unit and a selector. The confirmed rate calculator calculates a confirmed utterance rate which is an utterance rate of a confirmed part having already-confirmed text in the speech data. The candidate obtaining unit obtains multiple candidate character strings resulting from a speech recognition of an unconfirmed part having unconfirmed text in the speech data. The selector preferentially selects, from among the plurality of candidate character strings, a candidate character string whose utterance time consumed in uttering the candidate character string at the confirmed utterance rate is closest to an utterance time of the unconfirmed part of the speech data.