Speech Recognition Text Generator Confidence-Based Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional transcription assistance systems using speech recognition fail to provide accurate and efficient results, leading to increased operator burden due to insufficient adjustment capabilities for transcription accuracy and workload.

Innovation Solution

A text generator system comprising a recognizer, selector, and generation unit that selects recognized character strings based on confidence levels and transcription parameters, allowing for adjusted output and reduced operator workload by synchronizing character input with speech recognition results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech recognition systems are used to assist transcription work, then transcription speed is improved, but transcription accuracy deteriorates

Engineering Contradiction:
Improvetranscription speedVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments the speech recognition results into multiple candidate transcriptions with different confidence levels. The system divides the transcription task into automatic recognition (for speed) and manual verification (for accuracy), processing segments of speech separately and allowing operators to focus on correcting only the low-confidence segments rather than reviewing entire transcriptions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter of confidence level threshold to dynamically adjust between speed and accuracy. By setting different threshold values, the system can automatically determine which recognition results require manual review, thereby balancing transcription speed and accuracy based on the confidence levels of different speech segments.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If automatic speech recognition is used, then operator workload is reduced, but transcription quality deteriorates

Engineering Contradiction:
Improveoperator workloadVSAvoidtranscription quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system introduces an intermediary confidence level assessment mechanism between automatic recognition and final transcription. The confidence level acts as a mediator that determines which segments require operator intervention, allowing the system to automatically handle high-confidence segments (reducing workload) while directing operator attention to low-confidence segments (maintaining quality).

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by providing confidence level information to operators about each recognized segment. This feedback allows operators to efficiently allocate their attention and correction efforts only where needed, rather than uniformly reviewing all transcriptions, thus reducing overall workload while maintaining quality standards.

Inventive Principle:
Principle #23Feedback

3Loss of time

If speech recognition results are used directly, then transcription time is reduced, but correction workload increases

Engineering Contradiction:
Improvetranscription timeVSAvoidcorrection workload
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The system performs preliminary action by pre-assessing and ranking speech recognition results according to their confidence levels before presenting them to operators. This preliminary sorting allows operators to immediately identify which segments need correction without having to evaluate all segments equally, thereby reducing correction workload while maintaining fast transcription time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9460718B2Text generator, text generating method, and computer program product
Publication Date: 2016.10.04 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9460718B2 patent drawing
  • US9460718B2 patent drawing
  • US9460718B2 patent drawing

AI summary

According to an embodiment, a text generator includes a recognizer, a selector, and a generation unit. The recognizer is configured to recognize an acquired sound and obtain recognized character strings in recognition units and confidence levels of the recognized character strings. The selector is configured to select at least one of the recognized character strings used for a transcribed sentence on the basis of at least one of a parameter about transcription accuracy and a parameter about a workload needed for transcription. The generation unit is configured to generate the transcribed sentence using the selected recognized character strings.