Speech Recognition Text Generator Confidence-Based Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional transcription assistance systems using speech recognition fail to provide accurate and efficient results, leading to increased operator burden due to insufficient adjustment capabilities for transcription accuracy and workload.
Innovation Solution
A text generator system comprising a recognizer, selector, and generation unit that selects recognized character strings based on confidence levels and transcription parameters, allowing for adjusted output and reduced operator workload by synchronizing character input with speech recognition results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech recognition systems are used to assist transcription work, then transcription speed is improved, but transcription accuracy deteriorates
Solution Approach 1:
The patent segments the speech recognition results into multiple candidate transcriptions with different confidence levels. The system divides the transcription task into automatic recognition (for speed) and manual verification (for accuracy), processing segments of speech separately and allowing operators to focus on correcting only the low-confidence segments rather than reviewing entire transcriptions.
Solution Approach 2:
The system changes the parameter of confidence level threshold to dynamically adjust between speed and accuracy. By setting different threshold values, the system can automatically determine which recognition results require manual review, thereby balancing transcription speed and accuracy based on the confidence levels of different speech segments.
2Ease of operation
If automatic speech recognition is used, then operator workload is reduced, but transcription quality deteriorates
Solution Approach 1:
The system introduces an intermediary confidence level assessment mechanism between automatic recognition and final transcription. The confidence level acts as a mediator that determines which segments require operator intervention, allowing the system to automatically handle high-confidence segments (reducing workload) while directing operator attention to low-confidence segments (maintaining quality).
Solution Approach 2:
The system implements feedback by providing confidence level information to operators about each recognized segment. This feedback allows operators to efficiently allocate their attention and correction efforts only where needed, rather than uniformly reviewing all transcriptions, thus reducing overall workload while maintaining quality standards.
3Loss of time
If speech recognition results are used directly, then transcription time is reduced, but correction workload increases
Solution Approach 1:
The system performs preliminary action by pre-assessing and ranking speech recognition results according to their confidence levels before presenting them to operators. This preliminary sorting allows operators to immediately identify which segments need correction without having to evaluate all segments equally, thereby reducing correction workload while maintaining fast transcription time.
Data Source
AI summary
According to an embodiment, a text generator includes a recognizer, a selector, and a generation unit. The recognizer is configured to recognize an acquired sound and obtain recognized character strings in recognition units and confidence levels of the recognized character strings. The selector is configured to select at least one of the recognized character strings used for a transcribed sentence on the basis of at least one of a parameter about transcription accuracy and a parameter about a workload needed for transcription. The generation unit is configured to generate the transcribed sentence using the selected recognized character strings.


