Speech-to-text system with automatic error correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-to-text automation systems generate imprecise transcriptions, leading to erroneous word replacements during re-transcription attempts, which can result in additional charges due to repeated transmissions to third-party ASR services, lacking contextualization and efficient error correction.

Innovation Solution

A speech processing system that automatically receives a second utterance to correct an erroneously transcribed word, transmitting the audio file and location information to a speech recognition system for improved transcription, reducing the need for repeated transmissions and enhancing contextual understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If re-transmission is performed to correct erroneous transcriptions, then transcription accuracy is improved, but transmission costs increase due to repeated charges from third-party ASR services

Engineering Contradiction:
Improvetranscription accuracyVSAvoidtransmission cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent extracts only the erroneously transcribed word and its contextual information from the full audio file, transmitting minimal data to the ASR service for correction. This selective extraction approach maintains transcription accuracy improvement while significantly reducing re-transmission costs compared to sending entire audio files again.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary error identification and flags erroneous transcribed words before re-transmission. By pre-processing the transcription to identify specific errors and preparing contextual information in advance, the system optimizes the re-transmission process to achieve better accuracy with reduced computational and financial overhead.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If re-transmission is performed to correct erroneous transcriptions, then transcription accuracy is improved, but processing time increases due to multiple transcription attempts

Engineering Contradiction:
Improvetranscription accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts only the specific erroneous segments and their minimal context for re-transmission, rather than re-processing entire audio files. This targeted approach significantly reduces the time required for each correction cycle while maintaining the ability to improve transcription accuracy through focused re-transcription of problem areas.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If contextual information is utilized for error correction, then transcription accuracy is improved, but system complexity increases due to additional processing requirements

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies contextualization locally to only the erroneously transcribed words and their immediate surrounding text, rather than processing entire documents with complex contextual analysis. This localized approach improves accuracy at error points while keeping overall system complexity manageable by avoiding global re-processing.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11532308B2Speech-to-text system
Publication Date: 2022.12.20 ADEIA GUIDES INC
  • US11532308B2 patent drawing
  • US11532308B2 patent drawing
  • US11532308B2 patent drawing

AI summary

Systems and methods for processing speech transcription in a speech processing system are disclosed. A first transcription of a first utterance is received. In response to receiving an indication of an erroneous transcribed word in the first transcription, a control circuitry automatically activates an audio receiver for receiving a second utterance. In response to receiving the second utterance, an audio file of the second utterance and an indication of a location of the erroneous transcribed word within the first transcription is transmitted to a speech recognition system for a second transcription of the second utterance. Subsequently, the erroneous transcribed word in the first transcription is replaced with a transcribed word from the second transcription.