Real-Time Speech Transcription Using Hypothesis Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio transcription systems for communications often produce inaccurate and delayed transcriptions due to the challenges of phoneme recognition and contextual understanding, particularly in real-time conversations, leading to a lag between spoken words and their display.

Innovation Solution

The method involves using an automated speech recognition system to generate multiple hypothesis transcriptions, identifying consistent words across these transcriptions, and presenting them to the user before finalizing the transcription, thereby reducing the number of words displayed at once and minimizing the delay between spoken and displayed text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple hypothesis transcriptions are generated and analyzed to identify consistent words, then transcription accuracy is improved, but processing time increases

Engineering Contradiction:
Improvetranscription accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system generates multiple hypothesis transcriptions in advance and identifies consistent words before final presentation to the user. By performing the analysis of multiple hypotheses beforehand and presenting only the consistent words, the system prepares the most accurate transcription results in advance, reducing the perceived processing time when results are needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts only the consistent words that appear across multiple hypothesis transcriptions and presents these to the user, rather than presenting all possible transcriptions or waiting for complete processing. This extraction of essential consistent information provides accurate results faster by focusing only on the reliable portions of the transcription.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If complete transcriptions are processed before presentation, then transcription reliability is improved, but real-time responsiveness deteriorates

Engineering Contradiction:
Improvetranscription reliabilityVSAvoidreal-time responsiveness
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The system performs preliminary processing to generate multiple hypothesis transcriptions and identify consistent words as soon as sufficient audio data is available, rather than waiting for complete processing. This allows the system to present reliable consistent words in near real-time while continuing to process additional audio data in the background.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system presents transcription results based on partial processing of audio data as soon as consistent words can be identified from multiple hypotheses, rather than waiting for complete processing of all audio data. This partial action provides timely transcription output while maintaining reliability through the use of multiple hypothesis verification.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11600279B2Transcription of communications
Publication Date: 2023.03.07 SORENSON IP HOLDINGS LLC
  • US11600279B2 patent drawing
  • US11600279B2 patent drawing
  • US11600279B2 patent drawing

AI summary

A method to transcribe communications may include obtaining audio data originating at a first device during a communication session between the first device and a second device and providing the audio data to an automated speech recognition system configured to transcribe the audio data. The method may further include obtaining multiple hypothesis transcriptions generated by the automated speech recognition system. Each of the multiple hypothesis transcriptions may include one or more words determined by the automated speech recognition system to be a transcription of a portion of the audio data. The method may further include determining one or more consistent words that are included in two or more of the multiple hypothesis transcriptions and in response to determining the one or more consistent words, providing the one or more consistent words to the second device for presentation of the one or more consistent words by the second device.