Speech Recognition Transcription Multiplexing Quality Switch

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current transcription technologies for audio communications often face issues with quality fluctuations and word duplication or loss when switching between different speech recognition systems during a communication session, affecting the accuracy and reliability of real-time transcription display.

Innovation Solution

A method that involves multiplexing audio data to be processed by both an automatic speech recognition system and a re-voicing system, allowing for seamless transfer of transcription responsibility based on quality indications, ensuring continuous and improved transcription display without duplication or loss of words.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a single speech recognition system is used for transcription, then device complexity is reduced, but transcription quality becomes unreliable due to quality fluctuations

Engineering Contradiction:
Improvetranscription qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple speech recognition systems (first and second systems) into a unified transcription framework. The system manager coordinates both systems to process audio data simultaneously, merging their capabilities to achieve more reliable transcription quality while managing the complexity through structured coordination protocols.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system dynamically switches between the first and second speech recognition systems based on quality indications. The system manager monitors transcription quality in real-time and adjusts which system processes the audio data, creating a dynamic adaptation mechanism that maintains high reliability while managing complexity through selective activation.

Inventive Principle:
Principle #15Dynamics

2Reliability

If transcription responsibility is switched between different speech recognition systems, then transcription quality can be improved, but word duplication or loss occurs during switching

Engineering Contradiction:
Improvetranscription qualityVSAvoidword duplication or loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system performs preliminary actions by having the second speech recognition system generate a second transcript in advance while the first system is still active. The system manager compares transcripts from both systems and prepares for seamless switching by identifying quality degradation patterns before they cause word duplication or loss, enabling proactive quality maintenance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms where quality indications from the first speech recognition system trigger comparisons with the second system's transcript. This feedback loop allows the system manager to detect quality issues and switch systems appropriately, preventing word duplication or loss while maintaining improved transcription quality.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If audio data is processed by multiple speech recognition systems simultaneously, then transcription accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improvetranscription accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio processing task by dividing it between a first speech recognition system and a second speech recognition system. The system manager coordinates these segmented processing tasks, assigning audio data to appropriate systems based on quality indications, thereby improving transcription accuracy while managing processing complexity through structured task division.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11710488B2Transcription of communications using multiple speech recognition systems
Publication Date: 2023.07.25 SORENSON IP HOLDINGS LLC
  • US11710488B2 patent drawing
  • US11710488B2 patent drawing
  • US11710488B2 patent drawing

AI summary

A method may include obtaining audio data originating at a first device during a communication session between the first device and a second device and providing the audio data to a first speech recognition system to generate a first transcript based on the audio data and directing the first transcript to the second device. The method may also include in response to obtaining a quality indication regarding a quality of the first transcript, multiplexing the audio data to provide the audio data to a second speech recognition system to generate a second transcript based on the audio data while continuing to provide the audio data to the first speech recognition system and direct the first transcript to the second device, and in response to obtaining a transfer indication that occurs after multiplexing of the audio data, directing the second transcript to the second device instead of the first transcript.