Speech Recognition Transcription Multiplexing Quality Switch
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current transcription technologies for audio communications often face issues with quality fluctuations and word duplication or loss when switching between different speech recognition systems during a communication session, affecting the accuracy and reliability of real-time transcription display.
Innovation Solution
A method that involves multiplexing audio data to be processed by both an automatic speech recognition system and a re-voicing system, allowing for seamless transfer of transcription responsibility based on quality indications, ensuring continuous and improved transcription display without duplication or loss of words.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a single speech recognition system is used for transcription, then device complexity is reduced, but transcription quality becomes unreliable due to quality fluctuations
Solution Approach 1:
The patent combines multiple speech recognition systems (first and second systems) into a unified transcription framework. The system manager coordinates both systems to process audio data simultaneously, merging their capabilities to achieve more reliable transcription quality while managing the complexity through structured coordination protocols.
Solution Approach 2:
The system dynamically switches between the first and second speech recognition systems based on quality indications. The system manager monitors transcription quality in real-time and adjusts which system processes the audio data, creating a dynamic adaptation mechanism that maintains high reliability while managing complexity through selective activation.
2Reliability
If transcription responsibility is switched between different speech recognition systems, then transcription quality can be improved, but word duplication or loss occurs during switching
Solution Approach 1:
The system performs preliminary actions by having the second speech recognition system generate a second transcript in advance while the first system is still active. The system manager compares transcripts from both systems and prepares for seamless switching by identifying quality degradation patterns before they cause word duplication or loss, enabling proactive quality maintenance.
Solution Approach 2:
The system implements feedback mechanisms where quality indications from the first speech recognition system trigger comparisons with the second system's transcript. This feedback loop allows the system manager to detect quality issues and switch systems appropriately, preventing word duplication or loss while maintaining improved transcription quality.
3Measurement precision
If audio data is processed by multiple speech recognition systems simultaneously, then transcription accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent segments the audio processing task by dividing it between a first speech recognition system and a second speech recognition system. The system manager coordinates these segmented processing tasks, assigning audio data to appropriate systems based on quality indications, thereby improving transcription accuracy while managing processing complexity through structured task division.
Data Source
AI summary
A method may include obtaining audio data originating at a first device during a communication session between the first device and a second device and providing the audio data to a first speech recognition system to generate a first transcript based on the audio data and directing the first transcript to the second device. The method may also include in response to obtaining a quality indication regarding a quality of the first transcript, multiplexing the audio data to provide the audio data to a second speech recognition system to generate a second transcript based on the audio data while continuing to provide the audio data to the first speech recognition system and direct the first transcript to the second device, and in response to obtaining a transfer indication that occurs after multiplexing of the audio data, directing the second transcript to the second device instead of the first transcript.


