Speech Recognition Training With Hybrid ASR Transcription Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio transcription systems for hard-of-hearing individuals suffer from inaccuracies and delays due to reliance on human revoicing and speaker-dependent automatic speech recognition (ASR) systems, leading to increased costs and reduced effectiveness.
Innovation Solution
Implementing a hybrid system that combines speaker-dependent and speaker-independent ASR systems with real-time transcription fusion and human revoicing, utilizing revoiced and regular audio to generate accurate and synchronized transcriptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human revoicing is used for transcription, then transcription accuracy may be improved, but cost and time delay increase
Solution Approach 1:
The patent segments the transcription process into multiple parallel ASR systems (speaker-dependent and speaker-independent) that operate simultaneously on different audio inputs (regular audio and revoiced audio). This segmentation allows the system to process multiple transcription paths in parallel, reducing overall delay while maintaining accuracy through fusion of results.
Solution Approach 2:
The patent merges outputs from multiple ASR systems and audio sources through a fusion process. By combining transcriptions from speaker-dependent ASR, speaker-independent ASR, revoiced audio processing, and regular audio processing, the system achieves higher accuracy than any single method alone while reducing reliance on sequential human review.
2Measurement precision
If speaker-dependent ASR systems are used, then transcription accuracy improves, but system complexity and cost increase
Solution Approach 1:
The patent implements a universal transcription system that handles multiple speaker types and communication scenarios through a single integrated architecture. The system processes both speaker-dependent and speaker-independent ASR inputs, along with revoiced and regular audio, making it adaptable to various users and situations without requiring separate specialized systems for each case.
Solution Approach 2:
The patent introduces a fusion module as an intermediary that receives and integrates outputs from multiple ASR systems. This mediator component simplifies the overall system architecture by providing a centralized point for combining results, managing complexity, and producing a unified transcription output without requiring direct integration between all individual ASR components.
3Productivity
If multiple ASR systems are implemented with real-time fusion, then transcription accuracy and speed improve, but computational resources and system complexity increase
Solution Approach 1:
The patent performs preliminary processing of audio signals through revoicing before they reach the ASR systems. By pre-processing audio to improve speech clarity and consistency, the system reduces the computational burden on subsequent ASR processing stages, allowing faster and more accurate transcription with reduced overall computational requirements.
Data Source
AI summary
A method may include obtaining first audio data of a first communication session between a first and second device and during the first communication session, obtaining a first text string that is a transcription of the first audio data and training a model of an automatic speech recognition system using the first text string and the first audio data. The method may further include in response to completion of the training, deleting the first audio data and the first text string and after deleting the first audio data and the first text string, obtaining second audio data of a second communication session between a third and fourth device and during the second communication session obtaining a second text string that is a transcription of the second audio data and further training the model of the automatic speech recognition system using the second text string and the second audio data.


