Relay Captioning With Adaptive Voice Delay for ASR Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing relay systems for hearing-impaired individuals face challenges in providing accurate, real-time voice-to-text captioning due to limitations in automated transcription software, high costs associated with human transcribers, and the need for training the software to individual speakers, leading to delays and increased operational costs.
Innovation Solution
A hybrid semi-automated system that combines automated voice-to-text transcription with human assistance, utilizing a relay processor or the user's device to perform initial transcription, assess accuracy, and switch to automated text when quality is met, with voice models for improved accuracy and on-the-fly training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If automated voice-to-text transcription software is used, then operational costs are reduced, but transcription accuracy and real-time performance deteriorate
Solution Approach 1:
The system segments the transcription task into two parts: automated transcription for cost reduction and human review for accuracy assurance. The relay server automatically transcribes voice communications, and the hearing user can request human review when needed, allowing the system to operate in automated mode most of the time while providing human assistance on demand.
Solution Approach 2:
The system introduces an intermediary review mechanism where the hearing user acts as a mediator between the automated transcription system and the final output. The user can review transcribed text and request corrections, creating a feedback loop that improves accuracy without requiring full human transcription coverage.
2Measurement precision
If human transcribers are used, then transcription accuracy is improved, but operational costs increase
Solution Approach 1:
The hearing user serves themselves by reviewing and correcting transcribed text when needed, eliminating the need for continuous human transcription services. The system empowers the user to take control of accuracy for specific communications while relying on automated transcription for routine interactions.
Solution Approach 2:
The system discards the expensive human transcription service for most communications and recovers accuracy only when the hearing user requests it. This on-demand approach allows the system to operate cost-effectively in automated mode while providing human-level accuracy when needed.
3Measurement precision
If automated transcription software is trained to individual speakers, then transcription accuracy is improved, but system complexity and training time increase
Solution Approach 1:
The system uses a universal automated transcription model that can handle multiple speakers without individual training. This multi-functional approach allows the same transcription engine to serve all users and all speakers, eliminating the need for separate training processes while maintaining acceptable accuracy through post-transcription review.
Data Source
AI summary
A captioning method for presenting captions to an assisted user (AU) during communication with a hearing user (HU) where the assisted user uses a captioned device and the hearing user uses a hearing user's device to facilitate the communication, the captioned device including a display screen and a speaker for presenting captions and broadcasting the hearing user's voice signals, respectively, the method comprising the steps of during an ongoing call between the AU and the HU, using an automated speech recognition (ASR) engine to generate initial ASR captions associated with the HU's voice signal, assessing at least one caption quality factor associated with prior initial ASR captions generated during the ongoing call, delaying broadcast of HU voice signal to the AU and based on the at least one caption quality factor, adjusting a duration of the HU voice signal broadcast delay.


