Semi-Automated Relay Captioning for Faster Accurate Transcription
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing relay systems for hearing-impaired individuals face challenges with accuracy, speed, and cost due to reliance on human captioners, and automated voice-to-text transcription systems struggle with limited frequency ranges and noise in telephone communications, requiring extensive training and causing delays.
Innovation Solution
A semi-automated relay system that combines automated voice-to-text software with human captioners, using remote ASR servers and local device processors to enhance accuracy and efficiency, allowing for automated transcription with human intervention when necessary, and storing voice models for improved future recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated voice-to-text transcription software is used, then transcription speed increases, but accuracy deteriorates due to limited frequency ranges and noise in telephone communications
Solution Approach 1:
The system segments the transcription task between automated software (for speed) and human captioners (for accuracy). The automated software handles initial transcription while human captioners review and correct errors, creating a divided workflow that leverages both automated efficiency and human precision.
Solution Approach 2:
Human captioners serve as an intermediary between the automated transcription software and the final output. They review automated transcripts, correct errors caused by noise and frequency limitations, and ensure accuracy before final delivery, mediating between automated speed and reliable accuracy.
2Measurement precision
If human captioners are used for transcription, then accuracy improves, but cost and time consumption increase
Solution Approach 1:
Human captioners perform partial transcription work rather than complete transcription. They focus only on reviewing and correcting automated transcripts rather than transcribing everything from scratch, reducing their time and cost burden while maintaining accuracy.
Solution Approach 2:
Human captioners act as quality control intermediaries rather than primary transcribers. Their role is to review and correct automated output, which is more efficient than having them perform complete manual transcription while still ensuring high accuracy standards.
3Ease of manufacture
If automated transcription software is used, then cost decreases, but transcription accuracy deteriorates requiring extensive training
Solution Approach 1:
The system uses automated transcription software to handle the bulk of transcription work independently, providing cost-effective service. The software operates without human intervention for initial transcription, reducing operational costs while maintaining acceptable baseline accuracy.
Solution Approach 2:
Human captioners serve as periodic intermediaries to correct automated transcription errors, allowing the system to maintain low operational costs while achieving high accuracy. Their intervention is minimal and targeted rather than continuous, preserving cost efficiency.
4Extent of automation
If automated voice-to-text software is used, then the need for human captioners decreases, but transcription quality suffers due to noise and frequency limitations
Solution Approach 1:
The automated transcription software operates independently to handle transcription tasks, demonstrating high automation level. It processes voice signals and generates transcripts without continuous human involvement, reducing the need for human captioners while maintaining operational reliability.
Solution Approach 2:
Human captioners serve as quality assurance intermediaries who review automated transcripts to ensure reliability. Their periodic intervention corrects errors from noise and frequency limitations, maintaining transcription quality while preserving high automation levels.
Data Source
AI summary
A method to transcribe communications includes the steps of obtaining a plurality of hypothesis transcriptions of a voice signal generated by a speech recognition system, determining consistent words that are included in at least first and second of the plurality of hypothesis transcriptions, in response to determining the consistent words, providing the consistent words to a device for presentation of the consistent words to an assisted user, and presenting the consistent words via a display screen on the device, wherein a rate of the presentation of the words on the display screen is variable.


