Semi-Automated Relay Captioning with ASR-Human Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional telecommunication systems for hearing-impaired individuals face challenges in providing accurate, real-time voice-to-text captioning due to limitations in automated transcription software, high costs associated with human operators, and the need for training the software to individual speakers, leading to delays and increased operational expenses.
Innovation Solution
A semi-automated relay system that combines automated voice-to-text transcription with human assistance, using a hybrid approach where automated software transcribes voice messages when accurate and switches to human operators when necessary, leveraging remote ASR servers and local device processors for improved accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated voice-to-text transcription software is used, then transcription speed increases, but accuracy deteriorates
Solution Approach 1:
The system merges automated transcription software with human operator assistance in a hybrid architecture. The automated software handles initial transcription to maintain speed, while human operators review and correct errors to ensure accuracy. This combination allows the system to achieve both high transcription speed and high accuracy simultaneously.
2Measurement precision
If human operators are used for transcription, then accuracy improves, but operational costs increase
Solution Approach 1:
Human operators perform only partial transcription work rather than complete transcription. The system uses automated software for the majority of transcription tasks, with human operators intervening only when needed to correct errors or handle difficult cases. This partial human involvement maintains accuracy while significantly reducing operational costs compared to fully manual transcription.
3Measurement precision
If automated software is trained to individual speakers, then transcription accuracy improves, but processing time increases
Solution Approach 1:
The system performs speaker training in advance before transcription is needed. Voice prints and speaker characteristics are captured and stored in a database during off-peak times or initial setup phases. When transcription is required, the pre-trained models are immediately applied, eliminating training delays during actual transcription operations.
4Quantity of substance
If a hybrid system combining automated and manual transcription is used, then operational costs reduce, but system complexity increases
Solution Approach 1:
The system introduces a intelligent router as an intermediary component that automatically directs transcription tasks to either automated software or human operators based on task characteristics, confidence levels, and resource availability. This intermediary manages the complexity of coordinating multiple transcription methods, handling quality assurance, and optimizing resource allocation, thereby reducing operational costs while keeping the system architecture manageable.
Data Source
AI summary
A method and system for providing captioned telephone service, the method comprising the steps of, initiating a first captioned telephone service call, during the first captioned telephone service call, creating a first set of captions using a call assistant, simultaneous with creating the first set of captions using a call assistant, creating a second set of captions using an automated speech recognition engine, comparing the first set of captions and the second set of captions using a scoring algorithm based on errors between the first and second sets of captions to generate a score for the second set of captions, in response to the score being within a predetermined threshold range, continuing the call using only the automated speech recognition engine to generate text and in response to the score being outside of the predetermined threshold range, continuing the call using a call assistant to generate captions.


