ASR Caption Server for Hearing Impaired Live Broadcasts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current remote live broadcasting and teaching lack captions, making it difficult for hearing-impaired individuals to participate, as traditional listener-typists face fatigue and errors due to prolonged listening and typing.
Innovation Solution
A caption service system utilizing an automatic speech recognition (ASR) caption server, connected via network with live broadcast equipment, a listener-typist, and a live screen, which uses RTMP and an open-source speech recognition toolkit like Kaldi to convert audio to text, allowing real-time captioning and correction, and merging with video and audio for display on a live screen.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a listener-typist manually types the speaker's content in real-time, then captions are provided for hearing impaired individuals, but the listener-typist experiences fatigue and makes errors due to prolonged listening and typing
Solution Approach 1:
The patent introduces an ASR caption server as an intermediary between the speaker and the listener-typist. The server automatically transcribes speech to text, reducing the listener-typist's workload from complete manual transcription to only correcting and verifying ASR-generated captions. This mediator handles the bulk of the transcription work, improving both accuracy and sustainability.
Solution Approach 2:
The system enables self-service by allowing the listener-typist to review and correct captions at their own pace without the pressure of real-time complete transcription. The asynchronous nature of the correction process allows the listener-typist to work more efficiently and accurately without experiencing the same level of fatigue.
2Loss of information
If manual captioning by listener-typist is used, then captions are generated, but missed sentences and typos occur due to working hours exceeding capacity
Solution Approach 1:
The ASR caption server acts as a mediator that handles the time-consuming transcription task, allowing captions to be generated and corrected asynchronously. This eliminates the direct constraint between working hours and caption completeness, as the system can process and generate captions beyond human endurance limits.
Solution Approach 2:
The system performs preliminary transcription action automatically through ASR before the listener-typist needs to review and correct. This preliminary action by the server reduces the total time and effort required from the listener-typist while ensuring no content is missed during the correction phase.
3Productivity
If traditional manual captioning is used without automation, then real-time captions can be displayed, but the system complexity increases and efficiency decreases
Solution Approach 1:
The patent introduces an ASR caption server as a dedicated intermediary component that handles speech-to-text conversion. This modular architecture separates the complex ASR functionality from the simple caption review and display tasks, making the overall system more manageable despite the added complexity of automated recognition.
Solution Approach 2:
The system replaces the mechanical manual typing process with automated speech recognition technology. This substitution reduces the physical and cognitive burden on the listener-typist while improving captioning efficiency, even though it introduces computational complexity to the system.
Data Source
AI summary
The present invention provides a caption service system for remote speech recognition, which provides caption service for the hearing impaired. This system includes a speaker and a live broadcast equipment at A, a listener-typist and a computer at B, a hearing impaired and a live screen at C, and an automatic speech recognition (ASR) caption server at D. Connect the live broadcast equipment, the computer, the live screen and the ASR caption server with a network. The speaker's audio is sent to the automatic speech recognition (ASR) caption server to be converted into text, which is corrected by the listener-typist, and then the text caption is sent to the live screen of the hearing impaired together with the speaker's video and audio, so that the hearing impaired can see the text caption spoken by the speaker.


