ASR Caption Server for Hearing Impaired Live Broadcasts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current remote live broadcasting and teaching lack captions, making it difficult for hearing-impaired individuals to participate, as traditional listener-typists face fatigue and errors due to prolonged listening and typing.

Innovation Solution

A caption service system utilizing an automatic speech recognition (ASR) caption server, connected via network with live broadcast equipment, a listener-typist, and a live screen, which uses RTMP and an open-source speech recognition toolkit like Kaldi to convert audio to text, allowing real-time captioning and correction, and merging with video and audio for display on a live screen.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a listener-typist manually types the speaker's content in real-time, then captions are provided for hearing impaired individuals, but the listener-typist experiences fatigue and makes errors due to prolonged listening and typing

Engineering Contradiction:
Improvecaption accuracyVSAvoidlistener-typist endurance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces an ASR caption server as an intermediary between the speaker and the listener-typist. The server automatically transcribes speech to text, reducing the listener-typist's workload from complete manual transcription to only correcting and verifying ASR-generated captions. This mediator handles the bulk of the transcription work, improving both accuracy and sustainability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service by allowing the listener-typist to review and correct captions at their own pace without the pressure of real-time complete transcription. The asynchronous nature of the correction process allows the listener-typist to work more efficiently and accurately without experiencing the same level of fatigue.

Inventive Principle:
Principle #25Self-service

2Loss of information

If manual captioning by listener-typist is used, then captions are generated, but missed sentences and typos occur due to working hours exceeding capacity

Engineering Contradiction:
Improvecaption completenessVSAvoidlistener-typist working hours
Core Design Contradiction:
Loss of informationVSDuration of action of moving object

Solution Approach 1:

The ASR caption server acts as a mediator that handles the time-consuming transcription task, allowing captions to be generated and corrected asynchronously. This eliminates the direct constraint between working hours and caption completeness, as the system can process and generate captions beyond human endurance limits.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary transcription action automatically through ASR before the listener-typist needs to review and correct. This preliminary action by the server reduces the total time and effort required from the listener-typist while ensuring no content is missed during the correction phase.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If traditional manual captioning is used without automation, then real-time captions can be displayed, but the system complexity increases and efficiency decreases

Engineering Contradiction:
Improvecaptioning efficiencyVSAvoidsystem architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an ASR caption server as a dedicated intermediary component that handles speech-to-text conversion. This modular architecture separates the complex ASR functionality from the simple caption review and display tasks, making the overall system more manageable despite the added complexity of automated recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces the mechanical manual typing process with automated speech recognition technology. This substitution reduces the physical and cognitive burden on the listener-typist while improving captioning efficiency, even though it introduces computational complexity to the system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11735185B2Caption service system for remote speech recognition
Publication Date: 2023.08.22 NAT YANG MING CHIAO TUNG UNIV
  • US11735185B2 patent drawing
  • US11735185B2 patent drawing
  • US11735185B2 patent drawing

AI summary

The present invention provides a caption service system for remote speech recognition, which provides caption service for the hearing impaired. This system includes a speaker and a live broadcast equipment at A, a listener-typist and a computer at B, a hearing impaired and a live screen at C, and an automatic speech recognition (ASR) caption server at D. Connect the live broadcast equipment, the computer, the live screen and the ASR caption server with a network. The speaker's audio is sent to the automatic speech recognition (ASR) caption server to be converted into text, which is corrected by the listener-typist, and then the text caption is sent to the live screen of the hearing impaired together with the speaker's video and audio, so that the hearing impaired can see the text caption spoken by the speaker.