Live Video Captioning Automation via Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing live video streaming systems face challenges in accurately synchronizing captions with video content, leading to low accuracy and high delays due to manual caption insertion, which affects the overall streaming experience.

Innovation Solution

A method that involves obtaining audio stream data from live video streams, performing speech recognition to generate caption text with corresponding time information, and adding the captions to the video frames based on this information, thereby eliminating the need for manual insertion and reducing delays.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual caption insertion is used, then caption accuracy can be maintained, but streaming delay increases and productivity decreases

Engineering Contradiction:
Improvecaption synchronization accuracyVSAvoidcaptioning speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the manual mechanical process of caption insertion with an automated speech recognition system. The speech recognition module automatically converts audio stream data into caption text, eliminating the need for manual transcription and insertion operations. This substitution dramatically increases productivity while maintaining synchronization accuracy through automated time stamp generation and picture frame association.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual caption insertion is used, then caption quality can be controlled, but time consumption increases

Engineering Contradiction:
Improvecaption qualityVSAvoidcaptioning time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs speech recognition on the audio stream data in advance to generate caption text and time information before the actual captioning process. The system pre-processes the audio data, extracts speech content, generates corresponding time stamps, and prepares caption text ready for insertion. This preliminary action reduces the time required during live streaming while ensuring caption quality through systematic processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The manual process of quality control in captioning is replaced by an automated speech recognition system that consistently processes audio data through standardized algorithms. This substitution eliminates variability in manual transcription quality while maintaining high reliability through the systematic generation of time-synchronized caption text.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated speech recognition is implemented, then productivity increases, but system complexity increases

Engineering Contradiction:
Improvecaptioning speedVSAvoidprocessing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the captioning system into distinct functional modules: an audio stream data acquisition module, a speech recognition module, a caption generation module, and a picture frame association module. Each module performs a specific function independently, making the overall complex system manageable through modular design. This segmentation allows high productivity through automation while controlling system complexity through clear module boundaries and defined interfaces.

Inventive Principle:
Principle #1Segmentation

4Loss of time

If real-time speech recognition is performed, then streaming delay is reduced, but processing complexity increases

Engineering Contradiction:
Improvestreaming delayVSAvoidprocessing complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent replaces complex real-time manual captioning processes with automated speech recognition technology that processes audio stream data continuously. This substitution reduces streaming delay by eliminating manual transcription time while the automated system handles processing complexity through standardized algorithms and pre-established association rules between time information and picture frames.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11463779B2Video stream processing method and apparatus, computer device, and storage medium
Publication Date: 2022.10.04 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11463779B2 patent drawing
  • US11463779B2 patent drawing
  • US11463779B2 patent drawing

AI summary

A video stream processing method is provided. First audio stream data in live video stream data is obtained. Speech recognition is performed on the first audio stream data to generate speech recognition text. Caption data is generated according to the speech recognition text, the caption data including caption text and time information corresponding to the caption text. The caption text is added to a corresponding picture frame in the live video stream data according to the time information corresponding to the caption text to generate captioned live video stream data.