Real-time Caption Superposition in Video Communication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video communication systems lack real-time caption display functionality, requiring manual input and complex RF modulation, leading to poor real-time performance and disordered speech recognition in multipoint conferences.

Innovation Solution

Integrating speech recognition modules within video terminals and MCUs to directly convert speech signals to text and superpose them on video signals for real-time encoding and transmission, enabling seamless caption display during video conferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual input and RF modulation are used for caption display, then caption display functionality is achieved, but device complexity and time delay increase

Engineering Contradiction:
Improvecaption display functionalityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent extracts and removes the complex RF modulation components from the caption display system. Instead of using RF modulators to convert text signals to video baseband signals, the system directly processes and transmits caption data through simplified communication channels, eliminating unnecessary complexity while maintaining caption display functionality

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/electronic RF modulation process with a direct digital signal processing approach. Speech signals are converted to text through speech recognition, and the text is directly encoded and transmitted without requiring RF modulation stages, substituting a complex physical modulation system with a simpler digital processing system

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If manual input and pre-editing are used for caption display, then caption display functionality is achieved, but real-time performance deteriorates

Engineering Contradiction:
Improvecaption display functionalityVSAvoidreal-time performance
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent implements preliminary speech recognition processing during the speech signal reception phase. The speech recognition module continuously processes incoming speech signals and converts them to text in advance, so that when caption display is needed, the text is already prepared and can be immediately displayed without manual input or pre-editing delays

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service caption generation through automatic speech recognition. The speech recognition module automatically converts speech signals to text without requiring manual input, and the system automatically manages the caption display process, eliminating the need for human operators to edit or prepare caption content manually

Inventive Principle:
Principle #25Self-service

3Device complexity

If a single speech recognition module is used in multipoint conferences, then device complexity is reduced, but speech recognition accuracy and reliability deteriorate

Engineering Contradiction:
Improvemodule quantityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the speech recognition function into multiple independent speech recognition modules, each responsible for recognizing speech from specific participants in the multipoint conference. This segmentation allows each module to specialize in recognizing particular speech patterns and speakers, improving overall recognition accuracy while maintaining manageable system complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates speech recognition modules with multi-functionality that can handle different speech characteristics, accents, and languages. Each module is designed to universally process various types of speech inputs from different participants, allowing the system to reliably recognize multiple speakers simultaneously while maintaining flexibility and adaptability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP2154885B1A caption display method and a video communication control device
Publication Date: 2011.11.30 HUAWEI TECH CO LTD
  • EP2154885B1 patent drawingFigure 1
  • EP2154885B1 patent drawingFigure 2
  • EP2154885B1 patent drawingFigure 3

AI summary

A caption display method in a video communication includes the following steps. A video communication is established. Speech signals of a speaker are recognized and converted to text signals. The text signals and picture video signals that need to be received by and displayed to other conference participators are superposed and encoded, and are sent through the video communication. A video communication system and device are also described. Users directly decode display pictures and character information corresponding to a speech. The method is simple, and a real-time performance is high.