Closed Captioning Text Insertion via Terminal Tag Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing closed captioning technologies for video programming face accuracy issues due to signal-to-noise ratio, speech clarity, and pronunciation quality, which can be improved by bypassing indirect text generation methods.

Innovation Solution

A system that directly inputs known text into the caption content using explicit or implicit signals, allowing for precise control over the insertion of text into the caption stream, thereby enhancing accuracy by avoiding unnecessary speech-to-text conversions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If speech-to-text conversion is used to generate closed captions, then text can be automatically generated from audio, but accuracy deteriorates due to signal-to-noise ratio, speech clarity, and pronunciation quality issues

Engineering Contradiction:
Improveautomatic text generationVSAvoidcaption accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-synchronizing text segments with audio segments before the captioning process. Text segments are prepared and time-aligned with their corresponding audio portions in advance, using time codes to establish precise temporal relationships. This pre-synchronization eliminates the need for real-time speech-to-text conversion during playback, thereby maintaining high accuracy while achieving automatic caption generation.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If text segments are time-aligned with audio segments using time codes, then synchronization precision is improved, but system complexity increases due to the need for terminal tags and signal correspondence

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidsystem structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses copying by creating text segments that are direct representations or transcriptions of the audio content, rather than generating new text through complex speech-to-text algorithms. These text segments are copied or transcribed in advance and then time-aligned with the audio using simple time codes. This approach achieves high synchronization accuracy without requiring complex real-time processing systems, as the text is already prepared and matched to the audio timeline.

Inventive Principle:
Principle #26Copying

3Measurement precision

If known text is directly inserted into caption content, then accuracy is improved by avoiding speech-to-text errors, but control over caption generation becomes more complex

Engineering Contradiction:
Improvecaption accuracyVSAvoidcaption generation control
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent applies segmentation by dividing the caption content into discrete text segments, each corresponding to a specific audio segment. Each text segment is independently time-aligned and can be individually controlled or inserted. This segmentation allows for precise control over which text appears when, maintaining ease of operation through modular management while achieving high accuracy by using pre-prepared text segments rather than relying on speech-to-text conversion.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12256120B2Closed caption content generation
Publication Date: 2025.03.18 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12256120B2 patent drawing
  • US12256120B2 patent drawing
  • US12256120B2 patent drawing

AI summary

A system may include a memory and a processor in communication therewith configured to perform operations. The operations may include receiving an audio file and a text file related to the audio file, analyzing the audio file to produce an analysis, and determining a portion of the audio file is similar to a segment of the text file. The operations may include identifying a first terminal signal and corresponding the first terminal signal to a first terminal tag in the text file such that the first terminal tag is aligned with the first terminal signal; the first terminal signal identifies a first portion terminal end of the portion and the first terminal tag identifies a first segment terminal end of the segment. The operations may include generating a converted text from the analysis and inserting the segment into the converted text.