Caption Synchronization via Audio-Text Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Display apparatuses face issues with captions not being synchronized with images or sounds, leading to user confusion, especially in real-time broadcasting where captions may be delayed.

Innovation Solution

A display apparatus and method that includes a signal receiver, data extractors, a buffering section, a sound-text converter, and a synchronizer to extract and synchronize caption data with video frames based on sound timing information, ensuring accurate synchronization even if captions are initially delayed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If caption data is extracted directly from the received signal without synchronization processing, then the processing speed is fast, but the caption is delayed and not synchronized with the image or sound

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidcaption delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by buffering video frames before synchronization processing. The buffering section stores a plurality of video frames in advance, allowing the synchronizer to align caption data with the correct video frame based on timing information. This pre-buffering approach enables accurate synchronization without delaying the caption display, as the necessary video data is already available in the buffer when caption synchronization is performed.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If sound-text conversion is performed to verify caption synchronization, then the synchronization accuracy is improved, but the processing complexity increases

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses sound-text conversion as an intermediary verification mechanism. The synchronizer converts sound data to text and compares it with caption data to determine synchronization accuracy. This intermediary process provides an objective method for verifying whether caption data is properly synchronized with the video and audio content, enabling automated adjustment of caption timing without manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If multiple data extraction and synchronization processes are implemented, then the caption synchronization accuracy is improved, but the processing time increases

Engineering Contradiction:
Improvecaption synchronization accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent maintains continuous useful action by performing data extraction, buffering, and synchronization processes in a continuous pipeline. The buffering section continuously stores incoming video frames, the sound-text converter continuously processes sound data, and the synchronizer continuously aligns caption data with the buffered video frames based on timing information. This continuous operation ensures that synchronization is achieved without interrupting the overall processing flow, maintaining high productivity while achieving accurate caption synchronization.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS9319566B2Display apparatus for synchronizing caption data and control method thereof
Publication Date: 2016.04.19 SAMSUNG ELECTRONICS CO LTD
  • US9319566B2 patent drawing
  • US9319566B2 patent drawing
  • US9319566B2 patent drawing

AI summary

A display apparatus and a method of controlling the display apparatus are disclosed. The display apparatus includes: a signal receiver configured to receive a signal containing video data of a series of frames and corresponding sound data; a first data extractor configured to extract caption data from the signal; a second data extractor configured to extract the video data and the sound data from the signal; a buffering section configured to buffer the extracted video data; a sound-text converter configured to convert the extracted sound data into a text through sound recognition; a synchronizer configured to compare the converted text with the extracted caption data, and synchronize the caption data with frames corresponding to respective caption data among frames of the buffered video data; and a display configured to display the frames synchronized with the caption data.