Closed Caption Transport in Digital Video Streams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The transition to digital television broadcasts has increased complexity in handling closed captions, particularly in situations where MPEG-2 video is not present throughout every frame or when using video codecs other than MPEG-2, requiring a method to carry captioning data effectively in these scenarios.
Innovation Solution
A method of embedding closed caption data into a standard video syntax, allowing for temporal alignment with video frames, and encoding it as a background within the video stream, which can be applied to current and next-generation video streams, maintaining compatibility with existing decoder logic and minimizing overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If closed caption data is embedded into standard video syntax and encoded as background, then caption transport efficiency is improved and decoder compatibility is maintained, but the system cannot handle cases where video is not present or uses non-MPEG-2 codecs
Solution Approach 1:
The patent creates a universal caption transport mechanism that functions across multiple video scenarios: when video is present (MPEG-2 or other codecs) and when video is absent (audio-only modes). The caption data structure and transport method are designed to be format-agnostic, allowing the same infrastructure to serve both video-based and audio-only broadcast modes without requiring separate systems.
Solution Approach 2:
The patent segments the caption transport function from the video encoding function. By embedding caption data in a standardized syntax that can be independently processed, the system allows video and caption streams to be handled separately yet synchronized, enabling caption transport to work regardless of whether video is present or what codec is used.
2Adaptability or versatility
If a new caption transport method is developed for non-MPEG-2 codecs and audio-only modes, then adaptability to various formats is improved, but device complexity and implementation difficulty increase
Solution Approach 1:
The patent introduces an intermediary caption data structure that acts as a mediator between the video stream (or absence thereof) and the caption display system. This intermediary structure standardizes caption transport so that regardless of whether video is present or what codec is used, the caption data follows a predictable format that existing decoder logic can handle with minimal modification.
Solution Approach 2:
The patent changes the parameters of the caption transport system to be independent of video-specific parameters. By decoupling caption data timing and synchronization from video frame structures, the system allows caption transport to work in audio-only modes and with various video codecs without requiring complex video-aware logic in the decoder.
3Measurement precision
If closed caption data is carried in every video frame, then caption synchronization with video is improved, but overhead increases when video is not present or in audio-only modes
Solution Approach 1:
The patent implements periodic caption synchronization markers that are inserted at regular intervals in the caption stream, independent of video frame presence. These periodic markers provide sufficient synchronization information without requiring caption data in every single video frame, reducing overhead while maintaining adequate sync precision for caption display.
Solution Approach 2:
The patent uses partial action by providing caption synchronization information only when necessary (e.g., at interval markers or keyframes) rather than in every single frame. This partial approach maintains sufficient synchronization precision while significantly reducing the quantity of caption data that must be transmitted in modes where video is absent or present only intermittently.
Data Source
AI summary
A method and system for digital closed caption transport are provided. In one example, the method involves receiving closed caption data and a program feed having video content, embedding the closed caption data into a standard video syntax, and encoding the video content into the standard video syntax as a background, wherein the closed caption data and the video content are encoded into a closed caption program feed.


