Adaptable Video Captioning via Bitstream Repackaging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional captioning systems in video broadcasts often limit users to a single caption channel, making it difficult or impossible to select beyond the first channel, especially in devices like iPhones and cable boxes, which restricts accessibility and compliance with evolving disability and multilingual requirements.

Innovation Solution

An encoder and re-packager circuit system that generates bitstreams with a video portion, a subtitle placeholder channel, and multiple caption channels, allowing for the selection and swapping of caption channels during playback, reducing server overhead and enhancing usability across devices with limited multilingual capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional captioning systems use a single caption channel in the video stream, then device complexity is reduced and ease of operation is improved, but adaptability and accessibility for multilingual and disabled users deteriorate

Engineering Contradiction:
Improvecaption channel selectionVSAvoidcaption system structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments multiple caption channels (including different languages and accessibility formats) into separate data streams that are multiplexed within the video bitstream. Each caption channel is independently encoded and can be selectively extracted at the receiving end, allowing devices to access only the caption channel they need without processing all channels, thus maintaining device simplicity while providing extensive caption selection capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The video encoder is designed to handle multiple caption channels universally, supporting various caption formats (e.g., CEA-608, CEA-708, SCC) and languages within a single encoding framework. This multi-functional encoder can adapt to different captioning requirements without requiring separate encoding systems, thereby increasing adaptability while managing complexity through standardized processing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple caption channels are embedded in the video stream, then accessibility and multilingual support are improved, but bandwidth consumption and server overhead increase

Engineering Contradiction:
Improvemultilingual caption supportVSAvoiddata volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

Multiple caption channels are merged into a single video bitstream using efficient multiplexing techniques. Different caption channels (English, Spanish, French, accessibility captions, etc.) are combined in the same data stream with minimal overhead, allowing simultaneous transmission of multiple languages and formats without requiring separate video streams for each caption type, thus reducing overall bandwidth consumption.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

Caption data is nested within the video bitstream structure, with caption channels embedded as auxiliary data within the video encoding framework. This nesting allows caption information to be carried along with the video data without requiring separate transmission channels, efficiently utilizing the existing data structure to reduce overall data volume while maintaining multiple caption channel capability.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Adaptability or versatility

If caption channels are made selectable beyond the first channel, then user accessibility and compliance with disability laws are improved, but ease of operation deteriorates due to complex channel selection interfaces

Engineering Contradiction:
Improvecaption channel accessibilityVSAvoidcaption selection interface
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system performs preliminary organization of caption channels by assigning specific identifiers and metadata to each caption channel during encoding. Caption channels are pre-labeled with language codes, accessibility type indicators, and other metadata that enable automatic recognition and selection. This preliminary structuring allows receiving devices to automatically match and select appropriate caption channels based on user preferences or device capabilities without requiring complex manual navigation through multiple channels.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10666896B2Adaptable captioning in a video broadcast
Publication Date: 2020.05.26 AMAZON TECH INC
  • US10666896B2 patent drawing
  • US10666896B2 patent drawing
  • US10666896B2 patent drawing

AI summary

An encoder and a re-packager circuit. The encoder may be configured to generate one or more bitstreams each having (i) a video portion, (ii) a subtitle placeholder channel, and (iii) a plurality of caption channels. The re-packager circuit may be configured to generate one or more re-packaged bitstreams in response to (i) one of the bitstreams and (ii) a selected one of the plurality of caption channels. The re-packaged bitstream moves the selected caption channel into the subtitle placeholder channel.