SEI Text Metadata Packaging for Video Bitstream Comments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding schemes lack efficient methods to embed and retrieve textual data within video bitstreams, limiting the ability to add metadata and annotations that enhance video content utilization and user interaction.

Innovation Solution

Embed textual data within video bitstreams using supplemental enhancement information messages, specifically through the creation of a new SEI message type to carry textual data, allowing for the insertion and extraction of metadata such as production information, annotations, and user-generated comments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If video coding schemes use prediction and transform to achieve high compression efficiency, then compression ratio is improved, but the ability to embed and retrieve textual data within video bitstreams deteriorates

Engineering Contradiction:
Improvecompression efficiencyVSAvoidtextual data embedding capability
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent embeds textual data within the video bitstream by nesting it inside SEI message structures. The textual data is packaged within a SEI message payload, which itself is nested within the video bitstream structure. This allows compression-efficient video coding to maintain its structure while containing textual information in a nested manner, resolving the contradiction between compression efficiency and textual data embedding capability

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent introduces SEI (Supplemental Enhancement Information) messages as an intermediary structure to carry textual data. Rather than directly embedding text in the video data or sacrificing compression efficiency, the SEI message acts as a mediator that transports textual information alongside the compressed video stream, preserving both compression efficiency and textual data embedding capability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If existing video coding schemes are used without supplemental enhancement information messages, then device complexity is reduced, but the ability to add metadata and annotations deteriorates

Engineering Contradiction:
Improvecoding scheme simplicityVSAvoidmetadata embedding capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The SEI message structure serves multiple functions: it maintains compatibility with existing video coding schemes while simultaneously providing a universal container for various types of textual data including metadata, annotations, and comments. This multi-functionality allows the system to preserve coding simplicity while gaining enhanced adaptability for different data embedding needs

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the video bitstream into distinct components, separating the compressed video data from the supplemental textual information carried in SEI messages. This segmentation allows the main video coding scheme to remain simple while the textual data is handled in separate, standardized message structures, resolving the contradiction between device complexity and metadata embedding capability

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4633168A1Language and purpose information for text comments in a video bitstream using supplemental enhancement information message
Publication Date: 2025.10.15 INTERDIGITAL CE PATENT HOLDINGS SAS
  • EP4633168A1 patent drawingFigure 1~2
  • EP4633168A1 patent drawingFigure 3
  • EP4633168A1 patent drawingFigure 4

AI summary

A method and device allow to embed textual data into a bitstream that comprises encoded video by packaging the textual data as a SEI message and inserting it into the bitstream. Additional data in the SEI message specifies the language or the purpose of the textual data. The SEI message may comprise a persistence information indicating whether the textual data applies to a single picture or for subsequent pictures.