Temporal Video Fingerprint Metadata Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The SCTE 35 standard for digital television broadcasting is ambiguous, leading to confusion among content providers and distributors due to multiple valid interpretations, resulting in synchronization issues of metadata with audiovisual content, which causes incorrect triggering of events and poor viewer experience, especially during distribution via the internet.
Innovation Solution
A method and system that utilize temporal video fingerprints to synchronize metadata with audiovisual content by identifying significant visual transitions, generating timestamps, and creating metadata indices, allowing for precise insertion of metadata at correct temporal points, even after content processing and transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If SCTE 35 standard is used for metadata signaling, then commercial insertion and content replacement can be enabled, but synchronization accuracy deteriorates due to ambiguous interpretations and multiple valid configurations
Solution Approach 1:
The patent introduces temporal video fingerprints as an intermediary mechanism between the metadata and the video content. These fingerprints are generated by analyzing visual transitions in the video and creating unique temporal identifiers that can be matched with metadata timestamps, serving as a mediator to achieve precise synchronization without relying on ambiguous SCTE 35 interpretations
Solution Approach 2:
The patent changes the parameter used for synchronization from relying on SCTE 35 message timestamps (which are ambiguous) to using visual transition detection and temporal video fingerprints (which provide precise temporal reference points). This parameter change enables accurate metadata insertion regardless of the ambiguous metadata signaling standard
2Adaptability or versatility
If content is processed and transmitted through multiple distribution channels, then content reachability is improved, but metadata synchronization is lost due to processing and format conversions
Solution Approach 1:
The patent performs preliminary action by generating temporal video fingerprints from the video content before distribution. These fingerprints are created by detecting visual transitions and establishing temporal reference points in advance, so that even when content is processed and converted through multiple distribution channels, the synchronization information is already embedded and can be used to restore metadata alignment at the receiving end
Solution Approach 2:
The patent creates a copy of the temporal synchronization information in the form of temporal video fingerprints that can be independently transmitted and stored. These fingerprints are separate from the actual video data but can be matched against the video content at any point in the distribution chain to restore synchronization, enabling reliable metadata insertion across multiple distribution channels
3Measurement precision
If visual transition detection is performed on every frame, then synchronization precision is improved, but computational complexity increases
Solution Approach 1:
The patent extracts only the essential temporal synchronization information from the video content by detecting visual transitions at key points rather than analyzing every single frame in detail. By taking out only the necessary temporal reference points (visual transitions) and creating fingerprints from these extracted points, the system achieves sufficient synchronization precision without the computational burden of exhaustive frame-by-frame analysis
Data Source
AI summary
An example method comprises receiving, at a first digital device, video data, scanning video content of the video data for visual transitions within the video content between consecutive frames of the video data, each transition indicating significant visual transitions relative to other frames of the video data, timestamping each visual transition and create a first set of temporal video fingerprints, identifying items of metadata to be associated with the video data, identifying a location within the video data using the temporal video fingerprints for the identified items of metadata, generating a metadata index identifying each item of metadata and a location for each item of metadata relative to the video data using at least one of the temporal video fingerprints, and transmitting, at the first digital device, the video data, the first set of temporal video fingerprints, and the metadata index to a different digital device.


