Multiresolution Video Fingerprinting for Format Invariance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for uniquely identifying digital video objects, such as hash functions, are inadequate for content identification as they change with format and bitrate variations, lacking robustness and discriminability.
Innovation Solution
A method and system for generating a unique fingerprint for digital video objects by processing spatial and temporal signatures at multiple resolutions and frame rates, creating a robust and compact identifier that remains invariant to format, bitrate, and minor alterations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a hash function is used to identify digital video objects, then the identification is unique for each file, but the identifier changes when format or bitrate is changed
Solution Approach 1:
The video object is divided into multiple frames, and each frame is further divided into blocks. Spatial signatures are extracted from spatial relationships within frames, and temporal signatures are extracted from temporal relationships between frames. This segmentation allows the fingerprint to capture content-essential features that remain invariant across format changes.
Solution Approach 2:
The patent transitions from file-level identification (single hash value) to content-level identification by adding temporal dimension (multiple frames) and spatial dimension (block relationships). The fingerprint operates in a multi-dimensional space combining spatial signatures from multiple resolutions and temporal signatures from multiple frame rates, making it robust to format variations.
2Reliability
If traditional fingerprinting methods are used, then the identifier is compact, but it lacks robustness against format changes and minor alterations
Solution Approach 1:
Different parts of the video are processed with different signature extraction methods. Spatial signatures capture local spatial relationships within frames, while temporal signatures capture local temporal relationships between frames. This local quality approach ensures that essential content features are preserved while achieving robustness against format changes.
Solution Approach 2:
The fingerprint is constructed as a composite of multiple spatial signatures (at different resolutions) and multiple temporal signatures (at different frame rates). This composite structure combines the strengths of different signature types to achieve both robustness against format changes and preservation of content details.
3Measurement precision
If multiple signatures at different resolutions and frame rates are extracted, then the fingerprint becomes more robust and discriminating, but the processing complexity increases
Solution Approach 1:
The system dynamically adapts the number of spatial and temporal signatures based on the specific video object and application requirements. The method allows flexible configuration of signature quantities at different resolutions and frame rates, enabling optimization between precision and processing complexity for different scenarios.
Solution Approach 2:
The patent extracts multiple spatial signatures and multiple temporal signatures, which may seem excessive, but this partial redundancy provides robustness against format changes. The system can use all signatures for maximum precision or select a subset for lower complexity applications, achieving a balance between discriminability and processing efficiency.
Data Source
AI summary
A method and system for generating a spatial signature for a frame of a video object. The method includes obtaining a frame associated with a video object, and dividing the frame into a plurality of blocks. The plurality of blocks corresponds to a plurality of locations respectively, each of the plurality of blocks includes a plurality of pixels, and the plurality of pixels corresponds to a plurality of pixel values respectively. Additionally, the method includes determining a plurality of average pixel values for the plurality of blocks respectively. Each of the plurality of blocks corresponds to one of the plurality of average pixel values. Moreover, the method includes processing information associated with the plurality of average pixel values and determining a plurality of comparison values for the plurality of blocks respectively based on at least information associated with the plurality of average pixel values.


