Perceptual Hashing for Video Frame Reference Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video encoding techniques fail to efficiently encode and decode video frames when the temporal correlation is lost between frames acquired from different points of view, such as during camera switches in live sporting events, due to the time-consuming and costly process of comparing pixel blocks between frames.

Innovation Solution

The use of metadata, including statistical values and perceptual hashes, to identify and match frames from the same point of view, allowing for efficient encoding and decoding by determining whether a current frame matches a previously encoded frame based on similarity in perceptual hashes, and using a matching frame as a reference frame when a match is found.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional pixel block comparison methods are used to identify frames from the same point of view, then encoding accuracy is maintained, but computational complexity and processing time increase significantly

Engineering Contradiction:
Improveframe matching accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts essential visual features from video frames to create perceptual hashes, retaining only the most discriminative characteristics needed for frame identification. This extraction process removes redundant pixel data while preserving the ability to distinguish between frames from different points of view, thereby reducing computational complexity while maintaining matching accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates simplified copies of video frames in the form of perceptual hashes - compact representations that capture the essential visual content. These hash copies enable rapid comparison operations without requiring access to the full-resolution original frames, significantly reducing computational burden while preserving frame identification capability

Inventive Principle:
Principle #26Copying

2Measurement precision

If brute force pixel comparison is used to determine frame similarity, then accurate matching is achieved, but processing speed decreases

Engineering Contradiction:
Improveframe similarity detection accuracyVSAvoidframe comparison speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent replaces the mechanical pixel-by-pixel comparison process with a perceptual hashing system that uses mathematical transformations to generate compact frame representations. This substitution enables rapid hash-based comparison operations that are computationally much less intensive than traditional brute force pixel comparison, dramatically improving processing speed while maintaining accurate frame similarity detection

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If temporal correlation is assumed between successive frames, then encoding efficiency is improved, but reliability decreases when camera switches occur

Engineering Contradiction:
Improveencoding efficiencyVSAvoidframe matching reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary scene detection and perceptual hash generation for each frame before the actual encoding process. This preliminary action identifies camera switches and establishes reliable frame correspondences in advance, allowing the encoder to adapt its strategy based on detected scene changes. By preparing this information beforehand, the system maintains high encoding efficiency while ensuring reliability even when temporal correlation is broken by camera switches

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11956441B2Identifying long term reference frame using scene detection and perceptual hashing
Publication Date: 2024.04.09 ATI TECHNOLOGIES ULC
  • US11956441B2 patent drawing
  • US11956441B2 patent drawing
  • US11956441B2 patent drawing

AI summary

Methods and devices are provided for encoding a video stream which comprise encoding a plurality of frames of video acquired from different points of view, generating statistical values for the frames of video determined from values of pixels of the frames, generating, for each of the plurality of frames, a perceptual hash value based on statistical values of the frame and encoding a current frame comprising video acquired from a corresponding one of the different points of view using a previously encoded reference frame based on a similarity of perceptual hashes of the current frame and the previously encoded reference frame.