Video Identification via Temporal Pattern Signatures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video identification methods are inefficient in detecting near-duplicate videos due to sensitivity to editing operations, high computational costs, and memory requirements, particularly in consumer electronics environments.

Innovation Solution

A method and apparatus that process image sequences by generating descriptor elements for pixel neighborhoods, forming words from these elements, and comparing their frequency and temporal order to identify matching frames, providing a compact representation and robustness to common editing operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If key frames and key-point feature extraction are used for near-duplicate detection, then matching accuracy is improved, but computational cost and storage requirements increase significantly

Engineering Contradiction:
Improvematching accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The video sequence is divided into shot segments based on detected shot boundaries. Each shot is processed independently to generate a signature, enabling efficient comparison without analyzing entire video sequences. This segmentation reduces computational complexity while maintaining matching accuracy through localized feature analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention extracts only the essential temporal pattern information from video shots to form compact signatures. By extracting and comparing only the critical temporal relationships rather than full feature sets, the method achieves accurate near-duplicate detection with significantly reduced computational cost and storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

2Quantity of substance

If shot cuts and boundaries are used to form video signatures, then storage requirements are reduced, but detection reliability deteriorates on short sequences and becomes sensitive to shot-detection algorithms

Engineering Contradiction:
Improvestorage requirementsVSAvoiddetection reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The method performs preliminary shot boundary detection and segmentation before signature generation. By pre-organizing video content into shot segments with temporal markers, the system creates a robust foundation for subsequent comparison that works reliably on both short and long sequences, reducing sensitivity to variations in shot-detection algorithms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention changes the parameter representation from simple shot boundary markers to temporal patterns that capture the sequence and duration of shots. This parameter transformation enables reliable detection on short sequences by capturing essential temporal relationships while maintaining compact storage through efficient pattern encoding.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If visual vocabulary clustering is used for feature matching, then searching speed is improved, but the method over-fits to training data and fails to generalise

Engineering Contradiction:
Improvesearching speedVSAvoidgeneralisation capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

Instead of using learned visual vocabularies that may over-fit training data, the invention uses a fixed, predefined set of temporal pattern templates that are copied and applied universally across different video datasets. This approach eliminates over-fitting while maintaining fast searching through efficient template matching, improving generalisation capability without sacrificing speed.

Inventive Principle:
Principle #26Copying

4Productivity

If hash tables are used for fast searching in near-duplicate detection, then searching speed is improved, but memory requirements become prohibitively high for consumer electronics environments

Engineering Contradiction:
Improvesearching speedVSAvoidmemory requirements
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The video database is segmented into shot-based units with compact temporal signatures. This segmentation enables efficient indexing and searching without requiring large hash tables, as the segmented structure itself provides natural organization for fast retrieval. Consumer electronics devices can store and search these compact signatures with limited memory resources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention changes the representation parameters from high-dimensional feature vectors requiring large hash tables to compact temporal pattern signatures. This parameter transformation dramatically reduces memory requirements while maintaining fast searching capability through efficient pattern matching algorithms suitable for consumer electronics environments.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8699851B2Video identification
Publication Date: 2014.04.15 RAKUTEN GROUP INC
  • US8699851B2 patent drawing
  • US8699851B2 patent drawing
  • US8699851B2 patent drawing

AI summary

A method and apparatus for processing a first sequence of images and a second sequence of images to compare the first and second sequences is disclosed. Each of a plurality of the images in the first sequence and each of a plurality of the images in the second sequence is processed by (i) processing the image data for each of a plurality of pixel neighborhoods in the image to generate at least one respective descriptor element for each of the pixel neighborhoods, each descriptor element comprising one or more bits; and (ii) forming a plurality of words from the descriptor elements of the image such that each word comprises a unique combination of descriptor element bits. The words for the second sequence are generated from the same respective combinations of descriptor element bits as the words for the first sequence. Processing is performed to compare the first and second sequences by comparing the words generated for the plurality of images in the first sequences with the words generated for the plurality of images in the second sequence.