Compact Video Representation for Event Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video analysis solutions face challenges with variable-length feature lists, inaccuracy, especially in slow and smooth scene changes, and high dimensionality, leading to storage and processing issues, as well as inconsistency in representing video content.

Innovation Solution

The techniques generate a counting grid representation of videos, aggregate features from active locations, and apply normalization and dimension reduction to provide compact, spatially consistent feature representations, using a pre-trained counting grid model to ensure consistent spatial mapping of similar frames.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If appearance-based frame-level methods are used to extract video features, then video content can be represented, but the feature lists become variable-length and high-dimensional (thousands to tens of thousands of features), causing storage and processing issues

Engineering Contradiction:
Improvevideo content representation accuracyVSAvoidfeature list dimensionality
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the video into distinct events rather than processing at frame-level granularity. By identifying and separating meaningful event boundaries, the system converts the continuous stream of frames into discrete event units, dramatically reducing feature dimensionality while preserving important semantic information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a temporal dimension to the feature representation by organizing features into event-level temporal segments. This transforms the flat, high-dimensional frame-level feature space into a structured, multi-level representation where features are grouped by events, reducing overall dimensionality while maintaining discriminative power.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If frame-level features are extracted for key frames, then video features can be obtained, but the solutions tend to be inaccurate especially when scenes change slowly and smoothly

Engineering Contradiction:
Improvescene change detection accuracyVSAvoidscene change identification reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent implements dynamic event boundary detection that adapts to different scene change characteristics. Rather than using fixed thresholds or rigid frame-level analysis, the system dynamically identifies event boundaries based on content analysis, allowing it to accurately detect both rapid and gradual scene transitions while maintaining reliability across diverse video types.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If generic feature coding methods like fisher vector coding are used, then video features can be encoded, but storage, processing, and consistency issues arise due to high dimensionality (tens of thousands of features)

Engineering Contradiction:
Improvefeature coding generalityVSAvoidfeature dimensionality
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential and discriminative features from each event, rather than encoding all available frame-level features. By selectively extracting features that are most relevant for event representation and discarding redundant information, the system achieves compact feature codes with significantly reduced dimensionality while maintaining coding versatility.

Inventive Principle:
Principle #2Taking out (Extraction)

4Ease of manufacture

If hand-crafted pooling of frame-level features is used, then video-level features can be obtained under certain conditions, but the solutions do not generalize well

Engineering Contradiction:
Improvevideo-level feature generationVSAvoidsolution generalization
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary event segmentation and boundary detection before feature aggregation. By pre-identifying event structures and boundaries in a content-aware manner, the system creates a robust foundation for feature pooling that generalizes well across different video types and conditions, rather than applying fixed hand-crafted pooling operations directly to frames.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10986400B2Compact video representation for video event retrieval and recognition
Publication Date: 2021.04.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10986400B2 patent drawing
  • US10986400B2 patent drawing
  • US10986400B2 patent drawing

AI summary

Comprehensive, compact, and discriminative representations of videos can be obtained using a counting grid representation of the video and aggregating features associated with active locations of the counting grid to obtain a feature representation of the video. The feature representation can be used for video retrieval and/or recognition. In some examples, the techniques may include conducting normalization and dimension reduction on the aggregated features to obtain a further compact and discriminative feature representation. In some examples, the counting grid representation of the video is generated using a pre-trained counting grid model in order to provide spatially consistent feature representations of the videos.