Hierarchical Codebook Model for Video Event Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for human activity analysis in videos face challenges in real-world environments due to camera motion, cluttered backgrounds, occlusion, and variations in scale, viewpoint, and lighting, which affect the recognition and localization of human actions, and require pre-processing steps like background subtraction and tracking, limiting their effectiveness in unconstrained settings.

Innovation Solution

A multi-level and multi-scale hierarchical bag of video words structure is introduced, which generates a hierarchical codebook model of local spatio-temporal video volumes, allowing for action recognition and localization without prior knowledge of actions, background subtraction, or tracking, and is robust against spatial and temporal scale changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pre-processing steps like background subtraction and tracking are used, then action recognition accuracy is improved in controlled environments, but system complexity and computational cost increase

Engineering Contradiction:
Improveaction recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and eliminates the need for complex pre-processing steps (background subtraction, tracking) by directly using spatio-temporal video volumes. The invention takes out the problematic intermediate processing stages and works directly with the raw video data in a hierarchical codebook framework, thereby reducing system complexity while maintaining action recognition capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical pre-processing pipeline (background subtraction → tracking → action recognition) with a direct spatio-temporal volumetric analysis approach. By substituting the traditional multi-stage mechanical processing system with a unified hierarchical codebook model, the system achieves simpler architecture while handling real-world variations effectively.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If traditional action recognition methods are used, then performance is good in highly controlled environments, but reliability deteriorates in real-world environments with camera motion, cluttered background, occlusion, and scale variations

Engineering Contradiction:
Improveaction recognition reliabilityVSAvoidenvironmental adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transitions from analyzing 2D video frames sequentially to utilizing 3D spatio-temporal video volumes that incorporate temporal dimension explicitly. By stacking multiple frames to form volumetric representations, the system captures motion patterns and temporal relationships, enabling reliable action recognition in real-world environments with various challenges like occlusion and scale variations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent employs dynamic spatio-temporal volumetric analysis that adapts to changing environmental conditions. The hierarchical codebook model learns from temporal patterns in the video data, allowing the system to dynamically adjust to camera motion, lighting changes, and other environmental variations, thereby maintaining high reliability across diverse real-world scenarios.

Inventive Principle:
Principle #15Dynamics

3Ease of manufacture

If spatio-temporal volumetric representations are used, then pre-processing steps are eliminated, but the system becomes unable to handle scale variations due to local storage approach

Engineering Contradiction:
Improveease of implementationVSAvoidscale variation handling
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent segments the video data into hierarchical levels of spatio-temporal volumes. By organizing volumes at multiple scales and levels of abstraction (from local pixel-level volumes to higher-level action patterns), the system simultaneously achieves ease of implementation through structured processing while maintaining adaptability to handle scale variations through the multi-scale hierarchical architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested hierarchical structure where smaller spatio-temporal volumes are contained within larger contextual volumes. This nested organization allows the system to capture both fine-grained local patterns and coarse-grained global context, enabling effective handling of scale variations while maintaining a systematic and implementable framework.

Inventive Principle:
Principle #7Nested doll (Nesting)

4Loss of information

If dense spatio-temporal sampling is performed, then comprehensive video content coverage is achieved, but computational cost and data storage requirements increase

Engineering Contradiction:
Improvevisual information coverageVSAvoidcomputational cost
Core Design Contradiction:
Loss of informationVSPower

Solution Approach 1:

The patent performs preliminary action by pre-computing and storing hierarchical codebooks that capture the essential patterns of spatio-temporal volumes. By preparing these codebooks in advance through clustering and compression, the system reduces the computational burden during actual video analysis, as subsequent processing involves matching against the pre-computed codebooks rather than processing raw volumetric data from scratch.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the high-dimensional spatio-temporal volume data into a compressed parameter space using hierarchical codebooks. By changing the representation parameters from raw pixel values to clustered codebook indices, the system maintains comprehensive visual information coverage while significantly reducing computational cost and data storage requirements through dimensionality reduction and pattern compression.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10402655B2System and method for visual event description and event analysis
Publication Date: 2019.09.03 SPORTLOGIQ
  • US10402655B2 patent drawing
  • US10402655B2 patent drawing
  • US10402655B2 patent drawing

AI summary

A system and method are provided for analyzing a video. The method comprises: sampling the video to generate a plurality of spatio-temporal video volumes; clustering similar ones of the plurality of spatio-temporal video volumes to generate a low-level codebook of video volumes; analyzing the low-level codebook of video volumes to generate a plurality of ensembles of volumes surrounding pixels in the video; and clustering the plurality of ensembles of volumes by determining similarities between the ensembles of volumes, to generate at least one high-level codebook. Multiple high-level codebooks can be generated by repeating steps of the method. The method can further include performing visual event retrieval by using the at least one high-level codebook to make an inference from the video, for example comparing the video to a dataset and retrieving at least one similar video, activity and event labeling, and performing abnormal and normal event detection.