Self-Supervised Video Event Detection Using Behavioral Regime Changes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional video analytics techniques require large datasets of labeled examples for event detection, which is challenging for rare events, and fail to adapt to new types of events.

Innovation Solution

A self-supervised learning approach that represents spatial characteristics of objects in video data as timeseries, associates behavioral regimes, generates ground truth labels based on regime changes, and trains a model to detect events using these labels, allowing for automated training and updating.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional supervised learning with large labeled datasets is used for event detection, then detection accuracy for common events is improved, but the system fails to detect rare events and cannot adapt to new event types

Engineering Contradiction:
Improveevent detection accuracyVSAvoidadaptability to new event types
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs self-supervised learning by automatically generating labels from unlabelled video data through behavioral regime change detection. The model learns to detect events without human-annotated training data, enabling it to adapt to new event types autonomously while maintaining detection accuracy through unsupervised pattern recognition

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-processes video data by extracting spatial characteristics and detecting behavioral regime changes before formal model training. This preliminary analysis creates pseudo-labels that prepare the data for subsequent self-supervised learning, enabling the system to adapt to new events without requiring pre-collected labeled examples

Inventive Principle:
Principle #10Preliminary action

2Reliability

If large labeled datasets are collected and processed centrally, then comprehensive event detection is achieved, but computational complexity and data transmission requirements increase

Engineering Contradiction:
Improveevent detection capabilityVSAvoidcomputational resource requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments video analysis into distributed edge devices that independently process local video feeds. Each edge device runs its own self-supervised learning model, dividing the computational burden from centralized processing and enabling reliable event detection without requiring massive centralized computational resources or data transmission

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Edge devices perform self-supervised learning autonomously using their local video data, eliminating the need to transmit large datasets to centralized servers. This self-service capability reduces computational complexity at any single location while maintaining reliable event detection through distributed intelligence

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12608938B2Self-supervised learning for video analytics
Publication Date: 2026.04.21 CISCO TECHNOLOGY INC
  • US12608938B2 patent drawing
  • US12608938B2 patent drawing
  • US12608938B2 patent drawing

AI summary

In one embodiment, a device represents spatial characteristics of an object depicted in video data over time as one or more timeseries. The device associates different portions of the one or more timeseries with behavioral regimes of the object. The device generates ground truth labels for frames of the video data based on changes in the behavioral regimes of the object associated with those frames. The device trains a self-supervised model to detect an event depicted in the video data using the ground truth labels and their associated frames.