Self-Supervised Spatiotemporal Transformers for Group Activity Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing group activity recognition (GAR) methods rely heavily on ground-truth bounding boxes and substantial data labeling annotations, making them unworkable and limiting their application.

Innovation Solution

A self-supervised spatiotemporal transformer approach that generates varying temporal and spatial views of video clips, using a teacher-student framework to learn long-range dependencies without ground-truth bounding boxes or labeled data, by matching features across spatial and temporal dimensions in latent space.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If ground-truth bounding boxes and substantial data labeling annotations are used, then measurement precision of group activities is improved, but device complexity and loss of time increase

Engineering Contradiction:
Improvegroup activity recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-supervised learning by automatically generating supervision signals from the video data itself through temporal views and spatial transformations, eliminating the need for external ground-truth annotations and bounding boxes while maintaining recognition accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The method pre-processes video data by generating multiple temporal views and applying spatial augmentations before recognition, creating a robust feature representation that reduces dependency on labeled data and simplifies the overall system architecture

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If ground-truth bounding boxes and substantial data labeling annotations are used, then measurement precision of group activities is improved, but loss of time increases

Engineering Contradiction:
Improvegroup activity recognition accuracyVSAvoiddata labeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system generates its own supervision signals through self-supervised learning mechanisms, automatically creating temporal views and spatial transformations from raw video data without requiring manual annotation time

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The method performs preliminary data processing by generating multiple temporal views and applying spatial augmentations to create robust features before recognition, reducing the need for time-consuming labeled data preparation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250233958A1Spartan: self-supervised spatiotemporal transformers approach to group activity recognition
Publication Date: 2025.07.17 THE BOARD OF TRUSTEES OF THE UNIV OF ARKANSAS
  • US20250233958A1 patent drawing
  • US20250233958A1 patent drawing
  • US20250233958A1 patent drawing

AI summary

The present disclosure pertains to a computer-implemented method of predicting one or more motions of a video by (1) generating a plurality of temporal views of the video, where the temporal views of the video include a plurality of different video clips with varying motion characteristics; (2) varying spatial characteristics of the plurality of the video clips, where the varying includes generating local spatial fields and global spatial fields of the video clips; and (3) feeding the video clips, the local spatial fields, and the global spatial fields into an algorithm, where the algorithm matches varying views of the video clips across spatial and temporal dimensions in latent space to predict the one or motions of the video. Additional embodiments pertain to computer program products and systems for predicting one or more motions of a video.