Self-Supervised Data Augmentation for Skeleton-Based Action Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Skeleton-based action recognition requires significant pre-processing and data collection, which can lead to accuracy deterioration and reduced generality due to the need for body tracking equipment and diverse imaging conditions, limiting the effectiveness of recognition models.

Innovation Solution

A data augmentation apparatus and method using self-supervised learning that extracts feature vectors with object information from motion data to synthesize new motion data, enabling the generation of diverse action recognition data by deforming and combining existing data sets, thereby enhancing model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If skeleton-based action recognition is used, then data size is reduced and joint relationships are explicitly treated, but pre-processing complexity increases and data collection requirements increase

Engineering Contradiction:
Improvedata sizeVSAvoidpre-processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent uses skeleton data from existing datasets (like HMDB51) as copies or representations of action data, avoiding the need to collect and process large volumes of raw image data through complex pre-processing pipelines. The skeleton-based approach creates a simplified copy of the essential action information.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts skeleton joint coordinates and relationships from image data, separating the essential action information from the redundant visual data. This extraction process reduces data size while preserving the critical joint relationship information needed for action recognition.

Inventive Principle:
Principle #2Taking out (Extraction)

2Quantity of substance

If body tracking equipment and imaging places are required, then skeleton data can be collected, but the system complexity and resource requirements increase

Engineering Contradiction:
Improveskeleton data availabilityVSAvoidequipment and facility requirements
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent uses pre-existing skeleton datasets (such as HMDB51, which contains 5,000 skeleton sequences) as ready-to-use copies of action data. This eliminates the need to deploy body tracking equipment and imaging facilities, as the skeleton data is already extracted and prepared from public datasets.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary extraction and preparation of skeleton data from public datasets before the actual action recognition task. By pre-processing and storing skeleton sequences in advance, the system avoids the need for complex real-time body tracking equipment and imaging infrastructure during the recognition phase.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If skeleton data are not sufficiently provided, then recognition model accuracy deteriorates, but collecting more data requires more resources

Engineering Contradiction:
Improverecognition accuracyVSAvoiddata collection resources
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses a universal skeleton-based representation that can capture action information across different datasets and applications. The skeleton joint coordinates and relationships serve as a universal language that can be applied to various action recognition tasks without requiring dataset-specific processing, thereby improving accuracy without proportionally increasing collection resources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms and augments skeleton data by modifying parameters such as joint coordinates, skeleton configurations, and action labels. This parameter transformation allows the system to generate diverse training data from limited source data, improving recognition accuracy while avoiding the need to collect proportionally more raw data.

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If data augmentation is performed on imaging subjects and classes, then the number of images increases, but the complexity of processing and managing data increases

Engineering Contradiction:
Improvenumber of imagesVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent creates augmented data by copying and transforming existing skeleton sequences rather than processing large volumes of raw images. The augmentation process works on the simplified skeleton representation, generating varied skeleton sequences that can be directly used for training without requiring complex image processing pipelines.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent segments the action recognition task into separate components: skeleton extraction, action classification, and data augmentation. By working with segmented skeleton data rather than complete image sequences, the system reduces processing complexity while still achieving the desired increase in training data quantity and diversity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240212167A1Data augmentation apparatus and method for action recognition through self-supervised learning based on object
Publication Date: 2024.06.27 GWANGJU INST OF SCI & TECH
  • US20240212167A1 patent drawing
  • US20240212167A1 patent drawing
  • US20240212167A1 patent drawing

AI summary

The present invention relates to a data augmentation apparatus for action recognition through self-supervised learning based on objects, including: an image input unit for inputting image information; an information extraction unit for extracting a feature vector with object information from motion data of the inputted image information; and a motion information synthesis unit for synthesizing motion data taking a different action and the feature vector with the object information to generate new motion data.