Self-Supervised Data Augmentation for Skeleton-Based Action Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Skeleton-based action recognition requires significant pre-processing and data collection, which can lead to accuracy deterioration and reduced generality due to the need for body tracking equipment and diverse imaging conditions, limiting the effectiveness of recognition models.
Innovation Solution
A data augmentation apparatus and method using self-supervised learning that extracts feature vectors with object information from motion data to synthesize new motion data, enabling the generation of diverse action recognition data by deforming and combining existing data sets, thereby enhancing model performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If skeleton-based action recognition is used, then data size is reduced and joint relationships are explicitly treated, but pre-processing complexity increases and data collection requirements increase
Solution Approach 1:
The patent uses skeleton data from existing datasets (like HMDB51) as copies or representations of action data, avoiding the need to collect and process large volumes of raw image data through complex pre-processing pipelines. The skeleton-based approach creates a simplified copy of the essential action information.
Solution Approach 2:
The patent extracts skeleton joint coordinates and relationships from image data, separating the essential action information from the redundant visual data. This extraction process reduces data size while preserving the critical joint relationship information needed for action recognition.
2Quantity of substance
If body tracking equipment and imaging places are required, then skeleton data can be collected, but the system complexity and resource requirements increase
Solution Approach 1:
The patent uses pre-existing skeleton datasets (such as HMDB51, which contains 5,000 skeleton sequences) as ready-to-use copies of action data. This eliminates the need to deploy body tracking equipment and imaging facilities, as the skeleton data is already extracted and prepared from public datasets.
Solution Approach 2:
The patent performs preliminary extraction and preparation of skeleton data from public datasets before the actual action recognition task. By pre-processing and storing skeleton sequences in advance, the system avoids the need for complex real-time body tracking equipment and imaging infrastructure during the recognition phase.
3Measurement precision
If skeleton data are not sufficiently provided, then recognition model accuracy deteriorates, but collecting more data requires more resources
Solution Approach 1:
The patent uses a universal skeleton-based representation that can capture action information across different datasets and applications. The skeleton joint coordinates and relationships serve as a universal language that can be applied to various action recognition tasks without requiring dataset-specific processing, thereby improving accuracy without proportionally increasing collection resources.
Solution Approach 2:
The patent transforms and augments skeleton data by modifying parameters such as joint coordinates, skeleton configurations, and action labels. This parameter transformation allows the system to generate diverse training data from limited source data, improving recognition accuracy while avoiding the need to collect proportionally more raw data.
4Quantity of substance
If data augmentation is performed on imaging subjects and classes, then the number of images increases, but the complexity of processing and managing data increases
Solution Approach 1:
The patent creates augmented data by copying and transforming existing skeleton sequences rather than processing large volumes of raw images. The augmentation process works on the simplified skeleton representation, generating varied skeleton sequences that can be directly used for training without requiring complex image processing pipelines.
Solution Approach 2:
The patent segments the action recognition task into separate components: skeleton extraction, action classification, and data augmentation. By working with segmented skeleton data rather than complete image sequences, the system reduces processing complexity while still achieving the desired increase in training data quantity and diversity.
Data Source
AI summary
The present invention relates to a data augmentation apparatus for action recognition through self-supervised learning based on objects, including: an image input unit for inputting image information; an information extraction unit for extracting a feature vector with object information from motion data of the inputted image information; and a motion information synthesis unit for synthesizing motion data taking a different action and the feature vector with the object information to generate new motion data.


