Behavior Data Augmentation Using Spatiotemporal Characteristics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video data augmentation methods fail to consider correlations between classes and are not suitable for object-specific behavior recognition, as they do not account for spatiotemporal relationships in video data, leading to inefficient data augmentation and limited recognition capabilities.
Innovation Solution
A behavior data augmenting apparatus and method that extracts object regions from video data, defines spatiotemporal characteristics for each class, and augments data by considering temporal and spatial directionality, counterparts, and negative classes, using a learning algorithm to recognize object behaviors effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional data augmentation methods are used, then data quantity increases, but spatiotemporal relationships and class correlations are lost
Solution Approach 1:
The patent applies parameter changes by systematically transforming video data through temporal operations (forward/backward playback, speed variation, pausing) and spatial operations (flipping, rotating, cropping, scaling). These parameter transformations generate diverse augmented data while preserving the underlying spatiotemporal relationships and class correlations inherent in the original behavior data.
2Reliability
If video data is used for behavior recognition, then recognition capability improves, but data dimensionality increases making augmentation difficult
Solution Approach 1:
The patent segments video data into discrete actionable units with defined temporal boundaries and spatial characteristics. By dividing continuous video streams into segmented behavior clips with metadata annotations, the system reduces dimensional complexity while preserving essential recognition features, making the data more manageable for augmentation and learning processes.
Solution Approach 2:
The patent introduces additional dimensional structures by organizing video data with multiple layers of metadata including temporal timestamps, spatial coordinates, behavior class labels, and transformation parameters. This dimensional organization transforms high-dimensional raw video data into a structured format that facilitates systematic augmentation while maintaining recognition capability.
3Manufacturing precision
If class-specific augmentation is performed, then class representation improves, but correlation between classes is ignored
Solution Approach 1:
The patent implements a universal augmentation framework that simultaneously handles multiple behavior classes through a single systematic process. The same temporal and spatial transformation operations apply across all classes, ensuring consistent class representation while preserving inter-class correlations. The framework's multi-functionality allows it to augment diverse behavior types uniformly without isolating individual classes.
Data Source
AI summary
An embodiment behavior data augmenting apparatus includes a memory storing algorithms and data and a processor configured to execute the algorithms stored in the memory to extract an object region from video data, define a spatiotemporal characteristic for each class of behavior data by a behavior of an object in the object region, augment the behavior data, and perform learning to recognize the behavior of the object based on the augmented behavior data and a learning algorithm.


