Video Augmentation for Human Activity Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing human activity recognition systems face challenges such as camera motion, occlusion, background clutter, viewpoint variation, and the need for large datasets for training, especially when recognizing complex activities, which degrades their performance.

Innovation Solution

A system and method that utilize video augmentation techniques like image transformation, foreground synthesis, background synthesis, speed variation, motion variation, viewpoint variation, and obfuscation rendering to generate augmented videos, which are combined with original videos to create a diverse set of training videos, enabling deep learning methods for tasks like action classification and anomaly detection even with few demonstration videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a large number of training videos are used for training deep learning methods, then the performance of human activity recognition is improved, but the data collection and processing time increases

Engineering Contradiction:
Improvehuman activity recognition performanceVSAvoiddata collection and processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing a small set of demonstration videos into many augmented training videos before the actual training task. Video augmentation techniques (geometric transformations, color adjustments, speed variations) are performed in advance to create a large training dataset from limited demonstrations, eliminating the need for time-consuming data collection during deployment

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating multiple modified copies of the original demonstration videos through video augmentation. Each original video is transformed into numerous variant copies with different geometric properties, color spaces, and temporal characteristics, effectively multiplying the training data from a small set of demonstrations

Inventive Principle:
Principle #26Copying

2Loss of time

If few-shot learning approaches are used to reduce training data requirements, then data collection time is reduced, but they require large datasets for meta-training and fail on complex human activities

Engineering Contradiction:
Improvedata collection timeVSAvoidcomplex human activity recognition accuracy
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system performs preliminary video augmentation on demonstration videos to create diverse training samples before the actual learning task. This pre-processing step generates sufficient training data without requiring large meta-training datasets, enabling the model to learn complex activities from just a few demonstrations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes by systematically varying video parameters (geometric transformations, color space conversions, speed adjustments) during augmentation. These parameter modifications create diverse training samples that help the model generalize to complex human activities without requiring large datasets

Inventive Principle:
Principle #35Parameter changes

3Reliability

If synthetic 3D human models are used to improve human action recognition, then performance on simple actions is improved, but they cannot handle complex human activities involving objects

Engineering Contradiction:
Improvesimple action recognition performanceVSAvoidcomplex activity understanding capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

Instead of using synthetic 3D models, the patent creates 2D video copies with varied transformations that preserve the original scene's complexity. This approach maintains the ability to handle complex human-object interactions while still providing sufficient data variation for effective learning

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11941080B2System and method for learning human activities from video demonstrations using video augmentation
Publication Date: 2024.03.26 RETROCAUSAL INC
  • US11941080B2 patent drawing
  • US11941080B2 patent drawing
  • US11941080B2 patent drawing

AI summary

A system and method for learning human activities from video demonstrations using video augmentation is disclosed. The method includes receiving original videos from one or more data sources. The method includes processing the received original videos using one or more video augmentation techniques to generate a set of augmented videos. Further, the method includes generating a set of training videos by combining the received original videos with the generated set of augmented videos. Also, the method includes generating a deep learning model for the received original videos based on the generated set of training videos. Further, the method includes learning the one or more human activities performed in the received original videos by deploying the generated deep learning model. The method includes outputting the learnt one or more human activities performed in the original videos.