Video-Based Activity Labeling for Reliable HAR Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of generating reliable and accurate labels for training machine-learning models in Human Activity Recognition (HAR) is significant due to the time-consuming, costly, and error-prone nature of manual annotation, which affects the performance of these models.

Innovation Solution

A method involving the generation of a time-resolved reduced graph representation of human body or body parts from video data, using machine-learning models to automatically generate labels, which are then combined with sensor data to create input-output pairs for training datasets, thereby improving label accuracy and reducing manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used to generate labels for training datasets, then label quality can be ensured, but the process becomes time-consuming, costly, and error-prone

Engineering Contradiction:
Improvelabel qualityVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses video data as a copy or alternative representation of the same physical reality captured by sensors. By processing video data through automated pose estimation and activity recognition models, the system generates labels that replicate the quality of manual annotation without requiring human annotators to watch and label sensor data directly. This copying approach transfers the labeling task from one modality (sensor data) to another (video data) where automated tools are more effective.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of manual annotation with an automated computational system. Instead of human annotators manually reviewing and labeling sensor data, the system uses video-based pose estimation models and activity recognition algorithms to automatically generate labels. This substitution eliminates the time-consuming and error-prone manual process while maintaining or improving label quality through consistent automated application of recognition criteria.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual annotation is used to generate labels, then accurate training data can be obtained, but the cost and complexity increase significantly

Engineering Contradiction:
Improvetraining data reliabilityVSAvoidlabeling system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a parallel labeling system that processes video data alongside sensor data. By generating labels from video through automated pose estimation and activity recognition, the system produces a copy of the ground truth that can be used to train models on sensor data. This approach maintains training data reliability by using the same physical activity captured in both modalities, while avoiding the complexity of manual annotation processes.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces video data as an intermediary medium between the physical activity and the sensor data labeling process. Instead of directly annotating sensor data (which is complex and time-consuming), the system first captures the activity via video, processes it through automated recognition models to generate labels, and then uses these labels for sensor data training. This intermediary approach simplifies the overall system by leveraging well-established video analysis techniques.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If custom training data is collected for different HAR use cases, then model accuracy for specific applications improves, but the data collection and labeling effort increases

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates a universal labeling system based on video data that can serve multiple HAR use cases simultaneously. The same video-based pose estimation and activity recognition pipeline can generate labels for different applications (fall detection, gesture recognition, sport analysis, etc.) without requiring separate manual annotation processes for each use case. This multi-functional approach maintains model accuracy for specific applications while dramatically improving data collection efficiency through automated processing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses video data as a universal source that can be copied and processed for multiple different HAR applications. By capturing activities via video and generating standardized labels through automated models, the same data collection infrastructure supports diverse use cases. Researchers can collect video data once and generate labels for multiple different activity recognition tasks, eliminating the need to separately collect and annotate data for each application.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4682840A1Label generation for human activity recognition
Publication Date: 2026.01.21 INFINEON TECHNOLOGIES AG
  • EP4682840A1 patent drawingFigure 1
  • EP4682840A1 patent drawingFigure 2
  • EP4682840A1 patent drawingFigure 3

AI summary

Video data (121, 122, 123) depicting a human body or body part is obtained. Based on the video data, a time-resolved reduced graph representation (190) of the human body a body part is generated. Based on the time-resolved reduced graph representation, a label (129) associated with an activity of the human body or body part is determined. Such labels can be used to populate a training dataset that can later on be used for training a machine-learning model, e.g., for gesture classification, fall detection, or other tasks, based on sensor data that observes the human body or body part.