Fitness Action Recognition Model Using Depth Image Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional fitness action recognition technologies face challenges in accurately evaluating the capability value of training objects due to factors like background color similarity and presence of other viewers, leading to inaccurate action recognition and privacy concerns during fitness processes.
Innovation Solution
A fitness action recognition model incorporating an information extraction layer, pixel point positioning layer, feature extraction layer, vector dimensionality reduction layer, and feature vector classification layer, utilizing a random decision forest for position estimation and multidimensional feature vector classification, which improves key-point calibration accuracy and eliminates privacy risks by using depth images from a three-dimensional visual sensor.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional color image recognition is used, then the system can process images, but accuracy deteriorates when training object color is similar to background or when other viewers are present
Solution Approach 1:
The patent introduces depth information as an intermediary between the camera and the training object. By using depth images and depth-based feature extraction, the system can distinguish the training object from the background and other viewers based on spatial depth rather than color similarity, thereby resolving the interference problem while maintaining recognition accuracy.
Solution Approach 2:
The patent transitions from two-dimensional color image analysis to three-dimensional depth-based analysis. By extracting features from depth images and using depth information in the feature vector, the system adds a spatial dimension that enables accurate distinction between the training object and background elements regardless of color similarity.
2Measurement precision
If depth images are used to improve recognition accuracy, then measurement precision improves, but device complexity increases due to three-dimensional visual sensor requirements
Solution Approach 1:
The patent makes the depth image processing system multi-functional by using the same depth image data for multiple purposes: key-point detection, feature extraction, and action recognition. This universal approach justifies the device complexity by delivering multiple benefits from a single sensor type, including improved accuracy and privacy protection.
Solution Approach 2:
The patent changes the fundamental parameter used for image analysis from color information to depth information. By transforming the input from color images to depth images and adjusting the feature extraction parameters accordingly, the system achieves superior accuracy while the complexity increase is offset by the enhanced performance in challenging scenarios.
3Measurement precision
If color images are collected for training object analysis, then action recognition can be performed, but privacy leakage risk increases
Solution Approach 1:
The patent extracts only the necessary depth information from the three-dimensional visual sensor data while excluding color information and other unnecessary data. By taking out only the essential depth-based features needed for action recognition, the system maintains recognition capability while eliminating the privacy leakage risks associated with collecting and storing color images and personal identifiable information.
Solution Approach 2:
The patent creates a simplified copy of the visual data that contains only depth information rather than full color images. This copy retains the essential spatial structure needed for action recognition while removing sensitive color and personal information, thereby protecting privacy while preserving analytical capability.
Data Source
AI summary
A model including an information extraction layer that obtains image information of a training object in a depth image; a pixel point positioning layer that performs position estimation on a three-dimensional coordinate of human-body key points, defines a body part of the training object as a body component, and calibrates a three-dimensional coordinate of all human-body key points corresponding to the body component; a feature extraction layer that extracts a key-point position feature, a body moving speed feature, and a key-point moving speed feature for action recognition; a vector dimensionality reduction layer that combines the key-point position feature, the body moving speed feature, and the key-point moving speed feature as a multidimensional feature vector, and performs dimensionality reduction on the multidimensional feature vector; and a feature vector classification layer that classifies the multidimensional feature vector that is performed with dimensionality reduction, to recognize a fitness action of the training object.


