Fitness Action Recognition Model Using Depth Image Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional fitness action recognition technologies face challenges in accurately evaluating the capability value of training objects due to factors like background color similarity and presence of other viewers, leading to inaccurate action recognition and privacy concerns during fitness processes.

Innovation Solution

A fitness action recognition model incorporating an information extraction layer, pixel point positioning layer, feature extraction layer, vector dimensionality reduction layer, and feature vector classification layer, utilizing a random decision forest for position estimation and multidimensional feature vector classification, which improves key-point calibration accuracy and eliminates privacy risks by using depth images from a three-dimensional visual sensor.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional color image recognition is used, then the system can process images, but accuracy deteriorates when training object color is similar to background or when other viewers are present

Engineering Contradiction:
Improveaction recognition accuracyVSAvoidbackground color similarity and other viewers interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces depth information as an intermediary between the camera and the training object. By using depth images and depth-based feature extraction, the system can distinguish the training object from the background and other viewers based on spatial depth rather than color similarity, thereby resolving the interference problem while maintaining recognition accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from two-dimensional color image analysis to three-dimensional depth-based analysis. By extracting features from depth images and using depth information in the feature vector, the system adds a spatial dimension that enables accurate distinction between the training object and background elements regardless of color similarity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If depth images are used to improve recognition accuracy, then measurement precision improves, but device complexity increases due to three-dimensional visual sensor requirements

Engineering Contradiction:
Improvekey-point calibration accuracyVSAvoidthree-dimensional visual sensor system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the depth image processing system multi-functional by using the same depth image data for multiple purposes: key-point detection, feature extraction, and action recognition. This universal approach justifies the device complexity by delivering multiple benefits from a single sensor type, including improved accuracy and privacy protection.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the fundamental parameter used for image analysis from color information to depth information. By transforming the input from color images to depth images and adjusting the feature extraction parameters accordingly, the system achieves superior accuracy while the complexity increase is offset by the enhanced performance in challenging scenarios.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If color images are collected for training object analysis, then action recognition can be performed, but privacy leakage risk increases

Engineering Contradiction:
Improveaction recognition capabilityVSAvoidprivacy leakage risk
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent extracts only the necessary depth information from the three-dimensional visual sensor data while excluding color information and other unnecessary data. By taking out only the essential depth-based features needed for action recognition, the system maintains recognition capability while eliminating the privacy leakage risks associated with collecting and storing color images and personal identifiable information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a simplified copy of the visual data that contains only depth information rather than full color images. This copy retains the essential spatial structure needed for action recognition while removing sensitive color and personal information, thereby protecting privacy while preserving analytical capability.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11854306B1Fitness action recognition model, method of training model, and method of recognizing fitness action
Publication Date: 2023.12.26 NANJING SILICON INTELLIGENCE TECH CO LTD
  • US11854306B1 patent drawing
  • US11854306B1 patent drawing
  • US11854306B1 patent drawing

AI summary

A model including an information extraction layer that obtains image information of a training object in a depth image; a pixel point positioning layer that performs position estimation on a three-dimensional coordinate of human-body key points, defines a body part of the training object as a body component, and calibrates a three-dimensional coordinate of all human-body key points corresponding to the body component; a feature extraction layer that extracts a key-point position feature, a body moving speed feature, and a key-point moving speed feature for action recognition; a vector dimensionality reduction layer that combines the key-point position feature, the body moving speed feature, and the key-point moving speed feature as a multidimensional feature vector, and performs dimensionality reduction on the multidimensional feature vector; and a feature vector classification layer that classifies the multidimensional feature vector that is performed with dimensionality reduction, to recognize a fitness action of the training object.