Unified 2D and 3D Pose Recognition Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural network models for human body pose recognition are incompatible between 2D and 3D pose information, requiring separate training and consuming significant computing resources, resulting in low training efficiency.
Innovation Solution
A pose recognition model is developed that processes sample images to obtain both 2D and 3D key point parameters using integrated 2D and 3D models, constructing a target loss function to update the model, enabling compatibility and improving training efficiency by using a single set of training samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If separate neural network models are used for 2D and 3D pose recognition, then each model can be optimized for its specific task, but the training requires double the computing resources and time
Solution Approach 1:
The patent merges 2D and 3D pose recognition models into a single integrated neural network model. The model contains both 2D detection modules (for detecting key points in 2D space) and 3D reconstruction modules (for inferring 3D pose), allowing both functionalities to coexist in one model. This integration enables simultaneous optimization of both 2D and 3D pose recognition while reducing redundant computing resources, as the shared feature extraction layers serve both purposes.
Solution Approach 2:
The integrated model achieves multi-functionality by incorporating both 2D pose detection and 3D pose reconstruction capabilities within a single neural network architecture. The model can perform 2D key point detection on input images and simultaneously reconstruct 3D human pose from the same features, making it a universal pose recognition system that eliminates the need for separate specialized models.
2Reliability
If separate training processes are used for 2D and 3D models, then each model can be trained independently with tailored loss functions, but the total training time and resource consumption double
Solution Approach 1:
The patent combines multiple training objectives into a unified training process. The integrated model uses a composite loss function that includes both 2D detection loss (for key point accuracy) and 3D reconstruction loss (for pose accuracy), allowing simultaneous optimization of both functionalities in one training run. This unified approach maintains model performance while reducing training time from separate sequential training to a single parallel training process.
Solution Approach 2:
The model architecture is segmented into distinct functional modules (2D detection modules and 3D reconstruction modules) that can be independently optimized through their respective loss functions, while still being trained together in a unified framework. This modular segmentation within integration allows tailored optimization for each task while benefiting from shared feature representations.
3Productivity
If a single integrated model is used for both 2D and 3D pose recognition, then computing resources are saved, but the model complexity increases
Solution Approach 1:
The integrated model is segmented into distinct functional modules: feature extraction layers, 2D detection modules (including key point detection and affinity field prediction), and 3D reconstruction modules. This modular segmentation organizes the complexity into manageable components with clear interfaces, making the overall complex system easier to understand, train, and maintain while preserving the efficiency benefits of integration.
Data Source
AI summary
This application provides a method for training a pose recognition model performed at a computer device. The method includes: inputting a sample image labeled with human body key points into a feature map model included in a pose recognition model, to output a feature map of the sample image; inputting the feature map into a two-dimensional (2D) model included in the pose recognition model, to output 2D key point parameters used for representing a 2D human body pose; input a target human body feature map cropped from the feature map and the 2D key point parameter into a three-dimensional (3D) model included in the pose recognition model, to output 3D pose parameters used for representing a 3D human body pose; constructing a target loss function based on the 2D key point parameters and the 3D pose parameters; and updating the pose recognition model based on the target loss function.


