Unified 2D and 3D Pose Recognition Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current neural network models for human body pose recognition are incompatible between 2D and 3D pose information, requiring separate training and consuming significant computing resources, resulting in low training efficiency.

Innovation Solution

A pose recognition model is developed that processes sample images to obtain both 2D and 3D key point parameters using integrated 2D and 3D models, constructing a target loss function to update the model, enabling compatibility and improving training efficiency by using a single set of training samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If separate neural network models are used for 2D and 3D pose recognition, then each model can be optimized for its specific task, but the training requires double the computing resources and time

Engineering Contradiction:
Improvepose recognition accuracyVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent merges 2D and 3D pose recognition models into a single integrated neural network model. The model contains both 2D detection modules (for detecting key points in 2D space) and 3D reconstruction modules (for inferring 3D pose), allowing both functionalities to coexist in one model. This integration enables simultaneous optimization of both 2D and 3D pose recognition while reducing redundant computing resources, as the shared feature extraction layers serve both purposes.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The integrated model achieves multi-functionality by incorporating both 2D pose detection and 3D pose reconstruction capabilities within a single neural network architecture. The model can perform 2D key point detection on input images and simultaneously reconstruct 3D human pose from the same features, making it a universal pose recognition system that eliminates the need for separate specialized models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If separate training processes are used for 2D and 3D models, then each model can be trained independently with tailored loss functions, but the total training time and resource consumption double

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent combines multiple training objectives into a unified training process. The integrated model uses a composite loss function that includes both 2D detection loss (for key point accuracy) and 3D reconstruction loss (for pose accuracy), allowing simultaneous optimization of both functionalities in one training run. This unified approach maintains model performance while reducing training time from separate sequential training to a single parallel training process.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The model architecture is segmented into distinct functional modules (2D detection modules and 3D reconstruction modules) that can be independently optimized through their respective loss functions, while still being trained together in a unified framework. This modular segmentation within integration allows tailored optimization for each task while benefiting from shared feature representations.

Inventive Principle:
Principle #1Segmentation

3Productivity

If a single integrated model is used for both 2D and 3D pose recognition, then computing resources are saved, but the model complexity increases

Engineering Contradiction:
Improvetraining efficiencyVSAvoidmodel structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The integrated model is segmented into distinct functional modules: feature extraction layers, 2D detection modules (including key point detection and affinity field prediction), and 3D reconstruction modules. This modular segmentation organizes the complexity into manageable components with clear interfaces, making the overall complex system easier to understand, train, and maintain while preserving the efficiency benefits of integration.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11907848B2Method and apparatus for training pose recognition model, and method and apparatus for image recognition
Publication Date: 2024.02.20 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11907848B2 patent drawing
  • US11907848B2 patent drawing
  • US11907848B2 patent drawing

AI summary

This application provides a method for training a pose recognition model performed at a computer device. The method includes: inputting a sample image labeled with human body key points into a feature map model included in a pose recognition model, to output a feature map of the sample image; inputting the feature map into a two-dimensional (2D) model included in the pose recognition model, to output 2D key point parameters used for representing a 2D human body pose; input a target human body feature map cropped from the feature map and the 2D key point parameter into a three-dimensional (3D) model included in the pose recognition model, to output 3D pose parameters used for representing a 3D human body pose; constructing a target loss function based on the 2D key point parameters and the 3D pose parameters; and updating the pose recognition model based on the target loss function.