A 3D Human Pose Estimation Method Based on an Enhanced Topology-Aware Network

By introducing an enhanced topological perception network into the three-dimensional human pose estimation model, combining the space-time dual-branch Transformer and the hybrid constraint module, the problem of the model's poor performance in complex poses is solved, and a more accurate three-dimensional pose estimation is achieved.

CN119887928BActive Publication Date: 2025-06-20NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361017.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-20
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When modeling the global spatiotemporal correlation of joints, the existing three-dimensional human posture estimation model ignores the local topological connections and structural constraints of human skeletons, resulting in poor performance in complex postures.

Method used

The enhanced topology perception network is adopted to enhance the network's learning of human topology structure through space-time dual-branch Transformer and hybrid constraint module, and combine local topology constraints and global dependencies to generate a more accurate three-dimensional pose.

Benefits of technology

Through the enhanced topological perception network, the model can more accurately model the three-dimensional pose of the human body, especially in complex poses, improving the accuracy and robustness of the task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887928B_ABST
    Figure CN119887928B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional human pose estimation method based on an enhanced topology-aware network. The method includes obtaining a human motion capture data set; constructing an enhanced topology-aware network model, which includes a feature embedding block, an enhanced topology-aware module stacked repeatedly 5 times, and a regression head connected in sequence. The enhanced topology-aware module includes a spatio-temporal dual-branch Transformer and a hybrid constraint module; training the model using the data set to obtain a final enhanced topology-aware network model; inputting the human body picture or video to be detected into the final enhanced topology-aware network model to obtain the three-dimensional coordinates corresponding to each joint, and completing the estimation of the three-dimensional human pose. The three-dimensional pose coordinates generated by the present invention are closer to the real situation and have higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and particularly relates to a three-dimensional human pose estimation method based on an enhanced topology-aware network. Background Art

[0002] 3D HPE (3D Human Pose Estimation) is a core research topic in the field of computer vision, aiming to accurately recover the three-dimensional joint positions of the human body from two-dimensional images or videos. In recent years, thanks to the development of deep learning algorithms, the three-dimensional human pose estimation task has developed rapidly. The human body topology reflects the spatial and motion relationships between joints and is the essential feature of human poses. Accurately modeling this structure is crucial for improving the accuracy and robustness of the task. Transformer-based models have made significant progress in the three-dimensional human pose estimation task. These methods use the multi-head self-attention mechanism to model the global dependencies of joints and have advantages in capturing long-range spatio-temporal correlations between joints. However, the self-attention mechanism overemphasizes global relationships and ignores the local topological connections and structural constraints of the human body skeleton. This limitation makes it difficult for the model to fully utilize the prior knowledge of the human body skeleton structure, especially performing poorly in complex poses. Therefore, how to fully consider the human body topology while modeling the global spatio-temporal correlations of joints, strengthen the modeling of the topological dependencies between human joints, and generate more accurate three-dimensional poses has become a major problem. Summary of the Invention

[0003] The technical problem to be solved by the present invention is: to provide a three-dimensional human pose estimation method based on an enhanced topology-aware network, which solves the problem of insufficient modeling of the human body topology.

[0004] To solve the above technical problems, the present invention adopts the following technical solutions:

[0005] A three-dimensional human pose estimation method based on an enhanced topology-aware network includes the following steps:

[0006] S1. Obtain a human motion capture data set.

[0007] S2. Construct an enhanced topology-aware network model, and use the data set in step S1 to train the model to obtain the final enhanced topology-aware network model.

[0008] S3. Input the human body picture or video to be detected into the final enhanced topology-aware network model to obtain the three-dimensional coordinates corresponding to each joint, and complete the three-dimensional human pose estimation.

[0009] Further, in step S1, the two-dimensional coordinates, three-dimensional coordinates, and their ground truths of the joints are obtained from the large-scale motion capture datasets of Human3.6M and MPI-INF-3DHP.

[0010] Further, in step S2, the enhanced topology-aware network model includes a feature embedding block, an enhanced topology-aware module stacked repeatedly 5 times, and a regression head connected in sequence.

[0011] Among them, the enhanced topology-aware module includes a spatio-temporal dual-branch Transformer and a hybrid constraint module.

[0012] The dataset in step S1 is divided into a training set and a test set according to 7:3. The enhanced topology-aware network model is trained using the training set, and the trained enhanced topology-aware network model is tested using the test set to obtain the final enhanced topology-aware network model.

[0013] Further, in step S3, the estimation of the three-dimensional human pose includes the following:

[0014] The two-dimensional coordinates of the joints are obtained using a two-dimensional pose detector , where T and N represent the number of frames and the number of joints in the sequence respectively, and the number 3 represents the dimension, which includes the horizontal and vertical coordinates and the confidence score of the joint; the two-dimensional coordinates are input into the final enhanced topology-aware network model, and after passing through the feature embedding block, the two-dimensional coordinates are projected into a high dimension to obtain a preliminary high-dimensional feature , where C represents the dimension size; a set of tensors is added, and after adding it to the preliminary high-dimensional feature, it is input into the enhanced topology-aware module. The spatio-temporal global dependencies between joints are calculated using the spatio-temporal dual-branch Transformer to obtain a fused intermediate feature. According to the degrees of freedom of the joints and the limb categories to which they belong, the local topological constraints of different joints are obtained using the hybrid constraint module respectively, and the final hybrid topological constraint is obtained through adaptive fusion. The fused intermediate feature is structurally guided using this constraint to complete the operation of the enhanced topology-aware module;

[0015] The operations performed in the enhanced topology-aware module are repeated 5 times to obtain an enhanced feature of the human topological structure. After passing through the regression head, the final three-dimensional pose coordinates are predicted using a linear layer .

[0016] Further, the shapes of the tensors are N×C and T×1×C respectively and are initialized to 0.

[0017] Further, obtaining the fused intermediate feature includes the following:

[0018] The spatio-temporal dual-branch Transformer includes Transformer blocks stacked in a spatial-temporal order and Transformer blocks stacked in a temporal-spatial order; among them, the Transformer blocks stacked in a spatial-temporal order include a spatial encoder and a temporal encoder connected in sequence, and the Transformer blocks stacked in a temporal-spatial order include a temporal encoder and a spatial encoder connected in sequence.

[0019] The added tensor and the preliminary high-dimensional features pass through the spatio-temporal dual-branch Transformer to obtain corresponding intermediate features, and the specific expression is:

[0020] ;

[0021] ;

[0022] Among them, P1 represents the intermediate feature obtained through the Transformer blocks stacked in a spatial-temporal order, P2 represents the intermediate feature obtained through the Transformer blocks stacked in a temporal-spatial order, TTE represents the temporal encoder, and STE represents the spatial encoder.

[0023] Perform adaptive fusion on P1 and P2 to obtain the fused intermediate feature, and the specific expression is:

[0024] ;

[0025] ;

[0026] Among them, W represents the tensor after dimension transformation, FC represents the linear layer, Concat represents the concatenation operation, F represents the fused intermediate feature, and both W1 and W2 represent weights.

[0027] Furthermore, group the joints according to the degrees of freedom of the joints, and the specific expression is:

[0028] ;

[0029] ;

[0030] ;

[0031] Among them, DoF1, DoF2, and DoF3 all represent the degree-of-freedom grouping of joints, right_shoulder represents the right shoulder, left_shoulder represents the left shoulder, right_hip represents the right hip, left_hip represents the left hip, right_elbow represents the right elbow, left_elbow represents the left elbow, right_knee represents the right knee, left_knee represents the left knee, right_wrist represents the right wrist, left_wrist represents the left wrist, right_feet represents the right foot, and left_feet represents the left foot.

[0032] Group the joints according to the limb category to which they belong. The specific expression is:

[0033] ;

[0034] ;

[0035] ;

[0036] ;

[0037] Among them, , , , respectively represent the right arm, left arm, right leg, and left leg of the human body.

[0038] The expression for static joint grouping is:

[0039] ;

[0040] Among them, Static represents static joint grouping, head represents the head, neck represents the neck, thorax represents the chest, spine represents the spine, and hip represents the hip joint.

[0041] Furthermore, group the fused intermediate features according to the degree of freedom of the joints and the limb category to which they belong to obtain the degree-of-freedom grouping features and the limb-category grouping features to which they belong.

[0042] Perform feature dimension conversion on the grouped features to obtain the converted grouped features , where represents the i-th degree-of-freedom grouping feature, represents the j-th limb-category grouping feature to which it belongs, represents the static joint grouping feature.

[0043] Concatenate each degree-of-freedom grouping feature along the joint dimension to obtain the overall feature of the degree-of-freedom grouping , through a two-dimensional convolutional layer Conv2d with a kernel size of 4×3 D for feature extraction to obtain aggregated features containing each degree-of-freedom grouping , split along the joint dimension to obtain the local topological constraints of the i-th degree-of-freedom grouping feature , and the specific formula is:

[0044] ;

[0045] ;

[0046] ;

[0047] where represents the GELU activation function, Split represents splitting the feature along the joint dimension, represents the first degree-of-freedom grouping feature, represents the second degree-of-freedom grouping feature, represents the third degree-of-freedom grouping feature.

[0048] Concatenate the grouped features of each limb category along the joint dimension to obtain the overall feature of the grouped limb category , through a two-dimensional convolution Conv2d with a kernel size of 3×3 p for feature extraction to obtain aggregated features containing each grouped limb category , split along the joint dimension to obtain the local topological constraints of the j-th grouped limb category feature , and the specific formula is:

[0049] ;

[0050] ;

[0051] ;

[0052] where represents the first grouped limb category feature, represents the second grouped limb category feature, represents the third grouped limb category feature, represents the fourth grouped limb category feature.

[0053] The static joint grouped features pass through a two-dimensional convolution Conv2d with a kernel size of 5×3 s for feature extraction to obtain the static joint grouped aggregated features , and then obtain the local topological constraint of the k-th static joint grouping feature , and the specific formula is:

[0054] ;

[0055] .

[0056] Combine the local topological constraint of the grouping feature of the limb category to which it belongs with the local topological constraint of the grouping feature of different degrees of freedom, and based on the weight parameter, obtain the final hybrid topological constraint. The specific expression is:

[0057] ;

[0058] ;

[0059] ;

[0060] Among them, represents the hybrid feature of the i-th degree-of-freedom grouping feature in the j-th grouping feature of the limb category to which it belongs, represents the learnable parameter initialized to 0 corresponding to , R represents the final hybrid topological constraint, concat order represents the concatenation operation in order, and Y represents the result after adding the hybrid topological constraint.

[0061] Furthermore, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the three-dimensional human pose estimation method based on the enhanced topological perception network are implemented.

[0062] Furthermore, the present invention also proposes a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is run by a processor, the three-dimensional human pose estimation method based on the enhanced topological perception network is executed.

[0063] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0064] By designing a strengthened topological perception network structure, the present invention uses a hybrid topological constraint module to enhance the learning of the human topological structure by the network model, solves the problem of insufficient modeling of the human topological structure by the model, and generates more accurate three-dimensional pose coordinates. And the method proposed by the present invention is verified to be feasible and effective on the commonly used Human3.6M and MPI-INF-3DHP data sets. Description of the Drawings

[0065] Figure 1 It is the overall implementation flowchart of the present invention.

[0066] Figure 2 It is a schematic diagram of the joint grouping method in the embodiment of the present invention.

[0067] Figure 3 It is the operation diagram of the hybrid topology constraint module in the embodiment of the present invention.

[0068] Figure 4 It is the result visualization diagram of the embodiment of the present invention. Specific implementation manner

[0069] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.

[0070] To achieve the above object, the present invention proposes a three-dimensional human pose estimation method based on an enhanced topology-aware network, as Figure 1 shown, the specific steps are as follows:

[0071] S1. Obtain the two-dimensional coordinates, three-dimensional coordinates and their ground truths of the joints from the large-scale motion capture datasets of Human3.6M and MPI-INF-3DHP.

[0072] S2. Construct an ETANet model, and use the dataset in step S1 to train the model to obtain the final enhanced topology-aware network model. The specific content is as follows:

[0073] Construct an ETANet (Enhanced Topology-Aware Network) model through the PyTorch deep learning framework. First, perform environment configuration and data loading, use its built-in DataLoader method to process the dataset, and divide it into sequences with a specified number of frames (such as 243 frames). Then, build the network structure, define the forward propagation process and loss function of the model according to its specifications, and optimize the model.

[0074] The enhanced topology-aware network model includes a feature embedding block, an enhanced topology-aware module stacked repeatedly 5 times, and a regression head connected in sequence.

[0075] Among them, the enhanced topology perception module includes a spatio-temporal dual-branch Transformer and an MTC (Mixed Topological Constraints) based on human body bone topology. MTC is simple, efficient, flexible and easy to use. It is a plug-and-play module. By combining with different Transformer models, it significantly improves the accuracy and robustness of 3D pose estimation. MTC aims to learn the topological prior of the human body bones. By constraining the feature expressions of different joints, it guides the model to generate more reasonable 3D poses, further reducing errors and improving the model performance. The core of MTC is to group and mix according to different logics to obtain more powerful topological constraints.

[0076] Divide the dataset in step S1 into a training set and a test set according to 7:3. Use the training set to train the enhanced topology perception network model, and use the test set to test the trained enhanced topology perception network model. In the test phase, the parameters in the network model are no longer updated. This phase is only used to evaluate the performance of the network model. After the training and test phases, the final enhanced topology perception network model is obtained.

[0077] In this embodiment, the data of participants numbered 1, 5, 6, 7, and 8 are used for training, and the data of participants numbered 9 and 11 are used for testing. The NVIDIA RTX 4090 GPU is used to accelerate the model training.

[0078] In terms of model training, set the batch size to 4, the initial learning rate to 0.0005, train for a total of 100 epochs, and the learning rate decays by 0.99 after each epoch. Select the AdamW optimizer to optimize the model parameters and specify the weight decay to be 0.01.

[0079] S3. Input the human body picture or video to be detected into the final enhanced topology perception network model to obtain the 3D coordinates corresponding to each joint, and complete the estimation of the 3D human body pose. The specific content is as follows:

[0080] Use a 2D pose detector to obtain the 2D coordinates of the joints , where T and N respectively represent the number of frames of the sequence and the number of joints, and the number 3 represents the dimension, which includes the horizontal and vertical coordinates and the confidence score of the joint; input the 2D coordinates into the final enhanced topology perception network model, and through the feature embedding block, project the 2D coordinates into a high dimension to obtain preliminary high-dimensional features .

[0081] To address the lack of the model's ability to handle the order of elements in a sequence, it is also necessary to introduce the position information of each element in the sequence through positional embeddings. Add a set of tensors with shapes N×C and T×1×C respectively, initialized to 0, which can be learned by the model. After adding them to the preliminary high-dimensional features, input them into the enhanced topological awareness module, and use the spatio-temporal dual-branch Transformer to calculate the spatio-temporal global dependencies between joints to obtain the fused intermediate features. The specific content is as follows:

[0082] The spatio-temporal dual-branch Transformer includes Transformer blocks stacked in spatio-temporal order and Transformer blocks stacked in time-space order; among them, the Transformer blocks stacked in spatio-temporal order include a spatial encoder and a temporal encoder connected in sequence, and the Transformer blocks stacked in time-space order include a temporal encoder and a spatial encoder connected in sequence.

[0083] Both Transformer blocks consider the correlations between different joints within the same frame and the long-range dependencies of the same joint across frames.

[0084] The added tensors and the preliminary high-dimensional features pass through the spatio-temporal dual-branch Transformer to obtain the corresponding intermediate features. The specific expression is as follows:

[0085] ;

[0086] ;

[0087] Among them, P1 represents the intermediate feature obtained through the Transformer blocks stacked in spatio-temporal order, P2 represents the intermediate feature obtained through the Transformer blocks stacked in time-space order, TTE represents the temporal encoder, and STE represents the spatial encoder.

[0088] Perform adaptive fusion on P1 and P2. On the premise of keeping the feature dimension unchanged, fully integrate the temporal and spatial features to enhance the model's ability to model the global and local dependencies of three-dimensional postures, and obtain the fused intermediate feature. The specific expression is as follows:

[0089] ;

[0090] ;

[0091] Among them, W represents the tensor after dimension transformation; FC represents the linear layer; Concat represents the concatenation operation; F represents the fused intermediate feature; W1 and W2 both represent weights, which are taken from two values of the last dimension of W respectively.

[0092] The definition of joint degrees of freedom is the rotational degrees of freedom of a joint relative to its parent node. Since the joints at the ends of the limbs have a larger range of motion, higher degrees of freedom, and higher pose estimation errors, more refined modeling is required.

[0093] To cooperate with the grouping of limb categories and distinguish from the static joint grouping, only 12 joints including the four limbs are considered here. The 12 joints are grouped according to the degrees of freedom of the joints, as shown in (a) of Figure 2 The specific expression is:

[0094] ;

[0095] ;

[0096] ;

[0097] where DoF1, DoF2, and DoF3 all represent the degree-of-freedom grouping of joints, right_shoulder represents the right shoulder, left_shoulder represents the left shoulder, right_hip represents the right hip, left_hip represents the left hip, right_elbow represents the right elbow, left_elbow represents the left elbow, right_knee represents the right knee, left_knee represents the left knee, right_wrist represents the right wrist, left_wrist represents the left wrist, right_feet represents the right foot, and left_feet represents the left foot.

[0098] Each group contains 4 joints with the same degree of freedom. The larger the grouping subscript, the higher the DoF (Degree of Freedom) of the joints within the group.

[0099] Based on the basic principles of human kinematics and anatomy, the method of grouping the four limbs is used for 3D human pose estimation to adapt to the motion characteristics of different limbs. Similarly, to cooperate with the DoF grouping and distinguish from the static joint grouping, the 12 joints are grouped according to the limb categories to which they belong, as shown in (b) of Figure 2 The specific expression is:

[0100] ;

[0101] ;

[0102] ;

[0103] ;

[0104] where , , , respectively represent the four limbs of the right arm, left arm, right leg, and left leg of the human body.

[0105] Since the spine-related joints are relatively rigid and are usually used to represent the core structure and key nodes of the torso, while the limb joint groups are more flexible, and in order not to interfere with the mixing of the degrees of freedom and the two groupings of the limbs, the present invention separately considers the spine-related joints as a static joint grouping, which helps to reduce the movement conflicts between different joints, makes the movement of the spine more in line with the physiological reality, thereby avoiding unnatural pose prediction, and improving the accuracy and stability of the model. The expression of the static joint grouping is:

[0106] ;

[0107] where Static represents the static joint grouping, head represents the head, neck represents the neck, thorax represents the chest, spine represents the spine, and hip represents the hip joint.

[0108] So far, the joint groups after various groupings have been obtained. As Figure 3 shown, according to the degrees of freedom of the joints and the limb categories to which they belong, the local topological constraints of different joints are respectively obtained by using the hybrid constraint module, and the final hybrid topological constraint is obtained through adaptive fusion. Using this constraint to structurally guide the fused intermediate features completes the operation of the enhanced topological perception module; the specific content is:

[0109] Group the fused intermediate features according to the degrees of freedom of the joints and the limb categories to which they belong to obtain the degree-of-freedom grouped features and the limb-category grouped features to which they belong.

[0110] Perform feature dimension conversion on the grouped features to obtain the converted grouped features , where represents the i-th degree-of-freedom grouped feature, represents the j-th limb-category grouped feature to which it belongs, represents the static joint grouping feature. There is no subscript because there is only one joint group in this grouping method.

[0111] Concatenate each degree-of-freedom grouped feature along the joint dimension to obtain the overall feature of the degree-of-freedom grouping , passing through a two-dimensional convolutional layer Conv2d with a convolutional kernel size of 4×3, a padding setting of (0, 1), a stride setting of (4, 1), and keeping the number of feature channels unchanged DFeature extraction is performed so that while the convolution extracts each grouped feature, it captures the local temporal dependencies between adjacent three frames, complementing the global temporal dependencies captured by the Transformer to extract more comprehensive temporal correlation information and obtain the aggregated features containing each degree-of-freedom grouping , split along the joint dimension to obtain the local topological constraints of the grouped features of the i-th degree of freedom , and the specific formula is:

[0112] ;

[0113] ;

[0114] ;

[0115] where represents the GELU activation function, Split represents splitting the features along the joint dimension, represents the grouped features of the 1st degree of freedom, represents the grouped features of the 2nd degree of freedom, represents the grouped features of the 3rd degree of freedom.

[0116] Concatenate the grouped features of each limb category along the joint dimension to obtain the overall features of the grouped limb categories , and perform feature extraction through a 2D convolution Conv2d with a kernel size of 3×3, a padding setting of (0, 1), a stride setting of (3, 1), and keeping the number of feature channels unchanged p to obtain the aggregated features containing each grouped limb category , split along the joint dimension to obtain the local topological constraints of the grouped features of the j-th limb category , and the specific formula is:

[0117] ;

[0118] ;

[0119] ;

[0120] where represents the grouped features of the 1st limb category, represents the grouped features of the 2nd limb category, represents the grouped features of the 3rd limb category, represents the grouped features of the 4th limb category.

[0121] Since there is only one group in the static joint grouping, the feature splicing process is omitted. The static joint grouping feature undergoes a two-dimensional convolution Conv2d with a convolution kernel size of 5×3, a padding setting of (0, 1), a stride setting of 1, and the number of feature channels remains unchanged s for feature extraction to obtain the aggregated feature of static joint grouping , and then the local topological constraint of the k-th static joint grouping feature is obtained , and the specific formula is:

[0122] ;

[0123] .

[0124] After limb grouping, the 3 joints in each group exactly correspond to 3 different degrees of freedom. Therefore, the local topological constraint of the grouped feature of the limb category to which it belongs is combined with the local topological constraint of the grouped feature of different degrees of freedom, and based on the weight parameter, which is initialized to 0, to obtain a more refined and flexible final hybrid topological constraint, guiding the model to more comprehensively model the human body bone topological relationship and predict a more reasonable three-dimensional pose. The specific expression is:

[0125] ;

[0126] ;

[0127] ;

[0128] where represents the hybrid feature of the i-th degree-of-freedom grouped feature in the j-th grouped feature of the limb category to which it belongs, represents the learnable parameter initialized to 0 corresponding to , R represents the final hybrid topological constraint, concat order represents the concatenation operation in order, and Y represents the result after adding the hybrid topological constraint

[0129] The topological constraint formed by aggregating these joint features that contain both joint degree-of-freedom information and limb category information to which they belong is added to the input, and at the same time, a residual connection is applied to strengthen the model's learning of the human body topological structure

[0130] The operations performed in the enhanced topology perception module are repeated 5 times to obtain the enhanced feature of the human body topological structure. After passing through the regression head, the final three-dimensional pose coordinates are predicted using a linear layer .

[0131] Figure 4 is the visualization result comparison chart of the method of the present invention and the MixSTE method, where Figure 4Among them, (a) is the input two-dimensional human body posture diagram, Figure 4 and (b) is the three-dimensional human body posture diagram generated by the MixSTE method, Figure 4 and (c) is the three-dimensional human body posture diagram generated by the method of the present invention, Figure 4 and (d) is the true three-dimensional human body posture diagram corresponding to the input two-dimensional posture. It can be seen that the result of the MixSTE method does not well restore the posture of the human body with crossed feet, and the left shoulder joint of the human body is too high, and the prediction accuracy is poor. The method of the present invention restores the details at the left and right ankle joints and the left shoulder joint, and the prediction of the head joint is closer to the true value, and the result is more accurate compared with MixSTE. Therefore, the method of the present invention can achieve better results in the task of three-dimensional human body posture estimation, and the generated three-dimensional posture is more accurate.

[0132] Compare the performance of different three-dimensional human body posture estimation methods, as shown in Tables 1 and 2.

[0133] Table 1 Comparison results of the performance of different three-dimensional human body posture estimation methods on Human3.6M

[0134]

[0135] In Table 1, UGCN (U-shaped Graph Convolutional Network), PoseFormer, StridedFormer, MHFormer (Multi-Hypothesis Transformer), MixSTE (Mixed Spatio-Temporal Encoder), P-STMO (Pre-trained Spatial Temporal Many-to-One), GLA-GCN (Global-local Adaptive Graph Convolutional Network), POT (Pose-oriented transformer), STCFormer (Spatio-Temporal Criss-cross attention Transformer), KTPFormer (Kinematics and Trajectory Prior Knowledge-Enhanced Transformer), NC-RetNet (Non-Causal RetNet), and TMT (Two-step Mixed-Training strategy) are compared with the method of the present invention.

[0136] The number of frames refers to the number of video frames of the input data during training. The average joint error refers to the average Euclidean distance between the predicted 3D joint coordinates and the true coordinates. The joint error (ground truth) uses the ground truth of the 2D coordinates rather than the values predicted by the 2D detector as the input, and is the average Euclidean distance between the finally predicted 3D coordinates and the true values. The lower the value of the joint error, the better. When the number of frames is uniformly set to 243, compared with other methods in Table 1, the average joint error and the joint error (ground truth) in the method proposed by the present invention both achieve the minimum values, with improvements of 1.2 mm and 1.1 mm respectively. This shows that the pose predicted by the method proposed by the present invention is closer to the real situation, achieving the best model performance.

[0137] Table 2 Performance comparison results of different 3D human pose estimation methods on MPI-INF-3DHP

[0138]

[0139] In Table 2, VideoPose3D (Video Pose 3D), PoseFormer (Pose Transformer), MHFormer (Multi-Hypothesis Transformer), MixSTE (Mixed Spatio-Temporal Encoder), P-STMO (Pre-trained Spatial Temporal Many-to-One), D3DP (Diffusion-based 3D Pose estimation), GLA-GCN (Global-local Adaptive Graph Convolutional Network), PoseFormerV2 (Pose Transformer Version 2), STCFormer (Spatio-Temporal Criss-cross attention Transformer), and MotionAGFormer (Motion Attention-GCNFormer) are compared with the method of the present invention.

[0140] The correct joint percentage represents the accuracy of the predicted key points within a certain threshold range, and the area under the curve is the area of the correct joint percentage curve, which is used to measure the overall performance of the model at different thresholds. The larger the values of these two metrics, the higher the overall accuracy. Compared with other methods in Table 2, the method proposed in the present invention achieves the highest correct key point percentage and area under the curve, while having the smallest average joint error, which is improved by 0.9 mm compared to the best method. This indicates that the method proposed in the present invention achieves better performance than other methods when trained using the MPI-INF-3DHP dataset.

[0141] An embodiment of the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. It should be noted that when the processor executes the computer program, it corresponds to the specific steps of the method provided by the embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.

[0142] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. It should be noted that when the computer program is run by a processor, it corresponds to the specific steps of the method provided by the embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiment of the present invention.

[0143] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A three-dimensional human posture estimation method based on an enhanced topology-aware network, characterized in that: include: S1, obtain human motion capture dataset; S2: Construct an enhanced topology-aware network model, and train the model using the data set in step S1 to obtain the final enhanced topology-aware network model; the specific contents are: The enhanced topology-aware network model consists of sequentially connected feature embedding blocks, an enhanced topology-aware module and a regression head that are repeatedly stacked 5 times; Among them, the enhanced topology perception module includes a spatiotemporal dual-branch Transformer and a hybrid constraint module; The data set in step S1 is divided into a training set and a test set according to a ratio of 7:3, the training set is used to train the enhanced topology-aware network model, and the test set is used to test the trained enhanced topology-aware network model to obtain a final enhanced topology-aware network model; S3: Input the human body image or video to be detected into the final enhanced topology perception network model to obtain the three-dimensional coordinates corresponding to each joint and complete the three-dimensional human body posture estimation; the specific contents are as follows: Use the 2D posture detector to obtain the 2D coordinates P of the joint 2D ∈R T×N×3 , where T and N represent the number of frames and joints in the sequence respectively, and the number 3 represents the dimension, which includes the horizontal and vertical coordinates and confidence scores of the joints; the two-dimensional coordinates are input into the final enhanced topology-aware network model, and after the feature embedding block, the two-dimensional coordinates are projected to high dimensions to obtain the preliminary high-dimensional features P∈R T×N×C , where C represents the dimension size; Add a set of tensors, add them to the preliminary high-dimensional features and input them into the enhanced topology perception module. Use the spatiotemporal dual-branch Transformer to calculate the spatiotemporal global dependencies between joints to obtain the fused intermediate features. The specific content is: The spatiotemporal dual-branch Transformer includes Transformer blocks stacked in a spatial-temporal order and Transformer blocks stacked in a temporal-spatial order; wherein the Transformer blocks stacked in a spatial-temporal order include a spatial encoder and a temporal encoder connected in sequence, and the Transformer blocks stacked in a temporal-spatial order include a temporal encoder and a spatial encoder connected in sequence; The added tensor and preliminary high-dimensional features are passed through the spatiotemporal dual-branch Transformer to obtain the corresponding intermediate features. The specific expression is: P1 = TTE(STE(P)); P2 = STE(TTE(P)); Among them, P1 represents the intermediate features obtained by stacking Transformer blocks in space-time order, P2 represents the intermediate features obtained by stacking Transformer blocks in time-space order, TTE represents the time encoder, and STE represents the space encoder; Adaptively fuse P1 and P2 to obtain the intermediate features after fusion. The specific expression is: W = FC(Concat(P1, P2)); F = W1·P1+W2·P2; Among them, W represents the tensor after dimension conversion, FC represents the linear layer, Concat represents the concatenation operation, F represents the intermediate feature after fusion, and W1 and W2 both represent weights; According to the degrees of freedom of the joints and the limb categories they belong to, the hybrid constraint module is used to obtain the local topological constraints of different joints respectively, and the final hybrid topological constraints are obtained through adaptive fusion. The constraints are used to structurally guide the fused intermediate features to complete the operation of the enhanced topological perception module. The operation in the enhanced topology perception module is repeated 5 times to obtain the enhanced features of the human topology structure. After the regression head, the linear layer is used to predict the final 3D posture coordinates P 3D ∈R T×N×3 .

2. The method for 3D human posture estimation based on enhanced topology-aware network according to claim 1, characterized in that: In step S1, the two-dimensional coordinates, three-dimensional coordinates and their true values ​​of the joints are obtained from the Human3.6M and MPI-INF-3DHP large-scale motion capture datasets.

3. The method for 3D human posture estimation based on enhanced topology-aware network according to claim 1, characterized in that: The shapes of the tensors are N×C and T×1×C respectively and are initialized to 0.

4. The method for 3D human posture estimation based on enhanced topology-aware network according to claim 1, characterized in that: The joints are grouped according to their degrees of freedom. The specific expression is: DoF1={right_shoulder,left_shoulder,right_hip,left_hip}; DoF2={right_elbow,left_elbow,right_knee,left_knee}; DoF3={right_wrist,left_wrist,right_feet,left_feet}; Among them, DoF1, DoF2, and DoF3 all represent the degree of freedom grouping of joints, right_shoulder represents the right shoulder, left_shoulder represents the left shoulder, right_hip represents the right hip, left_hip represents the left hip, right_elbow represents the right elbow, left_elbow represents the left elbow, right_knee represents the right knee, left_knee represents the left knee, right_wrist represents the right wrist, left_wrist represents the left wrist, right_feet represents the right foot, and left_feet represents the left foot; The joints are grouped according to the limb category they belong to. The specific expression is: Part1={right_shoulder,right_elbow,right_wrist}; Part2={left_shoulder,left_elbow,left_wrist}; Part3={right_hip,right_knee,right_feet}; Part4={left_hip,left_knee,left_feet}; Among them, Part1, Part2, Part3, and Part4 represent the right arm, left arm, right leg, and left leg of the human body respectively; The expression for static joint grouping is: Static={head,neck,thorax,spine,hip}; Among them, Static represents static joint grouping, head represents head, neck represents neck, thorax represents chest, spine represents spine, and hip represents hip joint.

5. The method for 3D human posture estimation based on enhanced topology-aware network according to claim 1, characterized in that: The fused intermediate features are grouped according to the degrees of freedom of the joints and the limb categories to which they belong, and the degrees of freedom grouping features and the limb category grouping features are obtained; Perform feature dimension conversion on the grouping features to obtain the converted grouping features F S ,in represents the i-th degree of freedom grouping feature, represents the grouping feature of the jth limb category, F S Represents static joint grouping features; Concatenate the features of each degree of freedom group along the joint dimension to obtain the overall feature F of the degree of freedom group D , after a two-dimensional convolution layer Conv2d with a convolution kernel size of 4×3 D Perform feature extraction to obtain the corresponding aggregation features Will Split along the joint dimension to obtain the local topological constraints of the i-th degree of freedom grouping feature The specific formula is: Among them, σ represents the GELU activation function, Split represents the segmentation feature along the joint dimension, represents the first degree of freedom grouping feature, represents the second degree of freedom grouping feature, represents the third degree of freedom grouping feature, and Concat represents the concatenation operation; Concatenate the features of each limb category group along the joint dimension to obtain the overall feature F of the limb category group P , after a two-dimensional convolution Conv2d with a convolution kernel size of 3×3 p Perform feature extraction to obtain the corresponding aggregation features Will Split along the joint dimension to obtain the local topological constraints of the j-th limb category grouping feature The specific formula is: in, Indicates the first limb category grouping feature, Indicates the second limb category grouping feature, Indicates the third limb category grouping feature, Indicates the fourth limb category grouping feature; The static joint grouping features are processed by a two-dimensional convolution Conv2d with a convolution kernel size of 5×3. s Perform feature extraction to obtain static joint grouping aggregation features Then we get the local topological constraint of the kth static joint grouping feature: The specific formula is: The local topological constraints of the grouping features of the limb category are combined with the local topological constraints of the grouping features of different degrees of freedom, and based on the weight parameters, the final hybrid topological constraints are obtained. The specific expression is: Y = (F + R) + F; Among them, r i,j W represents the mixed feature of the i-th degree of freedom grouping feature in the j-th limb category grouping feature, i,j Represents r i,j The corresponding learning parameters are initialized to 0, R represents the final hybrid topology constraint, concat order represents the sequential concatenation operation, F is the intermediate feature after fusion, and Y represents the result after adding the hybrid topological constraint.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the three-dimensional human posture estimation method based on the enhanced topology-aware network described in any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the three-dimensional human posture estimation method based on an enhanced topology-aware network described in any one of claims 1 to 5 is executed.

Citation Information

Patent Citations

  • Attitude estimation method, related device and storage medium

    CN115273228A

  • Three-dimensional human body posture estimation method and system based on human body topology sensing network

    CN115908497A