3D human pose estimation method based on enhanced topology-aware network

WO2026199707A1PCT designated stage Publication Date: 2026-10-01NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/097329
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2025-05-27
Publication Date
2026-10-01

Smart Images

  • Figure CN2025097329_01102026_PF_FP_ABST
    Figure CN2025097329_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A 3D human pose estimation method based on an enhanced topology-aware network. The method comprises: acquiring a human motion capture dataset; constructing an enhanced topology-aware network model, wherein the model comprises a feature embedding block, an enhanced topology-aware module repeatedly stacked five times, and a regression head, which are connected in sequence, and the enhanced topology-aware module comprises a spatio-temporal dual-branch Transformer and mixed topological constraints; using the dataset to train the model to obtain a final enhanced topology-aware network model; and inputting a human picture or video that needs to be detected into the final enhanced topology-aware network model to obtain 3D coordinates corresponding to each joint, thereby completing 3D human pose estimation.
Need to check novelty before this filing date? Find Prior Art

Description

A 3D Human Pose Estimation Method Based on Enhanced Topology Sensing Network Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to a three-dimensional human pose estimation method based on an enhanced topological sensing network. Background Technology

[0002] 3D Human Pose Estimation (HPE), a core research topic in computer vision, aims to accurately recover the 3D joint positions of the human body from 2D images or videos. In recent years, thanks to the development of deep learning algorithms, the task of 3D human pose estimation has progressed rapidly. The human body's topology reflects the spatial and kinematic relationships between joints and is an essential feature of human pose; accurately modeling this structure is crucial for improving the accuracy and robustness of the task. Transformer-based models have made significant progress in 3D human pose estimation. These methods utilize multi-head self-attention mechanisms to model the global dependencies of joints, showing advantages in capturing long-distance spatiotemporal correlations between joints. However, self-attention mechanisms overemphasize global relationships, neglecting the local topological connections and structural constraints of the human skeleton. This limitation makes it difficult for the model to fully utilize prior knowledge of the human skeletal structure, especially under complex poses. Therefore, how to fully consider the human body's topology while modeling the global spatiotemporal correlations of joints, strengthen the modeling of topological dependencies between human joints, and generate more accurate 3D poses has become a major challenge. Summary of the Invention

[0003] The technical problem to be solved by this invention is to provide a three-dimensional human pose estimation method based on an enhanced topology-aware network, which solves the problem of insufficient modeling of human topology.

[0004] To solve the above technical problems, the present invention adopts the following technical solution:

[0005] A 3D human pose estimation method based on an enhanced topology-aware network includes the following steps:

[0006] S1. Obtain the human motion capture dataset.

[0007] S2. Construct an enhanced topology-aware network model. Train the model using the dataset from step S1 to obtain the final enhanced topology-aware network model.

[0008] S3. Input the human body image or video to be detected into the final enhanced topology sensing network model to obtain the three-dimensional coordinates of each joint and complete the estimation of the three-dimensional human body posture.

[0009] Furthermore, in step S1, the two-dimensional coordinates, three-dimensional coordinates, and their ground truth values ​​of the joints are obtained from the Human3.6M and MPI-INF-3DHP large motion capture datasets.

[0010] Furthermore, in step S2, the enhanced topology-aware network model includes a feature embedding block connected in sequence, an enhanced topology-aware module repeatedly stacked 5 times, and a regression head.

[0011] The enhanced topology awareness module includes a spatiotemporal dual-branch Transformer and a hybrid constraint module.

[0012] The dataset in step S1 is divided into a training set and a test set in a 7:3 ratio. The enhanced topology-aware network model is trained using the training set, and the trained enhanced topology-aware network model is tested using the test set to obtain the final enhanced topology-aware network model.

[0013] Furthermore, in step S3, the estimation of the three-dimensional human pose includes the following:

[0014] The two-dimensional coordinates P of the joint are obtained using a two-dimensional attitude detector. 2D ∈R T×N×3 Where T and N represent the number of frames and joints in the sequence, respectively, and the number 3 represents the dimension, which includes the horizontal and vertical coordinates of the joints and their confidence scores. This two-dimensional coordinate is input into the final enhanced topology-aware network model. After passing through a feature embedding block, the two-dimensional coordinate is projected into a higher dimension to obtain the preliminary high-dimensional feature P∈R. T×N×C Where C represents the dimension size; add a set of tensors, add them to the initial high-dimensional features and input them into the enhanced topology perception module. Use the spatiotemporal bi-branch Transformer to calculate the spatiotemporal global dependencies between joints to obtain the fused intermediate features. According to the degrees of freedom of the joints and the limb category they belong to, use the hybrid constraint module to obtain the local topological constraints of different joints respectively. Obtain the final hybrid topological constraints through adaptive fusion. Use this constraint to provide structured guidance for the fused intermediate features to complete the operation of the enhanced topology perception module.

[0015] The operation performed in the enhanced topology perception module was repeated 5 times to obtain enhanced features of the human body's topology. These features were then processed by a regression head and predicted using a linear layer to obtain the final 3D pose coordinates P. 3D ∈R T×N×3 .

[0016] Furthermore, the tensors have shapes of N×C and T×1×C and are initialized to 0.

[0017] Furthermore, the intermediate features obtained after fusion include the following:

[0018] The spatiotemporal dual-branch Transformer includes Transformer blocks stacked in a spatial-temporal order and Transformer blocks stacked in a temporal-spatial order; wherein, the Transformer blocks stacked in a spatial-temporal order include a spatial encoder and a temporal encoder connected in sequence, and the Transformer blocks stacked in a temporal-spatial order include a temporal encoder and a spatial encoder connected in sequence.

[0019] The tensor after addition and the initial high-dimensional features are passed through a spatiotemporal dual-branch Transformer to obtain the corresponding intermediate features, specifically expressed as: P1 = TTE(STE(P)); P2 = STE(TTE(P));

[0020] Where P1 represents the intermediate feature obtained by stacking Transformer blocks in spatial-temporal order, P2 represents the intermediate feature obtained by stacking Transformer blocks in temporal-spatial order, TTE represents the temporal encoder, and STE represents the spatial encoder.

[0021] Adaptive fusion is performed on P1 and P2 to obtain the fused intermediate features. The specific expression is: W = FC(Concat(P1,P2)); F = W1·P1 + W2·P2;

[0022] Where W represents the tensor after dimensionality transformation, FC represents the linear layer, Concat represents the concatenation operation, F represents the intermediate features after fusion, and W1 and W2 both represent weights.

[0023] Furthermore, the joints are grouped according to their degrees of freedom, with the following specific expressions: DoF1 = {right_shoulder, left_shoulder, right_hip, left_hip}; DoF2 = {right_elbow, left_elbow, right_knee, left_knee}; DoF3 = {right_wrist, left_wrist, right_feet, left_feet};

[0024] In this context, DoF1, DoF2, and DoF3 all represent the degrees of freedom grouping of the joints, right_shoulder represents the right shoulder, left_shoulder represents the left shoulder, right_hip represents the right hip, left_hip represents the left hip, right_elbow represents the right elbow, left_elbow represents the left elbow, right_knee represents the right knee, left_knee represents the left knee, right_wrist represents the right wrist, left_wrist represents the left wrist, right_feet represents the right foot, and left_feet represents the left foot.

[0025] Joints are grouped according to their limb category, as follows: Part1 = {right_shoulder, right_elbow, right_wrist}; Part2 = {left_shoulder, left_elbow, left_wrist}; Part3 = {right_hip, right_knee, right_feet}; Part4 = {left_hip, left_knee, left_feet};

[0026] Part 1, Part 2, Part 3, and Part 4 represent the right arm, left arm, right leg, and left leg of the human body, respectively.

[0027] The expression for static joint grouping is: Static = {head,neck,thorax,spine,hip};

[0028] Among them, Static represents static joint grouping, head represents head, neck represents neck, thorax represents chest, spine represents spine, and hip represents hip joint.

[0029] Furthermore, the fused intermediate features are grouped according to the degrees of freedom of the joints and the limb category to obtain the degree-of-freedom grouped features and the limb category grouped features.

[0030] Perform feature dimension transformation on the grouped features to obtain the transformed grouped features. F S ,in This represents the grouping feature of the i-th degree of freedom. F represents the grouping feature of the j-th limb category. S This indicates the characteristics of static joint grouping.

[0031] The features of each degree of freedom group are concatenated along the joint dimension to obtain the overall feature F of the degree of freedom group. DAfter passing through a 2D convolutional layer Conv2d with a kernel size of 4×3 D Feature extraction is performed to obtain aggregated features that include groupings for each degree of freedom. Will By splitting along the joint dimension, the local topological constraints of the i-th degree of freedom group feature are obtained. The specific formula is as follows:

[0032] Where σ represents the GELU activation function, and Split represents the feature segmentation along the joint dimension. This represents the grouping feature of the first degree of freedom. This represents the grouping feature of the second degree of freedom. This represents the grouping feature of the third degree of freedom.

[0033] The features of each limb category group are concatenated along the joint dimension to obtain the overall feature F of the limb category group. P After passing through a 2D convolution Conv2d with a kernel size of 3×3 p Feature extraction is performed to obtain aggregated features that include each limb category group. Will By splitting along the joint dimension, we obtain the local topological constraints of the grouping features of the j-th limb category. The specific formula is as follows:

[0034] in, This indicates the grouping feature of the first limb category. This indicates the grouping feature of the second limb category. This indicates the grouping feature of the third limb category. This indicates the grouping feature of the 4th limb category.

[0035] Static joint grouping features are processed by a 2D convolution Conv2d with a kernel size of 5×3. s Feature extraction is performed to obtain static joint grouping and aggregation features. This leads to the local topological constraints of the k-th static joint grouping features. The specific formula is as follows:

[0036] The local topological constraints of the grouping features belonging to the limb category are combined with the local topological constraints of the grouping features with different degrees of freedom, and based on the weight parameters, the final hybrid topological constraint is obtained, the specific expression of which is: Y = (F + R) + F;

[0037] Where, r i,jW represents the mixed feature of the i-th degree of freedom grouping feature in the j-th limb category grouping feature. i,j Indicates the relationship with r i,j The corresponding learnable parameters are initialized to 0, R represents the final hybrid topology constraint, and concat order Y represents the sequential splicing operation, and Y represents the result after adding hybrid topology constraints.

[0038] Furthermore, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the three-dimensional human pose estimation method based on an enhanced topology sensing network.

[0039] Furthermore, the present invention also proposes a computer-readable storage medium storing a computer program, which is executed by a processor to perform the described three-dimensional human pose estimation method based on an enhanced topology sensing network.

[0040] Compared with existing technologies, the present invention, employing the above technical solution, has the following technical effects: By designing an enhanced topology-aware network structure and utilizing a hybrid topology constraint module to enhance the network model's learning of human topology, the present invention solves the problem of insufficient modeling of human topology and generates more accurate 3D pose coordinates. Furthermore, the feasibility and effectiveness of the proposed method have been verified on commonly used Human3.6M and MPI-INF-3DHP datasets. Attached Figure Description

[0041] Figure 1 is a flowchart of the overall implementation of the present invention.

[0042] Figure 2 is a schematic diagram of the joint grouping method in an embodiment of the present invention.

[0043] Figure 3 is an operation diagram of the hybrid topology constraint module in an embodiment of the present invention.

[0044] Figure 4 is a visualization of the results of an embodiment of the present invention. Detailed Implementation

[0045] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0046] To achieve the above objectives, this invention proposes a three-dimensional human pose estimation method based on an enhanced topological sensing network, as shown in Figure 1. The specific steps are as follows:

[0047] S1. Obtain the 2D coordinates, 3D coordinates, and ground truth values ​​of the joints from the Human3.6M and MPI-INF-3DHP large motion capture datasets.

[0048] S2. Construct the ETANet model and train it using the dataset from step S1 to obtain the final enhanced topology-aware network model. The specific steps are as follows:

[0049] To build an ETANet (Enhanced Topology-Aware Network) model using the PyTorch deep learning framework, the process begins with environment configuration and data loading. The dataset is then processed using its built-in DataLoader method, dividing it into sequences of a specified number of frames (e.g., 243 frames). Next, the network structure is built, and the forward propagation process and loss function are defined according to the specifications to optimize the model.

[0050] The enhanced topology-aware network model consists of sequentially connected feature embedding blocks, enhanced topology-aware modules stacked five times, and a regression head.

[0051] The enhanced topology awareness module includes a spatiotemporal dual-branch Transformer and a human skeleton-based topology-based MTC (Mixed Topological Constraints) module. MTC is simple, efficient, flexible, and easy to use; it's a plug-and-play module that significantly improves the accuracy and robustness of 3D pose estimation by combining with different Transformer models. MTC aims to learn the topological priors of the human skeleton, constraining the feature representations of different joints to guide the model in generating more reasonable 3D poses, further reducing errors and improving model performance. The core of MTC lies in grouping and mixing components according to different logics to obtain stronger topological constraints.

[0052] The dataset from step S1 is divided into a training set and a test set in a 7:3 ratio. The enhanced topology-aware network model is trained using the training set, and the trained enhanced topology-aware network model is tested using the test set. During the testing phase, the parameters in the network model are not updated; this phase is only used to evaluate the performance of the network model. After the training and testing phases, the final enhanced topology-aware network model is obtained.

[0053] In this embodiment, participant data numbered 1, 5, 6, 7, and 8 were used for training, and participant data numbered 9 and 11 were used for testing. NVIDIA RTX 4090 GPU was used to accelerate model training.

[0054] For model training, the batch size was set to 4, the initial learning rate was set to 0.0005, and the model was trained for 100 epochs, with the learning rate decaying by 0.99 after each epoch. The AdamW optimizer was selected to optimize the model parameters, and the weight decay was specified to be 0.01.

[0055] S3. Input the human body image or video to be detected into the final enhanced topology-aware network model to obtain the 3D coordinates of each joint, thus completing the estimation of the 3D human body pose. The specific content is as follows:

[0056] The two-dimensional coordinates P of the joint are obtained using a two-dimensional attitude detector. 2D ∈R T×N×3 Where T and N represent the number of frames and joints in the sequence, respectively, and the number 3 represents the dimension, which includes the horizontal and vertical coordinates of the joints and their confidence scores. This two-dimensional coordinate is input into the final enhanced topology-aware network model. After passing through a feature embedding block, the two-dimensional coordinate is projected into a higher dimension to obtain the preliminary high-dimensional feature P∈R. T×N×256 .

[0057] To address the model's lack of ability to process the order of elements in a sequence, positional embedding is needed to introduce the positional information of each element within the sequence. A set of tensors with shapes N×C and T×1×C, initialized to 0, are added for the model to learn. These tensors are then added to the initial high-dimensional features and input into the enhanced topology-aware module. A spatiotemporal dual-branch Transformer is used to calculate the spatiotemporal global dependencies between joints, yielding the fused intermediate features. The specific details are as follows:

[0058] The spatiotemporal dual-branch Transformer includes Transformer blocks stacked in a spatial-temporal order and Transformer blocks stacked in a temporal-spatial order; wherein, the Transformer blocks stacked in a spatial-temporal order include a spatial encoder and a temporal encoder connected in sequence, and the Transformer blocks stacked in a temporal-spatial order include a temporal encoder and a spatial encoder connected in sequence.

[0059] The two Transformer blocks take into account both the correlation between different joints within the same frame and the long-distance dependency between the same joints across frames.

[0060] The tensor after addition and the initial high-dimensional features are passed through a spatiotemporal dual-branch Transformer to obtain the corresponding intermediate features, specifically expressed as: P1 = TTE(STE(P)); P2 = STE(TTE(P));

[0061] Where P1 represents the intermediate feature obtained by stacking Transformer blocks in spatial-temporal order, P2 represents the intermediate feature obtained by stacking Transformer blocks in temporal-spatial order, TTE represents the temporal encoder, and STE represents the spatial encoder.

[0062] Adaptive fusion of P1 and P2 is performed to fully integrate temporal and spatial features while keeping the feature dimensions unchanged, thereby enhancing the model's ability to model the global and local dependencies of 3D pose and obtaining the fused intermediate features. The specific expressions are: W = FC(Concat(P1,P2)); F = W1·P1 + W2·P2;

[0063] Where W represents the tensor after dimensional transformation; FC represents the linear layer; Concat represents the concatenation operation; F represents the fused intermediate features; W1 and W2 both represent weights, which are taken from the two values ​​of the last dimension of W respectively.

[0064] The degree of freedom of a joint is defined as the rotational degree of freedom of a joint relative to its parent node. Because the joints at the ends of the limbs have a large range of motion and a high degree of freedom, the pose estimation error is also higher, thus requiring more refined modeling.

[0065] To align with the grouping of limb categories and differentiate it from the grouping of static joints, we only consider 12 joints, including the limbs. These 12 joints are grouped according to their degrees of freedom, as shown in Figure 2(a). The specific expressions are: DoF1 = {right_shoulder, left_shoulder, right_hip, left_hip}; DoF2 = {right_elbow, left_elbow, right_knee, left_knee}; DoF3 = {right_wrist, left_wrist, right_feet, left_feet}.

[0066] In this context, DoF1, DoF2, and DoF3 all represent the degrees of freedom grouping of the joints, right_shoulder represents the right shoulder, left_shoulder represents the left shoulder, right_hip represents the right hip, left_hip represents the left hip, right_elbow represents the right elbow, left_elbow represents the left elbow, right_knee represents the right knee, left_knee represents the left knee, right_wrist represents the right wrist, left_wrist represents the left wrist, right_feet represents the right foot, and left_feet represents the left foot.

[0067] Each group contains 4 joints with the same degree of freedom. The larger the group index, the higher the degree of freedom (DoF) of the joints within the group.

[0068] Based on the fundamental principles of human kinematics and anatomy, the limbs are grouped for 3D human pose estimation to accommodate the movement characteristics of different limbs. Similarly, to complement DoF grouping and to differentiate static joint groups, the 12 joints are grouped according to their respective limb categories, as shown in Figure 2(b). The specific expressions are: Part1 = {right_shoulder, right_elbow, right_wrist}; Part2 = {left_shoulder, left_elbow, left_wrist}; Part3 = {right_hip, right_knee, right_feet}; Part4 = {left_hip, left_knee, left_feet}.

[0069] Part 1, Part 2, Part 3, and Part 4 represent the four limbs of the human body: the right arm, left arm, right leg, and left leg, respectively.

[0070] Because spine-related joints are relatively rigid, they are typically used to represent the core structure and key nodes of the trunk, while limb joints are more flexible. To avoid interfering with the mixing of degrees of freedom and limb joint groups, this invention considers spine-related joints as a separate static joint group. This helps reduce motion conflicts between different joints, making spinal motion more physiologically accurate, thus avoiding unnatural posture predictions and improving the model's accuracy and stability. The expression for the static joint group is: Static = {head,neck,thorax,spine,hip};

[0071] Among them, Static represents static joint grouping, head represents head, neck represents neck, thorax represents chest, spine represents spine, and hip represents hip joint.

[0072] At this point, various grouped joint groups have been obtained. As shown in Figure 3, based on the joint's degrees of freedom and its limb category, the hybrid constraint module obtains the local topological constraints for different joints. Through adaptive fusion, the final hybrid topological constraint is obtained. This constraint is then used to structurally guide the fused intermediate features, completing the operation of the enhanced topology awareness module. Specifically:

[0073] The fused intermediate features are grouped according to the degrees of freedom of the joints and the limb category to obtain the degree-of-freedom grouped features and the limb category grouped features.

[0074] Perform feature dimension transformation on the grouped features to obtain the transformed grouped features. F S ,in This represents the grouping feature of the i-th degree of freedom. F represents the grouping feature of the j-th limb category. S Indicates the characteristics of static joint grouping. F S There is no subscript because there is only one joint group under this grouping method.

[0075] Group features for each degree of freedom By splicing along the joint dimensions, we obtain the overall feature F of the degree-of-freedom grouping. D ∈R C×12×T The convolutional layer Conv2d with a kernel size of 4×3, padding set to (0,1), stride set to (4,1), and maintaining the same number of feature channels is passed through it. D Feature extraction is performed so that while the convolution extracts features for each group, it also captures the local temporal dependencies between three adjacent frames. This complements the global temporal dependencies captured by the Transformer, resulting in more comprehensive temporal correlation information and aggregated features containing each degree of freedom group. Will By splitting along the joint dimension, the local topological constraints of the i-th degree of freedom group feature are obtained. The specific formula is as follows:

[0076] Where σ represents the GELU activation function, and Split represents the feature segmentation along the joint dimension. This represents the grouping feature of the first degree of freedom. This represents the grouping feature of the second degree of freedom. This represents the grouping feature of the third degree of freedom.

[0077] The features of each limb category group are concatenated along the joint dimension to obtain the overall feature F of the limb category group. P The convolution is performed using a 3×3 kernel, padding set to (0,1), stride set to (3,1), and while maintaining the number of feature channels. p Feature extraction is performed to obtain aggregated features that include each limb category group. Will By splitting along the joint dimension, we obtain the local topological constraints of the grouping features of the j-th limb category. The specific formula is as follows:

[0078] in, This indicates the grouping feature of the first limb category. This indicates the grouping feature of the second limb category. This indicates the grouping feature of the third limb category. This indicates the grouping feature of the 4th limb category.

[0079] Since there is only one static joint group, the feature concatenation process is omitted. The static joint group features are processed by a 2D Conv2d convolution with a kernel size of 5×3, padding set to (0,1), stride set to 1, and the number of feature channels remaining unchanged. s Feature extraction is performed to obtain static joint grouping and aggregation features. This leads to the local topological constraints of the k-th static joint grouping features. The specific formula is as follows:

[0080] After grouping the limbs, the three joints in each group correspond to three different degrees of freedom. Therefore, the local topological constraints of the limb category grouping features are combined with the local topological constraints of the different degrees of freedom grouping features. Based on the weight parameters, which are initialized to 0, a more refined and flexible final hybrid topological constraint is obtained. This guides the model to more comprehensively model the topological relationships of the human skeleton and predict more reasonable 3D poses. The specific expression is as follows: Y = (F + R) + F;

[0081] Where, r i,j W represents the mixed feature of the i-th degree of freedom grouping feature in the j-th limb category grouping feature. i,j Indicates the relationship with r i,j The corresponding learnable parameters are initialized to 0, R represents the final hybrid topology constraint, and concat order Y represents the sequential splicing operation, and Y represents the result after adding hybrid topology constraints.

[0082] Topological constraints, which combine joint features that contain both joint degrees of freedom and limb category information, are added to the input. Residual connections are then applied to enhance the model's learning of human topology.

[0083] The operation performed in the enhanced topology perception module was repeated 5 times to obtain enhanced features of the human body's topology. These features were then processed by a regression head and predicted using a linear layer to obtain the final 3D pose coordinates P. 3D ∈R T×N×3 .

[0084] Figure 4 is a comparison of the visualization results of the method of this invention and the MixSTE method. Figure 4(a) shows the input 2D human pose diagram, Figure 4(b) shows the 3D human pose diagram generated by the MixSTE method, Figure 4(c) shows the 3D human pose diagram generated by the method of this invention, and Figure 4(d) shows the true 3D human pose diagram corresponding to the input 2D pose. It can be seen that the MixSTE method does not accurately reproduce the human pose with crossed legs, and the left shoulder joint is too high, resulting in poor prediction accuracy. The method of this invention reproduces details in the left and right leg joints and the left shoulder joint, and the prediction of the head joint is closer to the true value, resulting in more accurate results compared to MixSTE. Therefore, the method of this invention can achieve better results in the task of 3D human pose estimation, generating more accurate 3D poses.

[0085] The performance of different 3D human pose estimation methods is compared, as shown in Tables 1 and 2.

[0086] Table 1. Performance comparison of different 3D human pose estimation methods on Human 3.6M.

[0087] Table 1 lists the following network operators: UGCN (U-shaped Graph Convolutional Network), PoseFormer, StridedFormer, MHFormer (Multi-Hypothesis Transformer), MixSTE (Mixed Spatio-Temporal Encoder), P-STMO (Pre-trained Spatial Temporal Many-to-One), GLA-GCN (Global-local Adaptive Graph Convolutional Network), POT (Pose-oriented transformer), STCFormer (Spatio-Temporal Criss-cross attention Transformer), KTPFormer (Kinematics and Trajectory Prior Knowledge-Enhanced Transformer), and NC-RetNet (Non-Causal...). This invention compares RetNet (a non-causal preserving network) and TMT (Two-step Mixed-Training strategy) with the method of this invention.

[0088] Frame count refers to the number of video frames input during training. Average joint error refers to the average Euclidean distance between the predicted 3D joint coordinates and the ground truth coordinates. The true joint error is the average Euclidean distance between the predicted 3D coordinates and the ground truth values, using the ground truth values ​​of 2D coordinates instead of the values ​​predicted by the 2D detector as input. Lower joint error values ​​are better. When the frame count is uniformly set to 243, compared with other methods in Table 1, the method proposed in this invention achieves the minimum values ​​for both average joint error and true joint error, with improvements of 1.2 mm and 1.1 mm, respectively. This indicates that the pose predicted by the method proposed in this invention is closer to the real situation, achieving optimal model performance.

[0089] Table 2. Performance comparison results of different 3D human pose estimation methods on MPI-INF-3DHP

[0090] Table 2 compares VideoPose3D, PoseFormer (Pose Transformer), MHFormer (Multi-Hypothesis Transformer), MixSTE (Mixed Spatio-Temporal Encoder), P-STMO (Pre-trained Spatial Temporal Many-to-One), D3DP (Diffusion-based 3D Pose estimation), GLA-GCN (Global-local Adaptive Graph Convolutional Network), PoseFormerV2 (Pose Transformer Version 2), STCFormer (Spatio-Temporal Criss-cross attention Transformer), and MotionAGFormer (Motion Attention-GCNFormer) with the method of this invention.

[0091] The percentage of correct joints represents the accuracy of predicting keypoints within a certain threshold range, and the area under the curve (AUC) is the area of ​​the percentage of correct joints curve, used to measure the overall performance of the model at different thresholds. Higher values ​​for both metrics indicate higher overall accuracy. Compared to other methods in Table 2, the method proposed in this invention achieves the highest percentage of correct keypoints and AUC, while having the lowest average joint error, improving upon the best method by 0.9 mm. This demonstrates that, when trained using the MPI-INF-3DHP dataset, the method proposed in this invention achieves better performance than other methods.

[0092] This invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. It should be noted that when the processor executes the computer program, it corresponds to the specific steps of the method provided in this invention, possessing the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in this invention.

[0093] This invention also proposes a computer-readable storage medium storing a computer program. It should be noted that when the computer program is executed by a processor, it corresponds to the specific steps of the method provided in this invention, possessing the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in this invention.

[0094] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any modifications or equivalent changes made based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A three-dimensional human pose estimation method based on an enhanced topological sensing network, characterized in that, include: S1. Obtain the human motion capture dataset; S2. Construct an enhanced topology-aware network model. Train the model using the dataset from step S1 to obtain the final enhanced topology-aware network model. The specific steps are as follows: The enhanced topology-aware network model consists of sequentially connected feature embedding blocks, enhanced topology-aware modules stacked five times, and a regression head; The enhanced topology sensing module includes a spatiotemporal dual-branch Transformer and a hybrid constraint module. The dataset in step S1 is divided into a training set and a test set in a 7:3 ratio. The enhanced topology-aware network model is trained using the training set, and the trained enhanced topology-aware network model is tested using the test set to obtain the final enhanced topology-aware network model. S3. Input the human body image or video to be detected into the final enhanced topology-aware network model to obtain the 3D coordinates of each joint, thus completing the estimation of the 3D human body pose; the specific content is as follows: The two-dimensional coordinates P of the joint are obtained using a two-dimensional attitude detector. 2D ∈R T×N×3 Where T and N represent the number of frames and joints in the sequence, respectively, and the number 3 represents the dimension, which includes the horizontal and vertical coordinates of the joints and their confidence scores. This two-dimensional coordinate is input into the final enhanced topology-aware network model. After passing through a feature embedding block, the two-dimensional coordinate is projected into a higher dimension to obtain the preliminary high-dimensional feature P∈R. T×N×C , where C represents the dimension size; A set of tensors is added, summed with the initial high-dimensional features, and then input into the enhanced topology-aware module. A spatiotemporal dual-branch Transformer is used to calculate the spatiotemporal global dependencies between joints, yielding the fused intermediate features. The specific content is as follows: The spatiotemporal dual-branch Transformer includes Transformer blocks stacked in a spatial-temporal order and Transformer blocks stacked in a temporal-spatial order; wherein, the Transformer blocks stacked in a spatial-temporal order include a spatial encoder and a temporal encoder connected in sequence, and the Transformer blocks stacked in a temporal-spatial order include a temporal encoder and a spatial encoder connected in sequence. The summed tensor and the initial high-dimensional features are processed by a spatiotemporal dual-branch Transformer to obtain the corresponding intermediate features, the specific expression of which is: P1 = TTE(STE(P)); P2 = STE(TTE(P)); Where P1 represents the intermediate feature obtained by stacking Transformer blocks in spatial-temporal order, P2 represents the intermediate feature obtained by stacking Transformer blocks in temporal-spatial order, TTE represents the temporal encoder, and STE represents the spatial encoder. Adaptive fusion of P1 and P2 is performed to obtain the fused intermediate features, the specific expression of which is: W = FC(Concat(P1,P2)); F = W1·P1 + W2·P2; Where W represents the tensor after dimensionality transformation, FC represents the linear layer, Concat represents the concatenation operation, F represents the intermediate features after fusion, and W1 and W2 both represent weights; Based on the joint's degrees of freedom and the limb category it belongs to, the local topological constraints of different joints are obtained using the hybrid constraint module. The final hybrid topological constraints are obtained through adaptive fusion. The constraints are then used to guide the structured intermediate features after fusion, thus completing the operation of the enhanced topological perception module. The operation performed in the enhanced topology perception module was repeated 5 times to obtain enhanced features of the human body's topology. These features were then processed by a regression head and predicted using a linear layer to obtain the final 3D pose coordinates P. 3D ∈R T×N×3 .

2. The three-dimensional human pose estimation method based on enhanced topology sensing network according to claim 1, characterized in that, In step S1, the two-dimensional coordinates, three-dimensional coordinates, and their ground truth values ​​of the joints are obtained from the Human3.6M and MPI-INF-3DHP large motion capture datasets.

3. The three-dimensional human pose estimation method based on enhanced topology sensing network according to claim 1, characterized in that, The tensors have shapes of N×C and T×1×C and are initialized to 0.

4. The three-dimensional human pose estimation method based on enhanced topology sensing network according to claim 1, characterized in that, Joints are grouped according to their degrees of freedom, specifically as follows: DoF1 = {right_shoulder, left_shoulder, right_hip, left_hip}; DoF2 = {right_elbow, left_elbow, right_knee, left_knee}; DoF3 = {right_wrist, left_wrist, right_feet, left_feet}; In this context, DoF1, DoF2, and DoF3 all represent the degrees of freedom grouping of a joint, right_shoulder represents the right shoulder, left_shoulder represents the left shoulder, right_hip represents the right hip, left_hip represents the left hip, right_elbow represents the right elbow, left_elbow represents the left elbow, right_knee represents the right knee, left_knee represents the left knee, right_wrist represents the right wrist, left_wrist represents the left wrist, right_feet represents the right foot, and left_feet represents the left foot. The joints are grouped according to their respective limb categories, as shown in the following expression: Part1={right_shoulder,right_elbow,right_wrist}; Part2={left_shoulder,left_elbow,left_wrist}; Part3={right_hip,right_knee,right_feet}; Part4={left_hip,left_knee,left_feet}; Part 1, Part 2, Part 3, and Part 4 represent the right arm, left arm, right leg, and left leg of the human body, respectively. The expression for static joint grouping is: Static={head,neck,thorax,spine,hip}; Among them, Static represents static joint grouping, head represents head, neck represents neck, thorax represents chest, spine represents spine, and hip represents hip joint.

5. The three-dimensional human pose estimation method based on enhanced topology sensing network according to claim 1, characterized in that, The fused intermediate features are grouped according to the degrees of freedom of the joints and the limb category to obtain the degree-of-freedom grouped features and the limb category grouped features; Perform feature dimension transformation on the grouped features to obtain the transformed grouped features. F S ,in This represents the grouping feature of the i-th degree of freedom. F represents the grouping feature of the j-th limb category. S Indicates static joint grouping characteristics; The features of each degree of freedom group are concatenated along the joint dimension to obtain the overall feature F of the degree of freedom group. D After passing through a 2D convolutional layer Conv2d with a kernel size of 4×3 D Feature extraction is performed to obtain the corresponding aggregated features. Will By splitting along the joint dimension, the local topological constraints of the i-th degree of freedom group feature are obtained. The specific formula is as follows: Where σ represents the GELU activation function, and Split represents the feature segmentation along the joint dimension. This represents the grouping feature of the first degree of freedom. This represents the grouping feature of the second degree of freedom. This indicates the grouping feature of the third degree of freedom, and Concat represents the concatenation operation. The features of each limb category group are concatenated along the joint dimension to obtain the overall feature F of the limb category group. P After passing through a 2D convolution Conv2d with a kernel size of 3×3 p Feature extraction is performed to obtain the corresponding aggregated features. Will By splitting along the joint dimension, we obtain the local topological constraints of the grouping features of the j-th limb category. The specific formula is as follows: in, This indicates the grouping feature of the first limb category. This indicates the grouping feature of the second limb category. This indicates the grouping feature of the third limb category. This indicates the grouping feature of the 4th limb category; Static joint grouping features are processed by a 2D convolution Conv2d with a kernel size of 5×3. s Feature extraction is performed to obtain static joint grouping aggregation features. This leads to the local topological constraints of the k-th static joint grouping features. The specific formula is as follows: The local topological constraints of the grouping features belonging to the limb category are combined with the local topological constraints of the grouping features with different degrees of freedom, and based on the weight parameters, the final hybrid topological constraint is obtained, the specific expression of which is: Y = (F + R) + F; Where, r i,j W represents the mixed feature of the i-th degree of freedom grouping feature in the j-th limb category grouping feature. i,j Indicates the relationship with r i,j The corresponding learning parameters are initialized to 0, R represents the final hybrid topology constraint, and concat order This indicates a sequential splicing operation, where F represents the intermediate feature after fusion, and Y represents the result after adding hybrid topological constraints.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the three-dimensional human pose estimation method based on an enhanced topology sensing network as described in any one of claims 1 to 5.

7. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the three-dimensional human pose estimation method based on the enhanced topology sensing network as described in any one of claims 1 to 5.