A Symmetric Semantic Graph Convolutional Pose Estimation Method Based on Body Part Grouping

The symmetric semantic graph convolutional network with body part grouping enhances 3D pose estimation by addressing interference and improving generalization, achieving superior performance on the Human3.6M dataset.

CN115546888BActive Publication Date: 2025-07-15SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211084071.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-06
Publication Date
2025-07-15
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

The existing 3D pose estimation methods are easily affected by self-occlusion and environmental occlusion in the context of deep learning, and have poor generalization capabilities. The existing network models are not universal in human pose estimation.

Method used

The symmetric semantic graph convolutional pose estimation method based on body part grouping is adopted. By constructing a symmetric semantic graph convolutional layer and non-local layer of the joint node graph structure, combining body part grouping, a symmetric semantic graph convolutional pose estimation network model is constructed, and training is used using the Human3.6M dataset, symmetric semantic graph convolution and non-local layer are introduced to aggregate neighbor node information and enhance local feature learning.

Benefits of technology

It improves the accuracy and generalization ability of 3D pose estimation, can better handle occlusion and diversity of human postures, and improves the pose estimation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546888B_ABST
    Figure CN115546888B_ABST
Patent Text Reader

Abstract

The present invention discloses a symmetric semantic graph convolutional pose estimation method based on body part grouping, which comprises the following steps: S1, input two-dimensional human joint points and their connection relationships, and construct a symmetric semantic graph convolutional layer and a non-local layer of a joint point graph structure; S2, according to the body trunk, perform body part grouping to respectively obtain local and non-local features of each trunk and local and non-local features of the whole body, and perform fusion calculation on the obtained features; S3, based on the symmetric semantic graph convolutional layer, the non-local layer and the body part grouping, construct a symmetric semantic graph convolutional pose estimation network model for body part grouping; S4, use the Human3.6M dataset to train the symmetric semantic graph convolutional pose estimation network model, input the two-dimensional human joint points to be estimated into the trained symmetric semantic graph convolutional pose estimation network model, and output the estimated three-dimensional human joint points. The present invention can be applied to fields such as movie animation, virtual reality, sports action analysis, etc., and has better method effects and improved generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a symmetric semantic graph convolution pose estimation method based on body part grouping. Background Art

[0002] Human pose estimation has been widely applied in many computer vision tasks, such as virtual reality, human-computer interaction, and behavior recognition. Thanks to the rapid development of deep learning, the performance of estimating three-dimensional human poses from images has been significantly improved, becoming a current research hotspot.

[0003] There are two categories of existing 3D pose estimation methods. One is to directly predict 3D poses from images, and the other is to first predict 2D poses and then regress 3D poses. The first category of methods can obtain a large amount of information from images, but the model is greatly affected by factors such as image background and human clothing, and the learning features required by the model are complex. The second category of methods reduces the overall work complexity, and the network model is more likely to learn the mapping from 2D to 3D space. At the same time, due to the maturity of 2D pose estimation research, this type of model is more mainstream.

[0004] The three-dimensional human pose estimation method based on graph convolution network provided in "A Three-Dimensional Human Pose Estimation Method Based on Graph Convolution Network" (CN112712019A) has the advantages of improving the performance of three-dimensional human pose regression and reducing the use of network parameters, but the generalization ability of the model needs to be improved. Existing research shows that in the context of deep learning, human pose estimation algorithms are easily affected by self-occlusion, environmental occlusion, etc., and human poses are diverse, and the current model generalization ability is poor. Therefore, it is urgent to explore a more reasonable and more general network model to improve the pose estimation effect. Summary of the Invention

[0005] The purpose of the present invention is to solve the above problems in the prior art and provide a symmetric semantic graph convolution pose estimation method based on body part grouping.

[0006] The purpose of the present invention can be achieved by adopting the following technical solutions:

[0007] A symmetric semantic graph convolution pose estimation method based on body part grouping, as Figure 1 shown, the symmetric semantic graph convolution pose estimation method includes the following steps:

[0008] S1. Input the two-dimensional human joint points and their connection relationships in movie animations, virtual reality, or sports actions, and construct a symmetric semantic graph convolution layer and a non-local layer of the joint point graph structure;

[0009] S2. Group body parts according to the body trunk to obtain the local and non-local features of each trunk and the local and non-local features of the whole body, and perform fusion calculation on the obtained features;

[0010] S3. Based on the symmetric semantic graph convolutional layer, non-local layer, and body part grouping, construct a symmetric semantic graph convolutional pose estimation network model for body part grouping;

[0011] S4. Use the H u m a n3.6M dataset to train the symmetric semantic graph convolutional pose estimation network model, and input the two-dimensional human joint points to be estimated into the trained symmetric semantic graph convolutional pose estimation network model to output the estimated three-dimensional human joint points.

[0012] Furthermore, the process of constructing the symmetric semantic graph convolutional layer and non-local layer of the joint graph structure using the two-dimensional human joint points and their connection relationships in step S1 is as follows:

[0013] Let X (l) and X (l+1) represent the features of the nodes in the graph structure before and after the l-th layer of convolution respectively. Then the form of the symmetric graph convolution is:

[0014] X (l+1) =σ(WX (l) A sym ) (1)

[0015] where σ() represents the activation function, W represents the learnable weight parameter, and A sym is the matrix obtained by symmetric normalization of the adjacency matrix A of the graph, which is expressed as follows:

[0016]

[0017] where A is the adjacency matrix of the graph and D is the degree matrix. Symmetric normalization can better aggregate the information of neighboring nodes to obtain balanced node features;

[0018] By adding a learnable weighted matrix M on the basis of the symmetric graph convolution, the symmetric semantic graph convolutional layer is constructed. The calculation formula of the symmetric semantic graph convolutional layer is expressed as follows:

[0019] X (l+1) =σ(WX (l) ρ i (M⊙A sym )) (3)

[0020] where ρ i () is the Softmax non-linear function used to normalize the matrix of node i, and ⊙ represents the element-wise multiplication operation of the matrix;

[0021] To capture the global features between nodes in the graph, the concept of a non-local layer is introduced, and the operation of the non-local layer is defined as:

[0022]

[0023] where, W x represents the normalization factor of the learnable weight parameter W, K represents the number of nodes, i represents the index of the target node to be calculated, and j represents the index of the nodes other than i; represent the input features of nodes i and j respectively; represents the output feature of node i; f(,) is a learnable binary function used to calculate the similarity of two input features; g() is a learnable unary function that transforms the input feature.

[0024] Furthermore, in the step S2 for body part grouping, the human joint points are decomposed into a left limb group, a right limb group, and a whole body group. The joint points within each group have stronger correlations, and each group extracts features through an independent sub-network to enhance local relationships.

[0025] As Figure 4 shown, feature fusion adopts the late fusion method. First, the features in each group are learned, and then the features in each group are fused. The feature fusion is defined as:

[0026] f fuse = Concat(f left , f right , f all ) (5)

[0027] where, Concat(,,) represents the operation of concatenating features, f left is the feature of the left limb group, f right is the feature of the right limb group, f all is the feature of the whole body group, and f fuse is the feature obtained after fusion.

[0028] The implementation of body part grouping can learn the consistency of local joints while ensuring global pose consistency, and can better generalize to symmetric poses, as well as rare and occluded poses in the training data.

[0029] Furthermore, in the step S3, multiple symmetric semantic graph convolutional modules are constructed based on the symmetric semantic graph convolutional layer and the non-local layer. All symmetric semantic graph convolutional modules have the same structure. Each symmetric semantic graph convolutional module is composed of two symmetric semantic graph convolutional layers and a non-local layer connected in sequence. The local and global semantic relationships between nodes are obtained by alternating the symmetric semantic graph convolutional layer and the non-local layer;

[0030] In the symmetric semantic graph convolutional network, as Figure 3 shown, first use a symmetric semantic graph convolutional layer and a non-local layer to map the input to the latent space; then, through four sequentially connected symmetric semantic graph convolutional modules, the encoded features are obtained. After all the symmetric semantic graph convolutional layers in the symmetric semantic graph convolutional network, batch normalization and R e LU non-linear activation are performed;

[0031] The symmetric semantic graph convolutional pose estimation network model for body part grouping includes a first branch, a second branch, and a third branch, as Figure 2 shown, where the first branch, the second branch, and the third branch all use the symmetric semantic graph convolutional network for feature extraction: the left limb group is input to the first branch, and through the symmetric semantic graph convolutional network, the feature f left of the left limb is extracted; the right limb group is input to the second branch, and through the symmetric semantic graph convolutional network, the feature f right of the right limb is extracted; the whole body group is input to the third branch, and through the symmetric semantic graph convolutional network, the feature f all of the whole body is extracted; the fused feature f fuse is calculated according to formula (5), and then a symmetric semantic graph convolutional layer is used to project the encoded features into the output space.

[0032] Furthermore, the loss function L smoothl1 () defined by formula (6) is adopted in step S4 and trained on the Human3.6M dataset. The formula is as follows:

[0033]

[0034]

[0035] where X represents the difference between the true value and the predicted value, |·| represents the absolute value of the difference between the true value and the predicted value, J′ i represents the 3D joint coordinates of the predicted i-th node, and J i corresponds to the true value of the i-th node in the dataset. The L smoothl1 (J) loss function is insensitive to outliers and can control the magnitude of the gradient, enabling reasonable convergence during training.

[0036] Furthermore, the evaluation index commonly used in pose estimation is MPJPE (Mean Per Joint Position Error), and the formula is defined as (7):

[0037]

[0038] E MPJPE() The metric represents the mean of the L2 distances between the predicted values and the ground truth for each joint, and ||·||2 represents the L2 distance from the predicted value to the ground truth. When the evaluation metric MPJPE is small, the 3D human pose estimation result is considered to be better.

[0039] Furthermore, during the training process, the initial learning rate is 0.001, and batch processing with a size of 64 is used. The initial learning rate directly affects the convergence state of the model, while the batch size affects the generalization ability of the model. Using an initial learning rate of 0.001 is beneficial for model convergence, and using a batch size of 64 is beneficial for model generalization.

[0040] The present invention has the following advantages and effects compared with the prior art:

[0041] The symmetric semantic graph convolutional pose estimation network based on body part grouping proposed by the present invention introduces symmetric semantic graph convolution, which can better aggregate the information of neighboring nodes and obtain balanced node features; a body part grouping is designed, and the body is segmented into left / right torso according to parts, and these body part groups are learned through independent sub-networks to enhance local features. Compared with other methods on the H u m a n3.6M dataset, generally speaking, the effect of this method is better and the generalization ability is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0043] Figure 1 is a flowchart of a three-dimensional human pose transfer method driven by pose estimation disclosed by the present invention;

[0044] Figure 2 is a diagram of a symmetric semantic graph convolutional network model based on body part grouping in an embodiment of the present invention;

[0045] Figure 3 is a schematic diagram of a symmetric semantic graph convolutional module in an embodiment of the present invention;

[0046] Figure 4 is a schematic diagram of a body part grouping feature fusion module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0048] Example 1

[0049] A symmetric semantic graph convolutional pose estimation method based on body part grouping, as Figure 1 shown, the method includes the following steps:

[0050] S1. Input the two-dimensional human body joint points and their connection relationships in movie animations, virtual reality, or sports actions, and construct a symmetric semantic graph convolutional layer and a non-local layer of the joint point graph structure;

[0051] In step S1, the process of using the two-dimensional human body joint points and their connection relationships to construct a symmetric semantic graph convolutional layer and a non-local layer of the joint point graph structure is as follows:

[0052] Let X (l) and X (l+1) represent the features of the nodes in the graph structure before and after the l-th layer of convolution respectively. Then the form of the symmetric graph convolution is:

[0053] X (l+1) = σ(WX (l) A sym ) (1)

[0054] where σ() represents the activation function, W represents the learnable weight parameter, and A sym is the matrix obtained by symmetrically normalizing the adjacency matrix A of the graph, which is expressed as follows:

[0055]

[0056] where A is the adjacency matrix of the graph and D is the degree matrix;

[0057] By adding a learnable weighted matrix M on the basis of the symmetric graph convolution, a symmetric semantic graph convolutional layer is constructed, and the calculation formula of the symmetric semantic graph convolutional layer is expressed as follows:

[0058] X (l+1) = σ(WX (l) ρ i (M⊙A sym) ) (3)

[0059] where ρ i () is the Softmax non-linear function used to normalize the matrix of node i, and ⊙ represents the element-wise multiplication operation of the matrix;

[0060] To capture the global features between the nodes in the graph, the concept of the non-local layer is introduced, and the operation of the non-local layer is defined as:

[0061]

[0062] where Wx The normalization factor representing the learnable weight parameter W, K represents the number of nodes, i represents the index of the target node to be calculated, and j represents the index of the nodes other than i; respectively represent the input features of nodes i and j; represents the output feature of node i; f(,) is a learnable binary function for calculating the similarity of two input features; g() is a learnable unary function for transforming the input feature.

[0063] S2. According to the body trunk data in movie animations, virtual reality or motion actions, group the body parts, respectively obtain the local and non-local features of each trunk and the local and non-local features of the whole body, and perform fusion calculation on the obtained features;

[0064] In step S2 for body part grouping, the human joint points in movie animations, virtual reality or motion actions are decomposed into a left limb group, a right limb group, and a whole body group. The joint points within each group have stronger correlations, and each group extracts features through an independent sub-network to enhance the local relationship.

[0065] As Figure 4 shown, the feature fusion adopts the late fusion method. First, learn the features in each group, and then perform fusion calculation on the features in each group. The feature fusion is defined as:

[0066] f fuse = Concat(f left , f right , f all ) (5)

[0067] Among them, Concat(,,) represents the operation of connecting the features, f left is the feature of the left limb group, f right is the feature of the right limb group, f all is the feature of the whole body group, and f fuse is the feature obtained after fusion.

[0068] S3. Based on the symmetric semantic graph convolutional layer, non-local layer, and body part grouping, construct a symmetric semantic graph convolutional pose estimation network model for body part grouping;

[0069] In step S3, based on the symmetric semantic graph convolutional layer and non-local layer, construct multiple symmetric semantic graph convolutional modules. All symmetric semantic graph convolutional modules have the same structure, and each symmetric semantic graph convolutional module is composed of two symmetric semantic graph convolutional layers and a non-local layer connected in sequence;

[0070] In the symmetric semantic graph convolutional network, as Figure 3As shown, first use a symmetric semantic graph convolutional layer and a non-local layer to map the input into the latent space; then, through four symmetric semantic graph convolutional modules connected in sequence, obtain the encoded features. After all symmetric semantic graph convolutional layers in the symmetric semantic graph convolutional network, batch normalization and ReLU non-linear activation are performed.

[0071] The symmetric semantic graph convolutional pose estimation network model for body part grouping includes a first branch, a second branch, and a third branch. As Figure 2 shown, among them, the first branch, the second branch, and the third branch all use a symmetric semantic graph convolutional network for feature extraction: the left limb group is input into the first branch, and through the symmetric semantic graph convolutional network, the feature f of the left limb is extracted. left ; the right limb group is input into the second branch, and through the symmetric semantic graph convolutional network, the feature f of the right limb is extracted. right ; the whole body group is input into the third branch, and through the symmetric semantic graph convolutional network, the feature f of the whole body is extracted. all ; the fused feature f is calculated according to formula (5). fuse Then, use a symmetric semantic graph convolutional layer to project the encoded features into the output space.

[0072] S4. Use the Human3.6M dataset to train the symmetric semantic graph convolutional pose estimation network model. Input the two-dimensional human joint points in the movie animation, virtual reality, or motion action to be estimated into the trained symmetric semantic graph convolutional pose estimation network model, and output the three-dimensional human joint points in the estimated movie animation, virtual reality, or motion action.

[0073] In step S4, the loss function L smoothl1 () defined by formula (6) is used for training on the Human3.6M dataset. The formula is as follows:

[0074]

[0075]

[0076] Among them, X represents the difference between the ground truth and the predicted value of the movie animation, virtual reality, or motion action data, |·| represents the absolute value of the difference between the ground truth and the predicted value, and J′ i represents the 3D joint coordinates of the predicted i-th node, and J i corresponds to the ground truth of the i-th node in the dataset.

[0077] Among them, the evaluation index usually used for pose estimation is MPJPE (Mean Per Joint Position Error), and the formula is defined as (7):

[0078]

[0079] E MPJPE () The metric represents the mean of the L2 distances between the predicted values and the true values of each joint in the movie animation, virtual reality, or motion actions. ||·||2 represents the L2 distance from the predicted value to the true value. When the evaluation metric MPJPE is small, the 3D human pose estimation result is considered to be better.

[0080] During the training process, the initial learning rate is 0.001, and the batch size is 64.

[0081] Example 2

[0082] This example is based on a symmetric semantic graph convolutional pose estimation method based on body part grouping disclosed in Example 1. To verify the effectiveness of the present invention, experiments are carried out on the Human3.6M dataset, and the technical effects of the present invention are described in combination with the experimental results.

[0083] Human3.6M is one of the most widely used datasets for 3D pose estimation, covering 3.6 million images, collected in an indoor controlled environment. There are 11 experimenters in total, and a marked motion capture device is used to collect the body postures of the experimenters during their daily activities, including 15 actions.

[0084] Experimental configuration: Hardware environment: GPU RTX 2080Ti, video memory: 11GB, CPU 4-core Intel(R) Xeon(R) Silver 4110 CPU @ 2.10GHz, memory: 16GB. Software environment: Python v2.7, Pytorch v1.1.0, CUDA10.2. Operating system: Ubuntu18.04.

[0085] An ablation study is carried out on the method proposed by the present invention. With the above configuration, the pose estimation network proposed by the present invention includes two main modules: a symmetric semantic graph convolutional module and body part grouping. To verify their effectiveness, the following ablation experiments are set up: the first experiment only uses semantic graph convolution, the second experiment uses the symmetric semantic graph convolutional module, the second experiment uses body part grouping, and the third experiment uses the symmetric semantic graph convolutional module and body part grouping. The experimental results are shown in Table 1:

[0086] Table 1. Ablation experiment results table of the symmetric semantic graph convolutional pose estimation method based on body part grouping

[0087] Symmetric Semantic Graph Convolution Module Body Part Grouping MPJPE 41.47mm √ 40.68mm √ 40.53mm √ √ 39.93mm

[0088] Table 2 shows the comparative experiment results of the method of the present invention with the baseline method and the semantic graph convolution method under the MPJPE evaluation metric according to human action classification. The best method for each action is highlighted in bold.

[0089] Table 2. Comparative experimental results table of the symmetric semantic graph convolutional pose estimation method based on body part grouping

[0090]

[0091] It can be seen that the symmetric semantic graph convolutional pose estimation network based on body part grouping proposed by the present invention achieves better performance, which indicates that the model in this paper can effectively utilize the relationship between different joint groups in the graph.

[0092] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A symmetric semantic graph convolutional pose estimation method based on body part grouping, characterized in that, The symmetric semantic graph convolutional pose estimation method includes the following steps: S1. Input the two-dimensional human body joint points and their connection relationships in movie animations, virtual reality, or motion actions, and construct a symmetric semantic graph convolutional layer and a non-local layer of the joint point graph structure; S2. According to the body trunk, perform body part grouping to obtain the local and non-local features of each trunk and the local and non-local features of the whole body respectively, and perform fusion calculation on the obtained features; in step S2, the human body joint points are decomposed into a left limb group, a right limb group, and a whole body group. Each group enhances the local relationship through an independent sub-network, and then adopts a late fusion feature fusion method. First, learn the features in each group, and then fuse the features in each group. The feature fusion is defined as: f fuse = Concat(f left , f right , f all ) (5) Among them, Concat(,,) represents the operation of concatenating features, f left is the feature of the left limb group, f right is the feature of the right limb group, f all is the feature of the whole body group, f fuse is the feature obtained after fusion; S3. Based on the symmetric semantic graph convolutional layer, non-local layer, and body part grouping, construct a symmetric semantic graph convolutional pose estimation network model for body part grouping; in step S3, based on the symmetric semantic graph convolutional layer and non-local layer, construct multiple symmetric semantic graph convolutional modules. All symmetric semantic graph convolutional modules have the same structure. Each symmetric semantic graph convolutional module is composed of two symmetric semantic graph convolutional layers and a non-local layer connected in sequence; In the symmetric semantic graph convolutional network, first use a symmetric semantic graph convolutional layer and a non-local layer to map the input to the latent space; then, through four symmetric semantic graph convolutional modules connected in sequence, obtain the encoded features. After all symmetric semantic graph convolutional layers in the symmetric semantic graph convolutional network, batch normalization and ReLU non-linear activation are performed; The symmetric semantic graph convolutional pose estimation network model for the grouped body parts includes a first branch, a second branch, and a third branch. Among them, the first branch, the second branch, and the third branch all use a symmetric semantic graph convolutional network for feature extraction: the left limb group is input into the first branch, and through the symmetric semantic graph convolutional network, the feature f of the left limb is extracted left ; the right limb group is input into the second branch, and through the symmetric semantic graph convolutional network, the feature f of the right limb is extracted right ; the whole body group is input into the third branch, and through the symmetric semantic graph convolutional network, the feature f of the whole body is extracted all ; the fused feature f is calculated according to formula (5) fuse , and then a symmetric semantic graph convolutional layer is used to project the encoded features into the output space; S4. Use the Human3.6M dataset to train the symmetric semantic graph convolutional pose estimation network model, input the two-dimensional human body joint points to be estimated into the trained symmetric semantic graph convolutional pose estimation network model, and output the estimated three-dimensional human body joint points.

2. The symmetric semantic graph convolutional pose estimation method based on body part grouping according to claim 1, wherein The process of using the two-dimensional human body joint points and their connection relationships to construct the symmetric semantic graph convolutional layer of the joint point graph structure in step S1 is as follows: Let X (l) and X (l+1) represent the features of the nodes in the graph structure before and after the l-th layer of convolution respectively. Then the form of the symmetric graph convolution is as follows: X (l+1) = σ(WX (l) A sym ) (1) where, σ() represents the activation function, W represents the learnable weight parameter, and A sym is the matrix obtained by symmetrically normalizing the adjacency matrix A of the graph, which is expressed as follows: A sym = D -1 / 2 AD -1 / 2 (2) Among them, A is the adjacency matrix of the graph, and D is the degree matrix; A symmetric semantic graph convolutional layer is constructed by adding a learnable weighted matrix M on the basis of symmetric graph convolution. The calculation formula of the symmetric semantic graph convolutional layer is expressed as follows: X (l+1) = σ(WX (l) ρ i (M☉A sym )) (3) where ρ i () is the Softmax non - linear function, which is used to normalize the matrix of node i, and ⊙ represents the element - by - element multiplication operation of the matrix.

3. The symmetric semantic graph convolutional pose estimation method based on body part grouping according to claim 2, wherein The process of using the two-dimensional human body joint points and their connection relationships to construct the non-local layer of the joint point graph structure in step S1 is as follows: The operation of the non-local layer is defined as: Among them, W x represents the normalization factor of the learnable weight parameter W, K represents the number of nodes, i represents the index of the target node to be calculated, and j represents the index of the nodes other than i; respectively represent the input features of nodes i and j; represents the output feature of node i; f(,) is a learnable binary function used to calculate the similarity of two input features; g() is a learnable unary function used to transform the input features.

4. A symmetric semantic graph convolutional pose estimation method based on body part grouping according to claim 3, characterized in that, The loss function L defined by formula (6) is adopted in the step S4 smoothl1 (), and training is performed on the Human3.6M dataset. The formula is as follows: where X represents the difference between the true value and the predicted value, |·| represents the absolute value of the difference between the true value and the predicted value, and J′ i represents the 3D joint coordinates of the predicted i-th node, and J i corresponds to the true value of the i-th node in the dataset.

5. The symmetric semantic graph convolutional pose estimation method based on body part grouping according to claim 4, characterized in that The evaluation index adopted for the pose estimation is MPJPE, and the formula is defined as follows: E MPJPE () The metric represents the mean of the L2 distances between the predicted values and the ground truth for each joint, and ‖·‖2 represents the L2 distance from the predicted value to the ground truth.

6. The symmetric semantic graph convolutional pose estimation method based on body part grouping according to claim 4, wherein During the training process, the initial learning rate is 0.001, and a batch size of 64 is used.

Citation Information

Patent Citations

  • Three-dimensional human body posture estimation method based on graph convolution network

    CN112712019A

  • Two-dimensional to three-dimensional human body posture estimation method

    CN112232106A

  • Methods and systems for generating end-to-end model to estimate 3-dimensional(3-d) pose of object

    US20220222852A1