A 3D human pose estimation method and system based on the fusion of angular graphs and image features

By constructing the angle map feature comparison fusion module, the edge features are extracted using Gaussian Laplace operator and graph neural network, and combining convolutional neural networks to fusion multi-scale angle maps and image features, the depth fuzzy problem in 3D human pose estimation is solved, and the estimation accuracy and model expression ability are improved.

CN119888432BActive Publication Date: 2025-07-18ZHEJIANG UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510374245.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

In the existing 3D human posture estimation method, the 2D to 3D conversion has a problem of depth fuzziness, which is difficult to effectively solve in the existing technology, resulting in insufficient regression accuracy.

Method used

By constructing an angle map feature comparison fusion module, the Gaussian Laplace operator is used to extract edge features, combined with graph neural networks and convolutional neural networks, the multi-scale angle map features and image features are fused, and angle constraints are used to reduce depth blur and improve the accuracy of 3D pose estimation.

Benefits of technology

It effectively reduces the accuracy loss caused by depth blur, improves the accuracy of 3D human posture estimation, enhances the model's expression ability, and is upgradeable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888432B_ABST
    Figure CN119888432B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D human pose estimation method and system based on the fusion of angle maps and image features. The method includes: obtaining the 2D key point coordinates of a human body using a trained 2D human pose detector, and obtaining an original image with obvious edge features using the Laplacian of Gaussian operator; converting the topological map of the human body skeleton into two different angle maps; constructing an angle map feature contrast fusion module to fuse the angle map features of different scales through cross-contrast learning; fusing the image features with the angle map features, and finally regressing to obtain the 3D human pose information; repeating the training to obtain the final 3D human pose estimation model. The present invention introduces joint angle constraints into the graph network, reduces the influence brought by depth ambiguity, and at the same time fuses the image features, making the network have better expressive ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of 3D human pose estimation, and particularly relates to a 3D human pose estimation method and system based on the fusion of angle maps and image features. Background Art

[0002] Currently existing 3D human pose estimation methods can mainly be divided into one-stage methods and two-stage methods. One-stage methods directly regress 3D poses without a 2D estimator. In contrast, two-stage methods are based on 2D joint positions and concatenate 2D and 3D estimators. Thanks to the outstanding work in the field of 2D pose estimation in recent years, more and more researchers have devoted their energy to two-stage methods. The key to two-stage methods lies in the dimension elevation from 2D to 3D. However, due to the inherent depth ambiguity problem in the regression from 2D poses to 3D poses, that is, the same 2D pose can be obtained from multiple different 3D poses, there is still a large room for improvement in the regression accuracy of two-stage methods.

[0003] To reduce the accuracy impact caused by depth ambiguity, researchers have made many attempts. For example, the invention patent with the patent number CN202110546337.2 established an inverse projection network, which achieved motion constraints on different joints by grouping joint points, thereby reducing depth ambiguity to a certain extent; another example is the invention patent with the patent number CN202210142459.X, which also utilized a similar grouping idea to capture local and global features of human poses. However, the effect of such methods in overcoming the depth ambiguity problem is limited.

[0004] There are also methods that focus on the training data level. For example, the patent with the patent number CN202210460455.6 proposed a multi-view feature fusion method to obtain more robust depth features by extracting 2D pose expressions from multiple views. However, such methods face two problems: most existing datasets only provide single-view images, and there is little multi-view data for training, so their generalization is questionable in actual application scenarios; at the application level, providing information from multiple views often means higher costs. Therefore, the patent with the patent number CN202110569187.7 invented a method based on joint data augmentation and network training models, attempting to improve the quality of training data through data augmentation, but in essence, it still does not well solve the depth ambiguity problem. Summary of the Invention

[0005] The present invention aims to overcome the above-mentioned drawbacks of the prior art and proposes a 3D human pose estimation method and system based on the fusion of angle maps and image features to provide support for solving the problem of 2D to 3D human pose estimation.

[0006] To achieve the above object, the present invention provides the following solution: A 3D human pose estimation method based on the fusion of angle maps and image features, comprising the following steps:

[0007] S1: Construct an original dataset, obtain the 2D key point coordinates of the human body using a trained 2D human pose detector, and obtain an original image with obvious edge features using the Laplacian of Gaussian operator;

[0008] S2: Convert the original skeleton graph G with human joints as nodes into a line graph with bones as nodes;

[0009] S3: Convert the original skeleton graph G into a second-order graph that explicitly represents the angular relationship between joints;

[0010] S4: Construct an angle map feature contrast fusion module, and fuse angle map features of different scales through cross-contrast learning; The angle map feature contrast fusion module consists of a general graph neural network layer and an anti-pooling graph convolutional layer. The former is used to extract the features of the line graph and the second-order graph, and the latter makes the dimensions of the intermediate features of the line graph and the second-order graph consistent through anti-pooling operations, and finally completes feature fusion through learnable weight parameters;

[0011] S5: Train the angle map feature contrast fusion module to obtain an intermediate model: Input the processed image into a convolutional neural network model to obtain image features, fuse the image features with the angle map features, and finally regress to obtain 3D human pose information;

[0012] S6: Perform supervised fine-tuning on the intermediate model: Adjust the parameters of the regression head of the intermediate model, freeze the other parameters of the intermediate model, and obtain the final 3D human pose estimation model.

[0013] Preferably, step S1 specifically includes:

[0014] S1.1: The two-dimensional images in the Human3.6M dataset After passing through the cascaded pyramid network, obtain their normalized 16 2D key point coordinates , where ;

[0015] S1.2: Use the Laplacian of Gaussian operator to Perform edge sharpening processing to obtain , whose resolution is ;

[0016] Among them, the Laplacian of Gaussian operator is:

[0017] (1)

[0018] Among them, Represents the standard deviation, , respectively represent the coordinates of the pixel points, represents the natural logarithm.

[0019] Preferably, step S2 specifically includes:

[0020] S2.1: Convert the original skeleton graph with human joints as nodes into a line graph with human bones as nodes ;

[0021] wherein is represented as the node set and edge set of the original skeleton graph and the line graph , represented as and , the nodes of the line graph are the edges of the original skeleton graph, represented as , and the edges are represented as ; they are connected if and only if two edges in the original skeleton graph have a common node;

[0022] S2.2: Calculate the angle between two connected edges in the original skeleton graph as the edge weight of the line graph , and take the coordinates of the two end points of the edge of the original skeleton graph as the new input features of the line graph ; ;

[0023] wherein , the adjacency matrix of the line graph is represented as , where , can be obtained by the following formula:

[0024] (2)

[0025] where is a binary matrix, indicating whether the nodes in the node set of the original graph are the end points of the edges in the edge set , is represented as the identity matrix, and F(·) represents the function for calculating the joint angles in the original skeleton graph.

[0026] Preferably, the step S3 specifically includes:

[0027] S3.1: Convert the original skeleton graph into a second-order graph that directly represents the angular relationship between joint points;

[0028] wherein , representing the edges formed between nodes with a path length of 2 in the original skeleton graph.

[0029] S3.2: Calculate the second-order graph 's adjacency matrix and regenerate node features for the second-order graph ; ;

[0030] Among them, the adjacency matrix of the original skeleton graph is expressed as , where , , a value of 1 indicates that two nodes are connected, and 0 indicates that they are not connected; the adjacency matrix of the second-order graph representing the angle is expressed as:

[0031] (3)

[0032] The node features of the second-order graph , when and are second-order neighbors, is the angle value between and , when or and are not second-order neighbors, .

[0033] Preferably, the step S4 includes:

[0034] S4.1: Respectively perform feature extraction on the line graph and the second-order graph through ordinary graph neural network layers to generate intermediate features and , where is the feature dimension, and at the same time apply an ordinary graph neural network to the original graph to generate an intermediate feature ;

[0035] The ordinary graph neural network layer is expressed as:

[0036] (4)

[0037] The above formula is further described as:

[0038] (5)

[0039] (6)

[0040] Among them, denotes a learnable weight matrix, denotes a line graph adjacency matrix of, denotes an activation function, denotes a modulation matrix, denotes a Hadamard product, denotes the adjacency matrix obtained by symmetric normalization of the adjacency matrix of the second-order graph or the original graph ;

[0041] S4.2: Dimension conversion is performed on the intermediate features of the line graph through an anti-pooling graph neural network to obtain .

[0042] The anti-pooling graph neural network layer includes two separate graph convolutional networks, which are used to obtain an embedding matrix and an alignment matrix respectively:

[0043] (7)

[0044] (8)

[0045] Among them, denotes the embedding matrix, denotes the alignment matrix. Finally, the line graph obtains intermediate layer features with the same dimension as the second-order graph through the following formula:

[0046] (9)

[0047] Among them, denotes the intermediate features of the line graph after the anti-pooling operation.

[0048] S4.3: Cross-contrast learning is used to fuse multi-scale angular graph features. Specifically, feature learning is constrained by three different forms of loss functions.

[0049] Cross-contrast the original graph features with the line graph features to enhance the expression of local angular information and edge geometric properties. The calculation formula is:

[0050] (10)

[0051] Among them, denotes the cosine similarity, denotes the temperature parameter, , respectively denote the vectors of the nodes in the original graph and the line graph.

[0052] Cross-contrast the original graph features and second-order graph features , which is used to strengthen the capture of high-order angular relationships and global structure information. The calculation formula is:

[0053] (11)

[0054] where represents the eigenvector of node in the second-order graph.

[0055] The original graph features in the cross-comparison attribute module and the original graph features in the structure module , which is used to ensure the consistency of multi-view features. The calculation formula is:

[0056] (12)

[0057] Through the joint constraint of the above comparison losses, using the learnable weight parameter , the features of the three views are fused to generate intermediate fusion features . The fusion formula is:

[0058] (13)

[0059] where represents the concatenation operation along the feature dimension.

[0060] Preferably, step S5 includes:

[0061] S5.1: Training the angular graph feature comparison and fusion module to obtain an intermediate model, and optimizing the loss function through gradient descent . The overall comparison loss of the angular graph feature comparison and fusion module is:

[0062] (14)

[0063] where and represent hyperparameters used to balance each loss term.

[0064] S5.2: Inputting the processed image into a convolutional neural network model to obtain image features, and fusing the image features with the angular graph features. The convolutional neural network model is HRNet, and the image features of different scales in the four stages are fused with the angular graph features through a conversion module, where the conversion module uses a fully connected layer to reduce the two-dimensional image features to one-dimensional.

[0065] Preferably, in the intermediate model supervised fine-tuning process described in step S6, other parameters of the intermediate model are frozen, and the output head of the prediction module is fine-tuned using a weighted loss function. The specific formula is expressed as:

[0066] (15)

[0067] where represents the weighted sum of the mean squared error and the mean absolute error, and the weighting coefficient .

[0068] The present invention also provides a system for implementing the 3D human pose estimation method based on the fusion of the angle map and image features of the present invention, including:

[0069] A data preprocessing module, a graph network construction module, an angle map feature contrast and fusion module, an image feature fusion module, and an inference module;

[0070] The data preprocessing module obtains the human 2D key point coordinates from the original image and performs normalization processing. At the same time, the Laplacian of Gaussian operator is used to enhance the edge features of the original image;

[0071] The graph network construction module uses two different methods to construct a network using the known 2D pose information, so that the new graph network structure has the ability to represent joint angle information;

[0072] The angle feature contrast and fusion module makes the angle map features of different scales have consistent dimensions and performs feature fusion by cross-comparing the multi-scale angle map features;

[0073] The image feature fusion module uses a fully connected layer network to fuse two different features of the image and the graph at different stages of the model;

[0074] The inference module finally outputs the inference result of the model, that is, the 3D pose information of the human body;

[0075] The data preprocessing module, the graph network construction module, the angle map feature contrast and fusion module, the image feature fusion module, and the inference module are connected in sequence.

[0076] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the 3D human pose estimation method based on the fusion of the angle map and image features of the present invention is implemented.

[0077] The technical concept of the present invention is:

[0078] After obtaining the original image, first, the trained 2D pose estimator is used to detect the 2D joint coordinates of the human body, and the Laplacian of Gaussian operator is used for image processing to obtain an image with obvious edge features. Secondly, different angular maps are constructed using the known 2D joint coordinates, making full use of the joint angle information in the human body bone topology to overcome the depth blur problem. And by cross-comparing the features of multi-scale angular maps, the features of different-scale angular maps have consistent dimensions and are feature-fused. In addition, to ensure that the graph network does not lose the original information, a convolutional neural network model is used to extract image features and fuse them with the angular map features to improve the regression accuracy of the final 3D pose.

[0079] The beneficial effects of the present invention are as follows:

[0080] 1) The human body bone structure itself has topological properties, and the graph network can efficiently represent the features of human body joints and the edges between joints;

[0081] 2) Two different ways of constructing angular maps are proposed, introducing angular constraints into the graph network, thereby alleviating the accuracy loss caused by depth blur;

[0082] 3) The graph neural network and the convolutional neural network are used simultaneously in the same model. While using the graph neural network to extract the features of the human body bone structure, the convolutional neural network is used to extract its visual features;

[0083] 4) The backbone network used to extract image features in the present invention is not specific, and any convolutional network with the characteristics of a multi-stage model can be used in this method. Therefore, this method has scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 is the system framework diagram of the method of the present invention;

[0085] Figure 2 is the model structure diagram of the method of the present invention;

[0086] Figs. 3(a) to 3(f) are schematic diagrams of the local and global structures describing the angular relationship of the present invention. Among them, Fig. 3(a) is the angular schematic diagram at the center of the pelvis of the original skeleton diagram, Fig. 3(b) is the schematic diagram of converting the center of the pelvis of the original skeleton diagram into a second-order graph, in which the edges between the first-order neighbors are no longer retained (represented by dotted lines), and new edges are generated for the second-order neighbors (represented by solid lines), Fig. 3(c) is the mapping relationship diagram of each connecting edge at the center of the pelvis of the original skeleton diagram to the line graph structure, Fig. 3(d) is the original skeleton schematic diagram, Fig. 3(e) is the structural schematic diagram when the original skeleton diagram is reconstructed into a second-order graph, and Fig. 3(f) is the structural schematic diagram when the original skeleton is reconstructed into a line graph;

[0087] Figures 4(a) to 4(b) are schematic diagrams of the angle feature fusion model of the present invention. Among them, Figure 4(a) is the angle feature extraction module, and Figure 4(b) is the line graph intermediate feature alignment module. Detailed implementation manners

[0088] The following further describes in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings of the specification.

[0089] Example 1

[0090] Referring to Figure 1 ~ Figure 4(b), this embodiment relates to a 3D human pose estimation method based on the fusion of angle graphs and image features, and the steps are as follows:

[0091] S1: Construct an original data set, obtain the 2D key point coordinates of the human body using a trained 2D human pose detector, and obtain an original image with obvious edge features using the Laplacian of Gaussian operator;

[0092] S2: As shown in Figure 3(f), convert the original skeleton graph with human joints as nodes into a line graph with bones as nodes;

[0093] S3: As shown in Figure 3(e), convert the original skeleton graph into a second-order graph that explicitly represents the angular relationship between joints;

[0094] S4: Construct an angle graph feature contrast fusion module, and fuse angle graph features of different scales through cross-contrast learning; the angle graph feature contrast fusion module consists of a general graph neural network layer and an anti-pooling graph convolutional layer. The former is used to extract the features of the line graph and the second-order graph, and the latter makes the dimensions of the intermediate features of the line graph and the second-order graph consistent through anti-pooling operations, and finally completes feature fusion through learnable weight parameters;

[0095] S5: Train the angle graph feature contrast fusion module to obtain an intermediate model: input the processed image into a convolutional neural network model to obtain image features, fuse the image features with the angle graph features, and finally regress to obtain 3D human pose information;

[0096] S6: Perform supervised fine-tuning on the intermediate model: adjust the parameters of the regression head of the intermediate model, freeze the other parameters of the intermediate model, and obtain the final 3D human pose estimation model.

[0097] Step S1 specifically includes:

[0098] S1.1: As shown in Figure 3(d), pass the two-dimensional images in the Human3.6M data set through a cascaded pyramid network to obtain 16 normalized 2D key point coordinates , where ;

[0099] S1.2: Use the Laplacian of Gaussian operator to perform edge sharpening to obtain , whose resolution is ;

[0100] The Laplacian of Gaussian operator is:

[0101] (1)

[0102] where represents the standard deviation, , respectively represent the coordinates of the pixel point, represents the natural logarithm.

[0103] Step S2 specifically includes:

[0104] S2.1: Convert the original skeleton graph with human joints as nodes into a line graph with human bones as nodes ;

[0105] where represents the node set and edge set of the original skeleton graph and the line graph , denoted as and , as shown in Figure 3 (c), the nodes of the line graph are the edges of the original skeleton graph, denoted as , while the edges are denoted as ; they are connected if and only if two edges in the original skeleton graph have a common node;

[0106] S2.2: Calculate the angle between two connected edges in the original skeleton graph as the edge weight of the line graph , and take the coordinates of the two end points of the edge of the original skeleton graph as the node features of the line graph ;

[0107] where , the adjacency matrix of the line graph is denoted as , where , can be obtained by the following formula:

[0108] (2)

[0109] where is a binary matrix, representing the node set in the original graph Whether the nodes of are the endpoints of the connecting edges in is represented as the identity matrix, and F(·) is represented as a function for calculating the joint angles in the original skeleton graph.

[0110] Step S3 specifically includes:

[0111] S3.1: As shown in Fig. 3(b), convert the original skeleton graph into a second-order graph that directly represents the angular relationship between joint points ;

[0112] where , represents the connecting edges formed between nodes with a path length of 2 in the original skeleton graph.

[0113] S3.2: Calculate the adjacency matrix of the second-order graph and generate node features for the second-order graph ; ;

[0114] where the adjacency matrix of the original skeleton graph is represented as , where , , a value of 1 indicates that two nodes are connected, and 0 indicates that they are not connected; the adjacency matrix of the second-order graph representing angles is represented as:

[0115] (3)

[0116] As shown in Fig. 3(b), the node features of the second-order graph , when and are second-order neighbors, and is the angular value between or and are not second-order neighbors, .

[0117] Step S4 specifically includes:

[0118] S4.1: Respectively perform feature extraction on the line graph and the second-order graph through ordinary graph neural network layers to generate intermediate features and , where is the feature dimension, and at the same time apply an ordinary graph neural network to the original graph to generate an intermediate feature .

[0119] The ordinary graph neural network layer is expressed as:

[0120] (4)

[0121] The above formula is further described as:

[0122] (5)

[0123] (6)

[0124] Wherein, represents a learnable weight matrix, represents the line graph 's adjacency matrix, represents an activation function, represents a modulation matrix, represents the Hadamard product, represents the adjacency matrix obtained by symmetric normalization of the adjacency matrix of the second-order graph or the original graph ;

[0125] S4.2: Perform dimensional transformation on the intermediate features of the line graph through the anti-pooling graph neural network to obtain .

[0126] The anti-pooling graph neural network layer includes two separate graph convolutional networks, which are used to obtain the embedding matrix and the alignment matrix respectively:

[0127] (7)

[0128] (8)

[0129] Wherein, represents the embedding matrix, represents the alignment matrix. Finally, the line graph obtains the intermediate layer features with the same dimension as the second-order graph through the following formula:

[0130] (9)

[0131] Wherein, represents the intermediate features of the line graph after the anti-pooling operation.

[0132] S4.3: Achieve the fusion of multi-scale angular graph features through cross-contrast learning. Specifically, the feature learning is constrained by three different forms of loss functions.

[0133] Cross-contrast the original graph features With line graph features , which is used to enhance the expression of local angular information and edge geometric characteristics. The calculation formula is:

[0134] (10)

[0135] Among them, is expressed as cosine similarity, is expressed as temperature parameter, , respectively represent the vectors of the node in the original graph and the line graph.

[0136] Cross - compare the original graph features and second - order graph features , which is used to strengthen the capture of high - order angular relationships and global structure information. The calculation formula is:

[0137] (11)

[0138] Among them, is expressed as the node 's feature vector in the second - order graph.

[0139] Cross - compare the original graph features in the attribute module with the original graph features in the structure module , which is used to ensure the consistency of multi - view features. The calculation formula is:

[0140] (12)

[0141] Through the joint constraint of the above - mentioned contrast losses, using the learnable weight parameter , fuse the features of the three views to generate the intermediate fusion feature . The fusion formula is:

[0142] (13)

[0143] Among them, is expressed as the concatenation operation along the feature dimension.

[0144] Step S5 specifically includes:

[0145] S5.1: Train the angular graph feature contrast fusion module to obtain an intermediate model, and optimize the loss function through gradient descent . The overall contrast loss of the angular graph feature contrast fusion module is:

[0146] (14)

[0147] Among them, and Denoted as a hyperparameter, used to balance each loss term.

[0148] S5.2: Input the processed image into a convolutional neural network model to obtain image features, and fuse the image features with the angle map features. The convolutional neural network model is HRNet, and the image features at different scales in four stages are fused with the angle map features through a conversion module, where the conversion module uses a fully connected layer to reduce the two-dimensional image features to one-dimensional.

[0149] In the intermediate model supervised fine-tuning process of step S6, other parameters of the intermediate model are frozen, and the output head of the prediction module is fine-tuned using a weighted loss function. The specific formula is expressed as:

[0150] (15)

[0151] Where Denotes the weighted sum of the mean squared error and the mean absolute error, and the weighting coefficient .

[0152] Implementing the 3D human pose estimation system based on the fusion of angle map and image features of the present invention includes a data preprocessing module, a graph network construction module, an angle map feature contrast fusion module, an image feature fusion module, and an inference module;

[0153] The data preprocessing module obtains the human 2D key point coordinates from the original image and performs normalization processing. At the same time, the Laplacian of Gaussian operator is used to enhance the edge features of the original image. Specifically, it includes:

[0154] S1.1: The two-dimensional images in the Human3.6M dataset Pass through a cascaded pyramid network to obtain the normalized 16 2D key point coordinates , where ;

[0155] S1.2: Use the Laplacian of Gaussian operator to Perform edge sharpening processing to obtain , whose resolution is ;

[0156] The Laplacian of Gaussian operator is:

[0157] (1)

[0158] Where Denotes the standard deviation, , Respectively denote the coordinates of the pixel points, Denotes the natural logarithm.

[0159] The graph network construction module constructs a network in two different ways using the known 2D pose information, enabling the new graph network structure to have the ability to represent joint angle information, specifically including:

[0160] S2.1: Convert the original skeleton graph with human joints as nodes into a line graph with human bones as nodes ;

[0161] where denotes the original skeleton graph and the line graph 's node sets and edge sets, denoted as and , the nodes of the line graph are the edges of the original skeleton graph, denoted as , and the edges are denoted as ; they are connected if and only if two edges in the original skeleton graph have a common node;

[0162] S2.2: Calculate the angle between two connected edges in the original skeleton graph as the edge weight of the line graph , and take the endpoint coordinates of both ends of the edge of the original skeleton graph as the node features of the line graph ; ;

[0163] where , the adjacency matrix of the line graph is denoted as , where , can be obtained by the following formula:

[0164] (2)

[0165] where is a binary matrix, indicating whether the nodes in the node set of the original graph are the endpoints of the edges in the edge set , denotes the identity matrix, and F(·) represents the function for calculating the joint angles in the original skeleton graph.

[0166] S3.1: Convert the original skeleton graph into a second-order graph that directly represents the angular relationship between joint points ;

[0167] where , represents the edges formed between nodes with a path length of 2 in the original skeleton graph.

[0168] S3.2: Calculate the second-order graph ​adjacency matrix of, and for the second-order graph Generate node features ;

[0169] where the original skeleton graph The adjacency matrix is expressed as where , , a value of 1 indicates that two nodes are connected, and 0 indicates not connected; the second-order graph representing the angle The adjacency matrix of is expressed as:

[0170] (3)

[0171] Second-order graph The node features of , when and are second-order neighbors, and The angle value between, when or and are not second-order neighbors, .

[0172] The angle graph feature contrast and fusion module makes the angle graph features of different scales have consistent dimensions and perform feature fusion by cross-comparing multi-scale angle graph features, specifically including:

[0173] S4.1: Respectively extract features from the line graph and the second-order graph through ordinary graph neural network layers to generate intermediate features and where is the feature dimension, and at the same time apply an ordinary graph neural network to the original graph to generate an intermediate feature ;

[0174] The ordinary graph neural network layer is expressed as:

[0175] (4)

[0176] The above formula is further described as:

[0177] (5)

[0178] (6)

[0179] where, represents a learnable weight matrix, represents the line graph The adjacency matrix of represents the activation function, represents the modulation matrix, represents the Hadamard product, represents the adjacency matrix obtained by symmetric normalization of the adjacency matrix of the second-order graph or the original graph ;

[0180] S4.2: Perform dimensional transformation on the intermediate features of the line graph through the anti-pooling graph neural network to obtain .

[0181] The anti-pooling graph neural network layer includes two separate graph convolutional networks, which are used to obtain the embedding matrix and the alignment matrix respectively:

[0182] (7)

[0183] (8)

[0184] Among them, represents the embedding matrix, represents the alignment matrix. Finally, the line graph obtains the intermediate layer features with the same dimension as the second-order graph through the following formula:

[0185] (9)

[0186] Among them, represents the intermediate features of the line graph after the anti-pooling operation.

[0187] S4.3: Achieve the fusion of multi-scale angular graph features through cross-contrast learning. Specifically, the feature learning is constrained by three different forms of loss functions.

[0188] Cross-contrast the original graph features with the line graph features to enhance the expression of local angular information and edge geometric characteristics. The calculation formula is:

[0189] (10)

[0190] Among them, represents the cosine similarity, represents the temperature parameter, , respectively represent the vectors of the nodes in the original graph and the line graph.

[0191] Cross-contrast the original graph features and the second-order graph features , for enhancing the capture of high - order angular relationships and global structure information, the calculation formula is:

[0192] (11)

[0193] Among them, is represented as the eigenvector of node in the second - order graph.

[0194] The original graph features in the cross - comparison attribute module and the original graph features in the structure module , for ensuring the consistency of multi - view features, the calculation formula is:

[0195] (12)

[0196] Through the joint constraint of the above - mentioned contrast loss, using the learnable weight parameter , the features of the three views are fused to generate intermediate fusion features , and the fusion formula is:

[0197] (13)

[0198] Among them, represents the concatenation operation along the feature dimension.

[0199] Train the angular graph feature contrast fusion module to obtain an intermediate model, and use a fully - connected layer network in the image feature fusion module to fuse two different features of images and graphs at different stages of the model, specifically including:

[0200] S5.1: Train the angular graph feature contrast fusion module to obtain an intermediate model, and optimize the loss function through gradient descent , and the overall contrast loss of the angular graph feature contrast fusion module is:

[0201] (14)

[0202] Among them, and are represented as hyperparameters for balancing each loss term.

[0203] S5.2: Input the processed image into a convolutional neural network model to obtain image features, and fuse the image features with the angular graph features. The convolutional neural network model is HRNet, and the image features of different scales in the four stages are fused with the angular graph features through a conversion module, where the conversion module uses a fully - connected layer to reduce the two - dimensional image features to one - dimensional.

[0204] After the intermediate model is supervised and fine-tuned, the inference module finally outputs the inference result of the model, that is, the 3D pose information of the human body.

[0205] Embodiment 2

[0206] This embodiment relates to a system for implementing the 3D human pose estimation method based on the fusion of angle maps and image features in Embodiment 1, including:

[0207] A data preprocessing module, a graph network construction module, an angle map feature comparison and fusion module, an image feature fusion module, and an inference module;

[0208] The data preprocessing module obtains the 2D key point coordinates of the human body from the original image and performs normalization processing. At the same time, the Laplacian of Gaussian operator is used to enhance the edge features of the original image;

[0209] The graph network construction module uses two different methods to construct a network using the known 2D pose information, so that the new graph network structure has the ability to represent joint angle information;

[0210] The angle feature comparison and fusion module makes the angle map features of different scales have the same dimension and performs feature fusion by using the embedding matrix and the alignment matrix;

[0211] The image feature fusion module uses a fully connected layer network to fuse two different features of the image and the graph at different stages of the model;

[0212] The inference module finally outputs the inference result of the model, that is, the 3D pose information of the human body;

[0213] The data preprocessing module, the graph network construction module, the angle map feature comparison and fusion module, the image feature fusion module, and the inference module are connected in sequence.

[0214] Embodiment 3

[0215] This embodiment relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the 3D human pose estimation method based on the fusion of angle maps and image features in Embodiment 1.

[0216] The content described in the embodiments of this specification is only a list of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that those skilled in the art can think of according to the inventive concept of the present invention.

Claims

1. A 3D human pose estimation method based on the fusion of angular graphs and image features, characterized in that: It includes the following steps: S1: Construct an original dataset, obtain the 2D key point coordinates of the human body using a trained 2D human pose detector, and obtain an original image with obvious edge features using the Laplacian of Gaussian operator; S2: Convert the original skeleton graph with human joints as nodes into a line graph with bones as nodes; S3: Convert the original skeleton graph into a second-order graph that explicitly represents the angular relationships between joints; specifically including: S3.1: Convert the original skeleton diagram into a second-order graph that explicitly represents the angular relationship between joint points ; Among them , indicating the connecting edges formed between nodes with a path length of 2 in the original skeleton diagram; S3.2: Calculate the second-order graph 's adjacency matrix, and generate node features for the second-order graph ; ; Among them, the original skeleton diagram is represented by the adjacency matrix , where = 16, , the value of 1 indicates that two nodes are connected, and 0 indicates not connected; the second-order graph representing the angle The adjacency matrix of is expressed as: Second-order graph Node features When and are second-order neighbors, = the angle value between and or When and are not second-order neighbors, S4: Construct an angular graph feature contrast and fusion module to fuse angular graph features of different scales through cross-contrast learning; the angular graph feature contrast and fusion module consists of a general graph neural network layer and an anti-pooling graph convolutional layer. The former is used to extract the features of the line graph and the second-order graph, and the latter keeps the dimensions of the intermediate features of the line graph and the second-order graph consistent through anti-pooling operations, and finally completes feature fusion through learnable weight parameters; S5: Train the angular graph feature contrast and fusion module to obtain an intermediate model: input the processed image into a convolutional neural network model to obtain image features, fuse the image features with the angular graph features, and finally regress to obtain 3D human pose information; S6: Perform supervised fine-tuning on the intermediate model: adjust the parameters of the regression head of the intermediate model, freeze the other parameters of the intermediate model, and obtain the final 3D human pose estimation model.

2. The 3D human pose estimation method based on the fusion of angular graphs and image features according to claim 1, wherein: The specific steps of S1 include: S1.1: The two-dimensional images in the Human3.6M dataset are passed through a cascaded pyramid network to obtain the normalized coordinates of 16 2D key points , where ; S1.2: Use the Laplacian of Gaussian operator to perform edge sharpening to obtain , whose resolution is ; The Laplacian of Gaussian operator is: Among them, represents the standard deviation, , respectively represent the coordinates of the pixel point, represents the natural logarithm.

3. A 3D human pose estimation method based on the fusion of angular maps and image features according to claim 1, characterized in that: The specific steps of S2 include: S2.1: Convert the original skeleton diagram with human joints as nodes into a line diagram with human bones as nodes ; Among them is represented as the original skeleton diagram and the line graph of the node set and the edge set, represented as and , the nodes of the line graph are the edges of the original skeleton diagram, represented as , while the edges are represented as ; they are connected if and only if two edges in the original skeleton diagram have a common node; S2.2: Calculate the original skeleton graph The angle between two edges with common nodes in is used as the weight of the upper edge of the line graph, and the coordinates of the two end points of the bones at both ends of the edges of the original skeleton graph are used as the node features of the line graph ; Among them , the adjacency matrix representation of the line graph is expressed as , where is obtained by the following formula: Among them is a binary matrix, representing whether the nodes in the original skeleton graph are the endpoints of the edges in the edge set , is represented as an identity matrix, and F(·) is represented as a function for calculating the joint angles in the original skeleton graph.

4. The 3D human pose estimation method based on the fusion of angular graphs and image features according to claim 1, characterized in that: The specific steps of S4 include: S4.1: Respectively perform feature extraction on the line graph and the second-order graph to generate intermediate features and , where is the feature dimension. At the same time, apply a general graph neural network to the original graph to generate intermediate feature ; The general graph neural network layer is expressed as: The above formula is further described as: Among them, represents a learnable weight matrix, represents a line graph of the adjacency matrix, represents an activation function, represents a modulation matrix, represents a Hadamard product, represents the adjacency matrix obtained by symmetric normalization of the adjacency matrix of the second-order graph or the original graph ; S4.2: Perform dimensional transformation on the intermediate features of the line graph through the inverse pooling graph convolutional layer to obtain ; The anti-pooling graph convolutional layer includes two separate graph convolutional networks, which are used to obtain the embedding matrix and the alignment matrix respectively: Among them, represents the embedding matrix, represents the alignment matrix. Finally, the line graph obtains the intermediate features with the same dimension as the second-order graph through the following formula: Among them, represents the intermediate feature of the line graph after the anti-pooling operation; S4.3: Realize the fusion of angular graph features of different scales through cross-contrast learning; specifically, constrain feature learning through three different forms of loss functions; Intermediate features of the original figure in cross-comparison and intermediate features of the line graph , for enhancing the expression of local angular information and edge geometric characteristics, and the calculation formula is: Among them, is expressed as cosine similarity, is expressed as temperature parameter, are respectively expressed as feature vectors of the node in the original graph and the line graph; The intermediate features of the original graph in the cross-comparison and the intermediate features of the second-order graph , which are used to strengthen the capture of high-order angular relationships and global structural information, and the calculation formula is: Among them, is expressed as a node eigenvector in the second-order graph; The intermediate features of the original graph in cross-comparison And the intermediate features output after the original graph is processed by the second-order graph encoder , which is used to ensure the consistency of multi-view features, and the calculation formula is: Through the joint constraint of the contrastive loss, using learnable weight parameters , the features of the three views are fused to generate intermediate fused features , and the fusion formula is: Among them, represents a concatenation operation along the feature dimension.

5. A 3D human pose estimation method based on the fusion of angular graphs and image features according to claim 1, characterized in that: The specific steps of S5 include: S5.1: Train the angular graph feature contrast fusion module to obtain an intermediate model, and optimize the loss function through gradient descent , and the overall contrast loss of the angular graph feature contrast fusion module is as follows: Among them, and are expressed as hyperparameters and are used to balance each loss term; S5.2: Input the processed image into a convolutional neural network model to obtain image features, and fuse the image features with the angular graph features; the convolutional neural network model is HRNet, and the image features of different scales in the four stages are fused with the angular graph features through a conversion module, where the conversion module uses a fully connected layer to reduce the two-dimensional image features to one-dimensional.

6. The 3D human pose estimation method based on the fusion of angular graphs and image features according to claim 1, characterized in that: During the supervised fine-tuning process of the intermediate model described in step S6, other parameters of the intermediate model are frozen, and a weighted loss function is used to fine-tune the output head of the inference module. The specific formula is expressed as: Among them, represents the weighted sum of the mean squared error and the mean absolute error , with the weighting coefficient .

7. A system for implementing a 3D human pose estimation method based on the fusion of angular graphs and image features as described in claim 1, characterized in that: It includes a data preprocessing module, a graph network construction module, an angular graph feature contrast and fusion module, an image feature fusion module, and an inference module; The data preprocessing module obtains the 2D key point coordinates of the human body from the original image and performs normalization processing, and at the same time uses the Laplacian of Gaussian operator to enhance the edge features of the original image; The graph network construction module uses two different methods to construct a network using the known 2D pose information, so that the new graph network structure has the ability to represent joint angle information; specifically includes: S3.1: Convert the original skeleton graph into a second-order graph that explicitly represents the angular relationships between joint points ; Among them , represents the connecting edges formed between nodes with a path length of 2 in the original skeleton diagram; S3.2: Calculate the second-order graph 's adjacency matrix and generate node features for the second-order graph ; ; Among them, the original skeleton diagram The adjacency matrix representation is , where = 16, , the value of 1 indicates that two nodes are connected, and 0 indicates not connected; the second-order graph representing the angle The adjacency matrix of is expressed as: Second-order graph Node features of When and are second-order neighbors, = the angular value between and or When and are not second-order neighbors, The angular graph feature contrast and fusion module makes the angular graph features of different scales have consistent dimensions and performs feature fusion by cross-comparing the angular graph features of different scales; The image feature fusion module uses a fully connected layer network to fuse two different features of the original image at different stages of the model; The inference module finally outputs the inference result of the model, that is, the 3D pose information of the human body; The data preprocessing module, graph network construction module, angular graph feature comparison and fusion module, image feature fusion module, and inference module are connected in sequence.

8. A computer-readable storage medium, characterized in that, A program is stored thereon, and when the program is executed by a processor, the 3D human pose estimation method based on the fusion of angular graphs and image features according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • A Method for Establishing a 3D Human Pose Estimation Model Based on a Single-Frame Image and Its Application

    CN113192186B

  • 3D Human Pose Estimation Method Based on Joint Data Augmentation and Network Training Model

    CN113361570B

  • A 3D human pose estimation method integrating local and global features

    CN114565938B

  • Multi-view feature fusion method and system for 3D human posture estimation

    CN114758205B

  • Human body three-dimensional posture estimation method based on structural information

    CN110427877A