Three-dimensional human body posture estimation method based on multilevel dense posture generation
By introducing a multi-level dense pose generation method in three-dimensional human pose estimation, a dense 2D pose symbol generation model and decoder are used to generate dense 2D poses, and a three-dimensional pose estimation is performed through the Transformer network, the accuracy problem of sparse two-dimensional skeleton representation in complex scenarios is solved, and pose estimation with high accuracy and high robustness is achieved.
Patent Information
- Application Number
- CN202510209758.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
When the existing three-dimensional human posture estimation method is used to process sparse two-dimensional skeleton representations, it is difficult to effectively utilize the skeleton context information, resulting in a significant decrease in the accuracy of posture estimation in scenarios containing self-occlusion and complex actions.
A three-dimensional human pose estimation method based on multi-level dense pose generation is adopted. Through the autoregressive pose symbol generation model and decoder, dense two-dimensional poses are generated layer by layer, rich skeleton context information is introduced, and three-dimensional pose estimation is performed through the Transformer network.
It significantly improves the accuracy of posture estimation containing self-occlusion and complex actions, can obtain the accuracy matching the multi-frame video input under a single-frame image input, and effectively alleviates problems such as joint occlusion, partial limb occlusion and depth discontinuity.
Smart Images

Figure CN120047624A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular, to a three-dimensional human pose estimation method based on multi-level dense pose generation. Background Art
[0002] Three-dimensional human pose estimation aims to predict the three-dimensional spatial positions of human joint points through a single monocular image or video. Specifically, first, two-dimensional joints are estimated from the input image, and then the estimated two-dimensional joints are lifted to three-dimensional poses. As a hot topic in the field of computer vision, three-dimensional human pose estimation has broad research prospects and is also widely applied in various fields, such as abnormal behavior detection, human action recognition, etc. However, the lack or incompleteness of information on occluded parts greatly increases the difficulty of accurately restoring three-dimensional poses from two-dimensional images, resulting in a significant decrease in the accuracy of three-dimensional pose estimation.
[0003] The paper "Towards alleviating the modeling ambiguity of unsupervised monocular 3d human pose estimation" published by Radwan et al. in the "IEEE International Conference on Computer Vision" (ICCV2021) discloses a method that uses geometric constraints to enhance the structural understanding of joint correlations, but it has limited effectiveness in severe occlusion situations. The paper "Mixste: Seq2seq mixed spatio-temporal encoder for 3d human pose estimation in video" published by Zhang et al. in the "IEEE / CVF Conference on Computer Vision and Pattern Recognition" (CVPR 2022) discloses a method that improves the inference of occluded regions by enriching temporal information. The paper "Lifting by image – leveraging image cues for accurate 3d human pose estimation" published by Zhao et al. in the "AAAI Conference on Artificial Intelligence" (AAAI 2024) discloses a method that enhances the model's inference ability for occluded actions by introducing visual cues. However, these methods usually rely on complex temporal or visual encoders, increasing the number of model parameters and computational complexity. Essentially, the above methods focus on information augmentation in the lifting stage while ignoring the limitations of sparse two-dimensional skeleton representations. The sparse input limits the utilization of skeleton context information and fundamentally affects the two-dimensional to three-dimensional lifting effect. Specifically, current two-dimensional pose datasets usually use a small number of key points to represent the human body (such as 17 joint points in Human3.6M). This sparse input for lifting essentially limits the utilization of local context. Therefore, developing hierarchical dense two-dimensional skeleton representations is crucial for exploring complex skeleton contexts, thereby enhancing the robustness of two-dimensional to three-dimensional lifting. Summary of the Invention
[0004] Aiming at the deficiencies of sparse two-dimensional skeleton representations in the prior art, the present invention provides a three-dimensional human pose estimation method based on multi-level dense pose generation, which can introduce rich skeleton context information, solve the problem of sparse two-dimensional skeleton representations, and further improve the accuracy of pose estimation for poses with self-occlusions and complex actions.
[0005] According to a first aspect of the present invention, there is provided a three-dimensional human pose estimation method based on multi-level dense pose generation, including:
[0006] Obtain a two-dimensional human pose, where the two-dimensional human pose is the two-dimensional coordinates of human joint points;
[0007] Using an autoregressive pose token generation model, in the order from the central joint point of the human body to the edge joint points, successively obtain an input pose generation token and a dense pose generation token according to the two-dimensional human pose;
[0008] Through a decoder, generate a first-layer densified pose according to the dense pose generation token, and generate a second-layer densified pose according to the input pose generation token and the dense pose generation token;
[0009] Concatenate the two-dimensional human pose, the first-layer densified pose, and the second-layer densified pose, and input them into a Transformer to obtain three-dimensional pose estimation.
[0010] Preferably, the using an autoregressive pose token generation model, in the order from the central joint point of the human body to the edge joint points, successively obtaining an input pose generation token and a dense pose generation token according to the two-dimensional human pose includes:
[0011] According to the input two-dimensional human pose of J s joint points Use an action projector, a global causal self-attention module, and a prediction module to generate a central joint point generation token of the input two-dimensional human pose with index i = 0 and the corresponding dense pose generation tokens for r joint points
[0012] According to the central joint point generation token of the input two-dimensional human pose and the corresponding dense pose generation tokens for r joint points Based on an input pose codebook and a dense pose codebook Use a local self-attention module, a global causal self-attention module, and a prediction module to successively generate input pose generation tokens with index i = 1 to J s joint points and the corresponding dense pose generation tokens
[0013] Output J s input pose generation tokens for joint points and J d dense pose generation tokens for joint points where J d = rJs 。
[0014] Preferably, the input two-dimensional human pose of the s J joint points Using an action projector, a global causal self-attention module, and a prediction module, generate the central joint point generation token of the input two-dimensional human pose with index i = 0 and the corresponding dense pose r joint point generation tokens including:
[0015] Obtain the first intermediate feature g of the central joint point from the two-dimensional human pose using the action projector 0 ;
[0016] Extract the second intermediate feature h from the first intermediate feature g of the central joint point using the global causal self-attention module 0 0 ;
[0017] Obtain the central joint point generation token of the input two-dimensional human pose according to the second intermediate feature h using the prediction module 0 and the reconstruction feature
[0018] Obtain the corresponding dense pose r joint point generation tokens according to the reconstruction feature using the prediction module
[0019] Preferably, the central joint point generation token of the input two-dimensional human pose and the corresponding dense pose r joint point generation tokens Based on the input pose codebook and the dense pose codebook Using the local self-attention module, the global causal self-attention module, and the prediction module, sequentially generate the input pose generation tokens with index i = 1 to J s joint points and the corresponding dense pose generation tokens including cyclic:
[0020] Obtain the input pose reconstruction feature from the already generated input pose generation tokens according to the input pose codebook C s
[0021]
[0022] Obtain r dense pose reconstruction features from the already generated dense pose generation tokens according to the dense pose codebook C d
[0023] Reconstruct features from the input pose using the local self-attention module and the r dense pose reconstruction features Extract local intermediate features Average all r + 1 local intermediate features to obtain the first intermediate feature g of the i-th joint i , and use the global causal self-attention module to obtain the corresponding second intermediate features h from the first intermediate features g of the joints with indices from 0 to i 0 ,…,g i Extract the corresponding second intermediate features h 0 ,…,h i ;
[0024] Use the prediction module to obtain the generated token of the i-th joint of the input pose according to the second intermediate feature h i and the reconstruction feature and the reconstruction feature Use the prediction module to obtain the generated tokens of the r joints of the corresponding dense pose according to the reconstruction feature Obtain the generated tokens of the r joints of the corresponding dense pose
[0025] Preferably, the action projector includes multiple layers of linear projection layers, MLP-Mixer layers, and channel transpose layers, where the MLP-Mixer layer includes layer normalization operations and multi-layer perceptrons;
[0026] The global causal self-attention module includes multiple layers of causal self-attention layers and a feed-forward neural network layer;
[0027] The local self-attention module includes multiple layers of self-attention layers and a feed-forward neural network layer;
[0028] The prediction module includes multiple layers of self-attention layers and a feed-forward neural network layer.
[0029] Preferably, the decoder generates the first layer of densified pose according to the generated tokens of the dense pose, and generates the second layer of densified pose according to the generated tokens of the input pose and the generated tokens of the dense pose, including:
[0030] Obtain the reconstruction feature of the input pose according to the generated token p of the input pose s and the input pose codebook C s , and obtain the reconstruction feature of the input pose Generate the upsampled pose feature through the input pose decoder
[0031] Obtain the generated token p of the dense pose d and the dense pose codebook Cd , obtain dense pose reconstruction features Generate J through the first-layer dense pose decoder d The first-layer reconstructed dense pose of J joint points
[0032] According to the dense pose reconstruction features and the upsampled pose features Generate J through the second-layer dense pose decoder f The second-layer reconstructed dense pose of J joint points Where: J f > J d .
[0033] Preferably, the structures of the input pose decoder, the first-layer dense pose decoder, and the second-layer dense pose decoder are the same, and each includes multiple layers of linear projection layers, MLP-Mixer layers, and channel transpose layers, where the MLP-Mixer layer includes layer normalization operations and multi-layer perceptrons.
[0034] Preferably, the autoregressive pose token generation model is obtained through the following training process, specifically:
[0035] Obtain the two-dimensional human pose and the corresponding first-layer densified pose and second-layer densified pose;
[0036] According to the first-layer densified pose and the second-layer densified pose, extract input pose joint point tokens and dense pose joint point tokens, perform global and local joint point alignment, calculate the loss function, and update and optimize the input pose decoder, the first-layer dense pose decoder, the second-layer dense pose decoder, the output pose codebook, and the dense pose codebook;
[0037] According to the updated and optimized two-dimensional human pose, the output pose codebook, and the dense pose codebook, generate input pose joint point prediction tokens and the dense pose joint point prediction tokens, compare them with the input pose joint point tokens and the dense pose joint point tokens, and calculate the loss function for updating the autoregressive pose token generation model parameters.
[0038] Preferably, the extracting input pose joint point tokens and dense pose joint point tokens according to the first-layer densified pose and the second-layer densified pose, performing global and local joint point alignment, calculating the loss function, and updating and optimizing the input pose decoder, the first-layer dense pose decoder, the second-layer dense pose decoder, the output pose codebook, and the dense pose codebook includes:
[0039] The second-layer densified pose containing J f joint points Obtain J through the first-layer dense pose encoder d D-dimensional dense pose features of the joints wherein the first-layer pose encoder is a multi-layer MLP-Mixer;
[0040] Apply the dense pose feature z d Obtain J through the second-layer dense pose encoder s D-dimensional input pose features of the joints wherein the second-layer pose encoder is a multi-layer MLP-Mixer;
[0041] According to the input pose codebook C s , quantize the input pose feature z s into input pose code elements Use the value of each element in the input pose code element q s as an index to obtain the corresponding row vector in the input pose codebook C s as the reconstructed feature of the input pose feature z s
[0042] Apply the input pose estimation feature to generate an upsampled pose feature by inputting it into the pose decoder and add it to the dense pose feature z d to obtain a combined pose feature z' d ;
[0043] According to the dense pose codebook C d , quantize the combined pose feature z' d into dense pose code elements Use the value of each element in the dense pose code element q d as an index to obtain the corresponding row vector in the dense pose codebook C d as the reconstructed feature of the dense pose feature z d
[0044] According to the reconstructed feature of the dense pose feature z d generate a first-layer reconstructed dense pose through the first-layer dense pose decoder
[0045] According to the dense pose reconstruction feature and the upsampled pose feature generate a second-layer reconstructed dense pose through the second-layer dense pose decoder
[0046] The dense pose feature z d of the reconstruction feature and the upsampled pose feature are concatenated and input into an action projector to extract an action classification vector
[0047] Calculate the action classification vector and the cross-entropy CrossEntropy loss L A between the true action class y global to perform action-level global alignment:
[0048]
[0049] Calculate the average reconstruction feature of r joint points corresponding to the i-th joint point of the input pose in the first-layer dense pose estimation feature : Calculate the reconstruction feature of the i-th joint point of the input pose and of the local alignment loss:
[0050]
[0051] where: ·,· is the inner product of two vectors, τ is the temperature parameter, and J s is the number of joints of the two-dimensional human pose;
[0052] Calculate the joint position error between the second-layer reconstructed dense pose and the second-layer densified pose x f , and the joint position error between the first-layer reconstructed dense pose and the first-layer dense pose x d :
[0053]
[0054] where, · F is the Frobenious norm of the matrix;
[0055] Calculate the commitment error Commitmentloss between the input pose estimation feature and the input pose feature z s , and the commitment error between the first-layer dense pose reconstruction feature and the first-layer dense pose feature z d :
[0056]
[0057] Among them, sg(·) represents stopping gradient backpropagation, and β is a control parameter;
[0058] Calculate the training loss function: L 1 = L global + L local + L pos + L sg , and update the parameters and extract features using the backpropagation algorithm according to the training of the loss function L 1 ;
[0059] Repeat the above process until the training loss function converges to obtain the input pose decoder, the first-layer dense pose decoder, the second-layer dense pose decoder, the output pose codebook, and the dense pose codebook.
[0060] Preferably, generating the input pose joint point prediction code element and the dense pose joint point prediction code element according to the updated and optimized two-dimensional human pose, the input pose codebook, and the dense pose codebook, comparing with the input pose joint point code element and the dense pose joint point code element, and calculating the loss function to update the parameters of the autoregressive pose code element generation model, including:
[0061] According to the two-dimensional human pose and the input pose codebook C s and the dense pose codebook C d , obtain the input pose generation code element p s and the dense pose generation code element p d ;
[0062] Calculate the cross-entropy between the input pose generation code element p s and the input pose code element q s , and the cross-entropy between the dense pose generation code element p d and the dense pose code element q d , to obtain the training loss function:
[0063] L 2 = CrossEntropy(p s , q s ) + CrossEntropy(p d , q d )
[0064] According to the training of the loss function L 2 Use the backpropagation method to update the parameters of the autoregressive pose code element generation model and extract features;
[0065] Repeat the above process until the training loss function converges to obtain the autoregressive pose code element generation model.
[0066] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:
[0067] (1) In the three-dimensional human pose estimation method based on multi-level dense pose generation in the embodiments of the present invention, by introducing dense two-dimensional poses, the human joint point information is enriched, and the accuracy of three-dimensional human pose estimation is greatly improved. The accuracy of using the pose input of a single-frame image can be obtained to match that of the multi-frame video input.
[0068] (2) In the three-dimensional human pose estimation method based on multi-level dense pose generation in the embodiments of the present invention, by constructing the steps including S100 to S400, the input data features can be compactly extracted, and the number of network model parameters is saved.
[0069] (3) In the three-dimensional human pose estimation method based on multi-level dense pose generation in the embodiments of the present invention, by introducing dense two-dimensional poses, the accuracy of pose estimation for poses with self-occlusion and complex actions is significantly improved, and it has strong flexibility and scalability.
[0070] (4) The embodiments of the present invention have been verified for the collected virtual reality three-dimensional human poses. The results fully confirm its accurate estimation ability for typical human poses in various real environments, and effectively alleviate problems such as joint occlusion, partial limb occlusion, and depth discontinuity. The embodiments of the present invention can flexibly adopt single-frame images or multi-frame videos and be applied in application fields such as virtual reality and the metaverse to achieve real-time, high-precision, and high-robust human pose estimation for moving human poses, thereby promoting the progress of various downstream tasks (such as three-dimensional reconstruction, action recognition, etc.). BRIEF DESCRIPTION OF THE DRAWINGS
[0071] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent:
[0072] Figure 1 It is a flowchart of three-dimensional human pose estimation based on multi-level dense pose generation according to an embodiment of the present invention;
[0073] Figure 2 It is a flowchart of a training method for an autoregressive pose token generation model according to an embodiment of the present invention;
[0074] Figure 3 It is the MPJPE comparison data of various methods on the Human3.6M dataset in a specific embodiment of the present invention;
[0075] Figure 4 It is a schematic diagram of three-dimensional human pose estimation based on multi-level dense pose generation in a specific example of the present invention;
[0076] Figure 5 This is a schematic diagram of the autoregressive pose token generation model structure in a specific example of the present invention, which consists of an action projector, a local self-attention module, a global causal self-attention module, and a prediction module. Specific implementation manners
[0077] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several deformations and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention. The parts not described in detail below can be implemented using the prior art.
[0078] It should be understood that the terms "first", "second", etc. in the following embodiments are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0079] To solve the technical problems in the background art, the three-dimensional human pose estimation method based on multi-level dense pose generation proposed by the present invention can introduce rich skeleton context information, solve the problem of sparse two-dimensional skeleton representation, and thus improve the accuracy of pose estimation for poses with self-occlusion and complex actions.
[0080] As Figure 1 shown, it is a flowchart of the three-dimensional human pose estimation method based on multi-level dense pose generation according to an embodiment of the present invention. Please refer to Figure 1 , the three-dimensional human pose estimation method based on multi-level dense pose generation in this embodiment includes the following steps:
[0081] S100, obtain a two-dimensional human pose, where the two-dimensional human pose is the two-dimensional coordinates of human joint points;
[0082] S200, use the autoregressive pose token generation model to sequentially obtain an input pose generation token and a dense pose generation token according to the two-dimensional human pose in the order from the central joint point of the human body to the edge joint points;
[0083] S300, through a decoder, generate a first-layer densified pose according to the dense pose generation token obtained in S200, and generate a second-layer densified pose according to the input pose generation token and the dense pose generation token obtained in S200;
[0084] S400 splices the two-dimensional human pose of S100, the first-layer densified pose of S300, and the second-layer densified pose, and inputs them into the Transformer to obtain three-dimensional pose estimation.
[0085] In the above embodiment, by introducing the dense two-dimensional pose, the human joint point information is enriched, and the accuracy of three-dimensional human pose estimation is greatly improved, and the accuracy matching that of multi-frame video input can be obtained using the pose input of a single-frame image.
[0086] In a preferred embodiment of the present invention, S200 is implemented. Using the autoregressive pose token generation model, according to the order from the central joint point of the human body to the edge joint points, the input pose generation tokens and the dense pose generation codes are sequentially obtained according to the two-dimensional human pose (serialized pose). As Figure 5 shown, the autoregressive pose token generation model consists of an action projector 230, a local self-attention module 231, a global causal self-attention module 232, and a prediction module 233.
[0087] S200 can adopt steps S201 - S203. Specifically:
[0088] S201: Central joint point token generation. According to the input two-dimensional human pose of 16 joint points Using the action projector, the global causal self-attention module, and the prediction module, generate the central joint point generation token of the input two-dimensional human pose with index i = 0 and the corresponding dense pose 3 joint point generation codes
[0089] S202: Autoregressive joint point feature generation. According to the central joint point generation token of the input two-dimensional human pose and the corresponding dense pose 3 joint point generation codes Based on the input pose codebook and the dense pose codebook Using the local self-attention module, the global causal self-attention module, and the prediction module, sequentially generate the input pose generation tokens of indices i = 1 to 16 joint points and the corresponding dense pose generation codes
[0090]
[0091] S203: Output the input pose generation tokens of 16 joint points and the dense pose generation codes of 48 joint points
[0092] Furthermore, in a preferred embodiment, S201 is implemented, which can adopt steps S2011 - S2014:
[0093] S2011: Obtain the first intermediate feature g of the central joint point from the serialized poses using an action projector 0 , where the action projector consists of a 1-layer linear projection layer, an MLP-Mixer layer, and a channel transpose layer, and the MLP-Mixer layer includes a layer normalization operation and a multi-layer perceptron;
[0094] S2012: Extract the second intermediate feature h from the first intermediate feature g of the central joint point using a global causal self-attention module 0 0 , where the global causal self-attention module consists of 12 layers of causal self-attention layers and a feed-forward neural network layer;
[0095] S2013: Obtain the central joint point generation token and the reconstruction feature of the input two-dimensional human pose according to the second intermediate feature h using a prediction module 0 , where the prediction module consists of 4 layers of self-attention layers and a feed-forward neural network layer;
[0096] S2014: Obtain the corresponding dense pose 3-joint point generation tokens according to the reconstruction feature using a prediction module , where the prediction module consists of multiple layers of self-attention layers and a feed-forward neural network layer;
[0097] Furthermore, in a preferred embodiment, S202 is implemented, which may adopt steps S2021 - S2026:
[0098] S2021: Obtain the input pose reconstruction feature from the input pose generation tokens generated already according to the input pose codebook C s Specifically, C s contains 2048 learnable 128-dimensional row vectors. Using the value of the input pose token as an index, obtain the corresponding row vector in the input pose codebook C s as the input pose reconstruction feature
[0099] S2022: Obtain 3 dense pose reconstruction features from the dense pose generation tokens generated already according to the dense pose codebook C d Specifically, C d contains 2048 learnable 128-dimensional row vectors. Using the value of each element in the dense pose token as an index, obtain the corresponding row vector in the dense pose codebook C d as the dense pose reconstruction feature
[0100] S2023: Reconstruct features from the input pose using a local self-attention module and three dense pose reconstruction features Extract local intermediate features Average all four local intermediate features to obtain the first intermediate feature g of the i-th joint i , and the local self-attention module consists of one self-attention layer and one feed-forward neural network layer;
[0101] S2024: Use the global causal self-attention module to obtain the corresponding second intermediate features h 0 ,…,g i from the first intermediate features g of the joints indexed from 0 to i 0 ,…,h i ;
[0102] S2025: Use the prediction module to obtain the generated token i of the i-th joint of the input pose according to the second intermediate feature h and the reconstruction feature The prediction module consists of four self-attention layers and one feed-forward neural network layer;
[0103] S2026: Use the prediction module to obtain the generated tokens of three joints of the corresponding dense pose according to the reconstruction feature The prediction module consists of four self-attention layers and one feed-forward neural network layer.
[0104] In a preferred embodiment of the present invention, S300 is implemented. Through the decoder, the first layer of densified pose is generated according to the dense pose generated tokens obtained in S200, and the second layer of densified pose is generated according to the input pose generated tokens and the dense pose generated tokens obtained in S200, which can adopt steps S301 - S303.
[0105] S301: According to the input pose generated token p s and the input pose codebook C s , obtain the input pose reconstruction feature Generate the upsampled pose feature through the input pose decoder The input pose decoder consists of one linear projection layer, one MLP-Mixer layer, and one channel transpose layer, where the MLP-Mixer layer includes layer normalization operations and a multi-layer perceptron;
[0106] S302: According to the dense pose generated token p d and the dense pose codebook C d , obtain the dense pose reconstruction feature Generate the first-layer reconstructed dense pose of 48 keypoints through the first-layer dense pose decoder The first-layer dense pose decoder consists of a 1-layer linear projection layer, an MLP-Mixer layer, and a channel transpose layer, where the MLP-Mixer layer includes a layer normalization operation and a multi-layer perceptron;
[0107] S303: According to the dense pose reconstruction feature And the upsampled pose feature Generate the second-layer reconstructed dense pose of 96 keypoints through the second-layer dense pose decoder The second-layer dense pose decoder consists of a 1-layer linear projection layer, an MLP-Mixer layer, and a channel transpose layer, where the MLP-Mixer layer includes a layer normalization operation and a multi-layer perceptron.
[0108] In some specific embodiments, the Transformer architecture of S400 includes 12 layers of spatial self-attention layers and feed-forward neural network layers. Compared with other networks such as graph neural networks (GNNs) and convolutional neural networks (CNNs), the Transformer sequence modeling method is more convenient in introducing multi-level node end-to-end modeling. Using the Transformer network for 3D pose estimation shows its performance advantages.
[0109] In the above embodiments, steps S100 - S400 are implemented, which can compactly extract the input data features and save the number of network model parameters.
[0110] In order to improve the accuracy of the autoregressive pose token generation model, in a preferred embodiment, the model is trained. As Figure 2 shown, it is the flowchart of the training method of the autoregressive pose token generation model. Please refer to Figure 2 and Figure 5 the model structure diagram, the training of the autoregressive pose token generation model in this embodiment can adopt the following steps:
[0111] S500: Obtain the 2D human pose and the corresponding first-layer densified pose and second-layer densified pose;
[0112] S600: Multi-scale pose tokenization model training: According to the first-layer densified pose and the second-layer densified pose, extract the input pose keypoint tokens and dense pose keypoint tokens, perform global and local keypoint alignment, calculate the loss function, and use it to update the model parameters to obtain the input pose decoder, the first-layer dense pose decoder, the second-layer dense pose decoder, the output pose codebook, and the dense pose codebook for the 3D human pose estimation method based on multi-level dense pose generation in embodiments S100 - S400;
[0113] S700: Training of Autoregressive Pose Code Element Generation Model: Based on 2D human poses, output pose codebooks, and dense pose codebooks, generate predicted pose joint point code elements and dense pose joint point predicted code elements for the input poses, compare them with the input pose joint point code elements and dense pose joint point code elements, and calculate the loss function to update the model parameters.
[0114] Further, in a preferred embodiment, S600 is implemented, which can adopt steps S601 - S614. Specifically:
[0115] S601: Densify the second - layer pose containing 96 joint points Obtain 128 - dimensional dense pose features of 48 joint points through the first - layer dense pose encoder where the first - layer pose encoder is a 4 - layer MLP - Mixer;
[0116] S602: Pass the dense pose feature z d Through the second - layer dense pose encoder to obtain 128 - dimensional input pose features of 16 joint points where the second - layer pose encoder is a 4 - layer MLP - Mixer;
[0117] S603: According to the input pose codebook C s , quantize the input pose feature z s into input pose code elements Take the value of each element in the input pose code element q s as an index to obtain the corresponding row vector in the input pose codebook C s as the reconstruction feature of the input pose feature z s ;
[0118] S604: Pass the input pose estimation feature Through the input pose decoder to generate up - sampled pose features And add it to the dense pose feature z d to obtain the combined pose feature z' d ;
[0119] S605: According to the dense pose codebook C d , quantize the combined pose feature z' d into dense pose code elements Take the value of each element in the dense pose code element q d as an index to obtain the corresponding row vector in the dense pose codebook C d as the reconstruction feature of the dense pose feature z d ;
[0120] S606: According to the dense pose reconstruction feature Generate the first - layer reconstructed dense pose through the first - layer dense pose decoder
[0121] S607: According to the dense pose reconstruction feature and the up - sampled pose feature Generate the second - layer reconstructed dense pose through the second - layer dense pose decoder
[0122] S608: Concatenate the dense pose reconstruction feature and the up - sampled pose feature and input them into the action projector to extract the action classification vector
[0123] S609: Calculate the cross - entropy loss between the action classification vector and the true action category y A for action - level global alignment:
[0124]
[0125] S610: Calculate the average reconstruction feature of 3 joint points corresponding to the i - th joint point of the input pose in the first - layer dense pose estimation feature : Calculate the reconstruction feature of the i - th joint point of the input pose and and for the local alignment loss:
[0126]
[0127] where: ·,· is the inner product of two vectors, τ is the temperature parameter, which is 0.07;
[0128] S611: Calculate the joint - point position error between the second - layer reconstructed dense pose and the second - layer dense pose x f and the joint - point position error between the first - layer reconstructed dense pose and the first - layer dense pose x d :
[0129]
[0130] where, · F is the Frobenious norm of the matrix;
[0131] S612: Calculate the commitment error between the input pose estimation feature and the input pose feature z s and the commitment error between the first - layer dense pose reconstruction feature and the commitment error between the first - layer dense pose feature z d wherein, sg(·) represents stopping gradient backpropagation, β is a control parameter, and is 0.25;
[0132]
[0133]
[0134] S613: Calculate the training loss function: L 1 = L global + L local + L pos + L sg , and update the model parameters and extract features according to the training loss function L 1 using the backpropagation algorithm;
[0135] S614: Repeat the above process until the training loss function converges.
[0136] Furthermore, in another preferred embodiment, implement S700, which can adopt the following steps S701 - S204. Specifically:
[0137] S701: Adopt the autoregressive pose token generation model of Embodiment S200, and obtain the input pose generation token p s and the dense pose generation token p d according to the serialized pose, the input pose codebook C s and the dense pose codebook C d ;
[0138] S702: Calculate the cross - entropy between the input pose generation token p s and the input pose token q s , and the cross - entropy between the dense pose generation token p d and the dense pose token q d , and obtain the training loss function
[0139] L 2 = CrossEntropy(p s , q s ) + CrossEntropy(p d , q d )
[0140] S703: Update the model parameters and extract features according to the training loss function L 2 using the backpropagation algorithm;
[0141] S704: Repeat the above process until the training loss function converges.
[0142] Through the training of the autoregressive pose token generation model in the above embodiments, it can obtain accurate input pose generation tokens and dense pose generation tokens, providing a solid data foundation for subsequent 3D estimation.
[0143] To verify the feasibility and effectiveness of the 3D human pose estimation method based on multi-level dense pose generation in the above example, in a specific embodiment, the evaluation is divided into objective evaluation and subjective evaluation. The former includes statistical analysis of the virtual reality 3D human pose estimation results of the 3D human pose estimation based on multi-level dense pose generation to obtain indicators such as the mean per-joint position error (MPJPE); the latter includes visualizing the virtual reality 3D human pose estimation results of the 3D human pose estimation with action cues. In this embodiment, the virtual reality 3D human poses containing multiple actions are compared with the human pose estimation results of the original existing methods.
[0144] In terms of objective evaluation, the performance of the method of the embodiment of the present invention is compared with that of three types of publicly available optimal methods on the Human3.6M dataset. These methods use different additional information in the improvement stage. As Figure 3 shown: The pictures are divided into three groups: the first group uses the 2D poses detected by SH, the second group uses the 2D poses detected by CPN, and the third group uses the real label 2D poses as input. represents using temporal context, ★ represents using visual cues, and * represents using hierarchical information. f represents the number of frames used in the temporal method. For PoseFormer and MixSTE, f = 81; for VideoPose, f = 243; for MHFormer, f = 351. The best results are shown in bold, and the sub-optimal results are underlined. In the single-frame setting, the method of the embodiment of the present invention achieves the best performance among the methods based on hierarchical information (w / *) and visual information (w / ★). In addition, when using the results of the 2D pose detector as input, the method of the embodiment of the present invention is superior to the temporal-based methods in terms of both the MPJPE and PA-MPJPE metrics This advantage is closer to the requirements of the actual application scenario. Further data analysis proves that after adopting the method of the embodiment of the present invention, in the network construction method, additional dense 2D poses are considered, complex spatial context information is mined for the position-aware pose pattern of each action, and the pose features are refined by utilizing the correlation between the learnable pattern and the input pose sequence, enabling the model to well represent the structural information of the joint points in the case of complex actions and self-occlusions in the input 2D human poses, effectively improving the 3D human pose estimation results of the action cues for this action.
[0145] After adopting the method of this embodiment, the 3D human pose estimation effect of action prompts for complex actions and self-occlusion actions in virtual reality 3D human poses has been improved. Refer to Figure 4 the visualization results of
[0146] : The top two rows are two different input actions from top to bottom. The leftmost column is the 3D human pose result without adding a multi-level dense pose generation network. The middle column is the 3D human pose estimation result of the multi-level dense pose generation of this example. The rightmost column is the ground truth result, and the ground truth is the manually annotated semantic category. It can be seen that by introducing complex skeleton context information, the method of this embodiment helps to significantly improve the 3D human pose estimation accuracy of action prompts for complex actions and self-occlusion actions.
[0147] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.
Claims
1. A 3D human posture estimation method based on multi-level dense posture generation, characterized in that: include: Acquire a two-dimensional human body posture, where the two-dimensional human body posture is a two-dimensional coordinate of a human body joint point; Using an autoregressive gesture codeword generation model, in the order from the central joint point to the edge joint point of the human body, according to the two-dimensional human body posture, an input gesture generation codeword and a dense gesture generation codeword are sequentially obtained; Generate a first layer of densified gestures according to the dense gesture generation symbols, and generate a second layer of densified gestures according to the input gesture generation symbols and the dense gesture generation symbols by a decoder; The two-dimensional human body posture, the first-layer densified posture and the second-layer densified posture are concatenated and input into Transformer to obtain a three-dimensional posture estimation.
2. The 3D human posture estimation method based on multi-level dense posture generation according to claim 1, characterized in that: The method of using the autoregressive gesture codeword generation model to sequentially obtain input gesture generation codewords and dense gesture generation codewords according to the two-dimensional human body gesture in the order from the central joint point to the edge joint point of the human body includes: According to J s Input 2D human pose of joint points Using the action projector, the global causal self-attention module and the prediction module, the central joint point of the input 2D human posture with index i=0 is generated as a code element And the corresponding dense posture r joint points generate code elements Generate code elements according to the central joint points of the input two-dimensional human body posture And the corresponding dense posture r joint points generate code elements Based on input pose codebook and dense pose codebook Using the local self-attention module, the global causal self-attention module and the prediction module, we generate indexes i=1 to J in sequence s The input posture of the joint point generates code elements And the corresponding dense gesture generation codewords Output J s The input pose of the joint points generates code elements and J d Dense pose generation codewords for joint points Among them J d = rJ s .
3. The 3D human posture estimation method based on multi-level dense posture generation according to claim 2, characterized in that: According to J s Input 2D human pose of joint points Using the action projector, the global causal self-attention module and the prediction module, the central joint point of the input 2D human posture with index i=0 is generated as a code element And the corresponding dense posture r joint points generate code elements include: The first intermediate feature g of the central joint point is obtained from the two-dimensional human posture using the action projector 0 ; The global causal self-attention module is used to extract the first intermediate feature g of the central joint point. 0 Extract the second intermediate feature h 0 ; Using the prediction module according to the second intermediate feature h 0 Get the central joint point of the input two-dimensional human body posture to generate code elements and rebuild features Using the prediction module according to the reconstruction features Get the corresponding dense posture r joint points to generate code elements 4. The 3D human posture estimation method based on multi-level dense posture generation according to claim 2, characterized in that: Generating code elements according to the central joint points of the input two-dimensional human body posture And the corresponding dense posture r joint points generate code elements Based on input pose codebook and dense pose codebook Using the local self-attention module, the global causal self-attention module and the prediction module, we generate indexes i=1 to J in sequence s The input posture of the joint point generates code elements And the corresponding dense gesture generation codewords Including loops: According to the input posture codebook C s Generate codewords from generated input gestures Get input pose reconstruction features According to the dense gesture codebook C d Generate codewords from generated dense gestures Get r dense pose reconstruction features Reconstruct features from the input pose using the local self-attention module and the r dense pose reconstruction features Extracting local intermediate features Average all r+1 local intermediate features to obtain the first intermediate feature g of the i-th joint point i , using the global causal self-attention module to obtain the first intermediate feature g from the joint points indexed from 0 to i 0 ,…,g i Extract the corresponding second intermediate feature h 0 ,…,h i ; Using the prediction module according to the second intermediate feature h i Get the generated code element of the i-th joint point of the input posture and rebuild features Using the prediction module according to the reconstruction features Get the corresponding dense posture r joint points to generate code elements 5. The 3D human posture estimation method based on multi-level dense posture generation according to claim 2, characterized in that: The action projector includes a multi-layer linear projection layer, an MLP-Mixer layer and a channel transposition layer, wherein the MLP-Mixer layer includes a layer normalization operation and a multi-layer perceptron; The global causal self-attention module includes multiple layers of causal self-attention layers and feedforward neural network layers; The local self-attention module includes multiple layers of self-attention layers and feed-forward neural network layers; The prediction module includes multiple layers of self-attention layers and feed-forward neural network layers.
6. The 3D human posture estimation method based on multi-level dense posture generation according to claim 1, characterized in that: The method of generating a first layer of densified gestures according to the dense gesture generation symbols by a decoder, and generating a second layer of densified gestures according to the input gesture generation symbols and the dense gesture generation symbols, comprises: Generate a code element p according to the input gesture s and the input posture codebook C s , get the input posture reconstruction feature Generate upsampled pose features by inputting the pose decoder Generate codeword p according to the dense gesture d and the dense pose codebook C d , obtain dense pose reconstruction features Generate J through the first layer of dense pose decoder d The first layer of joints reconstructs the dense pose Reconstruct features based on the dense pose And the upsampled posture features Generate J through the second layer dense pose decoder f The second layer of the joints reconstructs the dense pose Among them: J f >J d .
7. The 3D human posture estimation method based on multi-level dense posture generation according to claim 6, characterized in that: The input posture decoder, the first layer dense posture decoder and the second layer dense posture decoder have the same structure, and all include a multi-layer linear projection layer, an MLP-Mixer layer and a channel transposition layer, wherein the MLP-Mixer layer includes a layer normalization operation and a multi-layer perceptron.
8. The 3D human posture estimation method based on multi-level dense posture generation according to claim 1, characterized in that: The autoregressive gesture codeword generation model is obtained through the following training process, specifically: Acquire the two-dimensional human body posture and the corresponding first-layer densified posture and second-layer densified posture; According to the first layer of densified posture and the second layer of densified posture, extract the input posture joint point codewords and the dense posture joint point codewords, perform global and local joint point alignment, calculate the loss function, update and optimize the input posture decoder, the first layer of dense posture decoder, the second layer of dense posture decoder, the output posture codebook and the dense posture codebook; According to the updated and optimized two-dimensional human body posture, the output posture codebook, and the dense posture codebook, input posture joint point prediction codewords and dense posture joint point prediction codewords are generated, and compared with the input posture joint point codewords and the dense posture joint point codewords, a loss function is calculated to update the autoregressive posture codeword generation model parameters.
9. The 3D human posture estimation method based on multi-level dense posture generation according to claim 8, characterized in that: According to the first layer of densified posture and the second layer of densified posture, input posture joint point codewords and dense posture joint point codewords are extracted, global and local joint point alignment is performed, a loss function is calculated, and an input posture decoder, a first layer of dense posture decoder, a second layer of dense posture decoder, an output posture codebook and a dense posture codebook are updated and optimized, including: Will include J f The second layer of densified pose of the joints Through the first layer of dense posture encoder, we get J d D-dimensional dense posture features of joint points The first-layer posture encoder is a multi-layer MLP-Mixer; The dense posture feature z d Through the second layer of dense posture encoder, we get J s D-dimensional input pose features of joint points The second-layer posture encoder is a multi-layer MLP-Mixer; According to the input posture codebook C s , the input posture feature z s Codeword conversion to input gesture codeword The input gesture codeword q s The value of each element in is used as an index to obtain the input posture codebook C s The corresponding row vector in is used as the input posture feature z s Reconstruction features The input pose estimation feature The upsampled pose features are generated by inputting the pose decoder and the dense posture feature z d Add together to obtain the combined posture feature z' d ; According to the dense gesture codebook C d , the combined posture feature z' d Dense gesture code The dense gesture codeword q d The value of each element in is used as an index to obtain the dense gesture codebook C d The corresponding row vector in is used as the dense posture feature z d Reconstruction features According to the dense posture feature z d The reconstruction features The first layer of dense pose decoder generates the first layer of reconstructed dense pose Reconstruct features based on the dense pose And the upsampled posture features The second layer dense pose decoder generates the second layer reconstructed dense pose The dense posture feature z d The reconstruction features And the upsampled posture features Concatenate and input into the action projector to extract the action classification vector Calculate the action classification vector and the real action category y A The cross entropy loss L between global , perform action-level global alignment: Calculate the first layer of dense pose estimation features In , the average reconstructed features of the r joint points corresponding to the i-th joint point of the input posture are: Calculate the reconstruction features of the i-th joint point of the input posture and Local alignment loss: Where: <·,·> is the inner product of two vectors, τ is the temperature parameter, J s is the number of joints in the 2D human pose; Calculate the second layer to reconstruct the dense pose and the second layer densified pose x f The joint position error between them, and the first layer reconstructs the dense pose and the first layer dense pose x d The position error between joints: Among them, ||·|| F is the Frobenious norm of the matrix; Compute input pose estimation features and input pose feature z s Commitment loss between the two layers, and the first layer of dense posture reconstruction features and the first layer dense posture feature z d The commitment error between: Where sg(·) means stopping gradient backpropagation, and β is the control parameter; Calculate the training loss function: L1 = L global +L local +L pos +L sg , according to the training loss function L1, the gradient backpropagation algorithm is used to update the parameters and extract the features; The above process is repeated until the training loss function converges, and the input pose decoder, the first layer dense pose decoder, the second layer dense pose decoder, the output pose codebook and the dense pose codebook are obtained.
10. The 3D human posture estimation method based on multi-level dense posture generation according to claim 8, characterized in that: The method generates input posture joint point prediction codewords and dense posture joint point prediction codewords according to the updated and optimized two-dimensional human body posture, the input posture codebook, and the dense posture codebook, compares the input posture joint point codewords and the dense posture joint point codewords, and calculates a loss function for updating the parameters of the autoregressive posture codeword generation model, including: According to the two-dimensional human body posture and the input posture codebook C s and the dense pose codebook C d , get the input gesture to generate the code element p s and dense gesture generation codeword p d ; Calculate the input gesture to generate code element p s With the input gesture symbol q s The cross entropy between the dense pose generation codeword p d With the dense gesture codeword q d The cross entropy between them gives the training loss function: L2=CrossEntropy(p s ,q s )+CrossEntropy(p d ,q d ) According to the training loss function L2, the gradient backpropagation method is used to update the autoregressive gesture codeword generation model parameters and extract features; Repeat the above process until the training loss function converges and obtain the autoregressive gesture codeword generation model.