3D human body posture estimation method based on diffusion model

By adopting a diffusion model-based method in 3D human pose estimation, combined with DCT and CCMA strategies, the robustness and accuracy of pose estimation in the prior art are solved, and efficient and accurate pose estimation on resource-constrained devices are achieved.

CN120088814AInactive Publication Date: 2025-06-03ZHEJIANG UNIV OF SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510077981.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as pattern collapse, limited expression ability and high computational complexity in 3D human posture estimation, making it difficult to achieve robustness and accuracy in complex scenarios.

Method used

A 3D human pose estimation method based on diffusion model is used to combine discrete cosine transform (DCT) for frequency domain feature extraction, and a multi-assumption aggregation (CCMA) strategy based on confidence and consistency is introduced to reduce the computational complexity and improve the accuracy and robustness of pose estimation.

Benefits of technology

It significantly improves the accuracy and robustness of pose estimation, is suitable for resource-constrained devices, can effectively deal with depth fuzzy and joint occlusion problems, and improves the ability to deal with multi-solving problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088814A_ABST
    Figure CN120088814A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D human body posture estimation method based on a diffusion model. According to the method, firstly, a continuous human body dynamic video is split into RGB images, two-dimensional coordinate information, embedded time steps and positions of all joint points of a human body are extracted from the RGB images through a 2D posture detector, then predicted 3D postures are output through a diffusion model, and low-frequency key information is reserved and high-frequency noise is filtered in combination with discrete cosine transform (DCT), so that the 3D postures are obtained. And the calculation complexity is greatly reduced. And based on a multi-hypothesis aggregation strategy of confidence and consistency evaluation, weighted fusion is performed on the predicted 3D postures, so that the precision and robustness of posture estimation are improved. The performance of the model in a dynamic scene is further enhanced by a time step embedding mechanism and time consistency scoring, so that the model adapts to a complex motion environment and a shielding condition. The method is low in computing resource demand, is suitable for being deployed on resource-constrained equipment in the fields of human-computer interaction, virtual reality, motion analysis, medical rehabilitation and the like, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and relates to a method for 3D human pose estimation in monocular RGB images, and particularly relates to a 3D human pose estimation method based on a diffusion model. Background Art

[0002] The task of pose estimation from real RGB images has a wide range of application scenarios in fields such as virtual reality, augmented reality, human-computer interaction, animation production, and medical rehabilitation.

[0003] In the human pose estimation task, previous methods mainly adopted adversarial generative network (GAN) or autoencoder (AE) techniques. However, GAN is prone to mode collapse, and the expressive ability of AE is limited, making it difficult to generate diverse and realistic 3D poses. These limitations affect the robustness and accuracy of the model in complex scenarios.

[0004] Currently, a new method has emerged that uses a diffusion model to generate diverse 3D human poses. The diffusion model reconstructs data from random noise through a step-by-step denoising process, has stronger generative ability, can effectively alleviate the mode collapse problem of GAN, and improve the robustness and generative diversity of the model. However, the diffusion model also has some drawbacks, such as high computational complexity, resulting in large time consumption, and its high number of parameters and resource requirements limit its application in resource-constrained devices. Summary of the Invention

[0005] The present invention proposes a 3D human pose estimation method based on a diffusion model, which combines discrete cosine transform (DCT) for frequency domain feature extraction and introduces a multi-hypothesis aggregation based on confidence and consistency (CCMA) strategy, significantly reducing the computational complexity while improving the accuracy and robustness of pose estimation, and is suitable for resource-constrained application scenarios.

[0006] A 3D human pose estimation method based on a diffusion model specifically includes the following steps:

[0007] Step 1: Image data processing:

[0008] Divide the real RGB video into frames of images containing human poses, and adjust the RGB images to a unified size.

[0009] Step 2: Extract 2D human pose coordinates:

[0010] Use a 2D pose detector to extract the two-dimensional coordinate information of each joint point of the human body from the RGB image. These two-dimensional coordinates contain the position information of each joint point of the human body, can describe the basic pose structure of the human body, and provide basic data input for subsequent 3D pose estimation.

[0011] Step 3: Pose Estimation Based on Diffusion Model:

[0012] s3.1. Concatenate the two-dimensional coordinate sequence extracted in Step 2 with pure Gaussian noise to form a richer input feature representation hx′:

[0013] x′ = Concat(x, y)

[0014] where F represents the number of RGB images, and J represents the number of human joint points extracted. Input the concatenated feature x′ into a linear transformation layer to map it to a high-dimensional embedding space, providing the input feature E′ for the subsequent Transformer module:

[0015] E′ = Linear(x′)

[0016] where d represents the dimension of the embedded feature. The embedding operation uniformly processes spatial and temporal features, facilitating efficient learning by the subsequent network.

[0017] s3.2. Input the feature E′ into a Transformer-based backbone network. The backbone network includes multiple spatial Transformer modules and temporal Transformer modules. Perform position embedding and time step embedding before the first spatial Transformer module:

[0018] E = E′ + Timestep Embedding + Position Embedding

[0019] where Timestep Embedding and Position Embedding represent the time step embedding information and position embedding information respectively.

[0020] s3.3. The spatial Transformer module uses a spatial attention mechanism to capture the spatial dependencies between joint points in the feature representation of each frame of RGB image, calculates queries, keys, and values, maps the input feature from the time domain to the frequency domain, and uses the discrete cosine transform to reduce the number of parameters:

[0021]

[0022] where D 1 is the discrete cosine transform (DCT) matrix of the spatial attention mechanism. It can concentrate the energy of the input data on the low-frequency components through frequency domain transformation, effectively reducing the interference of high-frequency noise and significantly reducing the computational complexity. The superscript T represents transpose, ⊙ represents element-wise multiplication, M Q , M K , MV is a learnable low-frequency mask that can retain the low-frequency components of the input features and filter out the high-frequency components. Calculate the frequency-domain attention score A fs :

[0023]

[0024] After performing the discrete cosine transform, the frequency-domain result is restored to the spatio-temporal domain to obtain the spatial attention score E fs :

[0025] E fs = IDCT(A fs )

[0026] Use the spatial attention score E fs to update the feature E to E s :

[0027] E s = E * E fs .

[0028] s3.4. The time transformer module performs temporal feature modeling after the spatial transformer module, which can model temporal dynamics across frames and extract the motion trajectory features of the same joint point from the feature representations of different frames:

[0029]

[0030] where D 2 is the discrete cosine transform matrix of the temporal attention mechanism, and calculate the frequency-domain attention score A ft :

[0031]

[0032] After the discrete cosine transform, the frequency-domain result needs to be restored to the spatio-temporal domain to obtain the spatio-temporal domain attention score E ft :

[0033] E ft = IDCT(A ft )

[0034] Use the spatio-temporal domain attention score E ft to update the feature E s to E T :

[0035] E T = E s * E ft

[0036] s3.4. To capture temporal dynamic information, add Timestep Embedding to each layer iteration:

[0037] E i = E T + Timestep Embedding

[0038] And perform recursive processing, using the updated feature E i as the input to the next layer of the spatial transformer module, ensuring that the time dynamic features run through the entire network. After multiple spatial and temporal transformations, the output is the pose after one iteration which is regarded as denoising the pure noise y of the input once. Continue to use y 1 as the input of the backbone network, and after repeating the denoising T times, the final pose y T is obtained.

[0039] Step 4: Generate multiple hypotheses and aggregate them:

[0040] Relying on the generation ability of the model in Step 3, generate a series of possible 3D poses {y 1 , y 2 , …, y H}, where H is the number of hypothesized poses generated for one frame of RGB image.

[0041] s4.1. First, project the generated 3D pose hypotheses {y 1 , y 2 , …, y H} back to the 2D plane and calculate the error between them and the true 2D coordinates Score the confidence of each joint point:

[0042]

[0043] where, is the predicted 2D coordinate of the j-th joint in the h-th hypothesis of the f-th frame of RGB image. is the true 2D coordinate, f ∈ {1, 2, ……, F}, h ∈ {1, 2, ……, H}, j ∈ {1, 2, ……, J}.

[0044] s4.2. Use the Softmax function to convert the error into a confidence score C f,h,j :

[0045]

[0046] s4.3. For each joint, calculate its motion consistency error in the time series:

[0047]

[0048] |||| represents L2 regularization.

[0049] s4.4. Then, use the Softmax function to convert the motion consistency error into a score T f,h,j :

[0050]

[0051] s4.5. Calculate the joint weight ω according to the joint confidence and temporal consistency f,h,j :

[0052] ω f,h,j = α·C f,h,j +(1 - α)·T f,h,j

[0053] where α is a hyperparameter that controls the weights of confidence and temporal consistency. Use the joint weight to perform weighted averaging on all hypotheses to obtain the weighted result of the j-th joint point

[0054]

[0055] Concatenate the weighted results of all joint points to form the final 3D pose

[0056]

[0057] Step 5: Model training and reconstruction

[0058] Use the L1 loss to calculate the loss between the pose estimation result in step 4 and the annotation of the input human pose RGB image. After averaging all the calculated losses, perform backpropagation to optimize the model parameters, save the weights with the best MPJPE metric, and complete the model training.

[0059] Input the RGB image that needs to be estimated for human pose into the trained model to obtain the 3D pose estimation result.

[0060] The present invention has the following beneficial effects:

[0061] The present invention generates multiple hypothesis 3D postures through a diffusion model and combines a multi-hypothesis aggregation method (CCMA) based on confidence and consistency, effectively solving the problems of depth ambiguity and joint occlusion, and significantly improving the accuracy of posture estimation; introducing discrete cosine transform (DCT) for frequency domain processing, reducing the computational complexity and the number of model parameters, making the present invention more suitable for deployment on resource-constrained devices; enhancing the estimation stability in dynamic scenarios through time step embedding and time consistency scoring; at the same time, the diversity generation ability of the diffusion model provides more solution spaces for complex scenarios, improving the processing ability for multi-solution problems. In addition, the present invention is compatible with mainstream 2D posture detection algorithms (such as OpenPose, CPN), can be seamlessly integrated into existing frameworks, and is widely applicable to fields such as human-computer interaction, virtual reality, motion analysis, and medical rehabilitation. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a schematic diagram of the main network for human pose estimation.

[0063] Figure 2 It is a self-attention mechanism diagram based on discrete cosine transform.

[0064] Figure 3 It is a schematic diagram of the time / space transformer.

[0065] Figure 4 It is the visualization result on the dataset Huma3.6M in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The following further explains and illustrates the present invention with reference to the accompanying drawings;

[0067] A 3D human pose estimation method based on a diffusion model, as Figure 1 shown, first extracts 2D joint points from the input RGB image, then performs a 2D to 3D lift, and finally obtains the 3D pose estimation of the person. The specific steps are as follows:

[0068] Step 1: Image data processing:

[0069] A series of human pose RGB images from a dynamic video are uniformly adjusted to a size of 1000*1000 pixels.

[0070] Step 2: Extract 2D human pose coordinates:

[0071] In this embodiment, the CPN model is used to extract the two-dimensional coordinate information of each joint point of the human body from the RGB image. The CPN model includes Backbone, GlobalNet, and RefineNet.

[0072] Among them, the Backbone (main network) usually uses ResNet-101 as the feature extraction network to extract multi-scale features from the input image and generate feature maps of different levels: C2, C3, C4, and C5, which correspond to the outputs of different stages of ResNet respectively. The GlobalNet (global network) is based on the architecture of the Feature Pyramid Network (FPN) to fuse the multi-scale features extracted by the Backbone and generate the initial heatmap of the key points. The feature map of each scale undergoes a 3×3 convolution to adjust the number of channels to match the number of key points. A bottom-up fusion strategy is adopted to combine the high spatial resolution of the low-level features with the rich semantic information of the high-level features. The RefineNet (refinement network) is used to refine the initial heatmap generated by the GlobalNet, paying particular attention to the key points that are difficult to detect. Integrate multi-level features from the GlobalNet, and through upsampling and concatenation operations, improve the spatial resolution of the feature map. Introduce the Online Hard Keypoints Mining (OHKM) strategy to focus on optimizing the key points that are difficult to detect.

[0073] Step 3: Pose estimation based on the diffusion model:

[0074] s3.1. Concatenate the input 2D key point sequence with pure Gaussian noise to form the input feature representation x′:

[0075] x′ = Concat(x, y)

[0076] where F represents the number of RGB images, and J represents the number of human joint points extracted. Input the concatenated feature x′ into the linear transformation layer and map it to E′:

[0077] E′ = Linear(x′)

[0078] where d represents the dimension of the embedded feature.

[0079] s3.2. Input the feature E′ into the backbone network based on the transformer. The backbone network includes d L consecutive spatial transformer modules and a temporal transformer module. Perform position embedding and time step embedding before the first spatial transformer module:

[0080] E = E′ + Timestep Embedding + Position Embedding

[0081] Among them, Timestep Embedding and Position Embedding represent timestep embedding information and position embedding information respectively.

[0082] S3.3, as Figure 2 shown, the spatial transformer module uses a spatial attention mechanism to capture the spatial dependencies between joint points in the feature representation of each frame of RGB image, calculates queries, keys, and values, maps the input features from the time domain to the frequency domain, and adopts a discrete cosine transform to reduce the number of parameters:

[0083]

[0084] where D 1 is the discrete cosine transform (Discrete Cosine Transform) matrix of the spatial attention mechanism. It can concentrate the energy of the input data on the low-frequency components through frequency domain transformation, effectively reducing the interference of high-frequency noise and significantly reducing the computational complexity. The superscript T represents the transpose, ⊙ represents the element-wise multiplication, M Q , M K , M V is a learnable low-frequency mask that can retain the low-frequency components of the input features and filter out the high-frequency components. The frequency domain attention score A fs is calculated as:

[0085]

[0086] After performing the discrete cosine transform again, the frequency domain result is restored to the spatio-temporal domain to obtain the spatial attention score E fs :

[0087] E fs = IDCT(A fs )

[0088] As Figure 3 shown, the feature E is updated to E fs using the spatial attention score E s :

[0089] E s = E * E fs .

[0090] S3.4, the temporal transformer module performs temporal feature modeling after the spatial transformer module, which can model temporal dynamics across frames and extract the motion trajectory features of the same joint point from the feature representations of different frames:

[0091]

[0092]

[0093] Among them, D 2 is the discrete cosine transform matrix of the temporal attention mechanism, which calculates the frequency-domain attention score A ft :

[0094]

[0095] After the discrete cosine transform, the frequency-domain result needs to be restored to the spatio-temporal domain to obtain the spatio-temporal domain attention score E ft :

[0096] E ft = IDCT(A ft )

[0097] Using the spatio-temporal domain attention score E ft to update the feature E s to E T :

[0098] E T = E s * E ft

[0099] s3.4. Perform Timestep Embedding on the output E T of the temporal transformer module:

[0100]

[0101] Take the updated feature as the input of the next spatial transformer module. After d L times of spatial transformation and temporal transformation, the output pose after one iteration is regarded as a denoising of the input pure noise y. Continue to use y 1 as the input of the backbone network. After repeating the denoising T times, the final pose y T is obtained.

[0102] Step 4: Generate multiple hypotheses and aggregate them:

[0103] Relying on the generation ability of the model in step 3, generate a series of possible 3D poses {y 1 , y 2 , …, y H}, where H is the number of hypothesized poses generated for one frame of RGB image.

[0104] s4.1. First, project the generated 3D pose hypotheses {y 1 , y 2 , …, y H} back to the 2D plane and calculate the error between them and the true 2D coordinates Score the confidence of each joint point:

[0105]

[0106] Among them, is the predicted 2D coordinate of the j-th joint in the h-th hypothesis of the f-th frame RGB image. is the true 2D coordinate, f ∈ {1, 2, ……, F}, h ∈ {1, 2, ……, H}, j ∈ {1, 2, ……, J}.

[0107] S4.2. Use the Softmax function to convert the error into the confidence score C f,h,j :

[0108]

[0109] S4.3. For each joint, calculate its motion consistency error in the time series:

[0110]

[0111] |||| represents L2 regularization.

[0112] S4.4. Then use the Softmax function to convert the motion consistency error into the score T f,h,j :

[0113]

[0114] S4.5. Calculate the joint weight ω according to the joint confidence and time consistency f,h,j :

[0115] ω f,h,j = α · C f,h,j + (1 - α) · T f,h,j

[0116] Among them, α is a hyperparameter that controls the weights of confidence and time consistency. Use the joint weight to perform weighted averaging on all hypotheses to obtain the weighted result of the j-th joint point

[0117]

[0118] Concatenate the weighted results of all joint points to form the final 3D pose

[0119]

[0120] Step 5: Model training and reconstruction

[0121] S5.1. Calculate the loss between the pose estimation result in step 4 and the annotation of the input human pose RGB image using the L1 loss, and perform backpropagation after averaging all the calculated losses.

[0122] S5.2. To better observe the training situation of the model, use the wandb tool to observe the decrease in loss. At the same time, add an early stopping mechanism. When the loss has not decreased after the model has been trained for 5 consecutive rounds, trigger the early stopping mechanism to terminate the model training.

[0123] S5.3. Save the weights with the best MPJPE metric to complete the model training. Input the original RGB image into the trained model to complete 3D human pose estimation.

[0124] To verify the beneficial effects of this method, it was compared with mainstream reconstruction methods on the Human3.6M dataset, and the Mean Per Joint Position Error (MPJPE) and Floating-Point Operations Per Second (FLOPs) were used as evaluation metrics. These two evaluation metrics respectively focus on the average Euclidean distance between the predicted joint points and their annotations and the number of floating-point operations that the model can perform within one second. The smaller the data, the higher the accuracy of the reconstruction.

[0125] The experiment was conducted on a server equipped with a single NVIDIA A6000 graphics card, and the training process was optimized based on the PyTorch framework and using the Adam optimizer. The total number of training rounds was 300, the initial learning rate was set to 0.0001, and the learning rate was decreased at the 100th and 120th rounds of training respectively. The overall training took approximately 60 hours. The comparison experiment results are shown in Table 1:

[0126] Table 1

[0127]

[0128]

[0129] The visualized 3D reconstruction results of this method on the Human3.6M dataset are as Figure 4 shown. Pictures with different poses were selected as inputs. It can be seen from the figure that under some relatively complex poses and dynamic actions, the model can still better restore the 3D human pose. Especially in the prediction of joint points, the model shows a high accuracy compared with the ground truth.

Claims

1. A 3D human posture estimation method based on a diffusion model, which splits a dynamic video into a series of RGB images by frame, uses a 2D posture detector to extract the two-dimensional coordinate information of each joint of the human body from the RGB image, and predicts the 3D human posture, characterized by: The following steps are also included: Step 1: Convert the 2D coordinate sequence output by the 2D posture detector With pure Gaussian noise Splice to x′: x′=Concat(x,y) Among them, F represents the number of RGB images, and J represents the number of human joint points extracted; The concatenated feature x′ is mapped to a high-dimensional embedding feature through a linear transformation layer Where d represents the feature dimension after embedding; Step 2: Construct a backbone network based on the spatial transformer module and the temporal transformer module; first embed the feature E′ by position and time step, then capture the spatial dependency between joint points in the spatial transformer module, and use the spatial attention score E fs Update the feature to E s ; Temporal feature modeling is performed in the temporal transformer module, using the attention score E in the spatiotemporal domain ft The feature E s Updated to E T After embedding the time step of the time transformer module, it is input into the next space transformer module. After multiple space transformations and time transformations, the posture after one iteration is output, which is regarded as one denoising. After repeating denoising T times, the final posture y is obtained. T ; Step 3: The backbone network generates H possible 3D poses {y1,y2,…,y H }, project it onto the 2D plane, and compare it with the real 2D coordinate information of each joint point, calculate the confidence and motion consistency error, and use them as weight information to weight the H possible 3D postures to form the final 3D posture Step 4: Calculate 3D pose using L1 loss The loss between the input human posture RGB image annotation is calculated and averaged for back propagation to optimize the model parameters, save the best weight of the MPJPE indicator, and complete the model training; the RGB image that needs to be estimated for human posture is input into the trained model to obtain the 3D posture estimation result.

2. A 3D human body posture estimation method based on a diffusion model as claimed in claim 1, characterized in that: The spatial transformer module uses the spatial attention mechanism to capture the spatial dependencies between joint points in the feature representation of each frame of RGB image, maps the input features from the time domain to the frequency domain, and uses discrete cosine transform to reduce the number of parameters to calculate the query Q fs , key K fs Sum value V fs : Q fs =(ED1 T )⊙M Q +b Q K fs =(ED1 T )⊙M K +b K V fs =(ED1 T )⊙M V +b V Where D1 is the discrete cosine transform matrix of the spatial attention mechanism, the superscript T represents transposition, ⊙ represents bitwise multiplication, and M Q ,M K ,M V is a learnable low-frequency mask; the frequency domain attention score A fs for: After discrete cosine transform, the frequency domain result is restored back to the spatiotemporal domain to obtain the spatial attention score E fs : AND fs =IDCT(A fs ) Using the spatial attention score E fs Update feature E to E s : AND s =E*E fs 。 3. A 3D human body posture estimation method based on a diffusion model as claimed in claim 1, characterized in that: The time transformer module outputs the feature E of the space transformer module. s Perform temporal feature modeling, combine attention mechanism and discrete cosine transform to reduce the number of parameters, and calculate the query Q ft , key K ft Sum value V ft : Q ft =(E s D2 T )⊙M Q +b Q K ft =(E s D2 T )⊙M K +b K V ft =(E s D2 T )⊙M V +b V Where D2 is the discrete cosine transform matrix of the temporal attention mechanism; the frequency domain attention score A ft for: After discrete cosine transform, we get the attention score E in the spatiotemporal domain. ft : AND ft =IDCT(A ft ) Using the attention score E in the spatiotemporal domain ft The feature E s Updated to E T : AND T =And s *AND ft 。 4. A 3D human body posture estimation method based on a diffusion model as claimed in claim 1, characterized in that: The confidence for: in, is the predicted 2D coordinate of the jth joint in the hth hypothesis of the fth frame RGB image; are real 2D coordinates, f∈{1,2,…,F}, h∈{1,2,…,H}, j∈{1,2,…,J}.

5. The 3D human body posture estimation method based on the diffusion model as claimed in claim 1, characterized in that: The motion consistency error for: in, is the predicted 2D coordinate of the jth joint in the hth hypothesis of the fth frame RGB image; are real 2D coordinates, |||| denotes L2 regularization; f∈{1,2,…,F}, h∈{1,2,…,H}, j∈{1,2,…,J}.

6. A 3D human body posture estimation method based on a diffusion model as claimed in any of claims 1, 4 and 5, characterized in that: The weight ω f,h,j for: oh f,h,j =α·C f,h,j +(1-a)·T f,h,j Where α is a hyperparameter that controls the weight of confidence and temporal consistency; C f,h,j represents the joint confidence score, T f,h,j Represents the temporal consistency score by Motion consistency error Using the Softmax function we get: Use the weights to perform weighted averaging on all hypotheses to obtain the weighted result of the jth joint point The weighted results of all joints are stitched together to form the final 3D pose 7. A 3D human posture estimation method based on a diffusion model as claimed in claim 1, characterized in that: Use the wandb tool to observe the decrease in loss, and add an early stopping mechanism. If the model has been trained for five consecutive rounds and the loss has not decreased, the early stopping mechanism is triggered to terminate the model training.

8. The 3D human body posture estimation method based on the diffusion model as claimed in claim 1, characterized in that: The 2D pose detector is CPN or OpenPose.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 5 or 7 and 8.

Citation Information

Cited By

  • Key point prediction method and system based on 2D-IDCT

    CN116895080A

  • A method and system for keypoint prediction based on 2d-idct

    CN116895080B

  • Explicit denoising-driven space-time fusion 3D human body posture estimation method and medium

    CN122049950A

  • Explicit denoising driven spatio-temporal fusion 3D human pose estimation method and medium

    CN122049950B