Three-dimensional human body posture estimation two-channel diffusion method based on noise self-adaption
Through the noise adaptive dual-channel diffusion method, combined with two-dimensional noise estimation and geometric reconstruction technology, the problem of joint node ambiguity in binocular three-dimensional human posture estimation is solved, and the accuracy and stability of three-dimensional posture estimation are improved.
Patent Information
- Application Number
- CN202510591137.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-01
AI Technical Summary
The existing binocular three-dimensional human posture estimation method has high ambiguity in the three-dimensional reconstruction area of the joint node, and diffusion model training requires modeling the uncertainty distribution of the initial three-dimensional posture, which is unknown, resulting in insufficient estimation accuracy.
The noise adaptive dual-channel diffusion method is adopted to introduce a two-dimensional point noise estimation network to adaptively estimate the noise interference degree of each two-dimensional point, and combined with geometric reconstruction technology, noise is gradually added and removed to achieve adaptive denoising of two-dimensional estimation points with different accuracy.
The accuracy of estimation of three-dimensional human postures is improved, the uncertainty areas of three-dimensional joint nodes are reduced, the shortcomings of statistical methods are avoided, and the influence of different perspective angles and baseline widths is adapted to the impact of estimation, which is enhanced.
Smart Images

Figure CN120410918A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically, relates to a dual-channel diffusion method and system for three-dimensional human pose estimation based on noise adaption. Background Art
[0002] Three-dimensional human pose estimation aims to locate the three-dimensional coordinates of human joints, which is widely used and important. So far, monocular three-dimensional human pose estimation has attracted much attention due to its convenience in practical applications, while multi-view (more than two cameras) three-dimensional human pose estimation is popular because it can achieve absolute coordinate positioning through geometric constraints. However, both of these methods have their own advantages and disadvantages: monocular estimation has the problem of depth ambiguity, while multi-view estimation is limited by strict acquisition scene requirements. The binocular setting combines the advantages of both, but has long been ignored by researchers.
[0003] Due to the existence of geometric constraints, the framework of binocular three-dimensional pose estimation is similar to that of general multi-view three-dimensional pose estimation, and usually consists of two steps: 1) Estimate the two-dimensional human poses in binocular images respectively; 2) Use the geometric reconstruction triangulation method to reconstruct the corresponding three-dimensional joint points through the binocular two-dimensional estimated points. By observing the uncertainty regions of the three-dimensional points reconstructed under different numbers of cameras, it can be found that although the binocular setting significantly reduces the uncertainty of three-dimensional reconstruction compared with the monocular setting, there is still a relatively high ambiguity compared with the multi-view setting.
[0004] The patent document "A Human Pose Positioning Method, System and Storage Medium Based on Binocular Vision" (CN113850865A) discloses using the human key points extracted by each of the binocular cameras, reconstructing the three-dimensional key point coordinates using the triangulation principle, predicting the possibly missing key points using Kalman filtering, and obtaining accurate information such as the accurate position of the key points and the accurate distance of the movement of the key points between each frame of images or within a period of time using this three-dimensional coordinate. Although the binocular camera setting significantly reduces the scene requirements and costs compared with multi-view, it does not solve the problem of higher ambiguity in the three-dimensional reconstruction region of joint points.
[0005] In previous monocular three-dimensional pose estimation studies, in order to reduce ambiguity, prior knowledge of human poses is usually introduced, such as joint angle limits and physical rationality. With the development of generative models in machine learning, many studies have been dedicated to using a large amount of data to model the real human pose distribution as a representation method of pose priors, including variational autoencoders, generative adversarial networks, normalizing flows, and diffusion models. Among them, diffusion models have attracted much attention due to their advantages such as simple training and flexible networks, which provides a new idea for reducing the ambiguity of the uncertainty region in binocular three-dimensional pose estimation.
[0006] The diffusion model consists of two processes: the forward diffusion process, which adds noise to perturb the real data distribution into a diffusion distribution; and the reverse denoising process, which denoises the noisy data to match the real distribution. In monocular 3D pose estimation, the diffusion model is used to recover the 3D pose from noisy data sampled from random noise or an initial 3D pose distribution with high uncertainty. The first type of method is excluded because it takes too much time to fail to narrow the search space. The second type of method starts from an initial 3D pose with high uncertainty (generated by existing 3D pose estimation methods) and converges to a more reasonable and accurate low-uncertainty pose distribution through the reverse denoising process of the diffusion model, with higher efficiency. However, the main problem is that the training of the diffusion model requires modeling the uncertainty distribution of the initial 3D pose, and this distribution is unknown. This problem can be solved by statistical methods, but this does not meet the actual application requirements because the 3D uncertainty distribution is closely related to the model ability, training dataset, etc. Once the test object does not conform to the distribution of the training dataset, the statistical result is not applicable.
[0007] However, different from the monocular setting, in the geometric framework, the error of 3D points originates from the error of binocular 2D estimated points. In other words, the 3D uncertainty region can be reconstructed based on the 2D uncertainty region because they are essentially interconnected by geometric constraints. Benefiting from mature pose estimation algorithms and large datasets, the uncertainty distribution of 2D estimation is relatively stable. Although the 2D estimation results are relatively stable, they are still affected by various factors. Occlusion is one of the main factors affecting the 2D estimation accuracy. Occlusion will cause the loss of the observed information of the joint points, thus increasing the uncertainty of the 2D estimation results of this joint. In addition, as the limb level (from the root node to the limb end node) increases, the uncertainty distribution of the 2D estimation of the joint points also gradually increases, which is mainly due to the small size and flexible movement of the end joint points, making it difficult to detect.
[0008] The diffusion models previously applied to 3D human pose estimation usually assume that a pose shares the same noise interference level, ignoring the differences in the uncertainty distributions of 2D estimations of different joint points. Therefore, a diffusion model-based method specifically for binocular 3D pose estimation, called dual-channel diffusion, is proposed, which can denoise the initial 2D and 3D uncertainty poses simultaneously.
[0009] The patent document "A 3D Reconstruction Method and Device Based on Binocular Vision and Diffusion Model" (CN116503553A) discloses reconstructing the 3D model of an object through a small number of images and optimizing it through a diffusion model to obtain a reconstruction effect with low cost, good effect, and high robustness. Although it also uses a diffusion model for optimization, its goal is to perform surface reconstruction of the 3D point cloud in the image acquisition scene, ignoring the problem that the uncertainty of the initial result is not accurately modeled and that there are different levels of uncertainty in different regions.
[0010] For this purpose, a two-channel diffusion method for noise-adaptive three-dimensional human pose estimation is proposed. Before the two-channel diffusion model, a two-dimensional point noise estimation network is introduced to adaptively estimate the noise interference degree of each two-dimensional point by using the continuity of temporal two-dimensional points, so as to realize adaptive denoising of two-dimensional estimation points with different precisions. Summary of the Invention
[0011] Aiming at the defects in the prior art, the purpose of the present invention is to provide a two-channel diffusion method for three-dimensional human pose estimation based on noise adaptability.
[0012] A method for training a two-channel diffusion system for three-dimensional human pose estimation based on noise adaptability according to the present invention includes:
[0013] Initializing noise step: Collecting the original image and two-dimensional pose u0, and adding noise to obtain the noisy two-dimensional pose u t ;
[0014] Noise estimation step: Estimating the noise of the noisy two-dimensional pose u t by the noise estimation network and superimposing it on the two-dimensional pose u0 to generate the intermediate two-dimensional pose u t ′;
[0015] Denosing projection step: Reconstructing the noisy three-dimensional pose y t and inputting it into the denoising network to obtain the true three-dimensional pose projected onto the two-dimensional plane to obtain the final two-dimensional pose
[0016] Preferably, the training objective is set to constrain the two-dimensional pose and the final two-dimensional pose to be consistent, and the training objective of the noise estimation network follows the L t loss:
[0017] L t =|t - t′|
[0018] The training objective of the denoising network follows the L θ loss:
[0019]
[0020] where t represents a random noise timestamp;
[0021] t′ represents the noise level of the joint points in u t ;
[0022] v represents the viewing angle;
[0023] Rpj represents reprojection;
[0024] f θ represents the denoising network;
[0025] E represents expectation.
[0026] Preferably, in the initialization noise step, it is collected through a public dataset or a pose acquisition device.
[0027] Randomly select a timestamp t ∈ [1, T], add noise to generate a noisy two-dimensional pose u t ;
[0028] Where T represents the total number of steps required to add noise to the ground truth data to the target noise distribution.
[0029] Preferably, in the noise estimation step, noises with different interference levels are generated for different joint points within the same pose. According to the previous two-dimensional pose estimation result and the current noisy two-dimensional pose, the joint point-level noise is estimated and added to the two-dimensional pose to generate an intermediate two-dimensional pose.
[0030] The perturbation level of the noise is determined according to the timestamp t and the root node z coordinate value of the noisy three-dimensional pose.
[0031] Preferably, the reconstruction is a triangulation method for geometrically reconstructing the noisy three-dimensional pose y t .
[0032] In the denoising projection step, the noisy three-dimensional pose y t is transformed into a root node-relative form through a baseline-normalization operation and input into a denoising network to generate a true three-dimensional pose through denoising
[0033] For the true three-dimensional pose perform a baseline-denormalization operation and project it onto a two-dimensional screen to generate a final two-dimensional pose
[0034] The baseline-normalization operation is to divide the three-dimensional pose by the baseline length scale factor, and the baseline-denormalization is to multiply the three-dimensional pose by the baseline length scale factor.
[0035] According to a dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptability provided by the present invention, trained by using the training method of the dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptability, including: a noise estimation network and a denoising network.
[0036] The two-dimensional pose generates a noisy two-dimensional pose through a noise estimation network, reconstructs a noisy three-dimensional pose, restores the true three-dimensional pose through a denoising network, and projects it onto a two-dimensional plane to form a final two-dimensional pose.
[0037] Preferably, the noise estimation network is a multi-layer perceptron structure, consisting of two fully connected layers and two activation layers.
[0038] Adaptively estimate the confidence of each two-dimensional point by using the continuity constraint of the temporal two-dimensional points.
[0039] For any joint, a noise timestamp t is randomly selected and the corresponding noise is added.
[0040] Using the two-dimensional pose estimation results of the previous five frames and the noisy two-dimensional pose as inputs, estimate the noise level t' of each joint point in the noisy two-dimensional pose.
[0041] Superimpose the noise at the t' level on the two-dimensional pose to generate the intermediate two-dimensional pose u t ′.
[0042] Preferably, the denoising network removes the noise in the noisy three-dimensional pose under different perturbation distributions, including a graph convolutional layer and an attention-graph convolutional layer.
[0043] The joint coordinates of the noisy three-dimensional pose are independently input into the graph convolutional layer, and feature fusion is performed using the defined bone graph structure.
[0044] Input the attention-graph convolutional layer to further extract the pose features.
[0045] Output the true three-dimensional pose through the graph convolutional layer.
[0046] According to a dual-channel diffusion method for three-dimensional human pose estimation based on noise adaption provided by the present invention, it is implemented by using the three-dimensional human pose estimation dual-channel diffusion system based on noise adaption, including:
[0047] Input the estimated noisy two-dimensional pose into the system, and associate the noisy two-dimensional pose with the three-dimensional pose mapping;
[0048] Reconstruct the noisy three-dimensional pose and sample it, denoise to obtain the true three-dimensional pose and re-project it onto the two-dimensional plane to obtain the final two-dimensional pose;
[0049] Iteratively denoise and output the iterated true three-dimensional pose and the final two-dimensional pose.
[0050] Preferably, the iterative denoising includes:
[0051] Use the DDIM strategy to generate the noisy two-dimensional pose u at a new timestamp t by diffusing the final two-dimensional pose t , as the input for the next iterative denoising.
[0052] The timestamps used for K iterations are all t' estimated by the noise estimation network according to the current noisy two-dimensional pose.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] 1. The noise - adaptive dual - channel diffusion model of the present invention learns the inherent prior of human postures through the training process of adding noise and denoising, and uses this prior to specifically narrow the uncertainty region of three - dimensional joint points, thereby improving the accuracy of the final three - dimensional pose estimation.
[0055] 2. By reasonably modeling the distribution of the initial three - dimensional pose, the present invention avoids the problem that the statistical results of statistical methods cannot meet the actual requirements due to changes in training data and models, reduces the influence of factors such as the depth of three - dimensional target points and the baseline width of binocular cameras, and thus narrows the uncertainty range.
[0056] 3. The noise - adding process of the present invention is carried out in the two - dimensional space, and a three - dimensional uncertainty distribution is constructed through a geometric reconstruction method, enabling the construction of the three - dimensional uncertainty region to get rid of the limitation of the limited number of three - dimensional pose datasets. The denoising process is still carried out in the three - dimensional space, so that a prior distribution that is not affected by perspective transformation can be learned.
[0057] 4. The noise of the dual - channel diffusion model of the present invention is at the joint point level, which well solves the problem that the uncertainty ranges of the two - dimensional estimation results of different joint points in different postures are not completely unified, and adaptively denoises them. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Other features, objectives, and advantages of the present invention will become more apparent by reading the detailed description of the non - restrictive embodiments with reference to the following drawings:
[0059] Figure 1 It is a schematic diagram of the framework of a dual - channel diffusion model for three - dimensional human pose estimation based on noise adaptation.
[0060] Figure 2 It is a schematic diagram of the training process of the noise - adaptive dual - channel diffusion model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0062] From a diffusion perspective, the uncertainty distribution of the initial two-dimensional (2D) estimated points can be clearly defined, usually modeled as a Gaussian distribution centered around the true value. However, the uncertainty of the initial three-dimensional (3D) pose remains unknown. From a denoising perspective, considering that the human pose prior is inherently 3D, it is better to recover the pose in 3D space rather than 2D space. The uncertainty region refers to the range within which the estimated points may fall, usually forming a certain error range around the true value point. For 2D points and 3D points, they are defined as 2D uncertainty region and 3D uncertainty region respectively.
[0063] To bridge the gap between diffusion and denoising in different spatial domains, according to the present invention, a dual-channel diffusion method for 3D human pose estimation based on noise adaption is provided, which uses geometric projection technology to associate the 2D plane with the 3D space. In the forward diffusion process, by gradually adding noise, poses are generated from the true 2D pose to conform to the initial pose distribution estimated by the 2D pose estimation algorithm. Subsequently, triangulation measurement technology is used to reconstruct the corresponding 3D pose, thereby defining the uncertainty distribution in 3D space. In the reverse denoising process, the noisy 3D pose is sampled, and a reasonable and accurate 3D pose is recovered through the denoising process.
[0064] Take Figure 1 as an example, in the forward diffusion process, noise is added step by step to the 2D pose true value, and finally, noise data conforming to the 2D pose estimated value distribution is generated. Geometric mapping is used to complete the conversion between 2D poses and 3D poses, connecting the diffusion and denoising processes. In the reverse denoising process, the noisy 3D pose is denoised step by step to restore the original true value.
[0065] In more preferred examples, a common practice in the common diffusion model is that one pose shares a timestamp t. However, considering that there are also different noise interference levels in the 2D estimation results of different joint points within the same pose, in order to more accurately reconstruct the uncertain range of the 3D noisy pose, it is necessary to accurately determine the noise uncertainty range at the joint point level of the 2D pose first. For this purpose, a noise estimation network for 2D points is introduced, which adaptively estimates the confidence of each 2D point by using the continuity constraint of sequential 2D points and reflects it in the degree of noise interference, thus realizing a noise-adaptive dual-channel diffusion model.
[0066] The dual-channel diffusion process of noise-adaptive 3D human pose estimation at two adjacent time instants takes Figure 2For example, starting from the true two-dimensional pose, a noisy two-dimensional pose is generated. Geometric reconstruction is used to reconstruct the noisy three-dimensional pose. The denoising network denoises it to restore the true three-dimensional pose and projects it back to the two-dimensional plane. The training objective is set to constrain the true two-dimensional pose to be consistent with the finally generated two-dimensional pose. It should be noted that the noise is at the joint point level, and different levels of noise interference are generated for different joint points within the same pose. The noise estimation network estimates the joint point-level noise based on the previous two-dimensional pose estimation result and the current noisy two-dimensional pose, and re-adds this noise to the true two-dimensional pose as the input for the subsequent steps.
[0067] The main objective of the denoising network of the dual-channel diffusion model is to remove the noise in the noisy three-dimensional pose under different perturbation distributions. The level of noise perturbation is determined by the timestamp t in the diffusion process and is reflected in the denoising network through timestamp embedding. Although t is directly related to the two-dimensional pose noise, the three-dimensional pose noise depends not only on the two-dimensional noise but also on the depth of the three-dimensional points and the baseline width of the binocular camera. To flexibly process three-dimensional noise data with different uncertainties at the same timestamp t, a depth embedding of the three-dimensional points is additionally introduced in the input of the denoising network, specifically the z coordinate value of the root node of the initial three-dimensional pose, that is, the noisy three-dimensional pose.
[0068] Furthermore, since the proportional relationship between the three-dimensional noise and the baseline width can be determined according to the triangulation calculation formula, a baseline-normalization operation is performed on the three-dimensional pose. The three-dimensional pose is divided by the baseline length scaling factor so that the three-dimensional noisy pose input to the denoising network is not affected by the baseline width, enabling the model to flexibly adapt to different baseline width settings. Correspondingly, baseline-de-normalization is used to multiply the three-dimensional pose by the baseline length scaling factor and apply it to the output of the denoising network to restore the original scale of the pose. The baseline length scaling factor is the length of the baseline. The baseline of the binocular camera refers to the distance between the two cameras, which is solved by the translation matrix and obtained when the camera parameters are acquired. It does not change after the acquisition device is fixed.
[0069] The implementation and application of the noise-adaptive dual-channel diffusion system involve two parts: training and inference.
[0070] Training process: As Figure 2 shown, diffusion data is generated for training and noise is added for noise estimation. The specific steps include:
[0071] Initialization: The original image and the true two-dimensional pose u0 are obtained through a public dataset or a pose acquisition device.
[0072] Adding noise: Randomly select a timestamp t ∈ [1, T], and add noise such as Gaussian noise to generate a noisy two-dimensional pose u t. T represents the total number of steps required to add noise to the true value data to the target noise distribution, preferably set to 25.
[0073] Noise Estimation: It should be noted that the noise is at the joint point level, generating different levels of interference noise for different joint points within the same pose. For any joint, a noise timestamp t is randomly selected, and the corresponding noise is added to this joint. The noise estimation network uses the two-dimensional pose estimation results of the previous five frames and u t as inputs to estimate the noise level t' of each joint point in u t . Then, the noise at the t' level is superimposed on the true value u0 again to generate the subsequent truly input noisy two-dimensional pose u t '.
[0074] The noise estimation network is a multi-layer perceptron structure, preferably three layers, composed of two fully connected layers and two activation layers (ReLU and Sigmoid). The training objective follows L t loss to estimate the true noise level t of each joint point in u t :
[0075] L t = |t - t'|
[0076] Triangulation: Use the triangulation method to reconstruct the noisy three-dimensional pose y t . Then, denoise it through the denoising network to restore the data. Specifically, it includes:
[0077] Denoising Three-Dimensional Pose: The denoising network takes the noisy three-dimensional pose y t as input and denoises it to restore the true three-dimensional pose The three-dimensional pose is converted to the root node-relative form before being processed by the denoising network. Baseline-normalization is applied before the denoising network, and then baseline-denormalization is applied and converted back to the original form.
[0078] Reprojection: Reproject the denoised three-dimensional pose back to the final two-dimensional pose The training objective follows L θ loss to facilitate the denoising network to be able to restore a two-dimensional pose consistent with the original input two-dimensional true value:
[0079]
[0080] where v represents the viewing angle, Rpj is the abbreviation of reprojection, f θ represents the denoising network, and E represents the expectation. The structure of the denoising network is as Figure 2As shown in the lower right part, the joint coordinates of the three-dimensional pose are independently input into the graph convolutional layer. First, the defined bone graph structure is used for basic feature fusion, and then five layers of attention-graph convolutional layers are input for further feature extraction. Finally, a three-dimensional pose is output through a graph convolutional layer.
[0081] Inference process: The binocular noisy two-dimensional pose u estimated by the two-dimensional pose estimation network ViTPose t is used as the input of the noise adaptive two-channel diffusion model. The reconstructed noisy three-dimensional pose y t is normalized and then input into the denoising network to generate a reasonable and accurate three-dimensional pose and their corresponding final binocular two-dimensional poses Then, the DDIM strategy is used to diffuse these poses to generate u at the new timestamp t t , which is used as the input for the next denoising. After K iterations of denoising, the final three-dimensional pose and two-dimensional pose
[0082] In more preferred examples, the timestamp used in each iteration is t′ estimated by the noise estimation network according to the current two-dimensional pose.
[0083] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily.
Claims
1. A training method for a dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation, characterized in that, including: Initialization noise step: Acquire the original image and the 2D pose u0, and add noise to obtain the noisy 2D pose u t ; Noise estimation step: Estimate the noise of the two-dimensional pose u of the noise through a noise estimation network, and superimpose it on the two-dimensional pose u0 to generate an intermediate two-dimensional pose u t ; t′ ; Denoising projection step: Reconstruct the noisy 3D pose y t , and input it into the denoising network to obtain the true 3D pose Project it onto the 2D plane to obtain the final 2D pose 2. The training method of the dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to claim 1, wherein The training objective is set to constrain the two-dimensional pose to be consistent with the final two-dimensional pose, and the training objective of the noise estimation network follows the L t loss: L t = |t - t'| The training objective of the denoising network follows L θ Loss: where t represents a random noise timestamp; t′ represents u t The noise level of the middle joint point; v represents the viewing angle; Rpj represents reprojection; f θ represents a denoising network; E represents expectation.
3. The training method of the dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to claim 1, wherein In the initializing noise step, it is collected through a public dataset or a pose acquisition device; Randomly select a timestamp \(t\in[1,T]\), and add noise to generate a noisy two-dimensional pose \(u\) t ; where T represents the total number of steps required to add noise to the ground truth data to the target noise distribution.
4. The training method of the dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to claim 1, wherein In the noise estimation step, noises with different interference levels are generated for different joint points within the same pose. Based on the previous two-dimensional pose estimation result and the current noisy two-dimensional pose, the joint point-level noise is estimated and added to the two-dimensional pose to generate an intermediate two-dimensional pose; The perturbation level of the noise is determined according to the timestamp t and the z coordinate value of the root node of the noisy three-dimensional pose.
5. The training method of the dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to claim 1, wherein, The reconstruction is a triangulation method for geometric reconstruction of the noise three-dimensional pose y t ; In the denoising projection step, the noisy 3D pose y t is transformed into the root - relative form through the baseline - normalization operation and input into the denoising network to generate the true 3D pose by denoising Perform baseline-denormalization operation on the true three-dimensional pose and project it onto a two-dimensional screen to generate the final two-dimensional pose The baseline-normalization operation is to divide the three-dimensional pose by the baseline length scale factor, and the baseline-denormalization is to multiply the three-dimensional pose by the baseline length scale factor.
6. A dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation, trained by the training method of the dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to any one of claims 1-5, characterized in that including: a noise estimation network and a denoising network; The two-dimensional pose generates a noisy two-dimensional pose through the noise estimation network, reconstructs the noisy three-dimensional pose, restores the true three-dimensional pose through the denoising network, and projects it onto the two-dimensional plane to form the final two-dimensional pose.
7. The dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to claim 6, characterized in that, The noise estimation network is of a multi-layer perceptron structure, consisting of two fully connected layers and two activation layers; adapting to estimate the confidence of each two-dimensional point using the continuity constraint of the temporal two-dimensional points; For any joint, a random noise timestamp t is selected and the corresponding noise is added; Using the two-dimensional pose estimation results of the previous five frames and the noisy two-dimensional pose as inputs, the noise level t′ of each joint point in the noisy two-dimensional pose is estimated; Add noise at the t' level to the two-dimensional pose to generate an intermediate two-dimensional pose u t′ .
8. The dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to claim 6, characterized in that, The denoising network removes the noise in the noisy three-dimensional pose under different perturbation distributions, including a graph convolutional layer and an attention-graph convolutional layer; The joint point coordinates of the noisy three-dimensional pose are independently input into the graph convolutional layer, and feature fusion is performed using the defined bone graph structure; Input into the attention-graph convolutional layer to further extract pose features; Output the true three-dimensional pose through the graph convolutional layer.
9. A dual-channel diffusion method for three-dimensional human pose estimation based on noise adaptation, which is implemented by using the dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation described in any one of claims 7-8, characterized in that including: Input the estimated noisy two-dimensional pose into the system, and associate the noisy two-dimensional pose with the three-dimensional pose mapping; Reconstruct the noisy three-dimensional pose and sample, denoise to obtain the true three-dimensional pose and re-project it onto the two-dimensional plane to obtain the final two-dimensional pose; Iteratively denoise and output the iterated true three-dimensional pose and the final two-dimensional pose.
10. The dual-channel diffusion system for three-dimensional human pose estimation based on noise adaptation according to claim 9, characterized in that, The iterative denoising includes: Use the DDIM strategy to obtain the final two-dimensional pose and diffuse to generate a noisy two-dimensional pose u at the new timestamp t t , which serves as the input for the next iteration of denoising; The timestamps used for K iterations are all t′ estimated by the noise estimation network according to the current noisy two-dimensional pose.
Citation Information
Patent Citations
Three-dimensional human body posture estimation method and system based on video sequence spatio-temporal context
CN118823833A
Method for establishing three-dimensional ultrasound image blind denoising model, and use thereof
WO2022222199A1