Human body posture reconstruction and tracking method based on neural network

By using neural networks and Transformer technologies to reconstruct human postures in long-distance scenarios, the problem of insufficient pose reconstruction accuracy and stability in the existing technology is solved, and higher pose reconstruction accuracy and stability are achieved.

CN120233316AInactive Publication Date: 2025-07-01MIRROR VISION (ZHEJIANG) TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510291813.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy and stability of human posture reconstruction in long-distance scenarios, especially because the signal-to-noise ratio of radar signals and the angular resolution are affected by beam broadening effects, resulting in loss of joint features and 'bone retraction' phenomenon.

Method used

Using a neural network-based human posture reconstruction and tracking method, the target radar echo data is obtained through ultra-wideband radar UWB, and combined with wavelet transform noise reduction and adaptive gain compensation, the visibility of long-distance target joint signals is improved. Target pose features are extracted using a Transformer-based feature extraction network, combined with the hybrid memory network HMN and the diffusion model pose completion network DiffusionModel for timing correlation modeling and pose optimization, and finally pose reconstruction is performed through self-supervised adversarial training SSAT pose correction network.

Benefits of technology

It significantly improves the accuracy and stability of long-distance human posture reconstruction, prevents distal joints from retracting to the trunk area, and improves the rationality and accuracy of posture completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233316A_ABST
    Figure CN120233316A_ABST
Patent Text Reader

Abstract

The invention discloses a human body posture reconstruction and tracking method based on a neural network, and relates to the technical field of posture reconstruction and tracking. Target echo data are acquired by adopting a multiple-input multiple-output (MIMO) antenna array, and a low-resolution three-dimensional radar image is constructed by adopting space-time weighting in combination with wavelet transform noise reduction and self-adaptive gain compensation; in the feature extraction stage, a Transform-based feature extraction network is introduced, short-distance joint details are processed in combination with a local convolutional neural network (CNN), the spatial correlation of a far-end joint is calculated by adopting a global self-attention mechanism, and a global feature enhancement (GFE) module performs feature completion on a low-confidence region; a hybrid memory network HMN is adopted to carry out short-time and long-time information fusion, a short-time memory unit STM is used to store current frame features, a long-time memory unit LTM is used to store historical key point data, a drift error is corrected in combination with time sequence alignment, and time sequence consistency of attitude estimation is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pose reconstruction and tracking, and in particular to a human pose reconstruction and tracking method based on a neural network. Background Art

[0002] Human pose reconstruction and tracking technologies are widely used in fields such as intelligent monitoring, computer vision, virtual reality, and telemedicine. Traditional methods mainly rely on optical sensors to accurately capture the positions of key points of the human body in a short-distance, high-resolution environment. However, the performance of these methods is extremely vulnerable to the influence of lighting conditions, target occlusion, and environmental complexity. In long-distance scenarios, such as sports event analysis, remote behavior monitoring, and security monitoring scenarios, the effectiveness of optical sensors drops significantly.

[0003] To overcome the above limitations, most solutions choose to introduce passive sensing technologies such as millimeter-wave radar and ultra-wideband (UWB) radar. Compared with optical methods, radar can work normally under harsh conditions such as night, haze, rain, and snow, and is suitable for more complex environments. However, different from optical sensors, the human pose data obtained by radar is not directly the joint point coordinates, but presented in the form of scattered point clouds, Doppler spectra, or low-resolution radar images. Therefore, deep learning methods are required to map radar echoes to the positions of human key points. However, in long-distance scenarios, >10m, there are still obvious deficiencies in the accuracy and stability of radar-based pose reconstruction.

[0004] The reason why it is difficult to reconstruct the pose of a long-distance target is mainly that the angular resolution of the radar is affected by the beam broadening effect, resulting in a reduced distinguishability between human key points, and the distance information of the target joints becomes blurred in long-distance scenarios. In addition, the signal-to-noise ratio (SNR) of the radar signal decreases as the target distance increases, and smaller joints, such as wrists and ankles, gradually disappear in the signal, eventually leading to the phenomenon of "bone retraction" in pose prediction, that is, the positions of the distal limb joints move closer to the torso, making the human pose show an unnatural "coalescence" trend.

[0005] Upon further analysis, this problem is also limited by the local feature extraction ability of existing deep learning networks. Currently, the mainstream radar-based pose estimation algorithms still take CNN as the core. When dealing with distant human targets, CNN only relies on local feature extraction, resulting in the gradual loss of distal joint features during multi-scale modeling. In addition, in existing research, some methods attempt to use super-resolution algorithms to improve the quality of radar images to enhance the resolution of distant targets, or fuse temporal information to accumulate data from multiple time frames to improve the stability of pose recognition. However, the super-resolution method is prone to introducing artifacts, affecting the stability of pose estimation. Although the temporal method can alleviate short-term pose jitter, during long-term tracking, there will still be pose accumulation errors, affecting the overall reconstruction accuracy. Therefore, how to solve the "bone retraction" problem in the distant scene and improve the accuracy and stability of radar pose reconstruction has become a key research direction. Summary of the Invention

[0006] In view of the above existing problems, the present invention is proposed.

[0007] The present invention provides a method for human pose reconstruction and tracking based on a neural network to solve the problems that the optical method is greatly affected by the environment, the radar method has a decreased accuracy at a long distance, and there are still error accumulation problems in the existing super-resolution and temporal fusion methods.

[0008] To solve the above technical problems, the present invention provides the following technical solutions:

[0009] An embodiment of the present invention provides a method for human pose reconstruction and tracking based on a neural network, which includes,

[0010] Step S1, obtaining distant human pose data, using an ultra-wideband radar UWB to obtain target radar echo data through a multi-input multi-output MIMO antenna array, and filtering the target distance information to form a low-resolution three-dimensional radar image;

[0011] In step S1, the obtained radar echo data is preprocessed, including noise removal and signal gain adjustment, to optimize the joint recognizability of distal targets;

[0012] Step S2, based on the radar image preprocessed in step S1, using a feature extraction network based on Transformer to extract target pose features;

[0013] Step S3, based on the target pose features extracted in step S2, using a hybrid memory network HMN for temporal correlation modeling to predict and generate a pose heatmap;

[0014] Step S4, based on the pose heatmap generated in step S3, using a diffusion model pose completion network DiffusionModel for pose optimization

[0015] Step S5: Based on the optimized pose joint information in Step S4, use the self-supervised adversarial training SSAT pose correction network to reconstruct the pose and output the final human pose skeleton structure.

[0016] As a preferred solution of the human pose reconstruction and tracking method based on neural network according to the present invention, wherein: the preprocessing of the obtained radar echo data, including noise removal and signal gain adjustment, and the step of optimizing the joint recognizability of the distal target are as follows

[0017] Model the target echo signal. After the radar echo signal is received by the antenna array, it is expressed as s raw (t):

[0018]

[0019] wherein, s raw (t) represents the unprocessed radar echo signal, A n represents the echo amplitude of the nth scattering point, f n represents the Doppler frequency of the nth scattering point, φ n represents the initial phase of the nth scattering point, N s represents the total number of scattering points of the target human body, and j is the imaginary unit.

[0020] Perform wavelet transform on the echo signal to remove high-frequency noise and retain the effective signal. The transform formula is:

[0021]

[0022] wherein, represents the radar echo signal after noise reduction, w m represents the wavelet decomposition coefficient of the mth layer, ψ m (t) represents the wavelet basis function corresponding to the mth layer, M w represents the number of layers of wavelet decomposition.

[0023] Aiming at the attenuation of the long-distance target signal, adopt adaptive gain compensation, and the compensation method is:

[0024] s gain (t) = s raw (t) · G(d),

[0025]

[0026] wherein, s gain(t) represents the radar echo signal after gain adjustment, G(d) represents the gain compensation factor related to the target distance d, d represents the target distance, d0 represents the reference distance for gain adjustment, α represents the control parameter for gain adjustment, and e represents the base of the natural logarithm.

[0027] Construct a low-resolution 3D radar image based on the preprocessed radar signal:

[0028]

[0029] Among them, I(x, y, z) represents the 3D radar image, W(x, y, z, t) represents the spatio-temporal weight distribution function, and T s represents the total number of time samplings.

[0030] As a preferred embodiment of the method for human pose reconstruction and tracking based on neural network according to the present invention, wherein: the feature extraction network based on Transformer includes:

[0031] A local convolutional neural network CNN module for extracting local features from the radar image to obtain short-distance detailed features of the target, including hand and foot information;

[0032] A global self-attention module that uses the Transformer structure to calculate the spatial correlation of distal joints, enhance the joint integrity of distant targets, and prevent the loss of target pose information;

[0033] A global feature enhancement module GFE for feature completion of low signal intensity regions, restoring the pose information of distant targets, and generating multi-scale human pose feature representations.

[0034] As a preferred embodiment of the method for human pose reconstruction and tracking based on neural network according to the present invention, wherein: the global self-attention module adopts the global feature enhancement GFE mechanism based on Transformer to calculate the spatial correlation between distal joints.

[0035] As a preferred embodiment of the method for human pose reconstruction and tracking based on neural network according to the present invention, wherein: the hybrid memory network HMN includes:

[0036] A short-term memory unit STM for storing and processing the pose feature information of the current frame and generating the heat map of key points of the current frame;

[0037] A long-term memory unit LTM for storing the position information of historical key points, and performing temporal alignment in combination with the features of the current frame to correct pose drift and cumulative error;

[0038] The hybrid memory network HMN combines a short-term memory unit STM and a long-term memory unit LTM. The short-term memory unit is used for pose prediction of the current frame, and the long-term memory unit is used to store the historical pose trend and correct pose drift through temporal alignment.

[0039] As a preferred embodiment of the method for human pose reconstruction and tracking based on a neural network according to the present invention, wherein: the step of predicting and generating a pose heatmap is as follows.

[0040] Perform pose feature extraction, and the extraction process is expressed as: F t = Transformer(I t ),

[0041] where F t represents the pose feature at time t, and I t represents the radar image input at time t.

[0042] Generate heatmap H t (x,y):

[0043]

[0044] where H t (x,y) represents the pose heatmap generated at time t, with a normalization range of [0,1], x i ,y i represent the coordinates of the i-th key point, K p represents the total number of human key points, and σ i represents the standard deviation of the Gaussian kernel corresponding to the i-th key point.

[0045] Perform temporal correlation modeling, and the modeling formula is:

[0046] M t = LTM(M t-1 , F t ) + STM(F t ),

[0047] where M t represents the temporal feature of the current frame t, M t-1 represents the temporal feature of the previous frame t-1, LTM(·) represents the long-term memory unit for storing long-term information, and STM(·) represents the short-term memory unit for storing the information of the current frame.

[0048] Correct the error, and the correction formula is:

[0049] P t = λM t + (1 - λ)P t-1 ,

[0050] Among them, P t represents the predicted result of the corrected key points, and P t-1 represents the predicted result of the corrected key points in the previous frame, and λ represents the temporal smoothing weight.

[0051] As a preferred solution of the method for human pose reconstruction and tracking based on neural network according to the present invention, wherein: in step S4, the pose optimization includes:

[0052] Completing the low-confidence key points through a probabilistic inference method;

[0053] Combining the kinematic prior PPP based on physical constraints to perform kinematic constraints on human joints, including limb length constraints and rotation angle limitations, to prevent distal joints from retracting into the torso area and optimize the pose rationality;

[0054] Adopting a generative adversarial network GAN to perform adversarial training on pseudo-data at different distances,

[0055] The Diffusion Model pose completion network uses the diffusion sampling method to perform probabilistic inference on low-confidence key points, and combines physical constraints to improve pose rationality and prevent the joint positions of distant targets from shrinking to the torso.

[0056] As a preferred solution of the method for human pose reconstruction and tracking based on neural network according to the present invention, wherein: the steps of using the Diffusion Model pose completion network to optimize the pose are as follows,

[0057] Using the Diffusion Model to complete the low-confidence key points in the pose heatmap. The Diffusion Model first perturbs the input key point heatmap with random noise to generate Gaussian noise data, which is expressed as:

[0058]

[0059] Among them, represents the key point heatmap at the diffusion step q, and H t represents the original key point heatmap, is the noise attenuation coefficient at step q, ∈ q represents the Gaussian noise at step q, represents a Gaussian distribution with a mean of 0 and a covariance matrix of the identity matrix I H of,

[0060] The Diffusion Model learns the reverse denoising process to recover the lost key point information, and this process is expressed as:

[0061]

[0062] Among them, represents the denoised key-point heat map at step size q - 1, and α q represents the noise attenuation coefficient at step size q, and ε θ (·) represents the noise prediction function trained based on a neural network, and σ q represents the noise scaling factor at step size q, and z q represents the Gaussian noise perturbation at step size q, represents a Gaussian distribution with a mean of 0 and a covariance matrix of the identity matrix I H ;

[0063] During the diffusion process, kinematic constraints are imposed on the key-point positions, and the constraint formula is:

[0064]

[0065] Among them, L kin represents the kinematic constraint loss, p j and p j-1 respectively represent the positions of the j-th and (j - 1)-th skeleton joints, K b represents the number of joints in the human skeleton, and L j,j-1 is the constraint length between the j-th and (j - 1)-th joints defined by human anatomy,

[0066] The diffusion model adopts a probabilistic inference method to complete the key points with low confidence, and the completion formula is:

[0067]

[0068] Among them, represents the prediction result of the r-th key point at diffusion step size q, Q r represents the total number of diffusion steps for completion, w q is the confidence-weighted factor at step size q, satisfying

[0069] As a preferred solution of the method for human pose reconstruction and tracking based on a neural network according to the present invention, wherein: the self-supervised adversarial training SSAT pose correction network includes:

[0070] A generative adversarial network GAN correction unit for correcting the pose of low-quality input data;

[0071] A self-supervised adaptive regularization SSAR module for enhancing the adaptability of the model to low signal-to-noise ratio inputs.

[0072] As a preferred solution of the human pose reconstruction and tracking method based on neural network described in the present invention, wherein: the steps of using the self-supervised adversarial training SSAT pose correction network for pose reconstruction and outputting the final human pose skeleton structure are as follows:

[0073] Use the GAN network to correct the pose of the low-quality input data and generate a pose structure. The correction process is expressed as:

[0074]

[0075] Among them, G θ (H t ) represents that the generator predicts the key point pose of the input heat map H t D φ (·) represents the discriminator,

[0076] Add cross-frame consistency constraints:

[0077]

[0078] Among them, L SSAR represents the cross-frame pose consistency loss, and respectively represent the position of the r-th key point at frame t and frame t-1, K b represents the number of observable joints in the human skeleton, T m represents the time window size for calculating consistency;

[0079] The optimized skeleton pose is expressed as:

[0080]

[0081] Among them, P final represents the final human pose skeleton, L GAN is the adversarial loss, L SSAR is the self-supervised regularization loss, L kin is the kinematic constraint loss, λ SSAR , λ kin are weight coefficients.

[0082] ​The beneficial effects of the present invention are as follows: In the present invention, a multi-input multi-output MIMO antenna array is used to obtain target echo data, and wavelet transform denoising and adaptive gain compensation are combined to improve the visibility of joint signals of distant targets. At the same time, spatio-temporal weighting is used to construct a low-resolution three-dimensional radar image to improve the structural integrity of distal targets; in the feature extraction stage, a feature extraction network based on Transformer is introduced, combined with a local convolutional neural network CNN to process short-distance joint details, and a global self-attention mechanism is used to calculate the spatial correlation of distal joints. The global feature enhancement GFE module complements the features in low-confidence regions to improve the stability and integrity of the postures of distant targets; in the temporal modeling stage, a hybrid memory network HMN is used for short-term and long-term information fusion. The current frame features are stored by the short-term memory unit STM, and the historical key point data is stored by the long-term memory unit LTM. The drift error is corrected by temporal alignment to ensure the temporal consistency of pose estimation; further, in the pose optimization stage, a diffusion model is introduced to complement low-confidence key points, and the pose information is repaired by the forward diffusion and reverse denoising processes. Combined with physical kinematic constraints, it prevents distal joints from erroneously retracting to the torso area, improving the rationality and accuracy of pose completion.

[0083] In the present invention, a generative adversarial network GAN is used to perform adversarial training on pseudo-data of targets at different distances to enhance the generalization ability of the model for the postures of distant targets. In the pose reconstruction stage, a self-supervised adversarial training network is used to finally correct the pose. The GAN correction unit is used to optimize low-quality inputs, and the self-supervised adaptive regularization is used to enhance the adaptability of the model to low signal-to-noise ratio environments. At the same time, cross-frame consistency constraints are added to improve the stability of the pose in the time series.

[0084] In summary, the present invention enables the human pose estimation of distant targets to still have high stability, robustness and accuracy in complex environments, significantly improving the practicality of distant human pose reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0086] Figure 1 It is a schematic flow chart of the neural network-based human pose reconstruction and tracking method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0087] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.

[0088] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0089] Secondly, as used herein, "one embodiment" or "an embodiment" refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.

[0090] Example 1, referring to Figure 1 , this example provides a method for human pose reconstruction and tracking based on a neural network, including the following steps:

[0091] Step S1, obtain long-distance human pose data. Use an ultra-wideband radar (UWB) to obtain target radar echo data through a multiple-input multiple-output (MIMO) antenna array, and perform filtering processing on the target distance information to form a low-resolution three-dimensional radar image;

[0092] In step S1, preprocess the obtained radar echo data, including noise removal and signal gain adjustment, to optimize the joint recognizability of the distal target;

[0093] The steps of preprocessing the obtained radar echo data, including noise removal and signal gain adjustment, to optimize the joint recognizability of the distal target are as follows:

[0094] Model the target echo signal. After the radar echo signal is received by the antenna array, it is expressed as s raw (t):

[0095]

[0096] Among them, s raw (t) represents the unprocessed radar echo signal, A n represents the echo amplitude of the nth scatter point, f n represents the Doppler frequency of the nth scatter point, φ n represents the initial phase of the nth scatter point, N s represents the total number of scatter points of the target human body, and j is the imaginary unit.

[0097] Perform wavelet transform on the echo signal to remove high-frequency noise and retain the effective signal. The transformation formula is as follows:

[0098]

[0099] Wherein, represents the radar echo signal after noise reduction, w m represents the wavelet decomposition coefficient of the m-th layer, ψ m (t) represents the wavelet basis function corresponding to the m-th layer, M w represents the number of layers of wavelet decomposition,

[0100] For the attenuation of the long-distance target signal, adopt adaptive gain compensation. The compensation method is as follows:

[0101] s gain (t) = s raw (t) · G(d),

[0102]

[0103] Wherein, s gain (t) represents the radar echo signal after gain adjustment, G(d) represents the gain compensation factor related to the target distance d, d represents the target distance, d0 represents the reference distance for gain adjustment, α represents the control parameter for gain adjustment, and e represents the base of the natural logarithm,

[0104] Construct a low-resolution three-dimensional radar image based on the preprocessed radar signal:

[0105]

[0106] Wherein, I(x, y, z) represents the three-dimensional radar image, and W(x, y, z, t) represents the spatio-temporal weight distribution function, T s represents the total number of time samplings;

[0107] Specifically, in this step, the signal is preprocessed to improve the quality of the radar echo data, making the key joint information of the long-distance target clearer; wavelet transform is used for noise reduction to remove high-frequency interference and prevent joint details from being masked by noise; adaptive gain adjustment improves the visibility of the long-distance joint signal and enhances weak echo signals; non-linear gain compensation is used to compensate for the attenuation of long-distance signals, improve the signal dynamic range, and prevent the loss of distal joint information; at the same time, spatio-temporal weighted reconstruction of the three-dimensional radar image is adopted to improve the integrity of distal targets;

[0108] Step S2, based on the radar image preprocessed in step S1, use a Transformer-based feature extraction network to extract the target pose features;

[0109] The Transformer-based feature extraction network includes:

[0110] A local Convolutional Neural Network (CNN) module for extracting local features from radar images to obtain short-distance detailed features of the target, including hand and foot information;

[0111] A global self-attention module that uses a Transformer structure to calculate the spatial correlation of distal joints, enhance the joint integrity of distant targets, and prevent the loss of target pose information;

[0112] A Global Feature Enhancement (GFE) module for feature completion in low signal strength regions, restoring the pose information of distant targets, and generating multi-scale human pose feature representations;

[0113] The global self-attention module uses a global feature enhancement (GFE) mechanism based on Transformer to calculate the spatial correlation between distal joints;

[0114] Step S3: Based on the target pose features extracted in step S2, a Hybrid Memory Network (HMN) is used for temporal correlation modeling to predict and generate a pose heatmap;

[0115] The Hybrid Memory Network (HMN) includes:

[0116] A Short-Term Memory (STM) unit for storing and processing the pose feature information of the current frame to generate the key point heatmap of the current frame;

[0117] A Long-Term Memory (LTM) unit for storing the position information of historical key points and performing temporal alignment in combination with the features of the current frame to correct pose drift and cumulative errors;

[0118] The Hybrid Memory Network (HMN) combines a Short-Term Memory (STM) unit and a Long-Term Memory (LTM) unit. The short-term memory unit is used for pose prediction of the current frame, and the long-term memory unit is used for storing the historical pose trend and correcting pose drift through temporal alignment;

[0119] The steps for predicting and generating a pose heatmap are as follows:

[0120] Perform pose feature extraction, and the extraction process is expressed as: F t = Transformer(I t ),

[0121] where F t represents the pose feature at time t, and I t represents the radar image input at time t,

[0122] Heatmap generation H t (x, y):

[0123]

[0124] Among them, H t (x, y) represents the pose heatmap generated at time t, with a normalized range of [0, 1], where x i , y i represents the coordinates of the i-th key point, and K p represents the total number of human key points, and σ i represents the standard deviation of the Gaussian kernel corresponding to the i-th key point;

[0125] Perform temporal correlation modeling, and the modeling formula is:

[0126] M t = LTM(M t-1 , F t ) + STM(F t ),

[0127] Among them, M t represents the temporal feature of the current frame t, M t-1 represents the temporal feature of the previous frame t - 1, LTM(·) represents the long-term memory unit that stores long-term information, and STM(·) represents the short-term memory unit that stores the information of the current frame.

[0128] Correct the error, and the correction formula is:

[0129] P t = λM t + (1 - λ)P t-1 ,

[0130] Among them, P t represents the corrected key point prediction result, P t-1 represents the corrected key point prediction result of the previous frame, and λ represents the time smoothing weight;

[0131] Specifically, based on the radar image data here, extract the pose features of the target human body, and use the temporal modeling method to reduce the pose drift and cumulative error of the distal target;

[0132] The Transformer feature extraction network combines local CNN to process short-distance information and avoid detail loss; the global self-attention mechanism strengthens the spatial correlation of distal joints and avoids joint information loss; Gaussian distribution modeling generates key point heatmaps to make pose prediction smooth and reduce key point jitter; the scale-adaptive standard deviation improves the key point localization accuracy of distal targets; significantly improves the temporal stability of distal human poses, avoids target drift, and improves the continuity of human poses across time steps;

[0133] Step S4: Based on the pose heatmap generated in Step S3, use the Diffusion Model pose completion network to optimize the pose.

[0134] In Step S4, the pose optimization includes:

[0135] Complete the low-confidence key points through a probabilistic inference method.

[0136] Combine the kinematic prior PPP based on physical constraints to perform kinematic constraints on human joints, including limb length constraints and rotation angle limits, to prevent distal joints from retracting into the torso area and optimize the pose rationality.

[0137] Use the generative adversarial network GAN to perform adversarial training on pseudo-data at different distances.

[0138] The Diffusion Model pose completion network performs probabilistic inference on low-confidence key points through the diffusion sampling method, and combines physical constraints to improve pose rationality and prevent the joint positions of distant targets from shrinking to the torso.

[0139] The steps of using the Diffusion Model pose completion network to optimize the pose are as follows:

[0140] Use the Diffusion Model to complete the low-confidence key points in the pose heatmap. The Diffusion Model first perturbs the input key point heatmap with random noise to generate Gaussian noise data, which is expressed as:

[0141]

[0142] Among them, represents the key point heatmap at the diffusion step q, H t represents the original key point heatmap, is the noise attenuation coefficient at step q, ∈ q represents the Gaussian noise at step q, represents a Gaussian distribution with a mean of 0 and a covariance matrix of the identity matrix I H of.

[0143] The Diffusion Model learns the reverse denoising process to recover the lost key point information. This process is expressed as:

[0144]

[0145] Among them, represents the denoised key point heatmap at step q - 1, α q represents the noise attenuation coefficient at step q, ε θ(·) represents the noise prediction function trained based on neural network, and σ q represents the noise scaling factor at step q, and z q represents the Gaussian noise perturbation at step q, represents the Gaussian distribution with mean 0 and covariance matrix being the identity matrix I H ;

[0146] During the diffusion process, kinematic constraints are imposed on the positions of key points, and the constraint formula is:

[0147]

[0148] where, L kin represents the kinematic constraint loss, p j and p j-1 respectively represent the positions of the j-th and (j - 1)-th skeleton joints, K b represents the number of joints in the human skeleton, L j,j-1 is the constraint length between the j-th and (j - 1)-th joints defined by human anatomy,

[0149] The diffusion model adopts a probabilistic inference method to complete the key points with low confidence, and the completion formula is:

[0150]

[0151] where, represents the prediction result of the r-th key point at diffusion step q, Q r represents the total number of diffusion steps for completion, w q is the confidence weighting factor at step q, satisfying

[0152] Specifically, the diffusion model Diffusion Model is used to complete the key points with low confidence, and combined with kinematic constraints to improve the pose rationality and prevent the distal joints from incorrectly retracting to the torso area;

[0153] Forward diffusion mixes the key point heatmap with Gaussian noise, so that the joint loss area can be reconstructed through noise learning. Inverse diffusion denoising gradually restores the pose structure to avoid the key points completed from deviating from the actual motion trajectory; The joint length constraint makes the pose still conform to the human biological structure after completion, avoiding the key points from overlapping or deviating too far; The rotation angle limit makes the relative position change between joints reasonable, avoiding exceeding the normal reach range of the human body;

[0154] Adopt a multi-step confidence weighting completion strategy, so that the key points with high confidence are corrected first, and then the key points with low confidence are interpolated and optimized to improve the accuracy of the completion result;

[0155] The attitude integrity of the distal target is significantly enhanced through step S4, effectively filling the information gap in the low-confidence region, making the attitude completion result more in line with the human motion law;

[0156] Step S5: Based on the attitude joint information optimized in step S4, use the self-supervised adversarial training SSAT attitude correction network to reconstruct the attitude and output the final human body attitude skeleton structure;

[0157] The self-supervised adversarial training SSAT attitude correction network includes:

[0158] The generative adversarial network GAN correction unit is used to correct the attitude of low-quality input data;

[0159] The self-supervised adaptive regularization SSAR module is used to enhance the adaptability of the model to low signal-to-noise ratio inputs;

[0160] The steps of using the self-supervised adversarial training SSAT attitude correction network to reconstruct the attitude and output the final human body attitude skeleton structure are as follows:

[0161] Use the GAN network to correct the attitude of low-quality input data and generate an attitude structure. The correction process is expressed as:

[0162]

[0163] Among them, G θ (H t ) represents the key point attitude predicted by the generator for the input heat map H t ; D φ (·) represents the discriminator,

[0164] Add cross-frame consistency constraints:

[0165]

[0166] Among them, L SSAR represents the cross-frame attitude consistency loss, and respectively represent the position of the r-th key point at frame t and frame t-1, K b represents the number of observable joints in the human skeleton, and T m represents the time window size used to calculate the consistency;

[0167] The optimized skeleton attitude is expressed as:

[0168]

[0169] Among them, P final represents the final human body attitude skeleton, L GAN is the adversarial loss, LSSAR is the self-supervised regularization loss, L kin is the kinematic constraint loss, λ SSAR , λ kin is the weight coefficient;

[0170] Specifically, based on the self-supervised adversarial training SSAT here, the complemented key points are further optimized, making the finally generated human body pose skeleton structure more natural and reasonable, and reducing the artifacts brought by data complementation.

[0171] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A human body posture reconstruction and tracking method based on a neural network, characterized in that: include, Step S1, obtaining long-distance human body posture data, using ultra-wideband radar UWB through a multiple-input multiple-output MIMO antenna array to obtain target radar echo data, and filtering the target distance information to form a low-resolution three-dimensional radar image; In step S1, the acquired radar echo data is preprocessed, including noise removal and signal gain adjustment, to optimize the joint recognizability of the remote target; Step S2, based on the radar image preprocessed in step S1, extracting target posture features using a Transformer-based feature extraction network; Step S3, based on the target posture features extracted in step S2, a hybrid memory network HMN is used to perform temporal association modeling to predict and generate a posture heat map; Step S4, based on the posture heat map generated in step S3, the posture optimization is performed using the Diffusion Model posture completion network. Step S5, based on the posture joint information optimized in step S4, the posture is reconstructed by using the self-supervised adversarial training SSAT posture correction network to output the final human posture skeleton structure.

2. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 1, characterized in that: The steps of preprocessing the acquired radar echo data, including noise removal and signal gain adjustment, and optimizing the joint recognizability of the remote target are as follows: Model the target echo signal. After the radar echo signal is received by the antenna array, it is expressed as s raw (t): Among them, s raw (t) represents the unprocessed radar echo signal, A n represents the echo amplitude of the nth scattering point, f n represents the Doppler frequency of the nth scattering point, φ n represents the initial phase of the nth scattering point, N s Represents the total number of scattering points of the target human body, j is an imaginary unit, Perform wavelet transform on the echo signal to remove high-frequency noise and retain the effective signal. The transformation formula is: in, represents the radar echo signal after noise reduction, w m represents the wavelet decomposition coefficient of the mth layer, ψ m (t) represents the wavelet basis function corresponding to the mth layer, M w represents the number of wavelet decomposition layers, In view of the attenuation of the long-distance target signal, adaptive gain compensation is adopted. The compensation method is: s gain (t)=s raw (t)·G(d), Among them, s gain (t) represents the radar echo signal after gain adjustment, G(d) represents the gain compensation factor related to the target distance d, d represents the target distance, d0 represents the reference distance for gain adjustment, α represents the control parameter for gain adjustment, e represents the base of the natural logarithm, Construct a low-resolution 3D radar image based on the preprocessed radar signal: Where I(x,y,z) represents the three-dimensional radar image, W(x,y,z,t) represents the spatiotemporal weight distribution function, and T s Indicates the total number of time samples.

3. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 2, characterized in that: The Transformer-based feature extraction network includes: The local convolutional neural network (CNN) module is used to extract local features from radar images and obtain short-range detail features of the target, including hand and foot information; The global self-attention module uses the Transformer structure to calculate the spatial correlation of distal joints; The global feature enhancement module GFE is used to complete the features of low signal strength areas, restore the posture information of distant targets, and generate multi-scale human posture feature representations.

4. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 3, characterized in that: The global self-attention module adopts the Transformer-based global feature enhancement GFE mechanism to calculate the spatial correlation between distal joints.

5. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 4, characterized in that: The hybrid memory network HMN comprises: The short-term memory unit STM is used to store and process the posture feature information of the current frame and generate the key point heat map of the current frame; Long-term memory unit LTM, used to store the location information of historical key points, and combine the current frame features for time alignment to correct attitude drift and accumulated errors; The hybrid memory network HMN adopts a combination of short-term memory units STM and long-term memory units LTM, wherein the short-term memory unit is used for posture prediction of the current frame, and the long-term memory unit is used to store historical posture trends and correct posture drift through time alignment.

6. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 5, characterized in that: The step of predicting and generating the posture heat map is: Perform posture feature extraction, and the extraction process is expressed as: F t = Transformer(I t ), Among them, F t represents the posture feature at time t, I t represents the radar image input at time t, Heatmap generation H t (x,y): Among them, H t (x, y) represents the posture heat map generated at time t, normalized to the range [0, 1], x i ,y i represents the coordinates of the i-th key point, K p represents the total number of key points of the human body, σ i Represents the Gaussian kernel standard deviation corresponding to the i-th key point; Perform time series correlation modeling, the modeling formula is: M t =LTM(M t-1 ,F t )+STM(F t ), Among them, M t represents the temporal characteristics of the current frame t, M t-1 represents the temporal features of the previous frame t-1, LTM(·) represents the long-term memory unit, which stores long-term information, and STM(·) represents the short-term memory unit, which stores the current frame information. The error is corrected, and the correction formula is: P t =λM t +(1-λ)P t-1 , Among them, P t represents the corrected key point prediction result, P t-1 represents the corrected key point prediction result of the previous frame, and λ represents the temporal smoothing weight.

7. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 6, characterized in that: In step S4, the posture optimization includes: Use probabilistic reasoning methods to complete low-confidence key points; Combined with the motion prior PPP based on physical constraints, kinematic constraints are imposed on human joints, including limb length constraints and rotation angle restrictions; The generative adversarial network (GAN) is used to perform adversarial training on pseudo data of different distances.

8. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 7, characterized in that: The steps of using the diffusion model posture completion network DiffusionModel to perform posture optimization are: The diffusion model DiffusionModel is used to complete the low-confidence key points in the posture heat map. The diffusion model first performs random noise perturbation on the input key point heat map to generate Gaussian noise data, which is expressed as: in, represents the key point heat map under the diffusion step length q, H t represents the original keypoint heat map, is the noise attenuation coefficient at step size q, ∈ q represents Gaussian noise at step size q, Indicates that the mean is 0 and the covariance matrix is ​​the identity matrix I H Gaussian distribution, The diffusion model learns the inverse denoising process to recover the lost key point information. The process is expressed as: in, represents the key point heat map after denoising at step size q-1, α q represents the noise attenuation coefficient at step size q, ∈ θ (·) represents the noise prediction function based on neural network training, σ q represents the noise scaling factor at step size q, z q represents the Gaussian noise perturbation at step size q, Indicates that the mean is 0 and the covariance matrix is ​​the identity matrix I H Gaussian distribution of During the diffusion process, kinematic constraints are imposed on the key point positions. The constraint formula is: Among them, L kin represents the kinematic constraint loss, p j and p j-1 Respectively represent the positions of the j-th and j-1-th skeleton joints, K b represents the number of joints in the human skeleton, L j,j-1 is the constraint length between the jth and j-1th joints defined by human anatomy, The diffusion model uses a probabilistic reasoning method to complete key points with low confidence. The completion formula is: in, Represents the prediction result of the rth key point at the diffusion step length q, Q r represents the total number of diffusion steps used for completion, w q is the confidence weighting factor at step length q, satisfying 9. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 8, characterized in that: The self-supervised adversarial training SSAT posture correction network includes: Generate adversarial network (GAN) correction unit, used to correct the posture of low-quality input data; The self-supervised adaptive regularization SSAR module is used to enhance the model's adaptability to low signal-to-noise ratio inputs.

10. A method for human posture reconstruction and tracking based on a neural network as claimed in claim 9, characterized in that: The steps of adopting self-supervised adversarial training of SSAT posture correction network to perform posture reconstruction and output the final human body posture skeleton structure are as follows: The GAN network is used to correct the posture of low-quality input data and generate the posture structure. The correction process is expressed as: Among them, G θ (H t ) represents the generator's response to the input heat map H t Predicted keypoint pose D φ (·) represents the discriminator, Add cross-frame consistency constraints: Among them, L SSAR represents the cross-frame pose consistency loss, and Respectively represent the position of the rth key point at frame t and frame t-1, K b represents the number of observable joints in the human skeleton, T m Indicates the time window size used to calculate consistency; The optimized skeleton pose is expressed as: Among them, P final represents the final human posture skeleton, L GAN To combat the loss, L SSAR is the self-supervised regularization loss, L kin is the kinematic constraint loss, λ SSAR ,λ kin is the weight coefficient.

Citation Information

Cited By

  • XR scene real-time human body posture tracking method and system

    CN120635151A

  • A method and system for real-time human pose tracking in XR scenes

    CN120635151B

  • Human body shape and posture evaluation system and method based on millimeter wave imaging technology

    CN121774488A

  • Articulated object attitude generation method based on physical perception and graph diffusion

    CN121810667A

  • A jointed object pose generation method based on physical perception and graph diffusion

    CN121810667B