Sparse human motion capture system based on diffusion probability model and model training method

CN118155286BActive Publication Date: 2026-09-25SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410340330.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2026-09-25
Estimated Expiration
2044-03-25

AI Technical Summary

Technical Problem

综上所述,当前稀疏人体姿态捕捉技术所面临的主要问题有:1、现有方法忽视了周边环境对人体动作的影响;2、现有方法不能很好地重建协调的上、下半身动

Benefits of technology

[0030]本发明的技术方案提出基于扩散概率模型的稀疏人体动作捕捉系统及模型训练方法,全面考虑了环境信息对人体动作的显著影响,鉴于人体动作与周围环境存在高度的关联性,通过引入环境信息来在没有直接下半身动作捕捉信号的情况下,有效地生成更加合理和真实的人体动作。此外,还重点关注人体上下半身运动的相互关联,并进行了系统的建模,通过运用周期动作特征提取器,能够精确提取上半身传感信号的运动特征以及相应生成的下半身动作特征。同时,通过应用正则化方法,减少了上下半身运动特征的特征向量差异,从而确保了生成的人体动作在上下半身之间保持协调一致。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118155286B_ABST
    Figure CN118155286B_ABST
Patent Text Reader

Abstract

The application discloses a sparse human motion capture system based on a diffusion probability model and a model training method, wherein the correlation between a scene and human motion is considered, scene information is introduced, and the quality of sparse human motion capture is improved. In addition, the application additionally introduces modeling of the correlation between upper body motion and lower body motion, so that the generated human motion is more realistic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, specifically to a sparse human motion capture system and model training method based on a diffusion probability model, which combines scene information and sparse tracking signals to achieve full-body human motion capture. Background Technology

[0002] With the rapid development of AR / VR technology, the demand for highly accurate and realistic full-body motion capture is increasing in applications such as virtual meetings, games, and mixed reality. Sparse motion capture technology plays a crucial role in this. This technology primarily uses 6-DOF tracking signals (including spatial rotation and translation) provided by head-mounted devices and controllers to reconstruct full-body motion. However, reconstructing full-body motion using these sparse sensor signals presents a challenging underconstrained one-to-many mapping problem. Currently, most solutions employ a data-driven approach, using large-scale motion capture datasets to address the ambiguity of this one-to-many problem.

[0003] However, these technologies face several challenges and limitations in practice. First, existing sparse human pose capture technologies often neglect the influence of the surrounding environment on human movements. In the real world, an individual's movements are not only a reflection of their own physiological structure but are also greatly influenced by the surrounding environment. For example, a person's movements in a crowded indoor environment will differ significantly from their movements in an open space. This neglect of environmental factors may lead to captured movements that do not match the actual situation, thereby reducing the realism of the virtual reality experience.

[0004] Secondly, existing motion capture methods have shortcomings in reconstructing coordinated upper and lower body movements. Since the sensor signals primarily originate from the wearer's upper body, this leads to a lack of coordination between upper and lower body movements during motion reconstruction. This incoordination not only affects the naturalness of the movements but may also cause deviations in the user's body perception within the virtual environment, thus impacting the overall interactive experience.

[0005] Therefore, to better adapt to the development of AR / VR technology and provide a more realistic user experience, sparse human motion capture technology needs to focus on and solve these problems in future research. This includes more accurately simulating and reproducing individual movements in different environments, as well as improving the coordination of upper and lower body movements. Through these improvements, we can expect to obtain a more natural and immersive experience in future virtual reality applications. In summary, the main problems currently faced by sparse human posture capture technology are: 1. Existing methods ignore the influence of the surrounding environment on human movements; 2. Existing methods cannot reconstruct coordinated upper and lower body movements well. Summary of the Invention

[0006] The purpose of this invention is to consider the influence of the surrounding environment on human movements and to model the coordination relationship between the upper and lower body in human movements, thereby generating more realistic human movements.

[0007] To solve the above-mentioned technical problems, the technical solution of the present invention is to provide a sparse human motion capture system based on a diffusion probability model, which is implemented based on the following steps:

[0008] During the initialization phase, given a sparse motion capture signal p and a scene point cloud S, a random Gaussian white noise latent vector z° ~ N(0,1) is sampled from a standard normal distribution.

[0009] For a sparse motion capture signal p and a random Gaussian white noise latent vector z, an initial human motion prior model is used to obtain the initial human motion. As initial values ​​for the diffusion probability model;

[0010] Periodic motion features f are extracted from the sparse motion capture signal p using a periodic motion feature extractor. Based on the positional information provided by the sparse motion capture signal p, scene geometric point cloud within a user-centered cube is extracted. The scene geometric point cloud is then encoded using the PointNet++ scene encoding network to obtain the scene embedding vector E. S ;

[0011] In the t-th round of back-diffusion in the diffusion probability model, based on the noisy human motion x obtained in the previous round... t+1 Initialize t = T. Where T is the preset number of back-diffusion steps, based on the sparse motion capture signal p, the periodic motion feature f, and the scene embedding vector E. S Noiseless human motion is generated using a diffusion probability model.

[0012]

[0013] In the formula, G represents the diffusion probability generation model; for the generated human actions... Regularization makes human movements more realistic and coordinated;

[0014]

[0015] In the formula, η is the learning rate. For parameters The gradient operator, where l is the human motion regularization term;

[0016] Regularized human movements Add random Gaussian noise for the next round of back diffusion:

[0017]

[0018] In the formula, Let be a predefined variance setting for the t-th round, and ∈ be Gaussian noise from random sampling;

[0019] Repeat the iteration, setting t = t-1 for each iteration, until t = 0. After the iteration ends, the final human motion x0 is output as the model.

[0020] Preferably, the sparse motion capture signal p is acquired by a wearable device.

[0021] The technical solution of this invention also provides a model training method for a sparse human motion capture system based on a diffusion probability model. The method employs the aforementioned sparse human motion capture system based on a diffusion probability model and includes the following steps:

[0022] A set of motion data x is randomly sampled from the human motion dataset, and the latent embedding vector z is obtained using the encoder of the variational autoencoder.

[0023] For the latent embedding vector z, the original action data is reconstructed using the decoder of the variational autoencoder. The variational autoencoder is trained using stochastic gradient descent until the specified number of training iterations is reached, and the model parameters θ are updated using the following formula;

[0024]

[0025] Where, θ + Here are the updated model parameters, and η is the learning rate. For the gradient operator with respect to the parameter θ, Represented as ||·|| is the L2 norm, and KL(·,z0) is the KL loss function, where z0 represents the standard normal distribution;

[0026] A set of action data x is randomly sampled from the human motion dataset, and the embedding vector E corresponding to the scene geometry within a 2×2×2 cube near the human motion trajectory is obtained. S The periodic motion feature f corresponding to the action;

[0027] Randomly sample backdiffusion rounds t, and add noise to the motion data x according to the variance setting to obtain noisy motion x. t The original action data was reconstructed using a probability diffusion model. The probability diffusion model was trained using stochastic gradient descent until the specified number of training iterations was reached, and the model parameters... Update using the following formula:

[0028]

[0029] in, These are the updated model parameters, where α is the learning rate. For parameters gradient operator, Represented as Where FK(·,·) represents the forward dynamic loss function.

[0030] This invention proposes a sparse human motion capture system and model training method based on a diffusion probability model. It comprehensively considers the significant impact of environmental information on human motion. Given the high correlation between human motion and the surrounding environment, environmental information is introduced to effectively generate more reasonable and realistic human motion even without direct lower body motion capture signals. Furthermore, it focuses on the interrelationship between upper and lower body movements and conducts systematic modeling. By using a periodic motion feature extractor, it can accurately extract the motion features of upper body sensor signals and the corresponding generated lower body motion features. Simultaneously, by applying regularization methods, the feature vector differences between upper and lower body motion features are reduced, thereby ensuring that the generated human motion remains coordinated between the upper and lower body. Attached Figure Description

[0031] Figure 1 This illustrates a sparse human motion capture framework based on a diffusion probability model.

[0032] Figure 2 This illustrates the framework of a priori model for human motion.

[0033] Figure 3 This is a flowchart of sparse human motion capture based on a diffusion probability model. Detailed Implementation

[0034] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0035] The sparse human motion capture system based on a diffusion probability model disclosed in this invention includes a human motion prior model, a periodic motion feature extractor, and a regularization process for generating realistic and coordinated upper and lower body movements. Specifically, it includes the following steps:

[0036] During the initialization phase, given a sparse motion capture signal p (returned by the wearable device) and a scene point cloud S, a random Gaussian white noise latent vector z° ~ N(0,1) is sampled from a standard normal distribution.

[0037] For a sparse motion capture signal p and a random Gaussian white noise latent vector z, an initial human motion prior model is used to obtain the initial human motion. Used as initial values ​​for the diffusion probability model.

[0038] A periodic motion feature extractor is used to extract the periodic motion features f = g(p) from the sparse motion capture signal p. The extraction process first compresses the time-domain motion signal into the latent space through one-dimensional convolution and then performs a Fourier transform to convert it to the frequency domain signal. The time-domain motion feature signal is then reconstructed based on the extracted frequency, amplitude, phase difference, and offset.

[0039] Based on the position information provided by the sparse motion capture signal p, the scene geometric point cloud within a 2×2×2 cube centered on the user is extracted. The scene geometric point cloud is then encoded using the PointNet++ scene encoding network to obtain the scene embedding vector E. S .

[0040] In the t-th round of back-diffusion in the diffusion probability model, based on the noisy human motion x obtained in the previous round... t+1 Initialize t = T. Where T is the preset number of back-diffusion steps, based on the sparse motion capture signal p, the periodic motion feature f, and the scene embedding vector E. S Noiseless human motion is generated using a diffusion probability model.

[0041]

[0042] In the formula, G represents the diffusion probability generation model. Then, the generated human motion... Regularization makes human movements more realistic and coordinated.

[0043]

[0044] In the formula, η is the learning rate. For parameters The gradient operator is l, which is the human motion regularization term. The regularization term includes: scene clipping regularization term, and upper and lower body coordination regularization term.

[0045] Finally, the regularized human movements Add random Gaussian noise for the next round of back diffusion:

[0046]

[0047] In the formula, Let be the predefined variance setting for the t-th round, and ∈ be the Gaussian noise from random sampling.

[0048] Repeat the iteration, setting t = t-1 for each iteration, until t = 0. After the iteration is complete, output the final human motion x0 as the model.

[0049] Another embodiment of the present invention provides a model training method for a sparse human motion capture system based on a diffusion probability model, comprising the following steps:

[0050] First, a set of motion data x is randomly sampled from the human motion dataset, and the latent embedding vector z is obtained using the encoder of the variational autoencoder.

[0051] For the latent embedding vector z, the original action data is reconstructed using the decoder of the variational autoencoder. The variational autoencoder is trained using stochastic gradient descent, and the model parameters θ are updated using the following formula:

[0052]

[0053] Where, θ + Here are the updated model parameters, and η is the learning rate. For the gradient operator with respect to the parameter θ, Represented as ||·|| is the L2 norm, and KL(·,z0) is the KL loss function, where z0 represents the standard normal distribution.

[0054] Repeat the above steps until the specified number of training iterations is reached.

[0055] A set of action data x is randomly sampled from the human motion dataset, and the embedding vector E corresponding to the scene geometry within a 2×2×2 cube near the human motion trajectory is obtained. S The periodic motion feature f corresponding to the action.

[0056] Randomly sample backdiffusion rounds t, and add noise to the motion data x according to the variance setting to obtain noisy motion x. t The original action data was reconstructed using a probability diffusion model. The probability diffusion model was trained using stochastic gradient descent, and the model parameters... Update using the following formula:

[0057]

[0058] in, These are the updated model parameters, where α is the learning rate. For parameters gradient operator, Represented as Where FK(·,·) represents the forward dynamic loss function.

[0059] Repeat the above steps until the specified number of training iterations is reached.

[0060] We call the model that obtains full-body human motion through the above steps, combining sparse human sensor data and scene geometry, S2Fusion. We then compared it with some popular sparse human motion capture models on the CIRCLE and GIMO human motion capture datasets. The experimental results are shown in Tables 1 and 2. As can be seen from the tables, S2Fusion outperforms other models in all cases.

[0061] Table 1: Comparison of S2Fusion with other models on the GIMO dataset

[0062]

[0063] Table 2: Comparison of S2Fusion with other models on the CIRCLE dataset

[0064]

[0065] This invention comprehensively considers the significant impact of environmental information on human movement. Given the high correlation between human movement and the surrounding environment, this invention introduces environmental information to effectively generate more reasonable and realistic human movements even without direct lower body motion capture signals. Furthermore, this invention focuses on the interrelationship between upper and lower body movements and performs systematic modeling. By employing a periodic motion feature extractor, this invention can accurately extract the motion features of upper body sensor signals and the corresponding generated lower body motion features. Simultaneously, by applying regularization methods, the feature vector differences between upper and lower body motion features are reduced, thereby ensuring that the generated human movements remain coordinated and consistent between the upper and lower body.

[0066] To further improve the generation speed of motion capture models and enhance the stability and robustness of human motion generation, this invention also introduces an innovative human motion prior model. This model, based on a variational autoencoder architecture, can provide reasonable initial values ​​for human motion to the diffusion probability model, thereby improving overall performance.

[0067] Compared with the prior art, the present invention has the following significant advantages:

[0068] (1) This invention is the first to explore the important role of scene information in sparse human motion capture, which effectively improves the generation quality of human motion data;

[0069] (2) This invention proposes a novel sparse human motion capture framework based on a diffusion probability model. This framework combines a human motion prior model, a periodic motion feature extractor, and a regularization process that can generate realistic and coordinated upper and lower body movements.

[0070] (3) This invention provides an efficient and easy-to-implement solution for sparse motion capture based on wearable devices, and demonstrates superior performance over other sparse motion capture methods in this field.

Claims

1. A sparse human motion capture system based on a diffusion probability model, characterized in that, This can be achieved based on the following steps: During the initialization phase, given a sparse motion capture signal p and a scene point cloud S, a random Gaussian white noise latent vector z° ~ N(0,1) is sampled from a standard normal distribution. For a sparse motion capture signal p and a random Gaussian white noise latent vector z, an initial human motion prior model is used to obtain the initial human motion. As initial values ​​for the diffusion probability model; Periodic motion features f are extracted from the sparse motion capture signal p using a periodic motion feature extractor. Based on the positional information provided by the sparse motion capture signal p, scene geometric point cloud within a user-centered cube is extracted. The scene geometric point cloud is then encoded using the PointNet++ scene encoding network to obtain the scene embedding vector E. S ; In the t-th round of back-diffusion in the diffusion probability model, based on the noisy human motion x obtained in the previous round... t+1 Initialize t = T. Where T is the preset number of back-diffusion steps, based on the sparse motion capture signal p, the periodic motion feature f, and the scene embedding vector E. S Noiseless human motion is generated using a diffusion probability model. In the formula, G represents the diffusion probability generation model; for the generated human actions... Regularization makes human movements more realistic and coordinated; In the formula, η is the learning rate. For parameters The gradient operator, where l is the human motion regularization term; For regularized human movements Add random Gaussian noise for the next round of back diffusion: In the formula, Let be a predefined variance setting for the t-th round, and ∈ be Gaussian noise from random sampling; Repeat the iteration, setting t = t-1 for each iteration, until t = 0. After the iteration ends, the final human motion x0 is output as the model.

2. The sparse human motion capture system based on a diffusion probability model as described in claim 1, characterized in that, The sparse motion capture signal p is acquired by the wearable device.

3. A model training method for a sparse human motion capture system based on a diffusion probability model, characterized in that, The sparse human motion capture system based on the diffusion probability model as described in claim 1 includes the following steps: A set of motion data x is randomly sampled from the human motion dataset, and the latent embedding vector z is obtained using the encoder of the variational autoencoder. For the latent embedding vector z, the original action data is reconstructed using the decoder of the variational autoencoder. The variational autoencoder is trained using stochastic gradient descent until the specified number of training iterations is reached, and the model parameters θ are updated using the following formula; Where, θ + Here are the updated model parameters, and η is the learning rate. For the gradient operator with respect to the parameter θ, Represented as ||·|| is the L2 norm, and KL(·,z0) is the KL loss function, where z0 represents the standard normal distribution; A set of action data x is randomly sampled from the human motion dataset, and the embedding vector E corresponding to the scene geometry within a 2×2×2 cube near the human motion trajectory is obtained. s The periodic motion feature f corresponding to the action; Randomly sample backdiffusion rounds t, and add noise to the motion data x according to the variance setting to obtain noisy motion x. t The original action data was reconstructed using a probability diffusion model. The probability diffusion model was trained using stochastic gradient descent until the specified number of training iterations was reached, and the model parameters... Update using the following formula: in, These are the updated model parameters, where α is the learning rate. For parameters gradient operator, Represented as Where FK(·,·) represents the forward dynamic loss function.