A High Physical Realism Human Motion Capture Method Based on Neural Motion Control

By adopting a neural motion control-based method in human motion capture technology, using sampling distribution prior network and scene contact constraints, the visual ambiguity of human motion capture and the error problems of traditional methods in single-view videos are solved, and a high physical reality human motion capture is achieved.

CN114550292BActive Publication Date: 2025-06-27SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210158059.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2025-06-27
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

The existing human motion capture technology has visual ambiguity in single-view videos, resulting in non-natural phenomena such as mold penetration, jitter and unreasonable foot sliding. The traditional methods have problems such as approximate error, sensitivity to environmental changes and time-consuming optimization of sampling distribution.

Method used

Using a method based on neural motion control, we optimize human body reference motion by proposing sampling distribution prior networks and scene contact constraints, and use the non-differentiated physics engine and sampling control to achieve high physical reality human motion capture.

Benefits of technology

It realizes high physical real-touching human motion capture in complex terrain and multi-variable shapes, avoids physical inreality and random errors in traditional methods, and improves the accuracy and stability of capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550292B_ABST
    Figure CN114550292B_ABST
Patent Text Reader

Abstract

The present invention proposes a high-physical-realism human motion capture method based on neural motion control. First, a sampling distribution prior network based on a physics engine is proposed to train an accurate sampling distribution prior. Secondly, a scene contact constraint is proposed, and a human reference motion is obtained through an optimization framework by inputting a single-view video. Finally, using the trained sampling distribution prior, the sampling distribution is estimated from the human reference motion and the current state of the physical character, and then the sampling control method is used to achieve high-physical-realism human motion capture in a non-differentiable physics engine. The human capture framework based on neural motion control proposed by the present invention uses a physics engine to provide hard physical constraints, avoiding physically unrealistic phenomena such as penetration and jitter in traditional human motion capture. It is convenient to collect, has a low cost, and is easy to implement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and computer graphics; in particular, it relates to a high-physical-realism human motion capture method based on neural motion control. Background Art

[0002] Human motion capture has various applications in aspects such as human-computer interaction, personal health management, and human behavior understanding. With the rapid development of markerless motion capture, the market's requirements for human motion capture are constantly increasing. A large number of existing works can capture accurate human postures kinematically from single-view videos and images through network regression or optimization methods. However, due to visual ambiguity, a series of unnatural phenomena will occur in single-view motion capture based on kinematics, such as penetration, jitter, and unreasonable sliding of footsteps. Some methods add physical constraints in the reconstruction framework through simple and differentiable physical models. However, such methods have large approximation errors; the remaining methods use non-differentiable physical engines and deep reinforcement learning methods for estimation. But training an ideal policy requires a complex configuration process, and the trained policy is very sensitive to environmental changes. In addition, traditional sampling control-based methods require accurate reference motions, and generating a reasonable sampling distribution is very time-consuming and has random errors. Therefore, adopting a sampling-based neural motion control method to achieve human motion capture in the case of complex scene terrains, diverse human body shapes, and various action postures is expected to help solve the pain points in the application of current human motion capture technologies. Summary of the Invention

[0003] To solve the above problems, the present invention discloses a high-physical-realism human motion capture method based on neural motion control. This method first proposes a sampling distribution prior network based on a physical engine and trains an accurate sampling distribution prior. Secondly, a scene contact constraint is proposed, and the input single-view video is used to obtain the human reference motion through an optimization framework. Finally, using the trained sampling distribution prior, the sampling distribution is estimated from the human reference motion and the current state of the physical character, and then the sampling control method is used to achieve high-physical-realism human motion capture in a non-differentiable physical engine.

[0004] The high-physical-realism human motion capture method based on neural motion control described in the present invention includes the following steps:

[0005] S1. Modeling of Human Kinematic Model and Human Dynamic Model: The human kinematic model is a skinnable human skeleton with 51 degrees of freedom of joints, and the skeleton pose is represented by joint rotations. The human dynamic model has the same structure as the human kinematic model, and each joint is equipped with dynamic parameters such as mass and friction, and can perform physical simulations in a non-differentiable physics engine. Given the human bone lengths, the human kinematic model and the human dynamic model can be automatically generated according to the kinematic tree structure. The dynamic parameters of the human dynamic model do not change with the human body shape.

[0006] S2. Definition of Human Reference Motion and Physical Character State: The human reference motion is obtained by fitting the human kinematic model to the human features in the input video. The human reference pose parameters of each frame after fitting form a complete human reference motion, where the human reference pose parameters of the t-th frame are represented as The human reference skeleton pose parameters include 3D model translation, 3D model global rotation, and 51D joint rotation, with a total dimension of 57D. The physical character state is the parameter read from the physics engine after the human dynamic model is simulated, which includes the simulated pose parameter q and the velocity parameter The simulated pose parameter has the same composition as the human reference skeleton pose parameter. The velocity parameter includes 3D root node linear velocity, 3D root node angular velocity, and 51D joint angular velocity, and the total dimension of the physical character state is 114D.

[0007] S3. Training of Sampling Distribution Prior Neural Network: Obtain data from the existing kinematic human dataset, run the covariance matrix adaptation evolution strategy offline to generate a sampling distribution for the pre-training of the supervised distribution encoder. After the pre-training is completed, connect to the two-branch decoder to further optimize and train the encoder network parameters until the network converges. Use the trained encoder network as the sampling distribution prior for subsequent human motion capture.

[0008] S4. Human Reference Motion and Human Body Shape Estimation: Use the existing human two-dimensional pose estimation method to obtain the human two-dimensional pose from the input video. Construct scene contact constraints, optimize the pose parameters of the human kinematic model, and obtain the human reference motion

[0009] S5. High Physical Realistic Human Motion Capture: Read the human reference pose of the t-th frame from the human reference motion Read the state s of the physical character of the (t - 1)-th frame t-1 , input the reference pose and the physical state into the sampling distribution prior, and output the sampling distribution Q(μ t , σ t ). Sample k correction values of the reference pose from the estimated distribution Respectively with the human reference posture Added together to form the human target posture Input k human target postures into the physics engine for simulation, evaluate each simulation result, select the simulation result with the smallest loss function value as the physical character state at time t, and repeat step S5 for motion capture of the next frame. After executing all video frames, the simulated posture parameters q of the physical character in each frame t Form a complete motion as the output.

[0010] Furthermore, the specific method of step S3 includes:

[0011] S31: Training data preparation: Select two consecutive frames from the existing kinematic human dataset: time t-1 and time t. By taking the difference of the selected postures, calculate the human joint velocity at time t-1, and then use the human posture and velocity at time t-1 as the state of the physical character, and use the human posture at the t-th frame as the reference posture. Construct a standard normal distribution as the initial sampling distribution, perform physical simulation on the corresponding physical character state, reference posture, and reference posture correction value sampled from the initial distribution, and use the covariance matrix adaptation strategy to optimize the initial distribution. The optimized distribution Process all the data in the dataset in sequence, and the estimated distribution is the supervised data for pre-training the network parameters; since the result obtained by taking the difference has a deviation from the real simulation, in order to enhance the authenticity of the data, add random noise to the physical character state to simulate the real situation;

[0012] Among them, the physical simulation includes the following steps, and the physical simulation mentioned later is the same as this step:

[0013] S311: Add the reference posture and the reference posture correction value sampled to obtain the target posture;

[0014] S312: Use the target posture as the set value, combine the current physical character state, and calculate the joint torque using PD control;

[0015] S313: Apply the joint torque to the joints of the physical character and execute physical simulation using the physics engine;

[0016] S314: Read the simulation result;

[0017] S32: Pre-training of the distribution encoding network: Construct a 6-layer fully connected encoding network, and the output is the mean and variance of 32 dimensions; form a normal distribution with the estimated mean and variance; use the distribution generated by the offline covariance matrix adaptation strategy as the supervision, and train the network until convergence. Use the Kullback-Leibler divergence as the constraint, and its loss function formula is:

[0018]

[0019] wherein is the distribution obtained by the covariance matrix adaptation evolution strategy, is the distribution estimated by the encoding network, and KL is the Kullback-Leibler divergence.

[0020] S33: After the training converges, due to the random error of the covariance matrix adaptation evolution strategy, there is noise in the generated annotations, and it is difficult for the pre-trained model to regress to the accurate sampling distribution. The present invention proposes a two-branch decoder including a pose decoding network and a physical simulation branch to assist the encoder in training to obtain accurate parameters. Among them, the pose decoding network has 4 fully connected layers to regress the simulated human pose from the target pose; the physical simulation branch also receives the target pose and obtains the simulated pose after simulation. Since the physical simulation branch is not differentiable, the obtained simulated pose is used as supervision and acts on the human pose estimated by the pose decoding network, prompting the pose decoding network to regress to the same result as the physical simulation. Its loss function is:

[0021]

[0022] wherein, is the human pose regressed by the pose decoding network, and j t are respectively the joint positions obtained by forward kinematics of the regressed human pose and the simulated human pose, and ‖·‖ 2 is the two-norm.

[0023] In addition, in order to make the encoder regress to the correct distribution, the reconstruction constraint is further used to constrain the estimated human pose:

[0024]

[0025] where is the joint position obtained by forward kinematics of the reference pose.

[0026] To avoid network overfitting, the regularization term is:

[0027]

[0028] where φ is the network parameter. The complete loss function after adding the two-branch decoder is:

[0029]

[0030] In this process, the weight coefficient λ of the Kullback-Leibler divergence loss is reduced to 0.2; the network is trained until it converges again. Take out the trained encoder and its parameters as the sampling distribution prior.

[0031] Furthermore, the specific method of step S4 includes:

[0032] S41: Obtain the human body's two-dimensional pose using existing human body two-dimensional pose estimation methods;

[0033] S42: Generate the signed distance field of the scene according to the scene grid for constructing contact constraints in the optimization fitting;

[0034] S43: Construct the optimization equation and constraint terms, and minimize the loss under each constraint term by optimizing the human kinematic skeleton parameters. The overall formula of the constraint terms is:

[0035]

[0036] where θ, are the pose parameters, global rotation, and translation of each frame of the person respectively; β is the human bone length parameter; T is the number of frames.

[0037] The formula of the data term is:

[0038]

[0039] where p, σ are the two-dimensional pose and the corresponding confidence. is the three-dimensional position of the joint points of the human kinematic model. Π is the camera projection that projects the three-dimensional joint points onto the two-dimensional plane.

[0040] The regularization term is:

[0041]

[0042] To better reconstruct the interaction between the human body and the scene, use the differentiable signed distance field generated in step S42 to further construct contact constraints. Pre-define the key points of the two feet's toes and ankles as the contact points, and sample using the pre-defined key points in the signed distance field to prompt the contact points to approach the scene surface:

[0043]

[0044] where is the position of the three-dimensional key point, and SDF is the sampling operation. To make this method conform to aerial motion, the Germa-McClure error function ρ is added to reduce the weight of the key points far from the scene grid. By optimizing the above equation, the reference pose can finally be obtained

[0045] Furthermore, the specific method of step S5 includes:

[0046] S51: Implement human motion capture based on neural motion control starting from the first frame of the video. Use the first and second frame postures of the human reference motion for differencing to initialize the physical character with the obtained angular velocity and the posture of the first frame.

[0047] S52: Take out the next frame of reference posture from the human reference motion, and feed the current state of the physical character and the reference posture into the trained distribution prior to estimate the sampling distribution;

[0048] S53: Sample k correction values of the reference posture from the estimated distribution, and add them to the estimated input reference posture to form the target posture;

[0049] S54: Use the physics engine to simulate each target posture in the current state of the physical character in turn;

[0050] S55: Use the motion consistency loss to evaluate the sampling results. Among them, the loss function of the simulated posture and the reference posture is used to measure the consistency of the posture and joint position:

[0051]

[0052] The dynamic characteristics of the human are particularly important for physics-based motion capture, so a dynamic loss is introduced to measure the speed consistency:

[0053]

[0054] Where and are the joint angular velocity and linear velocity.

[0055] To maintain the balance of the physical character, a balance term is added to adjust the center of mass:

[0056]

[0057] Where, d m = j m - j CoM | z=0 , is the two-dimensional vector from the end joint point of the limb to the center of mass; j CoM is the linear velocity of the center of mass; M is the number of ends.

[0058] In addition, further use image features to evaluate the sampling quality, add an image-level loss function, and enhance the robustness in occlusion scenarios:

[0059]

[0060] The overall loss function of the motion capture based on neural motion control is:

[0061]

[0062] S56: The simulation result with the minimum loss among all results is used as the starting state of the physical character in the next frame; Steps S52 - S56 are repeated until all video frames of the motion sequence are processed;

[0063] S57: The simulation result with the minimum motion loss among all frames of the motion sequence is output as the reconstruction result.

[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0065] 1. The human body capture framework based on neural motion control proposed by the present invention uses a physics engine to provide hard physical constraints, avoiding physically unrealistic phenomena such as penetration and jitter in traditional human motion capture. 2. The sampling distribution prior and its training method proposed by the present invention avoid the time-consuming distribution optimization process of the traditional covariance matrix adaptation optimization strategy, and eliminate the random error in the estimated distribution. 3. The contact constraint proposed by the present invention can obtain accurate human reference motion in an environment with a complex terrain scene, and further uses neural motion control to achieve human motion capture on complex terrains. 4. The input of this method is only a single-view RGB video, which is convenient to collect, has a low cost, and is easy to implement. Description of the Drawings

[0066] Figure 1 is the flowchart of the present invention;

[0067] Figure 2 is a schematic diagram of the prior training framework of the target pose distribution;

[0068] Figure 3 is a human motion capture framework diagram based on neural motion control;

[0069] Figure 4 is a physical human body modeling result diagram with different body proportions;

[0070] Figure 5 is a human pose reprojection result and 3D visualization result diagram. Detailed Embodiments

[0071] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. It should be noted that the terms "front", "rear", "left", "right", "upper" and "lower" used in the following description refer to the directions in the drawings, and the terms "inner" and "outer" refer to the directions towards or away from the geometric center of a specific component respectively.

[0072] The following details the implementation process of the present invention in conjunction with the embodiments and the drawings of the specification.

[0073] A high-physical-realism human motion capture method based on neural motion control in this embodiment includes the following steps:

[0074] (I) Construction of human body models with different body types

[0075] (1) Construct human body models with different body types according to the human bone lengths provided by the existing human kinematics dataset and the predefined kinematic tree, and the results are as Figure 4 shown.

[0076] (II) Prior training of sampling distribution

[0077] (2) Use the existing human kinematics dataset for prior training of sampling distribution. Sequentially select two consecutive frames of human body postures at time t - 1 and time t from the dataset. By differentiating the selected postures, calculate the human joint velocity at time t - 1, and then use the human body posture and velocity at time t - 1 as the state of the physical role, and use the human body posture at the t-th frame as the reference posture. Construct a standard normal distribution as the initial sampling distribution, perform physical simulation on the corresponding physical role state, reference posture, and the result of initial distribution sampling, and use the covariance matrix adaptation strategy to optimize the initial distribution. After processing all the data in the dataset in sequence, the estimated distribution is the supervised data for pre-training the network parameters; since the result obtained by differentiation has a deviation from the real-world simulation, in order to enhance the authenticity of the data, add random noise to the physical role state to simulate the real situation. After processing all the data in the dataset in sequence, the estimated distribution is the supervised data for pre-training the network parameters; since the result obtained by differentiation has a deviation from the real-world simulation, in order to enhance the authenticity of the data, add random noise to the physical role state to simulate the real situation.

[0078] (3) After completing the construction of the human body model and data preparation, construct a neural network training framework as shown in Figure 2 and pre-train the distribution encoding neural network using the generated data. Input the reference posture at time t and the physical state of the role at time t - 1 into the distribution encoding neural network to obtain the estimated sampling distribution Use the distribution optimized by the adaptive evolution strategy to supervise the encoder, that is, the distribution encoding neural network. Use the Kullback-Leibler divergence as the loss function to pre-train the distribution encoding neural network:

[0079]

[0080] where is the distribution obtained by the covariance matrix adaptation evolution strategy, is the distribution estimated by the encoding network, and KL is the Kullback-Leibler divergence. When the loss converges, the pre-training is completed

[0081] (4) After completing pre-training, since the covariance matrix adaptation evolution strategy has random errors, a two-branch decoder is introduced to optimize the network parameters network. As Figure 2 shown in the upper right branch, the pose decoding network estimates the human physical pose from the target pose; as Figure 2 shown in the lower right branch, the physical simulation branch also receives the target pose, obtains the simulated state after simulation, removes the joint point velocity, and obtains the human simulated pose q t . Since the physical simulation branch is not differentiable, the simulated pose obtained by it is used as supervision and acts on the human pose estimated by the pose decoder, prompting the pose decoding network to regress to the same result as the physical simulation. Its loss function is:

[0082]

[0083] where is the human pose regressed by the pose decoding network, and j t are the joint positions obtained by forward kinematics of the regressed human pose and the simulated human pose respectively, and ‖·‖ 2 is the two-norm.

[0084] In addition, in order to make the encoder regress to the correct distribution, the reconstruction constraint is further used to constrain the estimated human pose:

[0085]

[0086] where is the joint position obtained by forward kinematics of the reference pose.

[0087] A regularization term is added to avoid overfitting of the network:

[0088]

[0089] where φ are the network parameters. The complete loss function after adding the two-branch decoder is:

[0090]

[0091] In this embodiment, the weight coefficient λ of the Kullback-Leibler divergence loss is reduced to 0.2 during the training stage of introducing the two-branch decoder; the network is trained until it converges again. The trained encoder and its parameters are taken out and used as the prior of the sampling distribution.

[0092] (III) Human kinematic reconstruction

[0093] (5) Utilize existing human body two-dimensional pose estimation methods. Input a single-view video, where each frame is a human body RGB image. Use existing human body two-dimensional pose estimation methods to detect the bounding boxes of human bodies in the environment, and then detect the poses of each human body region to obtain the two-dimensional poses of the human body.

[0094] (6) Generate the signed distance field of the scene according to the scene grid for constructing contact constraints in the optimization fitting, that is Figure 1 the scene contact constraints in

[0095] (7) After completing the input preparation, perform the physical character modeling process as shown in Figure 1 . Construct the optimization equation and constraint terms, and minimize the loss under each constraint by optimizing the human kinematic skeleton parameters. The overall formula of the constraint terms in this embodiment is:

[0096]

[0097] where θ, are the pose parameters, global rotation, and translation of each frame of the person respectively; β is the human bone length parameter; T is the number of frames.

[0098] where the formula of the data term is:

[0099]

[0100] where p, σ are the two-dimensional poses of the human body estimated in step (5) and the corresponding confidence levels. is the three-dimensional position of the joint points of the human kinematic model. Π is the camera projection, which projects the three-dimensional joint points onto the two-dimensional plane.

[0101] The regularization term is:

[0102]

[0103] To better reconstruct the interaction between the human body and the scene, use the differentiable signed distance field generated in (6) to further construct contact constraints. Pre-define the key points of the two toes and ankles of both feet as the contact points, and sample using the pre-defined key points in the signed distance field to prompt the contact points to approach the scene surface:

[0104]

[0105] where is the position of the three-dimensional key points, and SDF is the sampling operation. To make this method conform to aerial motion, the Germa-McClure error function ρ is added to reduce the weight of the key points far from the scene grid. By optimizing the above equation, the reference pose of all frames can be obtained through the above process, and the reference motion of the human body can be obtained from the poses of all frames. That is, the physical character modeling is completed.

[0106] (4) Human motion capture based on neural motion control

[0107] (8) Starting from the first frame of the video, human motion capture based on neural motion control is implemented. The first and second frame postures of the human reference motion are used for differencing to obtain the angular velocity and the posture of the first frame, and the physical character is initialized with them, thus obtaining the physical character state of the first frame.

[0108] (9) Based on the physical character state of the (t - 1)-th frame and the reference posture of the t-th frame, the present invention uses the trained sampling distribution prior to implement neural motion control and obtains the physical character state of the t-th frame.

[0109] (10) As shown in Figure 3 , the current state of the physical character, i.e., the physical state at the (t - 1) moment, and the next frame taken from the human reference motion sequence, i.e., the reference posture at the t moment, are fed into the trained distribution prior neural network to estimate the sampling distribution Q(μ t , σ t ); it is sampled k times to obtain the corrected value of the reference posture which is respectively added to the human reference posture to form k human target postures, that is, that is Figure 3 the target posture correction sampling process in Figure 1 ; the k human target postures are input into the physics engine, and under the current state of the physical character, each target posture is simulated in turn, that is

[0110] (11) Motion - dynamics constraints obtained from the human two - dimensional posture are added to the simulation results; the simulation results are evaluated, and the quality of the samples is evaluated by the loss functions at the human level and the image level, and the simulation result with the minimum loss function value is selected as the physical character state at the t moment. The physical character is visually output as shown in Figure 5 . During the implementation of this example, the motion consistency loss is used to evaluate the sampling results. Among them, the loss function between the simulated posture and the reference posture is used to measure the consistency of the posture and the joint position:

[0111]

[0112] The dynamic characteristics of the human are particularly important for physics - based motion capture, so the dynamic loss is introduced to measure the speed consistency:

[0113]

[0114] where and are the joint angular velocity and linear velocity.

[0115] To maintain the balance of the physical character, a balance term is added to adjust the center of mass:

[0116]

[0117] where d m = j m - j CoM | z=0 , is the two-dimensional vector from the end joint point of the limb to the center of mass; j CoM is the linear velocity of the center of mass; M is the number of ends.

[0118] In addition, the image features are further used to evaluate the sampling quality, and an image-level loss function is added to enhance the robustness in the occlusion scenario:

[0119]

[0120] The overall loss function of the motion capture based on neural motion control is:

[0121]

[0122] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above specific embodiments. The above specific embodiments and the descriptions in the specification are only for further illustrating the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope claimed by the present invention is defined by the claims and their equivalents.

Claims

1. A high-physical-realism human motion capture method based on neuromotor control, characterized in that: It includes the following steps: S1. Modeling of human kinematic model and human dynamic model; S2. Definition of human reference motion and physical character state; S3. Training of sampling distribution prior neural network: Obtain data from the existing kinematic human dataset, run the covariance matrix adaptation evolution strategy offline to generate the sampling distribution for the pre-training of the supervised distribution encoder; After the pre-training is completed, connect to the two-branch decoder to further optimize and train the network parameters of the encoder until the network converges; Use the trained encoder network as the sampling distribution prior for subsequent human motion capture; The specific method of S3 includes: S31: Training data preparation: Select two consecutive frames from the existing kinematic human body dataset: frame t-1 and frame t. By differentiating the selected postures, calculate the human joint velocities of frame t-1, and then use the human posture and velocity of frame t-1 as the state of the physical character, and use the human posture of the t-th frame as the reference posture; construct a standard normal distribution as the initial sampling distribution, perform physical simulation on the corresponding physical character state, reference posture, and reference posture correction value sampled from the initial distribution, and use the covariance matrix adaptation strategy to optimize the initial distribution. After the optimization is completed, the distribution Process all the data in the dataset in sequence, and the estimated distribution is the supervised data used to pre-train the network parameters; since the results obtained by differentiation deviate from the real simulation, in order to enhance the authenticity of the data, add random noise to the physical character state to simulate the real situation; Among them, the physical simulation includes the following steps: S311: Add the reference pose and the corrected value of the sampled reference pose to obtain the target pose; S312: Use the target pose as the set value, combine with the current physical character state, and calculate the joint torque using PD control; S313: Apply the joint torque to the joints of the physical character and execute the physical simulation using the physics engine; S314: Read the simulation results; S32: Pre-training of the distribution encoding network: Construct a 6-layer fully connected encoding network, and the output is the mean and variance of 32 dimensions; Combine the estimated mean and variance to form a normal distribution; Use the distribution generated by the offline covariance matrix adaptation strategy as the supervision to train the network until it converges; Use the Kullback-Leibler divergence as the constraint, and its loss function formula is: where is the distribution obtained by the covariance matrix adaptation evolution strategy, is the distribution estimated by the encoding network, and KL is the Kullback-Leibler divergence; After the training converges, due to the random error of the covariance matrix adaptation evolution strategy, there is noise in the generated annotations, and it is difficult for the pre-trained model to regress to the accurate sampling distribution; A two-branch decoder including a pose decoding network and a physical simulation branch is proposed to assist the encoder in training to obtain accurate parameters; Among them, the pose decoding network has 4 fully connected layers to regress the simulated human pose from the target pose; The physical simulation branch also receives the target pose and obtains the simulated pose after simulation; Since the physical simulation branch is not differentiable, the obtained simulated pose is used as the supervision and acts on the human pose estimated by the pose decoding network to prompt the pose decoding network to regress to the same result as the physical simulation; Its loss function is: Among them, is the human pose regressed by the pose decoding network, and j t are the joint positions obtained by forward kinematics of the regressed human pose and the simulated human pose respectively, ‖·‖ 2 is the two-norm; In order to make the encoder regress to the correct distribution, the reconstruction constraint is further used to constrain the estimated human pose: Among them is the joint position obtained by forward kinematics of the reference pose; In order to avoid network overfitting, the regularization term is: Where φ is the network parameter; The complete loss function after adding the two-branch decoder is: In this process, the weight coefficient λ of the Kullback-Leibler divergence loss is reduced to 0.2; Train the network until it converges again; Take out the trained encoder and its parameters as the sampling distribution prior; S4. Human reference motion and human body shape estimation; S5. High-physical-reality human motion capture; The specific method of S5 includes: S51: Implement human motion capture based on neural motion control starting from the first frame of the video; Use the first and second frame poses of the human reference motion to perform differencing to obtain the angular velocity and use the pose of the first frame to initialize the physical character; S52: Extract the next reference pose from the human reference motion, and feed the current state of the physical character and the reference pose into the trained distribution prior to estimate the sampling distribution; S53: Sample k correction values of the reference pose from the estimated distribution, and add them to the estimated input reference pose to form the target pose; S54: Use the physics engine to simulate each target pose in the current state of the physical character in sequence; S55: Use the motion consistency loss to evaluate the sampling results; among them, the loss function between the simulated pose and the reference pose is used to measure the consistency of the pose and joint positions: The dynamic characteristics of the human are particularly important for physics-based motion capture, so a dynamic loss is introduced to measure the speed consistency: wherein and are the joint angular velocity and linear velocity; To maintain the balance of the physical character, a balance term is added to adjust the center of mass: where d m = j m - j CoM | z=0 , is the two-dimensional vector from the joint point at the limb end to the centroid; j CoM is the linear velocity of the centroid; M is the number of ends; In addition, the image features are further used to evaluate the sampling quality, and an image-level loss function is added to enhance the robustness in occlusion scenarios: The overall loss function of the motion capture based on neural motion control is: S56: The simulation result with the minimum loss among all results is used as the starting state of the next frame of the physical character; repeat steps S52 - S56 until all video frames of the motion sequence are processed; S57: The simulation result with the minimum motion loss among all frames of the motion sequence is output as the reconstruction result.

2. A high-physical-realism human motion capture method based on neural motion control according to claim 1, characterized in that: In the above S2: the human reference motion is obtained by fitting the human kinematic model to the human features in the input video. The human reference pose parameters of each frame after fitting constitute the complete human reference motion. The human reference pose parameters of the t-th frame are expressed as The human reference skeleton pose parameters include 3D model translation, 3D model global rotation, and 51D joint rotation, with a total dimension of 57D; the physical character state are the parameters read from the physics engine after the human dynamics model is simulated, which include the simulated pose parameters q and the velocity parameters The simulated pose parameters are the same as those of the human reference skeleton pose parameters. The velocity parameters include 3D root node linear velocity, 3D root node angular velocity, and 51D joint angular velocity. The total dimension of the physical character state is 114D.

3. A high-physical-realism human motion capture method based on neural motion control according to claim 1, characterized in that: The specific method of S4 includes: S41: Use the existing human two-dimensional pose estimation method to obtain the human two-dimensional pose; S42: Generate the signed distance field of the scene according to the scene grid, which is used to construct contact constraints in the optimization fitting; S43: Construct the optimization equation and constraint terms, and minimize the loss under each constraint term by optimizing the human kinematic skeleton parameters; the overall formula of the constraint terms is: where θ, τ are the pose parameters, global rotation, and translation of each frame of the person, respectively; β is the human bone length parameter; T is the number of frames; Among them, the formula of the data term is: where p, σ are two-dimensional postures and corresponding confidence levels; are the three-dimensional positions of the joint points of the human kinematic model; Π is the camera projection that projects the three-dimensional joint points onto the two-dimensional plane; The regularization term is: To better reconstruct the interaction between the human body and the scene, the differentiable signed distance field generated in step S42 is further used to construct contact constraints; key points of both feet's toes and ankles are predefined as contact points, and the predefined key points are sampled in the signed distance field to prompt the contact points to approach the scene surface: where is the position of the 3D key point, and SDF is the sampling operation; in order to make the method conform to the aerial motion, the Germa-McClure error function ρ is added to reduce the weight of the key points far from the scene grid; by optimizing the above equation, the reference pose can be finally obtained

Citation Information

Patent Citations

  • Virtual human movement simulation method based on graph neural network

    CN112017265A

  • Multi-person body model reconstruction method based on hidden space motion coding

    CN113379904A