Fine-grained Human Action Generation Method Based on Physical Laws and Skeletal Guidance
By using the Euler-Lagrangian equation to optimize the three-dimensional skeleton sequence based on physical laws and bone guidance, the existing generative model solves the problems of incoherence and visual distortion when generating severe human body deformation and significant movements in timing changes, and achieves a more natural and reasonable fine-grained human body movement generation.
Patent Information
- Application Number
- CN202510575275.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-06
AI Technical Summary
When the existing generative models generate fine-grained human movements with severe human body deformation and significant timing changes, there are problems of movement incoherence, visual distortion, lack of spatial consistency and mechanical rationality defects, especially when generating complex human body movements, it is difficult to effectively learn and apply physical laws.
Using a method based on physical laws and skeleton guidance, a two-dimensional skeleton sequence is extracted by a motion detector, a three-dimensional skeleton sequence is optimized by combining context learning and Euler-Lagrangian equation, a three-dimensional skeleton sequence is integrated with physical prediction and data-driven 3-dimensional skeleton sequence, and finally a fine-grained human action video is generated through 3D-UNet.
More natural and reasonable fine-grained human body movements are generated, which improves spatial perception and timing consistency, and significantly improves the physical credibility and quality of the movements.
Smart Images

Figure CN120088442B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and computer vision, and specifically to a fine-grained human action generation method based on physical laws and bone guidance. Background Art
[0002] In the field of generative model technology, breakthrough progress represented by diffusion models has significantly promoted the development of human action video generation technology. For example, the publication number CN119399332A proposes a human action generation method based on diffusion models and fine-grained text descriptions, which can generate zero-shot human actions beyond the scope of the original dataset and has good generalization ability.
[0003] However, the core elements involved in video temporal modeling - including camera motion trajectory control, dynamic background adaptation, and human action coherence - still pose important technical challenges. This challenge is particularly prominent in human action generation tasks, specifically manifested as the frequently occurring phenomena of inconsistent actions and visual distortion in the generated videos.
[0004] From the spatial dimension analysis, although the human body has strict anatomical structure constraints in the physical world, existing models often generate abnormal structural features that violate ergonomics during the processing. This lack of spatial consistency mainly stems from the insufficient ability of neural networks to represent complex human topological structures. At the time dynamics level, generating limb trajectories that conform to kinematic principles is still an urgent problem to be solved. The latest research shows that even the current most advanced generative models are difficult to effectively learn and apply basic physical laws including Newton's laws of motion during the generation process, which directly leads to obvious defects in the mechanical rationality of the generated actions.
[0005] The above problems lead to a two-dimensional challenge for existing methods in generating fine-grained human actions with significant human body deformation and temporal changes: in the spatial dimension, they need to cope with maintaining the topological structure under severe deformation, and in the time dimension, they need to satisfy the dynamic constraints of complex motion trajectories. For example, when generating gymnastic actions such as "turning 180° and swapping legs to jump", existing methods generally have unstable generation quality, including serious time inconsistency, obvious limb distortion, and abnormal human body structures, as Figure 1 shown, indicating that the models designed by existing methods have not established an effective mechanism for internalizing physical laws. Summary of the Invention
[0006] Aiming at the problems existing in the prior art, the present invention proposes a fine-grained human action generation method based on physical laws and bone guidance, abbreviated as FinePhys, which can complete the human action generation task with significant human body deformation and temporal changes, such as Figure 1As shown, FinePhys has shown excellent performance in generating physically plausible fine-grained human actions.
[0007] The technical solution of the present invention is as follows:
[0008] A method for generating fine-grained human actions based on physical laws and skeleton guidance, comprising the following steps:
[0009] Step 1: Obtain a text description of the fine-grained human action video to be output; ; According to the text description , obtain a human action video similar to the text description , extract a sampling frame sequence of length from the human action video , and use an action detector to perform two-dimensional pose estimation on the sampling frame sequence to generate a two-dimensional skeleton sequence ;
[0010] Step 2: Perform dimension elevation on the two-dimensional skeleton sequence obtained in Step 1 through context learning to obtain a data-driven three-dimensional skeleton sequence ; The parameters in the context learning process are parameters to be trained; ;
[0011] Step 3: Optimize the data-driven three-dimensional skeleton sequence obtained in Step 2 based on physical prediction to obtain a physically predicted three-dimensional skeleton sequence , and the specific process is as follows: ;
[0012] Step 3.1: For the data-driven three-dimensional skeleton sequence , respectively obtain the time state sequences of the overall sequence, the first three-frame sequence, and the last three-frame sequence in the three-dimensional skeleton sequence through an encoder, and fuse the time state sequence of the overall sequence with the time state sequence of the first three-frame sequence to obtain a forward fusion sequence ; Fuse the time state sequence of the overall sequence with the time state sequence of the last three-frame sequence to obtain a reverse fusion sequence ;
[0013] Step 3.2: Based on the Euler-Lagrange equation, perform forward update on the forward fusion sequence , perform reverse update on the reverse fusion sequence , where only the first three frames of the forward fusion sequence are subjected to forward update, only the last three frames of the reverse fusion sequence are subjected to reverse update, and the remaining frames take the average of the forward update result and the reverse update result;
[0014] The forward update process is as follows:
[0015] Input the forward fusion sequence into the forward physical parameter estimator to estimate the parameters of the forward updated Euler-Lagrange equation, including the forward updated generalized inverse inertia matrix , the forward updated generalized force and the forward updated joint constraints ; and introduce the forward updated noise parameter matrix and add it to the forward updated generalized inverse inertia matrix. The noise parameter matrix is also obtained through the forward physical parameter estimator; thus, according to the forward updated Euler-Lagrange equation:
[0016]
[0017] obtain the second derivative of the forward fusion sequence , and then use the second-order central difference formula
[0018]
[0019] to obtain the forward updated result, where is the time interval between adjacent frames; the forward updated result is:
[0020]
[0021] The reverse update process is as follows:
[0022] Input the reverse fusion sequence into the reverse physical parameter estimator to estimate the parameters of the reverse updated Euler-Lagrange equation, including the reverse updated generalized inverse inertia matrix , the reverse updated generalized force and the reverse updated joint constraints ; introduce the reverse updated noise parameter matrix and add it to the reverse updated generalized inverse inertia matrix. The noise parameter matrix is also obtained through the reverse physical parameter estimator; thus, according to the reverse updated Euler-Lagrange equation:
[0023]
[0024] obtain the second derivative of the reverse fusion sequence , and then use the second-order central difference formula
[0025]
[0026] to obtain the reverse updated result, where is the time interval between adjacent frames; the reverse updated result is:
[0027]
[0028] The parameters in both the forward physical parameter estimator and the inverse physical parameter estimator are parameters that need to be trained;
[0029] Step 3.3: Data-driven 3D skeleton sequence After each frame in the sequence is updated in Step 3.2, through pose decoding, a physically predicted 3D skeleton sequence is obtained; ;
[0030] Step 4: Fuse the data-driven 3D skeleton sequence obtained in Step 2 with the physically predicted 3D skeleton sequence obtained in Step 3 and project the fusion result back to 2D, and encode and fine-tune it into a multi-scale 2D skeletal heatmap sequence for guiding the denoising process of 3D-UNet; among them, the parameters of the fusion process, the mapping projection process, and the fine-tuning process are parameters that need to be trained;
[0031] Step 5: Normalize the text description through a pre-trained large language model to obtain a text sequence , and then encode the text sequence to obtain a text encoding vector , is a pre-trained text encoding model; extract a sampling frame sequence of length from the human action video to perform feature extraction to obtain a feature vector , and then perform noise addition processing on the feature vector to obtain a noise-added feature vector , where represents the number of noise addition steps; input the noise-added feature vector , the text encoding vector , and the number of noise addition steps together into a 3D-UNet model that combines the multi-scale 2D skeletal heatmap sequence output in Step 4; decode the output of the 3D-UNet model to obtain a fine-grained human action video; the parameters in the 3D-UNet model are parameters that need to be trained.
[0032] Furthermore, the parameters that need to be trained are realized through the following phased training process:
[0033] First, train the parameters that need to be trained in Step 2, Step 3.2, and the fusion process in Step 4. After the training is completed, freeze the parameters; the sample data used is a standard human 3D pose dataset, and the 3D sequence obtained after the fusion in Step 4 is denoted as , the loss function used in this stage of training is:
[0034]
[0035] Among them represents the th frame, represents the th human body joint; represents the three-dimensional coordinates of the th human body joint in the th frame of the three-dimensional sequence , and represents the true three-dimensional coordinates of the th human body joint in the th frame of the training sample data in this stage. , among which for the first three frames, adopts , and for the last three frames, adopts . For the remaining frames, adopts the mean value corresponding to and ;
[0036] Secondly, train the mapping projection parameters and fine-tuning parameters in step 4. After the training is completed, freeze the parameters. The sample data used is the fine-grained human action video. Let the multi-scale two-dimensional bone heat map sequence obtained after fine-tuning in step 4 be , then the loss function used in the training is:
[0037]
[0038] Among them is the two-dimensional coordinates of the th human body joint in the th frame, represents the true two-dimensional coordinates of the th human body joint in the th frame of the training sample data in this stage;
[0039] Finally, train the parameters that need to be trained in step 5. The sample data used is the fine-grained human action video. The loss function used in the training is:
[0040]
[0041] Among them is a random number that conforms to the normal distribution, is the set of noise addition step sizes, represents 's modulus, is the 3D-UNet model.
[0042] Furthermore, the specific process of step 2 is as follows:
[0043] Step 2.1: Using the skeleton dataset and the two-dimensional skeleton sequence obtained in step 1 , construct the prompt and the query ; where is a two-dimensional skeleton sequence randomly selected from the skeleton dataset, is the corresponding three-dimensional skeleton sequence in the skeleton dataset to ; is the average three-dimensional skeleton sequence, obtained by calculating the average of a large number of three-dimensional skeleton sequences selected from the skeleton dataset;
[0044] Step 2.2: Based on the prompt and the query, perform dimensional transformation to obtain a data-driven three-dimensional skeleton sequence .
[0045] Furthermore, in step 2.2, the dimensional transformation is realized by a bidirectional Transformer composed of spatio-temporal modules, and the parameters in the bidirectional Transformer are parameters to be trained.
[0046] Furthermore, in step 3.2, the forward physical parameter estimator and the inverse physical parameter estimator are each composed of four multi-layer perceptrons. Among them, the generalized force and the joint constraint are vectors, and each is obtained through a multi-layer perceptron; the generalized inverse inertia matrix is obtained in two steps. The first step is to obtain an upper triangular matrix through the third multi-layer perceptron, and then perform a symmetry operation on the upper triangular matrix to obtain a symmetric matrix. The second step is to add a noise parameter matrix to the symmetric matrix to obtain the generalized inverse inertia matrix; the noise parameter matrix is obtained in two steps. The first step is to obtain a Gaussian noise vector with a variance of 1 through the fourth multi-layer perceptron, and the second step is to stack the Gaussian noise vectors according to the set dimension and add random noise to each Gaussian noise vector to obtain the noise parameter matrix.
[0047] Furthermore, in step 5, during the downsampling and upsampling processes of the 3D-UNet model, a LoRA module that combines the multi-scale two-dimensional bone heat map sequence obtained in step 4 is used to update the weight matrices and in the downsampling and upsampling processes with two low-rank matrices , where is the initial weight matrix; the two low-rank matrices and are parameters to be trained.
[0048] In addition, the present invention also provides an electronic device and a readable storage medium:
[0049] An electronic device includes a processor and a memory, and the memory is used to store one or more programs;
[0050] When the one or more programs are executed by the processor, the above method is implemented.
[0051] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0052] Beneficial effects:
[0053] The FinePhys method architecture proposed by the present invention first obtains relevant action videos in an online manner, and extracts two-dimensional postures from the videos as prior information of the human physical structure; and realizes the transformation from 2D postures to 3D postures through a context learning method to enhance spatial perception, and obtains a data-driven three-dimensional skeleton sequence ; further designs a physical module PhysNet, uses the Euler-Lagrange formula to re-estimate human actions, and calculates joint accelerations bidirectionally in time series to generate physically predicted three-dimensional postures ; then, the physically estimated three-dimensional postures and the data-driven three-dimensional skeleton sequence are fused and converted into a 2D heatmap format to participate in 3D-UNets for video generation. The present invention is tested on 3 fine-grained action data subsets, and FinePhys far exceeds other video generation methods, and a large number of qualitative analyses further prove that FinePhys can generate more natural and reasonable fine-grained human actions.
[0054] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0055] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, wherein:
[0056] Figure 1 : Video generation result of the fine-grained human action "180° body turn and leg swap jump".
[0057] Figure 2 : Schematic diagram of the FinePhys architecture proposed by the present invention.
[0058] Figure 3 : Schematic diagram of the PhysNet module proposed by the present invention. Detailed Embodiments
[0059] Embodiments of the present invention will be described in detail below. The embodiments are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0060] To address the more challenging task of generating fine-grained human actions with drastic human deformations and significant temporal variations, this embodiment proposes a physical law and skeleton-guided fine-grained human action generation method, abbreviated as FinePhys, for fine-grained human action generation, which can complete the task of generating human actions with drastic human deformations and significant temporal variations.
[0061] As Figure 2 shown, FinePhys, as a physics-aware framework, obtains relevant action videos in an online manner according to the input text description, and extracts 2D poses from the videos as prior information of the human physical structure; then uses the context learning module to convert the 2D poses into 3D poses to enhance spatial perception, obtaining a data-driven three-dimensional skeleton sequence. . Since this data-driven 3D pose sequence often ignores the physical laws of motion, in order to incorporate the physical laws of motion, the PhysNet module is introduced in FinePhys, which uses the Euler-Lagrange equation of rigid body dynamics to encode Newtonian mechanics, and re-estimates the three-dimensional positions of each human joint by considering the forward and backward second-order time variations (i.e., acceleration), thereby generating physically predicted three-dimensional poses. . Subsequently, and are fused, projected back to the 2D space, encoded into multi-scale heatmaps, and integrated into 3D-UNets to guide the denoising process, and the final output is a fine-grained action video. .
[0062] In the entire perception framework of FinePhys, this embodiment adopts three strategies: observation bias (through data), inductive bias (through the network), and learning bias (through loss), and coherently integrates the physical laws into the learning process. Specifically:
[0063] 1. For observation bias, FinePhys takes the pose as an additional modality encoding the biophysical layout, and uses the average three-dimensional pose from the existing dataset as a pseudo-three-dimensional reference to achieve 2D-to-3D enhancement through context learning.
[0064] 2. For inductive bias, the Lagrangian rigid body dynamics is instantiated through the fully differentiable neural network module PhysNet, encoding stronger inductive bias into FinePhys and outputting the parameters in the Euler-Lagrange equation.
[0065] 3. For the learning deviation, FinePhys adopts a loss function that follows the underlying physical process.
[0066] Example 1:
[0067] In this example, a fine-grained human action generation method based on physical laws and skeleton guidance is proposed, which specifically includes the following steps:
[0068] Step 1: Obtain a text description of the fine-grained human action video that is desired to be output ;
[0069] According to the text description , obtain a human action video similar to the text description , for example, obtained through online search;
[0070] Extract from the human action video a sampling frame sequence of length as the prior information of the human physical structure;
[0071] Use an action detector to perform two-dimensional pose estimation on the sampling frame sequence to generate a two-dimensional skeleton sequence , where is the dimension of the bone joints, which is 17 in this example, represents two-dimensional, indicating that each joint point has two coordinate values. The two-dimensional skeleton sequence, as an additional mode, provides a compact and biologically structured reasonable description of human actions. The action detector is an existing technology in the art.
[0072] Step 2: In order to convert the two-dimensional layout of the human skeleton into a three-dimensional geometry, perform dimensionality lifting on the two-dimensional skeleton sequence obtained in Step 1 through context learning to obtain a data-driven three-dimensional skeleton sequence . The specific process is as follows:
[0073] Step 2.1: Since context learning requires some instances as prompt words, in this example, widely used skeleton datasets such as Human3.6M or AMASS are used to construct prompt words; the prompt words are represented by input-output pairs: prompt words , where is a two-dimensional skeleton sequence randomly selected from the skeleton dataset, and the is the three-dimensional skeleton sequence corresponding to in the skeleton dataset;
[0074] Construct a query word for the two-dimensional skeleton sequence :
[0075] Calculate the average of a large number of 3D skeleton sequences selected from the skeleton dataset, and use the obtained average 3D skeleton sequence as the pseudo 3D prior , in this embodiment, the average of all 3D skeleton sequences in the Human3.6M skeleton dataset is calculated. Then a query word is constructed ;
[0076] Step 2.2: Based on the prompt word and the query word, perform dimensional transformation to obtain a data-driven 3D skeleton sequence :
[0077]
[0078] Among them, the dimensional transformation is implemented by a bidirectional Transformer composed of spatio-temporal modules. The bidirectional Transformer includes input judgment and encoding, feature fusion and Transformer module processing, subsequent mapping and reconstruction, and output process, which are well-known technologies in the art. The parameters in the bidirectional Transformer are parameters to be trained.
[0079] Figure 2 The lower left shows the process of dimensional elevation of the 2D skeleton sequence through context learning. The entire 2D to 3D elevation process directly benefits from using the observed data and mastering the basic physical structures and rules, such as the relationships of limbs, the anatomical limitations of joints, the spatial layout of the 3D human body, etc. through a trainable bidirectional Transformer.
[0080] Step 3: Since most online estimators are difficult to accurately estimate 2D human actions, especially complex actions like gymnastics, noise is inevitably introduced. In addition, the data-driven 2D to 3D transformation lacks physical interpretability, resulting in the data-driven 3D skeleton sequence often having unreliable estimation results. To solve these problems, in this step, the second-order time difference change (i.e., acceleration) is calculated bidirectionally using the complete Euler-Lagrange equation to re-estimate the 3D motion dynamics, and the physical terms in the Euler-Lagrange equation are estimated to obtain the time series change amount, so as to realize the optimization of the data-driven 3D skeleton sequence obtained in Step 2 based on physical prediction, and obtain a physically predicted 3D skeleton sequence . As Figure 3 shown, the specific process is as follows
[0081] Step 3.1: For the data-driven 3D skeleton sequence , respectively obtain the 3D skeleton sequence The time state sequences of the overall sequence, the first three-frame sequence, and the last three-frame sequence, and fuse the time state sequence of the overall sequence with the time state sequence of the first three-frame sequence to obtain a forward fusion sequence ; fuse the time state sequence of the overall sequence with the time state sequence of the last three-frame sequence to obtain a reverse fusion sequence ; all subsequent calculations are carried out bidirectionally, with forward updates starting from the first three frames and reverse updates starting from the last three frames;
[0082] Step 3.2: Based on the Euler-Lagrange equation, perform forward updates on the forward fusion sequence and perform reverse updates on the reverse fusion sequence where the first three frames of the forward fusion sequence only perform forward updates, the last three frames of the reverse fusion sequence only perform reverse updates, and the remaining frames take the average of the forward update results and the reverse update results;
[0083] The forward update process is as follows:
[0084] Input the forward fusion sequence into the forward physical parameter estimator to estimate the Euler-Lagrange equation parameters for forward updates, including the forward updated generalized inverse inertia matrix , the forward updated generalized force and the forward updated joint constraints ; and since temporal motion often destroys the symmetry of the object, therefore, this embodiment also introduces a forward updated noise parameter matrix and add it to the forward updated generalized inverse inertia matrix, and the noise parameter matrix is also obtained through the forward physical parameter estimator; thus, according to the forward updated Euler-Lagrange equation:
[0085]
[0086] obtain the second derivative of the forward fusion sequence, and then use the second-order central difference formula
[0087]
[0088] to obtain the forward update result, where is the time interval between adjacent frames; the forward update result is:
[0089]
[0090] The reverse update process is as follows:
[0091] Input the reverse fusion sequence Input a reverse physical parameter estimator to estimate the parameters of the reverse updated Euler - Lagrange equation, including the reverse updated generalized inverse inertia matrix , the reverse updated generalized force and the reverse updated joint constraint ; During the reverse update process, a noise parameter matrix is also introduced and added to the reverse updated generalized inverse inertia matrix. The noise parameter matrix is also obtained through the reverse physical parameter estimator; thus, according to the reverse updated Euler - Lagrange equation:
[0092]
[0093] the second - order derivative of the reverse fusion sequence is obtained, and then the second - order central difference formula
[0094]
[0095] is used to obtain the reverse update result, where is the time interval between adjacent frames; the reverse update result is:
[0096]
[0097] In this embodiment, both the forward physical parameter estimator and the reverse physical parameter estimator are respectively composed of four multi - layer perceptrons. As Figure 3 shown, where the generalized force and the joint constraint are vectors, so they can be obtained respectively through one multi - layer perceptron; however, due to the high dimension, directly estimating the generalized inverse inertia matrix is very challenging. Therefore, based on the prior knowledge that the inertia tensor and the mass matrix are usually symmetric in the structural system, the generalized inverse inertia matrix is obtained in two steps. The first step is to obtain an upper - triangular matrix through a multi - layer perceptron, then perform a symmetry operation on the upper - triangular matrix to obtain a symmetric matrix, and then add the noise parameter matrix to the symmetric matrix to obtain the generalized inverse inertia matrix; since the noise parameter matrix is also a high - dimensional matrix, for this reason, the process of obtaining the noise parameter matrix is also divided into two steps. The first step is to obtain a Gaussian noise vector with a variance of 1 through a multi - layer perceptron, then stack the Gaussian noise vectors according to the required dimension, and add random noise to each Gaussian noise vector to obtain the noise parameter matrix; the parameters in the forward physical parameter estimator and the reverse physical parameter estimator are all parameters that need to be trained;
[0098] Step 3.3: After each frame in the data - driven three - dimensional skeleton sequence is updated through Step 3.2, through pose decoding, the physical prediction three - dimensional skeleton sequence is obtained.
[0099] Step 4: The data-driven three-dimensional skeleton sequence obtained in Step 2 is fused with the physically predicted three-dimensional skeleton sequence obtained in Step 3 The fused result is mapped and projected back to two dimensions, encoded, and fine-tuned into a multi-scale two-dimensional bone heat map sequence for guiding the 3D-UNet denoising process.
[0100] In this step, The fusion process with can be achieved by various methods well-known to those skilled in the art, such as weighted summation, etc., where the weights are training parameters; the mapping and projection parameters in the process of mapping and projecting the obtained fusion result back to two dimensions are also training parameters; the obtained two-dimensional mapping result is encoded to generate a two-dimensional reconstructed bone sequence, and further fine-tuned to obtain a multi-scale two-dimensional bone heat map sequence for guiding the 3D-UNet denoising process, where the fine-tuning parameters are training parameters;
[0101] Step 5: The text description is normalized by a pre-trained large language model to obtain a text sequence In this embodiment, the pre-trained large language model uses ChatGPT-4, and then the text sequence is text-encoded to obtain a text encoding vector , is a pre-trained text encoding model; the sampled frame sequence of length extracted from the human action video is feature-extracted to obtain a feature vector , and then the feature vector is noise-added to obtain a noise-added feature vector , where represents the number of noise-adding steps; the noise-added feature vector , the text encoding vector , and the number of noise-adding steps are jointly input into the 3D-UNet model; the output of the 3D-UNet model is decoded to obtain a fine-grained human action video;
[0102] In the downsampling and upsampling processes of the 3D-UNet model, a LoRA module that combines the multi-scale two-dimensional bone heat map sequence obtained in Step 4 is used to update the weight matrices and in the downsampling and upsampling processes with two low-rank matrices , where is the initial weight matrix; the two low-rank matrices and are training parameters.
[0103] When the above steps are implemented, all the parameters to be trained have been completed. The training process is carried out in stages:
[0104] First, train the parameters to be trained in step 2.3, step 3.2, and the fusion process of step 4. After training is completed, freeze the parameters; the sample data used is a dataset with large-scale standard human 3D postures such as Human3.6M and AMASS. The 3D sequence obtained after the fusion of step 4 is represented as , then the loss function used in this stage of training is:
[0105]
[0106] where represents the th frame, represents the th human joint; represents the 3D coordinate of the th th human joint in the 3D sequence , represents the true 3D coordinate of the th th human joint in the th frame of the training sample data in this stage, is used to limit the noise parameter matrix in step 3.2, where is used for the first three frames, is used for the last three frames, and for the remaining frames, uses the mean of the corresponding and .
[0107] Secondly, train the mapping projection parameters and fine-tuning parameters in step 4. After training is completed, freeze the parameters; the sample data used is the fine-grained human action videos provided by FineGym. Let the multi-scale 2D bone heatmap sequence obtained after the fine-tuning of step 4 be , then the loss function used for training is:
[0108]
[0109] where is the 2D coordinate of the th th human joint in , represents the true 2D coordinate of the th th human joint in the
[0110] Finally, train the parameters to be trained in step 5. The sample data used is the fine-grained human action video provided by FineGym. The loss function used for training is:
[0111]
[0112] where is a random number conforming to the normal distribution, is the set of noise addition step sizes, denotes the modulus of, and is the 3D-UNet model.
[0113] Example 2:
[0114] In this example, a fine-grained human action generation system based on physical laws and skeleton guidance is proposed, which specifically includes an action detector, a context learning module, a physical information module, a fusion guidance module, and a video generation module;
[0115] The input of the action detector is a sequence of sampled frames of length ; the sequence of sampled frames is extracted from the human action video , and the human action video is a human action video similar to the text , and the text is a text provided by the user describing the fine-grained human action video to be output .
[0116] The action detector performs two-dimensional pose estimation on the sequence of sampled frames to generate a two-dimensional skeleton sequence , where is the dimension of the bone joints, which is 17 in this example, represents two dimensions, indicating that each joint point has two coordinate values.
[0117] The input of the context learning module includes a query word composed of the two-dimensional skeleton sequence and the pseudo three-dimensional prior , as well as a prompt word ; where is a two-dimensional skeleton sequence randomly selected from the existing well-known skeleton dataset, and the is the corresponding three-dimensional skeleton sequence in the skeleton dataset to ; the is the average three-dimensional skeleton sequence obtained by calculating the average of a large number of three-dimensional skeleton sequences selected from the skeleton dataset.
[0118] The context learning module is composed of a bidirectional Transformer formed by a spatio-temporal module. The bidirectional Transformer includes input judgment and encoding, feature fusion and Transformer module processing, subsequent mapping and reconstruction, and an output process. The parameters in the context learning module are parameters to be trained.
[0119] Based on the query word and the prompt word , through the context learning module, the two-dimensional skeleton sequence is dimensionally enhanced to obtain a data-driven three-dimensional skeleton sequence .
[0120] The input of the physical information module is the data-driven three-dimensional skeleton sequence .
[0121] The physical information module includes an encoding fusion module, a physical parameter estimation module, an update module, and a decoding module.
[0122] The encoding fusion module encodes the entire sequence, the first three-frame sequence, and the last three-frame sequence of the three-dimensional skeleton sequence respectively, and obtains the time state sequences of the entire sequence, the first three-frame sequence, and the last three-frame sequence in the three-dimensional skeleton sequence respectively. Then it fuses the time state sequence of the entire sequence with the time state sequence of the first three-frame sequence to obtain a forward fusion sequence ; it fuses the time state sequence of the entire sequence with the time state sequence of the last three-frame sequence to obtain a reverse fusion sequence .
[0123] The physical parameter estimation module is divided into a forward physical parameter estimator and a reverse physical parameter estimator. Both the forward physical parameter estimator and the reverse physical parameter estimator are composed of four multi-layer perceptrons. Two of the multi-layer perceptrons respectively obtain a generalized force vector and a joint constraint vector, another multi-layer perceptron obtains a Gaussian noise vector with a variance of 1, then the Gaussian noise vectors are superimposed according to the required dimension, and random noise is added to each Gaussian noise vector to obtain a noise parameter matrix. The last multi-layer perceptron obtains an upper triangular matrix, and through a symmetric operation on the upper triangular matrix, a symmetric matrix is obtained. The noise parameter matrix is added to the symmetric matrix to obtain a generalized inverse inertia matrix. The parameters in the forward physical parameter estimator and the reverse physical parameter estimator are all parameters to be trained.
[0124] The update module is also divided into a forward update module and a reverse update module.
[0125] The forward update module uses the generalized inverse inertia matrix for forward update obtained by the forward physical parameter estimator , the generalized force vector and the joint constraint vector , through the forward updated Euler - Lagrange equation:
[0126]
[0127] obtain the second - order derivative of the forward fusion sequence , and then use the second - order central difference formula
[0128]
[0129] to obtain the forward update result, where is the time interval between adjacent frames; the forward update result is:
[0130]
[0131] The inverse update module uses the generalized inverse inertia matrix for inverse update obtained by the inverse physical parameter estimator , the generalized force vector and the joint constraint vector , through the inverse updated Euler - Lagrange equation:
[0132]
[0133] to obtain the second - order derivative of the inverse fusion sequence , and then use the second - order central difference formula
[0134]
[0135] to obtain the inverse update result, where is the time interval between adjacent frames; the inverse update result is:
[0136]
[0137] The first three frames of the forward fusion sequence only perform forward updates, and the last three frames of the inverse fusion sequence only perform inverse updates. For the remaining frames, take the average of the forward update result and the inverse update result.
[0138] The decoding module decodes the poses of the updated frames to finally obtain the physical prediction three - dimensional skeleton sequence .
[0139] The fusion guidance module includes a fusion module, a two - dimensional mapping module, and an encoding guidance module.
[0140] The fusion module combines the data - driven three - dimensional skeleton sequence obtained in step 2 Fuse with the physical prediction three-dimensional skeleton sequence obtained in step 3 The parameters of the fusion module are parameters to be trained.
[0141] The two-dimensional mapping module maps and projects the fusion result output by the fusion module back to two dimensions to obtain a two-dimensional mapping result. The parameters of the two-dimensional mapping module are parameters to be trained.
[0142] The two-dimensional mapping result obtained by the encoding guidance module generates a two-dimensional reconstructed skeleton sequence through encoding, and further obtains a multi-scale two-dimensional skeleton heat map sequence for guiding the 3D-UNet denoising process through fine-tuning, where the fine-tuning parameters are parameters to be trained.
[0143] The video generation module inputs the processing module, 3D-UNet, and output processing module.
[0144] The input processing module includes a text expander, a CLIP text encoder, and a video feature extraction and noise addition module; the text expander enhances the input text to make it easier for the model to understand. The enhanced data is encoded by the CLIP text encoder. In this embodiment, the text expander is a pre-trained large language model implemented using ChatGPT-4; the video feature extraction and noise addition module extracts features and adds noise to the sampled frame sequence of length extracted from the human action video and then inputs it together with the output of the CLIP text encoder into the 3D-UNet.
[0145] The 3D-UNet includes a downsampling module and an upsampling module. LoRA modules are embedded in both the downsampling module and the upsampling module. The LoRA module combines the multi-scale two-dimensional skeleton heat map sequence obtained in step 4 and updates the weight matrices and in the downsampling and upsampling processes with two low-rank matrices , where is the initial weight matrix; the two low-rank matrices and are parameters to be trained.
[0146] The output processing module decodes the output of the 3D-UNet to obtain a fine-grained human action video.
[0147] The parameters to be trained in each module are obtained through the following phased training process:
[0148] First, train the context learning module, the forward physical parameter estimator, the inverse physical parameter estimator, and the parameters to be trained in the fusion module. After training, freeze the parameters. The sample data used is a dataset with large-scale standard human three-dimensional postures such as Human3.6M and AMASS. The three-dimensional sequence obtained by the fusion module is represented as , then the loss function used in this stage of training is:
[0149]
[0150] where represents the th frame, represents the th human joint; represents the three-dimensional coordinates of the th th human joint in the three-dimensional sequence , represents the true three-dimensional coordinates of the th th human joint in the training sample data of this stage, .
[0151] Then, train the parameters to be trained in the two-dimensional mapping module and the fine-tuning parameters in the encoding guidance module. After training, freeze the parameters. The sample data used is the fine-grained human action videos provided by FineGym. Let the multi-scale two-dimensional skeleton heatmap sequence obtained after fine-tuning be , then the loss function used in training is:
[0152]
[0153] where is the two-dimensional coordinates of the th th human joint in , represents the true two-dimensional coordinates of the th th human joint in the training sample data of this stage.
[0154] Finally, train the parameters to be trained in the video generation module. The sample data used is the fine-grained human action videos provided by FineGym. The loss function used in training is:
[0155]
[0156] where is a random number that conforms to the normal distribution, is the set of noise addition step sizes, denotes the modulus of is the 3D-UNet model.
[0157] Experimental verification:
[0158] The experimental verification process extracted three fine-grained human action subsets - FX-JUMP, FXTURN, and FX-SALTO from FineGym and evaluated them. These subsets include challenging gymnastic movements performed by professional gymnasts.
[0159] The evaluation metrics are divided into automatic metrics and user studies:
[0160] Automatic metrics: PickScore is used to measure the alignment between video frames and text prompts, improved CLIP Domain and CLIP Smooth similarities are used to evaluate semantic similarity and embedding stability, and FréchetVideo Distance is used to evaluate video quality.
[0161] User study: The sensitivity and accuracy of the human body are used to evaluate the credibility of the movement, the retention of biological structures, and visual acceptability. Specifically, we showed a group of videos to the participants, including videos generated by our method and videos generated by the baseline method. For each video, the participants rated the consistency in the following aspects on a scale of 1 to 5: (1) text alignment; (2) domain consistency; (3) smooth stability. The final result is reported as the Mean Opinion Score (MOS).
[0162] Based on the above metrics, we evaluated FinePhys proposed in the present invention against other baseline methods, including Control-A-Video, VideoCrafter1 / 2, Text2Video-Zero, Latte, Follow-Your-Pose, and AnimateDiff, to generate fine-grained human actions. The results are shown in Table 1.
[0163] Table 1: Comparison with baseline methods on FineGym
[0164]
[0165] In Table 1, "T" in the input represents text prompts, while "P", "I", "D", and "C" represent pose, initial frame, depth map, and Canny edge detection respectively. It can be seen that in all cases, our FinePhys far outperforms various baselines under different conditions in terms of more reliable metrics, CLIPSIM*, and user studies. This indicates that FinePhys has excellent capabilities in understanding fine-grained human actions and generating more physically plausible actions.
[0166] To illustrate the importance of each module in FinePhys, ablation and analysis were carried out using the Human3.6M dataset and the FineGym dataset, as shown in Table 2:
[0167] Table 2: Importance of Each Module in FinePhys
[0168]
[0169] The metrics in Table 2 include mean per-joint position error (MPJPE), normalized MPJPE (N-MPJPE), and mean per-joint velocity error (MPJVE). To obtain the 2D skeleton sequence through the action detector; For the data-driven skeleton sequence obtained through the context learning module, corresponding to 3D and 2D respectively, where the 2D is obtained by projecting the 3D onto the 2D space; For the physical prediction skeleton sequence obtained through the physical information module, corresponding to 3D and 2D respectively, where the 2D is obtained by projecting the 3D onto the 2D space; For replacing the PhysNet module with a simple MLP.
[0170] As can be seen from Table 2, since the actions in Human3.6M mainly involve daily activities with moderate pose changes, the context learning module achieved good results ( . In addition, taking the average of and can reduce the estimation error, which indicates that the physically predicted pose can reduce the bias of data-driven estimation, thus verifying our design.
[0171] For the 2D space evaluation of the FineGym subset (without 3D pose annotations). By combining and Projected onto the 2D space, we obtain respectively and , and compare them with the online estimation result . As expected, performs poorly due to the drastic deformation of the human body and rapid temporal changes. Notably, exhibits higher accuracy, further verifying the necessity of the physical network module for fine-grained action understanding. Combining and yields the best results, highlighting the importance of each module in FinePhys.
[0172] In addition, the results of replacing PhysNet with a simple MLP are also given in Table 2, showing that such replacement leads to a significant performance degradation.
[0173] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.
Claims
1. A fine-grained human motion generation method based on physical laws and bone guidance, characterized in that: It includes the following steps: Step 1: Obtain a text description of a fine-grained human action video to be output ; According to the text description , obtain a human action video similar to the text description , extract a sampling frame sequence of length from the human action video , and use an action detector to perform two-dimensional pose estimation on the sampling frame sequence to generate a two-dimensional skeleton sequence ; Step 2: Perform dimensionality elevation on the two-dimensional skeleton sequence obtained in Step 1 through context learning to obtain a data-driven three-dimensional skeleton sequence ; the parameters in the context learning process are parameters to be trained; Step 3: Optimize the data-driven three-dimensional skeleton sequence obtained in Step 2 based on physical prediction to obtain a physically predicted three-dimensional skeleton sequence , and the specific process is as follows: Step 3.1: For the data-driven three-dimensional skeleton sequence , respectively obtain the temporal state sequences of the overall sequence, the first three-frame sequence, and the last three-frame sequence in the three-dimensional skeleton sequence through the encoder, and fuse the temporal state sequence of the overall sequence with the temporal state sequence of the first three-frame sequence to obtain the forward fusion sequence ; fuse the temporal state sequence of the overall sequence with the temporal state sequence of the last three-frame sequence to obtain the reverse fusion sequence ; Step 3.2: Based on the Euler-Lagrange equation, perform forward update on the forward fusion sequence and perform backward update on the backward fusion sequence where only the first three frames of the forward fusion sequence are forward updated, only the last three frames of the backward fusion sequence are backward updated, and the mean of the forward update result and the backward update result is taken for the remaining frames; Step 3.3: Data-driven three-dimensional skeleton sequence After each frame in is updated in Step 3.2, a physically predicted three-dimensional skeleton sequence is obtained through pose decoding; Step 4: The data-driven three-dimensional skeleton sequence obtained in Step 2 is fused with the physically predicted three-dimensional skeleton sequence obtained in Step 3 The fusion result is mapped and projected back to two dimensions, and encoded and fine-tuned into a multi-scale two-dimensional bone heatmap sequence for guiding the denoising process of 3D-UNet; Among them, the parameters of the fusion process, the mapping projection process, and the fine-tuning process are parameters to be trained; Step 5: Normalize the text description through a pre-trained large language model to obtain a text sequence , and then encode the text sequence to obtain a text encoding vector , is a pre-trained text encoding model; Extract the sampling frame sequence of length from the human action video for feature extraction to obtain a feature vector , and then perform noise addition processing on the feature vector to obtain a noise-added feature vector , where represents the number of noise addition steps; Input the noise-added feature vector , the text encoding vector , and the number of noise addition steps together into a 3D-UNet model that combines the multi-scale two-dimensional skeletal heat map sequence output in Step 4; Decode the output of the 3D-UNet model to obtain a fine-grained human action video; The parameters in the 3D-UNet model are parameters to be trained.
2. The method for generating fine-grained human actions based on physical laws and bone guidance according to claim 1, wherein: In step 3.2, The forward update process is as follows: Forward fusion sequence Input to the forward physical parameter estimator to estimate the parameters of the forward updated Euler-Lagrange equation, including the forward updated generalized inverse inertia matrix , the forward updated generalized force and the forward updated joint constraints ; And introduce the noise parameter matrix for forward update Added to the generalized inverse inertia matrix for forward update, the noise parameter matrix Is also obtained through the forward physical parameter estimator; thus, according to the forward updated Euler-Lagrange equation: Obtain the second derivative of the forward fusion sequence , and then use the second-order central difference formula Obtain the forward update result, where is the time interval between adjacent frames; the forward update result is: The backward update process is as follows: Input the reverse fusion sequence into the reverse physical parameter estimator to estimate the parameters of the reverse updated Euler-Lagrange equation, including the reverse updated generalized inverse inertia matrix , the reverse updated generalized force and the reverse updated joint constraint ; introduce the reverse updated noise parameter matrix and add it to the reverse updated generalized inverse inertia matrix. The noise parameter matrix is also obtained by the reverse physical parameter estimator; thus, according to the reverse updated Euler-Lagrange equation: Obtain the second derivative of the reverse fusion sequence , and then use the second-order central difference formula Obtain the reverse update result, where is the time interval between adjacent frames; the reverse update result is: The parameters in the forward physical parameter estimator and the backward physical parameter estimator are both parameters to be trained.
3. The method for generating fine-grained human body motions based on physical laws and skeleton guidance according to claim 2, characterized in that: The parameters to be trained are realized through the following phased training process: First, train the parameters required in Step 2, Step 3.2, and the fusion process of Step 4. After training is completed, freeze the parameters. The sample data used is a standard human body three-dimensional pose dataset, and the three-dimensional sequence obtained after the fusion in Step 4 is represented as , then the loss function used in this stage of training is: Among them represents the frame, represents the th human body joint; represents the th frame of the th human body joint three-dimensional coordinates in the three-dimensional sequence, represents the th th human body joint true three-dimensional coordinates of the frame in the training sample data of this stage, Among them, for the first three frames of adopt for the last three frames of adopt and for the remaining frames, adopt the mean value corresponding to and ; Next, train the mapping projection parameters and fine-tuning parameters in step 4, and freeze the parameters after training; the sample data used is fine-grained human action videos. Let the multi-scale two-dimensional bone heat map sequence obtained after fine-tuning in step 4 be , then the loss function used for training is: wherein is the two-dimensional coordinates of the th human body joint in the th frame, and represents the true two-dimensional coordinates of the th human body joint in the th frame in the training sample data at this stage; Finally, the parameters to be trained in step 5 are trained, and the sample data used is the fine-grained human body motion video, and the loss function used in the training is: wherein is a random number conforming to the normal distribution, is a set of noise addition step sizes, denotes the modulus of, is a 3D-UNet model.
4. The method for generating fine-grained human actions based on physical laws and bone guidance according to claim 1, wherein: The specific process of step 2 is: Step 2.1: Using the skeleton dataset and the two-dimensional skeleton sequence obtained in Step 1 , construct the prompt and the query ; where is a two-dimensional skeleton sequence randomly selected from the skeleton dataset, is the corresponding three-dimensional skeleton sequence in the skeleton dataset to ; is the average three-dimensional skeleton sequence, obtained by calculating the average of a large number of three-dimensional skeleton sequences selected from the skeleton dataset; Step 2.2: Based on the prompt and query words, perform dimensional transformation to obtain a data-driven three-dimensional skeleton sequence .
5. The method for generating fine-grained human actions based on physical laws and bone guidance according to claim 1, characterized in that: In step 2.2, the dimension transformation is realized by a bidirectional Transformer composed of spatio-temporal modules, and the parameters in the bidirectional Transformer are parameters to be trained.
6. The method for generating fine-grained human actions based on physical laws and bone guidance according to claim 2, characterized in that: In step 3.2, the forward physical parameter estimator and the backward physical parameter estimator are each composed of four multi-layer perceptrons. Among them, the generalized force and the joint constraint are vectors, and each is obtained through a multi-layer perceptron respectively; the generalized inverse inertia matrix is obtained in two steps. The first step is to obtain an upper triangular matrix through the third multi-layer perceptron, and then perform a symmetry operation on the upper triangular matrix to obtain a symmetric matrix. The second step is to add a noise parameter matrix to the symmetric matrix to obtain the generalized inverse inertia matrix; the noise parameter matrix is obtained in two steps. The first step is to obtain a Gaussian noise vector with a variance of 1 through the fourth multi-layer perceptron, and the second step is to stack the Gaussian noise vectors according to the set dimension and add random noise to each Gaussian noise vector to obtain the noise parameter matrix.
7. The method for generating fine-grained human motions based on physical laws and skeleton guidance according to claim 1, wherein: In step 5, during the downsampling and upsampling processes of the 3D-UNet model, the LoRA module that combines the multi-scale two-dimensional skeletal heat map sequence obtained in step 4 is used to update the weight matrices and in the downsampling and upsampling processes, where is the initial weight matrix; the two low-rank matrices and and are the parameters to be trained.
8. An electronic device, comprising a processor and a memory, the memory being used for storing one or more programs; characterized in that: When the one or more programs are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Human body action generation method based on diffusion model and fine-grained text description
CN119399332A
Method for tracing human body movement based on maximum geometric flow histogram
CN102663449A
Three-dimensional digital human driving method, medium and system
CN117237488A