Fine-grained human body action generation method based on physical law and skeleton guidance
Through the FinePhys method, using physical laws and bone guidance, FinePhys solves the problems of movement inconsistency and visual distortion of the human body movement generation model in the prior art when generating fine-grained human body movements, and achieves the generation of high consistency and physically credible fine-grained human body movement videos.
Patent Information
- Application Number
- CN202510575275.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
When the existing human movement generation model generates fine-grained human movements with severe human body deformation and significant timing changes, there are incoherence of movements, visual distortion, and abnormal structural characteristics that violate ergonomics, making it difficult to effectively learn and apply basic physical laws.
A fine-grained human body action generation method based on physical laws and bone guidance is proposed. By obtaining text descriptions and video skeleton sequences, context learning and physical prediction optimization are carried out, dynamic re-estimation is carried out in combination with the Euler-Lagrangian equation, a physically reliable three-dimensional skeleton sequence is generated, and fine-grained human body action video is generated through the 3D-UNet model.
The FinePhys method can generate fine-grained human body movement videos with high spatial and temporal consistency, significantly improving the generation quality of existing models on severe deformation and complex motion trajectories, and the generated actions are more natural and reasonable.
Smart Images

Figure CN120088442A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and computer vision, and specifically provides a fine-grained human action generation method based on physical laws and bone guidance. Background Art
[0002] In the field of generative model technology, breakthroughs represented by diffusion models have significantly promoted the development of human action video generation technology. For example, the published patent application number CN119399332A proposes a human action generation method based on a diffusion model and fine-grained text descriptions, which can generate zero-shot human actions beyond the scope of the original dataset and has good generalization ability.
[0003] However, the core elements involved in video temporal modeling - including camera motion trajectory control, dynamic background adaptation, and human action coherence - still pose important technical challenges. This challenge is particularly prominent in human action generation tasks, specifically manifested as the generated videos often exhibit inconsistent actions and visual distortions.
[0004] Analyzing from the spatial dimension, although the human body has strict anatomical structure constraints in the physical world, existing models often produce abnormal structural features that violate ergonomics during the processing. This lack of spatial consistency mainly stems from the insufficient ability of neural networks to represent complex human topological structures. At the time dynamics level, generating limb trajectories that conform to kinematic principles remains an urgent problem to be solved. The latest research shows that even the current most advanced generative models are difficult to effectively learn and apply basic physical laws including Newton's laws of motion during the generation process, which directly leads to obvious defects in the mechanical rationality of the generated actions.
[0005] These above problems cause existing methods to face dual-dimensional challenges when generating fine-grained human actions with significant human body deformations and temporal changes: in the spatial dimension, they need to cope with maintaining the topological structure under severe deformations, and in the time dimension, they need to meet the dynamic constraints of complex motion trajectories. For example, when generating gymnastic actions such as "turn 180° and exchange legs while jumping", existing methods generally have unstable generation quality, including serious time inconsistencies, obvious limb distortions, and abnormal human structures, as Figure 1 shown, indicating that the models designed by existing methods have not established an effective mechanism for internalizing physical laws. Summary of the Invention
[0006] Aiming at the problems existing in the prior art, the present invention proposes a fine-grained human action generation method based on physical laws and bone guidance, abbreviated as FinePhys, which can complete the human action generation task with significant human body deformations and temporal changes, such as Figure 1As shown, FinePhys has demonstrated excellent performance in generating physically plausible fine-grained human actions.
[0007] The technical solution of the present invention is as follows:
[0008] A method for generating fine-grained human actions based on physical laws and skeleton guidance, comprising the following steps:
[0009] Step 1: Obtain a text description of a fine-grained human action video to be output; ; According to the text description , obtain a human action video similar to the text description , extract a sampling frame sequence of length from the human action video , and use an action detector to perform two-dimensional pose estimation on the sampling frame sequence to generate a two-dimensional skeleton sequence ;
[0010] Step 2: Perform dimension elevation on the two-dimensional skeleton sequence obtained in Step 1 through context learning to obtain a data-driven three-dimensional skeleton sequence ; The parameters in the context learning process are parameters to be trained; ;
[0011] Step 3: Optimize the data-driven three-dimensional skeleton sequence obtained in Step 2 based on physical prediction to obtain a physically predicted three-dimensional skeleton sequence , and the specific process is as follows:
[0012] Step 3.1: For the data-driven three-dimensional skeleton sequence , respectively obtain the time state sequences of the overall sequence, the first three-frame sequence, and the last three-frame sequence in the three-dimensional skeleton sequence through an encoder, and fuse the time state sequence of the overall sequence with the time state sequence of the first three-frame sequence to obtain a forward fusion sequence ; Fuse the time state sequence of the overall sequence with the time state sequence of the last three-frame sequence to obtain a reverse fusion sequence ;
[0013] Step 3.2: Based on the Euler-Lagrange equation, perform forward update on the forward fusion sequence , and perform reverse update on the reverse fusion sequence , where only the first three frames of the forward fusion sequence are forward updated, only the last three frames of the reverse fusion sequence are reverse updated, and the average value of the forward update result and the reverse update result is taken for the remaining frames;
[0014] The forward update process is as follows:
[0015] Input the forward fusion sequence into the forward physical parameter estimator to estimate the parameters of the forward updated Euler - Lagrange equation, including the forward updated generalized inverse inertia matrix , the forward updated generalized force and the forward updated joint constraint ; and introduce the forward updated noise parameter matrix and add it to the forward updated generalized inverse inertia matrix. The noise parameter matrix is also obtained through the forward physical parameter estimator; thus, according to the forward updated Euler - Lagrange equation:
[0016]
[0017] obtain the second - order derivative of the forward fusion sequence , and then use the second - order central difference formula
[0018]
[0019] to obtain the forward updated result, where is the time interval between adjacent frames; the forward updated result is:
[0020]
[0021] The reverse update process is as follows:
[0022] Input the reverse fusion sequence into the reverse physical parameter estimator to estimate the parameters of the reverse updated Euler - Lagrange equation, including the reverse updated generalized inverse inertia matrix , the reverse updated generalized force and the reverse updated joint constraint ; introduce the reverse updated noise parameter matrix and add it to the reverse updated generalized inverse inertia matrix. The noise parameter matrix is also obtained through the reverse physical parameter estimator; thus, according to the reverse updated Euler - Lagrange equation:
[0023]
[0024] obtain the second - order derivative of the reverse fusion sequence , and then use the second - order central difference formula
[0025]
[0026] to obtain the reverse updated result, where is the time interval between adjacent frames; the reverse updated result is:
[0027]
[0028] The parameters in both the forward physical parameter estimator and the inverse physical parameter estimator are parameters to be trained;
[0029] Step 3.3: Data-driven three-dimensional skeleton sequence After each frame in the sequence is updated in Step 3.2, a physical prediction three-dimensional skeleton sequence is obtained through pose decoding; ;
[0030] Step 4: Fuse the data-driven three-dimensional skeleton sequence obtained in Step 2 with the physical prediction three-dimensional skeleton sequence obtained in Step 3 and project the fusion result back to two dimensions, and encode and fine-tune it into a multi-scale two-dimensional skeletal heatmap sequence for guiding the 3D-UNet denoising process; among them, the parameters in the fusion process, the mapping projection process, and the fine-tuning process are parameters to be trained;
[0031] Step 5: Normalize the text description through a pre-trained large language model to obtain a text sequence , and then encode the text sequence to obtain a text encoding vector , is the pre-trained text encoding model; extract a sampling frame sequence of length from the human action video to perform feature extraction to obtain a feature vector , and then perform noise addition processing on the feature vector to obtain a noise-added feature vector , where represents the number of noise addition steps; input the noise-added feature vector , the text encoding vector , and the number of noise addition steps together into a 3D-UNet model combined with the multi-scale two-dimensional skeletal heatmap sequence output in Step 4; decode the output of the 3D-UNet model to obtain a fine-grained human action video; the parameters in the 3D-UNet model are parameters to be trained.
[0032] Furthermore, the parameters to be trained are realized through the following phased training process:
[0033] First, train the parameters to be trained in Step 2, Step 3.2, and the fusion process in Step 4. After the training is completed, freeze the parameters; the sample data used is a standard human three-dimensional pose dataset, and the three-dimensional sequence obtained after the fusion in Step 4 is represented as , the loss function adopted in this stage of training is as follows:
[0034]
[0035] Among them represents the th frame, represents the th human body joint; represents the three-dimensional coordinates of the th human body joint in the three-dimensional sequence th frame, represents the true three-dimensional coordinates of the th human body joint in the th frame of the training sample data in this stage, Among them, for the first three frames, , adopts , for the last three frames, adopts , and for the remaining frames, adopts the mean value corresponding to and ;
[0036] Secondly, train the mapping projection parameters and fine-tuning parameters in step 4. After training is completed, freeze the parameters. The sample data used is the fine-grained human action video. Let the multi-scale two-dimensional bone heat map sequence obtained after fine-tuning in step 4 be , then the loss function adopted in the training is:
[0037]
[0038] Among them is the two-dimensional coordinates of the th human body joint in the th frame, represents the true two-dimensional coordinates of the th human body joint in the th frame of the training sample data in this stage, ;
[0039] Finally, train the parameters that need to be trained in step 5. The sample data used is the fine-grained human action video. The loss function adopted in the training is:
[0040]
[0041] Among them is a random number that conforms to the normal distribution, is the set of noise addition step sizes, represents modulus of, is the 3D-UNet model.
[0042] Further, the specific process of step 2 is as follows:
[0043] Step 2.1: Using the skeleton dataset and the two-dimensional skeleton sequence obtained in step 1 , construct a prompt and a query ; where is a two-dimensional skeleton sequence randomly selected from the skeleton dataset, is the corresponding three-dimensional skeleton sequence in the skeleton dataset to ; is the average three-dimensional skeleton sequence, obtained by calculating the average of a large number of three-dimensional skeleton sequences selected from the skeleton dataset;
[0044] Step 2.2: Based on the prompt and the query, perform dimensional transformation to obtain a data-driven three-dimensional skeleton sequence .
[0045] Further, in step 2.2, the dimensional transformation is implemented by a bidirectional Transformer composed of spatio-temporal modules, and the parameters in the bidirectional Transformer are parameters to be trained.
[0046] Further, in step 3.2, the forward physical parameter estimator and the inverse physical parameter estimator are each composed of four multi-layer perceptrons. Among them, the generalized force and the joint constraint are vectors, and each is obtained through a multi-layer perceptron; the generalized inverse inertia matrix is obtained in two steps. The first step is to obtain an upper triangular matrix through the third multi-layer perceptron, and then perform a symmetric operation on the upper triangular matrix to obtain a symmetric matrix. The second step is to add a noise parameter matrix to the symmetric matrix to obtain the generalized inverse inertia matrix; the noise parameter matrix is obtained in two steps. The first step is to obtain a Gaussian noise vector with a variance of 1 through the fourth multi-layer perceptron, and the second step is to stack the Gaussian noise vectors according to the set dimension and add random noise to each Gaussian noise vector to obtain the noise parameter matrix.
[0047] Further, in step 5, during the downsampling and upsampling processes of the 3D-UNet model, a LoRA module that combines the multi-scale two-dimensional bone heat map sequence obtained in step 4 is used to update the weight matrices and in the downsampling and upsampling processes with two low-rank matrices , where is the initial weight matrix; the two low-rank matrices and are parameters to be trained.
[0048] In addition, the present invention also provides an electronic device and a readable storage medium:
[0049] An electronic device includes a processor and a memory, and the memory is used to store one or more programs;
[0050] When the one or more programs are executed by the processor, the above method is implemented.
[0051] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0052] Beneficial effects:
[0053] The FinePhys method architecture proposed by the present invention first obtains relevant action videos in an online manner, and extracts 2D poses from the videos as prior information of the human physical structure; and realizes the transformation from 2D poses to 3D poses through a context learning method to enhance spatial perception, and obtains a data-driven three-dimensional skeleton sequence ; further designs a physical module PhysNet, uses the Euler-Lagrange formula to re-estimate human actions, and calculates joint accelerations bidirectionally in time series to generate physically predicted three-dimensional poses ; then, the physically estimated three-dimensional poses are fused with the data-driven three-dimensional skeleton sequence and converted into a 2D heatmap format to participate in 3D-UNets to realize video generation. The present invention is tested on 3 fine-grained action data subsets, and FinePhys far exceeds other video generation methods, and a large number of qualitative analyses further prove that FinePhys can generate more natural and reasonable fine-grained human actions.
[0054] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0055] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0056] Figure 1 : Video generation results of the fine-grained human action "180° body rotation and leg exchange jump".
[0057] Figure 2 : Schematic diagram of the FinePhys architecture proposed by the present invention.
[0058] Figure 3 : Schematic diagram of the PhysNet module proposed by the present invention. Detailed Embodiments
[0059] Embodiments of the present invention will be described in detail below. The embodiments are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0060] To address more challenging tasks: generating fine-grained human actions with drastic human deformations and significant temporal variations, this embodiment proposes a physical-law and bone-guided fine-grained human action generation method, abbreviated as FinePhys, for fine-grained human action generation, which can complete the task of generating human actions with drastic human deformations and significant temporal variations.
[0061] As Figure 2 shown, FinePhys, as a physics-aware framework, obtains relevant action videos in an online manner according to the input text description, and extracts 2D poses from the videos as prior information of the human physical structure; then uses a context learning module to convert the 2D poses to 3D poses to enhance spatial perception, obtaining a data-driven 3D skeleton sequence. . Since this data-driven 3D pose sequence often ignores the physical laws of motion, in order to incorporate the physical laws of motion, the PhysNet module is introduced in FinePhys, which uses the Euler-Lagrange equation of rigid body dynamics to encode Newtonian mechanics, and re-estimates the 3D positions of each human joint by considering the second-order time variations (i.e., accelerations) in the forward and backward directions, thereby generating physically predicted 3D poses. . Subsequently, and are fused, projected back to the 2D space, encoded into multi-scale heatmaps, and integrated into 3D-UNets to guide the denoising process, and the final output is a fine-grained action video. .
[0062] In the entire perception framework of FinePhys, this embodiment adopts three strategies: observation bias (through data), inductive bias (through the network), and learning bias (through loss), to coherently incorporate physical laws into the learning process. Specifically:
[0063] 1. For observation bias, FinePhys takes the pose as an additional modality encoding the biophysical layout, and uses the average 3D pose from existing datasets as a pseudo-3D reference to achieve 2D-to-3D enhancement through context learning.
[0064] 2. For inductive bias, the Lagrangian rigid body dynamics is instantiated through a fully differentiable neural network module PhysNet, encoding stronger inductive bias into FinePhys and outputting the parameters in the Euler-Lagrange equation.
[0065] 3. For the learning deviation, FinePhys adopts a loss function that follows the underlying physical process.
[0066] Example 1:
[0067] In this example, a fine-grained human action generation method based on physical laws and skeleton guidance is proposed, which specifically includes the following steps:
[0068] Step 1: Obtain a text description of a fine-grained human action video that is expected to be output. ;
[0069] According to the text description , obtain a human action video similar to the text description , for example, obtained through online search;
[0070] Extract from the human action video a sampling frame sequence of length as the prior information of the human physical structure;
[0071] Use an action detector to perform two-dimensional pose estimation on the sampling frame sequence to generate a two-dimensional skeleton sequence , where is the dimension of the bone joints, which is 17 in this example, represents two-dimensional, indicating that each joint point has two coordinate values. The two-dimensional skeleton sequence, as an additional mode, provides a compact and biologically structured reasonable description of human actions. The action detector is an existing technology in the art.
[0072] Step 2: In order to convert the two-dimensional layout of the human skeleton into a three-dimensional geometry, perform dimension elevation on the two-dimensional skeleton sequence obtained in Step 1 through context learning to obtain a data-driven three-dimensional skeleton sequence . The specific process is as follows:
[0073] Step 2.1: Since context learning requires some instances as prompt words, in this example, a widely used skeleton dataset, such as Human3.6M or AMASS, is used to construct prompt words; the prompt words are represented by input-output pairs: prompt words , where is a two-dimensional skeleton sequence randomly selected from the skeleton dataset, and the is the three-dimensional skeleton sequence corresponding to in the skeleton dataset;
[0074] Construct a query word for the two-dimensional skeleton sequence :
[0075] Calculate the average of a large number of three-dimensional skeleton sequences selected from the skeleton dataset, and use the obtained average three-dimensional skeleton sequence as the pseudo three-dimensional prior , in this embodiment, the average of all three-dimensional skeleton sequences in the Human3.6M skeleton dataset is calculated. Then a query word is constructed ;
[0076] Step 2.2: Based on the prompt word and the query word, perform dimensional transformation to obtain a data-driven three-dimensional skeleton sequence :
[0077]
[0078] Among them, the dimensional transformation is implemented by a bidirectional Transformer composed of spatio-temporal modules. The bidirectional Transformer includes input judgment and encoding, feature fusion and Transformer module processing, subsequent mapping and reconstruction, and output process, which are well-known technologies in the art. The parameters in the bidirectional Transformer are parameters to be trained.
[0079] Figure 2 The lower left shows the process of dimension elevation for the two-dimensional skeleton sequence through context learning. The entire two-dimensional to three-dimensional elevation process directly benefits from using the observed data and mastering the basic physical structures and rules, such as the relationships of limbs, the anatomical limitations of joints, the spatial layout of the three-dimensional human body, etc. through a trainable bidirectional Transformer.
[0080] Step 3: Since most online estimators are difficult to accurately estimate two-dimensional human actions, especially complex actions like gymnastics, noise is inevitably introduced. In addition, the data-driven two-dimensional to three-dimensional transformation lacks physical interpretability, resulting in the data-driven three-dimensional skeleton sequence often having unreliable estimation results. To solve these problems, in this step, the three-dimensional motion dynamics are re-estimated by bidirectionally calculating the second-order time difference change (i.e., acceleration) using the complete Euler-Lagrange equation, and the physical terms in the Euler-Lagrange equation are estimated to obtain the time-sequence variation amount, so as to realize the optimization of the data-driven three-dimensional skeleton sequence obtained in Step 2 based on physical prediction, and obtain a physically predicted three-dimensional skeleton sequence . As Figure 3 shown, the specific process is as follows
[0081] Step 3.1: For the data-driven three-dimensional skeleton sequence , respectively obtain the three-dimensional skeleton sequence through the encoder The time state sequences of the overall sequence, the first three-frame sequence, and the last three-frame sequence, and fuse the time state sequence of the overall sequence with the time state sequence of the first three-frame sequence to obtain a forward fusion sequence ; fuse the time state sequence of the overall sequence with the time state sequence of the last three-frame sequence to obtain a reverse fusion sequence ; all subsequent calculations are carried out bidirectionally, with forward updates starting from the first three frames and reverse updates starting from the last three frames;
[0082] Step 3.2: Based on the Euler-Lagrange equation, perform forward updates on the forward fusion sequence and perform reverse updates on the reverse fusion sequence , where only forward updates are performed on the first three frames of the forward fusion sequence, only reverse updates are performed on the last three frames of the reverse fusion sequence, and the average value of the forward update result and the reverse update result is taken for the remaining frames;
[0083] The forward update process is as follows:
[0084] Input the forward fusion sequence into the forward physical parameter estimator to estimate the parameters of the Euler-Lagrange equation for forward updates, including the forward-updated generalized inverse inertia matrix , the forward-updated generalized force , and the forward-updated joint constraint ; and since temporal motion often destroys the symmetry of the object, therefore, in this embodiment, a forward-updated noise parameter matrix is also introduced and added to the forward-updated generalized inverse inertia matrix, and the noise parameter matrix is also obtained through the forward physical parameter estimator; thus, the second derivative
[0085]
[0086] of the forward fusion sequence can be obtained, and then the forward update result can be obtained using the second-order central difference formula
[0087]
[0088] where is the time interval between adjacent frames; the forward update result is:
[0089]
[0090] The reverse update process is as follows:
[0091] Input the reverse fusion sequence Input a reverse physical parameter estimator to estimate the parameters of the reverse updated Euler-Lagrange equation, including the reverse updated generalized inverse inertia matrix , the reverse updated generalized force and the reverse updated joint constraint ; During the reverse update process, a noise parameter matrix is also introduced and added to the reverse updated generalized inverse inertia matrix. The noise parameter matrix is also obtained through the reverse physical parameter estimator; thus, according to the reverse updated Euler-Lagrange equation:
[0092]
[0093] the second derivative of the reverse fusion sequence can be obtained , and then the second-order central difference formula
[0094]
[0095] is used to obtain the reverse update result, where is the time interval between adjacent frames; the reverse update result is:
[0096]
[0097] In this embodiment, both the forward physical parameter estimator and the reverse physical parameter estimator are respectively composed of four multi-layer perceptrons, as shown in Figure 3 . Among them, the generalized force and the joint constraint are vectors, so they can be obtained respectively through one multi-layer perceptron; however, due to the high dimension, directly estimating the generalized inverse inertia matrix is very challenging. Therefore, based on the prior knowledge that the inertia tensor and the mass matrix are usually symmetric in the structural system in this embodiment, the generalized inverse inertia matrix is obtained in two steps. The first step is to obtain an upper triangular matrix through one multi-layer perceptron, then perform a symmetry operation on the upper triangular matrix to obtain a symmetric matrix, and then add a noise parameter matrix to the symmetric matrix to obtain the generalized inverse inertia matrix; since the noise parameter matrix is also a high-dimensional matrix, for this reason, the process of obtaining the noise parameter matrix is also divided into two steps. The first step is to obtain a Gaussian noise vector with a variance of 1 through one multi-layer perceptron, then stack the Gaussian noise vectors according to the required dimension, and add random noise to each Gaussian noise vector to obtain the noise parameter matrix; the parameters in the forward physical parameter estimator and the reverse physical parameter estimator are all parameters that need to be trained;
[0098] Step 3.3: After each frame in the data-driven three-dimensional skeleton sequence is updated through Step 3.2, physical prediction three-dimensional skeleton sequence is obtained through pose decoding.
[0099] Step 4: The data-driven three-dimensional skeleton sequence obtained in Step 2 is fused with the physically predicted three-dimensional skeleton sequence obtained in Step 3 The fused result is mapped and projected back to two dimensions, encoded, and fine-tuned into a multi-scale two-dimensional skeletal heatmap sequence for guiding the denoising process of 3D-UNet.
[0100] In this step, and The fusion process can be achieved by various methods well-known to those skilled in the art, such as weighted summation, etc., where the weights are parameters to be trained; the mapping and projection parameters in the process of mapping and projecting the obtained fusion result back to two dimensions are also parameters to be trained; the obtained two-dimensional mapping result is encoded to generate a two-dimensional reconstructed skeleton sequence, and further fine-tuned to obtain a multi-scale two-dimensional skeletal heatmap sequence for guiding the denoising process of 3D-UNet, where the fine-tuning parameters are parameters to be trained;
[0101] Step 5: The text description is normalized by a pre-trained large language model to obtain a text sequence In this embodiment, the pre-trained large language model uses ChatGPT-4, and then the text sequence is text-encoded to obtain a text encoding vector , is a pre-trained text encoding model; the sampled frame sequence of length extracted from the human action video is feature-extracted to obtain a feature vector , and then the feature vector is noise-added to obtain a noise-added feature vector , where represents the number of noise-adding steps; the noise-added feature vector , the text encoding vector and the number of noise-adding steps are jointly input into the 3D-UNet model; the output of the 3D-UNet model is decoded to obtain a fine-grained human action video;
[0102] In the downsampling and upsampling processes of the 3D-UNet model, a LoRA module combined with the multi-scale two-dimensional skeletal heatmap sequence obtained in Step 4 is used to update the weight matrices and in the downsampling and upsampling processes with two low-rank matrices , where is the initial weight matrix; the two low-rank matrices and are parameters to be trained.
[0103] When the above steps are implemented, all the parameters to be trained have been completed. The training process is carried out in stages:
[0104] First, train the parameters to be trained in step 2.3, step 3.2, and the fusion process of step 4. After training is completed, freeze the parameters; the sample data used is a dataset with large-scale standard human 3D postures such as Human3.6M and AMASS. The 3D sequence obtained after the fusion of step 4 is represented as , then the loss function used in this stage of training is:
[0105]
[0106] where represents the th frame, represents the th human joint; represents the 3D coordinates of the th th human joint in the 3D sequence , represents the true 3D coordinates of the th th human joint in the training sample data of this stage, , which is used to limit the noise parameter matrix in step 3.2. Among them, for the first three frames is adopted, for the last three frames is adopted, for the remaining frames and the mean value of is adopted.
[0107] Secondly, train the mapping projection parameters and fine-tuning parameters in step 4. After training is completed, freeze the parameters; the sample data used is the fine-grained human action video provided by FineGym. Let the multi-scale two-dimensional bone heat map sequence obtained after the fine-tuning of step 4 be , then the loss function used in the training is:
[0108]
[0109] where is the th th th two-dimensional coordinates of the human joint in, represents the true two-dimensional coordinates of the th th human joint in the training sample data of this stage.
[0110] Finally, train the parameters to be trained in step 5. The sample data used is the fine-grained human action video provided by FineGym. The loss function used in training is as follows:
[0111]
[0112] where is a random number that conforms to the normal distribution, is the set of noise addition step sizes, represents the modulus of, is the 3D-UNet model.
[0113] Example 2:
[0114] In this example, a fine-grained human action generation system based on physical laws and skeleton guidance is proposed, which specifically includes an action detector, a context learning module, a physical information module, a fusion guidance module, and a video generation module;
[0115] The input of the action detector is a sequence of sampling frames with a length of ; the sequence of sampling frames is extracted from the human action video , and the human action video is a human action video similar to the text , and the text is a text provided by the user describing the fine-grained human action video that is expected to be output .
[0116] The action detector performs two-dimensional pose estimation on the sequence of sampling frames to generate a two-dimensional skeleton sequence , where is the dimension of the bone joint, which is 17 in this example, represents two-dimensional, indicating that each joint point has two coordinate values.
[0117] The input of the context learning module includes the query word composed of the two-dimensional skeleton sequence and the pseudo-three-dimensional prior , as well as the prompt word ; where is a two-dimensional skeleton sequence randomly selected from the existing well-known skeleton datasets, and the is the three-dimensional skeleton sequence corresponding to in the skeleton dataset; the is the average three-dimensional skeleton sequence obtained by calculating the average of a large number of three-dimensional skeleton sequences selected from the skeleton dataset.
[0118] The context learning module is composed of a bidirectional Transformer formed by a spatio-temporal module. The bidirectional Transformer includes input judgment and encoding, feature fusion and Transformer module processing, subsequent mapping and reconstruction, and an output process. The parameters in the context learning module are parameters to be trained.
[0119] Based on the query word and the prompt word , through the context learning module, the two-dimensional skeleton sequence is dimensionally enhanced to obtain a data-driven three-dimensional skeleton sequence .
[0120] The input of the physical information module is the data-driven three-dimensional skeleton sequence .
[0121] The physical information module includes an encoding fusion module, a physical parameter estimation module, an update module, and a decoding module.
[0122] The encoding fusion module encodes the overall sequence, the first three-frame sequence, and the last three-frame sequence of the three-dimensional skeleton sequence respectively, and obtains the time state sequences of the overall sequence, the first three-frame sequence, and the last three-frame sequence in the three-dimensional skeleton sequence respectively. Then it fuses the time state sequence of the overall sequence with the time state sequence of the first three-frame sequence to obtain a forward fusion sequence ; it fuses the time state sequence of the overall sequence with the time state sequence of the last three-frame sequence to obtain a reverse fusion sequence .
[0123] The physical parameter estimation module is divided into a forward physical parameter estimator and a reverse physical parameter estimator. Both the forward physical parameter estimator and the reverse physical parameter estimator are composed of four multi-layer perceptrons. Two of the multi-layer perceptrons respectively obtain a generalized force vector and a joint constraint vector, another multi-layer perceptron obtains a Gaussian noise vector with a variance of 1, then the Gaussian noise vectors are superimposed according to the required dimension, and random noise is added to each Gaussian noise vector to obtain a noise parameter matrix. The last multi-layer perceptron obtains an upper triangular matrix, and through a symmetric operation on the upper triangular matrix, a symmetric matrix is obtained. The noise parameter matrix is added to the symmetric matrix to obtain a generalized inverse inertia matrix. The parameters in the forward physical parameter estimator and the reverse physical parameter estimator are all parameters to be trained.
[0124] The update module is also divided into a forward update module and a reverse update module.
[0125] The forward update module uses the generalized inverse inertia matrix for forward update obtained by the forward physical parameter estimator , the generalized force vector and the joint constraint vector , through the forward updated Euler - Lagrange equation:
[0126]
[0127] obtain the second - order derivative of the forward fusion sequence , and then use the second - order central difference formula
[0128]
[0129] to obtain the forward update result, where is the time interval between adjacent frames; the forward update result is:
[0130]
[0131] The reverse update module uses the generalized inverse inertia matrix for reverse update obtained by the reverse physical parameter estimator , the generalized force vector and the joint constraint vector , through the reverse updated Euler - Lagrange equation:
[0132]
[0133] to obtain the second - order derivative of the reverse fusion sequence , and then use the second - order central difference formula
[0134]
[0135] to obtain the reverse update result, where is the time interval between adjacent frames; the reverse update result is:
[0136]
[0137] The first three frames of the forward fusion sequence only perform forward updates, and the last three frames of the reverse fusion sequence only perform reverse updates. For the remaining frames, take the average of the forward update result and the reverse update result.
[0138] The decoding module decodes the poses of the updated frames to finally obtain the physical prediction three - dimensional skeleton sequence .
[0139] The fusion guidance module includes a fusion module, a two - dimensional mapping module, and an encoding guidance module.
[0140] The fusion module combines the data - driven three - dimensional skeleton sequence obtained in step 2 Fuse with the physical prediction three-dimensional skeleton sequence obtained in step 3 The parameters of the fusion module are parameters to be trained.
[0141] The two-dimensional mapping module maps and projects the fusion result output by the fusion module back to two dimensions to obtain a two-dimensional mapping result. The parameters of the two-dimensional mapping module are parameters to be trained.
[0142] The two-dimensional mapping result obtained by the encoding guidance module generates a two-dimensional reconstructed skeleton sequence through encoding, and further obtains a multi-scale two-dimensional skeleton heat map sequence for guiding the 3D-UNet denoising process through fine-tuning, where the fine-tuning parameters are parameters to be trained.
[0143] The video generation module inputs the processing module, 3D-UNet, and output processing module.
[0144] The input processing module includes a text expander, a CLIP text encoder, and a video feature extraction and noise addition module; the text expander enhances the input text to make it easier for the model to understand. The enhanced data is encoded by the CLIP text encoder. In this embodiment, the text expander is a pre-trained large language model implemented by ChatGPT-4; the video feature extraction and noise addition module extracts features and adds noise to the sampled frame sequence of length extracted from the human action video and then inputs it together with the output of the CLIP text encoder into the 3D-UNet.
[0145] The 3D-UNet includes a downsampling module and an upsampling module. LoRA modules are embedded in both the downsampling module and the upsampling module. The LoRA module combines the multi-scale two-dimensional skeleton heat map sequence obtained in step 4 and updates the weight matrices and in the downsampling and upsampling processes with two low-rank matrices , where is the initial weight matrix; the two low-rank matrices and are parameters to be trained.
[0146] The output processing module decodes the output of the 3D-UNet to obtain a fine-grained human action video.
[0147] The parameters to be trained in each module are obtained through the following staged training process:
[0148] First, train the context learning module, the forward physical parameter estimator, the inverse physical parameter estimator, and the parameters to be trained in the fusion module. After training, freeze the parameters. The sample data used is a dataset with large-scale standard human three-dimensional postures such as Human3.6M and AMASS. The three-dimensional sequence obtained by the fusion module is represented as , then the loss function used in this stage of training is:
[0149]
[0150] where represents the th frame, represents the th human joint; represents the three-dimensional coordinates of the th human joint in the th frame of the three-dimensional sequence , represents the true three-dimensional coordinates of the th human joint in the th frame of the sample data for this stage of training, .
[0151] Then, train the parameters to be trained in the two-dimensional mapping module and the fine-tuning parameters in the encoding guidance module. After training, freeze the parameters. The sample data used is the fine-grained human action videos provided by FineGym. Let the multi-scale two-dimensional bone heat map sequence obtained after fine-tuning be , then the loss function used in training is:
[0152]
[0153] where is the two-dimensional coordinates of the th human joint in the th frame, represents the true two-dimensional coordinates of the th human joint in the th frame of the sample data for this stage of training, .
[0154] Finally, train the parameters to be trained in the video generation module. The sample data used is the fine-grained human action videos provided by FineGym. The loss function used in training is:
[0155]
[0156] where is a random number that conforms to the normal distribution, is the set of noise addition step sizes, denotes the modulus of is the 3D-UNet model.
[0157] Experimental verification:
[0158] In the experimental verification process, three fine-grained human action subsets - FX-JUMP, FXTURN, and FX-SALTO were extracted from FineGym and evaluated within them. These subsets include challenging gymnastic movements performed by professional gymnasts.
[0159] The evaluation metrics are divided into automatic metrics and user studies:
[0160] Automatic metrics: PickScore is used to measure the alignment between video frames and text prompts, improved CLIP Domain and CLIP Smooth similarities are used to evaluate semantic similarity and embedding stability, and FréchetVideo Distance is used to evaluate video quality.
[0161] User study: The sensitivity and accuracy of the human body are utilized to evaluate the credibility of the movement, the retention of the biological structure, and visual acceptability. Specifically, we presented a set of videos to the participants, including videos generated by our method and videos generated by the baseline method. For each video, the participants rated the consistency in the following aspects on a scale of 1 to 5: (1) text alignment; (2) domain consistency; (3) smooth stability. The final result is reported as the Mean Opinion Score (MOS).
[0162] Based on the above metrics, we evaluated FinePhys proposed in the present invention against other baseline methods, including Control-A-Video, VideoCrafter1 / 2, Text2Video-Zero, Latte, Follow-Your-Pose, and AnimateDiff, to generate fine-grained human actions. The results are shown in Table 1.
[0163] Table 1: Comparison with baseline methods on FineGym
[0164]
[0165] In Table 1, "T" in the input represents text prompts, while "P", "I", "D", and "C" represent pose, initial frame, depth map, and Canny edge detection respectively. It can be seen that in all cases, our FinePhys far outperforms various baselines under different conditions in terms of more reliable metrics, CLIPSIM*, and user studies. This indicates that FinePhys has excellent capabilities in understanding fine-grained human actions and generating more physically plausible actions.
[0166] To illustrate the importance of each module in FinePhys, ablation and analysis were carried out using the Human3.6M dataset and the FineGym dataset, as shown in Table 2:
[0167] Table 2: Importance of Each Module in FinePhys
[0168]
[0169] The metrics in Table 2 include mean per-joint position error (MPJPE), normalized MPJPE (N-MPJPE), and mean per-joint velocity error (MPJVE). To obtain the 2D skeleton sequence through the action detector; The data-driven skeleton sequence obtained through the context learning module, corresponding to 3D and 2D respectively, where the 2D is obtained by projecting the 3D onto the 2D space; The physical prediction skeleton sequence obtained through the physical information module, corresponding to 3D and 2D respectively, where the 2D is obtained by projecting the 3D onto the 2D space; Replace the PhysNet module with a simple MLP.
[0170] As can be seen from Table 2, since the actions in Human3.6M mainly involve daily activities with moderate pose changes, the context learning module achieved good results ( . In addition, taking the average of and can reduce the estimation error, which indicates that the physical prediction pose can reduce the bias of data-driven estimation, thus verifying our design.
[0171] Perform 2D space evaluation on the FineGym subset (without 3D pose annotations). By combining and Projected onto the 2D space, we obtain and , respectively, and compare them with the online estimation result . As expected, performs poorly due to the drastic deformation of the human body and rapid temporal changes. Notably, exhibits higher accuracy, further validating the necessity of the physical network module for fine-grained action understanding. Combining and yields the best results, highlighting the importance of each module in FinePhys.
[0172] In addition, the results of replacing PhysNet with a simple MLP are also given in Table 2, showing that such a replacement leads to a significant performance degradation.
[0173] Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.
Claims
1. A fine-grained human motion generation method based on physical laws and skeleton guidance, characterized by: The following steps are involved: Step 1: Get a text description of the fine-grained human action video you want to output ; According to the text description , get a human action video that is close to the text description , from the human action video The length of the extraction is The sampling frame sequence is subjected to two-dimensional posture estimation by using an action detector to generate a two-dimensional skeleton sequence. ; Step 2: Use context learning to learn the 2D skeleton sequence obtained in step 1 Perform dimension enhancement to obtain a data-driven 3D skeleton sequence ; The parameters in the context learning process are parameters that need to be trained; Step 3: 3D skeleton sequence driven by data from step 2 based on physical prediction Optimize and obtain the physical prediction three-dimensional skeleton sequence , the specific process is: Step 3.1: For data-driven 3D skeleton sequence , and obtain the 3D skeleton sequence through the encoder The time state sequence of the overall sequence, the first three frame sequence and the last three frame sequence in the forward fusion sequence is obtained by fusing the time state sequence of the overall sequence with the time state sequence of the first three frame sequence. ; Fuse the time state sequence of the overall sequence with the time state sequence of the last three frame sequences to obtain the reverse fusion sequence ; Step 3.2: Based on the Euler-Lagrange equation, the forward fusion sequence Perform forward update to reverse fusion sequence Perform reverse update, where the first three frames of the forward fusion sequence are only forward updated, and the last three frames of the reverse fusion sequence are only reverse updated, and the remaining frames take the average of the forward update result and the reverse update result; Step 3.3: Data-driven 3D skeleton sequence After each frame in step 3.2 is updated, the physical prediction 3D skeleton sequence is obtained through posture decoding. ; Step 4: The data-driven 3D skeleton sequence obtained in step 2 The physical prediction 3D skeleton sequence obtained in step 3 Fusion is performed, the fusion result map is projected back to two dimensions, and encoded and fine-tuned into a multi-scale two-dimensional bone heat map sequence for guiding the 3D-UNet denoising process; The parameters of the fusion process, mapping projection process and fine-tuning process are the parameters that need to be trained; Step 5: Describe the text The text sequence is obtained by normalizing the pre-trained large language model , and then for the text sequence Perform text encoding to obtain the text encoding vector , is a pre-trained text encoding model; it will be extracted from the human action video The length of the extraction is The sampling frame sequence is used for feature extraction to obtain the feature vector , and then the eigenvector Perform noise processing to obtain the feature vector after noise addition ,in Represents the number of noise adding steps; the feature vector after noise addition , text encoding vector And the number of noise steps The 3D-UNet model that combines the multi-scale two-dimensional bone heat map sequence output from step 4 is jointly input; the output of the 3D-UNet model is decoded to obtain a fine-grained human action video; the parameters in the 3D-UNet model are the parameters that need to be trained.
2. According to claim 1, a fine-grained human motion generation method based on physical laws and skeleton guidance is characterized by: In step 3.2, The forward update process is: The forward fusion sequence Input the forward physical parameter estimator to estimate the parameters of the forward-updated Euler-Lagrange equations, including the forward-updated generalized inverse inertia matrix , the generalized force of forward update And the joint constraints of the forward update ; And introduce the noise parameter matrix of forward update Added to the generalized inverse inertia matrix of the forward update, the noise parameter matrix is also obtained through the forward physical parameter estimator; thus according to the forward updated Euler-Lagrange equation: Get the second-order derivative of the forward fusion sequence , and then use the second-order central difference formula Get the forward update result, where is the time interval between adjacent frames; the forward update result is: The reverse update process is: Reverse fusion sequence Input the inverse physical parameter estimator to estimate the parameters of the inverse updated Euler-Lagrange equation, including the inverse updated generalized inverse inertia matrix , the generalized force of reverse update And the joint constraints updated inversely ; Introduce the noise parameter matrix for inverse update Added to the generalized inverse inertia matrix of the inverse update, the noise parameter matrix It is also obtained by the inverse physical parameter estimator; thus according to the inverse updated Euler-Lagrange equation: Get the second-order derivative of the reverse fusion sequence , and then use the second-order central difference formula Get the reverse update result, where is the time interval between adjacent frames; the reverse update result is: The parameters in the forward physical parameter estimator and the reverse physical parameter estimator are parameters that need to be trained.
3. According to claim 2, a fine-grained human motion generation method based on physical laws and skeleton guidance is characterized by: The training parameters required are achieved through the following phased training process: First, train the parameters needed for step 2, step 3.2, and step 4 fusion, and freeze the parameters after training. The sample data used is the standard human 3D posture dataset. The 3D sequence obtained after fusion in step 4 is expressed as , then the loss function used in this stage of training is: in Indicates frame, Indicates individual body joints; Represents a three-dimensional sequence Middle Frame No. The three-dimensional coordinates of individual body joints, Indicates the first Frame No. The real 3D coordinates of individual body joints, , where the first three frames use , the last three frames use , the remaining frames Adopt corresponding and The mean of Secondly, the mapping projection parameters and fine-tuning parameters in step 4 are trained, and the parameters are frozen after the training is completed; the sample data used is a fine-grained human action video, and the multi-scale two-dimensional bone heat map sequence obtained after fine-tuning in step 4 is assumed to be , then the loss function used in training is: in for Middle Frame No. The 2D coordinates of individual body joints, Indicates the first Frame No. The real 2D coordinates of individual body joints; Finally, the training parameters required in step 5 are trained. The sample data used is fine-grained human action video, and the loss function used in training is: in is a random number that conforms to the normal distribution. is the noise step set, express The model, It is a 3D-UNet model.
4. The method for generating fine-grained human motion based on physical laws and skeleton guidance according to claim 1, characterized in that: The specific process of step 2 is: Step 2.1: Use the skeleton dataset and the 2D skeleton sequence obtained in step 1 , construct the prompt word and query terms ;in is a two-dimensional skeleton sequence randomly selected from the skeleton dataset, The skeleton data set is The corresponding three-dimensional skeleton sequence; is an average three-dimensional skeleton sequence, obtained by calculating the average value of a large number of three-dimensional skeleton sequences selected from the skeleton data set; Step 2.2: Based on the prompt words and query words, perform dimension transformation to obtain a data-driven 3D skeleton sequence .
5. The method for generating fine-grained human motion based on physical laws and skeleton guidance according to claim 1, characterized in that: In step 2.2, the dimensional transformation is achieved by a bidirectional Transformer composed of spatiotemporal modules, and the parameters in the bidirectional Transformer are parameters that need to be trained.
6. The method for generating fine-grained human motion based on physical laws and skeleton guidance according to claim 2, characterized in that: In step 3.2, the forward physical parameter estimator and the reverse physical parameter estimator are respectively composed of four multilayer perceptrons, where the generalized force and joint constraints are vectors, each of which is obtained by a multilayer perceptron; the generalized inverse inertia matrix is obtained in two steps, the first step is to obtain the upper triangular matrix through the third multilayer perceptron, and then perform symmetric operations on the upper triangular matrix to obtain a symmetric matrix, and the second step is to add a noise parameter matrix to the symmetric matrix to obtain the generalized inverse inertia matrix; the noise parameter matrix is obtained in two steps, the first step is to obtain a Gaussian noise vector with a variance of 1 through the fourth multilayer perceptron, and the second step is to superimpose the Gaussian noise vectors according to the set dimension, and add random noise to each Gaussian noise vector to obtain the noise parameter matrix.
7. The method for generating fine-grained human motion based on physical laws and skeleton guidance according to claim 1, characterized in that: In step 5, during the downsampling and upsampling process of the 3D-UNet model, the LoRA module combined with the multi-scale two-dimensional bone heat map sequence obtained in step 4 is used to generate two low-rank matrices and Update the weight matrix during downsampling and upsampling ,in is the initial weight matrix; two low-rank matrices and Requires training parameters.
8. An electronic device, comprising a processor and a memory, wherein the memory is used to store one or more programs; characterized in that: When the one or more programs are executed by the processor, the method described in any one of claims 1 to 7 is implemented.
9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Human body action generation method based on diffusion model and fine-grained text description
CN119399332A
Method for tracing human body movement based on maximum geometric flow histogram
CN102663449A
Three-dimensional digital human driving method, medium and system
CN117237488A
Fine-grained action recognition method based on cross-modal knowledge alignment
CN118196888A
Action Recognition Using Implicit Pose Representations
US20210073525A1