Text-aligned human motion generation method and system
By combining the quantitative variational autoencoder and the bidirectional mask Transformer model, semantic alignment loss is calculated, and the consistency problem at the motion sequence level in text-driven human motion generation is solved, achieving high-quality motion generation.
Patent Information
- Application Number
- CN202510417769.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
AI Technical Summary
In the text-driven human motion generation, it is difficult to maintain consistency between text input and generated motion at the entire motion sequence level, resulting in the generation results that are inconsistent with expectations.
The quantitative variational autoencoder model is used to combine the bidirectional mask Transformer model to optimize the parameter through the training framework, calculate the semantic alignment loss, and enhance the motion perception and generated semantic alignment.
The generated human movement is achieved more in line with the original data distribution, more consistent with the text description, and improves the quality of motion generation, especially when facing longer and complex text inputs.
Smart Images

Figure CN119941942A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and deep learning, and specifically relates to a method and system for generating human motion for text alignment. Background Art
[0002] The text-driven 3D human motion generation task can be defined as: for the input text , generating a length that is consistent with the semantic content of the text 3D human pose sequence ,, in , Represents the dimension of human pose data in a single frame. This task can be applied in multiple fields, including filmmaking, video games, AR / VR, etc. In previous text-driven human motion generation works such as T2M-GPT and MotionCLIP, pre-trained models such as CLIP trained on large-scale image and text datasets are usually used as text encoders to provide text guidance for human motion generation tasks. However, their supervision is usually performed at the level of a single frame, which cannot well reflect the consistency between the user's text input and the generated complete human motion at the entire sequence level. The lack of an overall understanding of the entire motion sequence often leads to the problem that the generated human motion is inconsistent with the text description, which greatly limits the model's ability to generate human motion consistent with the text. In addition, the current mainstream work using Transformer for human motion generation (such as T2M-GPT) usually adopts an autoregressive method to predict and generate human motion tokens in a unidirectional order. This method inevitably brings challenges to motion generation. On the one hand, in the unidirectional prediction process, the model can only predict the motion content based on the previous text rather than the global context, which limits the expressiveness of the model; on the other hand, this method will lead to the continuous accumulation of errors in the prediction generation process, and ultimately cause the generation results to be inconsistent with expectations. Summary of the invention
[0003] The purpose of the present invention is to solve the above problems existing in the prior art and to provide a method and system for generating human body motion for text alignment.
[0004] In a first aspect, the present invention provides a method for generating human motion for text alignment, comprising:
[0005] S1. Training a quantized variational autoencoder model for reconstructing human motion, wherein the quantized variational autoencoder model comprises an encoder network and a decoder network. An input human motion sequence is converted into a potential vector sequence through the encoder network, and then quantized using a codebook to obtain a first codeword sequence and a first tag sequence composed of codeword index tags. The first codeword sequence is input into the decoder network to reconstruct the human motion sequence.
[0006] S2, after freezing the parameters of the decoder network and the pre-trained text-to-motion cross-modal retrieval model including a text encoder and a motion encoder, they are trained in a training framework formed with a bidirectional masked Transformer model; during the training process, the input text is encoded into a text vector by the text encoder, and then input into the bidirectional masked Transformer model together with the first tag sequence after random masking, the second tag sequence is predicted and the tag prediction loss is calculated, and then the second tag sequence is converted into a second codeword sequence by using a codebook, and then input into the decoder network to reconstruct the human motion sequence, and then the motion vector is further obtained by the motion encoder, and the semantic alignment loss of the text vector and the motion vector is calculated, and finally the two losses are weighted summed and back-propagated to optimize the model parameters;
[0007] S3. The motion description text input by the user is first encoded into a text vector by the text encoder and input into the trained bidirectional masked Transformer model together with a fully masked tag sequence. A complete tag sequence is generated through a multi-step iterative strategy, and the complete tag sequence is converted into a codeword sequence using a codebook, and then the decoder network generates a human motion sequence.
[0008] As a preferred embodiment of the above-mentioned first aspect, the loss function adopted in the training process of the quantized variational autoencoder model is a weighted sum of reconstruction loss and quantization loss, the reconstruction loss is the L2 loss between the input human motion sequence and the reconstructed human motion sequence, and the quantization loss is the L2 loss between the latent vector sequence and the output sequence of the second codeword sequence after the stop gradient operation.
[0009] As a preferred embodiment of the first aspect, during the training process of the quantized variational autoencoder model, the codebook needs to be continuously updated through exponential moving average and codebook reset operations.
[0010] As a preferred embodiment of the above-mentioned first aspect, in the quantized variational autoencoder model, the encoder network is composed of a one-dimensional convolutional layer, a Relu activation layer, two combined modules consisting of a downsampling layer and a residual block, and a one-dimensional convolutional layer, which are cascaded in sequence; the decoder network is composed of a one-dimensional convolutional layer, two combined modules consisting of a residual block and an upsampling layer, a Relu activation layer, and a one-dimensional convolutional layer, which are cascaded in sequence.
[0011] As a preferred embodiment of the above-mentioned first aspect, when training the training framework, the tag prediction loss adopts the log-likelihood loss of the randomly masked tags in the first tag sequence and the second tag sequence, and the semantic alignment loss adopts the L1 loss of the text vector and the motion vector. The weighted sum of the two losses is used as the total loss and the gradient is calculated and then back-propagated, and only the parameters of the bidirectional masked Transformer model are optimized.
[0012] As a preferred embodiment of the above-mentioned first aspect, when the first marker sequence is randomly masked, the masked marker ratio γ each time is obtained by randomly obtaining a sampling value τ from a uniform distribution in the range of (0,1), and then calculated by a cosine function γ=cos(πτ / 2).
[0013] As a preferred embodiment of the above-mentioned first aspect, the specific method for generating a complete tag sequence through a multi-step iterative strategy is: input the text vector of the motion description text together with a completely masked tag sequence into the trained bidirectional masked Transformer model to obtain a predicted tag sequence with confidence, and then mask the part of the tags with the lowest confidence in the predicted tag sequence obtained in the previous iteration in a continuous iterative manner, and re-input the masked predicted tag sequence together with the text vector of the motion description text into the bidirectional masked Transformer model, and finally generate a complete tag sequence after iterating to a preset number of times.
[0014] As a preferred embodiment of the first aspect, when masking some of the labels with the lowest confidence in the predicted label sequence, the proportion of the masked labels is obtained by performing a cosine function transformation on the ratio of the current number of iterations to the total number of iterations.
[0015] In a second aspect, the present invention provides a text-aligned human motion generation system, comprising:
[0016] The first-stage training module is used to train a quantized variational autoencoder model for reconstructing human motion. The quantized variational autoencoder model includes an encoder network and a decoder network. The input human motion sequence is converted into a potential vector sequence through the encoder network, and then quantized using a codebook to obtain a first codeword sequence and a first tag sequence composed of codeword index tags. The first codeword sequence is input into the decoder network to reconstruct the human motion sequence;
[0017] The second-stage training module is used to freeze the parameters of the decoder network and the pre-trained text-to-motion cross-modal retrieval model, and then form a training framework with the bidirectional masked Transformer model for training; during the training process, the input text is encoded into a text vector by the text encoder, and then input into the bidirectional masked Transformer model together with the first tag sequence after random masking, and the second tag sequence is predicted and the tag prediction loss is calculated, and then the second tag sequence is converted into a second codeword sequence by using the codebook, and then input into the decoder network to reconstruct the human motion sequence, and then the motion encoder is used to obtain the motion vector, and the semantic alignment loss of the text vector and the motion vector is calculated, and finally the two losses are weighted summed and back-propagated to optimize the model parameters;
[0018] The inference generation module is used to encode the motion description text input by the user into a text vector through the text encoder and input it into the trained bidirectional masked Transformer model together with a fully masked tag sequence, generate a complete tag sequence through a multi-step iterative strategy, convert the complete tag sequence into a codeword sequence using a codebook, and then generate a human motion sequence by the decoder network.
[0019] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for generating human motion for text alignment as described in any one of the first aspects above is implemented.
[0020] In a fourth aspect, the present invention provides a computer electronic device comprising a memory and a processor;
[0021] The memory is used to store computer programs;
[0022] The processor is used to implement the human body motion generation method for text alignment as described in any one of the first aspects above when executing the computer program.
[0023] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0024] The present invention designs a text-aligned human motion generation method, which improves the semantic alignment of the generated motion and text input by improving the text encoder and adding semantic alignment loss, strengthens motion perception, and fully perceives the global context through a bidirectional masked Transformer, thereby achieving the effect that the generated human motion is more consistent with the original data distribution and more consistent with the text description. Verification results show that when faced with long and complex text input, the present invention can generate natural and text-description-compliant results. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Schematic diagram of the steps of the human motion generation method for text alignment;
[0026] Figure 2 This is a schematic diagram of the structure of the quantized variational autoencoder model;
[0027] Figure 3 It is the overall flow chart of two training stages of the present invention;
[0028] Figure 4 This is a flow chart of the reasoning phase of the present invention.
[0029] Figure 5 for the generated effect in the first instance;
[0030] Figure 6 for the resulting effect in the second instance. DETAILED DESCRIPTION
[0031] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.
[0032] In the description of the present invention, it is to be understood that when an element is considered to be "connected" to another element, it may be directly connected to the other element or indirectly connected, that is, there are intermediate elements. On the contrary, when an element is said to be "directly" connected to another element, there are no intermediate elements.
[0033] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features.
[0034] The present invention provides a text-aligned human motion generation method, which enhances the overall motion perception capability and the semantic alignment of generated motion through an improved text encoder and an added semantic alignment loss, as well as a bidirectional masked Transformer architecture, so that the generated human motion sequence is more consistent with the original data distribution and has more semantic consistency with the input motion description text, thereby improving the motion generation quality.
[0035] It should be noted that the motion description text mentioned in the present invention refers to a description of the motion sequence to be generated, and the human motion sequence mentioned in the present invention refers to a three-dimensional human posture sequence composed of a series of three-dimensional human posture frames in the process of human motion.
[0036] The specific implementation of the above-mentioned text-aligned human motion generation method of the present invention is described in detail below. The method includes two training stages and a reasoning stage when performing a specific three-dimensional human motion generation task after training.
[0037] like Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned text-aligned human motion generation method includes steps S1 to S3, and the specific implementation of each step is described in detail below.
[0038] S1. Train a vector quantized variational autoencoder (VQ-VAE) model for reconstructing human motion. The vector quantized variational autoencoder model includes an encoder network and a decoder network. The input human motion sequence is converted into a latent vector sequence through the encoder network, and then quantized (vector quantization) using a codebook to obtain a first codeword sequence and a first tag sequence composed of codeword index tags (code indices). The first codeword sequence is input into the decoder network to reconstruct the human motion sequence.
[0039] It should be noted that the VQ-VAE model belongs to the existing technology. It is a generative model that combines vector quantization (Vector Quantization) and variational autoencoder (VAE). Its basic network structure includes an encoder network and a decoder network. The core idea is to introduce a codebook-based quantization operation between the encoder network and the decoder network, and learn the structured representation of data through a discrete latent space (Latent Space).
[0040] It should be noted that the encoder network and decoder network used in the VQ-VAE model can be implemented by a combination of convolutional layers, residual blocks, etc., and can be designed according to actual task requirements. In an embodiment of the present invention, the encoder network is composed of a one-dimensional convolutional layer, a Relu activation layer, two combination modules consisting of a downsampling layer and a residual block, and a one-dimensional convolutional layer, and the decoder network is composed of a one-dimensional convolutional layer, two combination modules consisting of a residual block and an upsampling layer, a Relu activation layer, and a one-dimensional convolutional layer, and the overall VQ-VAE model structure is as follows: Figure 2 shown.
[0041] In an embodiment of the present invention, training a quantized variational autoencoder (VQ-VAE) model for human motion reconstruction specifically includes the following sub-steps:
[0042] S11. The encoder encodes a given input human motion sequence into a latent vector sequence.
[0043] Specifically, for a given input human motion sequence , is the length of the human motion sequence, that is, the total number of 3D human pose frames contained in it. For any i-th 3D human pose frame , Represents the dimension of a single frame of 3D human pose. First, use the encoder Convert it into a latent vector sequence , where n / N represents the downsampling rate and d represents the dimension of the latent vector.
[0044] S12, quantize the latent vector and update the codebook.
[0045] For the latent vector sequence Each latent vector in , from the initialized codebook Select and The closest codeword vector is used as the quantized vector , and record the codeword index mark corresponding to the closest codeword vector in the codebook , thereby minimizing the error, where , K represents the size of the codebook, and this process is called quantization , which can be expressed as:
[0046]
[0047] The entire latent vector sequence After quantization, each latent vector We get a quantized vector , the quantized vector Actually corresponds to the codeword index mark in the codebook The codeword under, so all quantized vectors The first codeword sequence can be formed . Record all at the same time The corresponding codeword index tag , which can constitute the first tag sequence .
[0048] VQ-VAE also needs to update the codebook during training. Since the codebook crash problem is often encountered during VQ-VAE training, this paper uses two methods, Exponential Moving Average (EMA) and codebook reset, to improve this problem. At the beginning of training, the codebook is initialized by directly using the result obtained by inputting the motion data into the encoder. Then, as the training process continues, the EMA method will be based on the formula Continuously update the codebook, where represents the corresponding codebook at the tth iteration, is the exponential shift parameter. If some codewords are found to be unused for a long time and inactive during training, the code reset technology is responsible for reallocating them. The above two technologies are used to ensure the successful update and iteration of the codebook during training.
[0049] S13, the decoder generates a reconstructed human motion sequence.
[0050] Finally, the first codeword sequence obtained by quantization Sent to the decoder , and thus is re-invested into the motion space to achieve motion reconstruction, and the reconstructed human motion sequence is obtained .
[0051] S14. Calculate the reconstruction loss and quantization loss, and optimize the model parameters through back propagation.
[0052] The training loss function of the motion reconstruction part of VQ-VAE mainly includes two parts: reconstruction loss and quantization loss. The total loss function is the weighted sum of reconstruction loss and quantization loss. The reconstruction loss is the input human motion sequence and reconstruction of human motion sequences The L2 loss between , and the quantization loss is the potential vector sequence With the second codeword sequence L2 loss between output sequences after stopping the gradient operation:
[0053]
[0054] The first part of the above loss represents the reconstruction loss, and the second part represents the quantization loss, where Represents the stop-gradient operator, is a hyperparameter that controls the quantization loss constraint. Represents the L2 loss calculation function.
[0055] Based on the pre-constructed training data set, the two loss functions are combined to calculate The gradient of is back-propagated to optimize the VQ-VAE model parameters to minimize By continuously iterating for the goal, the VQ-VAE training can be completed, and the training phase of the S1 step can be called the first phase.
[0056] S2. After freezing the parameters of the decoder network and the pre-trained Text-to-Motion Retrieval (TMR) model, they are trained together with the bidirectional masked Transformer model in a training framework. During the training process, the input text is encoded into a text vector by the text encoder, and then input into the bidirectional masked Transformer model together with the first tag sequence after random masking, and the second tag sequence is predicted and the tag prediction loss is calculated. The second tag sequence is then converted into a second codeword sequence using a codebook, and then input into the decoder network to reconstruct the human motion sequence, and the motion encoder is further used to obtain the motion vector, and the semantic alignment loss of the text vector and the motion vector is calculated. Finally, the two losses are weighted summed and back-propagated to optimize the model parameters.
[0057] It should be noted that the above-mentioned pre-trained text-to-motion cross-modal retrieval (TMR) model needs to be pre-trained on a text-human motion dataset. The TMR model belongs to the prior art and is used to achieve accurate matching of text descriptions with three-dimensional human motion sequences through deep learning technology. The TMR model includes a text encoder and a motion encoder, and both the text encoder and the motion encoder adopt the Transformer model architecture. The specific model structure and pre-training method of the TMR model belong to the prior art. The present invention can directly call the pre-trained TMR model and introduce the text encoder and motion encoder therein into the training framework of the present invention.
[0058] It should be noted that the model structure adopted by the bidirectional masked Transformer model of the present invention is directly the encoder module Encoder of the Transformer model, and the bidirectional mask refers to the introduction of a mask mechanism to the input tag sequence during training and subsequent reasoning. When the training framework is trained, the parameters of the decoder network and the TMR model are frozen, and only the parameters of the bidirectional masked Transformer model are optimizable. The total loss function of the training is the weighted sum of the tag prediction loss and the semantic alignment loss. The specific forms of the tag prediction loss and the semantic alignment loss need to be designed according to their respective loss characteristics. In an embodiment of the present invention, the above-mentioned tag prediction loss can adopt the log-likelihood loss of the randomly masked tags in the first tag sequence and the second tag sequence, and the above-mentioned semantic alignment loss can adopt the L1 loss of the text vector and the motion vector. Each round of training can use the weighted sum of the two losses as the total loss and calculate the gradient and then back propagate, but only the bidirectional masked Transformer model is optimized.
[0059] In order to ensure the training effect, when the first label sequence is randomly masked, the masked label ratio γ is obtained by randomly obtaining a sampling value τ from a uniform distribution in the range of (0,1), and then calculated by the cosine function γ = cos(πτ / 2).
[0060] In an embodiment of the present invention, the training of the bidirectional masked Transformer model responsible for modeling the input text and human motion markers in the above step S2 can be specifically implemented according to the following sub-steps.
[0061] S21. Randomly mask the first tag sequence quantized by the VQ-VAE encoder, and send it and the input text vector encoded by the text encoder of the pre-trained TMR model into the bidirectional masked Transformer model to predict the complete tag sequence.
[0062] Specifically, the input of each sample in the training dataset contains a human motion sequence and a motion description text. Using the VQ-VAE model trained in the first stage, we can Transformed into the corresponding first codeword sequence in the codebook At the same time, it is mapped to the first tag sequence according to its position information in the codebook , and then the first marker sequence Use a certain ratio to perform random masking to obtain The goal is to successfully predict the corresponding mark of the real data behind the mask mark based on the input text t and the unmasked mark. The mask ratio here is Follow the following cosine function calculation formula:
[0063]
[0064] in , represents the degree of damage to the marker sequence. , it means that the sequence has been completely destroyed.
[0065] During the training process, each training needs to start from a uniform distribution. Random sampling , then calculate the corresponding number of mask marks by the following formula:
[0066]
[0067] Where: Represents the length of the first tag sequence.
[0068] Thus, the number of marks that need to be masked in the compost can be calculated. For the first tag sequence Perform random masking to obtain a masked tag sequence. At the same time, the motion description text of each sample input in the training data set is encoded through the text encoder of the TMR model to obtain a text vector. The masked tag sequence and the text vector instrument are input into the bidirectional masked Transformer model to predict the masked part of the tag and obtain the second tag sequence .
[0069] S22. Calculate the loss function by comparing the predicted mark of the masked part with the real mark .
[0070] Considering that the codebook to which the motion marker belongs is discrete and finite in size, this can be simply understood as a classification problem. The mathematical optimization goal is to minimize the log-likelihood of the predicted target. Therefore, the calculation formula for the marker prediction loss is as follows:
[0071]
[0072] Where: t represents the input motion description text.
[0073] S23. The predicted index mark is indexed by the codebook in VQ-VAE to obtain a second codeword sequence.
[0074] Specifically, the index tag sequence predicted by Transformer, that is, the second tag sequence, will be used to index the codebook obtained in the first stage VQ-VAE training part, and the codeword vector corresponding to each codeword index tag in the second tag sequence is found from the codebook to obtain the corresponding codeword vector sequence, which is recorded as the second codeword sequence.
[0075] S24. Send the obtained second codeword sequence to the decoder network of the VQ-VAE model to generate a complete motion sequence, and calculate the semantic loss function at the same time.
[0076] Specifically, the codeword vector sequence obtained in the previous step is fed into the VQ-VAE decoder to obtain the complete motion sequence , and further input the motion encoder of the pre-trained TMR model to obtain the motion vector Finally, the motion vector The motion description text originally entered with the model The text vector obtained by TMR text encoder The loss function is calculated to further optimize the Transformer model parameters. It is worth noting that in this process, the trained VQ-VAE model and TMR text and motion encoder are in a frozen state to avoid destroying the results of the first stage of training. Here, the L1 loss function is used for the newly added semantic alignment loss, and the calculation formula is:
[0077]
[0078] The degree of semantic alignment between text and motion is measured by calculating the L1 distance between two vectors. The smaller the value, the more consistent the semantics of the two are.
[0079] S25. Back-propagation optimizes the parameters of the bidirectional masked Transformer model.
[0080] In summary, the loss function is composed of the label prediction loss and semantic alignment loss It consists of two parts:
[0081]
[0082] in, It is a hyperparameter that controls the semantic alignment loss. The larger the value, the more emphasis is placed on the semantic alignment of text and generated motion during training. Gradients are calculated and back-propagated to optimize the Transformer model parameters while keeping the rest of the model parameters fixed. To minimize the loss function With this as the goal, the training of the bidirectional masked Transformer model can be completed through continuous iterations. This training process can be recorded as the second stage.
[0083] Therefore, the above S1 and S2 constitute a two-stage model training mode, such as Figure 3As shown in the figure. In the first stage, VQ-VAE is trained to achieve human motion reconstruction. The goal of this stage is to learn to convert the complete input human motion sequence into discrete token representation in an encoded manner. In the second stage, a bidirectional masked Transformer model is trained to achieve human motion generation. The training goal of this stage is to model text vectors and motion tags, so that a more semantically consistent human motion tag sequence can be generated based on the input text in the inference stage, and then sent to the VQ-VAE motion decoder in the first stage to generate a complete human motion sequence.
[0084] In the above two-stage training, the present invention mainly proposes two improvements: first, the TMR model (Text-to-Motion Retrieval, TMR) pre-trained on a special text-human motion dataset is used as the text encoder and motion encoder to ensure the alignment of the input text and the generated human motion in the embedding space, enhance the model's ability to understand human motion at the sequence level, and add a new semantic alignment loss to constrain the generated human motion and input text to maintain consistency; second, a masked training strategy similar to BERT is adopted, and a bidirectional masked Transformer architecture is adopted, that is, a certain proportion of tokens are randomly masked for training and prediction. The model can fully perceive the global context content during training or inference, so that the predicted human motion is more consistent with the original data distribution, effectively avoiding the problem of error accumulation.
[0085] Therefore, after completing the two-stage training, the network parameters of the text encoder, bidirectional masked Transformer model, and decoder network can be fixed for the subsequent reasoning generation task in the S3 step.
[0086] S3. The motion description text input by the user is first encoded into a text vector by the text encoder and input into the trained bidirectional masked Transformer model together with a fully masked tag sequence. A complete tag sequence is generated through a multi-step iterative strategy, and the complete tag sequence is converted into a codeword sequence using a codebook, and then the decoder network generates a human motion sequence.
[0087] It should be noted that the present invention generates a complete tag sequence through a multi-step iterative strategy, which means that after a tag sequence is generated through a prediction, some tags with lower confidence need to be masked again, and then re-input for further prediction, and iterate continuously until an accurate tag sequence is obtained. It should be noted that the bidirectional masked Transformer model predicts the codeword index tag of each mask position in the predicted tag sequence, and the direct output is actually a probability distribution, corresponding to the probability value of each codeword in the codebook, and the codeword corresponding to the maximum probability value is recorded as the predicted tag of this mask position, and this maximum probability value is the confidence.
[0088] In an embodiment of the present invention, the specific method of generating a complete tag sequence through a multi-step iterative strategy is as follows: the text vector of the motion description text and a fully masked tag sequence are input into the trained bidirectional masked Transformer model to obtain a predicted tag sequence with confidence. Then, in a continuous iterative manner, the part of the tags with the lowest confidence in the predicted tag sequence obtained in the previous iteration is masked, and the masked predicted tag sequence is re-input into the bidirectional masked Transformer model together with the text vector of the motion description text, and the complete tag sequence is finally generated after iterating to a preset number of times.
[0089] It should be noted that when masking some of the lowest confidence marks in the predicted mark sequence, the masked mark ratio also needs to be reasonably optimized. In an embodiment of the present invention, at each iteration, the masked mark ratio is obtained by performing a cosine function conversion on the ratio of the current iteration number to the total iteration number.
[0090] In an embodiment of the present invention, in the above step S3, the process of generating a complete human body motion sequence based on input text reasoning can be specifically implemented through the following sub-steps.
[0091] S31. The input text is passed through the TMR text encoder to obtain a text vector, which is combined with the fully masked motion marker sequence and sent to the bidirectional masked Transformer model.
[0092] Specifically, assume that the bidirectional masked Transformer model needs to iterate L steps to generate a complete tag sequence of length n. At the beginning of reasoning, the motion description text t initially input by the user is obtained through the pre-trained TMR text encoder to obtain the text vector , which is the same as the initial fully masked motion marker sequence The long sequence is sent to the bidirectional masked Transformer model for prediction.
[0093] S32, multiple-step iterations to generate a complete motion sequence.
[0094] For the predicted tag sequence obtained in the previous step, it is re-masked at a certain ratio and then enters the next iteration. Specifically, given the The mask tag sequence generated by the iteration When , the predicted tag sequence will output the probability distribution of the masked position tag, and then generate the corresponding tag based on its sampling, and then among these newly predicted tags, the one with the lowest corresponding probability The markers will be re-masked, and the remaining generated markers will remain unchanged in the next iteration. The final generated marker sequence will be fed into the next iteration. The number of tokens that need to be re-masked here Depend on
[0095]
[0096] Get, where Represents the current iteration number, represents the total number of iterations required to generate, is the cosine function calculation formula of the mask ratio introduced above. After iterations, a complete motion marker sequence is finally generated.
[0097] S33. Index the motion marker sequence to obtain a codeword vector sequence.
[0098] Similar to the training process, the inference stage also uses the codebook obtained by VQ-VAE training for indexing. The complete tag sequence generated by the iteration uses the codebook to convert the complete tag sequence into a codeword sequence, that is, each codeword index tag in the complete tag sequence selects the corresponding codeword vector from the codebook according to the index, thereby converting the index tag sequence into a codeword sequence.
[0099] S34. Generate complete human motion through VQ-VAE decoder.
[0100] Finally, the converted codeword sequence is used by the decoder network with frozen parameters after the first stage VQ-VAE training to generate a complete human motion sequence that is semantically consistent with the input motion description text. Therefore, the human motion sequence generation process in the entire inference stage is as follows: Figure 4 shown.
[0101] In an embodiment of the present invention, in order to verify the specific generation effect of the text-aligned human motion generation method described in S1 to S3 above, two examples of human motion sequence generation based on a two-stage trained model are exemplified.
[0102] like Figure 5As shown in the figure, the input description text is "a person raises his right arm and pulls it back, prepares to throw the ball, and then catches the ball", and a human motion sequence of 196 frames in length is generated. The key 7 frames of human posture frames are selected here to dynamically display the corresponding semantic meaning in the above description text.
[0103] like Figure 6 As shown in the figure, the input description text is "a man sitting on something at waist height, with his arms hanging down to his thighs, then stretched to shoulder height, flat on both sides of his body", and a human motion sequence of 196 frames in length is generated. The key 7 frames of human posture frames are selected here to dynamically display the corresponding semantic meaning in the above description text.
[0104] It can be seen that the example samples presented above show that even when faced with long and complex text input, the present invention can generate results that are natural and consistent with the text description.
[0105] The above-described embodiments are only some preferred implementations of the present invention, but are not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.
Claims
1. A method for generating human motion for text alignment, characterized in that: include: S1. Training a quantized variational autoencoder model for reconstructing human motion, wherein the quantized variational autoencoder model comprises an encoder network and a decoder network. An input human motion sequence is converted into a potential vector sequence through the encoder network, and then quantized using a codebook to obtain a first codeword sequence and a first tag sequence composed of codeword index tags. The first codeword sequence is input into the decoder network to reconstruct the human motion sequence. S2, freezing the parameters of the decoder network, the pre-trained text-to-motion cross-modal retrieval model including the text encoder and the motion encoder, and then forming a training framework with the bidirectional masked Transformer model for training; During the training process, the input text is encoded into a text vector by the text encoder, and then input into the bidirectional masked Transformer model together with the first tag sequence after random masking, and the second tag sequence is predicted and the tag prediction loss is calculated. Then, the second tag sequence is converted into a second codeword sequence by using a codebook, and then input into the decoder network to reconstruct the human motion sequence, and then the motion vector is further obtained by the motion encoder, and the semantic alignment loss of the text vector and the motion vector is calculated. Finally, the two losses are weighted and summed and the model parameters are back-propagated to optimize. S3. The motion description text input by the user is first encoded into a text vector by the text encoder and input into the trained bidirectional masked Transformer model together with a fully masked tag sequence. A complete tag sequence is generated through a multi-step iterative strategy, and the complete tag sequence is converted into a codeword sequence using a codebook, and then the decoder network generates a human motion sequence.
2. The method for generating human motion for text alignment according to claim 1, characterized in that: The loss function used in the training process of the quantized variational autoencoder model is the weighted sum of reconstruction loss and quantization loss, the reconstruction loss is the L2 loss between the input human motion sequence and the reconstructed human motion sequence, and the quantization loss is the L2 loss between the latent vector sequence and the output sequence of the second codeword sequence after the stop gradient operation.
3. The method for generating human motion for text alignment according to claim 1, characterized in that: In the quantized variational autoencoder model, the encoder network is composed of a one-dimensional convolutional layer, a Relu activation layer, two combination modules consisting of a downsampling layer and a residual block, and a one-dimensional convolutional layer, which are cascaded in sequence; the decoder network is composed of a one-dimensional convolutional layer, two combination modules consisting of a residual block and an upsampling layer, a Relu activation layer, and a one-dimensional convolutional layer, which are cascaded in sequence.
4. The method for generating human motion for text alignment according to claim 1, characterized in that: When training the training framework, the tag prediction loss adopts the log-likelihood loss of the randomly masked tags in the first tag sequence and the second tag sequence, and the semantic alignment loss adopts the L1 loss of the text vector and the motion vector. The weighted sum of the two losses is used as the total loss and the gradient is calculated and then back-propagated, and only the parameters of the bidirectional masked Transformer model are optimized.
5. The method for generating human motion for text alignment according to claim 1, wherein: When the first marker sequence is randomly masked, the masked marker ratio γ each time is obtained by randomly obtaining a sampling value τ from a uniform distribution in the range of (0,1), and then calculated using the cosine function γ=cos(πτ / 2).
6. The method for generating human motion for text alignment according to claim 1, characterized in that: The specific method for generating a complete tag sequence through a multi-step iterative strategy is as follows: inputting the text vector of the motion description text together with a fully masked tag sequence into the trained bidirectional masked Transformer model to obtain a predicted tag sequence with confidence, and then masking the part of the tags with the lowest confidence in the predicted tag sequence obtained in the previous iteration in a continuous iterative manner, and re-inputting the masked predicted tag sequence together with the text vector of the motion description text into the bidirectional masked Transformer model, and finally generating a complete tag sequence after iterating to a preset number of times.
7. The method for generating human motion for text alignment according to claim 6, characterized in that: When masking some of the tags with the lowest confidence in the predicted tag sequence, the masked tag ratio is obtained by performing a cosine function transformation on the ratio of the current number of iterations to the total number of iterations.
8. A text-aligned human motion generation system, characterized in that: include: The first-stage training module is used to train a quantized variational autoencoder model for reconstructing human motion. The quantized variational autoencoder model includes an encoder network and a decoder network. The input human motion sequence is converted into a potential vector sequence through the encoder network, and then quantized using a codebook to obtain a first codeword sequence and a first tag sequence composed of codeword index tags. The first codeword sequence is input into the decoder network to reconstruct the human motion sequence; The second-stage training module is used to freeze the parameters of the decoder network and the pre-trained text-to-motion cross-modal retrieval model, and then form a training framework with the bidirectional masked Transformer model for training; During the training process, the input text is encoded into a text vector by the text encoder, and then input into the bidirectional masked Transformer model together with the first tag sequence after random masking, and the second tag sequence is predicted and the tag prediction loss is calculated. Then, the second tag sequence is converted into a second codeword sequence by using a codebook, and then input into the decoder network to reconstruct the human motion sequence, and then the motion vector is further obtained by the motion encoder, and the semantic alignment loss of the text vector and the motion vector is calculated. Finally, the two losses are weighted and summed and the model parameters are back-propagated to optimize. The inference generation module is used to encode the motion description text input by the user into a text vector through the text encoder and input it into the trained bidirectional masked Transformer model together with a fully masked tag sequence, generate a complete tag sequence through a multi-step iterative strategy, convert the complete tag sequence into a codeword sequence using a codebook, and then generate a human motion sequence by the decoder network.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating human motion for text alignment as described in any one of claims 1 to 7 is implemented.
10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the human motion generation method for text alignment as described in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Method for generating action sequence for driving virtual character to move according to text
CN116883555A
Motion sequence retrieval model training method and motion sequence retrieval method and device
CN118691784A
Cited By
Human motion generation method and system based on multi-token large language model
CN120597896A
A method and system for generating human motion based on a multi-token large language model
CN120597896B
Text and face collaborative restoration method based on cross-modal alignment
CN120655787A
Action migration method and device, electronic equipment, computer readable storage medium and program product
CN121354212A
Text-driven three-dimensional human body action generation method based on local generation and global fusion
CN121725115A