A method for generating a motion sequence for driving a virtual character motion according to a text
Patent Information
- Application Number
- CN202310840866.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-07-10
AI Technical Summary
[0008]首先,对于第一类技术,由于文本和人体运动数据在数据分布、数据属性、数据表示等各方面均存在较大差异,所以现有的技术很难学习到好的文本-运动的联合表示,导致该类技术生成的运动质量不佳,同时多样性不足
[0011]因此,本发明的目的在于克服上述现有技术的缺陷,提供一种根据文本生成驱动虚拟角色运动的动作序列的方法。
Smart Images

Figure CN116883555B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to the field of neural networks, and more specifically, to a method for generating action sequences that drive the movement of virtual characters based on text. Background Technology
[0002] Text-driven human motion generation technology aims to automatically generate realistic and natural 3D human motion sequences (virtual characters) based on a given text description. It has wide applications in fields such as intelligent animation, virtual reality, game development, and human-computer interaction. Existing text-driven human motion generation technologies are mainly divided into two categories:
[0003] (1) The first type of technique encodes both text and motion simultaneously, mapping them to the same latent space, aiming to learn a joint representation of text and motion. When generating action sequences from text, the text is directly encoded, and the text features are input into the motion decoder to generate the action sequences;
[0004] (2) The second type of technology uses a motion generation model as a priori, taking text as a constraint for motion generation, and constructing a conditional generation model to perform text-driven motion generation. Within this second type of technology, there are two technical approaches:
[0005] Technical approach (i) involves directly generating conditions within the posture space of human movement;
[0006] Technical approach (ii) involves first mapping motion to a low-dimensional latent space, and then performing conditional generation at the latent space level.
[0007] Both of these categories of technologies have their own drawbacks.
[0008] First, for the first type of technology, because text and human motion data differ greatly in terms of data distribution, data attributes, and data representation, existing technologies have difficulty learning good joint representations of text and motion, resulting in poor motion quality and insufficient diversity generated by this type of technology.
[0009] Secondly, for the second type of technology, since the posture space of human movement has extremely high dimensions and complex distribution, the technical route (i) is difficult to guarantee the consistency between the generated motion and the given text.
[0010] Finally, because the two main categories of techniques mentioned above consider only a single level of the relationship between text and motion, the degree of consistency between the motion generated by these techniques and the text description needs improvement. Therefore, existing techniques need to be improved. Summary of the Invention
[0011] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a method for generating motion sequences that drive the movement of virtual characters based on text.
[0012] The objective of this invention is achieved through the following technical solution:
[0013] According to a first aspect of the present invention, a method for training a motion generation model is provided, wherein the motion generation model generates action sequences that drive the movement of a virtual character based on text. The method includes: training a VQ-VAE model using multiple sample action sequences to reconstruct the sample action sequences, obtaining a trained VQ-VAE model including an encoder, a quantizer, a decoder, and a codebook; converting each sample action sequence of the multiple sample action sequences into a corresponding sample action index sequence according to the encoder, quantizer, and codebook of the trained VQ-VAE model; constructing multiple action pairs, each action pair including a sample action index sequence and sample text describing the action state indicated by the corresponding sample action sequence; training a text-to-motion index model based on a Transformer network to generate action index sequences based on text by using an attention mechanism to fuse local cross-attention and global conditional attention between sample text and sample action index sequences, obtaining a trained text-to-motion index model; and obtaining a trained motion generation model, which includes the trained VQ-VAE model and the trained text-to-motion index model.
[0014] Optionally, the text-to-motion index model includes: a first Transformer network for calculating cross-attention between each element of the first action encoding sequence and the word features of each word in the text, to extract a second action encoding sequence that integrates local information, wherein the first action encoding sequence is concatenated based on the embedding vectors obtained by looking up each index in the codebook according to the action index sequence; a second Transformer network for calculating conditional attention based on the local features and the text features corresponding to the whole text, to extract a third action encoding sequence that integrates local and global information; and a third Transformer network for generating a probability distribution sequence of action indices based on the third action encoding sequence, wherein each element of the probability distribution sequence contains the probability that the element is an index in the codebook.
[0015] Optionally, the trained text-to-motion index model is trained as follows: each action pair is input as training data into the text-to-motion index model, the probability distribution sequence of the action index is output, the cross-entropy loss between the action index sequence in the action pair and the probability distribution sequence of the output action index is determined, the gradient is calculated based on the cross-entropy loss, and the trainable parameters of the third Transformer network, the second Transformer network, and the first Transformer network are updated by backpropagation.
[0016] Optionally, the first, second, and third Transformer networks each contain multiple layers of Transformer subnetworks, with each Transformer subnetwork implemented based on a multi-head attention mechanism.
[0017] Optionally, the VQ-VAE model is trained as follows: multiple sample action sequences are acquired, where each frame of the sample action sequence contains control information of the virtual character's joints; the VQ-VAE model is trained using the loss function of the VQ-VAE model to reconstruct the multiple sample action sequences, resulting in a trained VQ-VAE model. During training, the encoder generates action feature sequences based on the sample action sequences, the quantizer obtains the quantized action feature sequences corresponding to the action feature sequences based on the codebook, and the decoder reconstructs the sample action sequences based on the quantized action feature sequences.
[0018] Optionally, the trained VQ-VAE model is trained as follows: Multiple sample action sequences are acquired, wherein each frame of the sample action sequence contains control information of the virtual character's joints; a motion relationship mask is acquired for the virtual character, the motion relationship mask being a matrix indicating the motion correlation between the character's control points, wherein the virtual character is divided into multiple limb segments, the motion correlation of control points located in the same limb segment is set to 0, and the motion correlation of control points located in different limb segments is set to negative infinity; a VQ-VAE model is acquired, whose encoder uses a fourth Transformer network, wherein the fourth Transformer network... The ORMER network is configured to: when calculating the single-head self-attention of the control information of each control point in each frame of the input action sequence, superimpose the QKT value in the attention mechanism with the motion relationship mask to shield attention interference between irrelevant limb segments; train the VQ-VAE model using the loss function of the VQ-VAE model to reconstruct multiple sample action sequences to obtain the trained VQ-VAE model, wherein, during training, the encoder generates an action feature sequence based on the sample action sequence, the quantizer obtains the quantized action feature sequence corresponding to the action feature sequence based on the codebook, and the decoder reconstructs the sample action sequence based on the quantized action feature sequence.
[0019] According to a second aspect of the present invention, a system for training a motion generation model based on the method of the first aspect is provided, comprising: a first training module configured to: reconstruct a VQ-VAE model using multiple sample action sequences to obtain a trained VQ-VAE model including an encoder, a quantizer, a decoder, and a codebook; an action pair construction module configured to: convert each sample action sequence of the multiple sample action sequences into a corresponding sample action index sequence according to the encoder, quantizer, and codebook of the trained VQ-VAE model; and construct multiple action pairs, each action pair including one of the sample actions. The system comprises: an index sequence and sample text describing the action states indicated by the corresponding sample action sequences; a second training module configured to: train a text-to-motion index model based on a Transformer network according to the multiple action pairs and the codebook, by using an attention mechanism to fuse local cross-attention and global conditional attention between the sample text and the sample action index sequence, and learn to generate action index sequences from the text, thus obtaining a trained text-to-motion index model; and a construction module configured to: obtain a trained motion generation model, which includes a trained VQ-VAE model and a trained text-to-motion index model.
[0020] According to a third aspect of the present invention, a method for generating an action sequence that drives the movement of a virtual character based on text is provided, comprising: acquiring action description text for describing the movement state of the virtual character; acquiring a trained motion generation model obtained according to the method of the first aspect or the system of the second aspect, which includes a trained VQ-VAE model and a trained text-to-motion index model; inputting the action description text into the trained text-to-motion index model to generate an action index sequence corresponding to the action description text; converting the action index sequence corresponding to the action description text into a quantized action feature sequence and inputting it into the decoder of the trained VQ-VAE model to generate an action sequence corresponding to the action description text, wherein the action sequence corresponding to the action description text can be used to drive the virtual character to move in order to generate animation.
[0021] According to a fourth aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions; wherein the one or more processors are configured to implement the steps of the methods described in the first and / or third aspects by executing the executable instructions. Attached Figure Description
[0022] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0023] Figure 1This is a flowchart illustrating a method for training a motion generation model according to an embodiment of the present invention.
[0024] Figure 2 This is a schematic diagram illustrating the principle of a motion relationship mask according to an embodiment of the present invention;
[0025] Figure 3 This is a schematic diagram illustrating the use of motion relationship masks to shield attentional interference between unrelated limb segments according to an embodiment of the present invention.
[0026] Figure 4 This is a schematic diagram illustrating the overall process of a method for training a motion generation model according to an embodiment of the present invention;
[0027] Figure 5 This is a schematic diagram of the processing flow inside the first Transformer network and the second Transformer network according to an embodiment of the present invention;
[0028] Figure 6 This diagram illustrates a comparison of the effects of the method according to an embodiment of the present invention and two prior art methods in generating action sequences. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.
[0030] As mentioned in the background section, existing technologies consider only a single level of the relationship between text and motion, resulting in a need to improve the consistency between the generated motion and the text description. In their research on text-driven motion generation, the inventors argued that this problem arises because there are local semantic correspondences and global sequential correspondences between character motion and text. Existing technologies have not fully explored this correspondence, leading to generated motion that does not accurately match the text description. Through research on text-driven motion generation tasks, the inventors determined that a text-to-motion indexing model based on a Transformer network can be trained to learn how to generate motion index sequences from text. This involves understanding both the local correspondences between elements of the sample action encoding sequence and words in the text, and the global correspondences between the sample action encoding sequence and the sample action sequence. Essentially, this involves representing local semantic correspondences by calculating the cross-attention between text word features and motion sub-segments, and representing global sequential correspondences by calculating the conditional self-attention between the overall text features and the entire motion segment, thereby improving the accuracy of the generated motion sequence's correspondence with the text.
[0031] Before describing the embodiments of the present invention in detail, some of the terms used therein are explained as follows:
[0032] The VQ-VAE (Vector-Quantised Variational AutoEncoder) model consists of an encoder, a quantizer, a decoder, and a codebook.
[0033] The encoder of the VQ-VAE model encodes the input of the VQ-VAE model and outputs an embedding vector.
[0034] The VQ-VAE model maintains a codebook (called EmbeddingSpace), which contains multiple indices and the corresponding embedding vector for each index. The number of indices in the codebook can be set by the implementer (e.g., 512, 1024, etc.), and the embedding vector for each index is generated during the training of the VQ-VAE model.
[0035] The quantizer of the VQ-VAE model is configured to calculate the distance between the encoder's output (embedded vector, or feature) and each embedded vector in the codebook, and then take the nearest embedded vector from the codebook to form a new embedded vector and pass it to the decoder.
[0036] The decoder of the VQ-VAE model reconstructs the input of the VQ-VAE model (that is, the input of the encoder of the VQ-VAE model) based on the output of the quantizer.
[0037] In their research on text-driven motion generation, the inventors discovered that, in addition to the problems mentioned in the background technology, the existing second-type technology (ii) does not consider the motion correlation between the character's limb structures when mapping motion to a low-dimensional latent space, resulting in unrealistic and unnatural generated motion. In summary, the inventors found that these shortcomings in the existing technology are caused by two aspects: (1) the existing technology does not fully consider the special spatial structure of the character when modeling the motion itself, resulting in unrealistic and unnatural generated motion; (2) there is a local semantic correspondence and a global sequential correspondence between the character's motion and the text, and the existing technology does not fully explore this correspondence, resulting in the generated motion not accurately conforming to the text description. Through research on text-driven motion generation tasks, the inventors found that defect (1) can be solved by introducing spatiotemporal feature extraction based on limb segment attention. Specifically, the character is divided into different parts according to limb segments, and spatial features are extracted from them using a spatial transformer. Then, all spatial features are input together into a temporal convolutional layer to extract temporal features. Finally, a vector quantization variational autoencoder is used to map the spatiotemporal features to a discrete low-dimensional latent space. Defect (2) can be solved by introducing global-local attention. Local semantic correspondence is represented by calculating the cross attention of text word features and motion sub-segments, while global sequential correspondence is represented by calculating the conditional self-attention of text sentence features and the entire motion. This makes the motion state of the character represented by the generated action sequence more consistent with the motion state of the character described in the text.
[0038] This invention is an improvement based on the technical route (ii) of the second type of technology. To facilitate understanding of the technical solution of the optimal embodiment that solves the above two defects, the following embodiments will be mainly described with reference to the solution that solves the above two defects. However, it should be understood that, in practice, the technical solution corresponding to the problem mentioned in the background technology can still produce positive technical effects when implemented alone.
[0039] To facilitate understanding of the structure of the motion generation model in this embodiment of the invention, a general description is provided below:
[0040] The motion generation model aims to generate motion sequences that drive the movement of characters (such as human bodies, virtual cats, virtual dogs, or other forms of movable characters; for simplicity, the human body will be used as an example below, but the working principle is similar for other forms of characters) based on text. This motion generation model includes the VQ-VAE model and the text-to-motion indexing model, where:
[0041] During training, the input of the VQ-VAE model is a sequence of sample actions, and the output is a sequence of generated sample actions. It is a self-supervised training model, which uses the sample action sequences themselves to guide the training.
[0042] During training, the input to the text-to-motion index model is a sequence of sample actions (used to guide training) and the corresponding text, and the output is the sequence of action indexes corresponding to the text.
[0043] During testing or application, the input to the text-to-motion indexing model is only text (which can be new text that has not been seen during model training, such as action description text), and the output is the action index sequence corresponding to the text. Based on the output of the text-to-motion indexing model, the corresponding embedding vector is extracted from the codebook, and then passed through the decoder of the VQ-VAE model to output the action sequence corresponding to the action description text, thereby driving the character to move to generate animation.
[0044] As can be seen from the foregoing summary, motion generation models need to be trained before application; after training, the trained motion generation model has the ability to generate action sequences based on text. The following text will primarily illustrate this using training methods as an example.
[0045] According to one embodiment of the present invention, a method for training a motion generation model is provided, wherein the motion generation model generates a sequence of actions that drive the movement of a virtual character based on text. (See also...) Figure 1 The method includes the following steps: A1, A2, A3, A4, and A5.
[0046] Step A1: Use multiple sample action sequences to train a VQ-VAE model to reconstruct the sample action sequences, and obtain a trained VQ-VAE model including an encoder, quantizer, decoder and codebook.
[0047] According to one embodiment of the present invention, the action sequence includes multiple frames, each frame containing control information of control points of a virtual character. It should be understood that the number of control points contained in each frame of the action sequence defined by the implementer may differ depending on the character. Furthermore, the control information may also take different forms of expression, resulting in different implementation methods. For example, some implementation methods may use the rotation angle of the control points as control information, while others may use the coordinate position of the control points, or a combination of both, or other forms defined by the implementer. This embodiment of the present invention does not limit these possibilities.
[0048] Sample action sequences are action sequences, referring to the action sequences in the dataset used to train the model. The dataset can be an existing one or a custom dataset defined by the implementer. An existing dataset could be the open-source HumanML3D dataset, which contains 14,616 3D human action sequences and 44,970 text descriptions (one action sequence may correspond to multiple text descriptions). When training the VQ-VAE model, only a portion of the action sequences (e.g., 80% or 70%) are used as sample action sequences. Since the VQ-VAE model is an autoencoder, its training task is to encode and reconstruct the action sequences from the input sequences; therefore, the text descriptions in the dataset are not used when training the VQ-VAE model.
[0049] To address the aforementioned deficiency (1), the trained VQ-VAE model is trained in the following manner:
[0050] A11. Obtain multiple sample action sequences, wherein each frame of the sample action sequence contains control information of the control points (such as joints) of a virtual human body.
[0051] A12. Obtain the motion relationship mask set for the human body. The motion relationship mask is a matrix indicating the motion correlation between control points of the human body. The human body is divided into multiple limb segments. The motion correlation of control points located in the same limb segment is set to 0, and the motion correlation of control points located in different limb segments is set to negative infinity.
[0052] For example, the human body can be divided into five limb segments: trunk, left hand, right hand, left leg, and right leg. Each limb segment contains its own joints, such as... Figure 2 As shown in Figure a, for ease of correspondence, the original English names of each joint in SMPL are used for explanation. The torso joints include head, neck, spine3, spine2, and spine1; the left hand joints include L_hand, L_wrist, L_elbow, L_shoulder, L_collar, and spine3; the right hand joints include R_hand, R_wrist, R_elbow, R_shoulder, R_collar, and spine3; the right leg joints include R_foot, R_ankle, R_knee, and R_hip; and the left leg joints include L_foot, L_ankle, L_knee, and L_hip. It can be seen that the spine3 joint is a shared joint, simultaneously classified in the torso, left hand, and right hand joints. Based on this classification, this embodiment of the invention defines a motion relation mask (or a new human body adjacency relation mask) M = m. {i,j} ∈R n*n , where n represents the number of control points in a frame of the motion sequence. If joints i and j are within the same limb segment, m{i,j} = 0, otherwise -∞. In some motion relationships, in addition to control points, other points may be added. For this, the mask can also add corresponding dimensions. For example, taking the motion sequence of the SMPL human body model specified in HumanML3D as an example, R... (n+3)*(n+3) Here, n represents the number of joints in the human body, a total of 23 joints. The additional 3 in n+3 correspond to the control information of the root joint, left foot contact point, and right foot contact point indicated in the motion sequence. The root joint, as the center of the human body, can be configured by the implementer to have a motion correlation of 0 with all other joints; the left foot contact point has a motion correlation of 0 with the left leg joint and -∞ with other joints; the right leg contact point has a motion correlation of 0 with the right leg joint and -∞ with other joints. The generated mask can be generated as follows: Figure 2 As shown in b, it should be understood that... Figure 2 b is for simplification only. Additionally, for other character types, the key points in the motion sequence will differ, and the mask can be adjusted accordingly.
[0053] A13. Obtain the VQ-VAE model, whose encoder uses a Transformer network (referred to as the fourth Transformer network here for distinction). The fourth Transformer network is configured to: when calculating the single-head self-attention of the control information of each control point in each frame of the input action sequence, integrate the QK... T The value is superimposed on the motion relationship mask to shield attentional interference between irrelevant limb segments;
[0054] According to one embodiment of the present invention, when the encoder calculates multi-head (the number of heads for attention can be customized by the implementer, and the same applies hereinafter) self-attention, it calculates the single-head self-attention of the control information between control points in each frame of the input action sequence in the following manner:
[0055]
[0056] Where sofmax is the sofmax activation function, Q, K, and V are vectors obtained by applying three different linear mappings of the attention mechanism to vectors composed of control information of each control point in each frame of the action sequence, D is the dimension of the Q, K, and V vectors, and M is the motion relationship mask. A related calculation illustration can be found in [reference needed]. Figure 3 , the action sequence j root j1, j2, ..., j n c f Convert to Q, K, V vectors, where j root This represents the control information of the root joints, j1, j2, ..., j nc represents the control information for the 1st to nth joints. f Control information indicating foot contact (containing control information for left and right foot contact, usually...) These represent the contact between the right ankle, right toe, left ankle, and left toe and the ground, respectively (0 indicates no contact, 1 indicates contact). After the self-attention of each single head is calculated, they are concatenated to obtain the final multi-head self-attention. For illustration, assuming four heads are set here, the self-attention of each single head is concatenated to obtain the final multi-head self-attention Att = {Att0, Att1, Att2, Att3}. The multi-head self-attention is then processed by the encoder's feedforward neural network (corresponding to...). Figure 3 From FFN), we obtain the spatial features F. s Spatial features F s The encoder's layer normalization layer (or temporal convolutional network, corresponding to...) Figure 3 TCN enc ), to obtain the spatiotemporal features F st The spatiotemporal features of each frame are concatenated to obtain the action feature sequence. It should be understood that... The formula is known to those skilled in the art. The improvement of this embodiment of the invention lies in incorporating QK into the attention mechanism. T The value is superimposed on the motion relationship mask. Since the motion relationship mask sets the motion correlation of control points located in the same limb segment to 0 and the motion correlation of control points located in different limb segments to negative infinity, based on the properties of the softmax function, it will shield the attention interference between unrelated limb segments, thus making the VQ-VAE model reconstruct action sequences more accurate.
[0057] A14. The VQ-VAE model is trained using the loss function of the VQ-VAE model to reconstruct multiple sample action sequences, resulting in a trained VQ-VAE model. The encoder is trained to generate action feature sequences based on the sample action sequences, the quantizer obtains the quantized action feature sequences corresponding to the action feature sequences based on the codebook, and the decoder is trained to reconstruct the sample action sequences based on the quantized action feature sequences.
[0058] Of course, without considering the defect (1), the existing method can also be used to train the VQ-VAE model in principle. According to an embodiment of the present invention, the VQ-VAE model is trained in the following manner: multiple sample action sequences are acquired, wherein each frame of the sample action sequence contains control information of the joints of the virtual character; the VQ-VAE model is trained using the loss function of the VQ-VAE model to reconstruct the multiple sample action sequences, thereby obtaining the trained VQ-VAE model, wherein during training, the encoder generates an action feature sequence based on the sample action sequence, the quantizer obtains the quantized action feature sequence corresponding to the action feature sequence based on the codebook, and the decoder reconstructs the sample action sequence based on the quantized action feature sequence.
[0059] Overall, training the VQ-VAE model corresponds to Figure 4 a and Figure 4 Phase 1 of b. During training, the total loss is calculated based on the existing loss function of the VQ-VAE model. The gradient is then calculated based on the total loss, and backpropagation is used to update the trainable parameters in the encoder T4 and decoder D of the VQ-VAE model, as well as the embedding vectors in the codebook (corresponding to vectors in the discrete motion latent space). The quantizer H uses the nearest neighbor algorithm and has no trainable parameters, so it does not require training. For a more intuitive illustration, refer to... Figure 4 At the top, there is a text: "A person is kicking with his right foot," corresponding to... Figure 4 The action sequence in the upper left corner of image a (for ease of understanding, the action sequence is shown graphically).
[0060] Step A2: Based on the encoder, quantizer, and codebook of the trained VQ-VAE model, convert each sample action sequence of the plurality of sample action sequences into a corresponding sample action index sequence.
[0061] According to one embodiment of the present invention, after the VQ-VAE model is trained, a trained VQ-VAE model is obtained. For each sample action sequence, the encoder of the VQ-VAE model is input to obtain the action feature sequence corresponding to the sample action sequence. The action feature sequence includes multiple action features. The quantizer obtains the index of the nearest embedding vector in the codebook for each action feature in the action feature sequence according to the nearest neighbor algorithm, and concatenates them in order to obtain the sample action index sequence. The concatenation of the embedding vectors corresponding to each index in the sample action index sequence in the codebook is the quantized action feature sequence.
[0062] Step A3: Construct multiple action pairs, each action pair including a sample action index sequence and sample text describing the action state indicated by the corresponding sample action sequence.
[0063] According to one embodiment of the present invention, action pairs can be established based on the relationship between text in the dataset and the action sequence corresponding to that text. For example, in the dataset, the action state described by text W can be represented by a sample action sequence Z, and the sample action index sequence corresponding to sample action sequence Z and text W can form an action pair.
[0064] Step A4: Based on the multiple action pairs and the codebook, train the text-to-motion index model built on the Transformer network. By using the attention mechanism to fuse local cross-attention and global conditional attention between sample text and sample action index sequences, the model learns to generate action index sequences from the text, thus obtaining the trained text-to-motion index model.
[0065] Overall, training a text-to-motion indexing model corresponds to Figure 4 a and Figure 4 Phase 2 in b.
[0066] According to one embodiment of the present invention, see Figure 4 b. The text-to-motion indexing model includes: a first Transformer network T1, used to calculate the cross-attention between each element of the first action encoding sequence and the word features of each word in the text, to extract a second action encoding sequence that integrates local information. The first action encoding sequence is concatenated based on the embedding vectors obtained by looking up each index in the codebook according to the action index sequence. A second Transformer network T2, used to calculate conditional attention based on the local features and the text features corresponding to the overall text, to extract a third action encoding sequence that integrates local and global information. A third Transformer network T3, used to generate a probability distribution sequence of action indices based on the third action encoding sequence. Each element of the probability distribution sequence contains the probability that the element is an index in the codebook. It should be understood that during training, the action index sequence can be the sample action index sequence during training; while during testing and application, the action index sequence can be the action index sequence predicted by the third Transformer network T3. To avoid ambiguity, the general term "action index sequence" is used here when describing the technical method in the text-to-motion indexing model.
[0067] According to one embodiment of the present invention, the word features of each word or the text features of the text can be extracted using a pre-trained language model. For example, one of the pre-trained CLIP model, BERT model, and GPT model. The following explanation uses the pre-trained CLIP model as an example.
[0068] For the first Transformer network T1, refer to Figure 4 a and Figure 5a. Using a pre-trained CLIP model, extract word features word by word from a given text. wi Assuming there are N words, then the corresponding... Figure 5 e in a w1 e w2 e w3 ..., e wN Then, it is combined with the low-dimensional c of the first action coding sequence (or the low-dimensional motion coding sequence c, assuming there are m frames, then they correspond to...). Figure 5 c1, c2, c3, ..., c in a m Calculate cross-attention to obtain the second action-encoding sequence c, which incorporates local information. ω (or a low-dimensional motion coding sequence c that incorporates local information) ω The illustrative steps are as follows:
[0069] Step A41: Given a text Ω consisting of N words, use the pre-trained CLIP model to extract word features e for each word. ω .
[0070] Step A42: Calculate e using the following formula ω Cross-attention with c:
[0071]
[0072] Where sofmax is the activation function, Q local Let K be the value of c after linear mapping. local V local e ω The values obtained after three different linear mappings through the attention mechanism, D is Q. local The dimension is then determined, and the multi-head attention (assuming 7 heads) is concatenated to obtain the final self-attention LocalAtt = {LocalAtt0, LocalAtt1, ..., LocalAtt7}. This final self-attention is then input into a feedforward neural network to obtain the intermediate feature f that integrates local information. local .
[0073] Step A43: If the first Transformer network T2 includes multiple layers (e.g., 2 or 3 layers), then the intermediate features f of the current layer are... local As the first action coding sequence c of the next layer, it is processed again according to step A42, and the feedforward neural network of the last layer will output the second action coding sequence c. ω .
[0074] For the second Transformer network T2, refer to Figure 4 b and Figure 5b. Using a pre-trained CLIP model, extract the text features of the entire given text, and then use them in conjunction with the second action encoding sequence. ω Calculate conditional self-attention to obtain the third action encoding sequence c. Ω (or a low-dimensional motion coding sequence that integrates global information c) Ω The illustrative steps are as follows:
[0075] Step A44: Given a text Ω consisting of N words, use the pre-trained CLIP model to extract text features e from the entire text. Ω and put e Ω spliced to c ω In the time dimension, the fourth action encoding sequence is obtained. (corresponding to) Figure 5 c in b w1 c w2 c w3 ... c wm e Ω (The sequence constituted)
[0076] Step A45: Calculate e using the following formula Ω and Self-attention:
[0077]
[0078] Where sofmax is the activation function, Q global K global V global for The values of D after three different linear mappings through the attention mechanism are Q. global The dimension is then used to concatenate the multi-head attention (assuming 15 heads, but other values such as 8, 12, etc.) to obtain the final self-attention GlobalAtt = {GlobalAtt0, GlobalAtt1, ..., GlobalAtt}. 15 Finally, the self-attention input is fed into the feedforward neural network to obtain the intermediate features f that fuse global information. global .
[0079] Step A46: If the second Transformer network T2 includes multiple layers (e.g., 9 or 10 layers), then the intermediate features f of the current layer... global c, as the second action encoding sequence of the next layer ω Following step A45 again, the last layer of the feedforward neural network will output the third action encoding sequence c. Ω .
[0080] For the third Transformer network T3, refer to Figure 4 b, Encode the third action sequence c Ω Input the third Transformer network, and output a probability distribution sequence of action indices. Each element of the probability distribution sequence contains the probability that the element is an index in the codebook. The index with the highest probability in each element is concatenated to obtain the generated action index sequence c. gen The illustrative steps are as follows:
[0081] Step A47: Calculate self-attention using the following formula:
[0082]
[0083] Where sofmax is the activation function, and Q, K, and V are the third action encoding sequences c. Ω After three different linear mappings of the attention mechanism, D is the dimension of Q. Then, the multi-head attention is concatenated to obtain the final self-attention GenAtt = {GenAtt0, GenAtt1, ..., GenAtt}. 15 Finally, the self-attention input is fed into the feedforward neural network to obtain the intermediate features f that fuse global information. gen ;
[0084] Step A48: In the case that the third Transformer network T3 includes multiple layers (such as 9 or 11 layers), the intermediate features f of the current layer are... gen c, as the third action encoding sequence of the next layer Ω Following step A47 again, the last layer of the feedforward neural network will output the action index sequence c. gen .
[0085] Step A5: Obtain the trained motion generation model, which includes the trained VQ-VAE model and the trained text-to-motion indexing model.
[0086] According to one embodiment of the present invention, step A5 is only to obtain the trained VQ-VAE model and the trained text-to-motion index model to constitute the trained motion generation model. Action index sequence c gen The quantized motion feature sequence, converted by the quantizer, is input into the decoder of the trained VQ-VAE model to generate the corresponding motion sequence.
[0087] According to an embodiment of the present invention, a system for training a motion generation model based on the method for training a motion generation model based on the foregoing embodiments is provided, comprising: a first training module configured to: reconstruct the sample action sequences by training a VQ-VAE model using multiple sample action sequences to obtain a trained VQ-VAE model including an encoder, a quantizer, a decoder, and a codebook; an action pair construction module configured to: convert each sample action sequence of the multiple sample action sequences into a corresponding sample action index sequence according to the encoder, quantizer, and codebook of the trained VQ-VAE model; and construct multiple action pairs, each action pair including a... The system comprises: a sample action index sequence and sample text describing the action states indicated by the corresponding sample action sequences; a second training module configured to: train a text-to-motion index model based on a Transformer network according to the multiple action pairs and the codebook, and learn to generate action index sequences from the text by using an attention mechanism to fuse local cross-attention and global conditional attention between the sample text and the sample action index sequence, thereby obtaining a trained text-to-motion index model; and a construction module configured to: obtain a trained motion generation model, which includes a trained VQ-VAE model and a trained text-to-motion index model.
[0088] According to one embodiment of the present invention, a method for generating a motion sequence that drives the movement of a virtual character based on text is also provided, comprising the steps of:
[0089] B1. Obtain the action description text used to describe the movement state of the virtual character;
[0090] B2. Obtain a trained motion generation model obtained by the method or system for training a motion generation model according to the foregoing embodiments, which includes a trained VQ-VAE model and a trained text-to-motion indexing model.
[0091] B3. Input the action description text into the trained text-to-motion index model to generate the action index sequence corresponding to the action description text;
[0092] B4. After converting the action index sequence corresponding to the action description text into a quantized action feature sequence, input it into the decoder of the trained VQ-VAE model to generate the action sequence corresponding to the action description text. The action sequence corresponding to the action description text can be used to drive the virtual character to move and generate animation. It should be noted that some implementers may worry that in this embodiment, since the corresponding action sequence is generated only from the action description text, without the action sequence corresponding to the text during training as input, the cross attention of the first Transformer network and the conditional attention of the second Transformer network lack an object for text computation. However, as those skilled in the art know, the input of the Transformer network during testing or application differs from that during training; during testing or application, it is not necessary to input the true value of the action sequence. When predicting the t-th frame, in addition to the text, the action sequence output by the Transformer network for frames 1 to t-1 is used as the first Transformer network and the second Transformer network. Furthermore, when predicting the 1st frame, the text features e of the entire text are used... Ω Enter directly (because here) Figure 5 {c in a ω1 ,…,c ωm If the first frame is empty (meaning no action has been generated yet), the second Transformer network generates the third action encoding sequence. The third Transformer network then generates a probability distribution sequence of action indices. The action index with the highest probability in the probability distribution sequence is taken as the first frame of the predicted action index sequence. Starting from the second frame, the first action encoding sequence can be obtained by looking up the codebook based on the previously predicted action index sequence. Based on the first action encoding sequence and the action description text, subsequent frames are predicted sequentially using the text-to-motion index model to obtain the final action index sequence. After converting the final action index sequence codebook into a quantized action feature sequence, it is input into the decoder of the trained VQ-VAE model to generate the action sequence corresponding to the action description text.
[0093] To demonstrate the effectiveness of the method of the embodiments of the present invention (which employs solutions to address defects (1) and (2)), see [link to relevant documentation]. Figure 6 a. This illustrates a graphical visualization of the action sequences generated from the text "a person quickly waves with his right hand." The existing MD method generates an action sequence where the person waves their left hand, while the existing T2M-GPT method generates an action sequence where the person waves their right hand, but not quickly enough. See also... Figure 6Figure b illustrates the graphical visualization of three methods for generating action sequences from the text "a person walk in a circle clockwise". In the existing MD method, the action sequence only shows the person walking, indicating the model does not understand the semantics of the text. The existing T2M-GPT method generates an action sequence where the person walks clockwise but does not form a complete circle. Therefore, compared to these two existing methods, the method of this embodiment significantly improves the accuracy of matching the action sequence with the text.
[0094] In addition, the inventors conducted comparative experiments with seven existing methods, and the results of the comparative experiments are shown in the table below:
[0095]
[0096] In the table, indicators marked with an upward ↑ indicate that higher values are better, indicators marked with a downward ↓ indicate that lower values are better, and indicators marked with a right → indicate that more dispersed distributions are better. As can be seen from the table, the method of this invention embodiment improves upon existing technologies in multiple metrics, achieving better performance.
[0097] Below are the references for the above seven existing technologies:
[0098] [1]Anindita Ghosh,Noshaba Cheema,Cennet Oguz,Christian Theobalt,andPhilipp Slusallek.Synthesis of compositional animations from textualdescriptions.In Proceedings of the IEEE / CVF international conference oncomputer vision,pages 1396–1406,2021.2,6;
[0099] [2]Chuan Guo,Xinxin Zuo,Sen Wang,and Li Cheng.Tm2t:Stochastic andtokenized modeling for the reciprocal generation of 3d human motions andtexts.In Computer Vision–ECCV 2022:17th European Conference,Tel Aviv,Israel,October 23–27,2022,Proceedings,Part XXXV,pages 580–597.Springer,2022.2,3,6;
[0100] [3]Chuan Guo,Shihao Zou,Xinxin Zuo,Sen Wang,Wei Ji,Xingyu Li,and LiCheng.Generating diverse and natural 3d human motions from text.InProceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition,pages 5152–5161,2022.2,3,4,5,6;
[0101] [4]Guy Tevet,Sigal Raab,Brian Gordon,Yoni Shafir,Daniel Cohen-or,andAmit Haim Bermano.Human motion diffusion model.In The Eleventh InternationalConference on Learning Representations,2023.2,3,5,6;
[0102] [5]Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondiffuse: Text-driven human motion generation with diffusion model.arXiv preprint arXiv:2208.15001,2022.2,3,5,6,7;
[0103] [6] Chen Xin, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, Jingyi Yu, and Gang Yu. Executing your commands via motion diffusion in latent space. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023.2, 3, 4, 5, 6;
[0104] [7] Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. T2m-gpt: Generating human motion from textual descriptions with discrete representations. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, 2, 3, 4, 5, 6, 7.
[0105] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.
[0106] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0107] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.
[0108] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for training a motion generation model, wherein the motion generation model generates a sequence of actions that drive the movement of a virtual character based on text, the method comprising: The VQ-VAE model is trained using multiple sample action sequences to reconstruct the sample action sequences, resulting in a trained VQ-VAE model including an encoder, quantizer, decoder, and codebook. Based on the encoder, quantizer, and codebook of the trained VQ-VAE model, each sample action sequence of the plurality of sample action sequences is converted into a corresponding sample action index sequence. Construct multiple action pairs, each action pair including a sample action index sequence and sample text describing the action state indicated by the corresponding sample action sequence; Based on the multiple action pairs and the codebook, a text-to-motion index model built on a Transformer network is trained. This model learns to generate action index sequences from the text by utilizing an attention mechanism to fuse local cross-attention and global conditional attention between sample text and sample action index sequences. The trained text-to-motion index model includes: a first Transformer network used to calculate cross-attention between each element of the first action encoding sequence and the word features of each word in the text, to extract a second action encoding sequence that fuses local information. The first action encoding sequence is concatenated based on the embedding vectors obtained from looking up each index in the codebook. A second Transformer network is used to calculate conditional attention based on local features and the overall text features, to extract a third action encoding sequence that fuses local and global information. A third Transformer network is used to generate a probability distribution sequence of action indices based on the third action encoding sequence. Each element of the probability distribution sequence contains the probability that the element corresponds to each index in the codebook. The indexes with the highest probability in each element are concatenated to obtain the generated action index sequence. A trained motion generation model is obtained, which includes a trained VQ-VAE model and a trained text-to-motion indexing model.
2. The method according to claim 1, characterized in that, The trained text-to-motion indexing model is trained in the following manner: Each action pair is input as training data into the text-to-motion index model, which outputs a probability distribution sequence of action indices. The cross-entropy loss between the action index sequence in the action pair and the probability distribution sequence of the output action index is determined. The gradient is calculated based on the cross-entropy loss and backpropagation is used to update the trainable parameters of the third Transformer network, the second Transformer network, and the first Transformer network.
3. The method according to claim 2, characterized in that, The first, second, and third Transformer networks each contain multiple layers of Transformer subnetworks, and each Transformer subnetwork is implemented based on a multi-head attention mechanism.
4. The method according to any one of claims 1-3, characterized in that, The trained VQ-VAE model was obtained in the following manner: Multiple sample action sequences are acquired, where each frame of the sample action sequence contains control information for the joints of the virtual character. The VQ-VAE model is trained using the loss function of the VQ-VAE model to reconstruct multiple sample action sequences, resulting in a trained VQ-VAE model. During training, the encoder generates action feature sequences based on the sample action sequences, the quantizer obtains the quantized action feature sequences corresponding to the action feature sequences based on the codebook, and the decoder reconstructs the sample action sequences based on the quantized action feature sequences.
5. The method according to any one of claims 1-3, characterized in that, The trained VQ-VAE model was obtained in the following manner: Multiple sample action sequences are acquired, where each frame of the sample action sequence contains control information for the joints of the virtual character. Obtain the motion relationship mask set for the virtual character. The motion relationship mask is a matrix that indicates the motion correlation between the control points of the character. The virtual character is divided into multiple limb segments. The motion correlation of control points located in the same limb segment is set to 0, and the motion correlation of control points located in different limb segments is set to negative infinity. The VQ-VAE model is obtained, whose encoder uses a fourth Transformer network. This fourth Transformer network is configured to: during single-head self-attention calculation of control information for each control point in each frame of the input action sequence, incorporate QK into the attention mechanism. T The value is superimposed on the motion relationship mask to shield attentional interference between irrelevant limb segments; The VQ-VAE model is trained using the loss function of the VQ-VAE model to reconstruct multiple sample action sequences, resulting in a trained VQ-VAE model. During training, the encoder generates action feature sequences based on the sample action sequences, the quantizer obtains the quantized action feature sequences corresponding to the action feature sequences based on the codebook, and the decoder reconstructs the sample action sequences based on the quantized action feature sequences.
6. A system for training a motion generation model based on the method described in any one of claims 1-5, characterized in that, include: The first training module is configured to: use multiple sample action sequences to train a VQ-VAE model to reconstruct the sample action sequences, and obtain a trained VQ-VAE model including an encoder, quantizer, decoder and codebook. The action pair construction module is configured to: convert each sample action sequence of the plurality of sample action sequences into a corresponding sample action index sequence according to the encoder, quantizer and codebook of the trained VQ-VAE model; and construct a plurality of action pairs, each action pair including a sample action index sequence and sample text describing the action state indicated by the corresponding sample action sequence. The second training module is configured to: train a text-to-motion index model based on a Transformer network according to the multiple action pairs and the codebook, and learn to generate action index sequences from the text by using an attention mechanism to fuse local cross attention and global conditional attention between sample text and sample action index sequences, thereby obtaining the trained text-to-motion index model. The building module is configured to obtain a trained motion generation model, which includes a trained VQ-VAE model and a trained text-to-motion indexing model.
7. A method for generating a sequence of actions to drive the movement of a virtual character based on text, characterized in that, include: Obtain the action description text used to describe the movement state of the virtual character; Obtain a trained motion generation model obtained by the method according to any one of claims 1-5 or the system according to claim 6, which includes a trained VQ-VAE model and a trained text-to-motion indexing model; The action description text is input into the trained text-to-motion index model to generate an action index sequence corresponding to the action description text; After converting the action index sequence corresponding to the action description text into a quantized action feature sequence, it is input into the decoder of the trained VQ-VAE model to generate the action sequence corresponding to the action description text. The action sequence corresponding to the action description text can be used to drive the virtual character to move in order to generate animation.
8. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method according to any one of claims 1 to 5 and 7.
9. An electronic device, characterized in that, include: One or more processors; as well as Memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1 to 5 and 7 by executing the executable instructions.