Text-driven action generation method, device, equipment, storage medium and program product
By combining the autoregressive motion transformer, vector quantization variational autoencoder and non-autoregressive motion transformer in the text-driven action generation method, the problem of insufficient accuracy of action generation in the prior art is solved, and higher accuracy and quality of action generation are achieved.
Patent Information
- Application Number
- CN202411847431.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-12-16
AI Technical Summary
The existing text-driven action generation methods have shortcomings in the accuracy of action generation, especially the autoregressive model is difficult to capture nonlinear characteristics, but the non-autoregressive model has a high model complexity and a low prediction accuracy.
A text-driven action generation method is proposed. By inputting the input text and the empty motion sequence into the autoregressive motion transformer, the basic motion sequence is obtained; then the basic motion sequence is processed by the vector quantization variational autoencoder to obtain the residual index; then the input text, the basic motion sequence and the residual index are inputted to the non-autoregressive motion transformer, and the residual motion sequence is obtained; finally, the residual motion sequence is processed by the decoder of the vector quantization variational autoencoder to obtain the target motion sequence.
By combining the autoregressive motion transformer, vector quantization variational autoencoder and non-autoregressive motion transformer, the accuracy of the text-driven action generation method is improved, jitter and cumulative errors are reduced, and the quality and fidelity of action generation are improved.
Smart Images

Figure CN119313785B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence, and particularly to a text-driven action generation method, a text-driven action generation device, a text-driven action generation device, a storage medium, and a computer program product. Background Art
[0002] In the fields of digital human video live broadcast, virtual special effects production, and game character driving, human body motion synthesis technology can provide strong support for the dynamic performance of 3D characters, effectively improving the realism and smoothness of motion.
[0003] In the current action generation technology, the text-driven action generation method provides an efficient action generation method. In practical applications, if the data shows an obvious linear relationship and the computing resources are limited, an autoregressive model may be used to implement text-driven action generation. If the data is complex and parallel computing is required, or long sequence data needs to be processed, a non-autoregressive model can be used to implement text-driven action generation.
[0004] However, in the current text-driven action generation method, there is still room for improvement in the accuracy of action generation. Among them, when using an autoregressive model to implement text-driven action generation, it is assumed that the generation process of the time series is linear, and it is difficult to capture non-linear characteristics; when using a non-autoregressive model to implement text-driven action generation, its model complexity is greater than that of the autoregressive model, and the prediction accuracy is lower than that of the autoregressive model.
[0005] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main purpose of the present application is to provide a text-driven action generation method, a text-driven action generation device, a text-driven action generation device, a storage medium, and a computer program product, aiming to solve the technical problem of poor accuracy of text-driven action generation.
[0007] To achieve the above object, the present application proposes a text-driven action generation method, and the text-driven action generation method includes:
[0008] Input the input text and an empty motion sequence into an autoregressive motion transformer to obtain a basic motion sequence;
[0009] Process the basic motion sequence through the encoder and quantizer in a vector quantization variational autoencoder in sequence to obtain a residual index;
[0010] Input the input text, the basic motion sequence, and the residual index into a non-autoregressive motion transformer to obtain a residual motion sequence;
[0011] Process the residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain a target motion sequence.
[0012] In one embodiment, before the step of sequentially processing the basic motion sequence through the encoder and the quantizer in the vector quantization variational autoencoder to obtain a residual index, it includes:
[0013] Add an inter-frame motion speed loss and an acceleration loss to the reconstruction loss function of the vector quantization variational autoencoder to be trained to obtain a new reconstruction loss function;
[0014] Train a vector quantization variational autoencoder based on the new reconstruction loss function.
[0015] In one embodiment, before the step of sequentially processing the basic motion sequence through the encoder and the quantizer in the vector quantization variational autoencoder to obtain a residual index, it further includes:
[0016] During the training of the vector quantization variational autoencoder, map different motions to corresponding vector quantization layers, where the number of vector quantization layers is the number of quantizers determined by uniform sampling of the quantizer.
[0017] In one embodiment, the step of processing the residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain a target motion sequence includes:
[0018] Add a sparse attention mask to the residual motion sequence to obtain a new residual motion sequence;
[0019] Process the new residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain a target motion sequence.
[0020] In one embodiment, the step of inputting the input text and an empty motion sequence into the autoregressive motion transformer to obtain a basic motion sequence includes:
[0021] In the latent space, based on the text embedding conditions extracted from the input text, predict the distribution of possible next indices through the motion autoregressive transformer and generate sub-motion sequences one by one until a complete sequence is obtained, and use the complete sequence as the basic motion sequence.
[0022] In one embodiment, the step of inputting the input text, the basic motion sequence, and the residual index into the non-autoregressive motion transformer to obtain a residual motion sequence includes:
[0023] Embed, as the input of a non-autoregressive motion transformer, the text embedding condition extracted from the input text, the base motion sequence, and the residual index serving as the residual layer indicator of the quantization layer, and predict, by the non-autoregressive motion transformer, the residual motion sequence of the residual quantization layer.
[0024] In addition, to achieve the above object, the present application also proposes a text-driven action generation device, which includes:
[0025] An autoregressive module, configured to input the input text and an empty motion sequence into an autoregressive motion transformer to obtain a base motion sequence;
[0026] An encoding quantization module, configured to sequentially process the base motion sequence through an encoder and a quantizer in a vector quantization variational autoencoder to obtain a residual index;
[0027] A non-autoregressive module, configured to input the input text, the base motion sequence, and the residual index into a non-autoregressive motion transformer to obtain a residual motion sequence;
[0028] A decoding module, configured to process the residual motion sequence through a decoder in the vector quantization variational autoencoder to obtain a target motion sequence.
[0029] In addition, to achieve the above object, the present application also proposes a text-driven action generation device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the text-driven action generation method as described above.
[0030] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the text-driven action generation method as described above.
[0031] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the text-driven action generation method as described above.
[0032] One or more technical solutions proposed by the present application have at least the following technical effects:
[0033] In this application, first, the input text and an empty motion sequence are input into an autoregressive motion transformer to obtain a basic motion sequence; then, the basic motion sequence is processed successively through the encoder and the quantizer in a vector quantization variational autoencoder to obtain residual indices; next, the input text, the basic motion sequence, and the residual indices are input into a non-autoregressive motion transformer to obtain a residual motion sequence; finally, the residual motion sequence is processed through the decoder in the vector quantization variational autoencoder to obtain the target motion sequence. Thus, by combining the advantages of using the autoregressive motion transformer to predict future sequences using the historical data of the motion sequence itself, the vector quantization variational autoencoder to capture and represent richer and more complex data distributions in the discrete latent space, and the non-autoregressive motion transformer to reduce the problem of error accumulation caused by sequential dependence, the accuracy of text-driven action generation is further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0035] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic flowchart provided for the first embodiment of the text-driven action generation method of this application;
[0037] Figure 2 It is a schematic diagram of double-layer motion modeling generation provided for the first embodiment of the text-driven action generation method of this application;
[0038] Figure 3 It is a schematic diagram of residual vector motion quantization provided for the first embodiment of the text-driven action generation method of this application;
[0039] Figure 4 It is a schematic diagram of sliding focus sparse attention provided for the first embodiment of the text-driven action generation method of this application;
[0040] Figure 5 It is a schematic diagram of the module structure of the text-driven action generation device according to the embodiment of this application;
[0041] Figure 6 It is a schematic diagram of the device structure of the hardware operating environment involved in the text-driven action generation method according to the embodiment of this application.
[0042] The realization of the purpose, functional features and advantages of this application will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners
[0043] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0044] In order to better understand the technical solutions of this application, the following will be described in detail in conjunction with the drawings of the specification and specific implementation manners.
[0045] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a text-driven action generation device, etc. that can implement the above functions. The following takes a text-driven action generation device as an example to illustrate this embodiment and the following embodiments.
[0046] Based on this, the embodiment of this application provides a text-driven action generation method, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the text-driven action generation method of this application.
[0047] In this embodiment, the text-driven action generation method includes steps S10 to S40:
[0048] Step S10, input the input text and an empty motion sequence into an autoregressive motion transformer to obtain a basic motion sequence;
[0049] In one embodiment, referring to Figure 2 , input the input text and an empty motion sequence into an autoregressive motion transformer to obtain a basic motion sequence. Among them, autoregressive movement is a prediction method based on historical data. It assumes that the current value depends on past values and is usually used to predict time series data with periodicity. In the autoregressive movement model, the current value of the time series is a linear combination of its own past values, plus an error term.
[0050] In a feasible implementation manner, step S10 includes:
[0051] In the latent space, based on the text embedding conditions extracted from the input text, predict the distribution of possible next indices through a motion autoregressive transformer and generate sub-motion sequences one by one until a complete sequence is obtained, and use the complete sequence as the basic motion sequence.
[0052] In one embodiment, in the latent space, each text embedding is obtained by extracting information from the prompt text, i.e., the input text, using CLIP. Given a number of previous indices, based on the text embedding conditions, the motion autoregressive transformer predicts the distribution of possible next indices. Then, the sub-motion sequences are generated autoregressively one by one, and finally the complete sequence is represented to obtain the basic motion sequence.
[0053] Step S20: The basic motion sequence is successively processed by the encoder and the quantizer in the vector quantization variational autoencoder to obtain the residual indices.
[0054] In one embodiment, referring to Figure 3 , in the residual vector quantization motion quantization, the residual vector quantization motion encoder is used to fit the discrete distribution of the human motion space to the latent space, and then the decoder is used to restore the original human motion.
[0055] In one embodiment, the residual vector quantization motion quantization adopts a VQ-VAE (Vector-Quantized Variational Autoencoder) variational autoencoder, which compresses and generates data by encoding the input data into a discrete latent space and then decoding it back to the original data space. Through the discrete latent space, VQ-VAE learns the complex distribution of the data and can generate high-quality samples when generating data. It includes the following execution processes: 1. Encoding process: The input data is mapped to a continuous latent space through the encoder, and the input data (such as an image) is encoded into a latent representation. 2. Quantization process: The continuous latent space representation is discretized into a set of fixed vectors, which are called codebooks. That is, in the vector quantization layer, the continuous latent representation output by the encoder is mapped to a predefined discrete latent vector space (corresponding to the embedding space in the figure) to achieve vector quantization. 3. Decoding process: The quantized discrete latent representation is received and the reconstructed input data is generated. In this way, the discrete latent representation is converted back to the original data space through the decoder.
[0056] That is to say, the residual VQ-VAE includes a motion encoder, a decoder, and a discrete codebook. The encoder maps the motion sequence to the latent vector space, performs vector quantization based on the codebook, and the decoder establishes a projection mapping between the motion sequence and the latent representation obtained from the codebook.
[0057] Among them, in the VQ-VAE, residual layers are usually used in the encoder and decoder to help the model better learn the feature representation of the data. The calculation method of the residual layer usually involves processing the input data through a series of convolutional layers, and then adding the processed result to the original input data to form a residual connection. This connection method can help the model better learn the feature representation of the data. Especially when dealing with deep networks, it can effectively alleviate the vanishing gradient problem. Through the residual connection, the model can more easily learn the local and global features of the data, thereby improving the performance of the model. In addition, the residual layer can also help the model converge faster during the training process and improve the training efficiency.
[0058] In a feasible implementation manner, before the step S20, it includes:
[0059] Adding the inter-frame motion speed loss and acceleration loss to the reconstruction loss function of the vector quantization variational autoencoder to be trained to obtain a new reconstruction loss function;
[0060] Training a vector quantization variational autoencoder based on the new reconstruction loss function.
[0061] Current motion quantization methods usually result in unintentional jitter in the synthesized actions, assuming that this jitter stems from sudden changes in inter-frame speed and acceleration. For motion sequences, traditional representations only fit and focus on the positions of the joints, but actual motion is related to both its speed and acceleration. Therefore, based on the residual quantization design, and by introducing a continuity guidance loss to alleviate the inter-frame discontinuity in the motion sequence, the problem caused by jitter is further improved.
[0062] Currently, the model is trained by minimizing the reconstruction error, embedding error, and constraint loss: In this step, the model is trained by minimizing the reconstruction error, embedding error, and constraint loss. Among them, the reconstruction error measures the difference between the original data and the output of the decoder, and the quantizer measures the difference between the latent vector and the closest prototype vector.
[0063] Among them, the decoder maps the discrete latent vector back to the high-dimensional continuous data space to generate the output of the model: In this step, the discrete latent vector is mapped back to the high-dimensional continuous data space through the decoder network to generate the output of the model.
[0064] During the training process, the loss function of the VQ-VAE includes the reconstruction loss (encouraging the decoded data to be close to the original input) and the regularization term (encouraging the latent representation to conform to a certain predefined distribution, usually the standard normal distribution). By optimizing these loss functions, the model can learn effective data representations and be able to reconstruct the original data from these representations.
[0065] In this embodiment, the loss function of the encoder is optimized during model training. A motion persistence guidance mechanism is added to the motion reconstruction loss function, and the motion speed loss and acceleration loss are used to alleviate jitter during the motion generation process.
[0066] Furthermore, to obtain a more continuous reconstruction effect based on motion continuity guidance, in motion processing, speed loss and acceleration loss are added to the reconstruction loss to prevent jitter during the motion generation process. The specific calculation formula is:
[0067]
[0068] where is an adjustable coefficient, represents the reconstructed motion sequence, represents the real motion sequence, is the speed representation of the motion sequence, is the acceleration representation of the motion sequence.
[0069] In a feasible implementation manner, before the step S20, it further includes:
[0070] During the training of the vector quantization variational autoencoder, different motions are mapped to the corresponding vector quantization layers, where the number of vector quantization layers is the number of quantizers determined by the uniform sampling quantizer.
[0071] During training, for each input instance, multiple residual quantizers are uniformly sampled, where the number of quantizers is obtained by uniform sampling. Only the information from the first layers is used during the training process. Motions with different granularities (different motion type categories, motion amplitudes) within the motion are mapped to different vector quantization layers. For example, some quantization layers are responsible for small-amplitude motions, and another part of the quantization layers is responsible for large-amplitude motions. By strengthening this hierarchical motion structure, the utilization rate of common motion patterns in the basic codebook can be increased, thus aligning with the double-layer generation modeling module.
[0072] Step S30, input the input text, the basic motion sequence, and the residual index into the non-autoregressive motion transformer to obtain the residual motion sequence;
[0073] In one embodiment, referring to Figure 2, the input text, the basic motion sequence, and the residual index are input into the non-autoregressive motion transformer to obtain the residual motion sequence. Among them, non-autoregressive movement is a prediction method that does not rely on historical data. It usually uses other variables or features to predict future values. In a non-autoregressive model, the current value of a time series is a function of other variables or features, rather than a function of its own past values.
[0074] Among them, the non-autoregressive motion transformer is used to model the tokens from the residual quantization layer to obtain the residual layer indicator, that is, to obtain Figure 2 the residual index in
[0075] In a feasible implementation manner, the step S30 includes:
[0076] Embed the text embedding condition extracted from the input text, the basic motion sequence, and the residual index as the residual layer indicator of the quantization layer into the input of the non-autoregressive motion transformer, and the non-autoregressive motion transformer predicts the residual motion sequence of the residual quantization layer.
[0077] In one embodiment, referring to Figure 2 , all the quantization layers and the tokens before them are embedded and aggregated into a token input. Using the input text, the residual layer indicator, and the basic motion sequence embedding as inputs, wherein the non-autoregressive motion transformer is trained to predict the motion sequence of the residual quantization layer, so that the non-autoregressive motion transformer predicts the residual motion sequence of the residual quantization layer.
[0078] Step S40, process the residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain the target motion sequence.
[0079] In one embodiment, referring to Figure 2 , process the residual motion sequence obtained by the non-autoregressive motion transformer through the decoder in the vector quantization variational autoencoder to obtain the target motion sequence.
[0080] In a feasible implementation manner, the step S40 includes:
[0081] Add a sparse attention mask to the residual motion sequence to obtain a new residual motion sequence;
[0082] Process the new residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain the target motion sequence.
[0083] In this embodiment, a transformer in the form of a masked decoder is designed in the self-attention mechanism. Feature embedding is performed for text conditions and is required to be continuously visible in the attention mask as conditional features, so as to continuously monitor the generation of motion sequences in the transformer. A sparse attention mask is added to the input motion sequence to control the local visibility of the motion subsequence when passing through the self-attention technique, reducing the mutual interference during the generation of long-sequence actions. For the autoregressive transformer, the motion subsequence can only see the current information, while for the non-autoregressive transformer, it can additionally see the future local information.
[0084] Referring to Figure 4 , based on fixed-condition sparse attention, the mask design in the attention calculation process is shown; the blue in the figure represents the fixed range of conditional tokens, and the red represents the sliding window range of adjacent motion tokens. It uniquely sets the text tokens to be globally visible while allocating the sliding window to each token. This configuration not only enhances the model's ability to reason about long sequences but also significantly improves the computational efficiency.
[0085] The sliding focus sparse attention mechanism restricts the motion tokens to two receptive fields: a wider one to access all text conditions and a narrower one to focus on the local context. In the training phase, the text tokens are used as the reference points for the autoregressive transformer, while for the non-autoregressive transformer, a combination of the non-autoregressive identifier (NAR ID) and the text tokens is used.
[0086] For the autoregressive transformer, its mask takes the form of a peak, allowing its sliding window to only view prior information; for the non-autoregressive transformer, it has a bidirectional context view and can view future information.
[0087] The text-driven action generation method in this application is applicable to fields such as digital human driving, virtual character animation, and robot control. In an application scenario, to overcome the problems of jitter and cumulative error existing in the existing action generation technology and achieve smoother, more natural, and accurate action generation, a two-layer generation model architecture based on a transformer is provided. In the latent space of motion residual vectorization, a two-layer generation modeling structure composed of an autoregressive motion transformer and a non-autoregressive motion transformer is used to reconstruct the discrete distribution of the residual space based on sliding focus sparse attention.
[0088] Specifically, in the latent space, each text embedding extracts information from the prompt text query, i.e., the input text, using CLIP. Given a number of previous indices, based on the text condition, the motion autoregressive transformer predicts the distribution of possible next indices. Then, the sub-motion sequences are generated autoregressively one by one, and finally the complete sequence is represented to obtain the base motion sequence. The non-autoregressive transformer is used to model the tokens from the residual quantization layer to obtain the residual motion sequence. During training, the number of quantizer layers is randomly selected. Subsequently, all the embeddings of these quantization layers and the previous tokens are aggregated as the token input. Among them, using the text condition, the residual layer indicator, and the motion sequence embedding as inputs, the non-autoregressive transformer is trained to predict the residual motion sequence of the residual quantization layer. Finally, the continuous motion guidance strategy is used to reduce the jitter in the generation process, and a decoder composed of a 6-layer residual convolutional neural network is used to obtain the reconstructed target motion sequence.
[0089] In this way, through the motion continuity guidance strategy and the hierarchical quantization screening strategy, the quality of the motion reconstruction process and the realism of the generated motion are improved. The conditional control and local understanding ability of the motion transformer are enhanced through double-layer generation modeling, thus realizing fine-grained motion generation driven by text description. This method reduces the jitter phenomenon and cumulative error in action synthesis and improves the generation quality of continuous action sequences.
[0090] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the text-driven action generation method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0091] This application also provides a text-driven action generation device. Please refer to Figure 5 , the text-driven action generation device includes:
[0092] The autoregressive module 10 is used to input the input text and the empty motion sequence into the autoregressive motion transformer to obtain the base motion sequence;
[0093] The encoding quantization module 20 is used to process the base motion sequence through the encoder and the quantizer in the vector quantization variational autoencoder in sequence to obtain the residual index;
[0094] The non-autoregressive module 30 is used to input the input text, the base motion sequence, and the residual index into the non-autoregressive motion transformer to obtain the residual motion sequence;
[0095] The decoding module 40 is used to process the residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain the target motion sequence.
[0096] In one embodiment, the encoding quantization module 20 is further used for:
[0097] Before the step of obtaining the residual index by successively processing the basic motion sequence through the encoder and the quantizer in the vector quantization variational autoencoder,
[0098] Add the inter-frame motion speed loss and the acceleration loss to the reconstruction loss function of the vector quantization variational autoencoder to be trained to obtain a new reconstruction loss function;
[0099] Train the vector quantization variational autoencoder based on the new reconstruction loss function.
[0100] In one embodiment, the encoding and quantization module 20 is further configured to:
[0101] Before the step of obtaining the residual index by successively processing the basic motion sequence through the encoder and the quantizer in the vector quantization variational autoencoder,
[0102] During the training of the vector quantization variational autoencoder, map different motions to the corresponding vector quantization layers, where the number of vector quantization layers is the number of quantizers determined by the uniform sampling quantizer.
[0103] In one embodiment, the decoding module 40 is further configured to:
[0104] Add a sparse attention mask to the residual motion sequence to obtain a new residual motion sequence;
[0105] Process the new residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain the target motion sequence.
[0106] In one embodiment, the autoregressive module 10 is further configured to:
[0107] In the latent space, based on the text embedding conditions extracted from the input text, predict the distribution of possible next indices through the motion autoregressive transformer and generate sub-motion sequences one by one until a complete sequence is obtained, and use the complete sequence as the basic motion sequence.
[0108] In one embodiment, the non-autoregressive module 30 is further configured to:
[0109] Embed the text embedding conditions, the basic motion sequence, and the residual index serving as the residual layer indicator of the quantization layer extracted from the input text as the input of the non-autoregressive motion transformer, and predict the residual motion sequence of the residual quantization layer by the non-autoregressive motion transformer.
[0110] The text-driven action generation device provided by this application adopts the text-driven action generation method in the above-mentioned embodiment, and can solve the technical problem of poor accuracy in text-driven action generation. Compared with the prior art, the beneficial effects of the text-driven action generation device provided by this application are the same as those of the text-driven action generation method provided by the above-mentioned embodiment, and other technical features in the text-driven action generation device are the same as the features disclosed in the method of the above-mentioned embodiment, which will not be elaborated here.
[0111] This application provides a text-driven action generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text-driven action generation method in the first embodiment above.
[0112] Refer to the following Figure 6 , which shows a schematic structural diagram of a text-driven action generation device suitable for implementing the embodiments of this application. The text-driven action generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Player: portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The text-driven action generation device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of this application.
[0113] As shown in Figure 6As shown, the text-driven action generation device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the text-driven action generation device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the text-driven action generation device to communicate with other devices wirelessly or wiredly to exchange data. Although the text-driven action generation device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.
[0114] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0115] The text-driven action generation device provided by the present application adopts the text-driven action generation method in the above embodiments, and can solve the technical problem of poor accuracy in text-driven action generation. Compared with the prior art, the beneficial effects of the text-driven action generation device provided by the present application are the same as those of the text-driven action generation method provided by the above embodiments, and other technical features in the text-driven action generation device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0116] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0117] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0118] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the text-driven action generation method in the above embodiments.
[0119] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0120] The above computer-readable storage medium can be included in the text-driven action generation device; or it can exist separately without being assembled into the text-driven action generation device.
[0121] The above computer-readable storage medium carries one or more programs, which, when executed by a text-driven motion generation device, cause the text-driven motion generation device to: input the input text and an empty motion sequence into an autoregressive motion transducer to obtain a basic motion sequence; sequentially process the basic motion sequence through the encoder and the quantizer in a vector quantization variational autoencoder to obtain a residual index; input the input text, the basic motion sequence, and the residual index into a non-autoregressive motion transducer to obtain a residual motion sequence; and process the residual motion sequence through the decoder in the vector quantization variational autoencoder to obtain a target motion sequence.
[0122] Computer program code for performing the operations of this application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through an Internet service provider using the Internet).
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0124] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0125] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned text-driven action generation method, which can solve the technical problem of poor accuracy in text-driven action generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the text-driven action generation method provided by the above embodiments, and will not be elaborated here.
[0126] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the text-driven action generation method as described above are implemented.
[0127] The computer program product provided by the present application can solve the technical problem of poor accuracy in text-driven action generation. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the text-driven action generation method provided by the above embodiments, and will not be elaborated here.
[0128] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A text-driven action generation method, characterized in that: The text-driven action generation method comprises: Input the input text and the empty motion sequence into the autoregressive motion transformer to obtain the basic motion sequence; The basic motion sequence is processed by an encoder and a quantizer in a vector quantized variational autoencoder in sequence to obtain a residual index; The text embedding condition based on the text embedding condition extracted from the input text, the base motion sequence and the residual index as the residual layer indicator of the quantization layer are embedded as the input of the non-autoregressive motion transformer, and the residual motion sequence of the residual quantization layer is obtained by the non-autoregressive motion transformer prediction; A sparse attention mask is added to the residual motion sequence to obtain a new residual motion sequence; the new residual motion sequence is processed by the decoder in the vector quantization variational autoencoder to obtain a target motion sequence.
2. The text-driven action generation method according to claim 1, characterized in that: The step of sequentially processing the basic motion sequence through an encoder and a quantizer in a vector quantized variational autoencoder to obtain a residual index includes: Adding the inter-frame motion speed loss and the acceleration loss to the reconstruction loss function of the vector quantized variational autoencoder to be trained to obtain a new reconstruction loss function; Based on the new reconstruction loss function, a vector quantized variational autoencoder is trained.
3. The text-driven action generation method according to claim 1, characterized in that: Before the step of sequentially processing the basic motion sequence through an encoder and a quantizer in a vector quantized variational autoencoder to obtain a residual index, the step further includes: During training of the vector quantized variational autoencoder, different motions are mapped to corresponding vector quantization layers, where the number of vector quantization layers is the number of quantizers determined by the uniform sampling quantizer.
4. The text-driven action generation method according to claim 1, characterized in that: The step of inputting the input text and the empty motion sequence into the autoregressive motion transformer to obtain the basic motion sequence comprises: In the latent space, based on the text embedding conditions extracted from the input text, the distribution of possible next indexes is predicted through a motion autoregressive transformer and sub-motion sequences are generated one by one until a complete sequence is obtained, and the complete sequence is used as the basic motion sequence.
5. A text-driven action generation device, characterized in that: The text-driven action generation device comprises: An autoregressive module, used for inputting input text and an empty motion sequence into an autoregressive motion transformer to obtain a basic motion sequence; An encoding and quantization module, used for sequentially processing the basic motion sequence through an encoder and a quantizer in a vector quantization variational autoencoder to obtain a residual index; A non-autoregressive module, used for embedding the residual index based on the text embedding condition extracted from the input text, the basic motion sequence and the residual layer indicator as the quantization layer as the input of the non-autoregressive motion transformer, and the non-autoregressive motion transformer predicts the residual motion sequence of the residual quantization layer; A decoding module is used to add a sparse attention mask to the residual motion sequence to obtain a new residual motion sequence; the new residual motion sequence is processed by the decoder in the vector quantization variational autoencoder to obtain a target motion sequence.
6. A text-driven action generation device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the text-driven action generation method according to any one of claims 1 to 4.
7. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the text-driven action generation method according to any one of claims 1 to 4 are implemented.
8. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the text-driven action generation method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Voice-driven whole body action generation method
CN118570344A
Facial animation synthesis method, storage medium and electronic equipment
CN119107391A