A text-to-video generation method, product, device, and storage medium
Through the optimized distillation strategy and multi-channel state space model, the problems of the existing Wensheng video model's calculation power consumption and inferred inference efficiency when generating high-resolution videos are solved, and efficient and fast high-resolution video generation is achieved.
Patent Information
- Application Number
- CN202510406455.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-02
AI Technical Summary
When generating high-resolution videos, the training process consumes huge computing power, and the inference efficiency is inefficient, making it difficult to quickly generate high-resolution videos.
The optimized distillation strategy is adopted, and the student model is denoised sequentially through the student model and the diffusion model based on the attention mechanism, and the student model parameters are updated based on the loss of the denoising result to generate a literary video model with high inference efficiency. Combined with the second diffusion model based on the multi-channel state space model, the video resolution and generation efficiency are further improved.
It significantly reduces the computing power consumption during the training process, improves the inference speed and resolution of Wensheng Video, and improves the user experience.
Smart Images

Figure CN119946378B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method, product, device, and storage medium for generating videos from text. Background Art
[0002] With the rapid development of artificial intelligence generated content (AIGC) technology, text-to-video models such as SoRA (a large video generation model proposed by OpenAI) have been widely used. This model can generate high-definition videos based on text descriptions, thus providing a new method for rapid prototyping and proof of concept in the creative content production, film, animation, game, and advertising industries.
[0003] However, current text-to-video models, such as Stable Diffusion Video, Imagen-Video, etc., have a large number of parameters. When using these models to generate high-resolution videos, they usually first perform preliminary training on small-resolution videos and then continuously train a super-resolution model to achieve high-resolution image generation. This cascaded training method consumes a huge amount of computing power and poses a great challenge to computing resources. In addition, current text-to-video models usually use diffusion models to denoise to complete the generation of video content, which leads to low inference efficiency. Especially when generating high-resolution images, the generation speed becomes very slow. Therefore, it is necessary to improve the inference speed for high-resolution videos.
[0004] In summary, how to avoid a cumbersome and computationally expensive training process in the training stage to obtain a model for high-resolution content generation, and how to quickly complete the generation of high-resolution video content in the inference stage are urgent problems to be solved currently. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide a method, product, device, and storage medium for generating videos from text, which can improve the resolution of videos generated from text and the efficiency of video generation from text. The specific solutions are as follows:
[0006] In the first aspect, this application discloses a method for generating videos from text, including:
[0007] Input the target text description and the first noise vector into the first text-to-video model to generate a video matching the target text description and the corresponding first video latent vector. The first text-to-video model is obtained by training a preset student model using a training set according to a preset distillation strategy. The preset distillation strategy is to denoise the historical high-resolution videos in the training set using the preset student model and the first diffusion model based on the attention mechanism in sequence, and update the model parameters of the preset student model based on the loss corresponding to the denoising results.
[0008] Upsample the first video latent vector, and concatenate the obtained first sampled vector and the second noise vector to obtain a first concatenated vector.
[0009] Input the first concatenated vector into the second text-to-video model to generate a target text-to-video matching the target text description. The second text-to-video model is obtained by training the second diffusion model based on the multi-way state space model using a training set.
[0010] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above text-to-video generation methods.
[0011] The present application also provides an electronic device, including a processor and a memory. When the processor executes the computer program stored in the memory, the above text-to-video generation method is implemented.
[0012] The present application also provides a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, the above text-to-video generation method is implemented.
[0013] It can be seen that the present application optimizes the traditional distillation strategy. The optimized distillation strategy undergoes two denoising processes. First, the student model performs the first denoising process on the historical high-resolution videos in the training set, and then the first denoising result is input into the first diffusion model based on the attention mechanism for the second denoising process. The model parameters of the student model are updated based on the loss corresponding to the two denoising results, so as to obtain the first text-to-video model with high inference efficiency. Moreover, based on the diffusion model based on the attention mechanism, the present application combines the second diffusion model based on the multi-way state space model. By reasoning the text-to-video output based on the attention mechanism again, the resolution of the text-to-video can be further improved. In addition, since the multi-way state space model can only process the features within the specified state space compared with the attention mechanism, the generation efficiency of the text-to-video is improved, thereby improving the user experience. Description of the Drawings
[0014] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained according to the provided drawings.
[0015] Figure 1 It is a flowchart of a text-to-video generation method disclosed in the present application;
[0016] Figure 2 It is a block diagram of a specific distillation training process disclosed in the present application;
[0017] Figure 3 It is a schematic diagram of a specific discriminator structure disclosed in the present application;
[0018] Figure 4 It is a schematic diagram of a specific discriminator sub-module structure disclosed in the present application;
[0019] Figure 5 It is a flowchart of a specific training process of the second text-to-video model disclosed in the present application;
[0020] Figure 6 It is a schematic diagram of a text-to-video model structure disclosed in the present application, where Figure 6 (a) is a traditional diffusion model based on the DIT diffusion module, Figure 6 (b) is for Figure 6 the diffusion model after optimizing the DIT diffusion module in (a);
[0021] Figure 7 It is a schematic diagram of a specific sequence processing unit structure disclosed in the present application;
[0022] Figure 8 It is a schematic diagram of a specific 4-way state space model scan disclosed in the present application;
[0023] Figure 9 It is a schematic diagram of the network structure of a specific state space model disclosed in the present application;
[0024] Figure 10 It is a flowchart of a specific text-to-video generation method disclosed in the present application;
[0025] Figure 11 It is a block diagram of a specific text-to-video generation method disclosed in the present application. Detailed implementation manners
[0026] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0027] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0028] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] The embodiments of the present application disclose a method for generating text-to-video. As shown in Figure 1 the method includes:
[0030] Step S11: Input the target text description and the first noise vector into the first text-to-video model to generate a video matching the target text description and the corresponding first video latent vector; the first text-to-video model is a model obtained by training a preset student model using a training set according to a preset distillation strategy, and the preset distillation strategy is to denoise the historical high-resolution videos in the training set using the preset student model and a first diffusion model based on the attention mechanism in sequence, and update the model parameters of the preset student model based on the loss corresponding to the denoising result.
[0031] In this embodiment, first, the text description to be used for text-to-video generation input by the current user (such as yellow and black tropical fish rapidly swimming through the sea) is obtained to get the target text description. Then, the target text description input by the user and the first noise vector are input into the first text-to-video model obtained by training a preset student model using a training set according to a preset distillation strategy, so as to generate a video matching the target text description and the corresponding first video latent vector through the first text-to-video model. Among them, the first noise vector can be a Gaussian noise vector.
[0032] In this embodiment, considering the correlation between the student model and the teacher model in the traditional distillation strategy, absolute decoupling cannot be achieved. For example, the initial weights and model structures of the student model and the teacher model need to be kept consistent, and the student model must load the initial weights of the teacher model. In this way, the model can only be updated by comparing the forward denoising vectors at similar time steps, resulting in a lack of robustness in the model obtained after distillation training. On the other hand, considering that the traditional distillation strategy generates a better student model through equivalent denoising and repeated distillation processes, this reduces the distillation efficiency and leads to a low generation quality of the student model that meets the inference efficiency. The synthesized samples look blurred and have obvious artificial traces, and this problem is more obvious with fewer sampling steps.
[0033] To address the above problems, this application optimizes the traditional distillation strategy. The optimized distillation strategy first uses a preset student model and a first diffusion model based on an attention mechanism (such as the self-attention mechanism, Self-Attention) to perform denoising processing on the historical high-resolution videos in the training set in sequence, and then updates the model parameters (such as model weights, etc.) of the preset student model based on the loss corresponding to the denoising result.
[0034] Specifically, using a preset student model and a first diffusion model based on an attention mechanism to perform denoising on the historical high-resolution videos in the training set in sequence and updating the model parameters of the preset student model based on the loss corresponding to the denoising result may include: adding noise to the historical high-resolution videos in the training set to obtain a first noise-added video; inputting the first noise-added video into the preset student model for denoising to obtain a first denoised result; the preset student model is a diffusion model with a different network structure from the first diffusion model; inputting the first denoised result and the historical high-resolution video into a preset discriminator for discrimination, and calculating the discrimination loss between the first denoised result and the historical high-resolution video; adding noise to the first denoised result to obtain a second noise-added video; using the first diffusion model as the teacher model and inputting the second noise-added video into the teacher model for denoising to obtain a second denoised result; calculating the distillation loss between the first denoised result and the second denoised result, and calculating the sum value of the distillation loss and the discrimination loss to obtain a target loss; guiding the training of the preset student model based on the target loss using the gradient descent method; the preset student model and the teacher model have different sampling step ranges. In this embodiment, see Figure 2As shown in the figure, the specific training process of the first text-to-video model obtained after training according to the preset distillation strategy is divided into two stages, namely the noise addition stage 1 and the noise addition stage 2. The specific execution process is as follows: First, collect a training set containing historical text descriptions and corresponding historical high-resolution videos, then perform noise addition processing on the historical high-resolution videos in the training set, and input the first noise-added video obtained after noise addition into the student model for the first denoising process to obtain the corresponding first denoised result; then, input the first denoised result and the historical high-resolution video into the discriminator for discrimination, and calculate the discrimination loss between the two; further, execute the noise addition stage 2, that is, perform noise addition processing on the first denoised result output by the student model to obtain the second noise-added video, and input the second noise-added video into the teacher model (i.e., the first diffusion model based on the attention mechanism) for the second denoising process to obtain the corresponding second denoised result; it should be noted that during the entire distillation process, the model parameters of the teacher model (i.e., the first diffusion model) are frozen, that is, the model parameters remain unchanged and are not updated. Finally, calculate the distillation loss between the two denoised results (i.e., the first denoised result and the second denoised result), and calculate the sum of the distillation loss and the discrimination loss to obtain a target loss, and then based on this target loss and use the gradient descent method (Gradient descent) to guide the training of the student model, that is, continuously optimize and update the model parameters. And, during the entire distillation process, the model weights of the discriminator also need to be trained.
[0035] In addition, the present application decouples the teacher model and the student model. The student model (a diffusion model of a preset type) and the teacher model (i.e., the first diffusion model) have different network structures, and the sampling step ranges of the student model and the teacher model are also different. For example, the network structure of the teacher model adopts the network structure of the Stable-Diffusion-Video version, while the student model adopts other types of network structures (no specific requirements here, and the network weights of the student model can be randomly initialized or the initial weights of any network can be loaded). And, the sampling step range of the student model is {250, 500, 750, 1000}, and the sampling step range of the teacher model is {1, …, 1000}.
[0036] It can be seen that the optimized distillation strategy performs denoising twice. Among them, the student model undergoes the first denoising, and the discriminative loss is obtained by discriminating the result of the first denoising with the real video. Then, the first denoised video output by the student model undergoes the second denoising through the teacher model to obtain the result of the second denoising. Based on the discriminative loss and the distillation loss between the two denoising results, the parameters of the student model are updated. Since the result obtained after the teacher model denoises has gone through two denoising processes, the obtained result is very accurate. At this time, calculating the distillation loss can better update the parameters of the student model, thereby obtaining a text-to-video model with high inference efficiency, improving the inference efficiency of text-to-video, and thus being well applied to the scenario of generating high-resolution video content. In addition, the optimized distillation strategy adopts an end-to-end noise-added distillation method, and a qualified student model can be obtained with only one distillation, thereby greatly improving the distillation efficiency, enabling the weights of the distilled model to complete accelerated inference, and obtaining inference results of the same quality with only one step. Moreover, through the above two-stage noise-added distillation process, the entire text-to-video inference process can obtain a high-resolution text-to-video with only two denoising steps, thereby greatly improving the inference efficiency.
[0037] In a specific embodiment, the result of the first denoising and the historical high-resolution video are input into a preset discriminator for discrimination, and the discriminative loss between the result of the first denoising and the historical high-resolution video is calculated. Specifically, it may include: inputting the result of the first denoising and the historical high-resolution video into a preset discriminator for discrimination; the preset discriminator includes multiple discriminative sub-modules; calculating the loss between the result of the first denoising and the historical high-resolution video through each discriminative sub-module in the preset discriminator to obtain multiple discriminative losses; calculating the sum of the multiple discriminative losses to obtain the adversarial loss; calculating the sum value of the distillation loss and the adversarial loss to obtain the target loss. In this embodiment, as shown in Figure 3 As shown, the network architecture of the discriminator includes 4 discriminative sub-modules, and each discriminative sub-module corresponds to a feature extraction network module of a stage. The structure of each discriminative sub-module can be seen in Figure 4 As shown, it includes a convolutional layer (Conv layer), a batch normalization layer (BatchNorm layer), and an activation function (such as Leaky ReLU). Through each discriminative sub-module in the preset discriminator, the loss between the result of the first denoising and the historical high-resolution video can be calculated to obtain multiple discriminative losses. Then, all the discriminative losses are superimposed to obtain an adversarial loss, and the sum value of the distillation loss and the adversarial loss is calculated to obtain the target loss for model training.
[0038] In a specific embodiment, the historical high-resolution videos in the training set can be first subjected to noise addition processing to obtain the first noise-added video , the specific calculation formula is:
[0039] ;
[0040] In the formula, s represents the diffusion steps of the student model, which is within the sampling step range {250, 500, 750, 1000} of the student model, and are hyperparameters for the diffusion steps s, is the standard Gaussian noise.
[0041] Next, input the first noise-added video into the student model for denoising to obtain the first denoised video , and calculate the discrimination loss between the first denoised video and the historical high-resolution video through the discriminator; among them, represents the model parameters of the student model. It should be noted that the first denoised video is a completely denoised video, rather than the denoised video in the traditional distillation strategy.
[0042] Furthermore, add noise to the first denoised video to obtain the second noise-added video , and the specific calculation formula is:
[0043] ;
[0044] In the formula, t represents the diffusion steps of the teacher model, which is within the sampling step range {1,…,1000} of the teacher model, and are hyperparameters for the diffusion steps t, is the standard Gaussian noise.
[0045] Next, input the second noise-added video into the teacher model for the second denoising operation to obtain the second denoised video ; represents the model parameters of the teacher model. It should be noted that the second denoised video is a completely denoised video, rather than the denoised video in the traditional distillation strategy.
[0046] Finally, calculate the distillation loss based on the denoised videos and obtained before and after, add the distillation loss and the discrimination loss, and complete the training of the student model based on the obtained target loss using the gradient descent method.
[0047] In a specific embodiment, when calculating the adversarial loss the multi-stage ensemble learning method can be adopted and the Hinge Loss function is used to calculate the losses of the features of the first denoised result and the historical high-resolution video corresponding to each stage of the discriminator sub-module respectively, so as to obtain multiple discriminant losses, and the sum of all discriminant losses is calculated. The specific calculation formula is:
[0048] ;
[0049] In the formula, k represents the number of discriminator sub-modules, represents the discriminator sub-module; represents the video feature extractor. Further, as shown in Figure 3 after passing through the serialized and four-stage feature extraction network module, the discriminator discriminates the output features of each feature extraction network module. By adopting the multi-stage ensemble learning method to calculate the discriminant loss, the accuracy of the discriminant loss calculation can be further improved.
[0050] Specifically, calculating the distillation loss between the first denoised result and the second denoised result may include: calculating the distillation loss between the first denoised result and the second denoised result based on the first denoised result, the second denoised result and the target hyperparameter and using the mean squared error loss function; wherein, the target hyperparameter is the hyperparameter for the diffusion steps, and the diffusion steps are within the sampling step range of the preset student model. In this embodiment, the mean squared error loss function (Mean Squared Error) can be used to calculate the distillation loss between the first denoised result and the second denoised result. The specific calculation formula is:
[0051] ;
[0052] In the formula, indicates that the higher the noise level, the smaller the proportion of the distillation loss; represents the mean squared error loss function.
[0053] Step S12: Upsample the first video latent vector, and splice the obtained first sampled vector and the second noise vector to obtain a first spliced vector.
[0054] In this embodiment, after generating the video and the corresponding first video latent vector that match the target text description, further, the first video latent vector is upsampled to obtain a first sampled vector, and then the first sampled vector and the second noise vector are spliced to obtain a first spliced vector. Among them, the second noise vector can be a Gaussian noise vector.
[0055] Step S13: Input the first spliced vector into the second text-to-video model to generate a target text-to-video that matches the target text description; the second text-to-video model is a model obtained by training a second diffusion model based on a multi-way state space model using a training set.
[0056] In this embodiment, after splicing the first sampled vector and the second noise vector to obtain the first spliced vector, the first spliced vector can be further input into the second text-to-video model obtained by training a second diffusion model based on a multi-way state space model using a training set, so as to generate a text-to-video that matches the target text description and obtain a target text-to-video with high resolution.
[0057] Specifically, the process of obtaining the second text-to-video model can specifically include: collecting historical text descriptions and corresponding historical high-resolution videos, and performing downsampling on the historical high-resolution videos to obtain historical low-resolution videos; respectively compressing the historical high-resolution videos and the historical low-resolution videos from the pixel space to the latent vector space through a video autoencoder to obtain compressed high-resolution latent vectors and compressed low-resolution latent vectors; splicing the compressed low-resolution latent vectors with a third noise vector to obtain a second spliced vector, and inputting the second spliced vector and the historical text description into the second diffusion model based on a multi-way state space model for model training to obtain the second text-to-video model. In this embodiment, first collect historical text descriptions and corresponding historical high-resolution videos (i.e., high-resolution video sequences) to obtain a training set, then perform downsampling on the historical high-resolution videos in the training set to obtain corresponding historical low-resolution videos; then, through a video autoencoder, such as VideoMAE (Video Masked Autoencoder), respectively compress the historical high-resolution videos and the historical low-resolution videos from the pixel space to the latent vector space to obtain corresponding compressed high-resolution latent vectors and compressed low-resolution latent vectors, and then splice the compressed low-resolution latent vectors with a preset noise vector to obtain a spliced vector, and input the spliced vector and the historical text descriptions in the training set into the diffusion model based on a multi-way state space model for model training, so as to obtain the second text-to-video model.
[0058] Specifically, splicing the compressed low-resolution latent vectors with a third noise vector to obtain a second spliced vector can include: performing upsampling on the compressed low-resolution latent vectors to generate a vector with the same spatial resolution size as the compressed high-resolution latent vectors to obtain a second sampled vector; splicing the second sampled vector with the third noise vector to obtain a second spliced vector; the third noise vector is a vector obtained by sampling from a Gaussian distribution. See Figure 5 As shown, for the high-resolution video sequence (with a dimension of ) Perform downsampling to obtain the corresponding low-resolution video sequence (with dimensions ); Next, compress the high-resolution video sequence and the low-resolution video sequence from the pixel space to the latent vector space through a video autoencoder. Here, the video autoencoder has both spatial resolution compression (from 256 to 32) and temporal dimension compression (from 16 to 2), that is, the compression ratio is 8 in both cases, obtaining the corresponding high-resolution latent vectors (with dimensions ) and low-resolution latent vectors (with dimensions ). The specific compression ratio can be set according to actual needs; Then, upsample the low-resolution latent vectors to generate vectors with the same spatial resolution size as the high-resolution latent vectors, and then concatenate the sampled vectors with noise vectors of the same size obtained through Gaussian distribution sampling to obtain the concatenated vectors; Among them, the features of the concatenated vectors and the high-resolution latent vectors are the concatenated vector features and the high-resolution latent vector features . Finally, input the concatenated vector features and the historical text description (such as at night, a spectacular tornado) into the second diffusion model based on the multi-channel state space model for training.
[0059] See Figure 5 As shown, the training process of the entire second text-to-video model can complete the end-to-end training of high-resolution text-to-video, thus avoiding the cascaded training method of traditional high-resolution text-to-video models and greatly reducing the computing power consumption.
[0060] It should be noted that during the training process of the second text-to-video model, it specifically also includes: obtaining the time component of the second text-to-video model; converting the compressed high-resolution latent vectors and the second concatenated vectors into vector features to obtain the target high-resolution latent space vector features and the concatenated vector features; based on the time component, the target high-resolution latent space vector features and the concatenated vector features, and using a preset loss function to calculate the loss of the second text-to-video model to obtain the diffusion loss. That is, first convert the compressed high-resolution latent vectors and the second concatenated vectors into vector features to obtain the corresponding target high-resolution latent space vector features (i.e., Figure 5 in ) and the concatenated vector features (i.e., Figure 5 in ); Then, based on the time component of the second text-to-video model, the target high-resolution latent space vector features (i.e., ) and the concatenated vector features (i.e., , with dimensions ), and using a preset loss function to calculate the loss of the second text-to-video model to obtain the diffusion loss.
[0061] Specifically, based on the time component, the target high-resolution latent space vector feature, and the concatenated vector feature, and using a preset loss function to calculate the loss of the second text-to-video model, the diffusion loss can include: performing a serialization operation on the concatenated vector feature to obtain a serialized feature; encoding the time component and the historical text description respectively to obtain a time step feature and a text feature, and adding the time step feature and the text feature to obtain a conditional feature; inputting the conditional feature and the serialized feature into a sequence processing unit including a multi-way state space model for state space processing to obtain a first output feature; the sequence processing unit includes a scale layer, a scale translation layer, and a multi-layer perceptron; the multi-layer perceptron is used to perform feature transformation on the conditional feature, the scale layer is used to perform scale transformation on the feature output by the multi-layer perceptron, and the scale translation layer is used to perform translation transformation on the feature output by the multi-layer perceptron; performing a normalization operation on the first output feature to obtain a second output feature, and performing a deserialization operation on the second output feature to obtain a predicted noise mean vector; the predicted noise mean vector is consistent with the dimension of the concatenated vector feature; calculating the loss of the second text-to-video model based on the predicted noise mean vector, the concatenated vector feature, and the target high-resolution latent space vector feature and using the mean square error loss function to obtain the diffusion loss. In this embodiment, referring to Figure 6 as shown in (b), first perform a serialization operation (including convolution and vector flattening operations) on the concatenated vector feature (i.e., the latent vector in Figure 6 ), to obtain a serialized feature x (with a dimension of ); N is the batch size, which can be specifically set to 64; then, encode the time component (i.e., the time step vector in ) and the historical text description (i.e., the text vector in Figure 6 ) respectively to obtain a time step feature and a text feature, and add the time step feature and the text feature to obtain a conditional feature c (with a dimension of Figure 6 ), where ; further, input the conditional feature c and the serialized feature into a sequence processing unit including a multi-way state space model for state space processing, and then perform a layer normalization operation (such as using RMS Norm normalization operation, i.e., root mean square error standardization operation) and a deserialization operation on the output feature in sequence, so as to obtain a predicted noise mean vector and a predicted noise standard deviation ; where the predicted noise mean vector and the predicted noise standard deviation are both consistent with the dimension of the concatenated vector feature . Finally, based on the predicted noise mean vector, the concatenated vector feature and the target high-resolution latent space vector feature , and the mean squared error loss function is used to calculate the loss of the second text-to-video model to obtain the diffusion loss. The specific calculation formula is as follows:
[0062] ;
[0063] In the formula, ; represents the second text-to-video model, represents the model parameters of the second text-to-video model; represents the mean squared error loss function; is equivalent to the predicted noise mean vector . Calculate the diffusion loss according to the above formula, thereby completing the training of the second text-to-video model.
[0064] It should be noted that in order to improve the training and inference speed of this application, the multi-head attention mechanism in the DIT diffusion module (a diffusion model based on the transformer architecture) in the traditional attention mechanism-based diffusion model is changed to a sequence processing unit that can accelerate the processing of long sequences. See Figure 6 as shown in Figure 6 (a) is a traditional diffusion model based on the DIT diffusion module, Figure 6 (b) is the diffusion model optimized for the Figure 6 DIT diffusion module in (a), specifically replacing the DIT diffusion module with a sequence processing unit.
[0065] Specifically, see Figure 7 as shown. The sequence processing unit includes a scale layer, a scale translation layer, and a multi-layer perceptron (MLP, Multilayer Perceptron); among them, the multi-layer perceptron is used to perform feature transformation on the conditional feature c, the scale layer is used to perform scale transformation on the features output by the multi-layer perceptron, and input the features after scale transformation into 4 one-dimensional convolutions. The outputs of the 4 one-dimensional convolutions are respectively connected to each state space model in the multi-way state space model (the forward state space model, the backward state space model, the lateral state space model, and the backward lateral state space model); at the same time, the output of the scale layer will also pass through a fully connected layer and an activation layer; the scale translation layer is used to perform translation transformation on the features output by the multi-layer perceptron, and can also perform translation transformation on the accumulated output features after convolution of the state space model and the activation layer. Finally, the output of the sequence processing unit is obtained through a fully connected layer . Among them, the output of the sequence processing unit has the same dimension as the serialized feature x. Specifically, the output of the scale layer and the output of the scale translation layer are calculated as follows:
[0066] ;
[0067] ;
[0068] In the formula, ; and are the inputs of the scale layer and the scale translation layer respectively.
[0069] It can be understood that the state space model is a recursively implemented model, similar to the RNN (Recurrent Neural Network) network layer. By performing four-way scanning state processing on the input features, the two-dimensional attention module based on the transformer can be transformed into a one-dimensional scanning module, thereby reducing the time complexity by one dimension. See Figure 8 as shown. Figure 8 shows that, compared with the attention mechanism, the forward state space model, the backward state space model, the lateral state space model, and the back-lateral state space model in the multi-way state space model only need to perform scanning operations on the state spaces in 4 directions. This 4-way state space model can reduce the time complexity from to .
[0070] Of course, any combination of models can also be selected from the forward state space model, the backward state space model, the lateral state space model, and the back-lateral state space model according to actual application requirements, or other multiple state space models can be selected.
[0071] Specifically, see Figure 9 as shown. The network structures of each state space model are as Figure 9 shown. Assuming the input is composed of a series of token vectors and can be expressed as , where L = 256; the output after passing through the state space vector model can be expressed as that each token can be calculated by the following formula:
[0072] ;
[0073] In the formula, ; ; ; ; A represents the state parameter matrix; ; and both represent fully connected layers, represents the activation layer; represents the current state vector; Represents the previous state vector.
[0074] Specifically, the conditional feature and the serialized feature are input into a sequence processing unit including a multi-path state space model for state space processing to obtain a first output feature, which may include: performing position encoding on the serialized feature by means of absolute position encoding to obtain an encoded feature; inputting the conditional feature and the encoded feature into a sequence processing unit including a multi-path state space model for state space processing to obtain a first output feature; the first output feature is consistent with the serialized feature in dimension. In this embodiment, referring to Figure 6 As shown, considering that Figure 6 other modules in do not include convolution operations, a position encoding vector is added after the serialization operation. Specifically, the serialized feature can be position-encoded by means of absolute position encoding (the specific position encoding method is not limited and can be selected according to the actual situation), and then the conditional feature c and the encoded feature are input into a sequence processing unit including a multi-path state space model for state space processing; among them, the output feature of the sequence processing unit is consistent with the serialized feature in dimension.
[0075] Specifically, the time component and the historical text description are respectively encoded to obtain a time step feature and a text feature, which may include: inputting the time component into a time embedder for feature encoding to obtain a time step feature; inputting the historical text description into a text embedder for encoding to obtain a text feature. For example, a time embedder in the same way as DIT is used to convert the time component into a time step feature (with a dimension of ), and a text embedder of the T5-XXL text encoder type is used to obtain a text feature with a dimension of through a fully connected layer.
[0076] It can be seen that the embodiments of the present application optimize the traditional distillation strategy. The optimized distillation strategy undergoes two denoising processes. First, the student model performs the first denoising process on the historical high-resolution videos in the training set, and then the first denoising result is input into the first diffusion model based on the attention mechanism for the second denoising process, and the model parameters of the student model are updated based on the losses corresponding to the two denoising results, so as to obtain a first text-to-video model with high inference efficiency; moreover, the embodiments of the present application combine a second diffusion model based on a multi-path state space model on the basis of the diffusion model based on the attention mechanism. By performing re-inference on the text-to-video output based on the attention mechanism, the resolution of the text-to-video can be further improved; in addition, since the multi-path state space model can process only the features within a specified state space relative to the attention mechanism, the generation efficiency of the text-to-video is improved, thereby improving the user experience.
[0077] An embodiment of the present application discloses a specific method for generating a video from text. Refer to Figure 10 as shown, the method includes:
[0078] Step S21: Input the target text description and the first noise vector into the first text-to-video model to generate a video matching the target text description and a corresponding first video latent vector; the first text-to-video model is a model obtained by training a preset student model using a training set according to a preset distillation strategy, and the preset distillation strategy is to denoise historical high-resolution videos in the training set using the preset student model and a first diffusion model based on an attention mechanism in sequence, and update the model parameters of the preset student model based on the loss corresponding to the denoising result.
[0079] Step S22: Upsample the first video latent vector, and splice the obtained first sampled vector and the second noise vector to obtain a first spliced vector.
[0080] Step S23: Input the first spliced vector into the second text-to-video model to generate a video matching the target text description and a corresponding second video latent vector; the second text-to-video model is a model obtained by training a second diffusion model based on a multi-way state space model using a training set.
[0081] In this embodiment, refer to Figure 11 as shown, after obtaining the first spliced vector, it can be input into the second text-to-video model obtained by training a second diffusion model based on a multi-way state space model using a training set, so as to generate a video matching the target text description and a corresponding second video latent vector.
[0082] Step S24: Decode the second video latent vector through a video decoder to obtain the target text-to-video corresponding to the target text description.
[0083] In this embodiment, refer to Figure 11 as shown, the second video latent vector can be decoded through a video decoder, so as to obtain a high-resolution target text-to-video corresponding to the target text description.
[0084] Among them, for more specific processing procedures of the above steps S21 and S22, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0085] It can be seen that by decoding the second video latent vector through the video decoder in the embodiments of the present application, a high-resolution text-to-video can be obtained. In addition, based on the diffusion model with an attention mechanism, the present application combines the diffusion model based on the multi-way state space model, and the diffusion model with an attention mechanism is trained using an optimized distillation strategy, which only requires two denoising steps to obtain a high-resolution text-to-video, thus greatly improving the inference efficiency and further enhancing the user experience.
[0086] Embodiments of the present application also provide a text-to-video generation device. For the description of the features in the corresponding embodiments of this device, reference can be made to the relevant descriptions in the corresponding embodiments of the process recovery method, which will not be elaborated here one by one.
[0087] Embodiments of the present application also provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the text-to-video generation method.
[0088] Embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the text-to-video generation method when running.
[0089] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc, and other media that can store computer programs.
[0090] Embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the text-to-video generation method.
[0091] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the text-to-video generation method.
[0092] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0093] The above has introduced in detail a text-to-video generation method, product, device, and storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope protected by this application.
Claims
1. A method for generating a Vincent video, characterized in that: include: Input the target text description and the first noise vector into the first cultural video model to generate a video matching the target text description and a corresponding first video latent vector; the first cultural video model is a model obtained by training a preset student model using a training set and according to a preset distillation strategy, wherein the preset distillation strategy is to use the preset student model and a first diffusion model based on an attention mechanism to sequentially denoise the historical high-resolution videos in the training set, and update the model parameters of the preset student model based on the loss corresponding to the denoising result; Upsampling the first video latent vector to obtain a first sampled vector, and concatenating the first sampled vector and the second noise vector to obtain a first concatenated vector; Inputting the first concatenated vector into a second language video model to generate a target language video matching the target text description; The second video model is a model obtained by training a second diffusion model based on a multi-path state space model using the training set.
2. The method for generating a Vincent video according to claim 1, characterized in that: Before inputting the target text description and the first noise vector into the first text-generated video model, the method further includes: Collecting historical text descriptions and corresponding historical high-resolution videos, and downsampling the historical high-resolution videos to obtain historical low-resolution videos; Compressing the historical high-resolution video and the historical low-resolution video from pixel space to latent vector space respectively by using a video autoencoder to obtain a compressed high-resolution latent vector and a compressed low-resolution latent vector; The compressed low-resolution latent vector is spliced with the third noise vector to obtain a second spliced vector, and the second spliced vector and the historical text description are input into a second diffusion model based on a multi-channel state space model for model training to obtain the second cultural video model.
3. The method for generating a Vincent video according to claim 2, characterized in that: The step of concatenating the compressed low-resolution latent vector with the third noise vector to obtain a second concatenated vector comprises: Upsampling the compressed low-resolution latent vector to generate a vector having the same spatial resolution as the compressed high-resolution latent vector, to obtain a second sampled vector; The second sampled vector is concatenated with the third noise vector to obtain a second concatenated vector; the third noise vector is a vector obtained after Gaussian distribution sampling.
4. The method for generating a Vincent video according to claim 2, characterized in that: The training process of the second video model also includes: Obtaining a time component of the second video model; Converting the compressed high-resolution latent vector and the second concatenated vector into vector features to obtain target high-resolution latent space vector features and concatenated vector features; Based on the time component, the target high-resolution latent space vector feature and the spliced vector feature, a loss of the second video model is calculated using a preset loss function to obtain a diffusion loss.
5. The method for generating a Vincent video according to claim 4, characterized in that: The method of calculating the loss of the second video model based on the time component, the target high-resolution latent space vector feature and the spliced vector feature by using a preset loss function to obtain the diffusion loss includes: Performing a serialization operation on the concatenated vector features to obtain serialized features; Encoding the time component and the historical text description respectively to obtain a time step feature and a text feature, and adding the time step feature and the text feature to obtain a conditional feature; The conditional features and the serialized features are input into a sequence processing unit including a multi-channel state space model for state space processing to obtain a first output feature; the sequence processing unit includes a scaling layer, a scaling translation layer and a multi-layer perceptron; the multi-layer perceptron is used to perform feature transformation on the conditional features, the scaling layer is used to perform scaling transformation on the features output by the multi-layer perceptron, and the scaling translation layer is used to perform translation transformation on the features output by the multi-layer perceptron; Normalizing the first output feature to obtain a second output feature, and deserializing the second output feature to obtain a predicted noise mean vector; the predicted noise mean vector has the same dimension as the concatenated vector feature; The loss of the second video model is calculated based on the predicted noise mean vector, the concatenated vector features and the target high-resolution latent space vector features using a mean square error loss function to obtain a diffusion loss.
6. The method for generating a Vincent video according to claim 5, characterized in that: The encoding of the time component and the historical text description respectively to obtain a time step feature and a text feature includes: Inputting the time component into a time embedder for feature encoding to obtain a time step feature; The historical text description is input into a text embedder for encoding to obtain text features.
7. The method for generating a Vincent video according to claim 5, characterized in that: The step of inputting the conditional features and the serialized features into a sequence processing unit including a multi-path state space model for state space processing to obtain a first output feature comprises: Position encoding is performed on the serialized features by means of absolute position encoding to obtain encoded features; The conditional features and the encoded features are input into a sequence processing unit including a multi-way state space model for state space processing to obtain a first output feature; the first output feature has the same dimension as the serialized feature.
8. The method for generating a Vincent video according to claim 2, characterized in that: The multi-path state space model includes any multiple models among a forward state space model, a backward state space model, a lateral state space model and a rear-lateral state space model.
9. The method for generating a Vincent video according to claim 1, characterized in that: The method of sequentially denoising the historical high-resolution videos in the training set by using the preset student model and the first diffusion model based on the attention mechanism, and updating the model parameters of the preset student model based on the loss corresponding to the denoising result, includes: Performing noise processing on the historical high-resolution video in the training set to obtain a first noisy video; Inputting the first noisy video into a preset student model for denoising to obtain a first denoising result; the preset student model is a diffusion model with a different network structure from the first diffusion model; Inputting the first denoised result and the historical high-resolution video into a preset discriminator for discrimination, and calculating the discrimination loss between the first denoised result and the historical high-resolution video; Performing noise processing on the first denoising result to obtain a second noisy video; Using the first diffusion model as a teacher model, and inputting the second noisy video into the teacher model for denoising, to obtain a second denoised result; Calculating a distillation loss between the first denoised result and the second denoised result, and calculating a sum of the distillation loss and the discrimination loss to obtain a target loss; The training of the preset student model is guided based on the target loss and using the gradient descent method; the preset student model and the teacher model have different sampling step ranges.
10. The method for generating a Vincent video according to claim 9, characterized in that: The step of inputting the first denoised result and the historical high-resolution video into a preset discriminator for discrimination, and calculating the discrimination loss between the first denoised result and the historical high-resolution video, comprises: Inputting the first denoising result and the historical high-resolution video into a preset discriminator for discrimination; the preset discriminator includes a plurality of discrimination submodules; Calculating the loss between the first denoising result and the historical high-resolution video through each of the discriminant submodules in the preset discriminator to obtain a plurality of discriminant losses; Calculate the sum of the multiple discriminant losses to obtain the adversarial loss; Accordingly, the calculation of the sum of the distillation loss and the discrimination loss to obtain the target loss includes: The sum of the distillation loss and the adversarial loss is calculated to obtain the target loss.
11. The method for generating a Vincent video according to claim 9, characterized in that: The calculating the distillation loss between the first denoised result and the second denoised result includes: Calculate the distillation loss between the first denoised result and the second denoised result based on the first denoised result, the second denoised result and the target hyperparameter and using a mean square error loss function; The target hyperparameter is a hyperparameter for the number of diffusion steps, and the number of diffusion steps is within the range of sampling steps of the preset student model.
12. The method for generating a Vincent video according to any one of claims 1 to 11, characterized in that: The step of inputting the first concatenated vector into a second language video model to generate a target language video matching the target text description includes: Inputting the first concatenated vector into a second text-generated video model to generate a video matching the target text description and a corresponding second video latent vector; The second video latent vector is decoded by a video decoder to obtain a target text video corresponding to the target text description.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the Vincent video generation method as claimed in any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the Vincent video generation method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the Vincent video generation method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Industrial defect image simulation method and device based on diffusion model
CN117649351A
Two-stage pre-training text-video retrieval method for realizing concept-level semantic alignment based on shared concept anchor points
CN118364284A