Method for accelerating video diffusion model by using synthetic data set

By constructing a synthetic data set and designing trajectory-based loss function and adversarial training strategy, students' models are subjected to knowledge distillation training, which solves the problem of inefficient video diffusion model generation and achieves efficient and high-speed video generation.

CN120297362AActive Publication Date: 2025-07-11SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Patent Information

Application Number
CN202510355770.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing video diffusion model requires multiple inference steps and a large number of computing resources when generating videos, and the distillation technology is not effective in the video diffusion model, resulting in inefficient generation.

Method used

By constructing a synthetic dataset, including synthetic video and denoising trajectories, using the pre-trained video diffusion model as the teacher model, designing trajectory-based loss functions and adversarial training strategies, and conducting knowledge distillation training on student models, reducing inference steps and improving generation quality.

Benefits of technology

It significantly accelerates the video generation process, reduces the number of inference steps, improves the quality and resolution of generated videos, and saves computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297362A_ABST
    Figure CN120297362A_ABST
Patent Text Reader

Abstract

The invention discloses a method for accelerating a video diffusion model by using a synthetic data set. The method comprises the following steps: generating a synthetic data set by using a pre-trained video diffusion model, wherein the synthetic data set comprises a synthetic video, a denoising track in a potential space and a corresponding text prompt; the pre-trained video diffusion model is used as a teacher model, a corresponding student model is constructed, and the student model and the teacher model share the same structure; based on the synthetic data set, knowledge distillation training is carried out on the student model, and in the knowledge distillation training process, the student model learns the denoising process of the teacher model and aligns data distribution generated by the teacher model until a set loss function standard is met; and taking the student model subjected to knowledge distillation training as a video generation model, and applying the video generation model to a video analysis task. According to the invention, videos with higher quality and higher resolution can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a method for accelerating video diffusion models using synthetic datasets. Background Art

[0002] Video diffusion models are widely used in video analysis tasks, including video generation, video editing, and various forms of video understanding tasks. In recent years, video diffusion models have achieved remarkable success. For example, VDM (Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. In NIPS, pages 8633–8646, 2022) extended the 2D U-Net used in image diffusion models to 3D U-Net and proposed a joint training method for images and videos. Another example is that some solutions adopt the latent space diffusion model (LDM) to learn the video data distribution in the latent space, significantly reducing the computational complexity. To further improve the resolution of the generated videos, some solutions adopt cascaded video diffusion models, which decompose the video generation process into multiple subtasks, such as key frame generation, frame interpolation, and super-resolution. With Sora (OpenAI. Sora, 2024. https: / / openai.com / sora) demonstrating powerful video generation capabilities, the diffusion transformer architecture (DiT) has gradually become the mainstream backbone network for video diffusion models.

[0003] In the prior art, video diffusion models are mainly accelerated from three aspects: efficient model architectures, high-compression variational autoencoders (VAEs), and distillation techniques. For example, LinGen adopts the Mamba2 module with linear complexity, significantly reducing the training computational cost. LTX-Video achieves a high compression ratio using a carefully designed Video-VAE, reducing the input dimension of the diffusion model and thus improving the generation speed. AnimateDiff-Lightning extends progressive adversarial distillation to video diffusion models. T2V-Turbo and T2V-Turbo-V2 use an improved consistency distillation loss to reduce the inference steps and improve the video quality through a reward model. CausVid and APT respectively use distribution matching loss and adversarial loss to distill pre-trained video diffusion models.

[0004] Upon analysis, existing video diffusion models need to iteratively denoise Gaussian noise to generate the final video, which usually requires dozens of inference steps. This process is both time-consuming and computationally resource-intensive. For example, HunyuanVideo (the Hunyuan video generator) takes 3234 seconds to generate a 5-second video with a resolution of 720×1280 and 24 frames per second on a single NVIDIA A100 GPU. Although progress has been made in accelerating image diffusion models through distillation techniques, these methods are difficult to directly apply to video diffusion models because they may use useless data points during distillation, resulting in a large amount of data and computational resources. Moreover, the useless data points can cause the teacher model to provide unreliable guidance to the student model, thus having an adverse impact on video quality. The spatio-temporal complexity of videos further exacerbates these problems. Summary of the Invention

[0005] The object of the present invention is to overcome the defects of the above-mentioned prior art and provide a method for accelerating video diffusion models using a synthetic dataset. The method includes the following steps:

[0006] Generate a synthetic dataset using a pre-trained video diffusion model, where the synthetic dataset includes synthetic videos, denoising trajectories in the latent space, and corresponding text prompts;

[0007] Use the pre-trained video diffusion model as the teacher model and construct a corresponding student model, where the student model and the teacher model share the same structure;

[0008] Based on the synthetic dataset, perform knowledge distillation training on the student model. During the knowledge distillation training process, the student model learns the denoising process of the teacher model and aligns the data distribution generated by the teacher model until the set loss function criterion is met;

[0009] Use the student model trained by distillation as a video generation model and apply it to video analysis tasks.

[0010] Compared with the prior art, the advantages of the present invention are that a novel and efficient distillation method is proposed, which reduces the inference steps by using a synthetic dataset, thereby accelerating the video diffusion model. For example, multiple effective denoising trajectories are generated using a pre-trained video diffusion model as the synthetic dataset, thus avoiding the use of useless data points during distillation. Based on this synthetic dataset, a trajectory-based loss function is designed to learn the mapping from Gaussian noise to video using key data points in the denoising trajectory, thereby generating videos with fewer inference steps. In addition, since the synthetic dataset captures the data distribution at each diffusion time step, an adversarial training strategy is introduced to align the output distribution of the student model with the distribution of the synthetic dataset, improving the quality of the generated videos.

[0011] Other features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0013] Figure 1 is a flowchart of a method for accelerating a video diffusion model using a synthetic dataset according to an embodiment of the present invention;

[0014] Figure 2 is an architecture diagram of a video diffusion model accelerated using a synthetic dataset according to an embodiment of the present invention;

[0015] Figure 3 is a schematic diagram of the process of constructing a synthetic video dataset according to an embodiment of the present invention;

[0016] Figure 4 is a schematic diagram of the layer-by-layer behavior of a feature extractor according to an embodiment of the present invention;

[0017] Figure 5 is a schematic diagram of the video generation result according to an embodiment of the present invention. DETAILED DESCRIPTION

[0018] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.

[0019] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present invention, its application, or uses.

[0020] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be considered as part of the specification.

[0021] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of exemplary embodiments may have different values.

[0022] It should be noted that: Like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.

[0023] Overall, in order to avoid using useless data points during the distillation process, the present invention proposes a novel and efficient video diffusion model distillation method, which can accelerate the video diffusion model through a synthetic dataset. Specifically, first, a synthetic video dataset (or called SynVid) is constructed, for example, containing 110,000 denoising trajectories, high-quality videos, and their corresponding fine-grained text descriptions. The data points on these denoising trajectories are intermediate results leading to the correct output, so they are all valid and meaningful. Then, a trajectory-based loss function is designed to select a small number of data points from each denoising trajectory to construct a shorter mapping path from noise to video, enabling the student model to generate videos with fewer inference steps. Further, in order to utilize the data distribution captured by the synthetic dataset, an adversarial training strategy is proposed to align the output distribution of the student model with the distribution of the synthetic dataset at each diffusion time step, thereby improving the quality of the generated videos.

[0024] Combined Figure 1 and Figure 2 As shown, the method for accelerating the video diffusion model using the synthetic dataset includes the following steps:

[0025] Step S110, using a pre-trained video diffusion model to construct a synthetic dataset, including synthetic videos, denoising trajectories in the latent space, and corresponding text prompts.

[0026] Multiple types of text-to-video diffusion models can be used as synthesizers to construct the synthetic dataset. As Figure 3 shown, HunyuanVideo is used as the generator to generate synthetic data. HunyuanVideo θ is a text-to-video diffusion model based on the DiT architecture, which is trained in the latent space and uses the flow matching loss for training. For example, the official code and settings can be used to generate the synthetic dataset.

[0027] Specifically, the synthetic dataset SynVid contains high-quality synthetic videos V i , denoising trajectories in the latent space , and their corresponding text prompts T i , where N = 110K represents the number of data, n = 50 represents the number of inference steps, and {t j |j ∈ [0, n]} represents the inference diffusion time steps. The text prompts T i are obtained by annotating high-quality videos on the Internet using a multimodal large model (such as InternVL2.5). For Gaussian noise , the denoising trajectories can be solved in the following way:

[0028]

[0029] Among them, i represents the data index in the dataset, n represents the number of inference steps, and j represents the time step index. t j-1 represents the (j - 1)-th diffusion time step.

[0030] The synthesized video V i can be obtained by decoding the clean latent code through a VAE (Variational Autoencoder). For the sake of simplicity, in the following description, the text prompt input T is omitted. It should be noted that in the distillation process of the present invention, only the synthesized dataset i is used. The data points in are intermediate results leading to the correct video, so they are all valid and meaningful. This helps with efficient distillation and significantly reduces the requirement for the amount of data.

[0031] In summary, the synthesized video dataset not only contains a large number of high-quality synthesized videos and their corresponding fine-grained text descriptions, but also contains complete denoising trajectories, which can promote the development of the video generation field.

[0032] Step S120: Use the pre-trained video diffusion model as the teacher model, construct the corresponding student model, and then perform knowledge distillation training on the student model based on the constructed synthesized dataset to learn the denoising process of the teacher model and align the data distribution generated by the teacher model.

[0033] Video diffusion models usually require a large number of inference steps to generate videos, which is a time-consuming process and requires a large amount of computing power. To accelerate the generation process, a student model S β is designed. This student model learns from the denoising trajectories generated by the pre-trained video diffusion model (i.e., the teacher model) and can generate videos with fewer inference steps. The student model shares the same architecture as the teacher model and is parameterized by β, and β is initialized using the parameters θ of the teacher model.

[0034] In one embodiment, a trajectory-based loss function is proposed to learn the denoising process of the teacher model in m steps. This loss selects m + 1 key diffusion time steps {t′ m = 1 > … > t0 = 0}, and obtains their corresponding latent codes on each denoising trajectory and then learns the mapping relationship from Gaussian noise to video through these key latent codes. For example, the trajectory-based loss function is expressed as:

[0035]

[0036] where k ∈ [0, m - 1], k represents the index of the k-th key time step, t′ mDenote the m-th critical diffusion time step, \(t'\) k+1 Denote the \((k + 1)\)-th critical diffusion time step. These critical latent encodings construct a shorter path from Gaussian noise to video latent encoding. By learning this shorter path, the student model significantly reduces the number of inference steps. For example, when \(m = 5\), the number of inference steps is reduced by 10 times compared with the teacher model, thus significantly accelerating the video generation process.

[0037] In addition to the one-to-one mapping between Gaussian noise and video latent encoding, the synthetic dataset also contains many latent encodings for each diffusion time step \(t'\) k of the data distribution These latent encodings implicitly represent the data distribution at this diffusion time step To fully release the knowledge in the synthetic dataset and improve the performance of the student model \(s\) β a adversarial training strategy is proposed to minimize the adversarial divergence where denotes the data distribution generated by the student model \(s\) β at the diffusion time step \(t'\) k The training objective follows the standard generative adversarial network. For example, the adversarial loss is set as:

[0038]

[0039] where denotes the distribution generated at the diffusion time step \(t'\) k belongs to D(.) represents the feature output by the discriminator, and \(D\) represents the discriminator parameterized by the parameter \(\varphi\). Since the discriminator needs to distinguish the noise latent encodings at different diffusion time steps, the designed discriminator \(D\) contains a noise-aware feature extractor and \(m\) diffusion-time-step-aware projection heads \(\{H_0,\cdots,H\) \( m-1 \}\),

[0040]

[0041] where \(H\) k denotes the \(k\)-th projection head, and \(v\) θ (.) represents the feature output by the feature extractor

[0042] Here, the frozen teacher model \(v\) θ can be used as the feature extractor. Figure 4 Shows the layer-by-layer behavior of \(v\) θ at different diffusion time steps. When \(t' k > 0\), the deeper network focuses on capturing high-frequency details, so the output features of the last layer (i.e., the 60th layer) can be selected for discrimination. When \(t' kWhen = 0, the output of the 40th layer is selected for discrimination. This layer contains fine-grained clues, such as the sticky notes on the table. Each projection head H k is a 2D convolutional neural network for outputting latent encodings and the source of (whether it is synthetic data). These projection heads have the ability to perceive diffusion time steps and can effectively process noisy features at different diffusion time steps, thus promoting adversarial learning.

[0043] For synthetic latent encodings s can be used β to solve iteratively, similar to formula (1). However, iterative solution takes a lot of time and is not feasible in actual training. Therefore, m + 1 queues {Q0,..., Q m} can be used, and these queues are used to maintain the latent encodings generated by s β at each diffusion time step. Through this design, it is possible to avoid denoising from Gaussian noise in each training step, thus saving training overhead.

[0044] It should be noted that the adversarial training strategy avoids the forward diffusion operation commonly used in existing methods, thus preventing the distillation of useless data points. This ensures that the teacher model can provide accurate guidance for the student model. In addition, the trajectory-based loss function improves the stability of adversarial learning without the need for complex regularization design.

[0045] The training parameters include β and the parameters of the projection head that perceives diffusion time steps. These parameters are trained through and The following Algorithm 1 shows the distillation process.

[0046] ***********************************************

[0047]

[0048] ***********************************************

[0049] In summary, the trajectory-based loss function can greatly accelerate the video generation process of the video diffusion model, reducing the requirement for computing power. And by avoiding using useless data points in the distillation process, it can reduce the training time. In addition, the adversarial training strategy can significantly improve the video generation performance of the student model by fully utilizing the data distribution knowledge captured in the synthetic dataset.

[0050] It should be understood that the pre-trained video diffusion model selects model architectures of various types or scales as needed. For example, a stronger pre-trained video diffusion model is used as the teacher model.

[0051] Step S130: Use the student model trained by knowledge distillation as a video generation model and apply it to video analysis tasks.

[0052] The student model trained by knowledge distillation has an efficient inference method. And because it has fully learned the denoising process of the teacher model and maintains the accuracy requirements, it can be used as a video generation model for actual video analysis tasks such as video generation, video editing, and various forms of video understanding tasks.

[0053] To further verify the effect of the present invention, experimental verification is carried out.

[0054] First, the inference time of the generated videos is compared. Table 1 shows the time required to generate videos measured by a single A100 GPU, including text encoding, VAE decoding, and diffusion time. The student model of the present invention only needs 5 steps of inference to generate videos. Compared with the teacher model, the generation speed is increased by 7.7 - 8.5 times. Even when the scale of existing models (such as CogVideoX-5B and CogVideoX1.5-5B) is only half of that of the model of the present invention, the model of the present invention is still 2.7 times faster. In addition, evaluations are also carried out on VBench, which is a comprehensive benchmark widely used to evaluate video quality. Specifically, VBench evaluates video generation models from 16 dimensions using 946 prompts. This benchmark is consistent with human perception. The quantitative results are shown in Tables 2 and 3. Except for the teacher model HunyuanVideo, the present invention exceeds all other comparative baseline methods at the same resolution level. Compared with the teacher model, the present invention achieves comparable performance at high resolutions and performs better at medium resolutions. Especially in terms of color and spatial relationships, excellent performance is achieved. The verification results show that the model of the present invention can effectively parse text prompts and generate videos consistent with the text. This not only proves the effectiveness but also highlights that the synthetic dataset contains high-quality text prompts and data. In addition, qualitative comparisons are also carried out, and the results are as Figure 5 shown. Compared with LTX-Video and PyramidFlow, the videos generated by the present invention have fewer artifacts. Compared with CogVideoX1.5, the videos generated by the present invention have higher fidelity and better backgrounds. Compared with HunyuanVideo, the present invention can better align with text prompts, such as vintage SUVs.

[0055] Table 1: Comparison of Inference Time for Generating Videos

[0056]

[0057] Table 2: Quantitative Results on VBench

[0058]

[0059] Table 3: Quantitative Results on VBench

[0060]

[0061] In summary, compared with the prior art, the present invention has the following advantages:

[0062] 1) The present invention proposes an efficient video diffusion model distillation method based on a synthetic video dataset. By using the teacher model to generate high-quality synthetic videos and denoising trajectories, a synthetic video dataset is constructed, which can avoid using useless data points during the distillation process, enabling the teacher model to provide accurate guidance for the student model and significantly reducing the requirements for data volume and training computing power.

[0063] 2) The present invention proposes a trajectory-based loss function to select key data points from the denoising trajectories and learn the mapping from Gaussian noise to video based on these data points, which can significantly reduce the inference steps for generating videos and thus accelerate the video generation process.

[0064] 3) The adversarial training strategy proposed by the present invention utilizes the data distribution at each diffusion time step captured by the synthetic dataset to further improve the quality of the videos generated by the student model and avoid the design of complex regularization.

[0065] 4) The present invention can generate videos of higher quality and resolution. A large number of experiments show that the model of the present invention is 8.5 times faster than the teacher model in terms of generation speed, while achieving comparable video generation effects. Compared with the prior art, the present invention can generate videos of higher quality and resolution, namely videos with a resolution of 720x1280, 24 frames per second, and a duration of 5 seconds.

[0066] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0067] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0068] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0069] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0070] Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of a method, apparatus (system), and computer program product according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0071] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data - processing apparatus, a device is created that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner. Thus, the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0072] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0073] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As is well known to those skilled in the art, implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.

[0074] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A method for accelerating video diffusion models using synthetic datasets, comprising the following steps: Generate a synthetic dataset using a pre-trained video diffusion model, the synthetic dataset including synthetic videos, denoising trajectories in the latent space, and corresponding text prompts; Use the pre-trained video diffusion model as a teacher model and construct a corresponding student model, the student model and the teacher model sharing the same structure; Based on the synthetic dataset, perform knowledge distillation training on the student model. During the knowledge distillation training process, the student model learns the denoising process of the teacher model and aligns the data distribution generated by the teacher model until a set loss function criterion is met; Use the student model trained by distillation as a video generation model and apply it to video analysis tasks.

2. The method according to claim 1, characterized in that The synthetic dataset is denoted as which contains the synthetic video V i , the denoising trajectory in the latent space and the corresponding text prompt T i , {t j |j ∈ [0, n]} represents the inference diffusion time steps, and for Gaussian noise the denoising trajectory is solved as follows: Among them, T i is a text prompt, and is Gaussian noise.

3. The method according to claim 1, wherein During the knowledge distillation training process, update the student model using a trajectory-based loss function, expressed as: where {t′ m = 1 > … > t0 = 0} are m + 1 diffusion time steps, are the latent encodings corresponding to the m + 1 critical diffusion time steps on each denoising trajectory, k ∈ [0, m - 1], S β denotes the student model, β are the parameters of the student model, t′ m denotes the m-th critical diffusion time step, t′ k+1 denotes the (k + 1)-th critical diffusion time step.

4. The method according to claim 1, characterized in that, During the knowledge distillation training process, the training objective follows a standard generative adversarial network, expressed as: Where: Among them, D represents the discriminator parameterized by the parameter φ. The discriminator D includes a noise-aware feature extractor and m diffusion time-step-aware projection heads {H0,..., H m-1}, v θ represents the teacher model as a noise-aware feature extractor, and θ is the parameter of the teacher model.

5. The method according to claim 1, characterized in that, During the distillation training process, initialize the parameters of the student model using the parameters of the teacher model.

6. The method according to claim 4, wherein During the knowledge distillation training process, for the synthetic latent codes Use m + 1 queues {Q0,..., Q m} to maintain the s β Latent codes generated at each diffusion time step.

7. The method according to claim 1, characterized in that, The text prompts are obtained by annotating videos on the Internet using a multimodal large model.

8. The method according to claim 1, wherein The synthetic videos are obtained by decoding latent codes through a variational autoencoder.

9. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, and characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image generation model compression and acceleration method and system based on diffusion model

    CN116542321A

  • Diffusion model distillation method based on cross-image pixel space relationship

    CN118587527A

  • User preference-based text video diffusion model training method and device

    CN119672696A

  • High-resolution video generation using image diffusion models

    US20240171788A1

Cited By

  • Video generation method and device, equipment, medium and program product

    CN121585879A

  • Diffusion model reasoning acceleration method based on optimal time step sequence search and knowledge distillation

    CN121809699A