Multi-stage flow matching diffusion model training method

By using a multi-stage flow matching diffusion model training method to gradually optimize the sampling trajectory, the problems of slow generation speed and insufficient model flexibility of the diffusion model are solved, and more efficient image generation is achieved.

CN120877015APending Publication Date: 2025-10-31SHANDONG ENERGY GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510138277.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing diffusion models are slow to generate and require a large number of iterations. Furthermore, existing optimization methods such as ReFlow and PeRFlow suffer from high optimization difficulty or insufficient model flexibility.

Method used

A multi-stage flow matching diffusion model training method is adopted. By gradually reducing the number of time windows and using a direction-guided loss function and corrective flow techniques, the sampling trajectory is optimized, and the model is gradually improved.

Benefits of technology

It improves the generation efficiency and image quality of the diffusion model, achieves more efficient few-step inference performance, and enhances the flexibility and applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_5
    Figure SMS_5
  • Figure QLYQS_1
    Figure QLYQS_1
Patent Text Reader

Abstract

The invention relates to a multi-stage flow matching diffusion model training method. The method comprises the following steps: S1, obtaining a training data set; s2, preprocessing the images in the training data set; s3, determining the inference step number of the pre-training diffusion model; s4, setting hyper-parameters; s5, dividing the sampling track into a plurality of time windows from a larger window number; and S6, gradually reducing the number of the time windows. The main innovation point of the method is that a multi-stage flow model training normal form is provided, and the learning target of the flow model is optimized. Through the multi-stage training method, the number of time windows is gradually reduced, and the learning difficulty of the flow matching diffusion model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and machine learning technology, and in particular relates to a multi-stage flow matching diffusion model training method. Background Technology

[0002] The diffusion model is a mainstream visual generative model that adds noise to an image through a forward diffusion process and then predicts the noise through a reverse process to generate a sharp image. However, traditional diffusion models are slow to generate images and require a large number of iterations, which limits their efficiency in practical applications.

[0003] To address the aforementioned issues, existing technologies have proposed two methods: ReFlow (corrective flow) and PeRFlow (segmented corrective flow). ReFlow directly applies flow matching techniques to optimize the pre-trained diffusion model, but optimization is difficult and performance is poor with few inference steps. PeRFlow reduces learning complexity by dividing the time window, but the fixed number of windows limits the model's flexibility and applicability. These limitations indicate that while existing technologies improve the sampling efficiency of diffusion models, further improvements are needed to meet the demands of practical applications. Summary of the Invention

[0004] (I) Purpose of the Invention

[0005] To overcome the above shortcomings, the purpose of this invention is to provide a multi-stage flow matching diffusion model training method to solve the above technical problems.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the technical solution provided in this application is as follows:

[0008] A multi-stage flow matching diffusion model training method includes the following steps:

[0009] S1 obtains the training dataset, which contains images and their corresponding text descriptions;

[0010] S2 preprocesses the images in the training dataset to make the image resolution meet the requirements of the pre-trained diffusion model;

[0011] S3 determines the number of inference steps for the pre-trained diffusion model to maintain the consistency of the sampling trajectory of the teacher model at different training stages;

[0012] S4 sets hyperparameters, including the number of windows and the hyperparameters of the loss function;

[0013] S5 starts with a large number of windows, divides the sampling trajectory into multiple time windows, and applies flow matching technology for training within each time window to obtain the initial polyline sampling trajectory;

[0014] S6 gradually reduces the number of time windows, uses the model obtained from the previous stage of training as initialization, and continues training with the new number of windows to gradually optimize the sampling trajectory until the final sampling trajectory is obtained.

[0015] Preferably, the flow matching technology is a corrected flow technology.

[0016] Preferably, a direction-guided loss function is used during the training process. This loss function includes mean squared error loss and cosine similarity loss, and the relative importance of the two is controlled by hyperparameters. The direction-guided loss function is expressed as follows:

[0017]

[0018] Where v represents the true value to be predicted. This represents the model's predicted value. α is a hyperparameter; the larger the value, the more emphasis is placed on directional alignment. When α is 0, this loss degenerates into the original MSE loss.

[0019] Preferably, the number of windows is set to "kk / 2-k / 4", that is, starting from k windows, gradually reducing to k / 2 windows, and then reducing to k / 4 windows.

[0020] Preferably, the training method is applied to a pre-trained diffusion model, wherein the diffusion model is StableDiffusion v1.5.

[0021] Preferably, the dataset used in the training process is the Laion-Art dataset or a subset thereof.

[0022] Preferably, the preprocessing includes adjusting the image resolution and applying center cropping.

[0023] A flow matching diffusion model trained based on the aforementioned multi-stage flow matching diffusion model training method.

[0024] A method for image generation using the aforementioned flow matching diffusion model includes the following steps:

[0025] Receive text description input from the user;

[0026] The flow matching diffusion model is used to generate the corresponding image based on the text description.

[0027] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, describes a multi-stage flow matching diffusion model training method.

[0028] Beneficial effects:

[0029] The main innovation of this invention lies in proposing a multi-stage flow model training paradigm and optimizing the learning objective of the flow model. This multi-stage training method gradually reduces the number of time windows, lowering the learning difficulty of the flow matching diffusion model and enabling it to learn the mapping relationship from noise to sharp images more efficiently. Simultaneously, by optimizing the learning objective, this invention further improves the performance of the flow matching diffusion model in low-step inference scenarios, significantly enhancing generation efficiency and image quality. These innovations collectively address the problems of high optimization difficulty and poor low-step inference performance in existing technologies, providing a more efficient and flexible solution for the practical application of diffusion models. Attached Figure Description

[0030] Figure 1 This illustrates the impact of hyperparameter settings on the generation effect in this embodiment of the invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of the invention.

[0032] This invention provides a multi-stage flow matching diffusion model training method, comprising the following steps:

[0033] S1 obtains the training dataset, which contains images and their corresponding text descriptions;

[0034] S2 preprocesses the images in the training dataset to make the image resolution meet the requirements of the pre-trained diffusion model;

[0035] S3 determines the number of inference steps for the pre-trained diffusion model to maintain the consistency of the sampling trajectory of the teacher model at different training stages;

[0036] S4 sets hyperparameters, including the number of windows and the hyperparameters of the loss function;

[0037] S5 starts with a large number of windows, divides the sampling trajectory into multiple time windows, and applies flow matching technology for training within each time window to obtain the initial polyline sampling trajectory;

[0038] S6 gradually reduces the number of time windows, uses the model obtained from the previous stage of training as initialization, and continues training with the new number of windows to gradually optimize the sampling trajectory until the final sampling trajectory is obtained.

[0039] Preferably, the flow matching technology is a corrected flow technology.

[0040] Preferably, a direction-guided loss function is used during the training process. This loss function includes mean squared error loss and cosine similarity loss, and the relative importance of the two is controlled by hyperparameters. The direction-guided loss function is expressed as follows:

[0041]

[0042] Where v represents the true value to be predicted. This represents the model's predicted value. α is a hyperparameter; the larger the value, the more emphasis is placed on directional alignment. When α is 0, this loss degenerates into the original MSE loss.

[0043] Preferably, the number of windows is set to "kk / 2-k / 4", that is, starting from k windows, gradually reducing to k / 2 windows, and then reducing to k / 4 windows.

[0044] Preferably, the training method is applied to a pre-trained diffusion model, wherein the diffusion model is StableDiffusion v1.5.

[0045] Preferably, the dataset used in the training process is the Laion-Art dataset or a subset thereof.

[0046] Preferably, the preprocessing includes adjusting the image resolution and applying center cropping.

[0047] A flow matching diffusion model trained based on the aforementioned multi-stage flow matching diffusion model training method.

[0048] A method for image generation using the aforementioned flow matching diffusion model includes the following steps:

[0049] Receive text description input from the user;

[0050] The flow matching diffusion model is used to generate the corresponding image based on the text description.

[0051] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-stage flow matching diffusion model training method.

[0052] Example 1

[0053] S1 involves obtaining the training dataset: This embodiment uses the Laion-Art dataset for fine-tuning. The Laion dataset is a large-scale, open-access image-text dataset developed by the Laion.ai community. It is characterized by its massive size and broad coverage, containing hundreds of millions of images and corresponding text descriptions. The Laion dataset collects images using web crawling technology, covering various types of images from everyday scenes to professional photography and artworks. The Laion-Art dataset used in this embodiment is a specific subset of the Laion dataset, focusing on images in the art field, and was obtained through a selection process from the Laion dataset.

[0054] It is important to note that since SD1.5 uses a fixed image resolution of 512*512 during the training phase, while the images in the Laion-Art dataset have varying resolutions, it is necessary to first adjust the shorter side of all images in the dataset to 512 resolution according to the original aspect ratio, and then apply center cropping to ensure that the final image resolution is also 512*512.

[0055] It should be noted that the dataset used in this embodiment is Laion-Art, but other datasets can also be used in actual applications without affecting the methods used subsequently.

[0056] S2 determines the number of inference steps for the pre-trained diffusion model: Due to the multi-stage training paradigm and the different initialization methods in each stage, we maintain the consistency of the sampling trajectory of the teacher model across different training stages by fixing the total number of DDIM steps to 32. Specifically, when the window size is 8, we use 4 DDIM steps within each window to derive the endpoint from the starting point. For the case of 4 windows, we use 8 DDIM steps within each window. This ensures that the sampling trajectory of the teacher model remains exactly the same across different training stages, thereby allowing for fair comparisons and achieving stable optimization.

[0057] S3 is for setting hyperparameters:

[0058] (1) This embodiment first explored the appropriate window number setting: taking the teacher model's 32 steps as the benchmark, it mainly explored the advantages and disadvantages of setting the window number to 16 and 8. Under the same number of iterations, it compared the performance of the models trained under different settings. The fid and clip scpre were calculated on COCO-5k as evaluation indicators, and the results are shown in Table 1. The results showed that in the two cases of total iterations of 10000 and 15000, the simplified window number gradually reduced to 4 was better than directly setting it to 4; and the effect of "16-8-4" was similar to "8-4" despite the additional training cost of 5000 iterations, indicating that more stages are not necessarily better.

[0059] Therefore, the number of windows was determined to be "8-4-2" in subsequent experiments.

[0060]

[0061] Table 1. Impact of different window number settings on model performance

[0062] (2) This embodiment first explores the setting of hyperparameters in the loss function on SD1.5. Experiments were conducted under two conditions: 8 windows and 8 inference steps, and 4 windows and 4 inference steps. Preliminary experiments with small batches and small iteration steps were performed first to compare the impact of different α settings on the generation effect under the same training cost. The experimental results are as follows: Figure 1 As shown, the improvement effect obtained by setting α=0.1 was relatively stable, so the setting of α=0.1 was used in subsequent experiments.

[0063] It should be noted that the hyperparameters for the number of exploration windows and the loss function in this embodiment are for the SD1.5 model, which resulted in the above settings. The same exploration method can also be transferred to other similar models.

[0064] It should be noted that this embodiment selects "8-4-2" when setting the number of windows, which can also be naturally extended to a window count of 1.

[0065] S4 is a multi-stage training program.

[0066] This embodiment uses an "8-4-2" window number setting and a loss function hyperparameter setting of α = 0.1.

[0067] First, starting with a window size of k (k=8), the unet of the pre-trained SD1.5 is used as the initialization of the student model at that stage. The sampling trajectory is divided into k windows, and each window contains 32 / k sampling steps.

[0068] A batch of data is randomly sampled. The latent space representations of the text and images are obtained by the text encoder and VAE of SD1.5, respectively. A time point t is randomly sampled and its time window is determined, denoted as [t0, t1].

[0069] Determine the start and end points of the time window. Based on the time schedule at the start point of SD1.5, determine the noise amplitude. Add noise to the latent space representation of the image based on this noise amplitude, denoted as latent_0. Input (t0, latent_0) into the teacher model, perform 32 / k DDIM steps for inference to time t1, and obtain latent_1.

[0070] The learning objective of the corrective flow, denoted as pred, is obtained by combining latent_1 and latent_0 with the SD1.5 training paradigm.

[0071] The loss term is calculated based on the loss function hyperparameters determined in this embodiment, and backpropagation is performed.

[0072] After convergence, an octave line sampling trajectory is formed; then the current stage's obtained UET model is used as the initialization of the next stage's student model, and the window number k=4 is adjusted, repeating (2) to (6) to obtain ProReflow-Ⅰ; then ProReflow-Ⅰ is used as the initialization of the next stage's student model, and the window number k=2 is set, repeating (2) to (6) to obtain ProReflow-Ⅱ.

[0073] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0074] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-stage flow matching diffusion model training method, characterized in that, Includes the following steps: S1 obtains the training dataset, which contains images and their corresponding text descriptions; S2 preprocesses the images in the training dataset to make the image resolution meet the requirements of the pre-trained diffusion model; S3 determines the number of inference steps for the pre-trained diffusion model to maintain the consistency of the sampling trajectory of the teacher model at different training stages; S4 sets hyperparameters, including the number of windows and the hyperparameters of the loss function; S5 starts with a large number of windows, divides the sampling trajectory into multiple time windows, and applies flow matching technology for training within each time window to obtain the initial polyline sampling trajectory; S6 gradually reduces the number of time windows, uses the model obtained from the previous stage of training as initialization, and continues training with the new number of windows to gradually optimize the sampling trajectory until the final sampling trajectory is obtained.

2. The multi-stage flow matching diffusion model training method according to claim 1, characterized in that, The flow matching technology mentioned is a corrected flow technology.

3. The multi-stage flow matching diffusion model training method according to claim 1, characterized in that, The training process uses a direction-guided loss function, which includes mean squared error loss and cosine similarity loss, and the relative importance of the two is controlled by hyperparameters. The direction-guided loss function is expressed as follows: ; Where v represents the true value to be predicted. This represents the model's predicted value. This is a hyperparameter; the larger it is, the more it emphasizes alignment in the direction. When it is 0, this loss degenerates into the original MSE loss.

4. The multi-stage flow matching diffusion model training method according to claim 1, characterized in that, The number of windows is set to "kk / 2-k / 4", that is, starting from k windows, gradually reducing to k / 2 windows, and then reducing to k / 4 windows.

5. The multi-stage flow matching diffusion model training method according to claim 1, characterized in that, The training method is applied to a pre-trained diffusion model, which is Stable Diffusion v1.

5.

6. The multi-stage flow matching diffusion model training method according to claim 1, characterized in that, The dataset used in the training process is the Laion-Art dataset or a subset thereof.

7. The multi-stage flow matching diffusion model training method according to claim 1, characterized in that, The preprocessing includes adjusting the image resolution and applying center cropping.

8. A flow-matching diffusion model trained based on the multi-stage flow-matching diffusion model training method according to any one of claims 1 to 7.

9. A method for image generation using the flow matching diffusion model of claim 8, characterized in that, Includes the following steps: Receive text description input from the user; The flow matching diffusion model is used to generate the corresponding image based on the text description.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-stage flow matching diffusion model training method according to any one of claims 1 to 7.