Diffusion model image generation method and device with fusion of structural disturbance guidance and consistent distillation
Through the method of structural perturbation guidance and consistent distillation fusion, the problems of high computational cost and insufficient quality when generating images in existing diffusion models are solved, and high-quality and efficient image generation are achieved.
Patent Information
- Application Number
- CN202510113704.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing diffusion models have problems with high computational cost and insufficient generation quality when generating images, especially in complex scenarios or detailed performances.
Using the method of fusion of structural perturbation guidance and consistent distillation, the original training data set is acquired and converted into latent space data sets, the student network and teacher network are initialized, and the classifier-free guidance after structural perturbation is introduced into the teacher network to perform state prediction and parameter updates, and the target model is obtained for image generation.
The quality and generation speed of generated images are improved, the calculation cost is reduced, and the generation process is more efficient, which can significantly accelerate the generation process of diffusion model while ensuring the generation quality.
Smart Images

Figure CN120047332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and in particular to a diffusion model image generation method and device that combines structural perturbation guidance and consistency distillation. Background Art
[0002] In recent years, significant progress has been made in the field of image generation. From the initial generative adversarial network (GAN) to the current diffusion model, generative models have gradually demonstrated stronger image generation capabilities. In the development process of diffusion models, various guidance methods have played a key role. Classifier Guidance (CG) is a method that provides semantic gradients through a pre-trained classifier to guide the generation process. However, due to the need for an additional classifier model, its computational overhead is relatively large. And the content generated by Classifier-Free Guidance (CFG) still has certain deficiencies in image structure, for example, it may lack consistency or stability in complex scenes or detail performance. In addition, existing latent space diffusion models need to perform denoising networks through dozens of iterations to gradually estimate the noise intensity at each step. There is a contradiction between this high computational cost and the user's demand for fast image generation in real-world scenarios. Summary of the Invention
[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a diffusion model image generation method and device that combines structural perturbation guidance and consistency distillation, in order to solve at least one of the existing technical problems. The present invention can improve the quality of generated images and accelerate the generation speed of the diffusion model.
[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a diffusion model image generation method that combines structural perturbation guidance and consistency distillation, and the method includes the following steps:
[0005] Obtain an original training data set, and convert the original training data set into a latent space data set;
[0006] Initialize the first student network through a diffusion model, and initialize the first teacher network;
[0007] Generate an initial latent space state according to the latent space data set and the noise scheduling function of the diffusion model;
[0008] Introduce classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network;
[0009] Generate an enhanced input of the first teacher network through state prediction by a solver according to the first teacher network, the second teacher network, and the initial latent space state;
[0010] Obtain a consistency loss according to the first student network, the first teacher network, and the enhanced input;
[0011] Update the parameters of the first student network according to the consistency loss to obtain a target model;
[0012] wherein, the target model is used for generating a target image.
[0013] In some embodiments, generating an initial latent space state according to the latent space dataset and the noise scheduling function of the diffusion model includes the following steps:
[0014] Randomly sample the latent space dataset to obtain a sampling sample;
[0015] Preset a time step range;
[0016] Randomly sample the time step range to obtain a sampled time step;
[0017] Generate the initial latent space state according to the noise scheduling function of the diffusion model, the sampling sample, and the sampled time step.
[0018] In some embodiments, introducing the classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network includes the following steps:
[0019] Perturb the first self-attention matrix of the first teacher network to generate a perturbed self-attention matrix;
[0020] Multiply the perturbed self-attention matrix by appearance information to obtain a perturbed self-attention module;
[0021] Replace the first self-attention module of the first teacher network with the perturbed self-attention module to obtain the second teacher network.
[0022] In some embodiments, the formula used for multiplying the perturbed self-attention matrix by appearance information to obtain a perturbed self-attention module includes:
[0023]
[0024] In the formula, PSA represents the perturbed self-attention module; Q t , K t , V t represent query, key, and value respectively; Shuffle(·) represents a perturbation operation; A t represents the first self-attention matrix; represents the perturbed self-attention matrix.
[0025] In some embodiments, the state prediction is performed by a solver according to the first teacher network, the second teacher network, and the initial latent space state to generate an enhanced input for the first teacher network, including the following steps:
[0026] Execute the solver based on the first teacher network to obtain first intermediate data;
[0027] Execute the solver based on the second teacher network to obtain second intermediate data;
[0028] According to the first intermediate data, the second intermediate data, and the initial latent space state, obtain the enhanced input for the first teacher network.
[0029] In some embodiments, the formula used for the state prediction by a solver according to the first teacher network, the second teacher network, and the initial latent space state to generate an enhanced input for the first teacher network includes:
[0030]
[0031] In the formula, represents the enhanced input for the first teacher network; z t+1 represents the initial latent space state; ω represents the guiding weight; t represents the sampling time step, t ∈ [0, T]; c represents the conditional information; φ represents the empty text; Ψ θ- represents the solver based on the first teacher network; represents the solver based on the second teacher network; θ - represents the parameters of the first teacher network; represents the parameters of the second teacher network.
[0032] In some embodiments, the obtaining of the consistency loss according to the first student network, the first teacher network, and the enhanced input includes the following steps:
[0033] Obtain a first prediction of the first student network;
[0034] According to the enhanced input of the first teacher network, obtain a distillation target output by the first teacher network;
[0035] According to the first prediction, the distillation target, and the metric function, obtain the consistency loss.
[0036] In some embodiments, the formula used for obtaining the consistency loss according to the first student network, the first teacher network, and the enhanced input includes:
[0037]
[0038] In the formula, represents the consistency loss; θ represents the parameters of the first student network; θ - represents the parameters of the first teacher network; Ψ θ- represents the solver based on the first teacher network; f θ (z t+1 , c, t + 1) represents the first prediction of the first student network; z t+1 represents the initial state of the latent space; c represents the conditional information; t represents the sampling time step, represents the distillation target output by the first teacher network; represents the enhanced input of the first teacher network; d(·, ·) represents the metric function; represents the average expected value of the random variables z, c, ω, and t.
[0039] To achieve the above object, on the other hand, an embodiment of the present invention proposes a diffusion model image generation device that fuses structural perturbation guidance and consistency distillation. The device includes:
[0040] The first module is used to obtain the original training data set and convert the original training data set into a latent space data set;
[0041] The second module is used to initialize the first student network through a diffusion model and initialize the first teacher network;
[0042] The third module is used to generate the initial state of the latent space according to the latent space data set and the noise scheduling function of the diffusion model;
[0043] The fourth module is used to introduce classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network;
[0044] The fifth module is used to perform state prediction through a solver according to the first teacher network, the second teacher network, and the initial state of the latent space to generate the enhanced input of the first teacher network;
[0045] The sixth module is used to obtain the consistency loss according to the first student network, the first teacher network, and the enhanced input;
[0046] The seventh module is used to update the parameters of the first student network according to the consistency loss to obtain a target model;
[0047] Wherein, the target model is used for generating target images.
[0048] To achieve the above object, on the other hand, an embodiment of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the diffusion model image generation method that combines structural perturbation guidance and consistency distillation described above.
[0049] To achieve the above object, on the other hand, an embodiment of the present invention provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, it implements the diffusion model image generation method that combines structural perturbation guidance and consistency distillation described above.
[0050] To achieve the above object, on the other hand, an embodiment of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions that are stored in a computer-readable storage medium. The processor of a computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device is caused to execute the diffusion model image generation method that combines structural perturbation guidance and consistency distillation described above.
[0051] The embodiments of the present invention at least include the following beneficial effects: The present invention provides a diffusion model image generation method and apparatus that combine structural perturbation guidance and consistency distillation. The solution includes obtaining an original training data set and converting the original training data set into a latent space data set; initializing a first student network through a diffusion model and initializing a first teacher network; generating an initial latent space state according to the latent space data set and the noise scheduling function of the diffusion model; introducing classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network; performing state prediction through a solver according to the first teacher network, the second teacher network, and the initial latent space state to generate an enhanced input for the first teacher network; obtaining a consistency loss according to the first student network, the first teacher network, and the enhanced input; and updating the parameters of the first student network according to the consistency loss to obtain a target model, where the target model is used for generating target images. The present invention simplifies the complex generation process into a simplified model, enabling the teacher model to be responsible for generating high-quality distillation targets, and the student model to approximate complex generation behaviors by learning these targets. In addition, classifier-free guidance after structural perturbation is introduced into the teacher network to further improve the quality of the generated images and make the generation process more efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0053] Figure 1 is a flowchart of a diffusion model image generation method that fuses structural perturbation guidance and consistency distillation provided by an embodiment of the present invention;
[0054] Figure 2 is a schematic diagram of the specific steps of diffusion model image generation that fuses structural perturbation guidance and consistency distillation provided by an embodiment of the present invention;
[0055] Figure 3 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0056] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present invention. They are only examples of devices and methods that are consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.
[0057] It should be noted that although functional module division is performed in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the system or the flowchart in the flowchart. The terms "first / S100" and "second / S200" in the specification, claims and the above accompanying drawings can be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information can also be called the second information, and similarly, the second information can also be called the first information. Depending on the context, the words "if" and "when" as used herein can be interpreted as "when" or "when" or "in response to a determination".
[0058] The terms "at least one", "a plurality", "each", "any one", etc. used in the present invention, at least one includes one, two or more, a plurality includes two or more, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used herein are for the purpose of describing embodiments of the invention only and are not intended to limit the invention.
[0060] In recent years, significant progress has been made in the field of image generation technology. From the initial Generative Adversarial Network (GAN) to the current diffusion models, generative models have gradually demonstrated stronger image generation capabilities. GAN generates realistic images through adversarial training, but there are certain limitations in terms of generation diversity and training stability. In contrast, diffusion models adopt a step-by-step denoising generation strategy, starting from random noise and gradually restoring clear images through multiple iterations. This generation mechanism can accurately model the data distribution, enabling the generated images to not only have higher quality but also possess a wider range of diversity. Therefore, it has gradually become the mainstream method in the field of image generation.
[0061] In the development process of diffusion models, various guidance methods have played a crucial role. Classifier Guidance (CG) is a method that provides semantic gradients through a pre-trained classifier to guide the generation process. However, due to the need for an additional classifier model, its computational overhead is relatively large. In contrast, Classifier-Free Guidance (CFG) simplifies the generation process by fusing conditional distributions with unconditional distributions, getting rid of the dependence on the classifier, and further improving the quality of the generated images. Therefore, CFG has become the current mainstream guidance method. However, there are still certain deficiencies in the image structure of the content generated by CFG. For example, there may be a lack of consistency or stability in complex scenes or detailed representations.
[0062] In addition to generation quality, generation speed is also one of the main challenges faced by diffusion models. Although latent space diffusion models reduce the computational complexity of the high-dimensional pixel space, improving computational efficiency and flexibility while ensuring high-quality image synthesis, existing latent space diffusion models still need to perform denoising network operations dozens of times to gradually estimate the noise intensity at each step. There is a contradiction between this high computational cost and the user's demand for quickly generating images in real-world scenarios.
[0063] In view of this, as Figure 1 shown, embodiments of the present invention provide a diffusion model image generation method that fuses structural perturbation guidance and consistency distillation, which may include but is not limited to steps S100 to S700:
[0064] Step S100, obtain an original training data set and convert the original training data set into a latent space data set;
[0065] Step S200: Initialize the first student network through a diffusion model and initialize the first teacher network;
[0066] Step S300: Generate an initial latent space state according to the latent space dataset and the noise scheduling function of the diffusion model;
[0067] Step S400: Introduce classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network;
[0068] Step S500: Perform state prediction through a solver according to the first teacher network, the second teacher network, and the initial latent space state to generate an enhanced input for the first teacher network;
[0069] Step S600: Obtain a consistency loss according to the first student network, the first teacher network, and the enhanced input;
[0070] Step S700: Update the parameters of the first student network according to the consistency loss to obtain a target model;
[0071] Among them, the target model is used for the generation of target images.
[0072] In step S100 of some embodiments, obtain the original training dataset D, and convert the original training dataset into a latent space dataset D through the encoder E(·) z , then there is D z = {(z, c)|z = E(x), (x, c) ∈ D}, where x is the original training data, z is the latent space encoding of the original training data x, and c is the corresponding conditional information. The representation through the latent space can significantly reduce the computational complexity and provide a more efficient basis for the subsequent training and inference of the model.
[0073] In step S200 of some embodiments, initialize the first student network using the parameters θ of the pre-trained diffusion model. At the same time, define the first teacher network, and the parameters θ of this first teacher network - can be updated using exponential moving average (EMA), and there is the following expression:
[0074] θ - ← μθ - + (1 - μ)θ
[0075] In the formula, θ - represents the parameters of the first teacher network; θ represents the parameters of the first student network; μ represents the weight hyperparameter of the exponential moving average. The exponential moving average (EMA) mechanism can effectively maintain the stability of the training process.
[0076] In some embodiments, the diffusion model is a generative model that generates samples step by step. It constructs the forward process by injecting Gaussian noise into the data and gradually recovers the data samples from the noise through reverse denoising.
[0077] (1) Forward diffusion process:
[0078]
[0079] where \(t\in[0,T]\) is the time step; \(x\) t represents the data state at time step \(t\); \(x\) 0 represents the original data state; \(\alpha(t)\), \(\sigma(t)\) represent the noise scheduling functions at time step \(t\); \(I\) represents the identity matrix; represents the Gaussian distribution; \(q(x\) t | \(x\) 0 ) represents the conditional probability distribution of the data state at time step \(t\) given the original data state. When \(t\) approaches 0, \(x\) t approaches the original data distribution \(p\) data (\(x\)); when \(t\) approaches \(T\), \(x\) t approaches the marginal distribution.
[0080] (2) Reverse generation process:
[0081] By training the denoising network \(\epsilon\) θ (\(x\) t , \(t\)), model \(x\) t and \(t\), evaluate the noise intensity, and gradually restore the marginal distribution to the data distribution. The network parameters \(\theta\) are optimized to minimize the denoising error, and the reverse generation formula is:
[0082]
[0083] where \(\mu\) θ represents the mean; \(\Sigma\) θ represents the covariance.
[0084] In some embodiments, step S300 may include but is not limited to steps S310 to S340:
[0085] Step S310, randomly sample the latent space dataset to obtain sampling samples;
[0086] Step S320, preset the time step range;
[0087] Step S330, randomly sample the time step range to obtain sampling time steps;
[0088] Step S340, generate the initial state of the latent space according to the noise scheduling function of the diffusion model, the sampling samples, and the sampling time steps.
[0089] In steps S310 to S340 of some embodiments, a sample pair (z, c) is randomly sampled from the latent space dataset D z where z is the latent space encoding of the original training data x and c is the conditional information; a random sample is taken from the time step range [1, T], so that the sampling time step t ∈ [1, T]; then, using the noise scheduling functions α(t) and σ(t) of the diffusion model, the initial latent space state z t+1 can be generated, and its distribution is a Gaussian distribution:
[0090]
[0091] This process simulates the forward diffusion process of the diffusion model, gradually introducing noise into the data.
[0092] In some embodiments, the self-attention module of the diffusion model models the structure and appearance of the image, where the query-key similarity is mainly responsible for structure modeling, and the value processes the appearance information. In the self-attention module at time step t, its output can be expressed as:
[0093]
[0094] where are the query, key, and value respectively; d represents the number of channels of the image; A t is the first self-attention matrix, which is used to model the relationship between different positions of the input image, especially the structural information of the image.
[0095] In some embodiments, consistency distillation is a method of approximating the generation behavior of a complex model by training a simplified model. In the diffusion model, its goal is to simplify the complex multi-step reverse generation process into fewer steps or even a single step generation while maintaining the generation quality. The core idea of consistency distillation is based on knowledge distillation, which condenses the complex generation behavior of the teacher model into an efficient generation process of the student model. Latent space consistency distillation further extends this distillation process to the latent space, thereby significantly improving the computational efficiency and inference speed. Given a solution trajectory {x t} t∈[0,T] of the diffusion model PF-ODE, the consistency function is defined as a function that maps from any time step (x t , t) to the initial state (x 0 , 0). The following are the core concepts of the latent space consistency distillation scheme:
[0096] 1) The consistency function has the following consistency properties
[0097] For the consistency function defined for any states (x t , t) and (x t′ , t′) on the same PF-ODE trajectory, its output results must be the same, that is
[0098]
[0099] 2) The boundary condition of the consistency function is the identity mapping
[0100] f(x 0 , 0) = x 0
[0101] This boundary condition is crucial during the training process of consistency distillation. For any trained diffusion model F θ (x, t), the consistency function can be defined in the following form to satisfy the boundary condition:
[0102] f θ (x, t) = c skip (t)x + c out (t)F θ (x, t)
[0103] where c skip (t) and c out (t) are differentiable functions of time t, satisfying c skip (0) = 1 and c out (0) = 0.
[0104] In traditional latent space consistency distillation, the teacher model predicts the latent space state corresponding to time step t from the initial latent space state z using an ODE solver (i.e., t+1 ) At this time, the ODE solver is executed indicating that its calculation is based on the teacher network, then generating the input of the teacher network There is the following expression:
[0105]
[0106] According to the above-generated input of the teacher network define the consistency loss as:
[0107]
[0108] where f θ (z t+1 , c, t + 1) represents the prediction of the student network; represents the distillation target output by the teacher network; d(·, ·) represents the metric function (such as Euclidean distance); Denotes the average expected value of the random variables z, c, and t. Based on the consistency loss, the parameters θ of the student model are optimized by gradient descent. At the same time, the parameters θ of the teacher network are updated using the EMA mechanism. - .
[0109] In some embodiments, classifier-free guidance (CFG) is a method to improve the generation quality of diffusion models, especially suitable for text-to-image tasks. It does not rely on a classifier but directly utilizes the capabilities of the conditional model itself for guidance, enhancing the generation quality and conditional consistency. By interpolating between unconditional generation and conditional generation, classifier-free guidance can effectively enhance the guiding effect of the condition on the generation task. Applying CFG in a text-to-image model can be expressed as:
[0110]
[0111] where x t represents the data state at time step t; c represents the conditional information; φ represents the empty text (unconditional); ω represents the guidance weight used to adjust the conditional guidance strength; ∈ θ represents the noise prediction network of the conditional model; represents the enhanced noise prediction network after classifier-free guidance. The enhanced noise prediction network after classifier-free guidance is composed of the outputs of two branch networks with exactly the same parameters, except for the different input conditions, one is c and the other is φ.
[0112] Classifier-free guidance can improve the matching degree of specific conditions in the generation task by directly utilizing the capabilities of the conditional model itself. Ordinary classifier-free guidance improves the generation quality by introducing the empty text (unconditional branch) in the diffusion model, but it mainly uses the conditional model itself for guidance without changing the model architecture. Although this method can improve the quality of image generation to a certain extent, there are still the following deficiencies:
[0113] 1) The structural integrity problem of unconditional generation: In classifier-free guidance, the sample structure of unconditional generation is relatively complete, but this cannot ensure that the model can effectively avoid potentially structurally collapsed samples during conditional generation.
[0114] 2) The limitation of the role of negative samples: Due to the lack of a clear negative sample interference mechanism, the finally generated images may still show structural degradation under conditional guidance.
[0115] In view of this, in steps S400 to S600 of some embodiments, a combination of a structure perturbation mechanism and classifier-free guidance is proposed. That is, through a structure perturbation guidance mechanism, the classifier-free guidance after structure perturbation is applied to the process of the teacher model generating the distillation target. By introducing negative samples of structure collapse, the unconditional generated samples not only differ from the conditional generated samples in terms of category, but also significantly collapse in structure, thereby strengthening the guiding role of the negative sample target. This combination method can more clearly guide the conditional model to stay away from collapsed samples during the generation process, further improving the structural consistency and quality of the finally generated images.
[0116] Optionally, the core idea of fusing structure perturbation with classifier-free guidance is to modify the conditional model of the unconditional branch. By introducing a structure perturbation mechanism, the self-attention matrix is disturbed to disrupt the structure of the unconditional generated samples. Its formula is expressed as:
[0117]
[0118] where c represents conditional information; φ represents empty text (unconditional); ω is the guidance weight used to control the intensity of conditional guidance; ∈ θ represents the noise prediction network of the conditional model; represents the unconditional branch model modified by structure perturbation. The difference from the ordinary classifier-free guidance formula is that the diffusion model parameters of the unconditional branch are modified from θ to to introduce structure perturbation.
[0119] Through the above formula, the embodiments of the present invention explicitly introduce a structure perturbation mechanism in the unconditional generation branch, making the generated negative sample target clearer. During the actual training process, these negative samples of structure collapse can significantly strengthen the learning objective of the conditional model, guiding it to avoid degenerate sample distributions, so as to better reconstruct high-quality image structures during the generation process.
[0120] In the embodiments of the present invention, the classifier-free guidance after structure perturbation is incorporated into the latent space consistency distillation. Combining the classifier-free guidance after structure perturbation with the latent space consistency distillation can significantly improve the quality and inference efficiency of the generated images. Distillation transforms the complex generation process into a learning process of a simplified model, enabling the teacher model to be responsible for generating high-quality distillation targets, while the student model approximates the complex generation behavior by learning these targets. In the latent space consistency distillation, the distillation targets generated by the teacher model are usually obtained through a multi-step inference process, and the classifier-free guidance after structure perturbation is introduced into the teacher network, thereby further improving the quality of the distillation targets generated by the teacher network, improving the quality of the generated images, and making the generation process more efficient.
[0121] In some embodiments, step S400 may include, but is not limited to, steps S410 to S430:
[0122] Step S410, perturb the first self-attention matrix of the first teacher network to generate a perturbed self-attention matrix;
[0123] Step S420, multiply the perturbed self-attention matrix with appearance information to obtain a perturbed self-attention module;
[0124] Step S430, replace the first self-attention module of the first teacher network with the perturbed self-attention module to obtain the second teacher network.
[0125] In steps S410 to S430 of some embodiments, the first self-attention matrix A of the first teacher network t is perturbed, that is, the rows of the first self-attention matrix A t are shuffled in order to generate a perturbed self-attention matrix thereby destroying the structural information of the image. Subsequently, the perturbed self-attention matrix is multiplied with the appearance information V t to obtain the output of the module. The first self-attention module after structural perturbation is rewritten as a perturbed self-attention module (PerturbedSelf-Attention, PSA) to obtain the second teacher network. Then the formula of the perturbed self-attention module is as follows:
[0126]
[0127] In the formula, PSA represents the perturbed self-attention module; Q t , K t , V t represent query, key, and value respectively; Shuffle(·) represents the perturbation operation; A t represents the first self-attention matrix; represents the perturbed self-attention matrix.
[0128] Through the structural perturbation guidance mechanism, that is, by interfering with the query-key similarity in the first self-attention matrix to generate negative samples with structural collapse, to guide the model to avoid generating degenerate samples, thereby improving the robustness of the denoising process.
[0129] In some embodiments, step S500 may include, but is not limited to, steps S510 to S530:
[0130] Step S510, execute the solver based on the first teacher network to obtain first intermediate data;
[0131] Step S520: Execute the solver based on the second teacher network to obtain second intermediate data;
[0132] Step S530: Obtain the enhanced input of the first teacher network according to the first intermediate data, the second intermediate data, and the initial state of the latent space.
[0133] In steps S510 to S530 of some embodiments, execute the solver Ψ based on the first teacher network θ- , and obtain the first intermediate data Ψ θ- (z t+1 , t + 1, t, c), execute the solver based on the second teacher network to obtain the second intermediate data According to the first intermediate data, the second intermediate data, and the initial state of the latent space, the formula for obtaining the enhanced input of the first teacher network is as follows:
[0134]
[0135] In the formula, represents the enhanced input of the first teacher network; z t+1 represents the initial state of the latent space; ω represents the guidance weight, which controls the unconditional guidance strength after structural perturbation; t represents the sampling time step, t ∈ [0, T]; c represents the conditional information; φ represents the empty text (unconditional); Ψ * represents that the ODE solver is calculated based on the diffusion model *, and is responsible for simulating the denoising process; Ψ θ- represents the solver based on the first teacher network; represents the solver based on the second teacher network; θ - represents the parameters of the first teacher network; represents the parameters of the second teacher network.
[0136] When the classifier-free guidance after structural perturbation acts on the teacher network, it means that the unconditional branch is explicitly perturbed to generate negative samples of structural collapse. These negative samples effectively help the teacher network avoid generating samples with structural degradation, thereby improving the quality of the generated images. In this way, the quality of the latent space state generated by the teacher network has been significantly improved, while strengthening the learning objective of the student model during the training process. This combination method ensures that the target generated by the teacher network is far from potential negative samples of structural collapse, so that the student model can better learn how to generate high-quality and structurally consistent images. Through this improvement, the student model can rely on fewer steps during inference, and even complete high-quality image generation through single-step inference without relying on the traditional multi-step inference process. This not only greatly improves the inference efficiency, but also ensures the quality and consistency of the finally generated images.
[0137] In some embodiments, step S600 may include, but is not limited to, steps S610 to S630:
[0138] Step S610, obtaining a first prediction of the first student network;
[0139] Step S620, obtaining a distillation target output by the first teacher network according to the enhanced input of the first teacher network;
[0140] Step S630, obtaining the consistency loss according to the first prediction, the distillation target, and a metric function.
[0141] In steps S610 to S630 of some embodiments, obtain a first prediction f θ (z t+1 , c, t + 1) of the first student network, and obtain a distillation target output by the first teacher network according to the enhanced input of the first teacher network Then, in combination with a metric function, the following expression for the consistency loss can be obtained:
[0142]
[0143] In the formula, represents the consistency loss; θ represents the parameters of the first student network; θ - represents the parameters of the first teacher network; Ψ θ- represents a solver based on the first teacher network; f θ (z t+1 , c, t + 1) represents the first prediction of the first student network; z t+1 represents the initial state of the latent space; c represents conditional information; t represents the sampling time step, represents the distillation target output by the first teacher network; represents the enhanced input of the first teacher network; d(·, ·) represents a metric function; represents the average expected value of the random variables z, c, ω, and t.
[0144] In step S700 of some embodiments, based on the consistency loss, the parameters θ of the first student model are optimized by gradient descent, and at the same time, the parameters of the first teacher network are updated using the EMA mechanism, then a target model for generating the target image can be obtained. Exemplarily, through the calculation of the consistency loss, the parameters of the first student model and the parameters of the first teacher network are updated, and it is judged whether the consistency loss meets the preset convergence condition. If not, the step of generating the initial state of the latent space according to the latent space dataset and the noise scheduling function of the diffusion model is returned. If it is satisfied, a target model for generating the target image can be obtained. The model obtained through consistency distillation training can complete high-quality generation tasks in fewer steps or even a single step, greatly improving the inference efficiency.
[0145] In summary, the processing flow of a diffusion model image generation method that fuses structural perturbation guidance and consistency distillation according to an embodiment of the present invention is as Figure 2 shown:
[0146] Step 1, Latent space data construction: The original training dataset is converted into a latent space dataset through an encoder;
[0147] Step 2, Model initialization and EMA update mechanism: The parameters of the pre-trained diffusion model are used to initialize the student model. At the same time, a teacher network is defined, and the parameters of the teacher network are updated using exponential moving average (EMA).
[0148] Step 3, Random sampling to generate the initial state of the latent space: In each training iteration, a sample pair is randomly sampled from the latent space dataset, a time step is randomly sampled from the time step range, and then the initial state of the latent space is generated using the noise scheduling function of the diffusion model. This process simulates the forward diffusion process of the diffusion model and gradually introduces noise into the data;
[0149] Step 4, Add a classifier-free guidance branch with structural perturbation, introduce the classifier-free guidance after structural perturbation into the teacher network, and based on the teacher network, use an ODE solver to predict the state corresponding to the time step from the initial state of the latent space to generate the input of the teacher network;
[0150] Step 5, Define the consistency loss using the prediction of the student network and the distillation target output by the teacher network. Based on the consistency loss, the parameters of the student model are optimized by gradient descent, and at the same time, the parameters of the teacher network are updated using the EMA mechanism;
[0151] Step 6, Convergence and model output: Repeat steps 3 to 5 until the loss function meets the preset convergence condition, the model converges, and a target model for generating the target image is obtained. The model obtained through consistency distillation training can complete high-quality generation tasks in fewer steps or even a single step, greatly improving the inference efficiency.
[0152] An embodiment of the present invention also provides a diffusion model image generation device that fuses structural perturbation guidance and consistency distillation, which can implement the above-mentioned diffusion model image generation method that fuses structural perturbation guidance and consistency distillation. The device includes:
[0153] A first module for obtaining an original training data set and converting the original training data set into a latent space data set;
[0154] A second module for initializing a first student network through a diffusion model and initializing a first teacher network;
[0155] A third module for generating an initial state of the latent space according to the latent space data set and the noise scheduling function of the diffusion model;
[0156] A fourth module for introducing classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network;
[0157] A fifth module for performing state prediction through a solver according to the first teacher network, the second teacher network, and the initial state of the latent space to generate an enhanced input for the first teacher network;
[0158] A sixth module for obtaining a consistency loss according to the first student network, the first teacher network, and the enhanced input;
[0159] A seventh module for updating the parameters of the first student network according to the consistency loss to obtain a target model;
[0160] Wherein, the target model is used for generating target images.
[0161] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0162] An embodiment of the present invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned diffusion model image generation method that fuses structural perturbation guidance and consistency distillation. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0163] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0164] ReferenceFigure 3 , Figure 3 illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0165] A processor 801, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0166] A memory 802, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802 and are called by the processor 801 to execute a diffusion model image generation method that combines structural perturbation guidance and consistency distillation;
[0167] An input / output interface 803, which is used to implement information input and output;
[0168] A communication interface 804, which is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0169] A bus 805, which transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);
[0170] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 achieve communication connections with each other inside the device through the bus 805.
[0171] The embodiments of the present invention also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned diffusion model image generation method that combines structural perturbation guidance and consistency distillation.
[0172] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0173] The embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the foregoing diffusion model image generation method that fuses structural perturbation guidance and consistency distillation.
[0174] In summary, the diffusion model image generation method and device that fuse structural perturbation guidance and consistency distillation in the embodiment of the present invention have the following advantages:
[0175] 1. In the classifier-free guidance framework, the embodiment of the present invention perturbs the image structure additionally to generate samples with degraded structures, and uses this guidance to drive the denoising process away from these degraded samples, thereby improving the structural quality and visual effects of the generated images (such as text-to-image), enhancing the structural consistency and detail representation ability of the generated images, and significantly optimizing the generation quality of complex scenes.
[0176] 2. The embodiment of the present invention introduces consistency distillation. By learning and optimizing the inference path in the latent space, it significantly reduces the number of inference steps. The number of steps required for inference is greatly reduced from dozens of steps in the traditional method to single digits, and even single-step generation is supported. While ensuring the generation quality, it significantly accelerates the generation process of the diffusion model, greatly improving the generation speed and meeting the actual needs of users for quickly generating images.
[0177] 3. By distilling the classifier-free guidance after structural perturbation into the model, the embodiment of the present invention further reduces the computational overhead of the classifier-free guidance branch. This optimization scheme significantly improves the generation efficiency of the model on the premise of ensuring the generation quality. Through the dual optimization of generation quality and efficiency, the innovative method in the embodiment of the present invention not only meets the requirements of high-quality image generation, but also greatly improves the generation speed, providing a more efficient and reliable solution for practical application scenarios.
[0178] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0179] In addition, while the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.
[0180] If the functions are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of the technical solution may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0181] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[0182] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.
[0183] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0184] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0185] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0186] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A diffusion model image generation method based on structural perturbation guidance and consistency distillation fusion, characterized in that: The following steps are involved: Acquire an original training data set, and convert the original training data set into a latent space data set; Initialize the first student network through the diffusion model, and initialize the first teacher network; generating an initial state of a latent space according to the latent space data set and a noise scheduling function of the diffusion model; Introducing the classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network; According to the first teacher network, the second teacher network and the initial state of the latent space, a state prediction is performed by a solver to generate an enhanced input of the first teacher network; Obtaining a consistency loss based on the first student network, the first teacher network, and the enhanced input; According to the consistency loss, updating the parameters of the first student network to obtain a target model; Wherein, the target model is used to generate a target image.
2. According to claim 1, the diffusion model image generation method based on structural perturbation guidance and consistency distillation fusion is characterized in that: The generating of the initial state of the latent space according to the latent space data set and the noise scheduling function of the diffusion model comprises the following steps: Randomly sampling the latent space data set to obtain sampling samples; Preset time step range; Randomly sampling the time step range to obtain a sampling time step; The initial state of the latent space is generated according to the noise scheduling function of the diffusion model, the sampling samples and the sampling time step.
3. The diffusion model image generation method of structural perturbation guidance and consistency distillation fusion according to claim 1 is characterized in that: The step of introducing the structure-perturbed non-classifier guidance into the first teacher network to obtain a second teacher network comprises the following steps: Perturbing the first self-attention matrix of the first teacher network to generate a perturbed self-attention matrix; Multiplying the perturbation self-attention matrix by the appearance information to obtain a perturbation self-attention module; The first self-attention module of the first teacher network is replaced by the perturbation self-attention module to obtain the second teacher network.
4. The method for generating diffusion model images by integrating structure perturbation guidance and consistency distillation according to claim 3, characterized in that: The perturbation self-attention matrix is multiplied by the appearance information to obtain the perturbation self-attention module, and the formula used includes: Where PSA stands for the perturbation self-attention module; Q t ,K t ,V t Represent query, key, and value respectively; Shuffle(·) represents the perturbation operation; A t represents the first self-attention matrix; Represents the perturbed self-attention matrix.
5. The diffusion model image generation method of structural perturbation guidance and consistency distillation fusion according to claim 1 is characterized in that: According to the first teacher network, the second teacher network and the initial state of the latent space, The state prediction is performed by the solver to generate the enhanced input of the first teacher network, including the following steps: Executing the solver based on the first teacher network to obtain first intermediate data; Executing the solver based on the second teacher network to obtain second intermediate data; The enhanced input of the first teacher network is obtained according to the first intermediate data, the second intermediate data and the initial state of the latent space.
6. The diffusion model image generation method of structural perturbation guidance and consistency distillation fusion according to claim 1 is characterized in that: The state prediction is performed by the solver according to the first teacher network, the second teacher network and the initial state of the latent space to generate the enhanced input of the first teacher network, and the formula used includes: In the formula, represents the augmented input of the first teacher network; z t+1 represents the initial state of the latent space; ω represents the guidance weight; t represents the sampling time step, t∈[0,T]; c represents the conditional information; φ represents the empty text; Ψ θ- represents the solver based on the first teacher network; represents the solver based on the second teacher network; θ - represents the parameters of the first teacher network; Represents the parameters of the second teacher network.
7. The diffusion model image generation method of structural perturbation guidance and consistency distillation fusion according to claim 1 is characterized in that: The step of obtaining a consistency loss according to the first student network, the first teacher network and the enhanced input comprises the following steps: Obtaining a first prediction of the first student network; According to the enhanced input of the first teacher network, obtaining a distillation target output by the first teacher network; The consistency loss is obtained according to the first prediction, the distillation target, and the metric function.
8. The diffusion model image generation method of structural perturbation guidance and consistency distillation fusion according to claim 1 is characterized in that: The consistency loss is obtained according to the first student network, the first teacher network and the enhanced input, and the formula used includes: In the formula, represents the consistency loss; θ represents the parameters of the first student network; θ - represents the parameters of the first teacher network; Ψ θ- represents the solver based on the first teacher network; f θ (z t+1 ,c,t+1) represents the first prediction of the first student network; z t+1 represents the initial state of the latent space; c represents the conditional information; t represents the sampling time step, t∈[0,T]; represents the distillation target of the output of the first teacher network; represents the augmented input of the first teacher network; d(·,·) represents the metric function; Represents the average expected value of the random variables z, c, ω and t.
9. A diffusion model image generation device based on structural perturbation guidance and consistency distillation fusion, characterized in that: include: The first module is used to obtain an original training data set and convert the original training data set into a latent space data set; The second module is used to initialize the first student network through the diffusion model and initialize the first teacher network; A third module is used to generate an initial state of the latent space according to the latent space data set and the noise scheduling function of the diffusion model; A fourth module is used to introduce the classifier-free guidance after structural perturbation into the first teacher network to obtain a second teacher network; A fifth module is used to perform state prediction through a solver according to the first teacher network, the second teacher network and the initial state of the latent space to generate an enhanced input of the first teacher network; A sixth module, configured to obtain a consistency loss according to the first student network, the first teacher network and the enhanced input; A seventh module is used to update the parameters of the first student network according to the consistency loss to obtain a target model; Wherein, the target model is used to generate a target image.
10. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
YOLOv5-based semi-supervised power quality disturbance identification and classification method, system and device, and storage medium
CN117034104A
Image generation method and device based on diffusion model
CN117291232A
Edge-end data security authentication method based on perceptual hash coding
CN118364862A
Diffusion model distillation method based on cross-image pixel space relationship
CN118587527A
One-step diffusion distillation via depth equilibrium model
CN119106708A
Cited By
Diffusion model sampling and distilling method for advertisement image material generation
CN120373356A
Unsupervised industrial defect detection network construction method
CN120953222A
Diffusion model reasoning acceleration method based on optimal time step sequence search and knowledge distillation
CN121809699A
Diffusion model inference acceleration method based on optimal time step sequence search and knowledge distillation
CN121809699B