Method and system for accelerating diffusion model
By buffering noise and using masks to control the denoising process, the problem of slow generation of diffusion model is solved, and a significant acceleration effect is achieved while maintaining the stability of the generation effect.
Patent Information
- Application Number
- CN202510195449.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-20
AI Technical Summary
The diffusion model is slow in generation due to multi-step iterative denoising. The existing acceleration methods have complex training and deployment processes, and different neural network structures cannot be applied.
The diffusion model is accelerated by buffering noise, and the specific steps include noise generation, relative change calculation, cache time step selection and mask control denoising process. This method quantifies noise similarity, selects the time step for reusable noise, and uses a mask to control the noise source during the denoising process.
The number of noise predictions during the diffusion model denoising process is reduced, and the generation speed is significantly improved, while avoiding additional time overhead and performance losses.
Smart Images

Figure CN120182113A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of generative models, and particularly to an acceleration method and system for diffusion models. Background Art
[0002] Diffusion models have shown powerful capabilities in the field of image generation. However, since they restore images by continuously denoising, they require many iterative denoising steps. In this case, the models generally have the problem of slow generation speed. Simply reducing the number of denoising steps of the diffusion model will cause the generated images to be distorted. Therefore, a new method needs to be designed to accelerate the diffusion model.
[0003] Currently, the acceleration methods for diffusion models can be roughly divided into three types. The first method aims to reduce the number of denoising steps to achieve the acceleration effect, such as optimizing the solution algorithm or distillation. The second method hopes to optimize the neural network to accelerate each denoising step, such as model pruning. However, these two methods require additional calculation steps and adjustments, which may complicate the training and deployment processes of the model. The third method hopes to store the intermediate calculations of some blocks in the model and reuse these cached results to reduce the computational burden during the denoising process. However, a key limitation of these methods is that they can only be applied to specific neural network architectures, such as U-Net or Transformer, making them inapplicable between different network structures.
[0004] Therefore, there is an urgent need to propose a strategy for accelerating diffusion models by caching noise to solve the problem of slow generation speed caused by multi-step iterative denoising of diffusion models. Summary of the Invention
[0005] To solve the problem of slow generation speed caused by multi-step iterative denoising of the diffusion model in the above-mentioned prior art, the present invention proposes a strategy for accelerating the diffusion model by caching noise.
[0006] In a first aspect, an embodiment of the present application provides an acceleration method for a diffusion model, the method including:
[0007] Noise generation step: Generate a batch of image data through a diffusion model, and obtain the noise mean value of each step of the image data;
[0008] Noise relative change amount calculation step: Calculate the relative change amount of adjacent-step noise according to the noise tensors of two consecutive time steps, and evaluate the similarity of adjacent noises based on the relative change amount;
[0009] Cached time step selection step: According to the similarity of adjacent noises, formulate a strategy to screen the cached time steps of reusable noises and generate corresponding masks; for the selected cached time steps, directly reuse the noises cached in the previous step;
[0010] Mask control denoising step: Use a mask to control the noise cache at each step during the denoising process.
[0011] In a specific embodiment of the present invention, the above-mentioned noise generation step includes:
[0012] For the denoising process of multiple time steps, generate a batch of image data through the diffusion model at multiple time steps, record the noise at each time step and calculate the average value.
[0013] In a specific embodiment of the present invention, the above-mentioned noise relative change amount calculation step includes:
[0014] Adopt the noise relative change amount L t to reflect the similarity degree of the noise, L t The calculation formula is as follows:
[0015]
[0016] where, ∈ t and ∈ t-1 are the noise tensors at consecutive time steps t and t - 1 respectively, ||.||1 represents the L1 norm, and the smaller the value of Lt, the smaller the difference between the two noises, and the more similar the two noises are.
[0017] In a specific embodiment of the present invention, the above-mentioned cache time step selection step includes:
[0018] Calculate the cache rate using the proportion of the time steps of the cached noise among all time steps, and set the target cache rate β.
[0019] In a specific embodiment of the present invention, the above-mentioned cache time step selection step further includes:
[0020] Set the threshold α of the relative change amount to screen the time steps using the noise cache;
[0021] Starting from the second time step, iterate through the entire time step sequence, check whether the relative change of the noise is less than the threshold of the relative change amount. If so, mark the noise in the current step as reusing the noise cached in the previous step; if not, the diffusion model will predict new noise;
[0022] Adjust the threshold α of the relative change amount to obtain the target cache rate.
[0023] In a specific embodiment of the present invention, the above-mentioned cache time step selection step further includes:
[0024] By continuously adjusting the threshold α of the relative change amount to traverse the relative change of the noise at all time steps, update the selection cache status of each time step through the judgment condition, and calculate the current cache rate until the target cache rate β is reached;
[0025] Calculate the maximum continuous caching step δ, where β is the target caching rate; the strategy for selecting δ is to make the maximum continuous caching step as small as possible while ensuring uniform caching;
[0026] When the continuous caching step exceeds the maximum continuous caching step δ, the diffusion model re-predicts the noise.
[0027] In the specific embodiment of the present invention, the above-mentioned mask-controlled denoising step includes:
[0028] During the denoising process of the diffusion model, each step checks the mask to determine whether to use the cached noise currently or predict new noise
[0029] In a second aspect, an embodiment of the present application provides an acceleration system for a diffusion model, which adopts the acceleration method of the diffusion model as described above. The system includes:
[0030] Noise generation module: Generate a batch of image data through the diffusion model and obtain the noise mean value of each step of the image data;
[0031] Noise relative change amount calculation module: Calculate the relative change amount of adjacent-step noise based on the noise tensors of two consecutive time steps, and evaluate the similarity of adjacent noise based on the relative change amount;
[0032] Caching time step selection module: According to the similarity of adjacent noise, formulate a strategy to screen the caching time steps of reusable noise and generate corresponding masks; for the selected caching time steps, directly reuse the noise cached in the previous step;
[0033] Mask-controlled denoising module: Use the mask to control the noise caching of each step during the denoising process.
[0034] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the acceleration method of the diffusion model are implemented.
[0035] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the acceleration method of the diffusion model as described above are implemented.
[0036] Compared with the related prior art, it has the following prominent beneficial effects:
[0037] 1) The method of the present invention proposes a noise caching mechanism; by caching and reusing noise, the number of noise prediction times in the denoising process of the diffusion model is reduced, thereby reducing the time overhead.
[0038] 2) The method of the present invention proposes a noise similarity measurement algorithm based on relative change; by calculating the relative change of adjacent noises to select the time step for using noise caching, so as to minimize the performance loss caused by caching.
[0039] 3) The method of the present invention proposes a mask control denoising process mechanism; according to the time step selection result, a corresponding mask is generated to control the noise source in the denoising process, and almost no additional time overhead is brought. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0041] Figure 1 Schematic diagram of the acceleration method of the diffusion model of the present invention Figure 1 ;
[0042] Figure 2 Schematic diagram of the acceleration method of the diffusion model of the present invention Figure 2 ;
[0043] Figure 3 Schematic diagram of the acceleration system of the diffusion model of the present invention;
[0044] Figure 4 Schematic diagram of the computer hardware of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0046] It should also be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0047] It should also be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0048] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0049] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0050] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0051] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, and other media that can store program codes.
[0052] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below and detailed descriptions are provided in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are merely for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0053] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0054] The method of the present invention aims to propose a framework for accelerating the diffusion model by caching noise. This framework controls the noise source at each time step during the denoising process through masking, and can achieve significant acceleration without introducing additional overhead. Compared with other current acceleration methods, the present invention not only has a simple deployment, but also can ensure better generation effects while accelerating the diffusion model.
[0055] When deploying the diffusion model, the present invention observes the changes in images and noise at different denoising steps, and finds that fewer denoising steps will cause more obvious changes in the image at each time step, resulting in a less smooth denoising process, which usually leads to distortion of the finally generated image. In a smoother denoising process, the noise between adjacent time steps usually shows a high degree of similarity. The present invention believes that this indicates that the quality of the generated image in the diffusion model depends on maintaining a sufficient number of sampling steps, and frequent noise updates may not be that important. Based on this discovery, the present invention proposes a framework for accelerating the diffusion model by caching noise. This framework evaluates the similarity by quantifying the relative change in noise between adjacent steps, and thus formulates a strategy to select the time steps where the noise can be reused. For these selected time steps, the noise prediction is bypassed by directly reusing the noise cached in the previous step. After evaluating the noise caching for all time steps, a masking mechanism is used to control the noise source at each step during the denoising process. This also means that this framework can be combined with acceleration methods that retain a certain number of sampling steps. Since the latency caused by noise prediction occupies the vast majority of the total latency of each time step, this caching mechanism can provide significant acceleration for the diffusion model.
[0056] Embodiment 1
[0057] As Figure 1 and Figure 2As shown, the embodiment of the present application provides an acceleration method for a diffusion model. First, a batch of data is generated through the diffusion model to obtain the intermediate noise at each step, then the relative change amount of the adjacent-step noise is calculated, and then a strategy is formulated based on the relative change amount to screen the noise cache time steps and generate corresponding masks. Finally, the masks are used to control the noise cache at each step in the denoising process. The method includes:
[0058] Noise generation step 101: Generate a batch of image data through the diffusion model, and obtain the noise mean value at each step of the image data;
[0059] Noise relative change amount calculation step 102: Calculate the relative change amount of the adjacent-step noise according to the noise tensors of two consecutive time steps, and evaluate the adjacent noise similarity based on the relative change amount;
[0060] Cache time step selection step 103: According to the adjacent noise similarity, formulate a strategy to screen the cache time steps of reusable noise and generate corresponding masks; for the selected cache time steps, directly reuse the noise cached in the previous step;
[0061] Mask control denoising step 104: Use the mask to control the noise cache at each step in the denoising process.
[0062] In a specific embodiment of the present invention, the above noise generation step 101 includes:
[0063] For the denoising process of multiple time steps, generate a batch of image data through the diffusion model for multiple time steps, record the noise at each time step, and calculate the average value.
[0064] In a specific embodiment of the present invention, for the denoising process of T steps, the present invention first generates a batch of data through the diffusion model for T steps, records the noise at each time step, and calculates the average value to reflect the expected noise level in the denoising process of the diffusion model.
[0065] In a specific embodiment of the present invention, the above noise relative change amount calculation step 102 includes:
[0066] Adopt the noise relative change amount L t to reflect the noise similarity degree, L t The calculation formula is as follows:
[0067]
[0068] where, ∈ t and ∈ t-1 are the noise tensors at consecutive time steps t and t - 1 respectively, ||.||1 represents the L1 norm, and the smaller the value of Lt, the smaller the difference between the two noises, and the more similar the two noises are.
[0069] In the specific embodiments of the present invention, it is observed that for diffusion models on different architectures, the relative change in noise is very small for most time steps, while significant relative changes are concentrated in the last few time steps. Therefore, our goal is to cache the results of early noise predictions while retaining the predictions for later noise.
[0070] In the specific embodiments of the present invention, the above-mentioned cache time step selection step 103 includes:
[0071] In the specific embodiments of the present invention, it is desired that the diffusion model predicts noise only at specific time steps, and at other time steps, it reuses the cached noise from the previous prediction step. Therefore, the present invention uses three variables to specifically select the time steps at which cached noise will be used in the denoising process.
[0072] Cache rate: We use β to represent the cache rate, which represents the proportion of time steps using cached noise among all time steps. A higher cache rate means fewer noise prediction times. In the present invention, the cache ratio needs to be set in advance.
[0073] Relative change threshold: We set a relative change threshold α ∈ (0, 1) to screen the time steps using noise caching. Starting from the second time step, iterate through the entire time step sequence T to check whether the relative change in noise is less than α. If so, mark the noise in this step as reusing the cached noise from the previous step. If not, the model will predict new noise. Adjusting the value of α can obtain the desired cache rate.
[0074] Maximum consecutive cache steps: The inventors found that using only α to select cache time steps may cause multiple consecutive steps to reuse cached noise, which can lead to the accumulation of noise differences and more serious distortion of the generated images. To alleviate this problem, the maximum consecutive cache steps δ are introduced. When the consecutive cache steps exceed δ, regardless of the current relative change in noise, the model needs to re-predict the noise. The strategy for the inventors to select δ is to make the maximum consecutive cache steps as small as possible without causing uniform caching. Therefore, it is necessary to calculate the length of uniform caching based on the target cache rate β, and then set δ to be just higher than this length:
[0075]
[0076] Calculate δ using the preset β, and then continuously adjust α to traverse the relative change in noise at all time steps. Update the cache status of each time step by judging whether the conditions of the relative change threshold and the maximum consecutive cache steps are met, and calculate the current cache rate until the target cache rate is reached.
[0077] In the specific embodiments of the present invention, the above-mentioned mask control denoising step 104 includes:
[0078] The caching steps selected according to the above criteria generate a mask, and then during the denoising process of the diffusion model, each step checks the mask to determine whether to use the cached noise or predict new noise at the current step.
[0079] Embodiment 2
[0080] As Figure 3 shown, an embodiment of the present application provides an acceleration system for a diffusion model, which adopts the acceleration method of the diffusion model as described above. The system includes:
[0081] Noise generation module 201: Generate a batch of image data through the diffusion model, and obtain the noise mean value of each step of the image data;
[0082] Noise relative change amount calculation module 202: Calculate the relative change amount of adjacent-step noise according to the noise tensors of two consecutive time steps, and evaluate the similarity of adjacent noises based on the relative change amount;
[0083] Caching time step selection module 203: According to the similarity of adjacent noises, formulate a strategy to screen the caching time steps of reusable noises and generate corresponding masks; for the selected caching time steps, directly reuse the noises cached in the previous step;
[0084] Mask control denoising module 204: Use the mask to control the noise caching of each step during the denoising process.
[0085] Embodiment 3
[0086] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the acceleration method of the diffusion model are implemented.
[0087] Embodiment 4
[0088] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the acceleration method of the diffusion model as described are implemented.
[0089] In addition, the self-supervised video representation learning method based on temporal correspondence described in Figure 1 the embodiments of the present application can be implemented by an electronic device, such as a computer device. Figure 4 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.
[0090] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 4 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.
[0091] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.
[0092] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.
[0093] By reading and executing the computer program instructions stored in the memory 82, the processor 81 implements the acceleration method of any one of the diffusion models in the above embodiments.
[0094] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0095] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for accelerating a diffusion model, characterized in that: The method comprises: Noise generation step: generating a batch of image data through a diffusion model, and obtaining the noise mean of each step of the image data; The step of calculating the relative noise variation: calculating the relative noise variation of adjacent steps according to the noise tensors of two consecutive time steps, and evaluating the adjacent noise similarity based on the relative noise variation; Cache time step selection step: according to the similarity of the adjacent noises, formulate a strategy to filter the cache time steps of the reusable noises, and generate a corresponding mask; for the selected cache time steps, directly reuse the noise cached in the previous step; Mask controls denoising steps: Use the mask to control the noise buffer at each step in the denoising process.
2. The method for accelerating the diffusion model according to claim 1, characterized in that: The noise generating step comprises: For the denoising process of multiple time steps, a batch of image data is generated by the diffusion model at multiple time steps, and the noise of each time step is recorded and averaged.
3. The method for accelerating the diffusion model according to claim 1, characterized in that: The step of calculating the relative noise variation comprises: The relative noise change L t Reaction noise similarity, L t The calculation formula is as follows: Among them, ∈ t and ∈ t-1 are the noise tensors at consecutive time steps t and t-1, respectively. ||.||1 represents the L1 norm. The smaller the Lt value, the smaller the difference between the two noises, and the more similar the two noises are.
4. The method for accelerating the diffusion model according to claim 1, characterized in that: The cache time step selection step comprises: The cache rate is calculated using the proportion of cached noise time steps in all time steps, and the target cache rate β is set.
5. The method for accelerating the diffusion model according to claim 4, characterized in that: The cache time step selection step further includes: Set the threshold α of the relative change amount to filter the time step of the noise cache; Starting from the second time step, iterate the entire time step sequence to check whether the relative change of the noise is less than the threshold of the relative change amount. If so, mark the noise in the current step as reusing the cached noise in the previous step; if not, the diffusion model predicts new noise; The target cache rate is obtained by adjusting the threshold α of the relative change amount.
6. The method for accelerating the diffusion model according to claim 5, characterized in that: The cache time step selection step further includes: By continuously adjusting the threshold α of the relative change amount to traverse the relative change of noise at all time steps, the selected cache state of each time step is updated according to the judgment conditions, and the current cache rate is calculated until the target cache rate β is reached; Calculate the maximum number of consecutive cache steps δ, Among them, β is the target cache rate; the strategy for selecting δ is to make the maximum number of consecutive cache steps as small as possible while ensuring uniform caching; When the number of consecutive cache steps exceeds the maximum number of consecutive cache steps δ, the diffusion model re-predicts the noise.
7. The method for accelerating the diffusion model according to claim 1, characterized in that: The mask control denoising step comprises: During the denoising process of the diffusion model, each step checks the mask to determine whether to use the current cached noise or predict new noise.
8. An acceleration system for a diffusion model, using the acceleration method for a diffusion model as claimed in any one of claims 1 to 7, characterized in that: The system comprises: Noise generation module: generates a batch of image data through a diffusion model, and obtains the noise mean of each step of the image data; Noise relative change calculation module: calculates the relative change of noise in adjacent steps according to the noise tensor of two consecutive time steps, and evaluates the similarity of adjacent noise based on the relative change; Cache time step selection module: formulate a strategy to filter cache time steps of reusable noise according to the similarity of adjacent noises, and generate corresponding masks; for the selected cache time step, directly reuse the noise cached in the previous step; Mask-controlled denoising module: Use the mask to control the noise buffer at each step in the denoising process.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the acceleration method of the diffusion model described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for accelerating the diffusion model according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Multi-modal data generation method and device, medium, equipment and program product
CN121580318A