Diffusion model sampling method and device, equipment and storage medium
By combining the multi-step solution strategies of PNDM and DPM-Solver and introducing a noise correction mechanism, the sampling process of the diffusion model is optimized, and the insufficient sampling speed and the trade-off between generation quality and speed are solved, achieving more efficient sample generation.
Patent Information
- Application Number
- CN202510272014.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-13
AI Technical Summary
The sampling technology of existing diffusion models has insufficient sampling speed, difficulty in weighing the quality and speed of generation, and lack of adaptability, which limits its potential in real-time generation, complex scenario generation, and large-scale applications.
By combining PNDM and DPM-Solver's multi-step solution strategy to optimize the reverse diffusion process, introduce a noise correction mechanism, use the accumulated noise of the historical time step to correct the noise prediction results of the current time step, and update the sample of the current time step based on the corrected noise amount to generate the results.
The noise accumulation error at each time step is reduced, the stability and consistency of sample generation is improved, the sample generation accuracy is maintained, and the calculation amount required for each time step is reduced, which speeds up the sampling speed of the diffusion model.
Smart Images

Figure CN120147779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and particularly to a sampling method, device, equipment and storage medium for a diffusion model. Background Art
[0002] In recent years, with the wide application of image generation models, diffusion models have made remarkable progress in the fields of image generation, audio synthesis, etc. A diffusion model is a generative model that uses the idea of physical thermodynamics diffusion. By gradually iterating, a correspondence between the Gaussian distribution and the data distribution is established. The final target sample is generated by forward denoising and backward denoising, and it has good performance in terms of the detail accuracy and richness of the sample generation results. Existing diffusion model sampling schemes mainly include PNDM (Pseudo Numerical Methods for Diffusion Models) and DPM-Solver (Denoising Diffusion Probabilistic Models Solver). Among them, PNDM is based on simple noise interpolation and step-by-step backward diffusion, which has high accuracy but slow speed. Especially when dealing with more complex samples, each step of the calculation is relatively cumbersome; DPM uses a step-by-step denoising process based on Markov chains, which has a faster speed, but in some cases, the quality of the generated samples will be sacrificed, especially when dealing with complex noise. Therefore, although the existing sampling schemes for diffusion models have made progress in sampling efficiency and quality to a certain extent, there are still deficiencies such as insufficient sampling speed, difficulty in balancing generation quality and speed, and lack of adaptability, which limit the potential of diffusion models in real-time generation, complex scene generation, and large-scale applications. Summary of the Invention
[0003] The present invention provides a sampling method, device, equipment and storage medium for a diffusion model to solve the technical problems of insufficient sampling speed, difficulty in balancing generation quality and speed, and lack of adaptability existing in the existing sampling technology of diffusion models.
[0004] In a first aspect, a sampling method for a diffusion model is provided, including:
[0005] Inputting a noise picture to be sampled into a trained diffusion model for sample generation, and predicting the noise amount at the current time step during the sample generation process;
[0006] Calculating the sum of the noise amounts at the first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps before;
[0007] Correct the amount of noise at the current time step according to the cumulative noise of the first number of previous historical time steps;
[0008] Update the sample generation result of the current time step according to the corrected amount of noise, and when reaching the last time step, output the finally generated target sample.
[0009] In a second aspect, a sampling device for a diffusion model is provided, including:
[0010] A noise prediction module: configured to input a noise picture to be sampled into a trained diffusion model for sample generation, and predict the amount of noise at the current time step during the sample generation process;
[0011] An accumulated noise calculation module: configured to calculate the total amount of noise of the first number of previous historical time steps of the current time step to obtain the cumulative noise of the first number of previous historical time steps;
[0012] A noise correction module: configured to correct the amount of noise at the current time step according to the cumulative noise of the first number of previous historical time steps;
[0013] A sample update module: configured to update the sample generation result of the current time step according to the corrected amount of noise, and when reaching the last time step, output the finally generated target sample.
[0014] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the sampling method of the above diffusion model are implemented.
[0015] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the sampling method of the above diffusion model are implemented.
[0016] In the solutions implemented by the above sampling method, device, computer device, and storage medium of the diffusion model, by combining the multi-step solution strategies of PNDM and DPM-Solver to optimize the reverse diffusion process, and by introducing a noise correction mechanism, using the cumulative noise of historical time steps to correct the noise prediction result of the current time step, and updating the sample generation result of the current time step according to the corrected amount of noise, the noise accumulation error of each time step is reduced, and the amount of calculation required for each time step can be reduced while maintaining the sample generation accuracy, thereby accelerating the sampling speed of the diffusion model. Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 is a schematic diagram of the architecture of a sampling system of a diffusion model in an embodiment of the present invention;
[0019] Figure 2 is a schematic flowchart of a sampling method of a diffusion model in the first embodiment of the present invention;
[0020] Figure 3 is a schematic flowchart of a sampling method of a diffusion model in the second embodiment of the present invention;
[0021] Figure 4 is a schematic structural diagram of a sampling device of a diffusion model in an embodiment of the present invention;
[0022] Figure 5 is a schematic structural diagram of a computer device in an embodiment of the present invention;
[0023] Figure 6 is another schematic structural diagram of a computer device in an embodiment of the present invention. Specific Embodiments
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0025] The sampling method of the diffusion model provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client communicates with the server through the network. The server can input the noise image to be sampled into the trained diffusion model for sample generation, and predict the amount of noise at the current time step during the sample generation process; calculate the sum of the amounts of noise at the first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps; correct the amount of noise at the current time step according to the cumulative noise of the first number of historical time steps; update the sample generation result at the current time step according to the corrected amount of noise, and when reaching the last time step, output the finally generated target sample and return the generated target sample to the client. In the present invention, in the insurance and banking and other businesses in the financial field, the above sampling method of the diffusion model can be used to generate customized personalized avatars and background images for users, which speeds up the generation speed of the diffusion model and improves the generation quality. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0026] Please refer to Figure 2 , which is a schematic flowchart of the sampling method of the diffusion model in the first embodiment of the present invention. The sampling method of the diffusion model provided by the first embodiment of the present invention includes the following steps:
[0027] S100: Input the noise image to be sampled into the trained diffusion model for sample generation, and predict the amount of noise at the current time step during the sample generation process;
[0028] In this step, in the diffusion model, the key to denoising lies in predicting the noise in the noise image. At each time step, the denoising network in the diffusion model will estimate the noise at that time step according to the input noise image and the current time step. Specifically, input the noise image at the current time step and the index of the current time step into the denoising network of the diffusion model. The denoising network sets the index of the current time step as the progress information of the denoising process and outputs a noise prediction value as the predicted value of the amount of noise at the current time step by the diffusion model. Among them, the noise image can be a certain noise version of the original image or an image under specific conditions (such as an image with a skeleton diagram or an edge diagram). Specifically, the prediction formula for the amount of noise at the current time step is:
[0029] ∈ t = Unet(x t ) (1)
[0030] Where ∈ t represents the predicted value of the amount of noise at the current time step t, Unet represents the diffusion model, and x tDenote the sample generated at the current time step t.
[0031] S110: Calculate the sum of the noise amounts of the first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps.
[0032] In this step, the cumulative noise is the total sum of the noise added / removed in all historical time steps, which is calculated through step-by-step iteration. The noise prediction not only depends on the noise image of the current time step but also helps reduce the error of noise estimation by introducing the noise information of historical time steps, especially in the case of strong noise or unstable denoising effect. Among them, the value of the first number K can be set according to the application scenario.
[0033] Specifically, the calculation process of the cumulative noise includes:
[0034] Obtain the first number of historical time steps before the current time step and the corresponding noise scheduling parameters for each historical time step.
[0035] According to the preset cumulative noise calculation formula, perform a cumulative multiplication calculation on the noise scheduling parameters corresponding to each historical time step to obtain the cumulative noise weight of the first number of historical time steps. Among them, the preset cumulative noise calculation formula is:
[0036]
[0037] Among them, represents the cumulative noise weight from time step 1 to time step t, and β t represents the noise scheduling parameter of time step t.
[0038] S120: Correct the noise amount of the current time step according to the cumulative noise of the first number of historical time steps.
[0039] In this step, at each time step of sample generation, not only the noise amount of the current time step is considered, but also a historical noise correction mechanism is introduced, that is, the predicted value of the noise amount of the current time step is integrated with the weighted information of the cumulative noise of the first number of historical time steps to further correct the predicted value of the noise amount of the current time step. The image sample of the current time step is generated using the corrected noise amount to continue the denoising process of the next time step and improve the denoising effect.
[0040] Specifically, the noise amount correction algorithm for the current time step includes:
[0041] S121: Obtain the number of steps of the historical time steps, as well as the noise estimation value and the corresponding first weight coefficient at each historical time step.
[0042] S122: Loop according to the number of steps in historical time steps, multiply the noise estimation value at each historical time step by the first weight coefficient, and accumulate the multiplication results to obtain a corrected noise prediction value; where the corrected noise prediction value is expressed as:
[0043]
[0044] where, ∈ pred represents the corrected noise prediction value, ω k represents the first weight coefficient, ∈ t―k represents the noise estimation value at the historical time step, and K represents the number of steps in the historical time step.
[0045] S123: Calculate the first difference between the corrected noise prediction value and the noise amount prediction value at the current time step;
[0046] S124: Correct the first difference through a preset first correction coefficient, and add the corrected first difference to the noise amount prediction value at the current time step to obtain the corrected noise amount at the current time step; where the corrected noise amount at the current time step is expressed as:
[0047] ∈′ t = ∈ t + α(∈ pred ― ∈ t ) (4)
[0048] where, ∈′ t represents the corrected noise amount at the current time step, and α is the correction coefficient.
[0049] S130: Update the sample generation result at the current time step according to the corrected noise amount, and output the finally generated target sample when reaching the last time step;
[0050] In this step, the sample update process at the current time step is specifically as follows:
[0051] S131: Use the corrected noise amount to remove the noise information in the sample generated at the current time step;
[0052] S132: Update the sample generation result at the current time step according to the result after removing the noise information, the preset random noise, the preset noise standard deviation, the noise scheduling parameter, and the cumulative noise weight; where the sample update formula is:
[0053]
[0054] where, x t―1 represents the sample generated at time step t - 1, is the random noise, and σ t is the noise standard deviation.
[0055] S133: Adopt a step-by-step update mechanism to optimize the sample generation result of the current time step according to the predicted noise amount and denoising weight of each time step;
[0056] S134: When reaching the last time step, calculate the finally generated target sample based on the sample generated at time step 1, the noise weight of time step 1, and the predicted noise amount of time step 1; wherein, the finally generated target sample is:
[0057]
[0058] wherein, x 0 represents the finally generated target sample, x 1 represents the sample generated at time step 1, represents the noise weight of time step 1, ∈ 1 represents the predicted noise amount of time step 1.
[0059] Based on the above, the present invention introduces a historical noise correction mechanism, uses the weighted average of the cumulative noise of historical time steps to correct the noise prediction result of the current time step, and updates the sample generation result of the current time step according to the corrected noise amount, thereby reducing the error accumulation in each time step and improving the stability and consistency of sample generation. By using the step-by-step update mechanism, the diffusion model can effectively recover samples with rich details and high quality from noise, wherein the number of iterations N can be set according to the actual application scenario.
[0060] It can be seen that in the above solution, the sampling method of the diffusion model provided by the embodiment of the present invention optimizes the reverse diffusion process by combining the multi-step solution strategies of PNDM and DPM-Solver. By introducing a noise correction mechanism, it uses the cumulative noise of historical time steps to correct the noise prediction result of the current time step, and updates the sample generation result of the current time step according to the corrected noise amount, thereby reducing the noise accumulation error in each time step. It can reduce the calculation amount required for each time step while maintaining the sample generation accuracy, and accelerate the sampling speed of the diffusion model.
[0061] Please refer to Figure 3 , which is a schematic flowchart of the sampling method of the diffusion model in the second embodiment of the present invention. The sampling method of the diffusion model provided by the second embodiment of the present invention includes the following steps:
[0062] S200: Set the initialization parameters of the diffusion model;
[0063] In this step, the diffusion model adopts the U-net model commonly used in existing generative models. The set initialization parameters include the total number of time steps T of the time step and the noise scheduling parameter β tand initial noise
[0064] S210: Pre-train the diffusion model with preset training sample data and initialization parameters, and optimize the initialization parameters according to the pre-training results;
[0065] In this step, the training process of the diffusion model includes a forward diffusion process and a reverse diffusion process. In the forward diffusion process, noise is added to a picture without any noise. The reverse diffusion process needs to start from a picture full of noise and gradually reduce the noise until the target sample without any noise is finally generated.
[0066] S220: Input the noise picture to be sampled into the trained diffusion model for sample generation, and predict the noise amount at the current time step during the sample generation process;
[0067] In this step, the prediction method of the noise amount at the current time step is the same as that in S100 of the first embodiment. To avoid redundancy, it will not be elaborated here.
[0068] S230: Calculate the sum of the noise amounts at the first number of historical time steps before the current time step to obtain the cumulative noise at the first number of historical time steps before the current time step;
[0069] In this step, the calculation method of the cumulative noise is the same as that in S110 of the first embodiment. To avoid redundancy, it will not be elaborated here.
[0070] S240: Correct the noise amount at the current time step according to the cumulative noise at the first number of historical time steps before the current time step;
[0071] In this step, the correction algorithm of the noise amount at the previous time step is the same as that in S120 of the first embodiment. To avoid redundancy, it will not be elaborated here.
[0072] S250: Update the sample generation result at the current time step according to the corrected noise amount;
[0073] In this step, the sample update method at the current time step is the same as that in S130 of the first embodiment. To avoid redundancy, it will not be elaborated here.
[0074] S260: Adjust the step size of the next time step according to the noise level at the current time step, and output the finally generated target sample when reaching the last time step:
[0075] In this step, in the diffusion model, the magnitude of the noise intensity determines the step size for updating the sample each time. In the early stage of the sample generation process, the noise intensity in the sample is large, and the model will select a larger step size to update the sample, thereby quickly removing the noise and increasing the denoising speed. Conversely, when entering the later stage of the sample generation process, the noise intensity in the sample is small. At this time, the step size for the next time step is reduced according to the noise level at the current time step to update the sample in a more refined manner. By this way of dynamically adjusting the step size, the model can adopt different denoising strategies at different denoising stages, making the sampling at each time step more delicate and avoiding over-denoising that affects the quality of sample generation. Among them, the step size adjustment parameter for the next time step can be set according to the actual application scenario. Specifically, the step size adjustment process is as follows:
[0076] Judge whether the predicted value of the noise amount at the current time step is greater than the preset noise threshold τ. If the predicted value of the noise amount at the current time step is greater than the noise threshold τ, adjust the step size for the next time step according to the preset large step size adjustment parameter;
[0077] If the predicted value of the noise amount at the current time step is less than or equal to the noise threshold τ, adjust the step size for the next time step according to the preset small step size adjustment parameter;
[0078] Among them, the step size adjustment formula is:
[0079]
[0080] Among them, Δ t represents the step size at the current time step, τ is the adjustable noise threshold, h large and h small are the large step size adjustment parameter and the small step size adjustment parameter respectively.
[0081] Based on the above, the present invention adjusts the step size for the next time step in real time according to the noise level at the current time step to accelerate the sampling process in the early stage of the sample generation process and perform more delicate sampling in the later stage of the sample generation process, thereby improving the sampling speed while ensuring the quality of sample generation.
[0082] It can be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0083] In one embodiment, a sampling device for a diffusion model is provided. The sampling device for the diffusion model corresponds one-to-one to the sampling method of the diffusion model in the above embodiment. As Figure 4As shown in the figure, the sampling device of the diffusion model includes a model parameter setting module 101, a pre-training module 102, a noise prediction module 103, an accumulated noise calculation module 104, a noise correction module 105, a sample update module 106, and a step size adjustment module 107. The detailed description of each functional module is as follows:
[0084] Parameter setting module 101: It is used to set the initialization parameters of the diffusion model. Among them, the diffusion model adopts the U-net model commonly used in existing generative models. The set initialization parameters include the total number of time steps T of the time step, the noise scheduling parameter β t and the initial noise
[0085] Pre-training module 102: It is used to pre-train the diffusion model with preset training sample data and initialization parameters, and optimize the initialization parameters according to the pre-training results. Among them, the training process of the diffusion model includes a forward diffusion process and a reverse diffusion process. In the forward diffusion process, noise is added to a picture without any noise. The reverse diffusion process needs to start from a picture full of noise and gradually reduce the noise until finally generating a target sample without any noise.
[0086] Noise prediction module 103: It is used to input the noise picture to be sampled into the trained diffusion model for sample generation, and predict the amount of noise at the current time step during the sample generation process. Among them, in the diffusion model, the key to denoising lies in predicting the noise in the noise picture. At each time step, the denoising network in the diffusion model will estimate the noise at that time step based on the input noise picture and the current time step. Specifically, the noise picture at the current time step and the index of the current time step are input into the denoising network of the diffusion model. The index of the current time step is set as the progress information of the denoising process through the denoising network, and a noise prediction value is output as the predicted value of the amount of noise at the current time step by the diffusion model. Among them, the noise picture can be a certain noise version of the original image, or an image under specific conditions (such as an image with a skeleton diagram or an edge diagram). Specifically, the prediction formula for the amount of noise at the current time step is:
[0087] ∈ t =Unet(x t ) (1)
[0088] where, where, ∈ t represents the predicted value of the amount of noise at the current time step t, Unet represents the denoising network, and x t represents the sample generated at the current time step t.
[0089] Cumulative Noise Calculation Module 104: It is used to calculate the sum of the noise amounts of the first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps; among them, the cumulative noise is the sum of the noises added / denoised in all historical time steps, which is obtained through step-by-step iterative calculation. Noise prediction not only depends on the noise image of the current time step, but also helps to reduce the error of noise estimation by introducing the noise information of historical time steps, especially in the case of strong noise or unstable denoising effect. Among them, the value of the first number K can be set according to the application scenario.
[0090] Specifically, the calculation process of the cumulative noise includes:
[0091] Obtain the first number of historical time steps before the current time step, and the corresponding noise scheduling parameters for each historical time step;
[0092] According to the preset cumulative noise calculation formula, perform a cumulative multiplication calculation on the corresponding noise scheduling parameters for each historical time step to obtain the cumulative noise weight of the first number of historical time steps; among them, the preset cumulative noise calculation formula is:
[0093]
[0094] Among them, represents the cumulative noise weight from time step 1 to time step t, and β t represents the noise scheduling parameter of time step t.
[0095] Noise Correction Module 105: It is used to correct the noise amount of the current time step according to the cumulative noise of the first number of historical time steps before; among them, at each time step of sample generation, not only the noise amount of the current time step is considered, but also a historical noise correction mechanism is introduced, that is, the predicted value of the noise amount of the current time step is integrated with the weighted information of the cumulative noise of the first number of historical time steps to further correct the predicted value of the noise amount of the current time step, and the corrected noise amount is used to generate the image sample of the current time step for continued denoising in the next time step to improve the denoising effect.
[0096] Specifically, the noise amount correction algorithm for the current time step includes:
[0097] Obtain the number of steps of the historical time steps, as well as the noise estimation value and the corresponding first weight coefficient at each historical time step;
[0098] Loop according to the number of steps of the historical time steps, multiply the noise estimation value and the first weight coefficient at each historical time step, and accumulate the multiplication results to obtain the corrected noise prediction value; among them, the corrected noise prediction value is expressed as:
[0099]
[0100] Among them, ∈ pred represents the corrected noise prediction value, ω k represents the first weight coefficient, ∈ t―k represents the noise estimation value at the historical time step, and K represents the number of steps at the historical time step.
[0101] Calculate the first difference between the corrected noise prediction value and the noise amount prediction value at the current time step;
[0102] Correct the first difference through a preset first correction coefficient, and add the corrected first difference to the noise amount prediction value at the current time step to obtain the corrected noise amount at the current time step; among them, the corrected noise amount at the current time step is expressed as:
[0103] ∈′ t = ∈ t + α(∈ pred ― ∈ t ) (4)
[0104] Among them, ∈′ t represents the corrected noise amount at the current time step, and α is the correction coefficient.
[0105] Sample update module 106: used to update the sample generation result at the current time step according to the corrected noise amount; among them, the sample update process at the current time step is specifically:
[0106] Use the corrected noise amount to remove the noise information in the sample generated at the current time step;
[0107] Update the sample generation result at the current time step according to the result after removing the noise information, the preset random noise, the preset noise standard deviation, the noise scheduling parameter, and the cumulative noise weight; among them, the sample update formula is:
[0108]
[0109] Among them, x t―1 represents the sample generated at time step t-1, is the random noise, σ t is the noise standard deviation.
[0110] Adopt a step-by-step update mechanism to optimize the sample generation result at the current time step according to the noise amount prediction value and the denoising weight at each time step; through the step-by-step update mechanism, the diffusion model can effectively recover detailed and high-quality samples from the noise, where the number of iterations N can be set according to the actual application scenario. Specifically, the sample update formula is:
[0111]
[0112] Among them, x t―1 represents the sample generated at time step t - 1, is random noise, and σ t is the standard deviation of the noise.
[0113] Based on the above, the present invention introduces a historical noise correction mechanism, uses the weighted average of the cumulative noise in historical time steps to correct the noise prediction result of the current time step, and updates the sample generation result of the current time step according to the corrected noise amount, thereby reducing the error accumulation in each time step and improving the stability and consistency of sample generation.
[0114] Step size adjustment module 107: It is used to adjust the step size of the next time step according to the noise level of the current time step, and output the finally generated target sample when reaching the last time step. Among them, in the diffusion model, the magnitude of the noise intensity determines the step size for updating the sample each time. In the early stage of the sample generation process, the noise intensity in the sample is large, so the model will select a larger step size for sample update to quickly remove the noise and improve the denoising speed. On the contrary, when entering the later stage of the sample generation process, the noise intensity in the sample is small. At this time, the step size of the next time step is reduced according to the noise level of the current time step to update the sample in a more refined manner. By this way of dynamically adjusting the step size, the model can adopt different denoising strategies at different denoising stages, making the sampling of each time step more delicate and avoiding excessive denoising from affecting the quality of sample generation. Among them, the step size adjustment parameter of the next time step can be set according to the actual application scenario. Specifically, the step size adjustment process is as follows:
[0115] Judge whether the predicted value of the noise amount at the current time step is greater than the preset noise threshold τ. If the predicted value of the noise amount at the current time step is greater than the noise threshold τ, then adjust the step size of the next time step according to the preset large step size adjustment parameter;
[0116] If the predicted value of the noise amount at the current time step is less than or equal to the noise threshold τ, then adjust the step size of the next time step according to the preset small step size adjustment parameter;
[0117] Among them, the step size adjustment formula is:
[0118]
[0119] Among them, Δ t represents the step size of the current time step, τ is an adjustable noise threshold, h large and h small are the large step size adjustment parameter and the small step size adjustment parameter respectively.
[0120] When reaching the last time step (i.e., time step 1), the finally generated target sample is:
[0121]
[0122] Among them, x 0 represents the finally generated target sample, and x 1 represents the sample generated at time step 1, α 1 represents the noise weight at time step 1, and ∈ 1 represents the predicted value of the noise amount at time step 1.
[0123] Based on the above, the present invention adjusts the step size of the next time step in real time according to the noise level of the current time step, so as to accelerate the sampling process in the early stage of the sample generation process and perform more delicate sampling in the later stage of the sample generation process, thereby improving the sampling speed while ensuring the sample generation quality.
[0124] It can be seen that in the above solution, the sampling device of the diffusion model provided by the embodiment of the present invention optimizes the reverse diffusion process by combining the multi-step solution strategies of PNDM and DPM-Solver. By introducing a noise correction mechanism, the cumulative noise of historical time steps is used to correct the noise prediction result of the current time step, and the sample generation result of the current time step is updated according to the corrected noise amount, thereby reducing the noise accumulation error of each time step. It can reduce the computational amount required for each time step and accelerate the sampling speed of the diffusion model while maintaining the sample generation accuracy. By adjusting the step size of the next time step in real time according to the noise level of the current time step, the sampling process is accelerated in the early stage of the sample generation process, and more delicate sampling is performed in the later stage of the sample generation process, thereby improving the sampling speed while ensuring the sample generation quality.
[0125] For the specific limitations of the sampling device of the diffusion model, reference can be made to the limitations of the sampling method of the diffusion model in the above text, which will not be elaborated here. Each module in the above sampling device of the diffusion model can be implemented in whole or in part by software, hardware and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0126] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 5As shown. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a sampling method of a diffusion model.
[0127] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 6 As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a sampling method of a diffusion model.
[0128] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0129] Input the noise image to be sampled into the trained diffusion model for sample generation, and predict the noise amount at the current time step during the sample generation process;
[0130] Calculate the sum of the noise amounts at the first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps before;
[0131] Correct the noise amount at the current time step according to the cumulative noise of the first number of historical time steps before;
[0132] Update the sample generation result at the current time step according to the corrected noise amount, and when reaching the last time step, output the finally generated target sample.
[0133] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented:
[0134] Input the noisy image to be sampled into the trained diffusion model for sample generation, and predict the amount of noise at the current time step during the sample generation process;
[0135] Calculate the sum of the amounts of noise at the first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps;
[0136] Correct the amount of noise at the current time step according to the cumulative noise of the first number of historical time steps before;
[0137] Update the sample generation result at the current time step according to the corrected amount of noise, and output the finally generated target sample when reaching the last time step.
[0138] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can achieve, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0139] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other storage medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0140] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0141] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A sampling method for a diffusion model, characterized in that: include: Input the noise image to be sampled into the trained diffusion model for sample generation, and predict the noise amount at the current time step during the sample generation process; Calculating the sum of noise amounts of a first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps before the current time step; Correcting the noise amount of the current time step according to the accumulated noise of the first number of historical time steps; The sample generation result of the current time step is updated according to the corrected noise amount, and when the last time step is reached, the finally generated target sample is output.
2. The sampling method of the diffusion model according to claim 1, characterized in that: The noise image to be sampled is input into the trained diffusion model to generate samples, and the noise amount at the current time step is predicted during the sample generation process, specifically: Inputting the noise image of the current time step and the index of the current time step into the denoising network of the diffusion model; Setting the index of the current time step as progress information of the denoising process through the denoising network; A noise prediction value is output as the noise amount prediction value of the diffusion model for the current time step.
3. The sampling method of the diffusion model according to claim 2, characterized in that: The calculation of the sum of the noise amounts of the first number of historical time steps before the current time step to obtain the cumulative noise of the first number of historical time steps before the current time step is specifically: Obtaining a first number of historical time steps before the current time step, and noise scheduling parameters corresponding to each of the historical time steps; According to a preset cumulative noise calculation formula, the noise scheduling parameters corresponding to each of the historical time steps are cumulatively multiplied to obtain the cumulative noise weights of the first first number of historical time steps.
4. The sampling method of the diffusion model according to claim 3, characterized in that: The noise amount of the current time step is corrected according to the accumulated noise of the first number of historical time steps, specifically: Obtaining the number of the historical time steps, as well as the noise estimation value and the corresponding first weight coefficient at each of the historical time steps; Looping according to the number of historical time steps, multiplying the noise estimation value and the first weight coefficient, and accumulating the results of the multiplication to obtain a corrected noise prediction value; Calculating a first difference between the corrected noise prediction value and the noise amount prediction value of the current time step; The first difference is corrected by presetting a first correction coefficient, and the corrected first difference is added to the noise amount prediction value of the current time step to obtain the corrected noise amount of the current time step.
5. The sampling method of the diffusion model according to claim 4, characterized in that: The sample generation result of the current time step is updated according to the corrected noise amount, and when the last time step is reached, the finally generated target sample is output, specifically: Using the corrected noise amount to remove noise information in samples generated at the current time step; Update the sample generation result of the current time step according to the result after removing the noise information, the preset random noise, the preset noise standard deviation, the noise scheduling parameter and the accumulated noise weight; Adopting a step-by-step updating mechanism to optimize the sample generation result according to the predicted value of the noise amount and the denoising weight at each time step; When the last time step is reached, the final target sample is calculated based on the sample generated at time step 1, the noise weight at time step 1, and the predicted value of the noise amount at time step 1.
6. The sampling method of the diffusion model according to any one of claims 1 to 5, characterized in that: The method further includes: inputting the noise image to be sampled into the trained diffusion model for sample generation, and predicting the noise amount at the current time step during the sample generation process. Setting initialization parameters of the diffusion model, wherein the initialization parameters include the total number of time steps, the noise scheduling parameters, and the initial noise; The diffusion model is pre-trained using preset training sample data and the initialization parameters, and the initialization parameters are optimized according to the pre-training results.
7. The sampling method of the diffusion model according to claim 6, characterized in that: After the sample generation result of the current time step is updated according to the corrected noise amount, the method further includes: Adjust the step size of the next time step according to the noise level of the current time step; The process of adjusting the step size of the next time step is: Determine whether the noise amount prediction value of the current time step is greater than a preset noise threshold, and if the noise amount prediction value of the current time step is greater than the noise threshold, adjust the step size of the next time step according to a preset large step size adjustment parameter; If the noise amount prediction value of the current time step is less than or equal to the noise threshold, the step size of the next time step is adjusted according to a preset small step size adjustment parameter.
8. A sampling device for a diffusion model, characterized in that: include: Noise prediction module: used to input the noise image to be sampled into the trained diffusion model for sample generation, and predict the noise amount at the current time step during the sample generation process; Cumulative noise calculation module: used to calculate the sum of the noise amounts of the first number of historical time steps before the current time step, and obtain the cumulative noise of the first number of historical time steps; Noise correction module: used for correcting the noise amount of the current time step according to the accumulated noise of the first number of historical time steps; Sample update module: used to update the sample generation result of the current time step according to the corrected noise amount, and output the final generated target sample when reaching the last time step.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the sampling method of the diffusion model according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the sampling method of the diffusion model according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Model adjustment method and device, video generation method and related equipment
CN120935381A