Diffusion model quantification method and electronic device
By grouping adjustment and parameter optimization of the time steps of the diffusion model, combined with low-rank adaptive LoRA algorithm and high-precision model guidance, the problem of cumulative error in diffusion model quantization is solved, and efficient quantization accuracy is achieved.
Patent Information
- Application Number
- CN202510407217.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
While the existing diffusion model quantization technology reduces computational costs and storage needs, it cannot effectively reduce cumulative errors, affecting the quality of output data.
By dividing the time steps of the diffusion model into multiple time step groups, weight adjustment and parameter fine-tuning are performed for each group, combined with teacher guidance from low-rank adaptive LoRA algorithm and high-precision models, quantized parameters are optimized to reduce errors.
Without significantly increasing the computational volume and storage requirements, the quantization accuracy of the diffusion model is improved, the cumulative error is reduced, and the output data quality is improved.
Smart Images

Figure CN120258070A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a quantization method in the field of deep learning, and more specifically, to a quantization method and an electronic device for a diffusion model. Background Art
[0002] Diffusion models have shown excellent performance in many fields that require the generation of high-quality samples. However, diffusion models have a huge computational cost in operation.
[0003] Post-training Quantization (PTQ) is a common model quantization algorithm that is applied after model training is completed. It aims to reduce the size of the model and improve the inference speed while maintaining the performance of the model as much as possible.
[0004] However, the performance of existing PTQ technology in the quantization of diffusion models does not meet expectations. Since the diffusion model needs to perform hundreds of denoising steps and the number of parameters of the diffusion model is huge, the cumulative error caused by the quantization of the diffusion model will seriously affect the quality of the output data of the diffusion model. Summary of the invention
[0005] In general, the present disclosure relates to a quantization method and an electronic device for a diffusion model. By grouping time steps based on the importance of the time steps and obtaining adjustment parameters for each time step group respectively, the cumulative error of the diffusion model caused by multiple denoising can be reduced. In addition, by further fine-tuning the adjustment parameters for two adjacent time step groups, the overfitting problem caused by multiple groups of adjustment parameters can be avoided. Therefore, the quantization accuracy of the diffusion model can be improved without introducing too much additional calculations and occupying less memory resources.
[0006] According to some embodiments, the present disclosure relates to a quantization method for a diffusion model, the quantization method comprising: quantizing a first diffusion model after training to obtain a second diffusion model; dividing the time steps of the second diffusion model into multiple time step groups; adjusting the weight of the second diffusion model for each time step group in the multiple time step groups to obtain a third diffusion model; using the output of the first diffusion model to train the third diffusion model to adjust the adjustment parameters of the third diffusion model, wherein the adjustment parameters of the third diffusion model reflect the change of the weight of the third diffusion model relative to the weight of the second diffusion model; and quantizing the adjusted adjustment parameters of the third diffusion model to obtain a quantized diffusion model.
[0007] The first diffusion model may be a pre-trained Unet model.
[0008] The step of dividing the time steps of the second diffusion model into multiple time step groups may include: obtaining the quantization error of the second diffusion model at the time step; based on the analysis of the quantization error at the time step, dividing the time steps of the second diffusion model into the multiple time step groups.
[0009] The step of obtaining the quantization error of the second diffusion model at the time step may include: at each time step of the second diffusion model, obtaining the outputs of the first diffusion model and the second diffusion model for the calibration dataset, where the calibration dataset is respectively input into the first diffusion model and the second diffusion model; determining the quantization error at each time step based on the output of the first diffusion model and the output of the second diffusion model.
[0010] The step of dividing the time steps of the second diffusion model into the multiple time step groups may include: evaluating the importance level of the time steps based on the quantization error, and grouping the time steps based on the importance level of the time steps.
[0011] The step of adjusting the weights of the second diffusion model may include: for each time step group among the multiple time step groups, for the time steps included in the corresponding time step group, performing training on the second diffusion model based on the calibration dataset; and obtaining the adjustment parameters for each time step group among the multiple time step groups.
[0012] The training of the second diffusion model may be performed using the Low-Rank Adaptation (LoRA) algorithm.
[0013] The adjustment parameters may include the low-rank parameter matrices of each layer of the second diffusion model.
[0014] The step of training the third diffusion model using the output of the first diffusion model may include: traversing two adjacent time step groups, using the difference between the output feature maps of the first diffusion model at two time steps and the output feature maps of the third diffusion model at the two time steps as the loss function to train the third diffusion model, where the two time steps respectively belong to the two adjacent time step groups.
[0015] The loss function may be calculated based on the following equation: , where and are two time steps respectively belonging to the two adjacent time step groups, represents the loss function for time step and , N represents the total number of the calibration dataset , represents the time step flow FTS matrix of the first diffusion model, The FTS matrix representing the third diffusion model, the FTS matrix of the first diffusion model and the FTS matrix of the third diffusion model are calculated based on the inner product of the output feature maps of the corresponding diffusion models at the two time steps.
[0016] The FTS matrix can be calculated based on the following equation: , where is the index of the layer of the corresponding diffusion model, is the -th layer of the corresponding diffusion model at time step output feature map, is the -th layer of the corresponding diffusion model at time step output feature map, is the height of the output feature map, is the width of the output feature map.
[0017] The quantization method may further include: dividing the layers of the second diffusion model into a plurality of blocks; wherein, the step of adjusting the weights of the second diffusion model includes: traversing each block in the plurality of blocks, and for each time step group in the plurality of time step groups, adjusting the weights of each block of the second diffusion model to obtain a third diffusion model, and the step of training the third diffusion model using the output of the first diffusion model includes: traversing each block in the plurality of blocks and training the third diffusion model using the output of each block of the first diffusion model.
[0018] The step of dividing the layers of the second diffusion model into the plurality of blocks may include: dividing the layers of the second diffusion model into residual blocks and attention blocks.
[0019] The diffusion model is configured to process at least one of image data, text data, and speech data.
[0020] According to some embodiments, the present disclosure relates to an electronic device, the electronic device includes: a memory configured to store a first diffusion model to be quantized; and a processor configured to: perform post-training quantization on the first diffusion model to obtain a second diffusion model, divide the time steps of the second diffusion model into a plurality of time step groups, for each time step group in the plurality of time step groups, adjust the weights of the second diffusion model to obtain a third diffusion model, train the third diffusion model using the output of the first diffusion model to adjust the adjustment parameters of the third diffusion model, wherein the adjustment parameters of the third diffusion model reflect the change of the weights of the third diffusion model relative to the weights of the second diffusion model, and quantize the adjusted adjustment parameters of the third diffusion model to obtain a quantized diffusion model.
[0021] The processor may be configured to: obtain the quantization error of the second diffusion model at a time step; and divide the time steps of the second diffusion model into the plurality of time step groups based on an analysis of the quantization error at the time step.
[0022] The processor may be configured to: evaluate the importance of a time step based on the quantization error, and group the time steps based on the importance of the time steps.
[0023] The processor may be configured to: for each time step group among the plurality of time step groups, traverse the time steps included in the time step group, perform training on the second diffusion model based on a calibration data set; and obtain adjustment parameters for each time step group.
[0024] The processor may be configured to: traverse two adjacent time step groups, and use the difference between the output feature maps of the first diffusion model at two time steps and the output feature maps of the third diffusion model at the two time steps as a loss function to train the third diffusion model, wherein the two time steps respectively belong to the two adjacent time step groups.
[0025] According to some embodiments, the present disclosure relates to a non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the quantization method described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Example embodiments will be understood more clearly through the following detailed description in conjunction with the drawings.
[0027] Figure 1 is a diagram showing an example of a diffusion model according to some embodiments.
[0028] Figure 2 is a diagram showing an example of the LoRA algorithm according to some embodiments.
[0029] Figure 3 is a block diagram showing an example of a quantization method of a diffusion model according to some embodiments.
[0030] Figure 4 is a block diagram showing an example of a method for grouping time steps according to some embodiments.
[0031] Figure 5A and Figure 5B is a diagram showing an example of the quantization error of a diffusion model at each time step according to some embodiments.
[0032] Figure 6 is a diagram showing an example of adjusting the weights of the second diffusion model for each time step group according to some embodiments.
[0033] Figure 7 FIG. is an illustration showing an example of training a quantized diffusion model using the output of an original diffusion model according to some embodiments.
[0034] Figure 8 FIG. is an illustration showing an example of a quantization method of a diffusion model according to some embodiments.
[0035] Figure 9 FIG. is an illustration showing an example of performing image processing using a quantized diffusion model according to some embodiments.
[0036] Figure 10 FIG. is a block diagram showing an example of an electronic device according to some embodiments. DETAILED DESCRIPTION
[0037] Hereinafter, example embodiments will be described in detail with reference to the accompanying drawings. The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent after understanding the disclosure of the present application. For example, the order of operations described herein is merely an example, except for operations that must occur in a specific order, and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application. In addition, descriptions of features known in the art may be omitted for greater clarity and brevity.
[0038] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. On the contrary, the examples described herein have been provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein that will be apparent after understanding the disclosure of the present application.
[0039] The structural or functional description of the examples disclosed herein is only intended for the purpose of describing the examples, and the examples may be implemented in various forms. The examples are not intended to be limiting, but rather are intended that various modifications, equivalents, and alternatives are also covered within the scope of the claims.
[0040] In the present disclosure, although the terms "first" or "second" are used to explain various components, the components are not limited to the terms. These terms should only be used to distinguish one component from another. For example, within the scope of the rights according to the concept of the present disclosure, a "first" component may be referred to as a "second" component, or similarly, a "second" component may be referred to as a "first" component.
[0041] In the present disclosure, it will be understood that when a component is referred to as being "connected to" another component, the component may be directly connected to or coupled to the other component, or there may be an intermediate component.
[0042] As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. It should also be understood that when the terms "comprises" and / or "comprising" are used in this specification, it indicates the presence of the stated features, integers, steps, operations, elements, components, or combinations thereof, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0043] Figure 1 is a diagram showing an example of a diffusion model according to some embodiments.
[0044] A diffusion model is a generative model that can generate desired data from Gaussian noise and is widely used in image processing (e.g., image denoising, image restoration, improving image resolution), text processing (e.g., text-to-image), and / or speech processing (e.g., speech enhancement), etc. In some embodiments, a series of face images can be used as a training set to train the diffusion model so that the diffusion model can generate new face images with various different features and expressions. In another exemplary embodiment, noisy speech can be used as a training set to train the diffusion model so that the diffusion model can predict the noise and thus generate clear speech without noise from the blurred speech.
[0045] The diffusion model obtains knowledge about the data distribution through a Markov chain process. The processing of the diffusion model generally includes two processes: the forward diffusion process and the reverse diffusion process. The forward diffusion process provides labeled training samples for the training of the diffusion model. Taking image processing as an example, in the forward diffusion process, noise is gradually added to a known image until the image completely becomes standard Gaussian noise. In the reverse diffusion process, with the added noise as the label, the noise estimation ability of the diffusion model is trained.
[0046] In Figure 1 , in the forward diffusion process, a Gaussian noise is iteratively added to a sample data according to a specific variable sampler, and T time steps are iterated to obtain a sequence of noise samples .
[0047]
[0048] In equation (1), represents the forward diffusion process; represents that the random noise conforms to a Gaussian distribution; I is the initial value of the variance of the Gaussian distribution, which is used to measure the intensity of the noise in the initial stage; is a predefined noise parameter, and the subscript t represents the time step.
[0049] Conversely, the reverse diffusion (denoising) process gradually generates high-quality data (e.g., images) by removing noise from noisy data. Since the true reverse process is not available, the diffusion model samples through a learnable conditional probability distribution.
[0050]
[0051] The mean can be inferred through the reparameterization trick, as shown in Equation (3) below.
[0052]
[0053] In Equation (3), . is a noise estimation model, usually using a UNet model.
[0054] In some embodiments, the UNet model can recover a clear image from a noisy image by performing multiple denoising processes. The input of the UNet model can include the noisy image and the current time step, and the output of the UNet model is the noise contained in the current image. The mean square error MSE (Mean-Square Error) between the noise predicted by the noise estimation model and the true noise added in the forward diffusion process can be used as the loss function to train the noise estimation model.
[0055] However, there is a huge computational cost in running the diffusion model. On the one hand, the diffusion model usually needs to perform hundreds of denoising steps to generate high-quality images, and each denoising step requires an inference of the noise estimation model. On the other hand, the structure of the noise estimation model is complex and has a very large number of parameters, requiring a large amount of memory resources.
[0056] Low-precision quantization is a model acceleration technique that compresses and accelerates neural networks by quantizing floating-point values of weights and / or activations into integer values with a specified bit width. Generally, low-precision quantization techniques can be divided into post-training quantization (PTQ) and quantization aware training (Quantization Aware Training, QAT).
[0057] PTQ obtains quantization parameters by applying quantization parameter reconstruction means in units of layers or blocks, using a specific optimization objective function and a small-scale calibration dataset without retraining the network (i.e., without updating the weights), thereby reducing quantization errors. Therefore, PTQ is more effective and practical compared to QAT which requires long-term training.
[0058] Existing PTQ research typically focuses on advanced reconstruction strategies for diffusion models. For example, Quantizing Diffusion (also known as Q-Diffusion) explores calibration dataset collection methods based on time steps. Temporal Feature Maintenance Quantization for Diffusion Models (also known as TFMQ-DM) focuses on individual reconstruction strategies for temporal feature modules. Post-training Quantization Diffusion (also known as PTQD) eliminates quantization errors by decomposing quantization noise into relevant and irrelevant parts. Accurate Data-free Post-training Quantization Framework of Diffusion Models (also known as ADP-DM) proposes grouped quantization parameters for the inference process at different time steps. Low-Rank Adaptation (also known as LoRA) helps to recover quantization errors by introducing low-rank matrix parameters.
[0059] Figure 2 A diagram showing an example of the LoRA algorithm according to some embodiments.
[0060] The LoRA algorithm freezes the weights of the pre-trained model and injects trainable low-rank decomposition matrices into each weight of the layer (e.g., Transformer layer), and greatly reduces the number of trainable parameters for downstream tasks.
[0061] In Figure 2 , W is the weight of the pre-trained model, with a size of d×d. During the training of applying the LoRA algorithm, W is frozen and not updated. Matrix A reduces the input data from d dimensions to r dimensions, and matrix B changes the data from r dimensions to d dimensions. During training, matrix A is initialized with random Gaussian noise, matrix B is zero at the start of training, and r is the rank of the matrix.
[0062] In some embodiments, the LoRA algorithm can be represented by the following equation (4). During training, the model parameters to be adjusted (i.e., the change in the weight matrix) can be represented as the product of two matrices B and A.
[0063]
[0064] In equation (4), represents the original weight matrix, A and B are low-rank parameter matrices, Input of the presentation layer Output of the presentation layer, where , , , 。
[0065] However, the existing PTQ algorithms cannot meet the expectations in low-precision quantized diffusion models. Since the diffusion model requires hundreds of denoising steps, and each denoising step requires an inference of the UNet model, the errors from the quantized UNet model will accumulate continuously, thus seriously affecting the quality of the finally generated output data.
[0066] The existing PTQ algorithms are difficult to effectively repair the accumulated quantization errors while occupying less memory resources. This is because the activation value distributions of each layer of the UNet are quite different at different time steps. If the same quantization parameters are used at all time steps, the accumulation of quantization errors will be exacerbated. If different quantization parameters are switched during the inference at different time steps, a large amount of memory resources will be occupied, thus bringing a very large additional burden to the actual deployment of the diffusion model.
[0067] Figure 3 is a block diagram showing an example of a quantization method for a diffusion model according to some embodiments.
[0068] In Figure 3 , in operation S310, the first diffusion model can be trained and quantized to obtain a second diffusion model.
[0069] In some embodiments, the first diffusion model can be a pre-trained original diffusion model and can have a high precision (e.g., floating-point 32-bit, FP32). Various diffusion models (e.g., including but not limited to GLIDE, DALL-E, DALL-E 2, Imagen, Stable Diffusion, etc.) can be used as the first diffusion model. In some embodiments, a training image set can be used to train the first diffusion model for image generation tasks. In some embodiments, noisy speech data can be used to train the first diffusion model for speech enhancement tasks.
[0070] For example, the first diffusion model can be a pre-trained Unet model. However, this is only an example, and the present disclosure is not limited thereto. Any other network structure can also be adopted as needed to implement the diffusion model.
[0071] Quantization parameters of the diffusion model can be obtained based on various PTQ algorithms to quantize the high-precision first diffusion model into a second diffusion model. The following equations (5) and (6) show the basic principles of the PTQ quantization algorithm.
[0072]
[0073]
[0074] In equations (5) and (6), and respectively represent the floating-point number before quantization and the fixed-point number after quantization, S and Z are two quantization parameters, where, S represents the scale factor, Z represents the zero point. clip represents the truncation function, round represents the rounding function. As shown in the following equations (7) and (8), the quantization parameters S and Z can be calculated based on the range of the fixed-point number ( , ) and the range of the floating-point number ( , ).
[0075]
[0076]
[0077] Therefore, when the first diffusion model is a pre-trained high-precision model, various PTQ algorithms can be used based on the calibration dataset to obtain the initial quantization parameters S and Z for each layer of the first diffusion model, and quantize the first diffusion model into the second diffusion model based on S and Z .
[0078] It can be seen from equations (6) and (8) that the rounding function round will cause a large quantization error. For this reason, optimization can be further performed after obtaining the initial quantization parameters S and Z to obtain the optimized quantization parameters S’ and Z’。
[0079] In some embodiments, various reconstruction optimization algorithms can be used to optimize the rounding function round to fine-tune the initial quantization parameters S and Z for each layer of the first diffusion model, so as to reduce the quantization error of the first diffusion model.
[0080] For example, the AdaRound algorithm can be used to optimize the initial quantization parameters of the weights and activations for each layer of the first diffusion modelS and Z are optimized, and the optimized quantization parameters S’ and Z’ are used to quantize the first diffusion model. However, this is only an example, and the present disclosure is not limited thereto. Any other rounding optimization algorithm can also be adopted to optimize the initial quantization parameters S and Z .
[0081] In operation S320, the time steps of the second diffusion model can be divided into multiple time step groups. The following will refer to Figure 4 for a detailed description of operation S320.
[0082] Figure 4 is a block diagram showing an example of a method for grouping time steps according to some embodiments.
[0083] In Figure 4 , in operation S410, the quantization error of the second diffusion model at each time step can be obtained.
[0084] At each time step, the outputs of the first diffusion model and the second diffusion model for the calibration data set as the input can be obtained respectively, and the quantization error of the second diffusion model at each time step can be determined based on the output of the first diffusion model and the output of the second diffusion model.
[0085] In some embodiments, the same calibration data set can be input into the high-precision first diffusion model and the quantized second diffusion model respectively, and the first diffusion model and the second diffusion model can be run at each time step respectively. The outputs of the first diffusion model and the second diffusion model at each time step can be compared. For example, the L2 norm (L2norm), cosine similarity, etc. can be adopted as the metric for the quantization error of the second diffusion model. However, this is only an example, and the present disclosure is not limited thereto. Any other metric can also be adopted to calculate the quantization error.
[0086] In operation S420, the quantization error at each time step can be analyzed to divide the time steps of the second diffusion model into multiple time step groups.
[0087] Figure 5A and Figure 5B are diagrams showing examples of the quantization error of the diffusion model at each time step according to some embodiments.
[0088] In some embodiments, referring to Figure 5A and Figure 5B , the first diffusion model is a FP32 Unet model, and the weights and activations of the second diffusion model are quantized to INT4 and INT8 respectively. Figure 5AShows the L2 norm between the output of the first diffusion model and the output of the second diffusion model. Figure 5B Shows the cosine similarity between the output of the first diffusion model and the output of the second diffusion model. From Figure 5A and Figure 5B it can be seen that as the time step gradually approaches t = 0, the quantization error of the second diffusion model increases significantly.
[0089] The importance of the time step can be evaluated based on the quantization error. The larger the quantization error, the higher the importance of the time step. The time steps can be grouped based on the importance of the time step. In some embodiments, consecutive time steps with relatively high importance (or relatively large quantization error) of the time step can be divided into more groups, and consecutive time steps with relatively low importance (or relatively small quantization error) of the time step can be divided into fewer groups. In another exemplary embodiment, one time step with a relatively large quantization error can form a time step group, and multiple consecutive time steps with relatively small quantization error can form a time step group.
[0090] In some embodiments, it is assumed that the number of time steps t of the second diffusion model is 50. The time steps of the second diffusion model can be divided into three groups based on the analysis of the quantization error of each time step. , where represents the time step range from t = 50 to t = 40, represents the time step range from t = 39 to t = 11, represents the time step range from t = 10 to t = 0.
[0091] By grouping the time steps through the analysis based on the quantization error, the quantization parameters of the second diffusion model can be adjusted in units of time step groups to reduce the accumulation of quantization errors caused by multiple denoising processes.
[0092] In Figure 3 , in operation S330, for each time step group among the multiple time step groups, the weights of the second diffusion model can be adjusted to obtain a third diffusion model.
[0093] In some embodiments, for each time step group, all the time steps included in the time step group can be traversed, and the second diffusion model can be trained based on the calibration dataset to obtain the adjustment parameters for each time step group. The adjustment parameters can reflect the change in the weights of the third diffusion model relative to the weights of the second diffusion model.
[0094] In some embodiments, the low-rank adaptation LoRA algorithm can be used to train the second diffusion model, and the adjustment parameters can be the low-rank parameter matrices A and B of each layer.
[0095] Figure 6 It is a diagram showing an example of adjusting the weights of the second diffusion model for each time step group according to some embodiments.
[0096] In Figure 6 , it is assumed that the time steps of the diffusion model have been divided into three time step groups based on the analysis of the quantization error. Different from the method of obtaining the same low-rank parameter matrix (A, B) for all time steps shown in Figure 2 , the quantization method according to the exemplary embodiments of the present disclosure can obtain respective low-rank parameter matrices (A1, B1), (A2, B2), and (A3, B3) for each time step group.
[0097] In some embodiments, for the time step group , calibration data sets including each time step in can be collected first. Then, from time step t = 50 to t = 40, the second diffusion model is run for each time step using the LoRA algorithm, and the second diffusion model is trained with the mean square root error MSE as the loss function to obtain the low-rank parameter matrix (A1, B1) for the time step group to . Next, for each of the time step groups
[0098] and and , the above process is repeated to obtain the low-rank parameter matrices (A2, B2) and (A3, B3) respectively.
[0099] By introducing multiple sets of low-rank parameter matrices based on time step groups, the quantization error of the diffusion model can be reduced. Since the LoRA algorithm introduces relatively small low-rank matrix parameters, even if multiple sets of low-rank parameter matrices are introduced, the increase in computational complexity and the occupation of memory resources are not significant.
[0100] To avoid the overfitting problem of the diffusion model during inference due to multiple sets of low-rank parameter matrices, a high-precision first diffusion model can also be used as a teacher model to train the quantized third diffusion model. The flow of the denoising ability of the diffusion model between different time steps can be defined as a kind of high-order distillable knowledge. By learning this knowledge of the high-precision model during the reconstruction process to optimize the low-rank parameter matrix of the quantization model, the calibration data sets of different time step groups can be effectively utilized to avoid the overfitting problem and improve the efficiency of the reconstruction process.
[0101] In Figure 3In operation S340, the output of the first diffusion model can be used to train the third diffusion model to adjust the adjustment parameters of the third diffusion model. In operation S350, the adjusted adjustment parameters of the third diffusion model can be quantized (e.g., quantized to an integer) to obtain a quantized diffusion model. In some embodiments, the adjusted adjustment parameters can be quantized to 4 bits.
[0102] In some embodiments, the knowledge of the denoising ability flow can be represented as the output feature maps of the diffusion model at two timesteps belonging to two timestep groups respectively. In some embodiments, the denoising ability flow of the diffusion model can be represented by a flow of timestep (FTS) matrix, and the FTS matrix of the first diffusion model with high precision is used as the learning target to train the third diffusion model to simultaneously fine-tune the low-rank parameter matrices of two timestep groups of the third diffusion model, thereby avoiding overfitting. However, this is only an example, and any other function can also be used to represent the knowledge of the denoising ability flow based on the output feature maps of the diffusion model at two timesteps belonging to two timestep groups respectively.
[0103] All adjacent two timestep groups can be traversed, and the difference between the output feature maps of the first diffusion model at two timesteps and the output feature maps of the third diffusion model at the two timesteps is used as the loss function to train the third diffusion model, wherein the two timesteps respectively belong to the adjacent two timestep groups.
[0104] The FTS matrix can be calculated based on the inner product of the two output feature maps of the diffusion model at two timesteps, and the loss function is constructed based on the difference between the FTS matrix of the first diffusion model and the FTS matrix of the third diffusion model.
[0105] In some embodiments, the FTS matrix of the diffusion model can be calculated based on the following equation (9).
[0106]
[0107] In equation (9), and are two timesteps belonging to two adjacent timestep groups respectively, is the index of the layer of the corresponding diffusion model, is the th layer of the corresponding diffusion model at timestep output feature map, is the th layer of the corresponding diffusion model at timestep output feature map, is the height of the output feature map, is the width of the output feature map.
[0108] After separately calculating the FTS matrix of the first diffusion model and the FTS matrix of the third diffusion model, a loss function can be constructed based on the difference between the two FTS matrices, so as to finely tune the adjustment parameters (e.g., low-rank parameter matrix) of two adjacent time step groups simultaneously.
[0109] In some embodiments, the loss function for training the third diffusion model can be calculated based on the L2 norm according to the following equation (10) . However, this is only an example of calculating the loss function, and the present disclosure is not limited thereto. Any other functional form can also be used to construct the loss function.
[0110]
[0111] In equation (10), and are two time steps belonging to two adjacent time step groups respectively, represents the loss function for time step and , N represents the total number of the calibration dataset , represents the time step stream FTS matrix of the first diffusion model, represents the FTS matrix of the third diffusion model.
[0112] Figure 7 is a diagram showing an example of training a quantization diffusion model using the output of an original diffusion model according to some embodiments.
[0113] When the weights of the second diffusion model are adjusted for three time step groups in step S330, three sets of low-rank parameter matrices (A1, B1), (A2, B2), and (A3, B3) can be obtained.
[0114] Then, in step S340, the low-rank parameter matrices (A1, B1) and (A2, B2) can be finely tuned simultaneously for time step groups and , and then the low-rank parameter matrices (A2, B2) and (A3, B3) can be finely tuned simultaneously for time step groups and .
[0115] Figure 7 shows an example of finely tuning the low-rank parameter matrices (A2, B2) and (A3, B3) simultaneously for time step groups and according to some embodiments. In Figure 7Among them, the original diffusion model can be a pre-trained high-precision quantization model (e.g., the first diffusion model), and the third diffusion model can be a diffusion model that applies the LoRA algorithm to each group of time steps to obtain multiple groups of low-rank parameter matrices. In Figure 7 Among them, to respectively represent the datasets input to the first diffusion model from time step t = 50 to time step t = 0, to respectively represent the datasets input to the third diffusion model from time step t = 50 to time step t = 0.
[0116] In Figure 7 among them, the time steps t = 20 in the group of time steps and the time step t = 10 in the group of time steps can be selected respectively, and the FTS matrices of the first diffusion model and the third diffusion model at these two time steps are calculated (in Figure 7 among them, respectively represented as FP32 FTS and Quan FTS), and the L2-norm between the FTS matrices is used as the loss function to perform backpropagation, so as to update the low-rank parameter matrices (A2, B2) and (A3, B3).
[0117] Figure 8 is a diagram showing an example of a quantization method of a diffusion model according to some embodiments.
[0118] In some embodiments, the diffusion model can also be structurally analyzed, the layers of the diffusion model are divided into multiple blocks, and the adjustment parameters for each group of time steps are calculated in units of blocks, and the adjustment parameters are further adjusted in units of blocks.
[0119] In Figure 8 among them, steps S810 and S820 can be the same as steps S310 and S320 described with reference to Figure 3 For the sake of brevity, their detailed descriptions will be omitted.
[0120] In step S830, the layers of the second diffusion model can be divided into multiple blocks. For example, the layers of the second diffusion model can be divided into residual blocks and attention blocks. In another example, each layer of the diffusion model can be used as a separate block. However, this is only an example, and the present disclosure is not limited thereto.
[0121] In step S840, each block among the multiple blocks can be traversed, and for each group of time steps among the multiple groups of time steps, the weights of each block of the second diffusion model are adjusted to obtain the third diffusion model. In some embodiments, for the first block of the diffusion model, the LoRA algorithm can be applied to each group of time steps to obtain adjustment parameters (e.g., low-rank parameter matrices).
[0122] In step S850, each of the plurality of blocks may be traversed, and the output of each block of the first diffusion model is used to train the third diffusion model. In some embodiments, after obtaining the adjustment parameters for each time step group of the first block of the diffusion model, for every two adjacent time step groups, a low-rank parameter matrix may be fine-tuned using a loss function based on FTS.
[0123] The above steps may be repeated for each of the remaining blocks to obtain the adjustment parameters for all blocks.
[0124] In step S860, the adjusted adjustment parameters of the third diffusion model may be quantized to obtain a quantized diffusion model.
[0125] Figure 9 FIG. is a diagram showing an example of performing image processing using a quantized diffusion model according to some embodiments.
[0126] In Figure 9 , at time steps t = 50 to t = 40, inference may be performed based on the first set of first low-rank parameter matrices (A1, B1), at time steps t = 39 to t = 11, inference may be performed based on the second set of first low-rank parameter matrices (A2, B2), and at time steps t = 10 to t = 0, inference may be performed based on the third set of first low-rank parameter matrices (A3, B3).
[0127] Figure 10 FIG. is a block diagram showing an example of an electronic device according to some embodiments.
[0128] The quantization method may be deployed on various electronic devices through software or hardware. The electronic device may include, for example, a desktop computer, a laptop computer, a tablet computer, a smart phone, a wearable electronic device (such as a smart bracelet, a smart watch, etc.), a virtual reality device, etc. However, the present disclosure is not limited thereto, and the electronic device according to the present disclosure may be any electronic device having a function of processing multimedia data (such as at least one of image data, text data, and voice data).
[0129] The quantization method may also be deployed on a server. The server may be located in the cloud or locally, and may be a physical device or a virtual device, such as a virtual machine, a container, etc. For example, the server may be located in the cloud and communicate with a terminal device. The server receives the original diffusion model sent by the terminal device, quantizes the original diffusion model using the quantization method deployed on the server, and returns the quantized target diffusion model to the terminal device.
[0130] In Figure 10 , the electronic device 1000 may include a memory 1010 and a processor 1020.
[0131] The memory 1010 may store the original diffusion model to be quantized (i.e., the first diffusion model). The original diffusion model may be a pre-trained high-precision model.
[0132] The processor 1020 may quantize the first diffusion model based on a quantization method. Specifically, the processor 1020 may perform post-training quantization on the first diffusion model and obtain a preliminarily quantized second diffusion model. The processor 1020 may divide the time steps of the second diffusion model into multiple time step groups, and for each time step group, adjust the weights of the second diffusion model to obtain a third diffusion model. Next, the processor 1020 may use the output of the first diffusion model to train the third diffusion model to adjust the adjustment parameters of the third diffusion model. Finally, the processor 1020 may quantize the adjusted adjustment parameters of the third diffusion model to obtain a finally quantized diffusion model.
[0133] The devices, units, modules, and other components described herein are implemented by hardware components. Examples of hardware components that can be used to perform the operations described in this application include, where appropriate: controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and / or any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements (such as logic gate arrays, controllers, arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, and / or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result). In one example, a processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) for performing the operations described in this application. The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For brevity, the singular terms "processor" or "computer" can be used in the description of the examples described in this application, but in other examples, multiple processors or computers can be used, or a processor or computer can include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component, or two or more hardware components, can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller, can implement a single hardware component, or two or more hardware components. The hardware components can have any one or more of different processing configurations, examples of different processing configurations include: single processor, independent processors, parallel processors, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and / or multiple instruction multiple data (MIMD) multiprocessing.
[0134] A method for performing the operations described in this disclosure is performed by computing hardware (e.g., by one or more processors or computers), which is implemented to execute instructions or software as described above to perform the operations performed by the method described in this application. For example, a single operation, or two or more operations, may be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations may be performed by one or more processors, or a processor and a controller, and one or more other operations may be performed by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller, may perform a single operation, or two or more operations.
[0135] Instructions or software for controlling a processor or computer to implement the hardware components and perform the method described above may be written as a computer program, code segment, instruction, or any combination thereof, to individually or jointly direct or configure the processor or computer to operate as a machine or special-purpose computer to perform the operations performed by the hardware components and method described above. In one example, the instructions and / or software include machine code (such as machine code generated by a compiler) that is directly executed by the processor or computer. In another example, the instructions or software include high-level code that is executed by the processor or computer using an interpreter. A person of ordinary skill in the art or a programmer can easily write the instructions and / or software based on the block diagrams and flowcharts shown in the drawings and the corresponding descriptions in the specification, and the block diagrams and flowcharts shown in the drawings and the corresponding descriptions in the specification disclose algorithms for performing the operations performed by the hardware components and method described above.
[0136] Instructions or software for controlling a processor or computer to implement the hardware components and execute the methods as described above, as well as any associated data, data files, and data structures, can be recorded, stored, or fixed in one or more non-transitory computer-readable storage media, or recorded, stored, or fixed on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-RLTH, BD-RE, Blu-ray or optical disc storage devices, hard disk drives (HDDs), solid state drives (SSDs), flash memory, cartridge memory (such as, multimedia cards or micro-cards (e.g., Secure Digital (SD) or Extreme Digital (XD))), magnetic tapes, floppy disks, magneto-optical data storage devices, optical data storage devices, hard disks, solid state disks, and / or at least one of any other device configured to store instructions or software and any associated data, data files, and / or data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and / or data structures to a processor or computer such that the processor or computer can execute the instructions.
[0137] Although the present disclosure contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, its equivalents, and the scope of the claims described later. The specific features described in the context of separate embodiments in the present disclosure can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although the features above may be described as acting in a particular combination, one or more features from the combination can be deleted in some cases, and the combination can be directed to a sub-combination or a variant of the sub-combination.
Claims
1. A quantization method for a diffusion model, comprising: Quantizing the first diffusion model after training to obtain a second diffusion model; Dividing the time steps of the second diffusion model into multiple time step groups; Adjusting the weights of the second diffusion model for each of the multiple time step groups to obtain a third diffusion model; Training the third diffusion model using the output of the first diffusion model to adjust the adjustment parameters of the third diffusion model, where the adjustment parameters of the third diffusion model reflect the change in the weights of the third diffusion model relative to the weights of the second diffusion model; Quantizing the adjusted adjustment parameters of the third diffusion model to obtain a quantized diffusion model.
2. The quantization method according to claim 1, wherein, The first diffusion model is a pre-trained Unet model.
3. The quantization method according to claim 1, wherein, The step of dividing the time steps of the second diffusion model into multiple time step groups includes: Obtaining the quantization error of the second diffusion model at the time step; Based on the analysis of the quantization error at the time step, dividing the time steps of the second diffusion model into the multiple time step groups.
4. The quantization method according to claim 3, wherein, The step of obtaining the quantization error of the second diffusion model at the time step includes: At each time step of the second diffusion model, obtaining the outputs of the first diffusion model and the second diffusion model for the calibration dataset, and the calibration dataset is respectively input into the first diffusion model and the second diffusion model; Determining the quantization error at each time step based on the output of the first diffusion model and the output of the second diffusion model.
5. The quantization method according to claim 4, wherein, The step of dividing the time steps of the second diffusion model into the multiple time step groups includes: Evaluating the importance of the time step based on the quantization error, and grouping the time steps based on the importance of the time step.
6. The quantization method according to claim 1, wherein The step of adjusting the weights of the second diffusion model includes: For each of the multiple time step groups, for the time steps included in the corresponding time step group, training the second diffusion model based on the calibration dataset; and Obtaining the adjustment parameters for each of the multiple time step groups.
7. The quantization method according to claim 6, wherein, Using the low-rank adaptation LoRA algorithm to train the second diffusion model.
8. The quantization method according to claim 7, wherein, The adjustment parameters include the low-rank parameter matrices of each layer of the second diffusion model.
9. The quantization method according to claim 1, wherein The step of training the third diffusion model using the output of the first diffusion model includes: Traversing two adjacent time step groups, and using the difference between the output feature maps of the first diffusion model at two time steps and the output feature maps of the third diffusion model at the two time steps as a loss function to train the third diffusion model, where the two time steps respectively belong to the two adjacent time step groups.
10. The quantization method according to claim 9, wherein, The loss function is calculated based on the following equation: Among them, and are two time steps respectively belonging to the two adjacent time step groups, represents the loss function for time steps and , N represents the total number of the calibration dataset , represents the time step flow FTS matrix of the first diffusion model, represents the FTS matrix of the third diffusion model, where the FTS matrix of the first diffusion model and the FTS matrix of the third diffusion model are calculated based on the inner product of the output feature maps of the corresponding diffusion models at the two time steps.
11. The quantization method according to claim 10, wherein, The FTS matrix is calculated based on the following equation: Among them, is the index of the layer of the corresponding diffusion model, is the -th layer of the corresponding diffusion model at time step output feature map, is the -th layer of the corresponding diffusion model at time step output feature map, is the height of the output feature map, is the width of the output feature map.
12. The quantization method according to claim 1, further comprising: Dividing the layers of the second diffusion model into multiple blocks; Among them, the steps of adjusting the weights of the second diffusion model include: traversing each of the plurality of blocks, and for each time step group in the plurality of time step groups, adjusting the weights of each block of the second diffusion model to obtain a third diffusion model. Among them, the steps of training the third diffusion model using the output of the first diffusion model include: traversing each of the plurality of blocks, and using the output of each block of the first diffusion model to train the third diffusion model.
13. The quantization method according to claim 12, wherein, The steps of dividing the layers of the second diffusion model into the plurality of blocks include: Dividing the layers of the second diffusion model into residual blocks and attention blocks.
14. The quantization method according to claim 1, wherein, The diffusion model is configured to process at least one of image data, text data, and voice data.
15. An electronic device, comprising: A memory configured to store a first diffusion model to be quantized; A processor configured to: Perform post-training quantization on the first diffusion model and obtain a second diffusion model; Divide the time steps of the second diffusion model into a plurality of time step groups; For each time step group in the plurality of time step groups, adjust the weights of the second diffusion model and obtain a third diffusion model; Train the third diffusion model using the output of the first diffusion model and adjust the adjustment parameters of the third diffusion model, where the adjustment parameters of the third diffusion model reflect the change in the weights of the third diffusion model relative to the weights of the second diffusion model; Quantize the adjusted adjustment parameters of the third diffusion model and obtain a quantized diffusion model.
16. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 14.
Citation Information
Cited By
Diffusion model reasoning acceleration method based on optimal time step sequence search and knowledge distillation
CN121809699A