A mixed-precision quantization method for diffusion model generation
By optimizing the quantization bit width of the diffusion model through single-path sampling and genetic algorithm, the problems of time step and mixed precision quantization are solved, achieving faster and better image generation.
Patent Information
- Application Number
- CN202410141913.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-02-01
AI Technical Summary
Existing technologies fail to effectively consider time steps and mixed-precision quantization in diffusion generation models, resulting in limited generation quality and speed, and unreasonable model quantization bit width allocation.
A mixed-precision quantization training strategy with single-path sampling and a genetic algorithm based on time steps are adopted to dynamically adjust the quantization bit width. The quantization parameters are updated through forward and backward propagation of the model, and the genetic algorithm is combined to search for the optimal configuration.
The generation speed and quality of the diffusion model are improved, the memory overhead is reduced, and it is suitable for resource-constrained scenarios, ensuring the quality and speed of generated images.
Smart Images

Figure CN117892792B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image generation, and in particular to a diffusion model mixed precision quantization method for generating images. Background Art
[0002] Existing engineering techniques primarily use knowledge distillation to distill a trained teacher model into a student model with fewer steps, thereby improving the speed and quality of diffusion-based generative models. However, this approach fails to consider the role of model quantization and the impact of different time step sizes on generation quality and speed.
[0003] The new research results mainly accelerate the diffusion model from two perspectives: efficient sampling and model quantization:
[0004] 1) Some methods significantly reduce the number of sampling steps by discarding the Markov property of the diffusion model and reformulating the generating equation. Other methods use differentiable search methods to search for time steps and accelerate sampler selection strategies to improve model performance. However, these methods do not simultaneously consider the issues of time step selection and mixed-precision quantization.
[0005] 2) In the unified precision quantization area, some methods design model quantization methods based on the UNet structure in the diffusion model and the activation quantization range that changes with time steps. In the mixed precision quantization area, some methods make bit width allocation decisions based on the signal-to-noise ratio of different bit widths. However, these methods cannot effectively capture the quantization sensitivity of different layers in the model, and the quantization bit width allocation for the model is also static.
[0006] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0007] The main purpose of the present invention is to overcome the defects of the above-mentioned background technology and provide a diffusion model mixed precision quantization method for generating images.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A mixed-precision quantization method for a diffusion model for generating an image, comprising:
[0010] Phase 1: Using image training data, the model is trained using a single-path sampling mixed-precision quantization training strategy. This strategy trains the model with a set quantization bit width and assigns different quantization bit widths to different layers based on their sensitivity to quantization. The corresponding quantization bit widths and quantization parameters are updated through forward and backward propagation of the model to minimize the loss error. After training, a mixed-precision quantization supernet is obtained.
[0011] The second stage: A time-step-based genetic algorithm is used to search on the trained mixed-precision quantization supernet. During the search process, the genetic algorithm dynamically adjusts the configuration of the quantization bit width according to the time step, and generates new candidate solutions through mutation and crossover operations. The candidate solutions contain different combinations of time steps and quantization bit widths. The best-performing candidate solution is maintained through a solution evaluation method. The best-performing candidate solution is used as the final solution to determine the optimal model configuration, thereby obtaining the final diffusion model for image generation.
[0012] Further:
[0013] In the first stage, a mixed precision model is obtained through multiple iterations of training using a single-path sampling method. The following steps are performed in sequence in each iteration round:
[0014] S1. Within each layer of the model, a path is selected for forward propagation based on a preset probability; the selected bit widths of weights and activations are determined and quantized; within each layer, forward propagation is performed based on the probability of the selected bit width to obtain the quantized output of this layer;
[0015] S2. After each layer samples the path according to step S1 and obtains the corresponding output, calculate the loss error between the quantized output of each layer and the actual output;
[0016] S3. Based on the loss error obtained in step S2, update the quantization parameter of the selected bit width by backpropagation using the gradient descent method.
[0017] Step S1 specifically includes:
[0018] In each layer of the model, a path is selected for forward propagation according to probability; the bit width selection set of weights is recorded as , for each element in the set, there is Represents the scaling factor and offset respectively; the i-th bit width weight To quantify, use the following formula:
[0019]
[0020] in Represents the rounding function; the activated bit width selection set is , for each element in the set, there is also a corresponding scaling factor and offset , the j-th bit width is the output of this layer To quantify, use the following formula:
[0021]
[0022] in It is both the output of this layer and the input of the next layer;
[0023] In each layer, the weights and activation widths taken in this round are determined according to the probability. Then, according to the above two formulas, forward propagation is performed in the layer to obtain the quantized output corresponding to this layer. .
[0024] In step S1, the selection probability of each bit width is inversely proportional to its own bit number. Specifically, the number of bits in each bit width is ,in , and its corresponding selection probability is ,in .
[0025] Step S2 and step S3 specifically include:
[0026] Step S2: After each layer follows the sampling path of step S1 and obtains the corresponding output, calculate the loss error of each layer:
[0027]
[0028] in, represents the mean square error;
[0029] Step S3: Based on the loss error obtained in step S2, update the selected bit width by backpropagation using the gradient descent method Quantization parameter as well as .
[0030] In the second stage, a genetic algorithm is used for searching, and the specific steps are as follows:
[0031] T1: Set the number of time steps to be sampled and the maximum amount of shaping operations;
[0032] T2: Randomly initialize multiple candidate strategies, calculate the generation performance of each strategy, and maintain the top strategies; among them, record the candidate strategies ,in Indicates which time steps are selected for sampling, Represents the set of quantization bit width selections used for each layer of the model at each time step;
[0033] T3: Perform crossover, mutation, and random initialization operations of the genetic algorithm during iteration, and calculate the corresponding FID to update the optimal strategy set;
[0034] T4: Generate new strategies through crossover operation of genetic algorithm;
[0035] T5: Generate new strategies through mutation operations of genetic algorithms;
[0036] T6: Generate new strategies through random initialization of genetic algorithm;
[0037] T7: Update the optimal strategy set based on FID.
[0038] In step T2, an approximate evaluation method based on the Kendall-tau correlation coefficient is used to calculate the generation performance of each strategy, and the FID indicator is used to measure the fidelity of the generated dataset;
[0039] Among them, the sampling Time-step paths ,in Represents a sampling path of length L. For these N time step paths, 50k and Samples and the target data set are used to calculate FID respectively, and two sets of FID evaluation data are obtained. and , calculate the Kendall-tau correlation coefficient between these two sets of data to meet the The minimum number of As the number of generated images for each strategy.
[0040] In step T4, the crossover operation specifically includes: randomly selecting two strategies from the candidate strategy set with the highest performance , from which new strategies are formed ,in Each dimension has a 50% probability The value of the corresponding dimension.
[0041] In step T4, the mutation operation specifically includes: randomly selecting a strategy from the candidate strategy set with the highest performance , from which new strategies are formed ,in 95% probability of each dimension and If the values of the corresponding dimensions are the same, there is a 5% chance that they will mutate into random values.
[0042] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the diffusion model mixed-precision quantization method.
[0043] The present invention has the following beneficial effects:
[0044] The present invention proposes a mixed-precision quantization method for a diffusion model used to generate images. Targeting the time step characteristics of the diffusion model, the method trains and searches within the context of mixed-precision quantization to obtain a model that dynamically allocates quantization bit widths across time steps and model layers, thereby improving the generation speed and quality of the diffusion model. The present invention can train a mixed-precision quantization supernet that performs well for different quantization strategies. Once this mixed-precision quantization supernet is obtained, the method can search for the optimal time step selection and mixed-precision quantization scheme for model generation speed and quality, determine the optimal model configuration, and ultimately obtain the diffusion model used to generate the image. The method of the present invention can reduce memory overhead for deployment, save related storage resources, and improve the generation speed and quality of the model.
[0045] The main advantages of the present invention are as follows:
[0046] 1. This invention, based on model quantization, allocates quantization bit widths to different layers based on their sensitivity to quantization, thereby accelerating the generation of diffusion models in a more rational and efficient manner. Through model quantization, the present invention's solution can reduce the number of model parameters in the diffusion model, thereby lowering deployment memory overhead and saving related storage resources, making it more applicable in resource-constrained scenarios.
[0047] 2. The present invention utilizes a single-path sampling method to train a mixed-precision quantization model. By using this concise method, the present invention can train a mixed-precision quantization supernet from a large training space, and can achieve relatively good performance under different quantization selection schemes.
[0048] 3. The genetic algorithm search architecture constructed by the present invention can perform efficient searches in the search space selected by the time step and quantization bit width. The efficiency of the genetic algorithm search is accelerated by the scheme evaluation acceleration method, thereby obtaining a mixed precision quantization model with excellent time step perception, thereby improving the generation speed and quality of the model.
[0049] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flowchart of a mixed-precision quantization method for a diffusion model used to generate an image according to an embodiment of the present invention;
[0051] Figure 2 This is an algorithm framework diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0052] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.
[0053] See Figure 1 , an embodiment of the present invention provides a diffusion model mixed precision quantization method for generating an image, comprising:
[0054] Phase 1: Using image training data, the model is trained using a single-path sampling mixed-precision quantization training strategy. This strategy trains the model with a set quantization bit width and assigns different quantization bit widths to different layers based on their sensitivity to quantization. The corresponding quantization bit widths and quantization parameters are updated through forward and backward propagation of the model to minimize the loss error. After training, a mixed-precision quantization supernet is obtained.
[0055] The second stage: A time-step-based genetic algorithm is used to search on the trained mixed-precision quantization supernet. During the search process, the genetic algorithm dynamically adjusts the configuration of the quantization bit width according to the time step, and generates new candidate solutions through mutation and crossover operations. The candidate solutions contain different combinations of time steps and quantization bit widths. The best-performing candidate solution is maintained through a solution evaluation method. The best-performing candidate solution is used as the final solution to determine the optimal model configuration, thereby obtaining the final diffusion model for image generation.
[0056] In an embodiment of the present invention, in the first stage, a mixed precision quantization training strategy for single-path sampling is proposed, and a quantization bit width selection strategy is sampled in each generation of training. The corresponding quantization bit width quantization parameters are updated through forward propagation and backward propagation of the model, and finally a mixed precision quantization supernet is obtained.
[0057] In an embodiment of the present invention, in the second stage, a time step and mixed precision quantization strategy search method based on a genetic algorithm is proposed. The mutation and crossover operations in the genetic algorithm are used on the trained mixed precision quantization supernet to form new candidate solutions. The set of best performing candidate solutions is maintained through a solution evaluation method. After the search is completed, the candidate solution with the best performance is the final solution.
[0058] The embodiment of the present invention proposes a mixed-precision quantization method for a diffusion model for generating images. Aiming at the time step characteristics of the diffusion model, the method trains and searches under the background of mixed-precision quantization to obtain a model that dynamically allocates quantization bit widths under different time steps and model layers, thereby improving the generation speed and quality of the diffusion model. The present invention can train a mixed-precision quantization supernet that can perform well for different quantization strategies, and when this mixed-precision quantization supernet is obtained, it can search for the time step selection and mixed-precision quantization scheme with the best model generation speed and quality, determine the optimal model configuration, and thus obtain the final diffusion model for generating images. The method of the present invention can reduce the memory overhead of deployment, save related storage resources, and improve the generation speed and quality of the model.
[0059] Specific embodiments of the present invention are further described below.
[0060] The embodiment of the present invention is divided into two stages of processing.
[0061] Phase 1
[0062] In the first stage (Stage 1), the present invention uses a single-path sampling method to obtain a mixed precision model through multiple iterative training. Steps S1 to S3 are performed in each training cycle (epoch). The specific steps are as follows:
[0063] In each iteration, execute S1, S2, and S3 in sequence:
[0064] Step S1: Select a path for forward propagation according to probability in each layer of the model. The bit width selection set of weights is , for each element in the set, there is Represents the scaling factor and offset respectively. The following formula is used for quantification:
[0065]
[0066] in Represents the rounding function; the activated bit width selection set is , for each element in the set, there is also a scaling factor and offset , when the j-th bit width is used to activate the output of this layer The following formula is used for quantification:
[0067]
[0068] in It is both the output of this layer and the input of the next layer.
[0069] In each layer, the weights and activation widths taken in this round are determined according to the probability. Then, according to the above two formulas, forward propagation is performed in the layer to obtain the quantized output corresponding to this layer. Here, the probability of selecting each bit width is inversely proportional to its own number of bits. This is because the loss of low bits is greater and more emphasis needs to be placed on training. Specifically, the number of bits in each bit width is ,in , then the corresponding selection probability is ,in .
[0070] Step S2: At each layer, sample the path according to step S1 and get the corresponding output After that, calculate each layer and The loss error between .
[0071]
[0072] Step S3: Based on the loss error obtained in step S2, update the selected bit width by backpropagation using the gradient descent method Quantization parameter as well as .
[0073] Phase II
[0074] In the second stage (Stage 2), the present invention uses a genetic algorithm to search. Within each training cycle (epoch), new candidate strategies are generated through crossover and mutation. The top 50 candidate strategies are then updated from low to high based on the FID metric. The specific steps are as follows:
[0075] Step T1: Initialize the strategy requirements and give the number of time steps to be sampled in advance And the maximum amount of plastic operations limit In the experiment, the amount of shaping operations for the entire 6-bit model was chosen as the upper limit.
[0076] Step T2: Initialize the candidate set and record the candidate strategy ,in Indicates which time steps are selected for sampling, Represents the set of quantization bit width selections used by each layer of the model at each time step. First, 50 candidate strategies are randomly initialized as the initial strategy, where All strategies are randomly selected. These strategies need to meet the computational constraints, otherwise they will be re-initialized. After obtaining these candidate strategies, an approximate evaluation method based on the Kendall-tau correlation coefficient is used to calculate the generation performance of each strategy. Here, the FID indicator is used to measure the fidelity of the generated dataset. The FID calculation formula is:
[0077]
[0078] in are the mean and variance of the actual image dataset and the generated image dataset, respectively. The lower the FID, the better the model performance.
[0079] The purpose of the approximate Kendall-tau correlation coefficient evaluation method is to reduce the computational overhead required to evaluate strategies. To obtain realistic FID performance, a large number of images (possibly up to 50,000) must be sampled for calculation. Performing a complete operation for each candidate strategy would be prohibitively expensive. To address this issue, in a preferred embodiment, the approximate evaluation method employs the following steps:
[0080] Sampling out Time-step paths ,in Represents a sampling path of length L. For these N time step paths, 50k and Samples and the target data set are used to calculate FID respectively, and two sets of FID evaluation data are obtained. and , calculate the Kendall-tau correlation coefficient between these two sets of data to meet the The minimum number of As the number of generated images for each strategy. For the cifar dataset, the number is 1000.
[0081] After obtaining the initial candidate strategies that meet the conditions and their corresponding FID performance, maintain the set of candidate strategies with the top 50 performance.
[0082] Step T3: In each iteration, execute steps T4, T5, T6, and T7 in sequence:
[0083] Step T4: Use crossover operation to obtain 15 new candidate strategies that meet the computational constraints. The crossover operation is as follows: randomly select two strategies from the Top 50 candidate strategy set. , from which new strategies are formed ,in Each dimension has a 50% probability The value of the corresponding dimension.
[0084] Step T5: Use mutation operation to obtain 20 new candidate strategies that meet the computational constraints. The mutation operation is as follows: Randomly select a strategy from the Top 50 candidate strategy set. , from which new strategies are formed ,in 95% probability of each dimension and For values of the corresponding dimensions that are the same, there is a 5% chance that they will mutate to a random value, as shown in the orange figure where the step size mutates from 85 to 92.
[0085] Step T6: Use the random initialization operation as in step T2 to obtain 15 new candidate strategies that meet the computational constraints.
[0086] Step T7: Update the top 50 candidate strategy set from the candidate strategies of steps T3, T4, T5, and T6 and the original top 50 candidate strategy set according to FID from low to high.
[0087] Experimental results
[0088] The present invention can improve performance in searching for timestep and further searching for bit-width.
[0089] Experiments are conducted on the Cifar dataset. Guided_diffusion is used, and DDIM is used as the sampler. The quantization code is based on the Q-diffusion code.
[0090] After two stages of training, the experimental results are as follows:
[0091] Quad means that the time step increases at a square rate, such as 0, 1, 2, 4, 9.
[0092] The experimental results show that increasing the search time step can improve performance. Adding mixed-precision quantization significantly exceeds the performance of using only 6 bits.
[0093] Compared with the traditional method, the main advantages of the present invention are as follows:
[0094] 1. This invention, based on model quantization, allocates quantization bit widths to different layers based on their sensitivity to quantization, thereby accelerating the generation of diffusion models in a more rational and efficient manner. Through model quantization, the present invention's solution can reduce the number of model parameters in the diffusion model, thereby lowering deployment memory overhead and saving related storage resources, making it more applicable in resource-constrained scenarios.
[0095] 2. The present invention utilizes a single-path sampling method to train a mixed-precision quantization model. By using this concise method, the present invention can train a mixed-precision quantization supernet from a large training space, and can achieve relatively good performance under different quantization selection schemes.
[0096] 3. The genetic algorithm search architecture constructed by the present invention can perform efficient searches in the search space selected by the time step and quantization bit width. The efficiency of the genetic algorithm search is accelerated by the scheme evaluation acceleration method, thereby obtaining a mixed precision quantization model with excellent time step perception, thereby improving the generation speed and quality of the model.
[0097] Application examples of the present invention:
[0098] User Scenario 1: In an instant image generation service, the generation speed affects different users differently. Simply changing the diffusion model generation step number to accommodate different user needs can reduce image quality. This invention pre-searches for appropriate time step selection and hybrid quantization strategies for different sampling step numbers, thereby adjusting the model sampling scheme to meet different needs and ensuring image quality.
[0099] User usage scenario 2: Due to limited computing and storage resources on edge servers or edge devices, the diffusion generation model cannot be directly deployed. The present invention can search for a suitable mixed-precision quantization scheme under the resource constraints of the edge, reduce the storage overhead of the diffusion generation model, and enable resource-constrained edge devices to run the diffusion generation model.
[0100] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.
[0101] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.
[0102] An embodiment of the present invention further provides a processor, which executes a computer program and at least performs the method described above.
[0103] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0104] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0105] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0107] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0108] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0109] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0110] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0111] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0112] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.
Claims
1. A mixed precision quantization method for a diffusion model for generating an image, characterized in that: include: Phase 1: Using image training data, the model is trained using a single-path sampling mixed-precision quantization training strategy. This strategy trains the model with a set quantization bit width and assigns different quantization bit widths to different layers based on their sensitivity to quantization. The corresponding quantization bit widths and quantization parameters are updated through forward and backward propagation of the model to minimize the loss error. After training, a mixed-precision quantization supernet is obtained. Phase II: A time-step-based genetic algorithm is used to search on the trained mixed-precision quantization supernet. During the search, the genetic algorithm dynamically adjusts the quantization bit width configuration based on the time step. New candidate solutions are generated through mutation and crossover operations. These candidate solutions contain different combinations of time steps and quantization bit widths. A solution evaluation method is used to maintain a set of best-performing candidate solutions. The best-performing candidate solution is selected as the final solution, and the optimal model configuration is determined, resulting in the final diffusion model used to generate the image. In the first stage, a mixed precision model is obtained through multiple iterations of training using a single-path sampling method. The following steps are performed in sequence in each iteration round: S1. Within each layer of the model, a path is selected for forward propagation based on a preset probability; the selected bit widths of weights and activations are determined and quantized; within each layer, forward propagation is performed based on the probability of the selected bit width to obtain the quantized output of this layer; S2. After each layer samples the path according to step S1 and obtains the corresponding output, calculate the loss error between the quantized output of each layer and the actual output; S3. Based on the loss error obtained in step S2, update the quantization parameter of the selected bit width by back propagation using the gradient descent method.
2. The mixed precision quantization method for the diffusion model according to claim 1, wherein: Step S1 specifically includes: In each layer of the model, a path is selected for forward propagation according to probability; the set of weight selection bits is recorded as , for each element in the set, there is Represents the scaling factor and offset respectively; the i-th bit width weight To quantify, use the following formula: ; in Represents the rounding function; the activated selection bit width set is , for each element in the set, there is also a corresponding scaling factor and offset , the j-th bit width is the output of this layer To quantify, use the following formula: ; in It is both the output of this layer and the input of the next layer; In each layer, the weights and activation widths taken in this round are determined according to the probability. Then, according to the above two formulas, forward propagation is performed in the layer to obtain the quantized output corresponding to this layer. .
3. The mixed precision quantization method for the diffusion model according to claim 2, wherein: In step S1, the selection probability of each bit width is inversely proportional to its own bit number. Specifically, the number of bits in each bit width is ,in , and its corresponding selection probability is ,in .
4. The mixed precision quantization method for the diffusion model according to claim 2, wherein: Step S2 and step S3 specifically include: Step S2: After each layer follows the sampling path of step S1 and obtains the corresponding output, calculate the loss error of each layer: ; in, represents the mean square error; Step S3: Based on the loss error obtained in step S2, update the weights and activations by backpropagation according to the gradient descent method. Quantization parameter as well as .
5. The diffusion model mixed precision quantization method according to any one of claims 1 to 4, characterized in that: In the second stage, a genetic algorithm is used for searching, and the specific steps are as follows: T1: Set the number of time steps to be sampled and the maximum amount of shaping operations; T2: Randomly initialize multiple candidate strategies, calculate the generation performance of each candidate strategy, and maintain multiple candidate strategies with the highest performance; among them, record the candidate strategy ,in Indicates which time steps are selected for sampling, Represents the set of quantization selection bit widths used for each layer of the model at each time step; T3: Perform crossover, mutation, and random initialization operations of the genetic algorithm during iteration, and calculate the corresponding FID to update the optimal strategy set; T4: Generate new strategies through crossover operation of genetic algorithm; T5: Generate new strategies through mutation operations of genetic algorithms; T6: Generate new strategies through random initialization of genetic algorithm; T7: Update the optimal strategy set based on FID.
6. The mixed precision quantization method for a diffusion model according to claim 5, wherein: In step T2, an approximate evaluation method based on the Kendall-tau correlation coefficient is used to calculate the generation performance of each strategy, and the FID indicator is used to measure the fidelity of the generated dataset; Among them, the sampling Time-step paths ,in Represents a sampling path of length L. For these N time step paths, 50k samples are sampled respectively. Samples and the target data set are used to calculate FID respectively, and two sets of FID evaluation data are obtained. and , calculate the Kendall-tau correlation coefficient between these two sets of data to meet the The minimum number of As the number of generated images for each strategy.
7. The diffusion model mixed precision quantization method according to claim 5, characterized in that: In step T4, the crossover operation specifically includes: randomly selecting two strategies from the candidate strategy set with the highest performance , from which new strategies are formed ,in Each dimension has a 50% probability The value of the corresponding dimension.
8. The mixed precision quantization method for a diffusion model according to claim 5, wherein: In step T5, the mutation operation specifically includes: randomly selecting a strategy from the candidate strategy set with the highest performance , from which new strategies are formed ,in 95% probability of each dimension and If the values of the corresponding dimensions are the same, there is a 5% chance that they will mutate into random values.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the diffusion model mixed precision quantization method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Runtime dynamic reasoning method for neural network quantization
CN115423071A
Sufficient training method for mixed bit width super network
CN115438784A