Dynamic quantification method and system of AIGC model
Through the dynamic quantization method, the network layer of the AIGC model is traversed layer by layer, and the problems of large complexity and accuracy losses of existing quantization tools are solved, and the efficient quantization of the AIGC model is realized, and the user operation is simplified.
Patent Information
- Application Number
- CN202510378954.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-24
AI Technical Summary
The existing quantitative tools have complexity, large accuracy losses and incompatible LoRA models for the quantification of AIGC models, which is difficult to meet the personalized needs of users.
The dynamic quantization method is adopted to obtain the final AIGC quantization model by traversing the network layer of the AIGC model layer by layer, performing quantization processing, and filtering according to structural similarity and calculation density, retaining or abandoning the quantization state.
It realizes automatic and efficient dynamic quantization of the AIGC model, reduces accuracy loss, is compatible with LoRA models, and does not require invasive code modification, simplifying user operations.
Smart Images

Figure CN120197502A_ABST
Abstract
Description
Background Art
[0002] The main tasks in the current AIGC field include generating images and generating videos. The mainstream of them all adopt diffusion-type models. Such models generally have a large number of parameters. For example, the number of parameters of StableDiffusion 2 is about 1 billion; the number of parameters of DALL-E 2 is about 3.5 billion; the number of parameters of SORA is speculated to be around 3 billion. The large number of parameters makes the model have a large amount of computation when deployed, and the disk and memory capacity required for model storage also increases accordingly. In addition, AIGC models are mostly trained in the form of fp16 data type, and the original models are also mostly released in the form of fp16.
[0003] LoRAs (Low-Rank Adaptations) is a large model fine-tuning technology that records the fine-tuning results of the large model in the results of decomposed low-dimensional matrix decomposition, obtaining a LoRA file that records the fine-tuning results. Compared with the Diffusion base model (from several GB to more than a dozen GB), LoRA fine-tuned models are often smaller (from several MB to several hundred MB). By using LoRA technology, small models with personalized styles can be trained at a relatively low cost and applied to inference. In the AIGC field, a rich LoRA community culture has been formed, along with a wealth of LoRA resources. Whether it is the enthusiast community or the industrial community, the method of "base Diffusion model + rich LoRA models" is generally adopted for model inference to meet various requirements for drawing styles.
[0004] In current typical AI chips, such as Nvidia GPUs, the computing power for int8-type data is higher than that for fp16 (usually int8 is twice as high as fp16). Therefore, there have been some attempts to convert fp16-type models into int8-type models through quantization for inference deployment. However, the existing quantization tools for diffusion models on the market are complex and not easy to use, and the accuracy loss of the quantized model is relatively large. The quality of the generated images has a significant gap compared with the unquantized model. Moreover, the existing quantization tools cannot well meet the requirements of LoRA. When switching LoRA models, very time-consuming quantization needs to be performed again. In addition, the existing quantization tools only consider the quantization problem of the model itself and do not consider the adaptation problem with the existing diffusion upper-layer inference framework. Before the quantized model is deployed online, it often requires invasive modification of the original workflow code.
[0005] Therefore, there is a desire to have a new quantization method and system that can facilitate users to automatically and efficiently perform dynamic quantization on the AIGC model, and simply meet the personalized needs of users. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a dynamic quantization method and system for the AIGC model, which can overcome all or part of the above disadvantages of existing quantization tools, and provide a one-stop solution for AIGC model quantization and inference.
[0007] To this end, according to the first aspect of the embodiments of the present disclosure, a dynamic quantization method for the AIGC model is provided, including: running the AIGC original model once to obtain the initial result of the AIGC original model; based on the AIGC original model, traversing each network layer of the AIGC original model one by one, performing quantization processing on the model parameters of the traversed network layer to obtain an intermediate AIGC quantization model; running the intermediate AIGC quantization model after quantization of the traversed network layer, obtaining an intermediate running result, and comparing the structural similarity between the intermediate running result and the initial result; and retaining the quantization status of the network layer with a structural similarity greater than or equal to a predetermined threshold, and abandoning the quantization status of the network layer with a structural similarity less than the predetermined threshold, thereby obtaining the final AIGC quantization model.
[0008] According to the dynamic quantization method for the AIGC model of the present disclosure, it further includes: when running the AIGC original model once to obtain the initial result of the AIGC original model, recording the original computational density of each network layer; when running the AIGC quantization model after quantization of the traversed network layer to obtain the intermediate running result, recording the total quantization network layer computational density of quantizing the activation value input to the quantized network layer, performing quantization using the model parameters of the quantized network layer, and dequantizing the quantized output result of the quantized network layer; and retaining the quantization status with a quantization network layer computational density lower than or equal to the original computational density of the original network layer of the quantized network layer, and abandoning the quantization status with a quantization network layer computational density greater than the original computational density of the original network layer of the quantized network layer.
[0009] According to the dynamic quantization method for the AIGC model of the present disclosure, it further includes: when retaining the quantization status of the network layer with a structural similarity greater than or equal to a predetermined threshold, only saving the calibration information required for dynamic quantization.
[0010] A dynamic quantization method for an AIGC model according to the present disclosure, wherein the AIGC model is a visual network structure, so that when the user traverses each network layer of the AIGC original model one by one, the user can directly click on the network layer pointed to by the network module in the visual network structure, and set the save file name of the calibration information and adjust the threshold of the quantization loss through the drop-down menu attached to the visual network layer, and visually compare the structural similarity before and after quantization through the visualized generation result, so that the user can determine whether to retain or discard the quantization status of the traversed network layer.
[0011] The dynamic quantization method for an AIGC model according to the present disclosure further includes: for the first final AIGC quantization model, traversing the unquantized network layers in the final AIGC quantization model one by one for the second time, performing quantization processing on the model parameters of the traversed network layers to obtain an intermediate AIGC quantization model; based on the quantization status of the upstream and downstream network layers of the traversed network layer, re-recording the total quantization network layer calculation density, and performing a second comparison on the traversed network layer; and for the second comparison result, retaining the quantization status of the quantization network layer whose calculation density is lower than or equal to the original calculation density of the original network layer of the quantization network layer, and discarding the quantization status of the quantization network layer whose calculation density is greater than the original calculation density of the original network layer of the quantization network layer.
[0012] According to another aspect of the present disclosure, there is provided a dynamic quantization system for an AIGC model, including: an initial operation component that receives input data and starts the AIGC original model to obtain an initial result of the AIGC original model; a traversal quantization component that traverses each network layer of the AIGC original model one by one based on the AIGC original model, and performs quantization processing on the model parameters of the traversed network layers to obtain an intermediate AIGC quantization model; a result comparison component that receives input data, runs the intermediate AIGC quantization model after quantization of the traversed network layer, obtains an intermediate operation result, and compares the structural similarity between the intermediate operation result and the initial result; and a quantization decision component that retains the quantization status of the network layer whose structural similarity is greater than or equal to a predetermined threshold, and discards the quantization status of the network layer whose structural similarity is less than the predetermined threshold, thereby obtaining a final AIGC quantization model.
[0013] A dynamic quantization system for an AIGC model according to the present disclosure, wherein when the initial operation component runs the AIGC original model once to obtain the initial result of the AIGC original model, it records the original calculation density of each network layer; when the result comparison component runs the AIGC quantization model after quantizing the traversed network layer to obtain the intermediate operation result, it records the total quantization network layer calculation density after quantizing the activation value input to the quantized network layer, performing quantization using the model parameters of the quantized network layer, and dequantizing the quantized output result of the quantized network layer; and the quantization decision component retains the quantization state where the quantization network layer calculation density is lower than or equal to the original calculation density of the original network layer of the quantized network layer, and abandons the quantization state where the quantization network layer calculation density is greater than the original calculation density of the original network layer of the quantized network layer.
[0014] The dynamic quantization system for an AIGC model according to the present disclosure further includes: when the quantization decision component retains the quantization state of the network layer where the structural similarity is greater than or equal to a predetermined threshold, it only saves the calibration information required for dynamic quantization.
[0015] The dynamic quantization system for an AIGC model according to the present disclosure further includes a workflow management component 150, which performs visual processing on the execution of the AIGC model to form a visual network structure of the AIGC model, so that when the user traverses each network layer of the AIGC original model one by one, the user can directly click on the network layer pointed to by the network module in the visual network structure, and set the save file name of the calibration information and adjust the threshold of the quantization loss through the drop-down menu attached to the visual network layer, and visually compare the structural similarity before and after quantization through the visualized generated result, so that the user can determine whether to retain or abandon the quantization state of the traversed network layer.
[0016] A dynamic quantization system for an AIGC model according to the present disclosure, wherein the traversal quantization component performs quantization processing on the model parameters of the network layer traversed for the first time on the final AIGC quantization model and the second time on the network layers in the final AIGC quantization model that have not been quantized, to obtain an intermediate AIGC quantization model; the result comparison component re-records the total quantization network layer calculation density based on the quantization states of the upstream and downstream network layers of the traversed network layer, and performs a second comparison; and the quantization decision component, for the second comparison result, retains the quantization state where the quantization network layer calculation density is lower than or equal to the original calculation density of the original network layer of the quantized network layer, and abandons the quantization state where the quantization network layer calculation density is greater than the original calculation density of the original network layer of the quantized network layer.
[0017] The dynamic quantization system and method of the AIGC model of the present disclosure adopt a network layer screening strategy based on the image generation effect to provide lossless quantization. That is, by traversing each network layer for layer-by-layer quantization screening, considering the differences in the sensitivity of different network layers to quantization and the degree of precision loss after quantization, the overall precision loss of the model after quantization is minimized. At the same time, by adapting to popular upper-layer inference frameworks such as diffusers, ComfyUI, and sd-webui and applying them to the original model framework in a plug-in manner, it is convenient for users to perform model quantization without invasive code modification.
[0018] It should be understood that... and the above general description and the following detailed description are only exemplary and explanatory and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0020] Figure 1 is a schematic diagram of a dynamic quantization system of an AIGC model according to a first exemplary embodiment of the present disclosure.
[0021] Figure 2 is a schematic diagram of a dynamic quantization system of an AIGC model according to a second exemplary embodiment of the present disclosure.
[0022] Figure 3 is a flowchart of a dynamic quantization method of an AIGC model according to a first exemplary embodiment of the present disclosure.
[0023] Figure 4 is a flowchart of a dynamic quantization method of an AIGC model according to a second exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0025] In addition, the drawings are only schematic illustrations of the present disclosure, and the same reference numerals in the drawings denote the same or similar parts, so repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0026] The example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0027] Figure 1 is a schematic diagram of a dynamic quantization system 100 of an AIGC model according to an exemplary embodiment of the present disclosure. As Figure 1 shown, the dynamic quantization system 100 of the AIGC model includes: an initial operation component 110, a traversal quantization component 120, a result comparison component 130, and a quantization decision component 140. The initial operation component 110 receives input data and starts the AIGC original model to obtain an initial result of the AIGC original model.
[0028] The traversal quantization component 120 traverses each network layer of the AIGC original model one by one based on the AIGC original model, performs quantization processing on the model parameters of the traversed network layer, and obtains an intermediate AIGC quantization model. According to the user's settings, it is possible to select to quantize the fully connected layer (linear) or the convolutional layer (conv). The convolutional layer (conv) extracts local features from the input data (such as an image) by sliding a convolutional kernel (Kernel), and it is a core component of computer vision tasks. The fully connected layer (linear) maps the input features to the output space through matrix multiplication. The fully connected layer (linear) or the convolutional layer (conv) are two basic computational layers in a neural network. Their quantization (Quantization) is a key step in model compression and accelerating inference.
[0029] The result comparison component 130 receives the input data, runs the intermediate AIGC quantization model after quantizing the traversed network layer, obtains the intermediate operation result, and compares the structural similarity of the intermediate operation result with the initial result. Generally speaking, it is to select one network layer at a time for model parameter quantization to form an intermediate AIGC quantization model, run this intermediate AIGC quantization model with the same initial output data, and obtain the intermediate operation result obtained by running this intermediate AIGC quantization model. The SSIM (structural similarity index) is used to compare the structures of each result. The structural similarity index (SSIM) is a comprehensive index for evaluating the visual similarity of two images by quantifying the differences in three aspects: brightness, contrast, and structural information. Based on the characteristic that the human eye is sensitive to the local structural features of images, it aims to be more in line with the perception law of the human visual system. This comparison method is a conventional technical method and will not be elaborated here.
[0030] Finally, the quantization decision-making component 140 retains the quantization status of the network layers with a structural similarity greater than or equal to a predetermined threshold and abandons the quantization status of the network layers with a structural similarity less than the predetermined threshold, thereby obtaining the final AIGC quantization model.
[0031] During the operation of the intermediate AIGC quantization model or the final AIGC quantization model, since there are unquantized network layers, the calculations between the quantized network layers and the unquantized network layers have different precisions. Therefore, it is necessary to dequantize the output activation values of the quantized network layers. This dynamic quantization process performs quantization and dequantization on the input and output, which is an inevitable overhead introduced by the quantization process. For some networks, the introduction of this overhead may result in the acceleration gain obtained after quantization being lower than the overhead introduced by quantization because the computational amount is relatively low. Quantization itself is to reduce the computational intensity. Therefore, the acceleration gain generated by reducing the computational intensity is less than the overhead generated by the quantization and dequantization processes, and the original meaning of quantization is lost. Therefore, the present disclosure uses the computational density before and after quantization to make a decision on quantization. Therefore, when the initial operation component 110 runs the AIGC original model once to obtain the initial result of the AIGC original model, it records the original computational density of each network layer. Computational density generally refers to the amount of operations completed per unit time, or the ratio of the amount of operations to the amount of memory accesses, that is, the ratio of computation to memory access, also known as computational intensity. The general calculation method of computational density is the amount of operations (FLOPs) divided by the amount of memory accesses (Bytes). For a layer of a neural network, such as a fully connected layer (linear) or a convolutional layer (conv), the amount of operations generally refers to the number of floating-point operations in the forward propagation, such as multiply-add operations. Subsequently, when the result comparison component 130 runs the AIGC quantization model after quantizing the traversed network layers to obtain the intermediate operation result, it records the total computational density of the quantized network layer, including quantizing the activation values input to the quantized network layer, performing quantization on the model parameters of the quantized network layer, and dequantizing the quantized output result of the quantized network layer. Finally, the quantization decision component 140 retains the quantization state where the computational density of the quantized network layer is lower than or equal to the original computational density of the original network layer of the quantized network layer, and abandons the quantization state where the computational density of the quantized network layer is greater than the original computational density of the original network layer of the quantized network layer.
[0032] For example, dynamic quantization calculation goes through the following stages: First, the input activation data (activations) is quantized, for example, from FP16 format to int8 format. This stage represents an increase in computational effort compared to the computational effort of the original network layer. The quantized activations are used as input for the calculation of the quantized network layer. This computational process represents a stage where the computational effort is reduced or the calculation is accelerated compared to the calculation of the original network layer. Subsequently, the obtained quantized output (int8 format) is dequantized into FP16 format and passed to the next layer of network calculation. This stage represents an increase in computational effort compared to the computational effort of the original network layer. The sum of the increased computational effort caused by the quantization of these two input and output activations and the computational effort of the quantized network layer, when compared with the computational effort of the original network layer, if the sum of the computational effort is greater than or equal to the computational effort of the original network layer, it means that quantization of this network layer does not bring beneficial acceleration effects and thus needs to be discarded. Otherwise, the quantized state of the network layer is retained.
[0033] In traditional quantization methods, the quantized model parameters are directly stored after offline quantization. Generally, the volume of the quantized model file is about 50% of the size of the original file. Thus, if the original model file is 8G, the size of the quantized model is about 4G. The volume of the quantized model file is still considerable. To further make the model lighter, when the quantization decision component 140 of the present disclosure retains the quantization state of network layers with structural similarity greater than or equal to a predetermined threshold, it only saves the calibration information required for dynamic quantization. Calibration information (Calibration Data) generally refers to auxiliary information used to statistically analyze the characteristics of input data. For example, the minimum value (min) and maximum value (max) of the input data; the mean (mean) and variance (variance) of the input data; the histogram or quantiles of the data distribution. In dynamic quantization, the purpose of saving this information is to accelerate the real-time calculation of quantization parameters, rather than directly storing the quantization parameters themselves. By adopting this method, the size of the stored file of the quantized model can be reduced to several hundred KB. More than 99.9% of the storage space is saved. Before inference, based on the original model and the quantization calibration information, a quantized model can be dynamically obtained, and the time taken is only a few seconds (2 - 4 seconds). This time taken and the saving of more than 99.9% of the storage space are obviously worthwhile. This saved storage space eliminates the need for a large-capacity GPU card and reduces the capacity of the GPU card required for model implementation, thus greatly reducing the equipment cost of large models.
[0034] The total computational density of the traversed quantized network layers generally includes the computational density for quantizing the activation values input to the quantized network layer, the computational density of the quantization process of the traversed network layer, the computational density when performing calculations using the quantized model parameters of the quantized network layer, and the computational density for dequantizing the quantized output results of the quantized network layer. The total computational density described here does not include the computational density for quantizing the activation values input to the quantized network layer if the previous input layer of the quantized layer being traversed is also in a retained quantized state. Similarly, the total computational density described here does not include the computational density for quantizing the activation values output from the quantized network layer being traversed if the next input layer of the selected quantized layer is also in a retained quantized state. Therefore, after a complete traversal of the AIGC original model, there may be a network layer that is in a non-quantized state during the first traversal because its next network layer has not yet been traversed, and its total computational density includes the computational density for quantizing the activation values it outputs, resulting in the total computational density of this network layer being greater than the computational density of this network layer and being discarded. To eliminate this wrongly discarded situation, a second traversal quantization process is restarted for the final AIGC quantized model after the first traversal, and the network layers in the retained quantized state are not traversed during the second traversal process, thereby eliminating the wrongly discarded situation. Therefore, in the dynamic quantization system 100 or 200 of the AIGC model according to the present disclosure, the traversal quantization component 120 or 220 traverses one by one the non-quantized network layers in the final AIGC quantized model for the first time, performs quantization processing on the model parameters of the traversed network layers, and obtains an intermediate AIGC quantized model; the result comparison component 130 or 230 re-records the total computational density of the quantized network layers based on the quantization states of the upstream and downstream network layers of the traversed network layer, and performs a second comparison. The quantization decision component 140 or 240 retains the quantization state where the computational density of the quantized network layer is lower than or equal to the original computational density of the original network layer of this quantized network layer according to the second comparison result, and abandons the retention of the quantization state where the computational density of the quantized network layer is greater than the original computational density of the original network layer of this quantized network layer.
[0035] Figure 2 is a schematic diagram of the dynamic quantization system 200 of the AIGC model according to the exemplary second embodiment of the present disclosure. Compared with Figure 1 the dynamic quantization system 100 of the AIGC model shown in Figure 1 this, the reference signs of the same parts are similar to Figure 1 those in
[0036] Relative to Figure 1As shown, such as Figure 2 shown Figure 2 In Figure 2 , the dynamic quantization system 200 of the AIGC model may further include a workflow management component 250, which performs visualization processing on the AIGC model to form a visual network structure of the AIGC model. So that when the user traverses each network layer of the AIGC original model one by one, the user can directly click on the network layer pointed to by the network module in the visual network structure, and set the save file name of the calibration information and adjust the threshold of the quantization loss through the drop-down menu attached to the visual network layer, and visually compare the structural similarity before and after quantization through the visualization generation result, so that the user can determine whether to retain or discard the quantization status of the traversed network layer.
[0037] The workflow management component 250 can be a node-based workflow management interface ComfyUI, which is commonly used for the visualization operation of Stable Diffusion. Other types of workflow management interfaces can be used, such as AIGCStudio. The core of the ComfyUI class workflow management interface is modular nodes and DAG workflows. Users can build an image generation process by connecting different nodes. Specifically, ComfyUI decomposes the image generation process of the AIGC model into independent nodes, such as model loading, quantization parameter configuration, accuracy calibration, and inference verification modules. This provides an infrastructure for the integration of the quantization process, so that users can build a complete quantization workflow by dragging nodes. This will not be elaborated here. In the visual network structure of the AIGC model, the quantization module of the traversal quantization component 120 is embedded into the nodes of the visual network structure. Quantization modules or tools, such as OneDiff, PyTorch Quantization, etc., will not be listed one by one here. The quantization engine of the quantization module or tool is integrated with ComfyUI through JIT compilation. In this way, through ComfyUI's support for custom nodes and plugins, which may include quantization-related nodes, users can configure quantization parameters by dragging and dropping nodes, such as selecting quantization precision (INT8, FP16, etc.). Users can build a complete quantization workflow by dragging nodes, and each node corresponds to a specific quantization operation interface of quantization modules such as OneDiff. The quantization engine of the quantization module or tool is integrated with ComfyUI through JIT compilation, so that when the user selects the node to be quantized through visual operations, a quantization-aware training module is automatically injected during the model loading phase, supporting the dynamic insertion of quantization operators, setting the save file name of the calibration information and adjusting the threshold of the quantization loss through the drop-down menu attached to the visual network layer, forming the user's own quantized AIGC model, and calling the computational graph of the quantized AIGC model optimized by the quantization module or tool during the inference execution phase to achieve low-precision operation acceleration. Figure 3It is a flowchart of the dynamic quantization method 300 of the AIGC model according to the first exemplary embodiment of the present disclosure. As Figure 3 shown, at step S310, the original AIGC model is run, the activation value of the initial input is received, and the initial result of the AIGC original model is obtained. Subsequently, at step S320, a network layer of the AIGC original model is traversed, and quantization processing is performed on the model parameters of the traversed network layer to obtain an intermediate AIGC quantization model. Subsequently, in step S330, in step S331, the intermediate AIGC quantization model after quantization of the traversed network layer is run based on the activation value of the initial input to obtain an intermediate operation result, and in step S332, the intermediate operation result is compared with the initial result for structural similarity. If the comparison result in step S332 shows that the structural similarity is greater than or equal to a predetermined threshold, then in step S341, the quantization state of the selected network layer is retained, otherwise, in step S342, the quantization state of the selected network layer is discarded. Thereafter, it is determined whether the selected network layer is the last selected network layer of the original AIGC model. If so, the final AIGC quantization model is output, and the model parameters of the network layers with the retained quantization state in the final AIGC quantization model are all in the quantized state. If it is not the last selected network layer, return to step S320 to select the next network layer, and repeat steps S320 - S340.
[0038] Figure 4 It is a flowchart of the dynamic quantization method 400 of the AIGC model according to the second exemplary embodiment of the present disclosure. Compared with Figure 3 the dynamic quantization method 300 of the AIGC model shown in Figure 3 , the reference numerals of the same parts are similar to those in Figure 3 , and the difference is only that it starts with the number "S4". In addition, step S433 is added. Therefore, the description of the same parts adopts the description for Figure 3 , and will not be repeated here one by one.
[0039] As Figure 4As shown, at step S410, when running the AIGC original model once to obtain the initial result of the AIGC original model, the original calculation density of each network layer is recorded. Subsequently, at step S420, a network layer of the AIGC original model is traversed, quantization processing is performed on the model parameters of the traversed network layer to obtain an intermediate AIGC quantization model, and the quantization calculation density when performing this quantization is recorded. Then, at step S431, when running the AIGC quantization model after quantizing the traversed network layer to obtain an intermediate running result, the quantization calculation density when quantizing the activation value input to the quantized network layer, the calculation density of running the quantized network layer, and the dequantization calculation density when the activation value input to the quantized network layer is dequantized are recorded, and at step S432, the intermediate running result is compared with the initial result in terms of structural similarity. If the comparison result at step S432 shows that the structural similarity is greater than or equal to a predetermined threshold, then at step S433, it is determined whether the total calculation density of the traversed quantized network layer is lower than the original calculation density of the original network layer of this quantized network layer.
[0040] The total computational density of the traversed quantized network layer generally includes the computational density of quantizing the activation values input to the quantized network layer, the computational density of the quantization process of the traversed network layer, the computational density when performing calculations using the quantized model parameters of the quantized network layer, and the computational density of dequantizing the quantized output result of the quantized network layer. The total computational density described here does not include the computational density of quantizing the activation values input to the quantized network layer when the previous input layer of the quantized layer being traversed is also in a retained quantized state. Similarly, the total computational density described here does not include the computational density of quantizing the activation values output from the quantized network layer being traversed when the next input layer of the selected quantized layer is also in a retained quantized state. Therefore, after a complete traversal of the AIGC original model, there may be a network layer that is in a non-quantized state because its next network layer has not yet been traversed during the first traversal. Its total computational density includes the computational density of quantizing the activation values it outputs, resulting in the total computational density of this network layer being greater than the computational density of this network layer and being discarded. To eliminate this wrongly discarded situation, a second traversal quantization process is restarted for the final AIGC quantized model after the first traversal, and the network layers in the retained quantized state are not traversed during the second traversal process, thereby eliminating the wrongly discarded situation. Therefore, according to the dynamic quantization method of the AIGC model of the present disclosure, a second quantization traversal is performed, that is, for the final AIGC quantized model of the first time, the non-quantized network layers in the final AIGC quantized model are traversed one by one for the second time, and the model parameters of the traversed network layer are subjected to quantization processing to obtain an intermediate AIGC quantized model. Then, based on the quantization states of the upstream and downstream network layers of the traversed network layer, the total computational density of the quantized network layer is re-recorded, and a second comparison is performed. Finally, for the result of the second comparison, the quantization state with the computational density of the quantized network layer lower than or equal to the original computational density of the original network layer of this quantized network layer is retained, and the quantization state with the computational density of the quantized network layer greater than the original computational density of the original network layer of this quantized network layer is discarded.
[0041] Finally, in step S440, if it is determined at step S433 that the total computational density of the traversed quantized network layer is lower than the original computational density of the original network layer of this quantized network layer, then the quantization state of the selected network layer is retained at step S441; otherwise, the quantization state of the network layer being traversed is discarded at step S442. In addition, if the comparison result at step S432 shows that the structural similarity is less than a predetermined threshold, then the quantization state of the network layer being traversed is discarded at step S442.
[0042] Optionally, when preserving the quantization state of the network layer, only the calibration information required for dynamic quantization is saved. Further, the AIGC model is a visual network structure, so that when the user traverses each network layer of the AIGC original model one by one, the user can directly click on the network layer pointed to by the network module in the visualization, and set the save file name of the calibration information and adjust the threshold of the quantization loss through the drop-down menu attached to the visual network layer, and visually compare the structural similarity before and after quantization through the visualization generation result, so that the user can determine whether to retain or discard the quantization state of the traversed network layer.
[0043] In summary, for the dynamic quantization system and method of the AIGC model of the present disclosure, in order to provide lossless quantization, a network layer screening strategy based on the rendering effect is adopted, that is, layer-by-layer quantization screening is performed by traversing each network layer, considering the different sensitivities of different network layers to quantization and the differences in the degree of precision loss after quantization, so as to minimize the overall precision loss of the model after quantization. At the same time, by adapting to popular upper-layer inference frameworks, ComfyUI, and diffuserssd-webui, and applying them to the original model framework in a plug-in manner, it is convenient for users to perform model quantization without invasive code modification.
[0044] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0045] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0046] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0047] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a portable hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0048] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium having a program product capable of implementing the above method of this specification. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0049] The program product for implementing the above method according to the embodiments of the present invention can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0050] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0051] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0052] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0053] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0054] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, and are not for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0055] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered exemplary, and the true scope and concept of the present disclosure are pointed out by the claims.
Claims
1. A dynamic quantization method of an AIGC model, comprising: Run the AIGC original model once to obtain the initial results of the AIGC original model; Based on the AIGC original model, traverse each network layer of the AIGC original model one by one, perform quantization processing on the model parameters of the traversed network layer, and obtain an intermediate AIGC quantization model; Running the intermediate AIGC quantization model after quantization of the traversed network layer, obtaining the intermediate running result, and comparing the intermediate running result with the initial result for structural similarity; as well as The quantization states of the network layers whose structural similarity is greater than or equal to a predetermined threshold are retained, and the quantization states of the network layers whose structural similarity is less than the predetermined threshold are abandoned, thereby obtaining the final AIGC quantization model.
2. The dynamic quantization method of the AIGC model according to claim 1, further comprising: When running the AIGC original model once to obtain the initial results of the AIGC original model, the original computational density of each network layer is recorded; When the intermediate AIGC quantization model after the traversal network layer quantization is run and the intermediate operation result is obtained, the total quantization network layer calculation density of quantizing the activation value input to the quantization network layer, the calculation after quantizing the model parameters of the quantized network layer using the quantization, and the dequantization of the quantization output result of the quantized network layer is recorded; as well as The quantization state of the quantized network layer whose computational density is lower than or equal to the original computational density of the original network layer of the quantized network layer is retained, and the quantization state of the quantized network layer whose computational density is greater than the original computational density of the original network layer of the quantized network layer is abandoned.
3. The dynamic quantization method of the AIGC model according to claim 1 or 2, further comprising: When retaining the quantization states of the network layers whose structural similarity is greater than or equal to a predetermined threshold, only the calibration information required for dynamic quantization is saved.
4. The dynamic quantization method of the AIGC model as claimed in claim 1 or 2, wherein the AIGC model is a visual network structure, so that when the user traverses each network layer of the AIGC original model one by one, he directly clicks on the network layer pointed to by the network module in the visual network structure, and sets the save file name of the calibration information and adjusts the threshold of the quantization loss through the drop-down menu attached to the visual network layer, and visually compares the structural similarity before and after quantization through the visualization of the generated results, so that the user can determine whether to retain or abandon the quantization state of the traversed network layer.
5. The dynamic quantization method of the AIGC model according to claim 2, further comprising: For the first final AIGC quantization model, the unquantized network layers in the final AIGC quantization model are traversed one by one for the second time, and the model parameters of the traversed network layers are quantized to obtain an intermediate AIGC quantization model; Based on the quantization states of the upstream and downstream network layers of the traversed network layer, the total quantized network layer calculation density is re-recorded, and a second comparison of the traversed network layers is performed; as well as For the second comparison result, the quantization state of the quantized network layer whose computational density is lower than or equal to the original computational density of the original network layer of the quantized network layer is retained, and the quantization state of the quantized network layer whose computational density is greater than the original computational density of the original network layer of the quantized network layer is abandoned.
6. A dynamic quantization system of an AIGC model, comprising: An initial operation component receives input data and starts the AIGC original model to obtain the initial results of the AIGC original model; Traversing the quantization component, based on the AIGC original model, traversing each network layer of the AIGC original model one by one, performing quantization processing on the model parameters of the traversed network layer, and obtaining an intermediate AIGC quantization model; A result comparison component receives input data, runs the intermediate AIGC quantization model after quantization of the traversed network layer, obtains the intermediate operation result, and compares the intermediate operation result with the initial result for structural similarity; as well as The quantization decision component retains the quantization state of the network layer whose structural similarity is greater than or equal to a predetermined threshold, and abandons the quantization state of the network layer whose structural similarity is less than the predetermined threshold, thereby obtaining the final AIGC quantization model.
7. The dynamic quantization system of the AIGC model as claimed in claim 6, wherein The initial operation component records the original calculation density of each network layer when running the AIGC original model once to obtain the initial result of the AIGC original model; The result comparison component records the total quantized network layer calculation density of quantizing the activation value input to the quantized network layer, the calculation after quantizing the model parameters of the quantized network layer, and dequantizing the quantized output result of the quantized network layer when the AIGC quantization model after the traversed network layer quantization is run to obtain the intermediate operation result; and The quantization decision component retains the quantization state of the quantized network layer whose computational density is lower than or equal to the original computational density of the original network layer of the quantized network layer, and abandons the quantization state of the quantized network layer whose computational density is greater than the original computational density of the original network layer of the quantized network layer.
8. The dynamic quantization system of the AIGC model as described in claim 6 or 7, wherein the quantization decision component only saves the calibration information required for dynamic quantization when retaining the quantization state of the network layer whose structural similarity is greater than or equal to a predetermined threshold.
9. The dynamic quantization system of the AIGC model according to claim 6 or 7, further comprising: The workflow management component visualizes the AIGC model to form a visualized network structure of the AIGC model, so that when the user traverses each network layer of the AIGC original model one by one, he can directly click on the network layer pointed to by the network module in the visualized network structure, and set the save file name of the calibration information and adjust the threshold of the quantization loss through the drop-down menu attached to the visual network layer. The visualization generates results to intuitively compare the structural similarities before and after quantization, so that the user can determine whether to retain or abandon the quantization status of the traversed network layer.
10. The dynamic quantization system of the AIGC model as claimed in claim 6, wherein The traversal quantization component traverses the unquantized network layers in the final AIGC quantization model one by one for the first final AIGC quantization model, performs quantization processing on the model parameters of the traversed network layers, and obtains an intermediate AIGC quantization model; A result comparison component re-records the total quantized network layer calculation density based on the quantized states of the upstream and downstream network layers of the traversed network layer, and performs a second comparison; and The quantization decision component, for the second comparison result, retains the quantization state of the original network layer whose calculation density of the quantized network layer is lower than or equal to the original calculation density of the original network layer of the quantized network layer, and abandons the quantization state of the original network layer whose calculation density of the quantized network layer is greater than the original calculation density of the original network layer of the quantized network layer.
Citation Information
Patent Citations
Algorithm training platform based on AIGC
CN116204325A
Neural network model quantification method and device, equipment and medium
CN117114075A
Neural network model processing method and device, electronic equipment and storage medium
CN118520921A
Image generation method and device, equipment and medium
CN119579718A
Neural Network Inference Acceleration Method, Target Detection Method, Device, and Storage Medium
US20240161474A1
Cited By
Video generation model acceleration method and system based on joint optimization
CN120897104A