Large model quantification method and system
By using technologies such as greedy grid search and reverse scaling in large-scale deep learning models, the quantitative configuration is optimized, and the problems of high computing efficiency and memory usage in the inference process of large-scale deep learning models are solved, and efficient quantization processing and precision retention are achieved.
Patent Information
- Application Number
- CN202510200979.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
AI Technical Summary
Large-scale deep learning models have problems of high computing efficiency and memory usage in the inference process. Traditional quantitative methods can lead to accuracy losses while increasing computing speed and reducing storage costs.
A quantization method for large models is proposed. Through greedy grid search, the optimal smoothing factor of each linear layer is obtained, reverse scaling and quantization processing is performed, and converted to INT8 or FP8 numerical formats are converted, and the quantization configuration is optimized to reduce quantization errors through chunking and automatic parameter search.
While ensuring the accuracy close to the BF16/FP16 baseline, it greatly improves inference throughput, reduces storage and computing costs, and adapts to the needs of different models and tasks through automated search and error feedback mechanisms.
Smart Images

Figure CN120124692A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer information processing, and more particularly, to a quantization method and system for large models, especially mixture-of-experts large models. Background Art
[0002] Deep learning is a very popular machine learning method at present. It uses neural networks for feature extraction to complete tasks such as classification and generation. In neural networks, the main computational complexity comes from the matrix multiplication of parameters and inputs, and the use of different data types for operations affects the computational efficiency and result accuracy. During the calculation process, common digital representation schemes generally fall into two categories: Floating Point: A way to represent numbers in the form of "mantissa + exponent". Common ones include FP32 (32-bit floating point), FP16 (16-bit floating point), FP8 (8-bit floating point), etc. They can represent a large numerical range and can also represent a very small precision range. FixedPoint: This is a way to represent numbers by fixing the integer and decimal places. The simplest example is ordinary integers (INT8, INT4), etc. The advantage of the fixed-point format is that the calculation speed may be faster and the data storage volume is also smaller, but the numerical range and precision it can represent are more restricted than floating-point numbers.
[0003] When a deep learning model is in inference (referring to using a trained model to process actual tasks), matrix multiplication and addition operations are frequently performed. In recent years, the scale of large language models has shown explosive growth, from millions to hundreds of millions of parameters in the early stage, all the way to the current model scales of tens of billions, hundreds of billions, or even trillions of parameters. These large models have shown amazing capabilities in language understanding, code generation, and logical reasoning, but at the same time, they have also brought huge computational and memory requirements. Although hardware resources have been continuously developing, for a medium-scale model, such operations still require a lot of computing power and time; and the current development trend is that the growth rate of the model scale has exceeded the development speed of hardware computing power, often with tens of billions, hundreds of billions, or even trillions of parameters, which makes the inference cost higher and the speed slower.
[0004] In addition to the inference cost and efficiency, the parameters of the model also occupy a very large memory space. The parameters need to be stored in global memory and are read from global memory to registers for numerical operations during the calculation process. However, the global memory size of current GPU graphics cards has limitations. If stored in a higher-precision floating-point format such as FP16 or FP32, the hardware storage and energy consumption will become bottlenecks and it will be difficult to bear such large-scale inference requirements.
[0005] To address the above challenges, quantization has become a very important technical direction in deep learning inference. That is, the originally high-precision (such as FP16, FP32) weight parameters and intermediate activation values are converted into a lower-precision numerical format (such as INT8 or lower). On the premise of maintaining relatively acceptable model accuracy, the storage and operation costs are significantly reduced. Therefore, quantization can make the model smaller (mainly referring to a significant decrease in the storage space required (or the video memory / memory occupancy)), making the calculation faster. However, traditional quantization also brings potential problems: due to the use of a coarser numerical scale, the model may experience a decrease in accuracy (precision). In previous medium-sized model sizes or insensitive scenarios, INT8 may be close to the accuracy of FP16, but for current large language models with tens of billions of parameters, there may be a significant loss of accuracy.
[0006] Therefore, there is a desire to obtain a quantization method and system for large models that can avoid significant accuracy losses during quantization and improve the calculation speed.
[0007] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0008] In view of this, the present disclosure provides a quantization method and system for large models, especially for mixture-of-experts large models, which can accelerate the inference of large models with tens of billions to hundreds of billions of parameters on AI hardware with a native INT8 instruction set, and can significantly improve the inference throughput while ensuring accuracy close to the BF16 / FP16 baseline. In addition, the present disclosure adopts automated search and calibration to quickly give the optimal quantization configuration to meet the needs of different models and tasks.
[0009] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.
[0010] According to one aspect of the present disclosure, a quantization method for a large model is proposed, including: for an offline large model, performing greedy grid search for each linear layer of the model to obtain the optimal smoothing factor corresponding to each linear layer; using the obtained optimal smoothing factor corresponding to each linear layer, performing reverse scaling on the weights and activation values of the corresponding linear layer in the channel dimension to keep the multiplication result between the scaled weights and activation values unchanged, thereby obtaining the smoothed weights and activation values; and performing quantization processing on the smoothed weights and activation values in the large model to convert them into a numerical format of INT8 or FP8 or a numerical format lower than INT8 or FP8.
[0011] A quantization method for a large model according to the present disclosure, wherein the large model is a mixture-of-experts large model.
[0012] According to the quantization method of the present disclosure, it further includes: after obtaining the smoothed weights and activation values and before performing the quantization process, dividing the weight tensor of each linear layer by row dimension and / or by column dimension, or after splitting the MoE expert layer by channel or expert and then dividing by row dimension and / or by column dimension to obtain a plurality of blocks; statistically calculating the weight values of each divided block to obtain the current maximum value and the current minimum value of the weights in each divided block; for each divided block, based on the current maximum value and the current minimum value, obtaining the truncation maximum value and the truncation minimum value of the truncation interval for the divided block for different predetermined truncation factors; directly clipping all or part of the weight values greater than the truncation maximum value that exceed the truncation interval in the divided block to the truncation maximum value of the truncation interval, and directly clipping all or part of the weight values less than the truncation minimum value that exceed the truncation interval in the divided block to the truncation minimum value of the truncation interval, thereby obtaining the corresponding truncated blocks of the divided block for different predetermined truncation factors; and calculating the corresponding quantization errors for the corresponding truncated blocks of different predetermined truncation factors, and taking the truncation factor adopted by the truncated block corresponding to the minimum quantization error as the optimal truncation factor of the divided block according to the quantization error measurement index used as the calibration set.
[0013] According to the quantization method of the large model of the present disclosure, it further includes: after sorting the weight values greater than the truncation maximum value that exceed the truncation interval in the divided block, directly clipping the largest 0.1 - 1% of the extreme values of the weight values greater than the truncation maximum value to the truncation maximum value of the truncation interval, and after sorting the weight values less than the truncation minimum value that exceed the truncation interval in the divided block, directly clipping the smallest 0.1 - 1% of the extreme values of the weight values less than the truncation minimum value to the truncation minimum value of the truncation interval.
[0014] The quantization method of the large model according to the present disclosure further includes: after obtaining the smoothed weights and activation values and before performing the quantization process, counting the smoothed weight values, dividing sub-tensors of a percentage greater than a predetermined percentage of the outliers with smoothed weight values higher than a predetermined weight value into high outlier blocks, and dividing other sub-tensors into smoothed blocks; counting the weight values of each divided block, so as to obtain the current maximum value and the current minimum value of the weights in each divided block; for each divided block, based on the current maximum value and the current minimum value, for different predetermined truncation factors, obtaining the truncation maximum value and the truncation minimum value of the truncation interval for the divided block; directly clipping all or part of the weight values greater than the truncation maximum value outside the truncation interval in the divided block to the truncation maximum value of the truncation interval, and directly clipping all or part of the weight values less than the truncation minimum value outside the truncation interval in the divided block to the truncation minimum value of the truncation interval, thereby obtaining the corresponding truncated blocks for different predetermined truncation factors of the divided block; and calculating the corresponding quantization errors for the corresponding truncated blocks for different predetermined truncation factors, and taking the truncation factor adopted by the truncated block corresponding to the minimum quantization error as the optimal truncation factor of the divided block according to the quantization error measurement index used as the calibration set.
[0015] The quantization method of the large model according to the present disclosure further includes: using a predefined offline calibration set to collect the original precision outputs and the output results after quantization of each block, and calculating the comparison errors before and after quantization; comparing the comparison errors with a preset error threshold, and marking the blocks with comparison errors exceeding the standard as the blocks that need to be rolled back; and performing a rollback operation on the blocks marked as needing to be rolled back, and adjusting their weight values and activation values to the original precision or a higher precision.
[0016] The quantization method of the large model according to the present disclosure further includes: before performing the greedy grid search, performing forward inference on the large model for data collection, so as to pre-collect the input and output data of each linear layer through a complete model forward inference, and establish a snapshot of the intermediate state of the entire model; and based on the data collected by the forward inference, instructing the quantization component to perform quantization processing on each linear layer in parallel and independently.
[0017] In the quantization method of the large model according to the present disclosure, the optimal smoothing factor corresponding to each linear layer is obtained by performing a greedy grid search, and grid search or iteration is performed within a predetermined smoothing factor s range through the following formula:
[0018] Among them, W represents the high-precision weights of each linear layer, X represents the high-precision activations of each linear layer, s is a smoothing factor used to magnify W by s times and shrink X by s times, and Q(·) represents the quantization operation on the corresponding tensor. To measure the gap between the output WX after applying the predetermined smoothing factor s and quantization and the original precision value, the L2 norm or other distance metrics are usually used. For Minimization is performed to obtain the smoothing factor that can minimize the quantization loss.
[0019] According to another aspect of the present disclosure, a quantization system for a large model is further provided, including: a smoothing factor search component that performs greedy grid search for each linear layer of the model for an offline large model to obtain the optimal smoothing factor corresponding to each linear layer; a parameter smoothing component that uses the obtained optimal smoothing factor corresponding to each linear layer to perform reverse scaling on the weights and activations of the corresponding linear layer in the channel dimension to keep the multiplication result between the scaled weights and activations unchanged, thereby obtaining the smoothed weights and activations; and a parameter quantization component that performs quantization processing on the smoothed weights and activations in the large model to convert them into a numerical format of INT8 or FP8 or a numerical format lower than INT8 or FP8.
[0020] The quantization system for the large model according to the present disclosure further includes: a pruning component, and the pruning component includes: a chunking module that chunks the weight tensor of each linear layer in the row dimension and / or column dimension after obtaining the smoothed weights and activations and before performing quantization processing, or after splitting the MoE expert layer by channel or expert and then chunking in the row dimension and / or column dimension to obtain multiple chunks; a statistics module that statistics the weight values of each divided chunk to obtain the current maximum and current minimum of the weights in each divided chunk; an interval calculation module that, for each divided chunk, based on the current maximum and current minimum, obtains the truncated maximum and truncated minimum of the truncated interval for the divided chunk for different predetermined truncation factors; a pruning change module that directly prunes all or part of the weight values greater than the truncated maximum that exceed the truncated interval in the divided chunk to the truncated maximum of the truncated interval, and directly prunes all or part of the weight values less than the truncated minimum that exceed the truncated interval in the divided chunk to the truncated minimum of the truncated interval, thereby obtaining the corresponding truncated chunks for the divided chunk for different predetermined truncation factors; and a truncation factor selection module that calculates the corresponding quantization error for the corresponding truncated chunks for different predetermined truncation factors, and uses the truncation factor adopted by the truncated chunk corresponding to the minimum quantization error as the best truncation factor for the divided chunk according to the quantization error measurement index used as the calibration set.
[0021] According to the quantization system of the large model of the present disclosure, in the clipping change module, after sorting the weight values greater than the truncation maximum value that exceed the truncation interval in the divided blocks by size, directly clip the largest 0.1 - 1% of the extreme values among the weight values greater than the truncation maximum value to the truncation maximum value of the truncation interval, and after sorting the weight values less than the truncation minimum value that exceed the truncation interval in the divided blocks by size, directly clip the smallest 0.1 - 1% of the extreme values among the weight values less than the truncation minimum value to the truncation minimum value of the truncation interval.
[0022] According to the quantization system of the large model of the present disclosure, it further includes: a clipping component, the clipping component includes: a block module, after obtaining the smoothed weights and activation values and before performing quantization processing, count the smoothed weight values, divide the sub-tensors with a proportion greater than a predetermined percentage of the outliers higher than the predetermined weight value among the smoothed weight values into high outlier blocks, and divide other sub-tensors into smoothed blocks; a statistics module, count the weight values of each divided block, so as to obtain the current maximum value and the current minimum value of the weights in each divided block; an interval calculation module, for each divided block, based on the current maximum value and the current minimum value, obtain the truncation maximum value and the truncation minimum value of the truncation interval for the divided block for different predetermined truncation factors; a clipping change module, directly clip all or part of the weight values greater than the truncation maximum value that exceed the truncation interval in the divided blocks to the truncation maximum value of the truncation interval, and directly clip all or part of the weight values less than the truncation minimum value that exceed the truncation interval in the divided blocks to the truncation minimum value of the truncation interval, thereby obtaining the corresponding truncated blocks for the divided blocks for different predetermined truncation factors; and a truncation factor selection module, calculate the corresponding quantization error for the corresponding truncated blocks for different predetermined truncation factors, and use the quantization error measurement index as the calibration set to take the truncation factor adopted by the truncated block corresponding to the minimum quantization error as the optimal truncation factor for the divided block.
[0023] According to the quantization system of the large model of the present disclosure, it further includes: an error comparison component, collect the original precision output and the output result after quantization of each block using a predefined offline calibration set, and calculate the comparison error before and after quantization; a rollback marking component, compare the comparison error with a preset error threshold, and mark the blocks with the comparison error exceeding the standard as the blocks that need to be rolled back; and a rollback component, perform a rollback operation on the blocks marked as needing to be rolled back, and adjust their weight values and activation values to the original precision or a higher precision.
[0024] The quantization system of the large model according to the present disclosure further includes: a data collection component that performs forward inference on the large model to collect data before performing greedy grid search, so as to pre-collect the input and output data of each linear layer through a complete model forward inference and establish a snapshot of the intermediate state of the entire model; and a parallel instruction component that based on the data instruction parameters collected by the forward inference, the quantization component independently performs quantization processing on each linear layer in parallel.
[0025] In the quantization system of the large model according to the present disclosure, the smoothing factor search component performs greedy grid search to obtain the optimal smoothing factor corresponding to each linear layer, and performs grid search or iteration within a predetermined smoothing factor s through the following formula:
[0026] Where, W represents the high-precision weight of each linear layer, X represents the high-precision activation value of each linear layer, s is the smoothing factor, which is used to perform s-fold magnification on W and s-fold reduction on X, and Q(·) represents the quantization operation on the corresponding tensor. To measure the gap between the output WX after applying the predetermined smoothing factor s and quantization and the original precision value, the L2 norm or other distance metrics are usually used. For By minimizing, the smoothing factor that can minimize the quantization loss is obtained.
[0027] In the quantization method and system using the large model according to the present disclosure, a layer-by-layer (linear-wise) scaling search strategy is introduced. During the offline calibration process, greedy grid search is performed on each linear layer to find the optimal smoothing factor that minimizes the quantization error of the layer, instead of using a unified fixed smoothing intensity for the entire model. And when performing grid search, a relatively wide search space can be set according to experience, and then the range can be dynamically reduced according to the degree of outliers. This adaptive optimization ensures that the activation outliers of each layer can be appropriately smoothed and migrated, fully reducing the quantization error and improving the 8-bit quantization accuracy of the entire model.
[0028] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objects, features and advantages of the present disclosure will become more apparent. The following described drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1It is a block diagram of the first embodiment of a quantization system for a large model shown according to an exemplary embodiment.
[0031] Figure 2 It is a block diagram of the second embodiment of a quantization system for a large model shown according to an exemplary embodiment.
[0032] Figure 3 Shown is a schematic diagram for reducing outliers in the output projection layer of MoE FFN in the present disclosure through Hadamard transform.
[0033] Figure 4 It is a flowchart of a quantization method for a large model shown according to an exemplary embodiment.
[0034] Figure 5 It is a block diagram of an electronic device shown according to an exemplary embodiment.
[0035] Figure 6 It is a block diagram of a computer-readable medium shown according to an exemplary embodiment. Detailed implementation manners
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repeated description will be omitted.
[0037] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0038] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0039] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily inclusive of all content and operations / steps, nor are they necessarily to be executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.
[0040] It should be understood that although terms such as first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below may be referred to as the second component without departing from the teachings of the present disclosure concept. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.
[0041] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily essential for implementing the present disclosure. Therefore, they cannot be used to limit the protection scope of the present disclosure.
[0042] Figure 1 is a block diagram of a first embodiment of a quantization system for a large model shown according to an exemplary embodiment. As Figure 1 shown, the quantization system 100 of the large model includes: a smoothing factor search component 110, a parameter smoothing component 120, and a parameter quantization component 130. The smoothing factor search component 110 performs a greedy grid search for each linear layer of an offline large model to obtain the optimal smoothing factor corresponding to each linear layer. The parameter smoothing component 120 uses the obtained optimal smoothing factor corresponding to each linear layer to perform inverse scaling on the weights and activation values of the corresponding linear layer in the channel dimension to keep the multiplication result between the scaled weights and activation values unchanged, thereby obtaining the smoothed weights and activation values. The parameter quantization component 130 performs quantization processing on the smoothed weights and activation values in the large model so as to convert them into a numerical format of INT8 or FP8 or a numerical format lower than INT8 or FP8.
[0043] Specifically, in the prior art, common quantization strategies can be divided into the following two types: post-training quantization strategy (Post-Training Quantization (PTQ)) and quantization strategy during model training (Quantization-Aware Training (QAT)). PTQ directly converts weights and activations to low-precision formats after the model training is completed. It is simple and straightforward, but the effect may not be ideal enough. During the original model training, low-precision operations are not simulated, and the model weights are accustomed to high-precision representations. Once suddenly reduced to 8-bit integers, the numerical distribution difference is too large, and the model doesn't know how to "adaptively" discretize the quantization range, resulting in large deviations in inference. In large models, activations are prone to very large outliers. If not specifically processed (such as smoothing, truncation), once quantized with a unified scaling interval, the outliers will widen the overall quantization scale, thus amplifying the precision error. QAT incorporates the concept of quantization during the model training process and uses simulated quantization to enable the model to learn to adapt to low precision. Usually, it can obtain better final precision, but it will make the training process more complex, make the training unstable, and introduce some potential problems. In extremely deep networks, tiny quantization errors may be amplified during the forward and backward propagation processes, making it difficult for the training to converge or requiring complex parameter tuning. In short, in high-difficulty training tasks, adding quantization simulation will further exacerbate the convergence difficulty and lead to unstable training. Considering the parameter scale and training difficulty of large models, using QAT to add a quantization simulation process during the training of large models will make the training even more unstable. In extremely deep networks, tiny quantization errors may be amplified during the forward and backward propagation processes, making it difficult for the training to converge or requiring complex parameter tuning. In short, in high-difficulty training tasks, adding quantization simulation will further exacerbate the convergence difficulty and lead to unstable training. Therefore, the mainstream solutions for large models are mainly based on PTQ.
[0044] Due to the problem of parameter scale, large models will have outlier activations at a certain model size, and when phase shift occurs, the values of outlier features will grow very fast, which makes traditional PTQ quantization schemes no longer effective. Therefore, the current PTQ quantization schemes for large models are mainly divided into two categories: weight quantization scheme (Weight-Only quantization) and SmoothQuant W8A8 quantization. Weight-Only quantization: Only quantizes the weights of the model without quantizing the activation values, reducing the storage overhead. During the model inference process, the weights are dequantized for normal-precision operations, so that the problem of outlier activation values doesn't need to be handled. SmoothQuant W8A8 quantization performs "reverse" scaling on the activation values and weights in the channel dimension, reducing the scale of outliers in the activations, and performing fixed-point operations during inference, while reducing the storage and computational overhead.
[0045] In the inference phase of the model, Weight-Only quantization is often seen as a way to improve efficiency, especially for models with large-scale parameters, because it only performs low-bit quantization on the model weights, while the activations still maintain high-precision floating-point operations, which can reduce storage and bandwidth overhead to a certain extent. However, under the actual large-scale concurrency conditions, the low-precision format of Weight-Only cannot fully utilize the native fixed-point operation instructions of AI chips or accelerators. For example, although hardware such as NVIDIA provides complete INT8 matrix multiplication instructions, if the activation is still BF16 / FP16, additional quantization and dequantization numerical operations are required before and after the calculation, which not only increases the operator overhead, but also cannot use the faster INT8 matrix multiplication instructions, resulting in a bottleneck in throughput. Therefore, when facing the deployment of ultra-large-scale models, if the bandwidth and computing power of the activation part cannot be optimized synchronously, it is difficult to take into account both high precision and high throughput, which becomes a major bottleneck in the actual implementation process. When W8A8 is used to quantize the weights and activation values of the model, the activation values often have large outliers, which will cause the accuracy of the model to drop sharply. For this reason, however, for different layers and modules of large models, especially multi-head attention in Transformer, MoE FFN, etc., the activation distribution often fluctuates significantly. At this time, if a dynamic smoothing strategy is not implemented, the fixed smoothing factor may not be compatible with different modules, resulting in unstable inference accuracy. In conventional smoothing operations, a globally unique and fixed smoothing hyperparameter s is often used to "smooth" or "migrate" the excessive amplitude of activations on certain channels to the corresponding weights, thereby reducing the degree of outliers in the activation distribution. However, in large-scale models, especially in complex MoE architectures, the activation distributions of each layer and each expert subnetwork are very different. If only a single fixed smoothing factor is used, it will cause insufficient smoothing for layers with serious outliers, and for layers with insignificant outliers, it will destroy the normal distribution and damage the accuracy. In addition, the MoE FFN module is very sensitive to outliers, which is significantly larger than other modules. At this time, after the smoothing strategy is adopted, there is still a part of the activation that is large. After quantization, the representation ability will be further compressed, thereby destroying the model's capture and aggregation of contextual information during reasoning; similarly, the gating layer in the MoE structure may also be more drastic for the gating layer due to extreme values, causing the numerical stage to be more severe.
[0046] To this end, the present disclosure proposes a layer-wise (linear-wise) scaling search strategy. During the offline calibration process, a greedy grid search is used for each linear layer to find the optimal smoothing factor that minimizes the quantization error of the layer, instead of using a fixed smoothing strength that is uniform for the entire model. Specifically, first, as mentioned above, the smoothing factor search component 110 performs a greedy grid search for each linear layer of the offline large model to obtain the optimal smoothing factor corresponding to each linear layer. A greedy grid search is performed to obtain the optimal smoothing factor corresponding to each linear layer, and a grid search or iteration is performed within a predetermined smoothing factor s range using the following formula:
[0047] Where W represents the high-precision weight of each linear layer, X represents the high-precision activation value of each linear layer, s is the smoothing factor, which is used to perform s-fold amplification on W and s-fold reduction on X, and Q(·) represents the quantization operation on the corresponding tensor. To measure the difference between the output WX after applying a predetermined smoothing factor s and quantization and the original precision value, the L2 norm or other distance metrics are usually used. To pass Minimize the smoothing factor that minimizes the quantization loss. Therefore, the search here adopts a hypothetical quantization process to obtain the product between the quantized weight and the activation value. The product of WX after quantization is theoretically the same as before quantization, but in practice there is a difference. .gap Usually the L2 norm or other distance metrics are used.
[0048] It should be pointed out that during grid search, a wider search space can be set based on experience, and then the range can be dynamically narrowed according to the degree of outliers. This adaptive optimization ensures that the activation outliers of each layer can be properly and smoothly migrated, fully reducing the quantization error and improving the 8-bit quantization accuracy of the entire model.
[0049] Since the optimal smoothing factor corresponding to each linear layer obtained according to the present disclosure The parameter smoothing component 120 uses the obtained optimal smoothing factor corresponding to each linear layer to reversely scale the weight and activation value of the corresponding linear layer in the channel dimension to keep the multiplication result between the scaled weight and activation value unchanged, thereby obtaining the smoothed weight and activation value. Since the scale of each linear layer is small, the difference in its activation distribution is relatively small, so there will be no problem of insufficient smoothing of the outlier-significant layer, and the normal distribution of the insignificant outlier layer will not be destroyed, so the accuracy will not be damaged.
[0050] Finally, the parameter quantization component 130 performs quantization processing on the smoothed weights and activation values in the large model to convert them into numerical formats of INT8 or FP8 or numerical formats lower than INT8 or FP8. For example, the linear quantization method is usually used to map the floating-point number x to the integer range through a "scaling factor" scale (and possibly a zero point zero_point). The typical approach is:
[0051] The simplest linear mapping has a floating-point number range of ([min, max]), and it is desired to map it to INT8(([-128, 127])). The specific steps are as follows: First, calculate scale = (max - min) / (127 - (-128)) ≈ (max - min) / 255.
[0052] At the same time, determine zero_point = -round(min / scale) (which aligns the minimum value of the floating-point interval to 0).
[0053] For any floating-point number x, its quantized value is:
[0054] After quantization, the obtained integer is stored as the int8 type; during inference execution, if calculation in the floating-point domain is required, "dequantization" can be performed:
[0055] When using INT8, the calculation throughput of a single matrix multiplication can usually be twice as high as that of FP16 and FP32. Taking the Tensor Core of a certain GPU as an example (the specific numbers are subject to the peak values announced by the official): FP16 mode: Assume that it can perform 128 half-precision floating-point multiply-adds in one instruction.
[0056] INT8 mode: It can perform 256 integer multiply-adds in one instruction.
[0057] In this way, when running large-scale matrix multiplications (such as when the batch size is large), INT8 can achieve approximately twice the theoretical peak throughput of FP16. In actual scenarios, factors such as bandwidth and scheduling may also need to be considered, but generally, INT8 has an obvious speed advantage in hardware. If the same GPU can reach 200 TFLOPS (200 trillion floating-point operations per second) at FP16 precision, then in INT8, the hardware can often reach 400 TOPS (400 trillion integer operations per second), and the 200 → 400 here reflects an acceleration improvement of about twice.
[0058] The present disclosure incorporates a linear-wise smoothing factor search strategy. Through online and offline calibration, a greedy grid search is adopted to automatically find the optimal smoothing factor for each linear layer, making it more targeted and minimizing the quantization error. It eliminates the defect of the traditional SmoothQuant method. Since it adopts a globally unique and fixed smoothing hyperparameter and uses the same smoothing intensity for all layers, it is difficult to simultaneously take into account the diversity of activation distributions in each layer.
[0059] Figure 2 It is a block diagram of a second embodiment of a quantization system 200 for a large model shown according to an exemplary embodiment. Compared with the Figure 1 first embodiment shown, in addition to also including a smoothing factor search component 110, a parameter smoothing component 120, and a parameter quantization component 130, it further includes a pruning component 210. Although in the first embodiment, linear layers are smoothed layer by layer, however, if the object of smoothing is more fine-grained, it will be more targeted. Hierarchical search finds its own smoothing factor for each layer (or linear operator). The pruning component 210 performs block pruning, so that within a certain layer, the weights are divided into blocks, and a truncation threshold or a scaling factor is determined for each block separately. Optionally, the pruning component 210 can be used alone to replace the smoothing factor search component 110 and the parameter smoothing component 120. This replacement is proposed here and can be simply replaced without further elaboration. In an actual system, usually, the smoothing factor of each layer is first determined, and then "block pruning" is performed on the weights of each layer to further process local outliers.
[0060] Specifically, the block module 211 of the pruning component 210, after obtaining the smoothed weights and activation values and before performing quantization processing, divides the weight tensor of each linear layer into blocks along the row dimension and / or along the column dimension, or after splitting the MoE expert layer by channel or expert and then dividing it into blocks along the row dimension and / or along the column dimension, to obtain multiple blocks. Block division: The weight tensor is decomposed into several small blocks (blocks) along a specific dimension or in several dimensions. The weight distribution within each block is more "locally consistent", and extreme values tend to be more concentrated. For the weights of a fully connected layer , it can be divided into blocks along the row dimension, dividing the d out dimension into blocks, and divided into blocks along the column dimension, dividing the d in dimension into several blocks. For example: If d in = 4096, d out= 128, first, it can be divided into chunks by rows, resulting in 128 chunks, each with a dimension of 4096; then each chunk can be further split into 64 blocks, with each block having a size of 64. In this way, each block is a sub - matrix of 64 elements. For different model structures, the following several chunking principles can be adopted or used in combination: One is chunking by row / column dimension as described above: for example, every 64 rows or 128 rows form a block, or every 64 columns or 128 columns form a block, which ensures that the shapes of all blocks are the same and is convenient for operation. Another way is to split by channels or experts: split the MoE expert layer, with each expert corresponding to a part of the weight parameters, and it can be divided by experts, and then perform row / column chunking within the experts. There is also an adaptive chunking method: after obtaining the smoothed weights and activation values and before performing quantization processing, count the smoothed weight values, divide the sub - tensors with a percentage of outliers higher than the predetermined weight value among the smoothed weight values into high - outlier chunks, and divide the other sub - tensors into smoothed chunks. Generally speaking, if some sub - tensors contain significant outliers while others are relatively smooth, smaller chunks can be divided for the former and larger chunks for the latter, so that the part with more serious problems can obtain higher quantization flexibility.
[0061] The statistical module 212 statistically analyzes the weight values of each divided chunk, thereby obtaining the current maximum and current minimum of the weights in each divided chunk. For example, max(W) / min(W) are respectively the maximum and minimum values of the weights within the current chunk (sub - matrix).
[0062] The interval calculation module 213, for each divided chunk, based on the current maximum max(W) and current minimum min(W), obtains the truncated maximum and truncated minimum of the truncated interval for the divided chunk for different predetermined truncation factors α. For the weights W within each chunk, introduce a clipping and scaling factor α, and let
[0063] Search for the truncation parameter α block by block: Different α values will be tried during the search to obtain different truncated intervals [Wmin, Wmax]. Thus, for each block, its optimal clipping / quantization parameter α is determined respectively, so as to obtain a more suitable quantization interval within a local range. α is a scaling factor. When it is less than 1, it can "tighten" the truncation range, and when it is greater than 1, it expands the range; Wmax / Wmin are the final actual upper and lower bounds used for quantization.
[0064] The clipping change module 214 directly clips all or part of the weight values greater than the clipping maximum value that exceed the clipping interval in the divided blocks to the clipping maximum value of the clipping interval, and directly clips all or part of the weight values less than the clipping minimum value that exceed the clipping interval in the divided blocks to the clipping minimum value of the clipping interval, thereby obtaining the corresponding clipped blocks of the divided blocks for different predetermined clipping factors. Through this process, most of the weight information is retained, and only a very small number of outliers within the block are clipped, which can avoid excessive information loss caused by global clipping, so as to retain the subtle differences of the weights as much as possible. Generally, methods such as statistical quantiles or absolute value rankings are used to judge the "very small number": for example, after sorting by size, only 0.1% of the extreme values are allowed to be clipped. Specifically, after sorting the weight values greater than the clipping maximum value that exceed the clipping interval in the divided blocks by the clipping change module 214, it directly clips the largest 0.1 - 1% of the extreme values of the weight values greater than the clipping maximum value to the clipping maximum value of the clipping interval, and after sorting the weight values less than the clipping minimum value that exceed the clipping interval in the divided blocks, it directly clips the smallest 0.1 - 1% of the extreme values of the weight values less than the clipping minimum value to the clipping minimum value of the clipping interval. The range of 0.1 - 1% of the extreme values can be smaller or larger, which will not be elaborated here one by one, as long as it has no substantial impact on the model accuracy / loss. Optionally, when the clipping change module 214 performs clipping, the specific ratio of clipping can also be measured by the calibration set, and then the impact on the model accuracy / loss is observed, and it is automatically optimized, or the 3σ principle (mean ± 3 times standard deviation) is used to clip the excess part.
[0065] Finally, the clipping factor selection module 215 calculates the corresponding quantization error for the corresponding clipped blocks of different predetermined clipping factors, and takes the clipping factor adopted by the clipped block corresponding to the minimum quantization error as the optimal clipping factor α of the divided block according to the quantization error measurement index used as the calibration set. Specifically, let X represent the input activation, which will subsequently be multiplied by the weight W, and Block(.) represent the operation process corresponding to this block, then the objective function is defined as follows:
[0066] Among them, represents the operation (such as matrix multiplication or convolution) performed on the input X and the block weight W in the floating-point domain, represents the obtained after quantizing the block W with the clipping range (adjusted by α) and performing the operation. That is to say, the value obtained after the assumed quantization. Minimize the difference between the output after the assumed quantization and the original floating-point output, and use the norm to measure.
[0067] Generally speaking, by traversing or optimizing α, the truncation interval is adjusted to the optimal value, so that the quantization output approximates the floating-point result. This means that subsequent quantization will be performed within the interval [Wmin, Wmax], and weight values exceeding this interval will be truncated to Wmin or Wmax. Among them, α may take a value different from 1, so there is a chance to "scale" or "compress" the maximum and minimum values of this block, making the quantization interval more compact. Values outside the interval [Wmin, Wmax] are truncated to the boundary. If α < 1, it means that the maximum and minimum value ranges are being compressed, and the specific value is selected through search.
[0068] Through block-level clipping and automatic parameter search, the present disclosure divides the weights into blocks and automatically searches for the optimal clipping and scaling factors for each block, making the local quantization interval more compact, effectively retaining local details, and reducing the overall error. This eliminates the defect in the prior art that generally uses a unified global clipping threshold to truncate the weights, which easily leads to excessive clipping in some layers or inadequate processing of local outliers, resulting in a large amount of information loss.
[0069] Return to Figure 2 , the quantization system 200 of the large model of the present disclosure further includes: an error comparison component 310, a rollback mark component 320, and a rollback component 330.
[0070] In a model with a large number of parameters, there are significant differences in the sensitivity of different sub-modules to accuracy. For example, in the MoE-Transformer model, the Attention module requires good numerical stability under low-bit quantization, while the soft operations (such as Softmax) of the MoE gating module are extremely sensitive to accuracy, and the feed-forward calculation of the expert layer (Experts) can be processed with low bits while maintaining a certain error tolerance to improve the calculation efficiency. In view of this characteristic, the present disclosure proposes an automatic mixed-precision strategy, which can evaluate the errors of each module in real time before and after quantization, and automatically determine whether to use a higher-precision calculation for a certain module according to the set accuracy tolerance threshold, so as to achieve the optimal balance between accuracy and performance.
[0071] In the overall quantization process, the layer-by-layer automatic mixed quantization method is adopted to obtain the optimal quantization configuration for each layer. First, offline sampling and preliminary quantization are performed to quantize the currently processed module into INT8. Then, the error comparison component 310 uses a predefined offline calibration set to collect the original-precision outputs and the outputs after quantization of each block, and calculates the comparison error before and after quantization. That is, it uses a predefined offline calibration set to collect the original-precision outputs and the outputs after quantization of each module, and then calculates the comparison error before and after quantization. Subsequently, the rollback marking component 320 compares the comparison error with a preset error threshold, and marks the blocks with comparison errors exceeding the standard as the blocks that need to be rolled back. That is, it compares the error index with the preset threshold to determine whether the error is within the acceptable range. For modules with excessive errors (such as Gate modules), they are automatically marked as modules that need to be rolled back. Subsequently, the rollback component 330 performs a rollback operation on the blocks marked as needing to be rolled back, and adjusts their weight values and activation values to the original precision or a higher precision. For example, it adjusts the INT8 module to FP16 (or a higher precision). Optionally, for some coupled modules, such as the Attention module, since a single linear layer Q is connected to multiple paths, it is not possible to only check the quantization error of Q, but it is necessary to regard the Attention as a whole for error confirmation.
[0072] Finally, through the layer-by-layer automatic search and error feedback mechanism, the system can finally determine different precisions for different modules, which not only ensures the precision of the model after quantization, but also can utilize the support of the underlying hardware for low-bit precision calculations to obtain a faster inference speed.
[0073] Through real-time error evaluation, the present disclosure designs an automatic mixed precision strategy, automatically selects appropriate precisions for different modules (such as Attention, MoE gating, feed-forward calculation) (for example, some modules are rolled back to FP16), while ensuring high-speed inference, it ensures the numerical precision of key modules and the overall stability of the model. This eliminates the defects of the prior art that mostly adopts the Weight-Only or unified INT8 quantization strategy, ignores the sensitivity of each module in the large model to the quantization precision, may lead to insufficient precision of key modules (such as Attention, Gate), and thus affects the overall performance of the model.
[0074] Return to see Figure 2, the quantization system 200 of the large model disclosed in the present invention includes a data collection component 410 and a parallel instruction component 420. The traditional quantization process usually adopts a layer-by-layer quantization method, that is, the activation parameters are obtained layer by layer based on the forward reasoning step to achieve layer-by-layer quantization. In order to improve the quantization result, the quantization steps described above in the present invention are adopted. Obviously, due to the numerous quantization steps described above in the present invention, the overall quantization time may be greatly increased. In order to shorten the quantization time of the present invention, parallel quantization means are adopted for the quantization process. Specifically, a method of centrally collecting reasoning data before smoothing is adopted so that each linear layer can be independently parallelized. Specifically, the data collection component 410 performs forward reasoning on the large model for data collection before performing greedy grid search, thereby pre-collecting the input and output data of each linear layer through a complete model forward reasoning, and establishing an intermediate state snapshot of the whole model. That is, forward reasoning is performed in advance for data collection, and before actual quantization, the input and output data of each linear layer are pre-collected through a complete model forward reasoning, and an intermediate state snapshot of the whole model is established for subsequent quantization calibration. The parallel instruction component 420 instructs the parameter quantization component to perform quantization processing in parallel and independently for each linear layer based on the data collected by forward reasoning. That is, each layer is independently quantized and executed in parallel, and based on the data collected by forward reasoning, each layer can independently perform quantization operations without waiting for other layers to complete in sequence. In this way, in a multi-core or GPU multi-threaded environment, quantization processing can be performed on all layers (or multiple blocks) at the same time, thereby greatly accelerating the overall quantization process. Through this efficient parallel quantization process, the present disclosure not only achieves parallelization in layer-by-layer quantization optimization (such as smoothing factor search and block stage threshold search), but also makes full use of the parallel computing resources of the underlying hardware to ensure that when facing large models with hundreds of billions of parameters, the entire quantization process can still be completed within an acceptable time, with good engineering practicality and efficiency.
[0075] This paper proposes a process of advance data collection and independent parallel quantization of each layer, effectively utilizing multi-core or GPU multi-threaded parallel execution, significantly accelerating the overall quantization process, and ensuring engineering practicality and high efficiency when faced with large models with tens of billions of parameters. This eliminates the defects of the traditional quantization process in the existing technology, which is often processed layer by layer. Due to the numerous steps, the data collection, calibration and search before and after quantization are time-consuming, making it difficult to meet the deployment requirements of ultra-large models.
[0076] Optionally, after smoothing, the residual local spikes in each dimension can also be scattered so that the residual local spikes are scattered into different dimensions. Specifically, for the output projection layer of the MoE FFN, there is a high probability of outliers because SmoothQuant mainly targets the outlier distribution at the channel level and smooths it by scaling some activation amplitudes to the corresponding weights, which cannot handle the extreme values inside the vector. Therefore, this disclosure introduces the Hadamard transform to scatter the residual local spikes in each dimension, so that the Expert weights clear the outliers. Thus, the two can cooperate to achieve all-round smoothing from "channel large scale" to "vector local", maximizing the quantization effect of the activation / weights.
[0077] It should be noted that after using the Hadamard transform, it needs to be dynamically inserted during the inference process, unlike SmoothQuant which only completes a one-time preprocessing during offline quantization. The reason is that the scaling factor of SmoothQuant is a diagonal matrix multiplication of the input activation, so it can be merged into the weights of the previous layer, while the Hadamard transform is an ordinary matrix multiplication and cannot be merged into the weights of the previous layer and must be explicitly calculated during the inference process. However, the Hadamard transform operation itself supports fast parallelization, and its complexity is usually O(nlogn). After being accelerated on the SIMD unit of the GPU, the overall overhead is relatively low and can be fully incorporated into the real-time inference process. In actual deployment, the online Hadamard transform can be inserted into the output projection process of the MoE FFN and fused with the quantization operation for optimization. Generally, it can be divided into the following steps: Activation generation: After applying SmoothQuant, the channel-level scaling of the input can be offline merged into the weights of the previous layer, so the output obtained from the operation of the previous layer has undergone the channel-level scaling operation; Perform the Hadamard transform: Perform a fast multiplication of the activation tensor after channel balancing with the Hadamard matrix to disperse the original outliers into different dimensions; INT8 quantization: Perform a dynamic INT8 quantization operation on the transformed activation values. Since the outliers have been dispersed into different dimensions, it has good accuracy retention for quantization, and the Hadamard transform and the input INT8 quantization can be fused into a parallel computing kernel for acceleration; and perform matrix multiplication with the weights on the INT8 tensor core or SIMD instruction, and then perform dequantization to obtain the output activation values.
[0078] Figure 3 Shown is a schematic diagram for reducing outliers in the output projection layer of the MoE FFN in this disclosure. As Figure 3As shown, the input features enter the model and first pass through two parallel fully connected layers (FC layers). The outputs of these two fully connected layers are R 1 -1 , R 1 -1 , W up and W gate . Among them, W up and W gate are trainable weight matrices, and R 1 -1 , R 1 -1 is a scaling operation. Swish activation function: The output of R 1 -1 W gate passes through a Swish activation function. Swish is a popular activation function in recent years, which helps the model better capture non-linear relationships. Hadamard transform: The output of the Swish activation function is subjected to a Hadamard product (element-wise multiplication) with the output of R 1 -1 W up . This operation helps to reduce outliers. Gating mechanism: Multiply by a gating parameter R 4 , which is used to control the amount of information flow passing through and further adjust the feature representation. Final processing: The features adjusted by the gating mechanism pass through another fully connected layer R 4 -1 , W up , and finally pass through another fully connected layer R1R1 to obtain the final output. Through this structure, the Hadamard transform combined with the gating mechanism can effectively reduce outliers, thereby improving the robustness and performance of the MoE FFN. In this disclosure, the Hadamard transform is introduced to scatter the local spikes in the Expert weights. This transform uses simple addition and subtraction operations to disperse the extreme value distribution. On the premise of keeping the overall energy unchanged, it achieves finer smoothing at the vector level, effectively improving the stability and accuracy after quantization. This eliminates the defect in the prior art that mainly relies on smoothing factors to process outliers at the channel level, and often has limited effectiveness for local outlier anomalies inside the expert layer in the MoE model and insufficient precision stability.
[0079] Figure 4 is a flowchart of a quantization method for a large model shown according to an exemplary embodiment. As Figure 4 shown As Figure 4As shown, first, at step S411, for the offline large model, a greedy grid search is performed for each linear layer of the model to select the optimal smoothing factor corresponding to each linear layer. Subsequently, at step S412, using the optimal smoothing factor corresponding to each linear layer obtained, inverse scaling is performed on the weights and activation values of the corresponding linear layer in the channel dimension to keep the multiplication result between the scaled weights and activation values unchanged, thereby obtaining the smoothed weights and activation values. Finally, at step S414, quantization processing is performed on the smoothed weights and activation values in the large model so as to convert them into a numerical format of INT8 or FP8 or a numerical format lower than INT8 or FP8. It should be noted that the large model is a mixture-of-experts large model, which makes the effect of this application more prominent and superior.
[0080] Optionally, an additional step S413 can be inserted between step S412 and step S414. At step S413, through block-level pruning and automatic parameter search, the weights are divided into blocks, and the optimal pruning and scaling factor is automatically searched for each block to make the local quantization interval more compact, effectively retain local details, and reduce the overall error.
[0081] Specifically, in step S413, at step S4131, each linear layer, channel, or expert is divided into blocks in the row and / or column dimension. Specifically, after obtaining the smoothed weights and activation values and before performing quantization processing, the weight tensor of each linear layer is divided into blocks in the row dimension and / or column dimension, or after splitting the MoE expert layer by channel or expert, it is then divided into blocks in the row dimension and / or column dimension to obtain multiple blocks. Or after obtaining the smoothed weights and activation values and before performing quantization processing, the smoothed weight values are counted, and the sub-tensors with a percentage greater than the predetermined weight value of the outliers higher than the predetermined weight value among the smoothed weight values are divided into high-outlier blocks, while the other sub-tensors are divided into smoothed blocks.
[0082] Subsequently, at step S4132, the current maximum and minimum values of the block weights are counted. That is, the weight values of each divided block are counted to obtain the current maximum and current minimum values of the weights in each divided block. Then, at step S4133, the truncated maximum and minimum values of the truncation factor are calculated. That is, for each divided block, based on the current maximum and current minimum values, for different predetermined truncation factors, the truncated maximum and truncated minimum values of the truncation interval for the divided block are obtained.
[0083] Subsequently, at step S4134, the extreme values outside the maximum and minimum values are cropped. That is, all or part of the weight values greater than the truncation maximum value that exceed the truncation interval in the divided blocks are directly cropped to the truncation maximum value of the truncation interval, and all or part of the weight values less than the truncation minimum value that exceed the truncation interval in the divided blocks are directly cropped to the truncation minimum value of the truncation interval, thereby obtaining the corresponding truncated blocks of the divided blocks for different predetermined truncation factors. Optionally, at step S4134, when performing the truncation and cropping process, the weight values greater than the truncation maximum value that exceed the truncation interval in the divided blocks can be sorted by size, and then the largest 0.1 - 1% of the extreme values of the weight values greater than the truncation maximum value are directly cropped to the truncation maximum value of the truncation interval, and the weight values less than the truncation minimum value that exceed the truncation interval in the divided blocks are sorted by size, and then the smallest 0.1 - 1% of the extreme values of the weight values less than the truncation minimum value are directly cropped to the truncation minimum value of the truncation interval. The range of the 0.1 - 1% extreme values can be smaller or larger, which will not be elaborated one by one here as long as it has no substantial impact on the model accuracy / loss. Optionally, when the cropping and modification module 214 performs truncation, the specific ratio of truncation can also be measured by the calibration set, and then the impact on the model accuracy / loss can be observed, and automatic optimization can be performed, or the 3σ principle (mean ± 3 times the standard deviation) can be used to truncate the excess part.
[0084] Finally, at step S4135, the optimal truncation factor of the block is selected. That is, the corresponding quantization error is calculated for the corresponding truncated blocks of different predetermined truncation factors, and according to the quantization error measurement index used as the calibration set, the truncation factor adopted by the truncated block corresponding to the minimum quantization error is used as the best truncation factor of the divided block.
[0085] Optionally, after step S413, step S415 can be additionally added. At step S415, different precisions are determined for different modules through a layer-by-layer automated search and error feedback mechanism, thereby ensuring the precision of the model after quantization, and obtaining a faster inference speed by leveraging the support of the underlying hardware for low-bit precision calculations. That is, the post-quantization check. If the quantization effect is poor, the poorly quantized blocks are restored to the original precision. Specifically, at step S4151, the original precision output and the output result after quantization of each block are collected through the error comparison component 310 using a predefined offline calibration set, and the comparison error before and after quantization is calculated. Subsequently, at step S4152, the comparison error is compared with a preset error threshold through the rollback marking component 320, and the blocks with a comparison error exceeding the standard are marked as blocks that need to be rolled back. That is, the error metric is compared with the preset threshold to determine whether the error is within an acceptable range. For modules with an error exceeding the standard (such as the Gate module), they are automatically marked as modules that need to be rolled back. Finally, at step S4153, the rollback component 330 performs a rollback operation on the blocks marked as needing to be rolled back, adjusting their weight values and activation values to the original precision or a higher precision. For example, an INT8 module is adjusted to FP16 (or a higher precision). Optionally, for some coupled modules, such as the Attention module, since a single linear layer Q has multiple connected paths, it is not possible to only check the quantization error of Q, but the Attention needs to be regarded as a whole for error confirmation.
[0086] To demonstrate the advantages of the present disclosure over existing quantization methods, after the design and implementation of the quantization method according to the present disclosure, experimental verification was carried out on a large-scale MoE-Transformer model. The experiments mainly focused on the following aspects: Model precision comparison: Compare the model precision differences at different precisions (such as FP16, BF16, INT8, etc.) during the inference stage. Pay attention to the degree of precision loss before and after quantization, and evaluate the impact of key technologies such as the proposed automated mixed precision strategy, SmoothQuant adaptive search, and Hadamard transform on the precision.
[0087]
[0088] Inference speed and throughput comparison: Measure the inference speed (such as single forward inference latency, number of tokens inferred per second, inference throughput, etc.) at INT8 and the original precision (such as BF16 / FP16). At the same time, count the proportion of FP16 and INT8 mixing in the entire model after using the automated mixed precision strategy, and the resulting inference acceleration effect.
[0089]
[0090] Memory Occupancy and Memory Access Bandwidth: Compare the model storage overheads (video memory / memory occupancy) under different precision parameters such as INT8, FP16 / BF16, etc. In the multi-path concurrent inference scenario, evaluate the changes in the bandwidth requirements of different quantized models and whether the memory access pressure is alleviated.
[0091] Model / Configuration Accuracy Parameter Scale (100 million) Single Model VRAM Occupancy (GB) Relative VRAM Savings (%) Baseline (FP16) FP16 6710 1342 The Present Invention (Mixed Precision) INT8 6710 805 ~60
[0092] Quantization Process Time Consumption Comparison: Record the execution time of the parallel quantization process and the acceleration ratio compared to the traditional layer-by-layer serial quantization process. Compare the quantization time consumption statistics of each small module (such as Attention, MoE Gate, MoE FFN, etc.) during automated search and block pruning to verify the engineering efficiency of parallel quantization.
[0093] Quantization Process Total Number of Parameters (100 million) Quantization Duration (min) Single Model VRAM Occupancy (GB) Speedup Ratio Traditional Serial PTQ 6710 1200 1342 Parallel Quantization (The Present Invention) 6710 300 805 4
[0094] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment.
[0095] The following refers to Figure 5 to describe the electronic device 500 according to this embodiment of the present disclosure. Figure 5 The shown electronic device 500 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0096] As Figure 5 shown, the electronic device 500 is presented in the form of a general-purpose computing device. The components of the electronic device 500 may include but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different system components (including the storage unit 520 and the processing unit 510), a display unit 540, etc.
[0097] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 510, so that the processing unit 510 executes the steps according to various exemplary embodiments of the present disclosure described in this specification. For example, the processing unit 510 can execute as Figure 1 , Figure 2 shown in the steps.
[0098] The storage unit 520 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 5201 and / or a cache storage unit 5202, and may further include a read-only storage unit (ROM) 5203.
[0099] The storage unit 520 may also include a program / utility 5204 having a set (at least one) of program modules 5205. Such program modules 5205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0100] The bus 530 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0101] The electronic device 500 may also communicate with one or more external devices 500' (such as a keyboard, a pointing device, a Bluetooth device, etc.), enabling the user to interact with the electronic device 500, and / or communicate with any device that allows the electronic device 500 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through the input / output (I / O) interface 550. Also, the electronic device 500 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 560. The network adapter 560 may communicate with other modules of the electronic device 500 through the bus 530. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0102] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, as Figure 6 shown, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiments of the present disclosure.
[0103] The software product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0104] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0105] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0106] The above computer-readable medium carries one or more programs, which, when executed by the device, cause the computer-readable medium to implement the following functions: identifying the resource requirements of the artificial intelligence model task for the GPU through a task awareness mechanism; during the execution of the artificial intelligence model task, monitoring the GPU load in real time; determining the resource requirements of the GPU according to the GPU load to dynamically mount GPU resources; executing the artificial intelligence model task through the GPU resources to obtain an output result; and dynamically unloading the GPU resources when the GPU recycling policy is met.
[0107] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are different from the present embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0108] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0109] The above specifically shows and describes the exemplary embodiments of the present disclosure. It should be understood that the present disclosure is not limited to the detailed structures, setting manners or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included in the spirit and scope of the appended claims.
Claims
1. A large model quantization method, comprising: For offline large models, greedy grid search is performed on each linear layer of the model to obtain the optimal smoothing factor corresponding to each linear layer; Using the obtained optimal smoothing factor corresponding to each linear layer, the weight and activation value of the corresponding linear layer are reversely scaled in the channel dimension to keep the multiplication result between the scaled weight and activation value unchanged, thereby obtaining the smoothed weight and activation value; as well as The smoothed weights and activation values in the large model are quantized to be converted into a numerical format of INT8 or FP8 or a numerical format lower than INT8 or FP8.
2. The large model quantization method as claimed in claim 1, wherein the large model is a mixed expert large model.
3. The large model quantization method according to claim 2, further comprising: After obtaining the smoothed weights and activation values and before performing the quantization process, the weight tensor of each linear layer is divided into blocks according to the row dimension and / or the column dimension, or after splitting the MoE expert layer by channel or expert, it is divided into blocks according to the row dimension and / or the column dimension to obtain multiple blocks; Count the weight values of each divided block, so as to obtain the current maximum value and the current minimum value of the weight in each divided block; For each divided block, based on the current maximum value and the current minimum value, for different predetermined truncation factors, obtain a truncation maximum value and a truncation minimum value of a truncation interval for the divided block; Directly clipping all or part of the weight values in the divided blocks that are larger than the truncation maximum value and exceed the truncation interval to the truncation maximum value of the truncation interval, and directly clipping all or part of the weight values in the divided blocks that are smaller than the truncation minimum value and exceed the truncation interval to the truncation minimum value of the truncation interval, thereby obtaining corresponding truncation blocks for different predetermined truncation factors of the divided blocks; as well as The corresponding quantization errors are calculated for the truncated blocks corresponding to different predetermined truncation factors, and according to the quantization error measurement index as the calibration set, the truncation factor adopted by the truncated block corresponding to the minimum quantization error is used as the optimal truncation factor of the divided block.
4. The large model quantization method according to claim 3, further comprising: After sorting the weight values in the divided blocks that exceed the truncation interval and are greater than the truncation maximum value, the largest 0.1-1% extreme values among the weight values greater than the truncation maximum value are directly clipped to the truncation maximum value of the truncation interval; and after sorting the weight values in the divided blocks that exceed the truncation interval and are less than the truncation minimum value, the smallest 0.1-1% extreme values among the weight values less than the truncation minimum value are directly clipped to the truncation minimum value of the truncation interval.
5. The large model quantization method according to claim 2, further comprising: After obtaining the smoothed weights and activation values and before performing the quantization process, the smoothed weight values are counted, and the sub-tensors with outliers greater than a predetermined percentage in the smoothed weight values that are higher than a predetermined weight value are divided into high outlier blocks, and the other sub-tensors are divided into smooth blocks; Count the weight values of each divided block, so as to obtain the current maximum value and the current minimum value of the weight in each divided block; For each divided block, based on the current maximum value and the current minimum value, for different predetermined truncation factors, obtain a truncation maximum value and a truncation minimum value of a truncation interval for the divided block; Directly clipping all or part of the weight values in the divided blocks that are larger than the truncation maximum value and exceed the truncation interval to the truncation maximum value of the truncation interval, and directly clipping all or part of the weight values in the divided blocks that are smaller than the truncation minimum value and exceed the truncation interval to the truncation minimum value of the truncation interval, thereby obtaining corresponding truncation blocks for different predetermined truncation factors of the divided blocks; as well as The corresponding quantization errors are calculated for the truncated blocks corresponding to different predetermined truncation factors, and according to the quantization error measurement index as the calibration set, the truncation factor adopted by the truncated block corresponding to the minimum quantization error is used as the optimal truncation factor of the divided block.
6. The large model quantization method according to any one of claims 3 to 5, further comprising: Use the predefined offline calibration set to collect the original precision output and quantized output results of each block, and calculate the comparison error before and after quantization; Compare the comparison error with the preset error threshold, and mark the blocks with the comparison error exceeding the threshold as blocks that need to be rolled back; as well as Perform a rollback operation on the blocks marked as needing rollback, and adjust their weights and activation values to the original precision or higher precision.
7. The large model quantization method according to any one of claims 1 to 5, further comprising: Before performing the greedy grid search, forward reasoning is performed on the large model to collect data, so that the input and output data of each linear layer are collected in advance through a complete model forward reasoning, and a snapshot of the intermediate state of the whole model is established; as well as The parameter quantization component is instructed to perform quantization processing independently and in parallel for each linear layer based on the data collected by forward reasoning.
8. The quantization method of a large model as claimed in claim 1, wherein the greedy grid search is performed to obtain the optimal smoothing factor corresponding to each linear layer, and the grid search or iteration is performed within the range of the predetermined smoothing factor s by the following formula: in, W represents the high-precision weight of each linear layer, X represents the high-precision activation value of each linear layer, s is the smoothing factor, which is used to perform s-fold amplification on W and s-fold reduction on X, and Q(·) represents the quantization operation on the corresponding tensor. To measure the difference between the output WX after applying a predetermined smoothing factor s and quantization and the original precision value, To pass Minimize the smoothing factor that minimizes the quantization loss.
9. A large model quantization system, comprising: The smoothing factor search component performs a greedy grid search on each linear layer of the offline large model to obtain the optimal smoothing factor corresponding to each linear layer; The parameter smoothing component uses the obtained optimal smoothing factor corresponding to each linear layer to reversely scale the weights and activation values of the corresponding linear layer in the channel dimension to keep the multiplication result between the scaled weights and activation values unchanged, thereby obtaining the smoothed weights and activation values; as well as The parameter quantization component performs quantization processing on the smoothed weights and activation values in the large model so as to convert them into a numerical format of INT8 or FP8 or a numerical format lower than INT8 or FP8.
10. The large model quantization system according to claim 9, further comprising: A cutting component, the cutting component comprising: A block module, after obtaining the smoothed weights and activation values and before performing quantization processing, blocks the weight tensor of each linear layer by row dimension and / or by column dimension, or after splitting the MoE expert layer by channel or expert, blocks by row dimension and / or by column dimension to obtain multiple blocks; A statistical module is used to count the weight values of each divided block, so as to obtain the current maximum value and the current minimum value of the weight in each divided block; An interval calculation module, for each divided block, based on a current maximum value and a current minimum value, and for different predetermined truncation factors, obtains a truncation maximum value and a truncation minimum value of a truncation interval for the divided block; A clipping and changing module clips all or part of the weight values in the divided blocks that are larger than the truncation maximum value and exceed the truncation interval directly to the truncation maximum value of the truncation interval, and clips all or part of the weight values in the divided blocks that are smaller than the truncation minimum value and exceed the truncation interval directly to the truncation minimum value of the truncation interval, thereby obtaining corresponding truncation blocks for different predetermined truncation factors of the divided blocks; and The truncation factor selection module calculates the corresponding quantization error for the truncated blocks corresponding to different predetermined truncation factors, and uses the truncation factor adopted by the truncated block corresponding to the minimum quantization error as the optimal truncation factor of the divided block according to the quantization error measurement index as the calibration set.
Citation Information
Cited By
Large language model weight and activation combined quantification method and system
CN120409566A
Urban rail transit engineering-oriented potential safety hazard identification model compression method and system
CN120893496A
Quantization method of large language model, related equipment and computer program product
CN121503701A