AI model weight quantification method and device, equipment and storage medium

By grouping and hybrid encoding of AI model weights, and processing outliers and ordinary values respectively, the problem of outliers affecting the quantization effect is solved, achieving higher quantization accuracy and less errors.

CN120430433APending Publication Date: 2025-08-05HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411144123.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-05
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

There is an outlier in the AI model that affects the quantization effect, resulting in large errors in the weight data, and it is difficult for existing quantization methods to take into account the accuracy requirements of outliers and ordinary values.

Method used

By grouping the weights of the AI model, distinguishing the isolated group value from the ordinary value, and quantizing them using different encoding methods to establish the encoding relationship between the outlier and the ordinary value, and reducing quantization errors.

Benefits of technology

Without reducing the quantization accuracy, the number of quantized weight coded values is increased, the quantization error is reduced, and the execution accuracy of the AI model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430433A_ABST
    Figure CN120430433A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an AI model weight quantification method and device, equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the steps that the weight of an AI model is obtained, and the precision of the weight is first precision; and determining an outlier and a common value included in the weight. And calculating a first quantized value corresponding to the outlier and a second quantized value corresponding to the common value, wherein the precision of the first quantized value and the second quantized value is the first precision. The first quantized value is subjected to first type coding to obtain a first coded value, the second quantized value is subjected to second type coding to obtain a second coded value, the precision of the first coded value and the precision of the second coded value are second precision, and the second precision is lower than the first precision. And obtaining a quantized weight of the AI model based on the first coded value and the second coded value. By adopting the method and the device, the weight of the AI model can be quantified to obtain an error.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202410171956.1, and the original application date is February 5, 2024. The entire content of the original application is incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for quantifying AI model weights. Background Art

[0003] With the development of artificial intelligence (AI) technology, AI models are increasingly being used. The number of parameters involved in AI models, such as weights and activation data, has become enormous, and as a result, the storage space occupied by AI models is also increasing.

[0004] In order to reduce the storage space of the AI model, the parameters of the AI model can be quantized, that is, high-precision parameters can be converted into low-precision parameters, thereby reducing the space occupied by the storage parameters.

[0005] However, there will be a small number of outliers in the weights, which will affect the quantization effect of the weight data, that is, the weights after quantization will have large errors. Summary of the Invention

[0006] The present application provides a method, apparatus, device, and storage medium for quantizing AI model weights, which can reduce the error of quantized weight data in the AI model. The corresponding technical solutions are as follows:

[0007] In a first aspect, a method for quantizing the weight of an artificial intelligence (AI) model is provided, the method comprising: obtaining the weight of the AI model, wherein the precision of the weight is a first precision. Determining the outliers and common values included in the weight. Calculating a first quantized value corresponding to the outlier and a second quantized value corresponding to the common value, wherein the precision of the first quantized value and the second quantized value is the first precision. Performing a first type of encoding on the first quantized value to obtain a first coded value, and performing a second type of encoding on the second quantized value to obtain a second coded value, wherein the precision of the first coded value and the second coded value is a second precision, and the second precision is lower than the first precision. Based on the first coded value and the second coded value, obtaining the quantized weight of the AI model.

[0008] Here, performing the first type of encoding on the first quantized value means establishing a correspondence between the first quantized value and each encoding value of the second precision, and each encoding value corresponding to the first quantized value is the first encoding value. Performing the second type of encoding on the second quantized value means establishing a correspondence between the second quantized value and each encoding value of the second precision, and each encoding value corresponding to the second quantized value is the second encoding value.

[0009] In the scheme shown in the present application, when the weights of the AI model are quantized, the outliers and ordinary values included in the weights can be distinguished, and then the outliers are encoded by the first type of encoding, and the ordinary values are encoded by the second type of encoding. In this way, by encoding the outliers and ordinary values separately, the encoding of the outliers and the encoding of the ordinary values can be made independent of each other, thereby reducing the error in quantizing the weights. In addition, since the encoding of the outliers and ordinary values is independent of each other, the quantized weights can be represented by two sets of encoding values. That is to say, when the quantization accuracy remains unchanged, the number of encoding values that can represent the weights after quantization can be increased, that is, when the weights are quantized, the number of selected quantization values (quantization points) can be increased, and thus the scheme shown in the present application can further reduce the error in quantizing the weights.

[0010] In one achievable embodiment, determining the outliers and normal values included in the weights includes grouping the values included in the weights to obtain multiple groups of values. Determining whether each group of values satisfies an outlier existence condition. For each group of values that satisfies the outlier existence condition, determining the outliers and normal values included in each group of values. For each group of values that does not satisfy the outlier existence condition, determining each value included in each group of values as a normal value.

[0011] In the solution described in this application, the values included in the weights can be first grouped to obtain multiple groups of values. Each group of values can then be quantized and compressed separately. This can reduce the processing load of quantization and compression compared to uniformly quantizing and compressing all the values included in the weights. In addition, since the weights only include a small number of outliers, such as 1% of outliers, after the weights are grouped, there may be a large number of groups that include outliers. This allows only outliers to be determined in groups that meet the outlier existence condition, which can reduce the processing load of determining outliers included in the weights.

[0012] In one implementable manner, the above-mentioned determination of outliers and normal values included in each group of values includes: determining values in each group of values that are greater than or equal to a specified threshold as outliers, and determining values in each group of weights that are less than the specified threshold as normal values.

[0013] Among them, the designated threshold can be determined from multiple candidate thresholds, and the candidate threshold can be 0.9 times, 0.8 times, 0.7 times, etc., of the maximum value in each group of values. In one example, the outliers and ordinary values in each group can be distinguished according to different candidate thresholds, and then the outliers are quantized according to the first type of coding, and the ordinary values are quantized according to the second type of coding to obtain the weight of the quantized AI model corresponding to each candidate threshold. Then, the error of the weight of the quantized AI model of each group of data relative to the unquantized weight can be determined. Finally, the candidate threshold with the smallest corresponding error can be determined as the designated threshold. In this way, determining the designated threshold from multiple candidate thresholds can reduce the error in quantizing the weights.

[0014] In one implementable manner, the above-mentioned determination of outliers and normal values included in each group of values may further include: determining the largest first number of values in each group of values as outliers, and determining the values in each group of values other than the outliers as normal values.

[0015] Among them, the first number is less than or equal to the number of numerical values that can be represented by the first coding value. In one example, the outliers and ordinary values in each group can be distinguished according to a plurality of candidate numbers that are less than the number of numerical values that can be represented by the first coding value, and then the outliers are quantized according to the first type of coding, and the ordinary values are quantized according to the second type of coding to obtain the weight of the quantized AI model corresponding to each candidate number. Then, the error of the quantized AI model corresponding to each candidate number relative to the unquantized weight can be determined. Finally, the candidate number with the smallest corresponding error can be determined as the first number. In this way, determining the first number from multiple candidate numbers can reduce the error in quantizing the weights.

[0016] In one feasible method, the weights are stored in the form of a matrix, and the weights after quantization of the AI model are obtained based on the first coding value and the second coding value, including: determining the weight matrix after quantization of the AI model based on the positions of outliers and ordinary values in the unquantized weight matrix of the AI model, the first coding value, and the second coding value.

[0017] In one achievable embodiment, the method further includes generating an identification matrix corresponding to the quantized weight matrix, wherein the size of the identification matrix is the same as the size of the quantized weight matrix, and each element in the identification matrix is used to indicate that the element at the same position in the quantized weight matrix is a first coded value. In this way, the identification matrix can be used to distinguish between the first coded value and the second coded value in the weight matrix, thereby achieving dequantization of the weights of the AI model and improving the execution accuracy of the AI model.

[0018] In one achievable embodiment, the method further includes: setting the element at the first position of the quantized weight matrix to an identification value of the second precision, where the first position is an adjacent position to the second position, the second position is the position of the first coded value in the quantized weight matrix, and the identification value is used to indicate that the element adjacent to the identification value is the first coded value. In this way, the identification value can be used to distinguish between the first coded value and the second coded value in the weight matrix, thereby achieving dequantization of the weights of the AI model and improving the execution accuracy of the AI model.

[0019] In one achievable embodiment, the method further includes: recording a first correspondence between the first quantized value and the first coded value and a second correspondence between the second quantized value and the second coded value. In response to an execution request of the AI model, based on the first correspondence and the second correspondence, the first coded value included in the quantized weight is decoded into a first quantized value, and the second coded value included in the quantized weight is decoded into a second quantized value. Based on the decoded first quantized value and second quantized value, the calculation processing in the AI model is performed.

[0020] In one achievable embodiment, the calculating of the first quantized value corresponding to the outlier and the second quantized value corresponding to the normal value includes: clustering the outliers to obtain a second number of cluster center values, wherein the second number is less than or equal to the number of numerical values that can be represented by the first coded value; determining the second number of cluster center values as the first quantized value corresponding to the outlier; clustering the normal values to obtain a third number of cluster center values, wherein the third number is less than or equal to the number of numerical values that can be represented by the second coded value; and determining the third number of cluster center values as the second quantized value corresponding to the normal value.

[0021] In a second aspect, a device for quantifying weights of an artificial intelligence (AI) model is provided, the device comprising:

[0022] An acquisition module is used to obtain the weight of the AI model, wherein the accuracy of the weight is a first accuracy.

[0023] A determination module is used to determine the outliers and normal values included in the weights.

[0024] The calculation module is used to calculate a first quantized value corresponding to the outlier value and a second quantized value corresponding to the normal value, where the precision of the first quantized value and the second quantized value is a first precision.

[0025] The encoding module is used to perform a first type of encoding on the first quantized value to obtain a first encoding value, and perform a second type of encoding on the second quantized value to obtain a second encoding value, wherein the precision of the first encoding value and the second encoding value is the second precision, and the second precision is lower than the first precision; based on the first encoding value and the second encoding value, the weight of the AI model after quantization is obtained.

[0026] In one achievable manner, the determination module is used to: group the values included in the weights to obtain multiple groups of values; determine whether each group of values meets the outlier existence condition; and for each group of values that meets the outlier existence condition, determine the outliers and normal values included in each group of values.

[0027] In one implementable manner, the determination module is further configured to: for each group of values that does not satisfy the outlier existence condition, determine each value included in each group of values as a normal value.

[0028] In one achievable manner, the determination module is configured to: determine values in each group of values that are greater than or equal to a specified threshold as outliers, and determine values in each group of weights that are less than the specified threshold as normal values.

[0029] In one achievable manner, the determination module is used to: determine a first maximum number of values in each group of values as outliers, and determine values other than outliers in each group of values as normal values, wherein the first number is less than or equal to the number of values that can be represented by the first coding value.

[0030] In one achievable method, the weight is stored in the form of a matrix, and the encoding module is used to determine the weight matrix after quantization of the AI model based on the positions of outliers and normal values in the unquantized weight matrix of the AI model, the first encoding value, and the second encoding value.

[0031] In one achievable method, the device also includes a generation module for: generating an identification matrix corresponding to the quantized weight matrix, the size of the identification matrix being the same as the size of the quantized weight matrix, and each element in the identification matrix being used to indicate that the element at the same position in the quantized weight matrix is a first coding value.

[0032] In one achievable embodiment, the device further includes a generation module for setting the element at the first position of the quantized weight matrix to an identification value of the second precision, the first position being the adjacent position of the second position, the second position being the position of the first coded value in the quantized weight matrix, and the identification value being used to indicate that the element adjacent to the identification value is the first coded value.

[0033] In one achievable manner, the device further includes an execution module for: recording a first correspondence between the first quantization value and the first coding value and a second correspondence between the second quantization value and the second coding value; in response to an execution request of the AI model, based on the first correspondence and the second correspondence, decoding the first coding value included in the quantized weight into a first quantization value, and decoding the second coding value included in the quantized weight into a second quantization value; and executing calculation processing in the AI model based on the decoded first quantization value and the second quantization value.

[0034] In one implementable manner, the computing module is used to: cluster the outliers to obtain a second number of cluster center values, wherein the second number is less than or equal to the number of numerical values that can be represented by the first coding value; determine the second number of cluster center values as the first quantitative values corresponding to the outliers; cluster the ordinary values to obtain a third number of cluster center values, wherein the third number is less than or equal to the number of numerical values that can be represented by the second coding value; and determine the third number of cluster center values as the second quantitative values corresponding to the ordinary values.

[0035] In a third aspect, a computing device is provided, comprising a processor and a memory. The processor is configured to execute instructions stored in the memory, so that the computing device performs the method for quantizing AI model weights as described in the first aspect and any one of the implementations of the first aspect.

[0036] In a fourth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computing device, the computing device executes the method for quantizing the AI model weights as described in the first aspect and any one of the implementation methods of the first aspect.

[0037] In a fifth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method for quantizing the AI model weights as described in the first aspect and any one of the implementation methods of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a structural diagram of a computing device provided in an embodiment of the present application;

[0039] Figure 2 This is a flow chart of a method for quantifying AI model weights provided in an embodiment of the present application;

[0040] Figure 3 This is a schematic diagram of a hybrid encoding provided by an embodiment of the present application;

[0041] Figure 4 is a schematic diagram of a weight matrix and an identification matrix provided in an embodiment of the present application;

[0042] Figure 5 is a schematic diagram of a weight matrix provided in an embodiment of the present application;

[0043] Figure 6 This is a flow chart of a method for quantifying AI model weights provided in an embodiment of the present application;

[0044] Figure 7 This is a schematic diagram of AI model reasoning provided by an embodiment of the present application;

[0045] Figure 8 This is a schematic diagram of the structure of a quantization processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0047] At the beginning of the design of the Artificial Intelligence (AI) model, in order to improve the accuracy of the AI model, the weights and activation data are set to higher-precision floating-point numbers, such as FP32 and FP16. Floating-point numbers occupy more bits, such as FP32 floating-point numbers occupy 32 bits, and FP16 floating-point numbers occupy 16 bits. Therefore, floating-point numbers require more storage and computing resources when stored and calculated. Especially in some large models, such as large language models (LLM), the amount of data corresponding to the weights is huge, and the storage of the weights alone requires hundreds of GB of storage space. Among them:

[0048] An AI model refers to a mathematical model that uses methods from fields such as mathematics, statistics, computer science, and machine learning to analyze, process, predict, and optimize data with certain regularity and predictability.

[0049] Weights are numerical values used in AI models to measure the impact of different features and parameters on the model's output. AI model weights can be obtained through training algorithms and data. In the actual operation of AI models, weights are generally represented as tensors, such as matrices.

[0050] Activation data can be the output data of an AI model's intermediate layers, related to the input data fed into the AI model. During the actual operation of the AI model, activation data can also be represented as tensors, often performing tensor operations with weights.

[0051] To improve the universality of AI models, quantization compression is used to reduce the storage space required to store the models and increase their running speed. Quantization compression of AI models involves converting the weights and activation data in the AI model from high-precision data types to low-precision data types, thereby reducing the storage space and increasing the running speed of the AI model. For example, the original FP16 weights and activation data in the AI model can be quantized to INT8. This reduces the storage of the weight values from 16 bits to 8 bits. However, converting floating-point numbers to integers reduces the precision of the values. The higher the precision of the floating-point number, the larger the corresponding value range, and the greater the loss of precision when converting the floating-point number to an integer. Therefore, quantization compression can have a significant impact on the accuracy of the AI model.

[0052] Because the numerical range of activation data in AI models is generally larger than the numerical range of weights, and there is a certain proportion of outliers in the activation data, the error caused by the quantization of activation data is larger, and the impact on the accuracy of the AI model is also greater than the quantization of weights. Among them, outliers, also known as escape values, refer to one or more values in a set of values that are significantly different from the other values, making these values far away from the overall numerical distribution.

[0053] In most AI models, the amount of weight data is much larger than the amount of activation data. Therefore, quantizing the weights can bring better quantization effects to the AI model, and the impact on the accuracy of the AI model is less than quantizing the activation data. Therefore, in order to ensure the accuracy of the AI model in related technologies, the activation data can be kept at a higher precision and the weight quantization value can be quantized to a lower precision. For example, the activation data is quantized to FP16 floating point numbers and the weights are quantized to INT4 integer numbers. During the inference process of the AI model, the weights can be dequantized to a value with the same precision as the activation data, and then calculated with the activation data.

[0054] However, there are also a small number of outliers in the weights, which can affect the quantization effect of the weights, that is, the weights after quantization will have large errors. In one example, most weights follow a probability distribution similar to the normal distribution, with a large number of normal weights concentrated near 0. Then there are a small number of outliers with particularly large values. These values will greatly affect the quantization accuracy of the majority of normal values. However, these outliers also have a high contribution to the accuracy of the AI model, making it difficult to quantize the weights to take into account both outliers and normal values.

[0055] In addition, the lower the precision of the weight, the fewer the number of numerical values that can be represented by the corresponding coding value of the weight after quantization. Due to the limited representation space (4 bits can only represent 16 numbers), using a fixed coding format will cause large precision errors and cannot achieve good generalization and compatibility. When using a non-fixed coding method (such as clustering 16 quantization points using a clustering algorithm), the problem of too long optimization time will be encountered.

[0056] The quantization method for AI model weights provided in the embodiment of the present application can quantize outliers and common values in the weights separately, and perform mixed encoding on the outliers and common values after quantization. This can reduce the impact of outliers on common values during quantization, and increase the number of weights that can be represented by the encoding of the quantized weights through mixed encoding, thereby reducing the error in quantizing the weights.

[0057] The AI model weight quantization method provided in the embodiment of the present application is implemented in a computing device. Figure 1 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. Figure 1 As shown, the computing device 100 may include: a bus 102, a processor 104, a memory 106, and optionally the computing device 100 may also include a communication interface 108. The processor 104, the memory 106 and the communication interface 108 communicate with each other through the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 100. The computing device 100 may be a device for running a model, a terminal or a server. When the computing device 100 is a terminal, the computing device 100 includes but is not limited to a desktop computer, a mobile phone, a notebook, a tablet computer, etc. When the computing device 100 is a server, the computing device 100 may be a separate server, such as a server in a server cluster, a physical machine, or a virtual machine or container virtualized by virtual technology.

[0058] The bus 102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 1 The fact that only one line is used in the figure does not mean that there is only one bus or only one type of bus. Bus 102 may include a path for transmitting information between various components of computing device 100 (eg, memory 106, processor 104, communication interface 108).

[0059] The processor 104 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP). The processor 104 may further include a decoding unit, a computing unit, and the like.

[0060] Memory 106 may include volatile memory, such as random access memory (RAM). Memory 106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). All of the above memories are referred to as global memory.

[0061] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the quantization method of the AI model weight provided in the embodiment of the present application, for example, obtaining the weight of the AI model, wherein the precision of the weight is the first precision. Determine the outliers and normal values included in the weight. Calculate the first quantization value corresponding to the outlier and the second quantization value corresponding to the normal value, and the precision of the first quantization value and the second quantization value is the first precision. Perform a first type of encoding on the first quantization value to obtain a first encoding value, and perform a second type of encoding on the second quantization value to obtain a second encoding value, wherein the precision of the first encoding value and the second encoding value is the second precision, and the second precision is lower than the first precision. Based on the first encoding value and the second encoding value, obtain the weight of the quantized AI model, etc.

[0062] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0063] Figure 2 This is a flow chart of a method for quantifying AI model weights provided in an embodiment of the present application. The method can be performed by Figure 1 The computing device shown executes, see Figure 2 , the method comprising:

[0064] Step 201: Obtain the weight of the AI model.

[0065] The precision of the obtained weight is the first precision, which can be the precision of the weight before the AI model is quantized. The precision can be reflected by the corresponding data type, and the higher the bit number, the higher the precision. For example, the precision corresponding to the FP32 value can be higher than the precision corresponding to the FP16 value.

[0066] In one example, the weights in the AI model are generally represented in matrix form, and the weights obtained in step 201 may be the weight matrices included in each layer of the AI model.

[0067] Step 202: Determine the outliers and normal values included in the weights.

[0068] After obtaining the weights of the AI model, the outliers and common values included in the weights can be determined. For example, the outliers and common values included in the weights can be determined by a threshold pre-set by a technician, that is, the elements in each weight matrix that are greater than or equal to the threshold are determined as outliers, and the elements in the weight that are less than the threshold are determined as common values. For another example, the outliers and common values included in the weights can be determined by numerical distribution, such as sorting the numerical values included in the weights from large to small and determining the values in the top 5% as outliers.

[0069] In one example, the amount of data corresponding to the weights of the AI model is huge. If all elements included in each weight matrix are directly quantized, the required amount of calculation may be relatively large. Therefore, in order to improve the quantization of the weights in subsequent steps 203-204, the weights can be grouped in step 202 to obtain multiple groups of values. For example, a row of elements or a column of elements in the weight matrix can be determined as a group of values, or the weight matrix can be divided into blocks, and the values included in each matrix block can be determined as a group of values. Among them, the number of values in each group of values can be pre-set by a technician, or the way of grouping the weight matrix can be pre-set, and the size is not limited in the embodiment of the present application.

[0070] Since the number of outliers included in the weights is small, after the weights are grouped, it is possible that most groups do not contain outliers. Therefore, after grouping the weights to obtain multiple groups of values, it is possible to first determine whether there are outliers in each group of values.

[0071] In one example, whether there are outliers in each group of values can be determined based on a set outlier existence condition. The outlier existence condition can be that the ratio of the maximum value to the average value in a group of values is greater than a set ratio threshold, or that the difference between the maximum value and the average value in a group of data is greater than a set difference threshold. If a group of values meets the outlier existence condition, it can be determined that there are outliers in the group of values, and the outliers and normal values included in the group of values can be further determined.

[0072] In one possible implementation, the process of determining the outliers and normal values included in each set of weights may include:

[0073] Processing 1: The values in each group of values that are greater than or equal to the specified threshold are determined as outliers, and the values in each group of values that are less than the specified threshold are determined as normal values.

[0074] Among them, the specified threshold can be pre-set by a technician. In one example, a technician can set multiple candidate thresholds, and execute the quantization method of the AI model weight provided in the embodiment of the present application once or multiple times according to each candidate threshold to obtain the weight after quantization. Then the mean square error (MSE) corresponding to the weight after quantization processing relative to the unquantized weight can be calculated for each group. Or perform AI model reasoning based on the obtained weight after quantization processing, and calculate the MSE of the reasoning result relative to the reasoning result of the unquantized AI model. Finally, the candidate threshold with the smallest corresponding MSE can be determined as the specified threshold. For setting multiple candidate thresholds, they can be 0.9 times, 0.8 times, 0.7 times, etc., respectively, of the maximum value in the group of values.

[0075] Processing 2: The first number of weights with the largest values in each group of values are determined as outliers, and the weights other than the outliers in each group of weights are determined as normal values.

[0076] Among them, the first number can be pre-set by a technician, and the first number is less than or equal to the number of values that can be represented by encoding the weights after quantization. For example, the encoding format of the weights after quantization is INT4, including 4 bits, and the number of values that can be represented is 16, then the corresponding first number is less than or equal to 16. Alternatively, the first number of weights with the largest quantization error corresponding to each value can be determined as outliers. The quantization error corresponding to each value refers to the difference between each value and the corresponding quantized value after uniform quantization of the outliers and ordinary values included in a set of data.

[0077] In one example, a technician may set multiple candidate numbers that are smaller than the number of numerical values that can be represented by the first coding value. Then, the outliers and normal values in each group can be distinguished according to each candidate number to execute the quantization method of the AI model weights provided in the embodiment of the present application to obtain the quantized weights. Then, the MSE corresponding to the quantized weights relative to the unquantized weights can be calculated for each group. Or, AI model reasoning can be performed based on the obtained quantized weights, and the MSE of the reasoning result relative to the reasoning result of the unquantized AI model can be calculated. Finally, the candidate number with the smallest MSE can be determined as the first number. Or, AI model reasoning can be performed based on the obtained quantized weights, and the MSE of the reasoning result relative to the reasoning result of the unquantized AI model can be calculated. Finally, the candidate number with the smallest MSE can be determined as the first number.

[0078] In the embodiment of the present application, after the weights are grouped, the numerical range corresponding to each group of values may be narrowed, so quantizing each group of values separately can further reduce the quantization error.

[0079] Step 203: Calculate a first quantized value corresponding to the outlier value and a second quantized value corresponding to the normal value, where the precision of the first quantized value and the second quantized value is the first precision.

[0080] In step 202, the weights are divided into two parts: an outlier part and a normal value part. In steps 203 and 204, the outlier part and the normal value part can be quantized separately. In this way, the quantization of the normal value will not be affected by the outliers, thereby reducing the error of the quantized weights.

[0081] Determining a first quantization value of a first precision corresponding to an outlier refers to compressing the number of values corresponding to the outlier to a second number. The second number of outliers is the first quantization value, which may also be referred to as a first quantization point. The second number is less than or equal to the number of values that can be represented by the first encoding value. For example, when the first encoding value is INT4, the second number may be 16 or 15.

[0082] In one example, clustering may be performed on the outliers to obtain a second number of cluster center values, and the second number of cluster center values may be determined as a second number of quantization points.

[0083] Determining the second quantization value of the first precision corresponding to the outlier value means compressing the number of values corresponding to the normal value to a third number. This third number of outliers is the second quantization value, which can also be called a second quantization point. The third number is less than or equal to the number of values that can be represented by the second encoding value. For example, when the second encoding value is INT4, the third number can be 16 or 15.

[0084] In one example, clustering processing may be performed on the common values to obtain a third number of cluster center values, and the third number of cluster center values may be determined as a third number of quantization points.

[0085] The clustering process can be implemented by a k-means clustering algorithm, a density-based spatial clustering algorithm (DBSCAN), etc., which will not be described in detail in the embodiments of the present application. In addition, since each weight is grouped in step 202, when clustering the outliers or common values included in each group of data, the computational complexity of the clustering process can be reduced, thereby improving the efficiency of the clustering process.

[0086] Step 204: Perform a first type of encoding on the first quantized value to obtain a first encoded value, and perform a second type of encoding on the second quantized value to obtain a second encoded value.

[0087] The precision of the first and second coded values is the second precision, which is lower than the first precision. For example, when the data type of the value corresponding to the first precision is FP16, the data type of the value corresponding to the second precision may be INT4 or IN8.

[0088] Performing the first type of encoding on the first quantized value means establishing a correspondence between the first quantized value and each coded value of the second precision, where each coded value corresponding to the first quantized value is the first coded value. Performing the second type of encoding on the second quantized value means establishing a correspondence between the second quantized value and each coded value of the second precision, where each coded value corresponding to the second quantized value is the second coded value.

[0089] In step 204, the first quantization value and the second quantization value may be mixedly encoded to obtain a code value corresponding to the first quantization value and a code value corresponding to the second quantization value. The first code value and the second code value may be the same code value. That is, in an embodiment of the present application, the same code value may be the code value of the first quantization value or the code value corresponding to the second quantization value. In one example, two coding tables may be generated for each set of weights in step 202. In one coding table, the correspondence between the first quantization value and the code value may be recorded, and in another coding table, the correspondence between the second quantization value and the code value may be recorded.

[0090] Since the first quantized value and the second quantized value have the same precision in the code values, that is, the first code value and the second code value have the same bit positions, the first code value and the second code value represent the same numerical range and the same code value. Therefore, in this application, a separate code can be set to distinguish whether the code value belongs to the first code value (outlier) or the second code value (normal value).

[0091] Figure 3 This is a schematic diagram of a hybrid coding provided by an embodiment of the present application. Figure 3 As shown, by specifying a threshold (thres), the weight can be divided into two parts, one part is the outlier value and the other part is the normal value. For the outlier value, 15 first quantization points can be determined in the outlier value, and the correspondence between the 15 first quantization points and the 4-bit encoding value can be recorded through the outlier value coding table. For the normal value, 15 second quantization points can be determined in the normal value, and the correspondence between the 15 second quantization points and the 4-bit encoding value can be recorded through the normal value coding table. Among them, "0000" can be used as a flag bit to distinguish whether the 4-bit encoding value is used to encode an outlier value or a normal value. The distinction method is not explained here, and detailed description is provided below. In implementation, for each group of weights in the weight matrix corresponding to each layer in the AI model, two corresponding coding tables can be generated. When executing the AI model, the weights of each group can be dequantized according to the coding table corresponding to the weights of each group.

[0092] In addition, since the quantized weights can be mapped through two sets of coding tables. That is to say, when the quantization accuracy remains unchanged, the number of coded values that can represent the weights after quantization can be increased, that is, when the weights are quantized, the number of selected quantization values (quantization points) can be increased. For example, the data type of the quantized weights is INT4. If the outliers and ordinary values are mapped respectively through two sets of coding tables, the quantized weights can be represented by up to 32 numerical values. That is, when each set of weights is quantized, 32 quantization points can be selected. Therefore, the embodiment of the present application can also increase the number of quantization points, which can further reduce the error of quantizing the weights.

[0093] Step 205: Obtain the quantized weight of the AI model based on the first coding value and the second coding value.

[0094] After obtaining the first coded value corresponding to the first quantized value and the second coded value corresponding to the second quantized value, the quantized weight of the AI model can be generated based on the first coded value and the second coded value. That is, the outliers and normal values included in the weight are encoded into the corresponding first coded value and the second coded value respectively.

[0095] In one example, the weight is a weight matrix, and the weight matrix after quantization of the AI model is determined based on the positions of outliers and normal values in the unquantized weight matrix of the AI model, the first coding value, and the second coding value.

[0096] In implementation, after obtaining the first coding value corresponding to the first quantization value and the second coding corresponding to the second quantization value, the various numerical values included in the weight can be encoded, the outliers included in the weight matrix can be encoded as the corresponding first coding values, and the ordinary values included in the weight matrix can be encoded as the corresponding second coding values, and then the quantized weight matrix can be obtained.

[0097] After determining the weight matrix after quantization processing, an identification matrix corresponding to the weight matrix after quantization processing can also be generated. The size of the identification matrix is the same as the size of the weight matrix after quantization processing. Each element in the identification matrix is used to indicate whether the coding value at the same position in the weight matrix after quantization processing is an outlier.

[0098] Figure 4 This is a schematic diagram of a weight matrix and an identification matrix provided in an embodiment of the present application. Figure 4 As shown in the figure, after obtaining the quantized weight matrix, an identification matrix can be generated based on the positions of the coded values corresponding to outliers and the positions of the coded values corresponding to normal values in the weight matrix. The identification matrix is a sparse matrix. An element "1" in the identification matrix indicates that the element at the same position in the weight matrix is the code for the outlier value, and an element "0" in the identification matrix indicates that the element at the same position in the weight matrix is the code for the normal value.

[0099] Alternatively, after determining the weight matrix after quantization processing, the numerical value of the specified position of the weight matrix after quantization processing can also be set to an identification value of the second precision, where the specified position is the adjacent position of the coding value corresponding to the outlier in the weight matrix after quantization processing, and the identification value is used to indicate the position of the coding value corresponding to the outlier in the weight matrix after quantization processing.

[0100] Figure 5 This is a schematic diagram of another weight matrix and identification matrix provided in the embodiment of the present application. Figure 5As shown, after obtaining the weight matrix after quantization processing, the position of the code value corresponding to the outlier in the weight matrix can be determined, and the element on the left or right of the code value in the weight matrix can be set to an identification value, such as "0000". Since the amount of data of ordinary values is much larger than that of outliers, setting the codes corresponding to ordinary values on the left and right of the outliers as identification values will not affect the accuracy of the AI model. In one example, for outliers located in odd-numbered columns, such as outliers in the 1st, 3rd, and n+1th columns in the weight matrix, the element adjacent to the right of the outlier can be set to an identification value. For outliers located in even-numbered columns, such as outliers in the 2nd, 4th, and 2nth columns in the weight matrix, the element adjacent to the left of the outlier can be set to an identification value.

[0101] After obtaining the quantized AI model, the quantized AI model can be sent to the computing device that executes the AI model. In addition, quantization information can also be sent to the corresponding computing device. For example, the quantization information may include a coding table, weight grouping information, etc. The computing device that executes the AI model can also be Figure 1 The computing device shown.

[0102] In implementation, the computing device, in response to an AI model execution request, can obtain quantization information from the AI model. Before executing the calculation of activation data and weights corresponding to each layer, the weights can be dequantized based on the quantization information. That is, the first and second correspondences mentioned above are used to convert each encoded value in the quantized weights into first and second quantization values. Calculations are then performed based on the high-precision first and second quantization values and the high-precision activation data to complete the inference of the AI model.

[0103] In the embodiment of the present application, by performing mixed encoding on the outliers and common values in the weights, the influence of the outliers on the quantization of the common values can be avoided, and the number of quantization points to be selected can be increased, thereby reducing the error of the selected quantization points. In this way, during the inference process of the AI model, after the first encoded value and the second encoded value are respectively dequantized into the corresponding first quantization value and the second quantization value, the inference of the AI model can be completed based on the first quantization value and the second quantization value with higher precision and smaller error, which can improve the accuracy of the AI model.

[0104] Figure 6 This is a flow chart of a method for quantizing weights and anti-AI model weights provided in an embodiment of the present application. Figure 6 As shown, this method can include the model quantization process and the model inference process, see Figure 6 , the method comprising:

[0105] Model quantization process:

[0106] Step 601: Divide the original tensors in the AI model into fine-grained groups.

[0107] The original tensor refers to the tensor corresponding to the unquantized weights in the AI model, for example, it can be a weight matrix. The granularity of the segmentation of the original tensor can be pre-set by the technician. For example, the weight matrix can be divided into blocks to obtain multiple matrix blocks, and the elements included in each matrix block are a group. Among them, if the size of the block for the weight matrix is too small, the access and computing performance of the computing device may be wasted. If the size of the group for the weight matrix is too large, the subsequent processing amount for quantizing the weights may increase, and the quantization accuracy will also be reduced. Therefore, in one example, the technician can set the size of the block for the weight matrix according to the computing performance of the computing unit included in the processor in the computing device, such as the computing size supported by the matrix operation unit or the vector calculation unit.

[0108] Step 602: Search for quantization point information for each group independently based on the segmented joint optimization algorithm.

[0109] For each group including outliers, quantization point information corresponding to the outliers and normal values in each group may be determined. The quantization point information may include coding tables corresponding to the outliers and normal values in the group.

[0110] The following takes the data type of the quantized weight as INT4 as an example to illustrate the implementation methods of the two segmented joint optimization algorithms.

[0111] Method 1: An optimization algorithm based on distinguishing the designated thresholds corresponding to the cluster values and the common values. The optimization algorithm may include the following steps:

[0112] Step 1: Divide the original tensor into reasonable sub-blocks (group).

[0113] Step 2: For the value of each group, set the threshold thres1 / thres2, and divide the data [min, max] in the group into two parts. The first part is called outlier or outlier, which is [thres2, max]. The value of this part is large, but the amount of data is small; the second part is called normal or ordinary value, which is [min, thres1]. The distribution of the values in this part is more concentrated, and the amount of data is larger.

[0114] Among them, thres1 can be the value closest to the candidate threshold among the values in the group that are smaller than the above candidate threshold, and thres2 can be the value closest to the candidate threshold among the values in the group that are larger than the above candidate threshold.

[0115] Step 3-1: A clustering algorithm is used to optimize the high-precision quantization points for outliers. The number of quantization points is set to 15 (a total of 16 in the 4-bit encoding space), and a 4-bit encoding value is assigned to each quantization point. In this way, one encoding value can be left to distinguish between outliers and normals.

[0116] Step 3-2: For normal values, optimize the non-uniform encoding scheme. Since normal data can be approximated to a normal distribution, or can be approximated to a normal distribution through scaling, existing mainstream non-uniform encoding schemes such as NF4 can be used. Alternatively, if the values within some groups are relatively evenly distributed, the int4 encoding scheme can be used.

[0117] Step 4: Use the method of minimizing the MSE optimization to find reasonable thresholds and coding schemes within each segment.

[0118] Here, mse refers to mean square error, a reasonable threshold value may be the threshold value specified in the above embodiment, and the coding scheme within each segment refers to the mixed coding corresponding to each group including outliers.

[0119] Step 4 refers to setting multiple candidate thresholds and executing Step 2, Step 3-1, and Step 3-2 once or multiple times for each candidate threshold. After obtaining the quantized weights in Step 2, Step 3-1, and Step 3-2 corresponding to each candidate threshold, calculate the mse corresponding to the quantized weights. Here, mse can be input mse or output mse. Input mse represents the quantization error of the weight, that is, the mean square error of the quantized weight relative to the unquantized weight. Output mse represents the mean square error of the calculation result obtained after matrix calculation of the quantized weight and the corresponding activation data relative to the calculation result corresponding to the unquantized weight.

[0120] After obtaining the mse corresponding to the candidate threshold of each group, the candidate threshold of the minimum mse corresponding to each group can be determined as the specified threshold used by each group to distinguish between the cluster value and the normal value, thereby further reducing the error in quantizing the weight.

[0121] Method 2: An optimization algorithm based on distinguishing outliers from normal values by specifying the number of outliers. The optimization algorithm may include the following steps:

[0122] Step 1: Divide the original tensor into reasonable sub-blocks (group).

[0123] Step 2: For each group value, select n values as outlier values.

[0124] Here, n is the number of candidates, n≤16, for example, it can be 8 or 15.

[0125] Step 3-1: A clustering algorithm is used to optimize the high-precision quantization points for outliers. The number of quantization points is set to 15 (a total of 16 in the 4-bit encoding space), and a coding value is assigned to each quantization point. In this way, one coding value can be left to distinguish between outliers and normals.

[0126] Step 3-2: For normal data, optimize existing non-uniform coding schemes. Since normal data can be approximated to a normal distribution, or can be approximated to a normal distribution through scaling, existing mainstream non-uniform coding schemes such as NF4 can be used. Alternatively, if the values within some groups are relatively evenly distributed, the int4 coding scheme can be used.

[0127] Step 4: Use the method of minimizing the MSE optimization to find reasonable thresholds and coding schemes within each segment.

[0128] Step 4 refers to setting multiple values of n, and executing the above Step 2, Step 3-1, and Step 3-2 once or multiple times according to each value. After completing the quantized weights in Step 2, Step 3-1, and Step 3-2 corresponding to each value, calculate the mse corresponding to the quantized weights. Here, mse can be input mse or output mse. Input mse represents the quantization error of the weight, that is, the mean square error of the quantized weight relative to the unquantized weight. Output mse represents the mean square error of the calculation result obtained after matrix calculation of the quantized weight and the corresponding activation data relative to the calculation result corresponding to the unquantized weight.

[0129] After obtaining the mse corresponding to multiple candidate numbers of each group, the candidate number of the minimum mse corresponding to each group can be determined as the first number used by each group to distinguish between cluster values and common values, thereby further reducing the error in quantizing the weights.

[0130] Step 603: quantize the original tensor based on the quantization point information and set a flag.

[0131] Among them, for the setting of the flag bit, refer to the above Figure 4 The content will not be repeated here.

[0132] Model inference process:

[0133] Step 604: Configure the quantization point information into the inference device.

[0134] The inference device can be a different device from the computing device that performs steps S1 to S3. The quantized AI model and the corresponding grouping information of the quantized point information weights can be stored in the inference device. In practice, upon receiving an execution request for the AI model, the inference device can perform the following steps.

[0135] Step 605: Read the quantized weights and perform an inverse quantization decoding operation using the quantization point information of the group to implement an inverse quantization operation from a low-bit coded value to a high-bit quantized value.

[0136] In implementation, the quantized weights can be read in groups according to the weight grouping information to obtain multiple groups of encoding values. Figure 7 As shown in the AI model reasoning diagram, after obtaining multiple sets of coding values, the processor of the inference device can input each set of coding values (W) and the corresponding quantization point information (W-info) in the memory into the corresponding decoding unit. The decoding unit is a hardware unit included in the processor in the inference device, which can be used to decode each set of coding values according to the quantization point information to obtain a high-precision quantization value corresponding to each set of coding values. In another example, the decoding of the quantized weights can also be implemented by software.

[0137] Step 606: Read activation data and perform inference calculation process.

[0138] Continue to refer to Figure 7 After the decoding unit completes decoding of each set of data, the decoded quantized value and the corresponding activation data (A) input value can be used to calculate the weight and activation data matrix. Steps 605 and 606 can be repeated until all inference calculations in the AI model are completed, and the obtained inference results can be returned to the memory.

[0139] In an embodiment of the present application, by performing mixed encoding of outliers and common values in the weights, the influence of outliers on the quantization of common values can be avoided, and the number of quantization points to choose from can be increased. During the inference process of the AI model, the quantized weights can be decoded according to the quantization point information to obtain the corresponding first and second quantization values. The AI model can then be further inferred based on the high-precision, less-error first and second quantization values, thereby improving the accuracy of the AI model.

[0140] Figure 8 This is a schematic diagram of the structure of a quantization device for an artificial intelligence AI model weight provided in an embodiment of the present application, such as Figure 8 As shown, the device includes:

[0141] The acquisition module 810 is used to obtain the weight of the AI model, wherein the accuracy of the weight is the first accuracy, which can be specifically used to implement the acquisition function of the above step 201 and other implicit steps.

[0142] The determination module 820 is used to determine the outliers and normal values included in the weights, and can be used to implement the determination functions of the above step 202 and other implicit steps.

[0143] The calculation module 830 is used to calculate the first quantization value corresponding to the outlier and the second quantization value corresponding to the normal value. The precision of the first quantization value and the second quantization value is the first precision. Specifically, it can be used to implement the calculation function of the above step 203 and other implicit steps.

[0144] The encoding module 840 is used to perform a first type of encoding on the first quantized value to obtain a first encoding value, and to perform a second type of encoding on the second quantized value to obtain a second encoding value, wherein the precision of the first encoding value and the second encoding value is the second precision, and the second precision is lower than the first precision; based on the first encoding value and the second encoding value, the weight of the AI model after quantization is obtained, which can be specifically used to implement the encoding functions of the above steps 204 and 205 and other implicit steps.

[0145] In one achievable manner, the determination module 820 is used to: group the values included in the weights to obtain multiple groups of values; determine whether each group of values satisfies the outlier existence condition; and for each group of values that satisfies the outlier existence condition, determine the outliers and normal values included in each group of values.

[0146] In one achievable manner, the determination module 820 is further configured to: for each group of values that does not satisfy the outlier existence condition, determine each value included in each group of values as a normal value.

[0147] In one achievable manner, the determination module 820 is configured to: determine values in each group of values that are greater than or equal to a specified threshold as outliers, and determine values in each group of weights that are less than the specified threshold as normal values.

[0148] In one achievable manner, the determination module 820 is used to: determine the largest first number of values in each group of values as outliers, and determine the values in each group of values other than the outliers as normal values, and the first number is less than or equal to the number of values that can be represented by the first coding value.

[0149] In one achievable method, the weight is stored in the form of a matrix, and the encoding module 840 is used to determine the weight matrix of the AI model after quantization based on the positions of the outliers and the normal values in the unquantized weight matrix of the AI model, the first encoding value, and the second encoding value.

[0150] In one achievable method, the device also includes a generation module for: generating an identification matrix corresponding to the quantized weight matrix, the size of the identification matrix being the same as the size of the quantized weight matrix, and each element in the identification matrix being used to indicate that the element at the same position in the quantized weight matrix is a first coding value.

[0151] In one achievable embodiment, the device further includes a generation module for setting the element at the first position of the quantized weight matrix to an identification value of the second precision, the first position being the adjacent position of the second position, the second position being the position of the first coded value in the quantized weight matrix, and the identification value being used to indicate that the element adjacent to the identification value is the first coded value.

[0152] In one achievable manner, the device further includes an execution module for: recording a first correspondence between the first quantization value and the first coding value and a second correspondence between the second quantization value and the second coding value; in response to an execution request of the AI model, based on the first correspondence and the second correspondence, decoding the first coding value included in the quantized weight into a first quantization value, and decoding the second coding value included in the quantized weight into a second quantization value; and executing calculation processing in the AI model based on the decoded first quantization value and the second quantization value.

[0153] In one implementable manner, the calculation module 830 is used to: cluster the outliers to obtain a second number of cluster center values, wherein the second number is less than or equal to the number of numerical values that can be represented by the first coding value; determine the second number of cluster center values as the first quantitative values corresponding to the outliers; cluster the ordinary values to obtain a third number of cluster center values, wherein the third number is less than or equal to the number of numerical values that can be represented by the second coding value; determine the third number of cluster center values as the second quantitative values corresponding to the ordinary values.

[0154] The division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, the functional modules in the various embodiments of the present application can be integrated into one processor, or they can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. In addition, the quantization device for the artificial intelligence AI model weights provided in the above embodiments and the quantization method embodiment for the artificial intelligence AI model weights belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0155] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions for enabling a terminal device (which can be a personal computer, mobile phone, or network device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.

[0156] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method for quantizing the weights of an artificial intelligence (AI) model provided in the present application.

[0157] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the quantization method of the artificial intelligence AI model weight provided in the embodiment of the present application.

[0158] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is there a limit on quantity and execution order. It should also be understood that although the following description uses the terms "first", "second", etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. In this application, the term "at least one" means one or more, and the term "plurality" means two or more.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for quantifying the weight of an artificial intelligence (AI) model, characterized in that: The method comprises: Obtaining a weight of the AI model, wherein the precision of the weight is a first precision; determining outliers and common values included in the weights; Calculating a first quantized value corresponding to the outlier and a second quantized value corresponding to the normal value, where the precision of the first quantized value and the second quantized value is the first precision; Performing a first type of encoding on the first quantized value to obtain a first encoded value, and performing a second type of encoding on the second quantized value to obtain a second encoded value, wherein the precision of the first encoded value and the second encoded value is a second precision; Based on the first coding value and the second coding value, the quantized weight of the AI model is obtained.

2. The method according to claim 1, characterized in that Determining the outliers and normal values included in the weights includes: Grouping the values included in the weights to obtain multiple groups of values; Determine whether each group of values meets the conditions for the existence of outliers; For each group of values that meets the outlier existence condition, outliers and normal values included in each group of values are determined.

3. The method according to claim 2, characterized in that The method further comprises: For each group of values that does not satisfy the outlier existence condition, each value included in the group of values is determined to be a normal value.

4. The method according to claim 2 or 3, characterized in that Determining the outliers and normal values included in each group of values includes: The values in each group of values that are greater than or equal to a specified threshold are determined as outliers, and the values in each group of weights that are less than the specified threshold are determined as normal values.

5. The method according to claim 2 or 3, characterized in that Determining the outliers and normal values included in each group of values includes: A first maximum number of values in each group of values is determined as outliers, and values in each group of values other than the outliers are determined as normal values, wherein the first number is less than or equal to the number of values that can be represented by the first coding value.

6. The method according to any one of claims 1 to 5, characterized in that The weight is stored in the form of a matrix, and obtaining the quantized weight of the AI model based on the first coding value and the second coding value includes: The quantized weight matrix of the AI model is determined based on the positions of the outlier and the normal value in the unquantized weight matrix of the AI model, the first coding value, and the second coding value.

7. The method according to claim 6, characterized in that The method further comprises: Generate an identification matrix corresponding to the quantized weight matrix, where the size of the identification matrix is the same as the size of the quantized weight matrix, and each element in the identification matrix is used to indicate that the element at the same position in the quantized weight matrix is the first coding value.

8. The method according to claim 6, characterized in that The method further comprises: The element at the first position of the quantized weight matrix is set to an identification value of the second precision, where the first position is an adjacent position of the second position, the second position is the position of the first coded value in the quantized weight matrix, and the identification value is used to indicate that the element adjacent to the identification value is the first coded value.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Recording a first correspondence between the first quantized value and the first encoded value and a second correspondence between the second quantized value and the second encoded value; In response to an execution request of the AI model, based on the first correspondence and the second correspondence, decoding the first encoded value included in the quantized weight into a first quantized value, and decoding the second encoded value included in the quantized weight into a second quantized value; Based on the first quantization value and the second quantization value obtained by decoding, calculation processing in the AI model is performed.

10. The method according to any one of claims 1 to 9, characterized in that The calculating the first quantized value corresponding to the outlier and the second quantized value corresponding to the normal value includes: Performing clustering processing on the outliers to obtain a second number of cluster center values, wherein the second number is less than or equal to the number of numerical values that can be represented by the first coded value; Determining the second number of cluster center values as first quantized values corresponding to the outliers; Performing clustering processing on the common values to obtain a third number of cluster center values, wherein the third number is less than or equal to the number of numerical values that can be represented by the second coded value; The third number of cluster center values is determined as the second quantized values corresponding to the common values.

11. The method according to any one of claims 1 to 10, characterized in that The second precision is a floating-point precision lower than the first precision, or the second precision is an integer precision lower than the first precision.

12. A device for quantifying the weight of an artificial intelligence (AI) model, characterized in that: The device comprises: An acquisition module, configured to acquire a weight of the AI model, wherein the precision of the weight is a first precision; A determination module, configured to determine outliers and common values included in the weights; a calculation module, configured to calculate a first quantized value corresponding to the outlier value and a second quantized value corresponding to the normal value, wherein the precision of the first quantized value and the second quantized value is the first precision; An encoding module is used to perform a first type of encoding on the first quantized value to obtain a first encoding value, and to perform a second type of encoding on the second quantized value to obtain a second encoding value, wherein the precision of the first encoding value and the second encoding value is the second precision; based on the first encoding value and the second encoding value, obtain the quantized weight of the AI model.

13. The device according to claim 12, characterized in that The determining module is configured to: Grouping the values included in the weights to obtain multiple groups of values; Determine whether each group of values meets the conditions for the existence of outliers; For each group of values that meets the outlier existence condition, outliers and normal values included in each group of values are determined.

14. The device according to claim 13, characterized in that The determining module is further configured to: For each group of values that does not satisfy the outlier existence condition, each value included in the group of values is determined to be a normal value.

15. The device according to claim 13 or 14, characterized in that The determining module is configured to: The values in each group of values that are greater than or equal to a specified threshold are determined as outliers, and the values in each group of weights that are less than the specified threshold are determined as normal values.

16. The device according to claim 13 or 14, characterized in that The determination module is used to: determine a first maximum number of values in each group of values as outliers, and determine values in each group of values other than the outliers as normal values, wherein the first number is less than or equal to the number of values that can be represented by the first coding value.

17. The device according to any one of claims 12 to 16, characterized in that The weights are stored in the form of a matrix, and the encoding module is used to: The quantized weight matrix of the AI model is determined based on the positions of the outlier and the normal value in the unquantized weight matrix of the AI model, the first coding value, and the second coding value.

18. The device according to claim 17, characterized in that The device further includes a generating module, configured to: Generate an identification matrix corresponding to the quantized weight matrix, where the size of the identification matrix is the same as the size of the quantized weight matrix, and each element in the identification matrix is used to indicate that the element at the same position in the quantized weight matrix is the first coding value.

19. The device according to claim 17, characterized in that The device further includes a generating module, configured to: The element at the first position of the quantized weight matrix is set to an identification value of the second precision, where the first position is an adjacent position of the second position, the second position is the position of the first coded value in the quantized weight matrix, and the identification value is used to indicate that the element adjacent to the identification value is the first coded value.

20. The device according to any one of claims 12 to 19, characterized in that The device further includes an execution module, configured to: Recording a first correspondence between the first quantized value and the first encoded value and a second correspondence between the second quantized value and the second encoded value; In response to an execution request of the AI model, based on the first correspondence and the second correspondence, decoding the first encoded value included in the quantized weight into a first quantized value, and decoding the second encoded value included in the quantized weight into a second quantized value; Based on the first quantization value and the second quantization value obtained by decoding, calculation processing in the AI model is performed.

21. The device according to any one of claims 12 to 20, characterized in that The computing module is configured to: Performing clustering processing on the outliers to obtain a second number of cluster center values, wherein the second number is less than or equal to the number of numerical values that can be represented by the first coded value; Determining the second number of cluster center values as first quantized values corresponding to the outliers; Performing clustering processing on the common values to obtain a third number of cluster center values, wherein the third number is less than or equal to the number of numerical values that can be represented by the second coded value; The third number of cluster center values is determined as the second quantized values corresponding to the common values.

22. The device according to any one of claims 12 to 21, characterized in that The second precision is a floating-point precision lower than the first precision, or the second precision is an integer precision lower than the first precision.

23. A computing device, characterized in that The computing device includes a processor and a memory; The processor is configured to execute instructions stored in the memory, so as to enable the computing device to perform the method according to any one of claims 1 to 11.

24. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 11.

25. A computer-readable storage medium, characterized in that The method comprises computer program instructions, and when the computer program instructions are executed by a computing device, the computing device performs the method according to any one of claims 1 to 11.