Large model mixing precision quantification method and system based on dynamic tree quantification and application

By adopting a hybrid precision quantization method of dynamic tree quantization in large models, the problems of high memory usage and large quantization error in the existing technology are solved, and higher quantization accuracy and better generalization capabilities are achieved.

CN119962594APending Publication Date: 2025-05-09SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410174036.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing large-scale model quantization scheme relies on calibration data sets, and has poor generalization capabilities and cannot effectively reduce the memory usage in model training and inference. In addition, traditional fixed-point quantization will lead to large quantization errors when the data distribution is uneven.

Method used

A hybrid precision quantization method based on dynamic tree quantization is adopted, and the exponential digits are reduced through non-standard floating-point representation, the decimal expression ability is increased, the quantized bit allocation is dynamically adjusted, the data distribution characteristics are adapted, and the outlier and nonlinear operator data are maintained in a high-precision format, and only the linear operator data is dynamically quantized.

Benefits of technology

It effectively reduces the memory usage during the training and inference of large models, reduces quantization errors, improves the quantization accuracy, and does not rely on calibration data sets, which improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004701714500000081
    Figure BDA0004701714500000081
  • Figure HDA0004701714510000011
    Figure HDA0004701714510000011
  • Figure HDA0004701714510000021
    Figure HDA0004701714510000021
Patent Text Reader

Abstract

The invention discloses a large model mixing precision quantification method based on dynamic tree quantification. The method comprises the following steps: step 1, dividing input feature points into outliers and non-outliers; 2, extracting an outlier from the matrix sequence of the feature points, and extracting a row where a weight corresponding to the outlier is located; or processing and quantizing the non-outlier data used for the linear operator; non-outlier data applied to a nonlinear operator is not quantized; 3, performing matrix multiplication on the outlier data and the corresponding weight of the outlier; or, performing inverse quantization on non-outlier data stored in the linear operator in a quantized manner, and performing matrix multiplication operation; directly performing matrix multiplication on non-outlier data in the nonlinear operator; and 4, combining a non-outlier matrix multiplication operation result and an outlier matrix multiplication operation result through addition, and outputting a final result. The invention further discloses application and system of the method and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large model quantization, and relates to a large model mixed precision quantization method, system and application based on dynamic tree quantization. Background Art

[0002] Currently, large models based on the transformer architecture have become the most important research direction in the current AI field. Due to the large number of parameters of large models, a large amount of memory space is required during training and reasoning, which poses a great challenge to hardware performance. In order to save the memory required during training and reasoning, quantization is a commonly used technology. In the convolutional neural network (CNN) quantization scheme, it is mainly divided into training-aware quantization (QAT) and post-training quantization (PTQ). Due to factors such as the high training overhead of large models and the difficulty in collecting pre-training data, the quantization of large models mainly adopts the post-training quantization scheme.

[0003] After the convolutional neural network is quantized, a certain calibration data set is required to determine the data distribution, such as KLD [1] , AdaRound [2] , QDrop [3] For large models, this often brings two problems: 1. Calibration data sets are often difficult to obtain. Since large models use a large number of data sets for pre-training, it is often difficult for users to obtain all the data to obtain calibration data sets; 2. Calibration for certain specific data sets often only performs well on that data set, while performance will degrade on other data sets.

[0004] Among the quantization schemes currently used in large models, there is a type of scheme that only quantizes weights, such as GPTQ [4] 、QLoRA [5] Although such solutions can effectively reduce the size of model parameters, they cannot reduce the size of intermediate data during runtime. Another solution is to quantize the intermediate data into 8-bit fixed-point numbers (Int8), such as llm.int8 [6] The disadvantage of this method is that it will evenly quantize the floating point range of the data into 256 parts, but in reality, the data distribution is often not evenly distributed, which is bound to cause a large quantization error. However, since floating point numbers are inherently non-uniform, it can bring better quantization effects.

[0005] In each layer of the neural network, there are often some data points that are far away from the main distribution of data, which are called outliers. In CNN, the outlier method (relative entropy KLD, percentage Percentile) is often used to avoid the impact of outliers, but in large models, outliers often have a greater impact on data results. [6] , so outliers cannot be simply ignored.

[0006] Traditional floating-point representation has a large number of bits used to represent the exponent of the floating-point number, but the mantissa is often not precise enough.

[0007] Specifically, the current quantization schemes have the following problems: (1) They rely heavily on calibration datasets and have poor generalization capabilities; (2) Many large model schemes only quantize weights, but do not quantize the intermediate data features (featuremaps) in the network, resulting in a large memory footprint; (3) Traditional fixed-point quantization uses uniform quantization, which can cause large quantization errors when data distribution varies greatly; (4) Traditional low-bit floating-point numbers need to use some bits for the exponent, resulting in limited accuracy in representing decimals. When data distribution is mainly concentrated in decimal places, using low-bit floating-point numbers for quantization will result in large quantization losses. Summary of the invention

[0008] In order to solve the shortcomings of the prior art, the purpose of the present invention is to provide a large model mixed precision quantization method, system and application based on dynamic tree quantization. The present invention aims to reduce the high memory usage during large model training and reasoning through a quantization method.

[0009] The method of the present invention adopts a dynamic tree quantization method and uses a non-standard floating point representation to characterize data, which reduces the number of exponent bits, increases the decimal expression capability, and improves accuracy. Dynamic tree quantization can produce lower quantization errors when processing both small and large values. Unlike data types with fixed exponents and fractions, dynamic tree quantization uses a data type with dynamic exponents and fractions that can change with each number. It consists of four parts, such as Figure 1 As shown in the figure: 1. The first bit of the data type is used to indicate the sign, 1 indicates a negative sign, and 0 indicates a positive sign; 2. The number of subsequent zero bits indicates the size of the exponent; 3. The first bit set to 1 is the indicator bit, indicating that all subsequent values ​​are used for linear quantization; 4. Linear quantization. Dynamic tree quantization has better absolute and relative errors for non-uniform distributions. [7] .

[0010] In the past two years, the decode-only Transformer model has gradually become the mainstream in practical applications. Therefore, the method of the present invention mainly considers the decode-only model architecture. For the encoding-only and encoding-decoding model structures, the method of the present invention can be easily promoted due to the high consistency of the internal operators and the overall model architecture.

[0011] The present invention provides a large model mixed precision quantization method based on dynamic tree quantization, the quantization method comprising the following steps:

[0012] Step 1: Divide the input feature points into outlier points and non-outlier points;

[0013] Step 2: extracting the outlier points from the matrix sequence of the feature points, and extracting the row where the weight corresponding to the outlier points is located; or,

[0014] The non-outlier data applied to linear operators are processed and quantified and stored; the non-outlier data applied to nonlinear operators are not quantified;

[0015] Step 3: Perform matrix multiplication on the outlier data and the weights corresponding to the outliers; or,

[0016] Dequantize the non-outlier point data stored in the linear operator and perform matrix multiplication. Perform matrix multiplication directly on the non-outlier point data in the nonlinear operator.

[0017] Step 4: Combine the matrix multiplication results of the non-outlier points and the matrix multiplication results of the outlier points by addition, and output the final result.

[0018] The mixed precision in the present invention means that the quantization precision in the quantization process is not a single fixed one. For example, the combination of fp16 and int8 in the present invention is two forms of precision, namely the mixed precision mentioned above. In the quantization process, the model uses multiple data types with different bit numbers at the same time, which can speed up the operation and reduce memory usage.

[0019] In step 1, the outlier points refer to data points that are not within the preset data distribution range and will affect the data results of the large model; the non-outlier points refer to data points that are included in the preset data range.

[0020] In step 2, the outlier data is not quantized, and the outliers are extracted from the feature point matrix by dimension; or,

[0021] For non-outlier data applied to linear operators, the feature map is divided into blocks to obtain the maximum absolute value of each block, and normalized by block. The floating-point data of each block is converted into fixed-point numbers through dynamic tree quantization and stored.

[0022] No quantization is performed on non-outlier data applied to nonlinear operators.

[0023] In a specific embodiment of the present invention, the outlier data is kept in fp16 format and is not quantized; the non-outlier data format applied to linear operators is converted from fp16 to int8 for storage, and then dequantized and converted back to fp16 during operation; the non-outlier data applied to nonlinear operators is kept in fp16 format and is not quantized.

[0024] The dynamic tree quantization representation method includes four parts, including a sign representation bit, an exponent representation bit, an indicator bit, and a linear quantization bit; the sign representation bit represents the positive and negative signs of the data, 1 represents a negative sign, and 0 represents a positive sign; the exponent representation bit represents the size of the exponent through the number of 0 bits; the indicator bit is the first 1 that appears after the exponent representation bit, and is used to indicate that all values ​​after the indicator bit are used for linear quantization; the linear quantization bit indicates that the corresponding numerical value is represented by a linear quantization method.

[0025] In step three, the dequantization refers to remapping the fixed-point numbers corresponding to the non-outlier data obtained in step two into floating-point numbers through a pre-constructed mapping table, and performing matrix multiplication operations.

[0026] The quantized fixed-point number is related to the selected quantization bit number. After the quantization bit number is determined, each original floating-point number will be mapped to a fixed fixed-point number, and a mapping table is constructed accordingly.

[0027] The present invention also provides an application of the above-mentioned mixed precision quantization method in reducing the memory consumption of large model training inference.

[0028] The present invention also provides a hardware system for implementing the above-mentioned mixed precision quantization method, and the hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned method is implemented.

[0029] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned mixed precision quantization method is implemented.

[0030] The beneficial effects of the present invention include:

[0031] The method of the present invention does not require data calibration like traditional quantization methods. It adopts a dynamic tree quantization method, utilizes non-standard floating-point representation and dynamically adjusted quantization strategy, so that the quantization process is independent of the specific distribution of the data, thereby effectively avoiding the phenomenon that the quantization effect is inconsistent on different data sets.

[0032] The present invention solves the influence of outliers in large models on quantization errors and adopts a mixed precision quantization solution. Non-outliers and outliers that are input as nonlinear operators are not quantized and are directly processed in a high-precision format. At the same time, a dynamic tree quantization method is used for non-outliers, thereby effectively reducing the quantization error.

[0033] The present invention adopts dynamic tree quantization for the feature points of each layer in the neural network, and can more accurately adapt to the distribution characteristics of the data by dynamically adjusting the bit allocation of quantization. In particular, for those feature data with uneven distribution, dynamic tree quantization can reduce the number of exponent bits and increase the number of mantissa bits through its non-standard floating point representation method, thereby providing higher quantization accuracy under the same bit limit, and can effectively reduce the quantization error compared with fp8. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 It is a schematic diagram of the dynamic tree quantization data representation method of the present invention.

[0036] Figure 2 It is the overall data flow chart of the present invention.

[0037] Figure 3 It is a flow chart of outlier processing of the present invention.

[0038] Figure 4 This is the flow chart of converting fp16 into int8 of the present invention.

[0039] Figure 5 This is the flowchart of converting int8 into fp16 of the present invention. DETAILED DESCRIPTION

[0040] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.

[0041] The present invention aims to solve the pain points of large memory consumption and difficulty in quantizing models in the field of large model deployment. It innovatively abandons the traditional floating-point representation scheme and adopts a method based on dynamic tree quantization to quantize the feature map of the large model. The main technical points are: (1) Mixed precision quantization: In order to solve the influence of large model outliers on quantization errors, a mixed precision quantization scheme is adopted to effectively reduce the quantization error. (2) Normalization: In order to reduce the problem of insufficient quantization accuracy caused by excessive data range, the method of the present invention first normalizes the data and scales it to a certain range, thereby ensuring that a limited number of bits are used to represent a larger original data range. (3) Dynamic tree quantization: In view of the fact that most of the normalized data is relatively concentrated, a dynamic tree quantization method is adopted. Better results are achieved than traditional floating-point representation.

[0042] The dynamic tree quantization representation method of the present invention includes four parts, including a sign representation bit, an exponent representation bit, an indication bit, and a linear quantization bit; the sign representation bit represents the positive and negative signs of the data, 1 represents a negative sign, and 0 represents a positive sign; the exponent representation bit represents the size of the exponent through the number of 0 bits; the indication bit is the first 1 that appears after the exponent representation bit, and is used to indicate that all values ​​after the indication bit are used for linear quantization; the linear quantization bit indicates that the corresponding numerical value is represented by a linear quantization method.

[0043] by Figure 1 For example: Figure 1 In the figure, the first bit is the sign bit, which is 1, indicating a negative sign; the 2nd and 3rd bits are the exponent bits, including two 0 bits, that is, the exponent is -2 times; the 4th bit is 1, indicating that starting from this bit, the subsequent values ​​are used for linear quantization; the 5th to 8th bits are the binary number 1001, which is converted to decimal as 9. The 5th to 8th bits are 4 bits in total, one bit represents 2 to the power of n, and 4 bits are 2^4=16, starting from 0, that is, 0-15, and the linear quantization is 9 / 15=0.6.

[0044] In the present invention, there are mainly two types of operators in the decode only large model: nonlinear operators and linear operators (mainly including matrix multiplication and element-wise multiplication (ElementwiseMul)). Since quantizing nonlinear operators often causes large quantization losses, the nonlinear operators in the entire large model are quantized in fp16 format. This type of operator includes: rotational position encoding (ROPE), Reshape, Softmax, Silu, normalization (Norm), etc. For linear operators, when the output of the operator on the previous layer of the operator is a floating point number, the output is first quantized into int8 for storage, and then the operand is converted from int8 to floating point number for calculation before the actual calculation. In summary, in a specific implementation, the data flow of the entire decode module is as follows: Figure 2 As shown in the figure, when the data encounters a linear operator (Linear layer, matrix multiplication MatMul, ElementwiseMul), the data is first converted from fp16 to int8, and then converted back to fp16 for calculation during the calculation process.

[0045] Figure 2 middle,

[0046] Input: The input data is 32-bit floating point numbers (fp32), which is the original data format of the large model.

[0047] Dynamic tree quantization representation: Use non-standard floating point representation to represent data, including sign bit, exponent bit, indicator bit and linear quantization part. This step is the initial stage of quantization process, and its purpose is to adapt to the non-uniform distribution of data in dynamic tree quantization technology.

[0048] Nonlinear and linear operator processing:

[0049] Nonlinear operators (such as ROPE, Softmax, SiLu, etc.) are directly processed in fp16 format to avoid the loss caused by quantization.

[0050] Linear operators (such as matrix multiplication MatMul, linear layer Linear, etc.) first quantize the fp16 format data into int8 format for storage, and then convert it back to fp16 format before actual operation.

[0051] Processing of outliers and non-outliers:

[0052] The outliers are extracted directly from the matrix without quantization, keeping its original fp16 format.

[0053] The non-outlier data is quantized into int8 format according to the dynamic tree quantization method. After necessary calculations, the results are converted back to fp16 format.

[0054] Normalization and quantization: Normalize the data in the feature map to ensure that the data can be represented using a limited number of bits when quantized. Then convert the normalized fp16 format data into int8 format using dynamic tree quantization technology.

[0055] Data conversion: Finally, the conversion of int8 format data to fp16 format is realized, and a pre-established lookup table is used for fast data format conversion while processing outlier data.

[0056] Output: Finally, add the processing results of outliers and non-outliers to get the final quantization model output, which is still in fp16 format.

[0057] In order to solve the outlier problem mentioned above, the method of the present invention adopts a mixed precision solution for processing. Specifically, the quantization precision in the quantization process of the present invention is not single and fixed. The combined use of fp16 and int8 can speed up the operation speed and reduce memory usage.

[0058] i. For outlier data, the method adopted by the present invention is not to quantify, but to extract the outliers from the original matrix by column, find the row where the corresponding weight is located, and extract it as well.

[0059] ii. For non-outlier data, the present invention quantizes according to the following fp16 to int8 quantization method, performs matrix multiplication, and obtains the fp16 operation result.

[0060] iii. Finally, add the results of i and ii to get the final result. Figure 3 shown.

[0061] Figure 3 middle,

[0062] Feature point processing: The input feature points (feature map) are first used as the beginning of this process. In large models, the feature map is the output of the front layer of the neural network and contains a large number of data points.

[0063] Dimension separation: This process extracts outlier data from the matrix sequence. Outliers are data points that are not within the expected data distribution range and may have a significant impact on the final model output. These points are identified and separated from the main data stream.

[0064] Normalization: Non-outlier data is normalized, that is, through the division (Div) operation, the data range is adjusted to facilitate the subsequent quantization process. Normalization helps maintain the accuracy of the data during the quantization process because it ensures that the data values ​​are scaled to a uniform range.

[0065] Dynamic tree quantization: The normalized non-outlier data is then quantized using the dynamic tree quantization method to convert floating point numbers into fixed point numbers. Dynamic tree quantization is a quantization technique that takes data distribution into account. It dynamically adjusts the quantization bit allocation to more accurately adapt to the distribution characteristics of the data.

[0066] Direct outlier processing: Outlier data is processed directly without quantization to maintain its original data format and accuracy.

[0067] Merge results: Finally, the processing results of non-outlier data are merged with the processing results of outliers through addition operations to produce the final output. This step ensures the integrity of the entire data set and the accuracy of the quantification process.

[0068] Compared with traditional methods, the conversion of traditional float data into fixed-point data often requires the use of calibration data sets for reasoning first, and by collecting the data distribution of feature points at each layer, a scaling ratio and an offset are determined, and then the corresponding fixed-point value is obtained by formula (1), where s represents the scaling ratio, b represents the offset, f represents the floating value before quantization, and q represents the result after quantization. Traditional methods often bring disadvantages such as weak generalization ability and difficulty in obtaining data sets. Therefore, the present invention does not adopt a method similar to formula (1), but uses dynamic tree quantization to convert fp16 into fixed-point numbers.

[0069] q=clip(-128,round(f*s)+b,127) (1)

[0070] Where s: represents the scaling factor, which is used to adjust the range of floating-point values ​​before quantization so that they can be mapped into the representation range of fixed-point numbers. The scaling factor is determined based on the distribution characteristics of the data in order to preserve the information of the original data to the greatest extent possible.

[0071] b: represents the bias, which is a fixed value that may need to be added or subtracted during the quantization process to adjust the quantized value to more accurately represent the original value.

[0072] f: represents the floating-point value before quantization, which is the original data that needs to be quantized.

[0073] q: represents the quantized value, that is, the value converted into a fixed-point number after scaling and possible bias adjustment.

[0074] clip(-128,round(f*s)+b,127): This function represents the entire process of quantization. First, the floating point value f is multiplied by the scaling factor s, then the bias b is added, and then the result is rounded (round() function), and finally the value is limited to a specific range through the clip() function, usually the range that can be represented by fixed-point numbers, for example, for 8-bit integers, the range is -128 to 127.

[0075] In the method of the present invention, the specific process of converting the non-outlier point data of the input linear operator from fp16 to int8 is as follows: (1) taking the maximum absolute value of the feature map in a blockwise manner; (2) normalizing the feature map according to the maximum value taken, as shown in formula (2); (3) converting the floating point number into a fixed point number using dynamic tree quantization. The complete process is as follows Figure 4 shown.

[0076]

[0077] The above formula is the normalization of vector data, where num represents a vector data, max() takes the maximum value in the data, and num[i] takes the i-th element in the data.

[0078] Figure 4 middle,

[0079] Input: The process takes feature maps as input, which are the raw data in the large model, usually a multidimensional array calculated by the front layer of the neural network.

[0080] Normalization: Calculate the maximum value of all elements in the feature map, and then perform a division operation (Div) on the feature map. This is the normalization step, the purpose of which is to scale the feature value to a range between 0 and 1. This helps maintain data accuracy in the quantization step.

[0081] Outlier Detection: Outlier detection is also performed, which identifies values ​​that are significantly different from the majority of the data.

[0082] Dynamic tree quantization: The normalized data is fed into the dynamic tree quantization step, where the data is converted from fp16 format to int8 format. Dynamic tree quantization is a quantization method that takes data distribution into account and can produce lower quantization errors when dealing with small and large values.

[0083] Outlier processing: At the same time, the outlier data is directly used in its fp16 format for matrix multiplication without going through the quantization process.

[0084] Result output: The quantized data and the outlier processing results are combined (through addition operation) to produce the final output. This can maintain the integrity of the data and reduce the quantization errors that may be caused by outliers.

[0085] In the method of the present invention, the fixed-point number obtained by dynamic tree quantization has no difference with the actual data distribution of the feature points, and is only related to the number of quantized bits. Therefore, when the number of bits required for quantization is determined, the quantized fixed-point number and the floating-point number obtained by inverse quantization are determined, and they have a one-to-one correspondence. Therefore, the present invention uses a table lookup method to obtain the fp16 value corresponding to int8. The specific process of converting int8 to fp16 can be obtained, as shown in Figure 5 shown.

[0086] Figure 5 middle,

[0087] Input: The process starts with data in int8 format, which means the data has been quantized into 8-bit integers.

[0088] Table lookup: int8 data is converted to fp16 format through table lookup. This table lookup process is pre-calculated to quickly map from quantized integer values ​​back to floating-point values.

[0089] Matrix Multiplication (MatMul): The converted fp16 data participates in matrix multiplication operations, which is a common operation in neural networks to calculate information passed between layers.

[0090] Outlier processing: At the same time, another branch in the flowchart is the processing of outliers. Outliers are data points that are significantly different from the main distribution of the data. Outliers are not quantized and table-looked up, but directly participate in matrix multiplication in fp16 format.

[0091] Result merging (add): The results of the above two matrix multiplications are added and output. This is done to merge the calculation results of the outliers into the main data stream to maintain the integrity of the data.

[0092] In the dynamic tree quantization method, the determination of the quantized value (fixed-point number) depends only on the number of quantization bits selected, and has nothing to do with the distribution of the original feature data. In other words, the quantization process is not affected by how the original data is distributed; once it is determined how many bits are used to quantize the data, each original floating-point value will be mapped to a fixed fixed-point value.

[0093] This quantization process can be viewed as a mapping function, where the input is the original floating-point value, the output is a fixed-point value, and the mapping rules are completely defined by the number of quantization bits. This mapping is one-to-one, meaning that for a given number of quantization bits, each floating-point value has a unique fixed-point value corresponding to it, and vice versa. Therefore, the process is deterministic, and once the number of bits is determined, a lookup table can be pre-calculated and built to quickly convert values ​​during quantization and dequantization.

[0094] In order to verify the correlation error, the present invention adopts two groups of experiments for verification: (1) a certain number of [-1, 1] random numbers (fp32) are selected for testing, and the input random numbers are quantized using fp8 and dynamic tree quantization methods respectively, and the final effect is measured by evaluating the quantization error (the calculation method is shown in Formula 3, q ​​represents the result of quantization and dequantization, and fp represents the original floating point number). The final results are shown in Table 1. The quantization error of the dynamic tree quantization method is 0.01, and the quantization error of the traditional fp8 quantization method is 0.04. It can be seen from the results that the quantization error of the scheme using dynamic tree quantization is significantly better than the result of quantization using fp8; (2) The present invention uses this scheme to quantize the feature points of llama7B and verifies it on the mmlu data set. The results are shown in Table 2. It can be seen from the results that the effect of using the dynamic tree quantization scheme is also better than that of fp8; The results show the performance of models using different methods on the data set, among which the error of dynamic tree quantization 8bit (34.81) is lower than the error of fp8 (35.04), and is also close to the original performance of floating point numbers (Float32) (35.16), which verifies that the dynamic tree quantization method is not only effective in random number testing, but also provides better quantization effects in actual model applications and reduces quantization errors.

[0095] The quantization error err is calculated by formula (3), which measures the absolute error between the quantized value q and the original floating-point value fp before quantization, relative to the size of the original value. This error definition takes into account the absolute value of the error and normalizes it relative to the scale of the original value, making it a more fair reflection of the accuracy of quantization.

[0096] err=abs(q-fp) / abs(fp+1e-7) (3)

[0097] Table 1

[0098] network Quantization Error -1 to 1 random number test (fp8) 0.04 -1 to 1 random number test (dynamic tree quantization 8bit) 0.01

[0099] Table 2

[0100] Quantitative solution Quantization Error Float32 35.16 fp8 35.04 Dynamic tree quantization 8bit 34.81

[0101] Example

[0102] This embodiment mainly quantifies the feature points of each layer of the large model, and the specific implementation steps are as follows:

[0103] (1) All matrix multiplications (including Linear and MatMul) in the model are replaced, that is, mixed precision quantization is performed using the quantization method of the present invention.

[0104] (2) Quantify the model weights.

[0105] (3) Prepare the data set and perform inference on the model.

[0106] Taking llama 7B (hugging face version) as an example, the following layers need to be replaced:

[0107] self_attn.q_proj

[0108] self_attn.k_proj

[0109] self_attn.v_proj

[0110] self_attn.o_proj

[0111] MatMul of query_states and key_states in self_attn

[0112] MatMul of attn_weights and value_states in self_attn

[0113] Specifically, the model layer replacement mentioned in this embodiment is for the llama 7B model, which is a large neural network model commonly used for natural language processing tasks. In the mixed precision quantization method, specific layers in the model need to be replaced to achieve quantization, including the following layers:

[0114] self_attn.q_proj: This is the query projection layer in the self-attention module. It is responsible for projecting the input data into the query space, usually through a linear transformation.

[0115] self_attn.k_proj: Also in the self-attention module, the key projection layer performs a similar operation to project the input data into the key space.

[0116] self_attn.v_proj: The value projection layer is also part of the self-attention module, which projects the input data into the value space.

[0117] self_attn.o_proj: The output projection layer is used to project the data after self-attention calculation to the output space, completing the final step of a self-attention operation.

[0118] MatMul of query_states and key_states in self_attn: This is a step in the self-attention calculation, where the query state and the key state are multiplied by matrix multiplication (MatMul) to calculate the attention score.

[0119] MatMul of attn_weights and value_states in self_attn: After calculating the attention weights (attn_weights), these weights are multiplied with the value states through matrix multiplication to get the weighted sum of the self-attention output.

[0120] In the original model, these layers and operations are calculated with standard floating point precision. During quantization, these layers are replaced with quantized versions, which means that their calculations will involve the use of mixed precision, where part of the data may be calculated with low precision (such as int8) while the other part (outlier data) maintains the original high precision (such as fp16).

[0121] Replacing these layers with mixed precision quantization can reduce the model's memory footprint and improve computational efficiency, especially at inference time. This approach is particularly useful for large models that need to be deployed on resource-constrained devices. Through the dynamic tree quantization method, the quantization error can be effectively reduced without sacrificing model performance, ensuring the generalization ability of the model.

[0122] References

[0123] [1]8-bit Inference with TensorRT s7310-8-bit-inference-with-tensorrt.pdf(gputechconf.com)

[0124] [2]Up or Down? Adaptive Rounding for Post-Training Quantization[2004.10568]Up or Down? Adaptive Rounding for Post-Training Quantization(arxiv.org)

[0125] [3]QDrop:Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization[2203.05740]QDrop:Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization(arxiv.org)

[0126] [4]GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers arxiv.org / pdf / 2210.17323.pdf

[0127] [5]QLoRA: Efficient Finetuning of Quantized LLMs 2305.14314.pdf(arxiv.org)

[0128] [6]LLM.int8():8-bit Matrix Multiplication for Transformers at Scale

[0129] [7]8-BIT OPTIMIZERS VIA BLOCK-WISE QUANTIZATION 2110.02861.pdf(arxiv.org)

[0130] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A large model mixed precision quantization method based on dynamic tree quantization, characterized in that: The quantification method comprises the following steps: Step 1: Divide the input feature points into outlier points and non-outlier points; Step 2: extracting outliers from the matrix sequence of feature points, and extracting the row where the weight corresponding to the outliers is located; or, The non-outlier data of the input linear operator are processed and quantized and stored; the non-outlier data of the input non-linear operator are not quantized; Step 3: Perform matrix multiplication on the outlier data and the weights corresponding to the outliers; or, Dequantize the non-outlier point data stored in the linear operator and perform matrix multiplication. Perform matrix multiplication directly on the non-outlier point data in the nonlinear operator. Step 4: Combine the matrix multiplication results of the non-outlier points and the matrix multiplication results of the outlier points by addition, and output the final result.

2. The mixed precision quantization method according to claim 1, characterized in that: The outlier data is kept in fp16 format and is not quantized. The non-outlier data used for linear operators is converted from fp16 to int8 for storage, and then dequantized and converted back to fp16 during operation. The non-outlier data used for nonlinear operators is kept in fp16 format and is not quantized.

3. The mixed precision quantization method according to claim 1, characterized in that: In step 1, the outlier points refer to data points that are not within the preset data distribution range and will affect the data results of the large model; the non-outlier points refer to data points that are included in the preset data range.

4. The mixed precision quantization method according to claim 1, characterized in that: In step 2, the outlier data is not quantized, and the outliers are extracted from the feature point matrix by dimension; or, For non-outlier data applied to linear operators, the feature map is divided into blocks to obtain the maximum absolute value of each block, and normalized by block. The floating-point data of each block is converted into fixed-point numbers through dynamic tree quantization and stored. No quantization is performed on non-outlier data applied to nonlinear operators.

5. The mixed precision quantization method according to claim 4, characterized in that: The representation method of the dynamic tree quantization is a non-standard floating-point representation method, including four parts, including a sign representation bit, an exponent representation bit, an indicator bit, and a linear quantization bit; the sign representation bit represents the positive and negative signs of the data, 1 represents a negative sign, and 0 represents a positive sign; the exponent representation bit represents the size of the exponent through the number of 0 bits; the indicator bit is the first 1 that appears after the exponent representation bit, and is used to indicate that all values ​​after the indicator bit are used for linear quantization; the linear quantization bit indicates that the corresponding numerical value is represented by a linear quantization method.

6. The mixed precision quantization method according to claim 1, characterized in that: In step three, the dequantization refers to remapping the fixed-point numbers corresponding to the non-outlier data obtained in step two into floating-point numbers through a pre-constructed mapping table, and performing matrix multiplication operations.

7. The mixed precision quantization method according to claim 6, characterized in that: The quantized fixed-point number is related to the selected quantization bit number. After the quantization bit number is determined, each original floating-point number will be mapped to a fixed fixed-point number, and a mapping table is constructed accordingly.

8. Application of the quantization method as described in any one of claims 1 to 7 in reducing the memory consumption of large model training inference.

9. A hardware system for implementing the method according to any one of claims 1 to 7, characterized in that: The hardware system comprises: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Method and device for realizing matrix processing in large model reasoning

    CN121561240A

  • Method and apparatus for implementing matrix processing in large model inference

    CN121561240B