Dynamic quantification method and system for low bit weight and activation value of large language model and application

By using 4-bit normal floating point quantization and 8-bit dynamic tree quantization in large language models, the weights and activation values ​​are quantized without calibration of data, and the problem of calibration of low-bit quantization in the existing technology is solved, and efficient dynamic quantization is achieved without affecting the performance and generalization ability of the model.

CN119993134APending Publication Date: 2025-05-13SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410807170.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing low-bit quantization schemes require calibration of data, resulting in a decline in generalization capability, and dynamic quantization is difficult to ensure the performance of the model after quantization, especially in low-bit cases.

Method used

The weights are quantized by 4-bit normal floating point quantization, and the 8-bit dynamic tree quantization quantizes the activation value without calibration data. Quantification target encoding is generated by calculating quantile points and exponential bit lengths.

Benefits of technology

Dynamic quantization with a weight of 4 bits/8 bits of activation value is realized, without calibration data, and almost no quantization loss is generated, ensuring the generalization ability and inference efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004905116040000045
    Figure BDA0004905116040000045
  • Figure FDA0004905116030000011
    Figure FDA0004905116030000011
  • Figure FDA0004905116030000013
    Figure FDA0004905116030000013
Patent Text Reader

Abstract

The invention discloses a dynamic quantization method for a low-bit weight and an activation value of a large language model, and the method comprises the following steps: 1, selecting a quantization data type according to different distribution characteristics of the weight and the activation value, employing 4-bit normal floating point quantization for the weight, and employing 8-bit dynamic tree quantization for the activation value; and 2, generating a quantization target code for the 4-bit weight and the 8-bit activation value, and obtaining the quantization target code by calculating quantile and / or exponential bit length. In addition, in the quantization process, the quantization precision can be improved by partitioning the data to be quantized and performing quantization and / or inverse quantization processing. The invention further discloses a dynamic quantization system for implementing the dynamic quantization method and application of the dynamic quantization method or system, and the dynamic quantization method or system has wide application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large language models, and relates to a dynamic quantization method, system and application of low-bit weights and activation values ​​of a large language model. Background Art

[0002] Pretrained Large Language Model (LLM) has become the hottest research topic in the field of artificial intelligence in the past two years. It is trained on massive amounts of text data and can perform well on a wide range of downstream language tasks with little or no fine-tuning. Large language models are generally composed of many transformer models, and their powerful capabilities mainly rely on their huge parameter scale (usually billions of parameters). In order to quickly train, fine-tune, infer, and deploy large language models (the target device for deployment may not be able to support high-precision storage and communication of huge parameters), model quantization technology is used to process the parameters of large language models.

[0003] In a narrow sense, model quantization usually refers to representing model parameters stored in floating point format (usually 32-bit full-precision floating point numbers or 16-bit half-precision floating point numbers) with integer data (usually 8-bit integers). Because the bit width of data storage is reduced, the storage space occupied by the model and the overhead during loading and calculation will be reduced, which will improve the efficiency of training and inference. Of course, the most primitive quantization (without any optimization and special processing) will also bring some loss in model performance. In order to reduce this loss, other data types are used instead of integer data as the target code for quantization, such as 8-bit floating point numbers, 4-bit floating point numbers, etc. Among them, Dettmers T et al. [1] A data type of normal floating point number (NormalFloat) is proposed. When it is used as the target encoding for quantization, the stored integer is still an integer, but during dequantization, the dequantized floating point number is obtained by looking up a table (that is, the stored quantized integer is used as the index value, and each index value corresponds to a normalized floating point number). This data type can achieve the optimality in information theory. Dettmers T [2] The proposed Dynamic Tree Quantization is also a similar data type, which performs quantization and dequantization through table lookup. Because it dynamically defines the length of the exponent bit, it has stronger expression capabilities for numbers with large and small absolute values ​​and smaller quantization loss.

[0004] Existing quantization methods for large language models can be divided into static quantization and dynamic quantization according to whether calibration is required. Static quantization methods require a certain amount of calibration data to calibrate the model in order to achieve effective quantization. And due to the diversity and complexity of downstream language tasks, calibration may lead to a decrease in generalization capability. Dynamic quantization is real-time quantization performed during the inference process of the model. Although no offline work is required, it is generally difficult to guarantee the performance of the quantized model, especially in some low-bit quantization (generally refers to a quantization bit width of less than or equal to 4 bits), which is more sensitive to quantization errors and more likely to affect model performance.

[0005] The existing low-bit quantization scheme requires the introduction of calibration data for calibration in order to achieve low-bit quantization with less loss, so dynamic quantization cannot be achieved; and the existence of calibration data may affect the generalization performance of the quantized model;

[0006] Existing dynamic quantization schemes can be divided into two categories. One category can quantize both weights and activation values ​​at the same time, but the quantization bit width is generally 8 bits; the other category only performs low-bit quantization on the weights, while the activation values ​​are retained at 16 bits.

[0007] Therefore, both the existing static quantization and dynamic quantization methods have their shortcomings. Summary of the invention

[0008] In view of the above problems in the prior art, the purpose of the present invention is to provide a method, system and application for dynamic quantization of low-bit weights and activation values ​​of a large language model. The method of the present invention can achieve dynamic quantization of 4-bit weights / 8-bit activation values ​​without any calibration data and almost no quantization loss.

[0009] Specifically,

[0010] The present invention provides a method for dynamic quantization of low-bit weights and activation values ​​of a large language model, comprising the following steps:

[0011] Step 1: Select the quantization data type according to the different distribution characteristics of weights and activation values. The weights are quantized using 4-bit normal floating point quantization, and the activation values ​​are quantized using 8-bit dynamic tree quantization.

[0012] Step 2: Generate a quantized target code for the 4-bit weight and 8-bit activation value, and obtain the quantized target code by calculating the quantile point and / or exponent bit length.

[0013] In step 1, weights and activation values ​​use different quantization data types due to their different distribution characteristics:

[0014] Generally speaking, when using quantile quantization, it is necessary to estimate the quantile of the data based on the empirical cumulative distribution function (ECDF) of the data to be quantized, so as to theoretically ensure that each quantized data can represent an equal amount of data to be quantized, so that each quantization bit width can be fully utilized. According to theoretical experience and experimental research, in pre-trained large language models, weights generally present a zero-mean normal distribution, which allows all weights to share the same quantile by scaling the standard deviation (σ) of different weight distributions.

[0015] In step 1, the weights present a zero-mean normal distribution, with a mean of 0, bilateral symmetry, and a normal distribution;

[0016] The weight data adopts the quantile quantization method, the quantization bit width is 4 bits, and the quantization target code has 16 values;

[0017] and / or,

[0018] The activation values ​​are normally distributed, a dynamic tree quantization method is adopted, the quantization bit width is 8 bits, and the quantization target code has 256 values.

[0019] In step 1, the weight data is scaled according to the standard deviation (σ) by equivalently implementing the process of quantization to nf4;

[0020] In the present invention, a normal floating point number (NormalFloat, NF) is used as a target code to quantize the weights in a large language model, the quantization bit width used is 4 bits, and the quantization target code has 16 values.

[0021] In NF4 quantization, data standardization is part of the quantization step. This means that the original weight data has been subtracted from the mean and divided by the standard deviation during the quantization process, automatically converted to a standard normal distribution.

[0022] The quantile calculation and distribution transformation in NF4 quantization make each quantized data point represent an equal amount of original data, thus achieving the effect of standardization and scaling.

[0023] NF4 quantization obtains the quantiles under the standard normal distribution by looking up a table, and then applies these quantiles directly to the original data. The target encoding of NF4 quantization directly includes these uniform quantiles, thus simplifying the calculation steps.

[0024] For the activation values ​​in the large language model, although it is found through experiments that some of them also present a normal distribution, considering the larger dynamic range brought by matrix multiplication, it is necessary to express data with larger absolute values ​​and smaller absolute values ​​more accurately, so more quantization bit widths should be allocated to these data; therefore, in the quantization method of the activation value, data of different ranges are represented by dynamically defining the length of the exponent bit.

[0025] In a specific embodiment, the dynamic range is set to be from 10 -6 to 10 6 .

[0026] In the present invention, 8-bit dynamic tree quantization is used as the target encoding to quantize the activation value of the large language model.

[0027] In step 2, the generation of the quantization target code includes the following steps:

[0028] Based on the different data distribution characteristics of weights and activation values, it is proposed to use different quantization data types as target codes. The following will explain how to generate the corresponding target codes:

[0029] Step 2.1, generation of 4-bit normal floating point number NF4:

[0030] Since the quantization bit width is 4 bits, the quantization target code has a total of 16 values, which will be generated as follows: [3] :

[0031] Step 2.1.1, Setup

[0032] Step 2.1.2, calculate the probability value corresponding to the negative semi-axis target encoding: set the minimum probability value p 1 =δ, maximum probability value p 8 = 0.5, take 6 probability values ​​p at equal intervals between the minimum and maximum probability values 2 , p 3 ,…,p 7 , thus obtaining the probability value p corresponding to the negative semi-axis quantization target encoding 1 , p 2 ,…,p 7 ;

[0033] Step 2.1.3, get the negative half axis target code: use the inverse function of the standard normal distribution cumulative distribution function (CDF) Φ, that is, the percentile function (PPF) Φ -1 , get the quantiles corresponding to the 7 probability values ​​in step 2.1.2

[0034] Step 2.1.4, calculate the probability value corresponding to the positive half axis target code: set the minimum probability value p 8 =0.5, maximum probability value p 16 =1-δ, take 7 probability values ​​p at equal intervals between the minimum and maximum probability values 9 , p1 0 ,…,p 15 , thus obtaining the positive semi-axis quantization target code and the probability value p corresponding to the zero point 8 ,p 9 ,…,p 16 ;

[0035] Step 2.1.5: Get the positive half axis target code: Also use the percentile function Φ -1 To calculate, we can get the quantiles corresponding to the 9 probability values ​​in step d

[0036] Step 2.1.6, Normalization: By dividing Will get Normalized to between -1 and 1, that is

[0037] The above gives us the 4-bit target encoding in the form of a normal floating point number.

[0038] Step 2.2, Generation of 8-bit dynamic tree quantization:

[0039] Because the quantization bit width is 8 bits, the quantization target code has a total of 256 values, which will be generated as follows:

[0040] Step 2.2.1, let the range of the number of exponent bits be e∈0,6, where e is an integer;

[0041] Step 2.2.2: For each exponent digit e, calculate the number of interval divisions f e =2 e and the exponential weight w e =10 e -6 ;

[0042] Step 2.2.3: Divide the number of parts into f according to the interval e , divide the interval [0.1,1] equally into f e Take the midpoint value of each interval As the mantissa part;

[0043] Step 2.2.4: Get the target code for quantization under the current condition of e

[0044] Step 2.2.5: Repeat steps 2.2.2-2.2.4 to find the union of all quantized target codes corresponding to e Then, two values ​​0 and 1 are added to the union and the entire set is sorted from small to large to obtain the final quantized target coding set.

[0045] In the present invention, the quantization accuracy can be further improved and the quantization loss can be reduced by block processing, quantization and dequantization during dynamic quantization, which includes the following steps:

[0046] Step i, block division: In order to improve quantization accuracy and reduce quantization loss, the data to be quantized can be divided into blocks, and then the same quantization parameters are used for quantization in each data block. For example, for the weight tensor W to be quantized (the shape of W is 4096×4096), it is expanded into a one-dimensional vector by row and divided by block size B=64, resulting in 262144 data blocks W b ,b=1,…,262144. For each data block W b , find its maximum absolute value N b =max(|W b |) as a normalization constant, and then all elements of the data block are normalized using this normalization constant, that is,

[0047] Step ii, quantization: After the weights are processed by blocks, the quantization target encoding of the weights is selected according to the nearest neighbor principle. The value q that is closest to the i as the quantized value (a 4-bit integer); and / or,

[0048] After the activation value is processed by block, the quantization target encoding of the activation value is selected according to the nearest neighbor principle The value q that is closest to the i The index i of is taken as the quantized value (an 8-bit integer).

[0049] and / or,

[0050] Step iii, dequantization: When dequantizing, first get the corresponding value q according to the index i i Then, according to the normalization constant N of each saved data block b , get the dequantized floating point number

[0051] The present invention also provides a dynamic quantization system for low-bit weights and activation values ​​of a large language model, the dynamic quantization system comprising: an input module, a quantization selection module, a weight quantization module, an activation value quantization module, and a block processing module;

[0052] The input module is used to receive low-bit weight and activation value data of a large language model;

[0053] The quantization selection module is used to select the quantization data type according to different distribution characteristics of weights and activation values;

[0054] The weight quantization module is used to perform normal floating point quantization on the 4-bit weight data to obtain a quantized target code;

[0055] The activation value quantization module is used to perform dynamic tree quantization on the 8-bit activation value data to obtain a quantized target code;

[0056] The block processing module is used to process the quantized data in blocks.

[0057] The present invention also provides the application of the above-mentioned dynamic quantization method or the above-mentioned dynamic quantization system in large model applications of terminal devices, large language model reasoning, etc.

[0058] The present invention also provides a hardware system for implementing the above-mentioned dynamic quantization method, and the hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned dynamic quantization method is implemented.

[0059] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned dynamic quantization method is implemented.

[0060] The beneficial effects of the present invention include:

[0061] Compared with the existing quantification method that requires calibration, the method proposed by the present invention does not require any calibration data for calibration and can achieve online real-time dynamic quantification;

[0062] Because no calibration data is required, the generalization ability of the quantized model is not affected;

[0063] Compared with the existing dynamic quantization method, the method proposed in the present invention can quantize weights and activation values ​​at the same time, and realize low-bit quantization (4 bits) of weights; because the different distribution characteristics of weights and activation values ​​are fully considered, it has almost no effect on the inference time, and can ensure that the performance of the quantized model is almost not lost. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0065] Figure 1 This is a schematic diagram of a 4-bit normal floating point number.

[0066] Figure 2 A comparison diagram of the target encoding of 8-bit dynamic tree quantization and 8-bit normal floating point numbers. DETAILED DESCRIPTION

[0067] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.

[0068] The present invention provides a dynamic quantization method for low-bit weights and activation values ​​of a large language model, and the dynamic quantization method comprises the following steps: step 1, selecting the quantization data type according to the different distribution characteristics of the weights and activation values, the weights are quantized using 4-bit normal floating point, and the activation values ​​are quantized using 8-bit dynamic tree; step 2, generating a quantization target code for the 4-bit weights and 8-bit activation values, and obtaining the quantization target code by calculating the quantile point and / or exponent bit length. In the quantization process, the quantization accuracy can also be improved by dividing the data to be quantized into blocks, and performing quantization and / or dequantization processing. The present invention also provides a dynamic quantization system for implementing the above-mentioned dynamic quantization method, and the application of the dynamic quantization method or system, which has wide application value.

[0069] The method of the present invention adopts different quantization data types (weights are quantized by 4-bit normal floating point and activation values ​​are quantized by 8-bit dynamic tree) according to the different distribution characteristics of low-bit weights and activation values ​​of a large language model, and quantizes the data to be quantized in blocks. It can realize dynamic quantization without any calibration, and has almost no effect on the inference time, and almost no loss to the performance of the model.

[0070] Specifically, the present invention provides a dynamic quantization method for low bit weights and activation values ​​of a large language model, characterized in that the dynamic quantization method comprises the following steps:

[0071] Step 1: Select the quantization data type according to the different distribution characteristics of weights and activation values;

[0072] Step 2: Generate quantized target code for 4-bit weight and 8-bit activation value.

[0073] In the present invention, the quantization accuracy can be improved by dividing the data to be quantized into blocks and performing quantization and / or inverse quantization processing.

[0074] In step 1, the weights present a zero-mean normal distribution, with a mean of 0, bilateral symmetry, and a normal distribution;

[0075] The weight data adopts the quantile quantization method, the quantization bit width is 4 bits, and the quantization target code has 16 values;

[0076] and / or,

[0077] The activation values ​​are normally distributed, a dynamic tree quantization method is adopted, the quantization bit width is 8 bits, and the quantization target code has 256 values.

[0078] In step 1, the weight data is scaled according to the standard deviation (σ) by equivalently implementing the process of quantization to nf4;

[0079] and / or,

[0080] The activation value quantization method dynamically defines the length of the exponent bit to represent data of different ranges; the setting size of the dynamic range is 10 -6 to 10 6 .

[0081] In step 2, the quantized data encoding generation of the 4-bit weight includes the following steps:

[0082] Step 2.1.1, set the minimum probability value

[0083] Step 2.1.2, calculate the probability value corresponding to the negative semi-axis target encoding: set the minimum probability value p 1 =δ, maximum probability value p 8 = 0.5, take 6 probability values ​​p at equal intervals between the minimum and maximum probability values 2 , p 3 ,…,p 7 , probability value p 1 , p 2 ,…,p 7 The probability value corresponding to the negative semi-axis quantization target encoding;

[0084] Step 2.1.3: Use the percentile function Φ -1 Get the quantile of the probability value corresponding to the negative semi-axis target encoding

[0085] Step 2.1.4, calculate the probability value corresponding to the positive half axis target code: set the minimum probability value p 8=0.5, maximum probability value p 16 =1-δ, take 7 probability values ​​p at equal intervals between the minimum and maximum probability values 9 , p 10 ,…,p 15 , thus obtaining the positive semi-axis quantization target code and the probability value p corresponding to the zero point 8 , p 9 ,…,p 16 ;

[0086] Step 2.1.5: Use the percentile function Φ -1 Get the quantile of the probability value corresponding to the positive semi-axis target encoding

[0087] Step 2.1.6: By dividing the quantile by Normalized to [-1, 1], that is The target code with 4-bit weight is obtained

[0088] like Figure 1 In the figure, the solid line is the cumulative distribution function of the standard normal distribution, the dotted line is the probability density function of the standard normal distribution, the horizontal axis represents the value of the random variable, and the vertical axis is the probability value of the cumulative distribution function. Figure 1 The values ​​of the positive half axis of a 4-bit normal floating-point number are given in the figure. It can be seen that the vertical axis is sampled at equal intervals (that is, the probability value of each interval is the same), which ensures that the number of data before quantization represented by each quantized target code is the same. The q in the figure is the value of the random variable at the quantile corresponding to each sampling probability value.

[0089] In step 2, the quantized data encoding generation of the 8-bit activation value includes the following steps:

[0090] Step 2.2.1, let the range of the number of exponent bits be e∈[0,6], where e is an integer;

[0091] Step 2.2.2: For each exponent digit e, calculate the number of interval divisions f e =2 e and the exponential weight w e =10 e -6 ;

[0092] Step 2.2.3: Divide the interval [0.1, 1] into fe equal parts according to the number of parts fe, and take the midpoint value of each interval. As the mantissa part;

[0093] Step 2.2.4: Get the target code for quantization under the current condition of e

[0094] Step 2.2.5: Repeat steps 2.2.2-2.2.4 to find the union of all quantized target codes corresponding to e Then add 0 and 1 to the union and sort the entire set from small to large to obtain the final quantized target coding set.

[0095] Figure 2 In the figure, the solid line is an 8-bit normal floating point number, the dotted line is an 8-bit dynamic tree quantization, the horizontal axis is the index value (0 to 255) corresponding to the target code after quantization, and the vertical axis is the normalized target value (-1 to 1) after quantization. It can be seen that compared with normal floating point numbers, the target code of dynamic tree quantization has more bits to represent numbers with large and small absolute values, that is, in the target value ranges of -1 to -0.75, -0.25 to 0.25, and 0.75 to 1 in the figure, dynamic tree quantization has more values ​​(corresponding to more index values) than normal floating point numbers. This feature makes dynamic tree quantization more suitable for quantization of activation values.

[0096] In the present invention, the quantization is processed by quantization block processing and quantization and inverse quantization, including the following steps:

[0097] Step i: Block division: divide the data to be quantized into blocks according to the block size, and normalize each data block;

[0098] Step ii, quantization: using the quantization target code calculated in step 2, mapping the normalized data to the index of the closest quantization value; and / or,

[0099] Step iii, dequantization: restore the quantized value to an approximate original floating-point value according to the quantization index and the saved normalization constant.

[0100] The method of the present invention can be applied to many fields, including but not limited to the following:

[0101] 1. Mobile devices and embedded systems

[0102] Mobile devices (e.g., smartphones, tablets) and embedded systems (e.g., IoT devices, smart home devices) usually have limited computing power and storage space.

[0103] Speech Recognition: Implement efficient speech recognition applications on smartphones, reduce model size and computing requirements through dynamic quantization, and improve real-time response speed;

[0104] Image processing: Use dynamic quantization to run image recognition and processing algorithms in embedded systems, reducing power consumption and memory usage.

[0105] 2. Cloud computing and edge computing

[0106] Cloud computing and edge computing environments need to process large amounts of data and require efficient computing resource management.

[0107] Real-time video analysis: Deploy video analysis models on edge computing devices, reduce model size through dynamic quantization, and improve processing speed and efficiency.

[0108] Large-scale language model reasoning: Run large-scale language models on cloud computing platforms, reduce computing and storage requirements through dynamic quantization, and provide efficient reasoning services.

[0109] 3. Autonomous driving and intelligent transportation systems

[0110] Autonomous driving and intelligent transportation systems require processing large amounts of sensor data and making decisions in real time.

[0111] Real-time path planning: Use dynamic quantization methods to run path planning algorithms in autonomous vehicles to improve computational efficiency and ensure real-time response.

[0112] Traffic monitoring and management: Deploy dynamic quantitative models in intelligent transportation systems to predict and manage traffic flow, optimize traffic signal control and vehicle scheduling.

[0113] 4. Large-scale machine learning model training

[0114] Training large-scale machine learning models usually requires a lot of computing resources and storage space.

[0115] Model training acceleration: In a distributed computing environment, intermediate calculation results are quantized through dynamic quantization methods to reduce data transmission and storage requirements and accelerate the model training process.

[0116] Model compression and transmission: In a multi-node distributed training system, dynamic quantization methods are used to compress model parameters, improve data transmission efficiency, and reduce training time.

[0117] 5. Online Recommendation System

[0118] Online recommendation systems need to process large amounts of user data in real time and provide personalized recommendations.

[0119] Real-time recommendation algorithm: Use dynamic quantization methods to run the recommendation algorithm in the online recommendation system to improve computing efficiency and response speed, and provide efficient personalized recommendation services.

[0120] User behavior analysis: During the user behavior analysis process, dynamic quantification is used to reduce model computing requirements, analyze user behavior in real time, and adjust recommendation strategies.

[0121] Example

[0122] In this embodiment, given a pre-trained large language model M to be quantized, it has several weights W 1 ,W 2 ,…,W N , and some activation value A 1 ,A 2 ,…,A M (In addition to the output of the activation function, the activation value can also refer to the output of any linear layer in the model or the output of any function, such as softmax, etc.). Dynamic quantization during inference can be performed as follows:

[0123] Step 1. For any weight W i , divide the method into blocks, normalize it, and then quantize it, and then dequantize it according to the dequantization technology to obtain

[0124] Step 2. Weight W i Input I i is a floating point number obtained through dequantization, Perform matrix multiplication to get output A i ; Divide it into blocks and normalize it, then select 8-bit dynamic tree quantization according to the nearest neighbor principle The value q that is closest to the i The index i is taken as the quantized value (an 8-bit integer); and then the dequantization is performed according to the method described in the dequantization technology of the present invention to obtain Can be passed backward as input to subsequent modules;

[0125] Step 3. For any other activation value A j (The output of the nonlinear layer weight calculation), also follow the steps in step 2 for A i The method of processing is to obtain And passed backward as input to subsequent modules.

[0126] Experiments were conducted using the LLaMA7B model, which was quantized according to the method described above and evaluated using the mmlu, wiki, and c4 datasets. The results are as follows:

[0127] Data Types mmlu Wiki-ppl C4-ppl FP16 (unquantized) 0.35 5.37 7.08 NF4+DTQ8 0.3498 5.49 7.28

[0128] It can be seen that the quantized model (NF4+DTQ8) has almost no loss compared to the baseline model using half-precision floating point numbers (FP16).

[0129] References

[0130] [1]DettmersT,PagnoniA,HoltzmanA,etal.Qlora:Efficientfinetuningofquantizedllms[J].arXivpreprintarXiv:2305.14314,2023.

[0131] [2]DettmersT.8-bitapproximationsforparallelismindeeplearning[J].arXivpreprintarXiv:1511.04561,2015.

[0132] [3]YoshidaD.NF4Isn'tInformationTheoreticallyOptimal(andthat'sGood)[J].arXivpreprintarXiv:2306.06965,2023.

[0133] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A method for dynamic quantization of low-bit weights and activation values ​​of a large language model, characterized in that: The dynamic quantization method comprises the following steps: Step 1: Select the quantization data type according to the different distribution characteristics of weights and activation values. The weights are quantized using 4-bit normal floating point quantization, and the activation values ​​are quantized using 8-bit dynamic tree quantization. Step 2: Generate a quantized target code for the 4-bit weight and 8-bit activation value, and obtain the quantized target code by calculating the quantile point and / or exponent bit length.

2. The dynamic quantization method according to claim 1, characterized in that: In step 1, the weights present a zero-mean normal distribution, with a mean of 0, bilateral symmetry, and a normal distribution; The weight data adopts the quantile quantization method, the quantization bit width is 4 bits, and the quantization target code has 16 values; and / or, The activation values ​​are normally distributed, a dynamic tree quantization method is adopted, the quantization bit width is 8 bits, and the quantization target code has 256 values.

3. The dynamic quantization method according to claim 2, characterized in that: The quantization method of the weight data is to equivalently implement the scaling according to the standard deviation in the process of quantization to nf4; and / or, The activation value quantization method dynamically defines the length of the exponent bit to represent data of different ranges; the setting size of the dynamic range is 10 -6 to 10 6 .

4. The dynamic quantization method according to claim 1, characterized in that: In step 2, the quantized data encoding generation of the 4-bit weight includes the following steps: Step 2.1.1, set the minimum probability value Step 2.1.2, calculate the probability value corresponding to the negative semi-axis target coding: set the minimum probability value p1 = δ, the maximum probability value p8 = 0.5, and take 6 probability values ​​p2, p3, ..., p7 at equal intervals between the minimum and maximum probability values. The probability values ​​p1, p2, ..., p7 are the probability values ​​corresponding to the negative semi-axis quantization target coding; Step 2.1.3: Use the percentile function Φ -1 Get the quantile of the probability value corresponding to the negative semi-axis target encoding Step 2.1.4, calculate the probability value corresponding to the positive half axis target code: set the minimum probability value p8 = 0.5, the maximum probability value p 16 =1-δ, take 7 probability values ​​p9,p at equal intervals between the minimum and maximum probability values 10 ,…,p 15 , thus obtaining the positive semi-axis quantization target code and the probability values ​​corresponding to the zero point p8, p9, ..., p 16 ; Step 2.1.5: Use percentile function Φ -1 Get the quantile of the probability value corresponding to the positive semi-axis target encoding Step 2.1.6: By dividing the quantile by Normalized to [-1,1], that is The target code with 4-bit weight is obtained 5. The dynamic quantization method according to claim 1, characterized in that: In step 2, the quantized data encoding generation of the 8-bit activation value includes the following steps: Step 2.2.1, let the range of the number of exponent bits be e∈[0,6], where e is an integer; Step 2.2.2: For each exponent digit e, calculate the number of interval divisions f e =2 e and the exponential weight w e =10 e-6 ; Step 2.2.3: Divide the number of parts into f according to the interval e , divide the interval [0.1,1] equally into f e Take the midpoint value of each interval As the mantissa part; Step 2.2.4: Get the target code for quantization under the current condition of e Step 2.2.5: Repeat steps 2.2.2-2.2.4 to find the union of all quantized target codes corresponding to e Then add 0 and 1 to the union and sort the entire set from small to large to obtain the final quantized target coding set.

6. The dynamic quantization method according to claim 1, characterized in that: Improving the quantization accuracy by processing the quantization includes the following steps: Step i: Block division: divide the data to be quantized into blocks according to the block size, and normalize each data block; Step ii, quantization: using the quantization target code calculated in step 2, the normalized data is mapped to the index of the closest quantization value; and / or, Step iii, dequantization: restore the quantized value to an approximate original floating-point value according to the quantization index and the saved normalization constant.

7. A dynamic quantization system for low-bit weights and activation values ​​of large language models, characterized in that The dynamic quantization system includes: an input module, a quantization selection module, a weight quantization module, an activation value quantization module, and a block processing module; The input module is used to receive low-bit weight and activation value data of a large language model; The quantization selection module is used to select the quantization data type according to different distribution characteristics of weights and activation values; The weight quantization module is used to perform normal floating point quantization on the 4-bit weight data to obtain a quantized target code; The activation value quantization module is used to perform dynamic tree quantization on the 8-bit activation value data to obtain a quantized target code; The block processing module is used to process the quantized data in blocks.

8. Application of the dynamic quantization method as described in any one of claims 1 to 6, or the dynamic quantization system as described in claim 7 in large model applications of terminal devices and large language model reasoning.

9. A hardware system for implementing the dynamic quantization method according to any one of claims 1 to 6, characterized in that: The hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the dynamic quantization method according to any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the dynamic quantization method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Large model mixing precision quantification method and system based on LSTM dynamic prediction

    CN121189379A