Quantization method and reasoning method of large language model and electronic equipment
By dividing the linear layer channels of the large language model into normal and outlier channels, and using a hybrid precision quantization solution, the problem of large memory usage of large language model is solved, and efficient reduction of video memory and maintenance of inference computing power is achieved.
Patent Information
- Application Number
- CN202510118310.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-22
AI Technical Summary
Due to the huge amount of parameters, large language models consume a large amount of computing resources and memory. Although the existing quantitative solution reduces computing resources, the video memory is still large, making it difficult to further optimize.
The channels of the linear layer of the large language model in the hidden layer dimension are divided into normal channels and outlier channels, and the normal channels are quantized INT8 and INT4 quantized, combined with the second-order information of the Hessian matrix for error compensation, and the outlier channels are not quantized or INT8 quantized to form a hybrid precision quantization scheme.
On the basis of ensuring quantitative accuracy, the video memory usage of large language models is significantly reduced, efficient compression of the model is achieved, and the inference computing power is not affected.
Smart Images

Figure CN120354895A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of large language models, and particularly relates to a quantization method, an inference method, and an electronic device for large language models. Background Art
[0002] Due to the excellent processing capabilities demonstrated by large language models (LLMs) during task processing, they have become the most concerned language processing models currently. At the same time, large language models also face various challenges during implementation due to their huge number of parameters, such as high computational resource consumption and large memory occupancy.
[0003] In response to this, model quantization schemes such as W8A8 quantization have been proposed in related technologies to reduce computational resource consumption and lower memory occupancy. However, taking the W8A8 quantization scheme provided in related technologies as an example, although the model quantization accuracy is acceptable, the video memory occupied by the model is still relatively large. Therefore, how to reduce the video memory occupied by large language models is still a technical problem urgently to be solved in this field. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a quantization method, an inference method, and an electronic device for large language models, which can reduce the video memory occupied by large language models.
[0005] In a first aspect, a quantization method for large language models is provided, including: for each linear layer to be quantized in the large language model, dividing the channels in the hidden layer dimension of the linear layer into normal channels and outlier channels; performing INT8 quantization on the first activation matrix corresponding to the normal channels in the token dimension of the tokens to obtain a second activation matrix, and performing INT4 quantization on the first weight matrix corresponding to the normal channels according to the output channels to obtain a second weight matrix; determining the output result of the linear layer based on the second activation matrix, the second weight matrix, the third activation matrix corresponding to the outlier channels, and the third weight matrix corresponding to the outlier channels.
[0006] In a second aspect, an inference method for large language models is provided. The large language model is quantized using the method described in the first aspect. The inference method includes: obtaining the quantization configuration information corresponding to the large language model; performing model loading and inference according to the quantization configuration information; where the quantization configuration information includes at least one of the following: the first information, used to indicate the bit number of weight quantization corresponding to the linear layer in the large language model; the second information, used to indicate whether the large language model is a floating-point model.
[0007] In a third aspect, a quantization device for a large language model is provided, including: a first processing module configured to divide channels of each linear layer to be quantized in a large language model into normal channels and outlier channels in a hidden layer dimension; a second processing module configured to perform INT8 quantization on a first activation matrix corresponding to the normal channels in a token dimension to obtain a second activation matrix, and perform INT4 quantization on a first weight matrix corresponding to the normal channels by output channels to obtain a second weight matrix; and a third processing module configured to determine an output result of the linear layer according to the second activation matrix, the second weight matrix, a third activation matrix corresponding to the outlier channels, and a third weight matrix corresponding to the outlier channels.
[0008] In a fourth aspect, an inference device for a large language model is provided. The large language model is quantized by using the method described in the first aspect. The inference device includes: a fourth processing module configured to obtain quantization configuration information corresponding to the large language model; and a fifth processing module configured to perform model loading and inference according to the quantization configuration information. The quantization configuration information includes at least one of the following: a first piece of information for indicating a weight quantization bit number corresponding to a linear layer in the large language model; and a second piece of information for indicating whether the large language model is a floating-point model.
[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor, and computer executable instructions stored on the memory and executable on the processor. When the computer executable instructions are executed by the processor, the steps of the method described in the first aspect or the second aspect are implemented. In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium for storing computer executable instructions. When the computer executable instructions are executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented. In an embodiment of the present application, by quantizing the activation and weight corresponding to the normal channels in a large language model by using a mixed-precision quantization scheme based on W4A8, the video memory occupied by the large language model can be effectively reduced on the basis of ensuring quantization accuracy. Description of the Drawings In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0010] Figure 1It is one of the schematic flowcharts of the quantization method of the large language model provided by an exemplary embodiment of the present application.
[0011] Figure 2 It is a schematic diagram of the quantization process provided by an exemplary embodiment of the present application.
[0012] Figure 3a It is the second schematic flowchart of the quantization method of the large language model provided by an exemplary embodiment of the present application.
[0013] Figure 3b It is the third schematic flowchart of the quantization method of the large language model provided by an exemplary embodiment of the present application.
[0014] Figure 4 It is the fourth schematic flowchart of the quantization method of the large language model provided by an exemplary embodiment of the present application.
[0015] Figure 5 It is the fifth schematic flowchart of the quantization method of the large language model provided by an exemplary embodiment of the present application.
[0016] Figure 6 It is the first schematic flowchart of the inference method of the large language model provided by an exemplary embodiment of the present application.
[0017] Figure 7a It is the first schematic flowchart of the loading method of the quantized large language model provided by an exemplary embodiment of the present application.
[0018] Figure 7b It is the second schematic flowchart of the loading method of the quantized large language model provided by an exemplary embodiment of the present application.
[0019] Figure 8 It is a schematic diagram of the inference process of each linear layer in the large language model provided by an exemplary embodiment of the present application.
[0020] Figure 9 It is a schematic diagram of the INT type conversion provided by an exemplary embodiment of the present application.
[0021] Figure 10 It is the second schematic flowchart of the inference method of the large language model provided by an exemplary embodiment of the present application.
[0022] Figure 11 It is a schematic diagram of the structure of the quantization device of the large language model provided by an exemplary embodiment of the present application.
[0023] Figure 12 It is a schematic diagram of the structure of the inference device of the large language model provided by an exemplary embodiment of the present application.
[0024] Figure 13It is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] The quantization system and the inference system mentioned below can be deployed on the same electronic device or on different electronic devices, which is not limited here.
[0027] In addition, "WxAy (Weight x-bit Activation y-bit)" mentioned below means that the weight is quantized using INTx and the activation is quantized using INT y.
[0028] Below, in conjunction with the accompanying drawings, the technical solutions provided by the embodiments of the present application will be described in detail through some embodiments and their application scenarios.
[0029] Figure 1 Fig. 18 shows a schematic flowchart of a quantization method 100 for a large language model provided by an embodiment of the present application. The method 100 can be executed by a quantization system installed in an electronic device. Among them, the electronic device can be a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. As Figure 1 shown, the method 100 may include the following steps.
[0030] Step S110, for each linear layer to be quantized in the large language model, divide the channels of the linear layer in the hidden dimensions into normal channels and outlier channels.
[0031] The large language model mentioned in this embodiment can be understood as being obtained by stacking decoders, and most of the calculations in the large language model are concentrated in the linear layers in the decoder. Therefore, the quantization of the large language model in this embodiment can be understood as quantizing the linear layers in the decoder.
[0032] In addition, for the large language model, considering that in some activated channels, regardless of how the tokens change, there will continuously appear outliers with very large amplitude values, which can also be called anomalies, and the quantization processing method for outliers determines the accuracy of the large language model after quantization. Therefore, in this embodiment, when quantizing the large language model, the channels of the linear layer in the hidden layer dimension can be divided into normal channels and outlier channels according to the outliers. Among them, the normal channels can be understood as the channels corresponding to normal points or ordinary points, and the outlier channels can be understood as the channels corresponding to outliers or anomalies. Finally, the weights are quantized according to the differences between the normal channels and the outlier channels, achieving the purpose of quantizing the large language model and ensuring the accuracy of the large language model after quantization.
[0033] In this embodiment, the normal points can be, but are not limited to, activation values whose amplitude values (or absolute values) meet specific conditions. Correspondingly, the outlier points can be, but are not limited to, activation values whose amplitude values do not meet specific conditions. Among them, the specific conditions can include, but are not limited to, the amplitude value being less than a predetermined threshold, etc.
[0034] In some embodiments, considering that for each linear layer, the channels where the outliers of the weights are located are determined by activation and correspond to the positions of the activated channels. Therefore, in this embodiment, when dividing the channels of the linear layer in the hidden layer dimension into normal channels and outlier channels in step S110, the channels of the linear layer in the hidden layer dimension can be divided into normal channels and outlier channels according to activation statistics (such as the maximum value of the amplitude values or absolute values of the activated statistics).
[0035] That is, the process of dividing the channels of the linear layer in the hidden layer dimension into normal channels and outlier channels in step S110 can include: dividing the channels of the linear layer in the hidden layer dimension into the normal channels and the outlier channels according to the importance of each channel of the linear layer in the hidden layer dimension. Among them, the importance of the channel is determined according to the activation statistics corresponding to the channel. For example, the maximum value of the absolute values of the activations of the large language model can be statistically calculated for each channel through a calibration dataset (such as PILE). The greater the absolute value of the activation value, the higher the importance of the corresponding channel.
[0036] Exemplarily, please refer to Figure 2 , assuming that for the linear layer to be quantized, the calculation of the linear layer can be expressed as , its activation , weight , where M is the number of word segments on the input channel, K is the number of channels in the hidden layer dimension, and N is the number of output channels. Then, when dividing the channels of the linear layer in the hidden layer dimension, the maximum value among the absolute values of each activation value X corresponding to each channel of the linear layer in the hidden layer dimension can be first counted; then, the channels can be sorted in ascending order according to the maximum value corresponding to each channel to obtain a sorting result, as shown in Figure 2 in (2) below; finally, a predetermined number of channels ranked at the end are selected as the outlier channels according to the sorting result, and the other channels except the outlier channels are used as the normal channels. Among them, by rearranging the activations by channel, the channels where the outliers are located can be arranged at the end of the matrix, as shown in Figure 2 in (3) below. The part without background color represents the normal channels, that is, the weights corresponding to the normal points and the channels corresponding to the activations. The part shaded with left diagonal lines represents the abnormal channels, that is, the weights corresponding to the outliers and the activations. Figure 2 In (3) below, the gray color blocks illustrate the quantization granularity of the activations corresponding to the normal points and the quantization granularity of the weights corresponding to the normal points. In this embodiment, when performing channel rearrangement, the reorder_index can be recorded in the linear layer for dynamic rearrangement and quantization of activations during subsequent model inference.
[0037] Among them, the predetermined number can be, but is not limited to, 256, etc.
[0038] Step S120, perform INT8 quantization on the first activation matrix corresponding to the normal channels in the token dimension (per-token) to obtain a second activation matrix, and perform INT4 quantization on the first weight matrix corresponding to the normal channels per output channel to obtain a second weight matrix.
[0039] In some embodiments, the quantization methods for the third activation matrix corresponding to the outlier channels and the third weight matrix corresponding to the outlier channels may include, but are not limited to: not quantizing or using INT8 quantization for the third activation matrix and the third weight matrix to ensure the accuracy of the large language model after quantization.
[0040] Among them, if the non-quantization method is adopted, then the third activation matrix and the third weight matrix are floating-point matrices, such as floating-point data matrices in FP16 / BF16 format.
[0041] If the INT8 quantization method is adopted, then the third activation matrix can be understood as the matrix after INT8 quantization in the token dimension, and the third weight matrix is the matrix after INT8 quantization in the output channel dimension.
[0042] Step S130, determine the output result of the linear layer according to the second activation matrix, the second weight matrix, the third activation matrix corresponding to the outlier channels, and the third weight matrix corresponding to the outlier channels.
[0043] In this embodiment, by quantizing the activations and weights corresponding to the normal channels in the large language model using the W4A8-based mixed precision quantization scheme, the video memory occupied by the large language model can be effectively reduced while ensuring the quantization accuracy. In some embodiments, the process of performing INT4 quantization on the first weight matrix corresponding to the normal channels in step S120 according to the output channels may include, but is not limited to: performing INT4 GPTQ (Generative Pretrained Transformer Quantization) quantization on the first weight matrix corresponding to the normal channels using the sorted Hessian matrix according to the output channels to obtain the second weight matrix.
[0044] In this embodiment, when performing INT4 GPTQ quantization on the first weight matrix corresponding to the normal channels using the rearranged Hessian matrix according to the output channels, the columns where the first weight matrix W is located are iterated. For each column, all elements are quantized simultaneously. After each column is quantized, the other unquantized columns are adjusted. Since the same quantization order is used for each row in the first weight matrix W, by performing block quantization on the weights column by column, the information of the inverse Hessian matrix stored in the Cholesky decomposition can be used to update the unquantized weights in the block to compensate for the loss caused by quantization. That is, in this embodiment, when quantizing the first weight matrix W based on GPTQ quantization, the second-order information of the Hessian matrix can be used to compensate for the introduced quantization error, thereby further improving the quantization accuracy of the first weight matrix corresponding to the normal channels.
[0045] At the same time, when the GPTQ process for the first weight matrix corresponding to the normal channels is completed, the introduced quantization error is further compensated to the outlier weight part (such as Figure 2 the left diagonal shaded part of the W matrix in (2) in Figure 2 3 columns are exemplified), that is, the third weight matrix corresponding to the outlier channels. As described above, in the embodiments of the present application, the third weight matrix corresponding to the outlier channels can be quantized using non-quantization or INT8 quantization, and can be represented with higher precision to ensure the quantization accuracy of the third weight matrix, thereby improving the quantization accuracy of the language model.
[0046] In some embodiments, the process of determining the output result of the linear layer according to the second activation matrix, the second weight matrix, the third activation matrix corresponding to the outlier channel, and the third weight matrix in step S130 may include but is not limited to Figure 3a Steps S1301 - S1303 shown below.
[0047] Step S1301: Determine a first intermediate matrix according to the second activation matrix and the second weight matrix, and perform dequantization on the first intermediate matrix to obtain a floating - point data matrix.
[0048] Optionally, the first intermediate matrix may be, but is not limited to, the multiplication result of multiplying the second activation matrix and the second weight matrix, such as Figure 2 the shape (M, N) in
[0049] In this embodiment, considering that the second weight matrix is quantized to INT4 according to the output channels, and the second activation matrix is quantized to INT8 in the token dimension, that is, the data types of the second activation matrix and the second weight matrix are different, and matrix multiply - and - accumulate (MMA) can only be performed under the same data type. Therefore, in this embodiment, before determining the first intermediate matrix based on MMA, the second weight matrix can be converted from INT4 type to INT8 type, and then the second activation matrix of INT8 type is multiplied by the second weight matrix to obtain the first intermediate matrix.
[0050] In some embodiments, when converting the second weight matrix from INT4 type to INT8 type, reference can be made to Figure 3b as shown, shift the INT4 - type data in the second weight matrix 4 bits to the left logically to obtain INT8 - type data, where the 4 zeros in the dotted line are all obtained by left - shifting.
[0051] Step S1302: Determine a second intermediate matrix according to the third activation matrix and the third weight matrix corresponding to the outlier channel.
[0052] In some embodiments, the second intermediate matrix may be, but is not limited to, the multiplication result of multiplying the third activation matrix and the third weight matrix, such as Figure 2 the shape (M, N) in
[0053] Step S1303: Determine the output result of the linear layer according to the floating - point data matrix and the second intermediate matrix.
[0054] In some embodiments, depending on the difference in the second intermediate matrix, the determination method of the output result may be different. For example, if the second intermediate matrix is a floating-point data matrix, then the sum of the floating-point data matrix and the second intermediate matrix can be used as the output result; if the second intermediate matrix is a quantization matrix of INT8 type, then the second intermediate matrix can be dequantized into a floating-point data matrix first, and then the sum of the dequantized second intermediate matrix and the floating-point data matrix can be used as the output result.
[0055] In some embodiments, after the quantization of the large language model is completed through the foregoing embodiments, the large language model after the quantization of each linear layer, that is, the floating-point model, can be directly saved, and the quantization configuration information corresponding to the large language model can be saved for the inference system to load and infer the quantized large language model.
[0056] Among them, the quantization configuration information may include at least one of the first information and the second information. The first information is used to indicate the quantization bit number of the weights corresponding to the linear layers in the large language model, such as the quantization bit number {w_bit: 4} corresponding to the second weight matrix.
[0057] The second information is used to indicate whether the large language model is a floating-point model, so as to enable the inference system to determine the loading method of the large language model. For example, when the second information is {from_float: true}, the inference system can determine that the quantized large language model is a floating-point model, then the large language model can be loaded from the floating-point model. In this embodiment, for the method of loading the large language model from the floating-point model, it is especially friendly to scenarios where only quantization once is required to view the quantization effect, such as quantization exploration.
[0058] For another example, when the second information is {from_float: false}, the inference system can determine that the quantized large language model is a non-floating-point model, then the large language model can be loaded from the non-floating-point model. Among them, the non-floating-point model refers to that the quantization system has performed quantization compression on the obtained floating-point model.
[0059] In some embodiments, for the case where the second information indicates that the large language model is a non-floating-point model, after the quantization system dequantizes the first intermediate matrix in step S1301 described above to obtain a floating-point data matrix (that is, the quantized large language model is a floating-point model), it is also necessary to further perform quantization compression on the floating-point model to obtain a non-floating-point model, which can also be called a compressed model.
[0060] In this embodiment, the process of further quantizing and compressing the floating-point model to obtain a non-floating-point model may include Figure 4 the steps S140 - S150 shown below.
[0061] Step S140: For each of the linear layers after quantization, perform INT4 round-to-nearest (RTN) symmetric quantization on the weight matrix corresponding to the normal channels in the linear layer by output channel to obtain an INT4 weight matrix, pack the INT4 weight matrix into a predetermined data format, and save the first quantization parameter corresponding to the INT4 weight matrix, such as FP16 / BP16.
[0062] Among them, the predetermined data format may be, but is not limited to, the UINT32 data format. In this embodiment, "packing the INT4 weight matrix into the UINT32 data format" can be understood as having 8 INT4s in one UINT32.
[0063] Step S150: Perform INT8 RTN symmetric quantization on the weight matrix corresponding to the outlier channels in the linear layer by output channel to obtain an INT8 weight matrix, pack the INT8 weight matrix into the predetermined data format, and save the second quantization parameter corresponding to the INT8 weight matrix, such as FP16 / BP16.
[0064] Among them, the predetermined data format may be, but is not limited to, the UINT32 data format. In this embodiment, "packing the INT4 weight matrix into the UINT32 data format" can be understood as having 4 INT8s in one UINT32.
[0065] In this embodiment, through the compression quantization in steps S140 - S150, the inference system can directly load from the compressed and quantized non-floating-point model when loading the large language model, so as to improve the model loading efficiency.
[0066] Based on the quantization method of the large language model provided in the foregoing method embodiment 100, the following combines Figure 2 and Figure 5 to give an exemplary description of the quantization process of the large language model provided in the embodiments of the present application, and the content is as follows.
[0067] Step S501: Obtain activation statistics.
[0068] (1) Load the large language model to be quantized into a computing unit, such as a central processing unit (CPU), loop to obtain all the linear layers of the large language model, and register a hook function for the forward pass of each linear layer. (2) Extract a subset from the calibration dataset as the input to the large language model. For example, a subset consisting of 512 samples can be extracted from the calibration dataset, and the maximum length of each sample is 512.
[0069] (3) Load the large language model layer by layer onto the Graphics Processing Unit (GPU), and use the hook function to obtain the activation values corresponding to each channel of each linear layer in the hidden layer dimension.
[0070] (4) Statistically calculate the maximum value among the absolute values of the activation values corresponding to each channel of the linear layer in the hidden layer dimension. For example, traverse the absolute values of each activation value in sequence. Assume that the maximum value of the absolute value of the next sample (i.e., activation) in this channel is larger than the previous sample, then the larger value can be used as the statistical value for this channel, and so on, to obtain the activation statistics of each channel of the input of the linear layer under the subset, that is, the maximum value among the absolute values of the activation values corresponding to each channel.
[0071] (5) For the linear layer, after completing the activation statistics of this linear layer, move this linear layer out of the GPU and onto the CPU, and continue to repeat steps (3)-(4) to obtain the activation statistics of each channel of each linear layer in the hidden layer dimension.
[0072] (6) Store the activation statistics results in a dictionary (dict) structure, and the key when storing is the unique name of each linear layer in the large language model.
[0073] Step S502: Sort the channels in ascending order according to the maximum value corresponding to each channel to obtain the sorting result, record the reordered channel index, such as reorder_index, and store the sorting result reorder_index in a dictionary with the linear layer name as the key.
[0074] Step S503: Determine whether the sorting results of all linear layers in the large language model have been obtained, such as the obtaining of reorder_index. If so, execute step S504; otherwise, for the linear layers for which the sorting results have not been obtained, execute step S502 again until the sorting results of all linear layers in the large language model have been obtained.
[0075] Among them, the large language model records the rearrangement index of the channels (or activations) of each linear layer, which is used to rearrange the activations of each linear layer during model inference to ensure that the channel indices match the weights.
[0076] Step S504: For each linear layer, select a predetermined number (e.g., 256) of channels with lower ranks as the outlier channels according to the sorting result, and use the other channels except the outlier channels as the normal channels.
[0077] Step S505: Determine whether the channel partitioning of each linear layer is completed. If so, execute Step S506; otherwise, for the linear layer that has not been channel-partitioned, execute Step S504 again until the channel partitioning of each linear layer is completed.
[0078] Step S506: For each linear layer, perform INT8 quantization on the first activation matrix corresponding to the normal channels in the token dimension to obtain a second activation matrix, and perform INT4 GPTQ quantization on the first weight matrix corresponding to the normal channels according to the sorted Hessian matrix by output channels to obtain the second weight matrix. Do not quantize the third weight matrix and the third activation matrix corresponding to the outlier channels, or perform INT8 quantization on the third weight matrix and the third activation matrix corresponding to the outlier channels.
[0079] Among them, for all linear layers to be quantized, it can be understood as a decoding layer (DecoderLayer) with a block structure. In this embodiment, DecoderLayer can be loaded into the GPU for quantization.
[0080] Step S507: Determine whether the quantization of all linear layers is completed. If so, execute Step S508; otherwise, for the linear layer that has not been quantized, execute Step S506 again until the quantization of all linear layers is completed.
[0081] Step S508: Dequantize the INT-type weights in each quantized linear layer into floating-point weights to obtain a floating-point large language model. Save the large language model after the quantization of each linear layer, that is, the floating-point model, and save the quantization configuration information corresponding to the large language model. Among them, the second information in the quantization configuration information can be {from_float: true} to indicate that the large language model is a floating-point model.
[0082] Optionally, after obtaining the floating-point model, it can also be compressed and quantized through Step S140 - Step S150 and then saved. Based on this, the second information in the quantization configuration information can be {from_float: false} to indicate that the large language model is a non-floating-point model after compressed quantization.
[0083] The quantization scheme of the large language model provided in this embodiment loads the model into the GPU in chunks and rearranges the weights and activations according to the activation statistics to obtain a normal point part (corresponding to normal channels) and an outlier point part (corresponding to outlier channels). Then, dynamic INT8 quantization is performed on the first activation matrix corresponding to the normal channels, and error compensation is performed on the first weight matrix corresponding to the normal channels layer by layer using the second-order information of the matrix. Then, INT4 quantization is performed according to the output channels, and the third activation matrix and the third weight matrix corresponding to the outlier channels are quantized using non-quantization or INT8 quantization methods. To achieve the purpose of quantizing the large language model, the aforementioned quantization process provided in this embodiment adopts a mixed-precision quantization scheme based on W4A8, which can complete the model quantization process through a single GPU, reduce the video memory occupancy during the quantization process, and make the size of the quantized model not only compressed to 1 / 4 of the FP16 model, achieving the quantization compression effect of W4A16, but also restoring the accuracy of the quantized model and having the inference calculation ability of INT8×INT8.
[0084] Figure 6 FIG. 600 shows a schematic flowchart of an inference method 600 for a large language model provided by an embodiment of the present application. The method 600 can be executed by an inference system installed in an electronic device, where the electronic device can be a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. As Figure 6 shown, the method 600 may include the following steps.
[0085] Step S610, obtaining quantization configuration information corresponding to the large language model.
[0086] Among them, the large language model is a model quantized by each step in the foregoing method embodiment 100, and the large language model stores rearrangement indexes of channels in the hidden layer dimension of each linear layer.
[0087] In some embodiments, the quantization configuration information corresponding to the large language model can be obtained according to the model address provided by the user.
[0088] Step S620, performing model loading and inference according to the quantization configuration information.
[0089] Among them, the quantization configuration information may include at least one of the first information and the second information. The first information is used to indicate the quantization bit number of the weights corresponding to the linear layers in the large language model, such as the quantization bit number {w_bit: 4} corresponding to the second weight matrix.
[0090] The second piece of information is used to indicate whether the large language model is a floating-point model, so that the inference system can determine the loading method of the large language model. For example, when the second piece of information is {from_float: true}, the inference system can determine that the quantized large language model is a floating-point model. Then, the large language model can be loaded from the floating-point model.
[0091] For another example, when the second piece of information is {from_float: false}, the inference system can determine that the quantized large language model is a non-floating-point model. Then, the large language model can be loaded from the non-floating-point model. The non-floating-point model refers to that the quantization system performs quantization compression on the obtained floating-point model.
[0092] In some embodiments, for the case of loading from a floating-point model, the process of loading the model according to the quantization configuration information in step S620 may include, but is not limited to Figure 7a steps S6201 - S6204 shown below, and the content is as follows.
[0093] Step S6201, determine that the second piece of information (such as {from_float: true}) in the quantization configuration information indicates that the large language model is a floating-point model.
[0094] Step S6202, for each linear layer, perform INT4 RTN symmetric quantization on the weight matrix corresponding to the normal channels in the linear layer according to the number of quantization bits of the weights indicated by the first piece of information, obtain an INT4 weight matrix, pack the INT4 weight matrix into a predetermined data format, and save the first quantization parameter corresponding to the INT4 weight matrix.
[0095] Step S6203, perform INT8 RTN symmetric quantization on the weight matrix corresponding to the outlier channels in the linear layer according to the output channels, obtain an INT8 weight matrix, pack the quantized INT8 weight matrix into the predetermined data format, and save the second quantization parameter corresponding to the INT8 weight matrix.
[0096] Among them, in the foregoing S6202 - S6203, the outlier channels in the linear layer are a predetermined number of channels that are ranked later in the hidden layer dimension of the linear layer, and the normal channels are the channels in the linear layer except the outlier channels in the hidden layer dimension. In this embodiment, the predetermined number is the same as the predetermined number (such as 256) used by the foregoing quantization system for quantizing the large language model, so as to ensure the consistency of the channel division results between the inference system and the quantization system, and further ensure the accuracy of the inference results.
[0097] Step S6204: Load the large language model into a computing unit, such as a GPU, according to the INT4 weight matrix, the INT8 weight matrix, the first quantization parameter, and the second quantization parameter in the said predetermined data format.
[0098] In this embodiment, the model loading scheme from the floating-point model through steps S6201 - S6204 is that the inference system first reads the floating-point model into the memory of the CPU, quantizes the floating-point model, and automatically packages it into a compressed format (such as the aforementioned predetermined data format), and then loads the compressed quantized model into the GPU video memory. Thus, users do not need to concern themselves with the specific format of the model compression of the large language model and can directly use the inference system for inference.
[0099] In addition, for the case of loading from a non-floating-point model, the process of loading the model according to the quantization configuration information in step S620 may include, but is not limited to Figure 7b Steps S6205 - S6206 shown below.
[0100] Step S6205: Determine that the second information in the quantization configuration information indicates that the large language model is a non-floating-point model.
[0101] Step S6206: Load the large language model into the computing unit according to the predetermined parameters corresponding to the large language model.
[0102] Among them, the predetermined parameters include the INT4 weight matrix, the INT8 weight matrix in the predetermined data format, the first quantization parameter corresponding to the INT4 weight matrix, and the second quantization parameter corresponding to the INT8 weight matrix.
[0103] In the model loading method provided by S6205 - S6206 in this embodiment, the step of quantizing the floating-point model and automatically packaging it into a compressed format (such as the aforementioned predetermined data format) has been executed in the quantization system. That is, this embodiment directly loads from a non-floating-point model (from_float = false), or in other words, the inference system will load the compressed quantized non-floating-point model. This loading method does not require weight quantization during model loading (or deployment), saving model deployment time, but requires users to have an in-depth understanding of the compression format of the quantized model of this system or to use the compression and packaging tools provided by this system.
[0104] For the two model loading methods provided in the foregoing step S6201 - step S6204 and step S6205 - step S6206, regardless of which model loading method is used, due to the quantization scheme of the large language model in the foregoing Method Embodiment 100, when the large language model is used for model loading and inference, its video memory occupancy can be reduced to 1 / 4 of the FP16 / BF16 type model, while ensuring that the accuracy of the quantized model is almost lossless.
[0105] In some embodiments, when performing inference based on the large language model after loading is completed, the activation of each linear layer can be dynamically quantized in the token dimension according to the activation quantization method in the foregoing quantization system. For example, the activation of the normal channel can be quantized to INT8 in the token dimension, and the activation of the outlier channel can be quantized to INT8 in the token dimension or not quantized (consistent with the quantization method of the weight corresponding to the abnormal channel).
[0106] In some embodiments, considering that during model inference, the activation needs to be dynamically quantized, while the weight has been statically quantized by channel in the quantization system. Therefore, in order to minimize the inference latency as much as possible, in this embodiment, the operations of rearrangement and quantization of the activation can also be fused into the existing operators or newly added operators of the large language model to improve the inference speed during model inference. The fusion method will be described below.
[0107] (1) A first operator is fused into the precursor layer of at least one linear layer in the large language model, where the first operator is used to perform at least one of the rearrangement operation and the quantization operation on the activation of the at least one linear layer, and the rearrangement operation and / or quantization operation implemented by the first operator is static.
[0108] In some embodiments, the precursor layer can be an existing layer or a newly added layer in the large language model. Of course, when fusing the rearrangement operation and the quantization operation in this embodiment, the rearrangement of the corresponding activation in the current layer can be preferably fused into the operator in the previous layer (precursor layer), and new operators can be inserted where fusion is not possible. Among them, if the rearrangement is fused into the operator of the precursor layer, the rearrangement is static; if the rearrangement is only fused into the operation of the previous layer, the rearrangement is dynamic.
[0109] Exemplarily, assuming that the first operator is used to perform the rearrangement operation on the activation of the at least one linear layer, then, please combine Figure 8As shown, the activations of q_proj, k_proj, and v_proj are the same. The rearrangement of the activations of q_proj, k_proj, and v_proj can be fused into the LayerNorm layer in front of LlamaAttention to achieve dynamic rearrangement. The rearrangement of the activation of o_proj cannot be incorporated into the previous layer, so a reorder_and_requant operator can be added before o_proj to rearrange and quantize the activation of o_proj. The activations of gate_proj and up_proj are the same, and the rearrangement of the activations of gate_proj and up_proj can be fused into the LayerNorm after LlamaAttention. The rearrangement of the activation of down_proj is fused into the linear layers gate_proj and up_proj in front of it to perform a static rearrangement of the output channels of the weights of gate_proj and up_proj (output channels static reorder).
[0110] For another example, assuming that the first operator is used to perform a quantization operation on the activation of the at least one linear layer, then, please combine Figure 8 As shown, the quantization of the activations of q_proj, k_proj, and v_proj is fused into the LayerNorm layer in front of LlamaAttention. The quantization of the activation of o_proj cannot be incorporated into the previous layer, so a reorder_and_requant operator is added before o_proj to quantize the activation of o_proj. The quantization of the activations of gate_proj and up_proj is fused into the LayerNorm after LlamaAttention. The quantization of the activation of down_proj is fused into the activation in front of it.
[0111] (2) A second operator is fused into at least one linear layer in the large language model, and the second operator is used to perform at least one of an inverse quantization operation and an INT type conversion operation on the output of the linear layer.
[0112] For example, the INT type conversion operation can be fused into the linear layer operators corresponding to q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj.
[0113] Through the above fusion operations, the functions of the operators during the final inference are as Figure 8As shown. In addition to performing layer normalization operations, LayerNorm also incorporates rearrangement and quantization (the corresponding operators are shown in gray), and this operator needs to record the reorder_index. Before o_proj, a reorder_and_requant operator needs to be inserted specifically for rearrangement and quantization. The reorder index reorder_index can be recorded in the LlamaAttention where it is located because the reorder_and_requant operator only needs to exist in the inference system, and when implementing the quantization algorithm, its effect can be simulated. For all linear layer operators, in addition to the MMA that the linear layer needs to perform, an INT type conversion operation and an inverse quantization operation are also incorporated (the corresponding operators are shown in right diagonal shading). The activation operator before down_proj incorporates a quantization operation, which is shown in left diagonal shading.
[0114] In some embodiments, taking the LLaMA model in the large language model as an example, please refer to again Figure 8 As shown, considering that the LLaMA model is stacked by LlamaDecoderLayer, in order to ensure the accuracy of the quantized large language model, during quantization, the activations of all linear layers in LlamaDecoderLayer can be statistically sorted according to the specified metrics (such as the absolute value amplitude) through the calibration dataset to obtain the reorder_index. Then, during inference, the activations can be rearranged according to the reorder_index recorded in the large language model, as Figure 2 shown, the outliers are arranged to the end of the matrix and then calculations are performed; the rearrangement of the weights is performed statically and completed during model quantization, corresponding to Figure 8 the static reorder of the hidden dimensions in Figure 8 If the rearrangement of the activations can be incorporated into the weights of the previous layer operator, static rearrangement is also used, corresponding to
[0115] Based on this, after adopting the quantization and inference solutions provided by this application, the input of LlamaDecoderLayer is FP16 / BF16, and the input data of all linear layers in LlamaDecoderLayer is of INT8 type, such as q_proj, k_proj, v_proj, o_proj in the LlamaAteention structure, and gate_proj, up_proj, down_proj in LlamaMLP. Among them, for each linear layer, before the data is input into the linear layer, it is necessary to quantize the activation in the token dimension. Considering that this application uses dynamic INT8 quantization for the activation corresponding to the normal channels in the linear layer, static INT4 quantization for the weights corresponding to the normal channels, and INT8 quantization (or no quantization) for the activation and weights corresponding to the outlier channels, the output result of the linear layer is FP16 / BF16.
[0116] Among them, please refer to Figure 2 As shown, considering that the normal channels and outlier channels are quantized separately. And for the input of matrix multiplication involved in the outlier channels, all are INT8, and MMA can be performed. While for the MMA involved in the normal channels, the inputs are INT4 and INT8 respectively. Therefore, before MMA, it is necessary to convert INT4 to INT8, and then INT8 MMA can be performed. Finally, the operation result of MMA is dequantized to FP16 / BF16.
[0117] In some embodiments, this embodiment also provides a conversion method when converting INT4 to INT8. Among them, as Figure 9 shown, considering that the INT4 weights are packed in UINT32 format as signed INT4 as required, that is, there are 8 INT4s in one UINT32. Before the INT4 weights corresponding to the normal channels in the linear layer and the INT8 activation perform MMA, it is necessary to convert the INT4 weights compressed and packed in UINT32 to INT8 format. As can be seen from Figure 9 the left side, when converting INT4 to INT8, in fact, the original INT4 numbers are logically left-shifted by 4 bits, and the 4 zeros in the dotted line are all obtained by left-shifting.
[0118] Exemplarily, when converting INT4 weights in UINT32 to INT8 format, a fast INT type conversion operation can be performed as shown on the right side of 9. For example, assuming that 8 INT4 format weights are stored in UINT32 format, represented by W7, W6... W0 respectively, then, the weights in UINT32 format can be bitwise ANDed with 0xf0f0f0f0 to obtain the INT8 format corresponding to the odd-bit INT4 format weights; or, the weights in UINT32 format can be bitwise ANDed with 0x0f0f0f0f and then logically left-shifted by 4 bits to obtain the INT8 format corresponding to the even-bit INT4 format weights.
[0119] Based on the description of the foregoing method embodiment 600, the following combines Figure 10 to illustrate the inference process of the large language model provided by this application, the content is as follows.
[0120] Step S1001, obtain the quantized large language model and quantization configuration information according to the model address of the large language model provided by the user.
[0121] Step S1002, determine whether the large language model is a floating-point model according to the second information in the quantization configuration information. If so, first execute step S1003 to compress the quantized floating-point model, and then execute step S1004. Otherwise, directly execute step S1004.
[0122] Step S1003, load the floating-point model into the memory of the CPU, and then layer by layer load the weights of all linear layers in LlamaDecoderLayer into the GPU and perform quantization layer by layer. Among them, the weights corresponding to the normal channels are symmetrically quantized to INT4 RTN by output channel, the quantized INT4 weights are packed into the data format of UINT32, and the first quantization parameter is saved. The weights corresponding to the outlier channels (such as 256 channels in the hidden layer dimension) are symmetrically quantized to INT8 RTN by output channel, the quantized INT8 weights are packed into the data format of UINT32, and the second quantization parameter is saved. Step S1004, load the compressed quantized model (i.e., non-floating-point model) into the video memory of the GPU.
[0123] Among them, the inference system adopts symmetric per-channel quantization during implementation. Therefore, the parameters loaded by the inference system at this time include: the INT4 weight matrix corresponding to the normal channels and its corresponding first quantization parameter, and the INT8 weight matrix corresponding to the outlier channels and its corresponding second quantization parameter. The INT4 weights in the INT4 weight matrix corresponding to the normal channels are stored in UINT32 data format in groups of 8, and the INT8 weights in the INT8 weight matrix corresponding to the outlier channels are stored in UINT32 data format in groups of 4. The quantization parameter range (scales) is FP16 / BF16.
[0124] Step S1005, the inference system performs inference using the large language model that has been loaded into the GPU video memory according to the user's input.
[0125] Among them, taking the LLaMA model in the large language model as an example, when performing inference using the large language model that has been loaded into the GPU video memory, the workflow during inference can be seen Figure 8 as shown. For example, for each linear layer, perform INT8 dynamic quantization on the first activation matrix corresponding to the normal channels in the token dimension, etc., and perform INT8 dynamic quantization or no quantization on the third activation matrix corresponding to the abnormal channels in the token dimension, etc. The quantization method can refer to the relevant description in the aforementioned quantization system and will not be elaborated here.
[0126] In the inference scheme of the large language model provided in the embodiments of this method, on the one hand, when users use the quantized large language model, they do not need to understand the specific format of model compression, and only need to load the model according to the second information in the quantization configuration information; on the other hand, during model inference, for the activations and weights corresponding to the normal channels, the INT4 weights can be quickly converted to INT8 weights during inference. Therefore, whether it is the activations and weights corresponding to the normal channels or the activations and weights corresponding to the outlier channels, INT8×INT8 Tensor Core calculations are ultimately used, which can improve the inference speed.
[0127] Figure 11The figure shows a schematic structural diagram of a quantization device for a large language model provided by an embodiment of the present application. The device 1100 includes: a first processing module 1110 that divides the channels of each linear layer to be quantized in the hidden layer dimension of the large language model into normal channels and outlier channels; a second processing module 1120 that performs INT8 quantization on the first activation matrix corresponding to the normal channels in the token dimension to obtain a second activation matrix, and performs INT4 quantization on the first weight matrix corresponding to the normal channels according to the output channels to obtain a second weight matrix; and a third processing module 1130 that determines the output result of the linear layer according to the second activation matrix, the second weight matrix, the third activation matrix corresponding to the outlier channels, and the third weight matrix corresponding to the outlier channels.
[0128] In some embodiments, the dividing the channels of the linear layer in the hidden layer dimension into normal channels and outlier channels includes: dividing the channels of the linear layer in the hidden layer dimension into the normal channels and the outlier channels according to the importance of each channel of the linear layer in the hidden layer dimension; wherein, the importance of the channel is determined according to the activation statistics corresponding to the channel.
[0129] In some embodiments, the dividing the channels of the linear layer in the hidden layer dimension into the normal channels and the outlier channels according to the importance of each channel of the linear layer in the hidden layer dimension includes: statistically calculating the maximum value among the absolute values of the activation values corresponding to each channel of the linear layer in the hidden layer dimension; sorting the channels in ascending order according to the maximum value corresponding to each channel to obtain a sorting result; selecting a predetermined number of channels with a lower ranking as the outlier channels according to the sorting result, and using the other channels except the outlier channels as the normal channels.
[0130] In some embodiments, the performing INT4 quantization on the first weight matrix corresponding to the normal channels according to the output channels to obtain a second weight matrix includes: performing INT4 GPTQ quantization on the first weight matrix corresponding to the normal channels according to the output channels using the sorted Hessian matrix to obtain the second weight matrix.
[0131] In some embodiments, the determining the output result of the linear layer according to the second activation matrix, the second weight matrix, the third activation matrix corresponding to the outlier channels, and the third weight matrix corresponding to the outlier channels includes: determining a first intermediate matrix according to the second activation matrix and the second weight matrix, and performing dequantization on the first intermediate matrix to obtain a floating-point data matrix; determining a second intermediate matrix according to the third activation matrix and the third weight matrix corresponding to the outlier channels; and determining the output result of the linear layer according to the floating-point data matrix and the second intermediate matrix.
[0132] In some embodiments, determining the first intermediate matrix according to the second activation matrix and the second weight matrix includes: converting the second weight matrix from INT4 type to INT8 type; multiplying the second activation matrix by the second weight matrix to obtain the first intermediate matrix.
[0133] In some embodiments, the floating-point data matrix is a floating-point data matrix in FP16 / BF16 format.
[0134] In some embodiments, the third activation matrix and the third weight matrix are floating-point matrices; or, the third activation matrix is a matrix quantized to INT8 in the token dimension, and the third weight matrix is a matrix quantized to INT8 in the output channel dimension.
[0135] In some embodiments, after dequantizing the first intermediate matrix to obtain a floating-point data matrix, the third processing module 1130 is further configured to: for each of the linear layers after quantization, perform INT4 RTN symmetric quantization on the weight matrix corresponding to the normal channel in the linear layer by output channel to obtain an INT4 weight matrix, and pack the INT4 weight matrix into a predetermined data format, and save the first quantization parameter corresponding to the INT4 weight matrix; perform INT8 RTN symmetric quantization on the weight matrix corresponding to the outlier channel in the linear layer by output channel to obtain an INT8 weight matrix, and pack the INT8 weight matrix into the predetermined data format, and save the second quantization parameter corresponding to the INT8 weight matrix.
[0136] In some embodiments, the apparatus 1100 further includes: a model saving module, configured to save the large language model after quantizing each of the linear layers, and save the quantization configuration information corresponding to the large language model; wherein, the quantization configuration information includes at least one of the following: a first piece of information for indicating the number of quantization bits of the weights corresponding to the linear layers in the large language model; a second piece of information for indicating whether the large language model is a floating-point model.
[0137] The apparatus 1100 provided in the embodiments of the present application can execute the various methods described in the foregoing method embodiment 100, and implement the functions and beneficial effects of the various methods described in the foregoing method embodiment 100, which will not be elaborated herein.
[0138] Figure 12The structural schematic diagram of the quantization device 1200 of the large language model provided by the embodiment of the present application is shown. The device 1200 includes: a fourth processing module 1210, configured to obtain quantization configuration information corresponding to the large language model; a fifth processing module 1220, configured to perform model loading and inference according to the quantization configuration information; wherein, the large language model is quantized by using the methods described in the method embodiment 100, and the quantization configuration information includes at least one of the following: a first piece of information, used to indicate the weight quantization bit number corresponding to the linear layer in the large language model; a second piece of information, used to indicate whether the large language model is a floating-point model.
[0139] In some embodiments, the performing model loading according to the quantization configuration information includes: determining that the second piece of information in the quantization configuration information indicates that the large language model is a floating-point model; for each of the linear layers, performing INT4 RTN symmetric quantization on the weight matrix corresponding to the normal channels in the linear layer according to the weight quantization bit number indicated by the first piece of information for each output channel to obtain an INT4 weight matrix, and packing the INT4 weight matrix into a predetermined data format, and saving the first quantization parameter corresponding to the INT4 weight matrix; performing INT8 RTN symmetric quantization on the weight matrix corresponding to the outlier channels in the linear layer for each output channel to obtain an INT8 weight matrix, and packing the quantized INT8 weight matrix into the predetermined data format, and saving the second quantization parameter corresponding to the INT8 weight matrix; loading the large language model into a computing unit according to the INT4 weight matrix, the INT8 weight matrix, the first quantization parameter, and the second quantization parameter in the predetermined data format; wherein, the outlier channels in the linear layer are a predetermined number of channels with a later order in the hidden layer dimension of the linear layer, and the normal channels are the channels in the linear layer except the outlier channels in the hidden layer dimension.
[0140] In some embodiments, the performing model loading according to the quantization configuration information includes: determining that the second piece of information in the quantization configuration information indicates that the large language model is a non-floating-point model; loading the large language model into a computing unit according to the predetermined parameters corresponding to the large language model; wherein, the predetermined parameters include an INT4 weight matrix, an INT8 weight matrix, a first quantization parameter corresponding to the INT4 weight matrix, and a second quantization parameter corresponding to the INT8 weight matrix in a predetermined data format.
[0141] In some embodiments, a first operator is fused in the precursor layer of at least one linear layer in the large language model, and the first operator is used to perform at least one of a rearrangement operation and a quantization operation on the activation of the at least one linear layer.
[0142] In some embodiments, the precursor layer is an existing layer or a newly added layer in the large language model.
[0143] In some embodiments, a second operator is integrated in at least one linear layer of the large language model, and the second operator is used to perform at least one of an anti-quantization operation and an INT type conversion operation on the output of the linear layer.
[0144] The apparatus 1200 provided in the embodiments of the present application can execute the various methods described in the foregoing method embodiments, and implement the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated herein.
[0145] Figure 13 The figure shows a schematic hardware structure diagram of an electronic device provided in the embodiments of the present application. Referring to this figure, at the hardware level, the electronic device includes a processor, and optionally, an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0146] The processor, the network interface, and the memory can be interconnected through an internal bus, and the internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a bidirectional arrow is used in this figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0147] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0148] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, and forms a device for locating a specified user at the logical level. The processor executes the program stored in the memory, and is specifically used to execute: Figures 1-6 The methods disclosed in the illustrated embodiments and implement the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated herein.
[0149] As described above in the present application Figures 1-6 The method disclosed in the above embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in the form of software. The above processor may be a general-purpose processor, including GPU, CPU, Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0150] The electronic device can also execute each method described in the foregoing method embodiments and implement the functions and beneficial effects of each method described in the foregoing method embodiments, which will not be elaborated here.
[0151] Of course, in addition to the software implementation, the electronic device of the present application does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.
[0152] The embodiments of the present application also propose a computer-readable storage medium. The computer-readable storage medium stores one or more programs. When the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute Figures 1-6 the method disclosed in the foregoing embodiments and implement the functions and beneficial effects of each method described in the foregoing method embodiments, which will not be elaborated here.
[0153] Among them, the computer-readable storage medium includes a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disc, etc.
[0154] The embodiments of the present application also provide a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the following processes are implemented: Figures 1-6 The method disclosed in the illustrated embodiment realizes the functions and beneficial effects of each method described in the foregoing method embodiments, and will not be elaborated herein.
[0155] A computer-readable storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable storage medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0156] In summary, the above are only the preferred embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0157] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0158] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an …" does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0159] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
Claims
1. A quantization method for large language models, comprising: For each linear layer to be quantized in the large language model, dividing the channels of the linear layer in the hidden layer dimension into normal channels and outlier channels; Performing INT8 quantization on the first activation matrix corresponding to the normal channels in the token dimension to obtain a second activation matrix, and performing INT4 quantization on the first weight matrix corresponding to the normal channels according to the output channels to obtain a second weight matrix; Determining the output result of the linear layer according to the second activation matrix, the second weight matrix, the third activation matrix corresponding to the outlier channels, and the third weight matrix corresponding to the outlier channels.
2. The method according to claim 1, wherein, The step of dividing the channels of the linear layer in the hidden layer dimension into normal channels and outlier channels includes: Dividing the channels of the linear layer in the hidden layer dimension into the normal channels and the outlier channels according to the importance of each channel of the linear layer in the hidden layer dimension; Wherein, the importance of the channel is determined according to the activation statistics corresponding to the channel.
3. The method according to claim 2, wherein The step of dividing the channels of the linear layer in the hidden layer dimension into the normal channels and the outlier channels according to the importance of each channel of the linear layer in the hidden layer dimension includes: Statistical maximum value of the absolute values of the activation values corresponding to each channel of the linear layer in the hidden layer dimension; Ascendingly sorting the channels according to the maximum value corresponding to each channel to obtain a sorting result; Selecting a predetermined number of channels with lower rankings as the outlier channels according to the sorting result, and taking the other channels except the outlier channels as the normal channels.
4. The method according to claim 1, wherein, The step of performing INT4 quantization on the first weight matrix corresponding to the normal channels according to the output channels to obtain a second weight matrix includes: Performing INT4 GPTQ quantization on the first weight matrix corresponding to the normal channels using the sorted Hessian matrix according to the output channels to obtain the second weight matrix.
5. The method according to claim 1, wherein, The step of determining the output result of the linear layer according to the second activation matrix, the second weight matrix, the third activation matrix corresponding to the outlier channels, and the third weight matrix corresponding to the outlier channels includes: Determining a first intermediate matrix according to the second activation matrix and the second weight matrix, and performing dequantization on the first intermediate matrix to obtain a floating-point data matrix; Determining a second intermediate matrix according to the third activation matrix and the third weight matrix corresponding to the outlier channels; Determining the output result of the linear layer according to the floating-point data matrix and the second intermediate matrix.
6. The method according to claim 5, wherein The step of determining a first intermediate matrix according to the second activation matrix and the second weight matrix includes: Converting the second weight matrix from INT4 type to INT8 type; Multiplying the second activation matrix and the second weight matrix to obtain the first intermediate matrix.
7. The method according to claim 5, wherein, The floating-point data matrix is a floating-point data matrix in FP16 / BF16 format.
8. The method according to claim 5, wherein The third activation matrix and the third weight matrix are floating-point matrices; Alternatively, the third activation matrix is a matrix quantized to INT8 in the token dimension, and the third weight matrix is a matrix quantized to INT8 in the output channel dimension.
9. The method according to claim 5, wherein, After the dequantization of the first intermediate matrix to obtain a floating-point data matrix, the method further includes: For each of the linear layers after quantization, the weight matrix corresponding to the normal channels in the linear layer is symmetrically quantized to INT4 RTN in the output channel to obtain an INT4 weight matrix, and the INT4 weight matrix is packed into a predetermined data format, and the first quantization parameter corresponding to the INT4 weight matrix is saved; The weight matrix corresponding to the outlier channels in the linear layer is symmetrically quantized to INT8 RTN in the output channel to obtain an INT8 weight matrix, and the INT8 weight matrix is packed into the predetermined data format, and the second quantization parameter corresponding to the INT8 weight matrix is saved.
10. The method according to any one of claims 5-9, wherein, The method further includes: Saving the large language model after quantization of each of the linear layers, and saving the quantization configuration information corresponding to the large language model; Wherein, the large language model records the rearrangement index of the channels corresponding to each linear layer, and the quantization configuration information includes at least one of the following: The first information, which is used to indicate the bit number of weight quantization corresponding to the linear layer in the large language model; The second information, which is used to indicate whether the large language model is a floating-point model.
11. An inference method for a large language model, where the large language model is quantized by using the method according to any one of claims 1 to 10, and the inference method includes: Obtaining the quantization configuration information corresponding to the large language model; Performing model loading and inference according to the quantization configuration information; Wherein, the large language model records the rearrangement index of the channels corresponding to each linear layer, and the quantization configuration information includes at least one of the following: The first information, which is used to indicate the bit number of weight quantization corresponding to the linear layer in the large language model; The second information, which is used to indicate whether the large language model is a floating-point model.
12. The method according to claim 11, wherein, The performing model loading according to the quantization configuration information includes: Determining that the second information in the quantization configuration information indicates that the large language model is a floating-point model; For each of the linear layers, the weight matrix corresponding to the normal channels in the linear layer is symmetrically quantized to INT4 RTN in the output channel according to the bit number of weight quantization indicated by the first information to obtain an INT4 weight matrix, and the INT4 weight matrix is packed into a predetermined data format, and the first quantization parameter corresponding to the INT4 weight matrix is saved; The weight matrix corresponding to the outlier channels in the linear layer is symmetrically quantized to INT8 RTN in the output channel to obtain an INT8 weight matrix, and the quantized INT8 weight matrix is packed into the predetermined data format, and the second quantization parameter corresponding to the INT8 weight matrix is saved. Load the large language model into the computing unit according to the INT4 weight matrix, the INT8 weight matrix, the first quantization parameter, and the second quantization parameter in the said predetermined data format; Among them, the outlier channels in the said linear layer are a predetermined number of channels that are ranked at the end in the hidden layer dimension of the linear layer, and the normal channels are the channels in the hidden layer dimension of the linear layer except the outlier channels.
13. The method according to claim 11, wherein The model loading according to the quantization configuration information includes: Determine that the second information in the quantization configuration information indicates that the large language model is a non-floating point model; Load the large language model into the computing unit according to the predetermined parameters corresponding to the large language model; Among them, the said predetermined parameters include the INT4 weight matrix, the INT8 weight matrix in the predetermined data format, the first quantization parameter corresponding to the INT4 weight matrix, and the second quantization parameter corresponding to the INT8 weight matrix.
14. The method according to any one of claims 11-13, wherein, A first operator is fused in the precursor layer of at least one linear layer in the large language model, where the first operator is used to perform at least one of a rearrangement operation and a quantization operation on the activation of the at least one linear layer.
15. The method according to claim 14, wherein, The precursor layer is an existing layer or a newly added layer in the large language model.
16. The method according to any one of claims 11-13, wherein, A second operator is fused in at least one linear layer in the large language model, and the second operator is used to perform at least one of an inverse quantization operation and an INT type conversion operation on the output of the linear layer.
17. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the method according to any one of claims 1-16 are implemented.
18. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the method according to any one of claims 1-16 are implemented.
Citation Information
Patent Citations
Model quantification method and device, equipment and storage medium
CN114936619A
Quantization method and reasoning method and device of large language model, equipment and medium
CN118036755A
Model quantification method and device, electronic equipment and storage medium
CN119337045A
Cited By
Quantization method and system of large language model and electronic equipment
CN121031661A
Mixing precision quantification method, device, equipment, medium and program product
CN122173753A
Shared quantum kernel-based large model compression training method and apparatus, and electronic device
CN122389998A
Quantization method and inference method for large language model, and electronic device
WO2026158311A1