Model quantization method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202610813654.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-15
AI Technical Summary
[0003]相关技术中,通过标量量化将权重参数从高精度的浮点格式转换为固定位宽的整型格式,减小了模型的存储空间需求,但是现有技术的性能随着量化精度的下降呈指数级衰减,模型性能损失严重
Smart Images

Figure CN122759232A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to fields such as artificial intelligence technology and computer vision technology, and in particular to a model quantization method, apparatus, device, and storage medium. Background Technology
[0002] Large Language Models (LLMs) are deep learning models trained on large amounts of text data. They possess strong versatility, adaptability, and reasoning capabilities in tasks such as natural language processing, multimodal understanding, and intelligent interaction. To facilitate the migration of LLMs to edge and end-device devices, they typically require quantization to reduce storage requirements and accelerate inference.
[0003] In related technologies, scalar quantization converts weight parameters from high-precision floating-point format to fixed-width integer format, reducing the model's storage space requirements. However, the performance of existing technologies degrades exponentially with decreasing quantization precision, resulting in severe performance loss for the model. Summary of the Invention
[0004] To address the aforementioned technical problems, embodiments of this disclosure provide a model quantization method, apparatus, device, and storage medium.
[0005] According to one aspect of the embodiments of this disclosure, a model quantization method is provided, comprising:
[0006] The calibration dataset is input into the large language model to be quantized for forward computation to obtain the first activation value of each linear layer.
[0007] Based on the symmetric ladder quantizer and the first activation value, the expected mean square loss of the symmetric ladder quantization process is determined. The expected mean square loss includes the rounding loss and underflow loss of the activation value.
[0008] Based on the rounding loss and underflow loss, determine the optimal rotation matrix;
[0009] Based on the first activation value and the optimized rotation matrix, the large language model to be quantized is subjected to multi-stage quantization processing to obtain the quantized target model, which is then deployed to a hardware device. The quantized target model can be called by the hardware device to process corresponding tasks.
[0010] In some optional implementations, determining the expected mean square loss of the symmetric gradient quantization process based on the symmetric gradient quantizer and the first activation value includes:
[0011] Using a standard rotation matrix, the first activation values of each linear layer are initially rotated to obtain the second activation values of each linear layer;
[0012] The expected mean square loss is determined based on the symmetric gradient quantizer and the second activation value.
[0013] In some optional implementations, determining the optimized rotation matrix based on the rounding loss and underflow loss of the activation value includes:
[0014] Differentiable processing is performed on the rounding loss and the underflow loss respectively to obtain the differentiable expected loss;
[0015] Based on the differentiable expected loss, the optimal rotation matrix is determined.
[0016] In some optional implementations, the multi-stage quantization process of the large language model to be quantized based on the first activation value and the optimized rotation matrix includes:
[0017] Based on the optimized rotation matrix, the first activation value of each linear layer is geometrically reshaped to obtain the third activation value of each linear layer.
[0018] Based on the optimized rotation matrix, the first Hessian matrix is determined;
[0019] The large language model to be quantized is quantized based on the third activation value and the first Hessian matrix.
[0020] In some alternative implementations, determining the first Hessian matrix based on the optimized rotation matrix includes:
[0021] Based on the optimized rotation matrix, the large language model to be quantized is transformed to obtain the transformed large language model.
[0022] The calibration dataset is input into the transformed large language model for forward computation to obtain the first Hessian matrix.
[0023] In some alternative implementations, it also includes:
[0024] When the calibration dataset is input into the large language model to be quantized for forward computation, the second Hessian matrix is obtained.
[0025] Determining the first Hessian matrix based on the optimized rotation matrix includes:
[0026] The first Hessian matrix is determined based on the second Hessian matrix and the optimized rotation matrix.
[0027] In some optional implementations, quantizing the large language model to be quantized based on the third activation value and the first Hessian matrix includes:
[0028] The third activation value is quantized to obtain the fourth activation value;
[0029] Based on the third activation value, the fourth activation value, the first Hessian matrix, and the original weights of the large language model to be quantized, the weight optimization objective is determined.
[0030] Determine the autocorrelation matrix of the fourth activation value, and the cross-correlation matrix of the third and fourth activation values;
[0031] Based on the autocorrelation matrix, the cross-correlation matrix, and the original model weights, a weight offset center is determined. The weight offset center is used to compensate for the cumulative distortion error of the activation values of each linear layer.
[0032] The large language model to be quantized is quantized based on the weight offset center, the first Hessian matrix, and the weight optimization objective.
[0033] In some optional implementations, the quantization of the large language model to be quantized based on the weight offset center and the first Hessian matrix includes:
[0034] Transform the first Hessian matrix into a lower triangular matrix;
[0035] The transformed lower triangular matrix is normalized to obtain a strictly lower triangular matrix.
[0036] Based on the strict lower triangular matrix and the weight optimization objective, the weights of the large language model to be quantized are iteratively quantized.
[0037] In response to the fulfillment of the iteration termination condition, the model quantization weights of the large language model to be quantized are obtained.
[0038] According to another aspect of the present disclosure, a model quantization apparatus is provided, comprising:
[0039] The forward computation module is used to input the calibration dataset into the large language model to be quantized and perform forward computation to obtain the first activation value of each linear layer.
[0040] The first determining module is used to determine the expected mean square loss of the symmetric gradient quantization process based on the symmetric gradient quantizer and the first activation value. The expected mean square loss includes the rounding loss and underflow loss of the activation value.
[0041] The second determining module is used to determine the optimized rotation matrix based on the rounding loss and underflow loss;
[0042] The quantization module is used to perform multi-stage quantization processing on the large language model to be quantized based on the first activation value and the optimized rotation matrix to obtain the quantized target model, so as to deploy the quantized target model to the hardware device, and the quantized target model can be called by the hardware device to process corresponding tasks.
[0043] In some optional implementations, the first determining module includes:
[0044] The rotation transformation submodule is used to perform initial rotation processing on the first activation values of each linear layer using a standard rotation matrix to obtain the second activation values of each linear layer.
[0045] The first determining submodule is used to determine the expected mean square loss based on the symmetric middle gradient quantizer and the second activation value.
[0046] In some optional implementations, the second determining module includes:
[0047] The first processing submodule is used to perform differentiable processing on the rounding loss and the underflow loss respectively to obtain the differentiable expected loss;
[0048] The second determining submodule is used to determine the optimized rotation matrix based on the differentiable expected loss.
[0049] In some alternative implementations, the quantization module includes:
[0050] The second processing submodule is used to perform geometric reshaping processing on the first activation value of each linear layer based on the optimized rotation matrix to obtain the third activation value of each linear layer.
[0051] The third determining submodule is used to determine the first Hessian matrix based on the optimized rotation matrix;
[0052] The quantization submodule is used to quantize the large language model to be quantized based on the third activation value and the first Hessian matrix.
[0053] In some optional implementations, the third determining submodule is used to transform the large language model to be quantized based on the optimized rotation matrix to obtain the transformed large language model; and input the calibration dataset into the transformed large language model for forward computation to obtain the first Hessian matrix.
[0054] In some alternative implementations, it also includes:
[0055] The acquisition module is used to acquire the second Hessian matrix when the calibration dataset is input into the large language model to be quantized for forward computation;
[0056] The third determining submodule is used to determine the first Hessian matrix based on the second Hessian matrix and the optimized rotation matrix.
[0057] In some optional implementations, the quantization submodule is used to quantize the third activation value to obtain a fourth activation value; determine a weight optimization objective based on the third activation value, the fourth activation value, the first Hessian matrix, and the original model weights of the large language model to be quantized; determine the autocorrelation matrix of the fourth activation value, and the cross-correlation matrix of the third and fourth activation values; determine a weight offset center based on the autocorrelation matrix, the cross-correlation matrix, and the original model weights, the weight offset center being used to compensate for the cumulative distortion error of the activation values of each linear layer; and quantize the large language model to be quantized based on the weight offset center, the first Hessian matrix, and the weight optimization objective.
[0058] In some optional implementations, the quantization submodule is used to transform the first Hessian matrix into a lower triangular matrix; normalize the transformed lower triangular matrix to obtain a strict lower triangular matrix; perform iterative quantization on the weights of the large language model to be quantized based on the strict lower triangular matrix and the weight optimization objective; and obtain the model quantization weights of the large language model to be quantized in response to the fulfillment of the iteration termination condition.
[0059] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which stores computer program instructions that, when executed, implement the above-described method.
[0060] According to another aspect of the present disclosure, an electronic device is provided, the electronic device comprising:
[0061] Memory, used to store computer program products;
[0062] A processor is configured to execute a computer program product stored in memory, and when the computer program product is executed, to implement the above method.
[0063] According to another aspect of the present disclosure, a computer program product is provided, including computer program instructions that, when executed by a processor, implement the above-described method.
[0064] Based on the embodiments of this disclosure, for scenarios requiring quantization of large language models, a calibration dataset can be first input into the large language model to be quantized for forward computation to obtain the first activation values of each linear layer. Then, based on the symmetric gradient quantizer and the first activation values, the expected mean square loss of the symmetric gradient quantization process is determined. This expected mean square loss includes the rounding loss and underflow loss of the activation values. Based on the rounding loss and underflow loss, an optimized rotation matrix is determined. Finally, based on the first activation values and the optimized rotation matrix, the large language model to be quantized undergoes multi-stage quantization processing to obtain the quantized target model. The quantized target model can then be deployed to a hardware device, where it can be called by the hardware device to perform corresponding tasks. This resolves the coupling contradiction between outlier suppression and increased underflow error caused by introducing an orthogonal rotation matrix to rotate the activation values, achieving globally optimal control of quantization error. Furthermore, through the multi-stage model quantization process, the weights can be adaptively quantized based on the cumulative distortion input of the activation values, adapting to the cumulative distortion input and effectively achieving a balance between high precision and high efficiency under bit quantization.
[0065] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0066] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0067] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description, wherein:
[0068] Figure 1 This is a schematic diagram of the multi-stage model quantization process in the model quantization method disclosed herein;
[0069] Figure 2 This is a flowchart illustrating an embodiment of the model quantization method disclosed herein;
[0070] Figure 3 This is a flowchart illustrating step 203 in the model quantization method of this disclosure;
[0071] Figure 4 This is a flowchart illustrating step 204 in the model quantization method of this disclosure;
[0072] Figure 5 This is a flowchart illustrating step 243 in the model quantization method of this disclosure;
[0073] Figure 6 This is a schematic diagram of the activation value reshaping process in the model quantization method of this disclosure;
[0074] Figure 7 This is a schematic diagram of the structure of one embodiment of the model quantization device of this disclosure;
[0075] Figure 8 This is a schematic diagram of the structure of yet another embodiment of the model quantization device disclosed herein;
[0076] Figure 9 This is a structural diagram of an electronic device provided as an illustrative embodiment of the present disclosure. Detailed Implementation
[0077] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0078] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0079] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0080] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0081] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0082] Furthermore, in this embodiment, the term "and / or" merely describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this embodiment, the character " / " generally indicates that the preceding and following associated objects have an "or" relationship.
[0083] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0084] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0085] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0086] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0087] In order to accurately describe the technical content in this disclosure, and to accurately understand the embodiments of this disclosure, the terms used in the embodiments of this disclosure are explained or defined as follows:
[0088] 1) SmoothQuant: By adjusting the scaling ratio of activation values and weights, the dynamic range of activation values is reduced, thereby reducing quantization error. However, SmoothQuant does not consider the cumulative effect of quantization error in multi-layer propagation, and its performance improvement is limited in low-bit quantization scenarios.
[0089] 2) GPTQ (Generative Pretrained Quantization): This is a quantization technique used for generative pretrained Transformer models. GPTQ uses an iterative optimization approach, quantizing the weights layer by layer. It optimizes the quantization parameters by minimizing the mean squared error (MSE) between the quantized weights and the original weights. However, GPTQ assumes that the quantized activation inputs are clean inputs (i.e., inputs not contaminated by the quantization errors of previous layers), which does not match the cumulative distortion of activation inputs in real-world quantization scenarios, leading to a disconnect between the optimization objective and actual needs.
[0090] 3) SpinQuant and QuaRot: These methods introduce orthogonal rotation matrices to rotate model parameters such as activation values or weights, adjusting the distribution characteristics of the model parameters to make them more suitable for quantization and thus reducing quantization error. These methods primarily focus on suppressing outliers through rotation, but they do not address the "dead-zone" underflow problem that occurs during rotation. This causes some small values to be mapped to zero after rotation, introducing additional underflow error and creating a coupled contradiction between "outlier suppression and increased underflow error," making it difficult to achieve globally optimal control of quantization error.
[0091] OSTQuant and DartQuant: As relatively advanced quantization methods in the target field, they have made improvements in rotation optimization and error suppression, but they have not yet established a systematic compensation mechanism for accumulated distortion. Furthermore, during the weight quantization process, they have not fully considered the impact of output sensitivity on quantization error, which leads to quantization noise being concentrated in channels that are sensitive to model task loss, thus affecting the performance of the quantized model.
[0092] Furthermore, the symmetric mid-tread quantization method used in related technologies has an inherent "dead zone" characteristic (values within the interval (-Δ / 2, Δ / 2) are mapped to zero), which leads to underflow errors. At the same time, in the optimization process of the rotation matrix, the use of non-differentiable objective functions (such as error calculation based on maximum value or discrete count) results in variance in gradient estimation, making the optimization process unstable and affecting the rotation effect.
[0093] Rounding loss: When high-precision floating-point numbers (such as FP32) caused by model quantization are converted to low-bit integers (such as INT8, INT4), continuous weights or activation values are forcibly mapped to discrete quantization levels, resulting in information loss.
[0094] Underflow loss: In model quantization, the rotation operation may further compress the originally non-zero small amplitude components to an extremely small range, which cannot be represented in low-precision formats and is thus truncated to zero, resulting in a loss.
[0095] Based on the above-mentioned technical problems in related technologies, the present disclosure provides a model quantization method that solves the coupling contradiction between outlier suppression and increased underflow error caused by introducing an orthogonal rotation matrix to rotate the activation value. It achieves global optimal control of quantization error, and through a multi-stage model quantization process, it can adaptively quantize the weights based on the cumulative distortion input of the activation value, adapt to the cumulative distortion input, and effectively achieve a balance between high precision and high efficiency under bit quantization.
[0096] Application Overview
[0097] Figure 1 This diagram illustrates the multi-stage model quantization process in the model quantization method disclosed herein. Figure 1As shown, in the model quantization method disclosed herein, the calibration dataset is input into the large language model to be quantized, and during the forward computation process, the activation values (including input activation values and output activation values) of network layers such as the embedding layer, multi-head self-attention layer (MHSA), feed-forward network layer (FFN), and language model head (LM head) in the large language model are collected. Then, the activation values (first activation values) are rotated by the initial rotation matrix R to obtain the rotated activation values (second activation values). In the first stage of model quantization, the rotated activation values are quantized using a symmetric gradient quantizer to determine the expected mean square loss of symmetric gradient quantization. This loss includes rounding loss and underflow loss of the activation values. By differentiating the rounding loss and underflow loss respectively, the cure loss is obtained. This cure loss is a differentiable expected loss. The rotation matrix R can be optimized using the differentiable expected loss to obtain an optimized rotation matrix. The optimized rotation matrix is then used to geometrically reshape the activation values to obtain the third activation value. In the second stage of model quantization, a distortion-compensated rounding (DCR) mechanism is used to construct a weight optimization objective function based on output sensitivity to address the cumulative distortion situation. The weight optimization objective is to minimize the Hessian matrix weighted error between the quantized output and the original output. The quantization noise is directed to a channel insensitive to task loss, and the optimal offset center is determined based on the autocorrelation matrix of the distorted input and the cross-correlation matrix of the original and distorted inputs. Then, based on the offset optimal center, the input Hessian matrix and the output Hessian matrix are decomposed into lower triangular matrices to obtain the lower triangular matrices of the input and output Hessian matrices. Based on the offset optimal center and the lower triangular matrices of the input and output Hessian matrices, model weight quantization and weight iterative rounding update are performed to obtain the final quantized weights.
[0098] In this disclosed technical solution, a "two-stage quantization" framework is adopted. In the first stage, the activation value is geometrically reshaped through a differentiable coupling error optimization (CURE) mechanism to balance outlier suppression and underflow error. In the second stage, the weight is quantized through a distortion compensated rounding (DCR) mechanism to adapt to the accumulated distortion input and consider the output sensitivity, ultimately achieving a balance between high precision and high efficiency under low bit quantization.
[0099] Exemplary methods
[0100] Figure 2This is a flowchart illustrating an embodiment of the model quantization method of this disclosure; the model quantization method can be applied to terminal devices (such as computer devices), such as... Figure 2 As shown, the model quantization method includes the following steps 201 to 204. Each step is explained below.
[0101] In step 201, the calibration dataset is input into the large language model to be quantized for forward computation to obtain the first activation value of each linear layer.
[0102] Among them, the large language model to be quantized refers to a deep learning model trained on a large amount of text data, enabling the model to generate natural language text or understand language text. It can process various types of text content from the input text data, including conversational AI, chatbots, marketing content, and code assistants. A multimodal model is a model trained by combining text, images, video, audio, and other multimodal information. The input and output of a multimodal model can include multiple forms, such as text, images, audio, and video.
[0103] The calibration dataset for model quantization refers to a small set of sample data used to determine the quantization parameters (such as scaling factor, zero point, threshold, etc.) of the weights and activation values of each layer of the model during the quantization process of a large language model to be quantized.
[0104] In this embodiment, forward computation includes calculating layer by layer from the input layer of the network until the result of the output layer is calculated, that is, the process of calculating the output given a set of inputs. During the process of calculating layer by layer, for each calibration data, each linear layer can collect activation values, including input activation values and output activation values, which are also the first activation values (original activation values) of the model.
[0105] In step 202, based on the symmetric ladder quantizer and the first activation value, the expected mean square loss of the symmetric ladder quantization process is determined. The expected mean square loss includes the rounding loss and underflow loss of the activation value.
[0106] Among them, the symmetric ladder quantizer is a quantization method in which the quantization level of the symmetric ladder quantizer is symmetrically distributed about zero, and the zero value is located at the midpoint of two adjacent quantization intervals.
[0107] In this disclosure, when quantizing the activation values using a symmetric mid-ladder quantizer, a standard rotation matrix can be used to perform initial rotation processing on the first activation values of each linear layer to obtain the second activation values of each linear layer. Based on the symmetric mid-ladder quantizer and the second activation values, the expected mean square loss is determined.
[0108] For any vector (d is the vector dimension), the model quantization bit width is b, and a symmetric gradient quantizer is used. The definition is as shown in equation (1).
[0109] Equation (1)
[0110] In equation (1), To quantize the step size, To quantize the maximum value; symmetric gradient quantizers have an inherent "dead zone". Values falling into this range will be mapped to zero, resulting in underflow error.
[0111] To quantify the aforementioned underflow error, this disclosure derives an upper bound for the expected mean square loss (MSE) of symmetric in-gradient quantization: It is obtained by multiplying the original activation value by the standard rotation matrix.
[0112] Equation (2)
[0113] In equation (2), , The second activation value The dynamic range is determined by outliers. This is the number of elements that fall into the "dead zone" (underflow count). For vector dimensions, The rounding loss of the activation value, The underflow loss of the activation value.
[0114] The upper bound reveals the coupling contradiction between rounding loss and underflow loss: reducing (Suppressing outliers) reduces the quantization step size. ,lead to Increase (increase underflow error), increase (Suppressing outliers) increases the quantization step size. ,lead to Reduce (underflow error is reduced).
[0115] In step 203, the optimized rotation matrix is determined based on rounding loss and underflow loss.
[0116] In this embodiment, the optimized rotation matrix is a matrix used to optimize the standard rotation matrix. By determining the optimized rotation matrix to perform rotation transformation on the model activation values, it can be ensured that the expected mean square loss composed of rounding loss and underflow loss is smaller.
[0117] In this embodiment, the specific implementation method for determining the optimized rotation matrix can be found in [reference needed]. Figure 3 The embodiments shown are not described in detail here.
[0118] In step 204, based on the first activation value and the optimized rotation matrix, the large language model to be quantized is subjected to multi-stage quantization processing to obtain the quantized target model, which is then deployed to the hardware device. The quantized target model can be called by the hardware device to process corresponding tasks.
[0119] In this embodiment, after determining the optimized rotation matrix, the activation values in the large language model to be quantized can be rotated and transformed using the optimized rotation matrix in the first stage, that is, the activation values are geometrically reshaped. Then, in the second stage, the geometrically reshaped activation values are used to achieve high-precision quantization of weights through the DCR mechanism by adapting the input after cumulative distortion and incorporating the output sensitivity.
[0120] In this embodiment, the specific implementation process of multi-stage quantization of the model can be found in [reference needed]. Figure 4 The embodiments shown are not described in detail here.
[0121] Through steps 201 to 204 above, for scenarios requiring quantization of large language models, the calibration dataset can be input into the large language model to be quantized for forward computation to obtain the first activation value of each linear layer. Then, based on the symmetric gradient quantizer and the first activation value, the expected mean square loss of the symmetric gradient quantization process is determined. This expected mean square loss includes the rounding loss and underflow loss of the activation value. Based on the rounding loss and underflow loss, the optimized rotation matrix is determined. Finally, based on the first activation value and the optimized rotation matrix, the large language model to be quantized undergoes multi-stage quantization processing to obtain the quantized target model. The quantized target model can then be deployed to a hardware device, where it can be called by the hardware device to perform corresponding tasks. This solves the coupling contradiction between outlier suppression and increased underflow error caused by introducing an orthogonal rotation matrix to rotate the activation value, achieving globally optimal control of quantization error. Moreover, through the multi-stage model quantization process, the weights can be adaptively quantized based on the cumulative distortion input of the activation value, adapting to the cumulative distortion input and effectively achieving a balance between high precision and high efficiency under bit quantization.
[0122] Figure 3 This is a flowchart illustrating step 203 of the model quantization method disclosed herein. Figure 3 As shown above, in the above Figure 2 Based on the illustrated embodiment, the specific process of determining the optimized rotation matrix based on rounding loss and underflow loss may include steps 231-232. Each step is explained below.
[0123] In step 231, the rounding loss and underflow loss are made differentiable to obtain the differentiable expected loss.
[0124] In this embodiment, since the dynamic range of the second activation value v in the rounding error and the underflow count in the underflow loss are both non-differentiable functions and cannot be directly used for gradient optimization, this embodiment designs a method for differentiable transformation of the rounding loss and a method for differentiable transformation of the underflow loss.
[0125] Specifically, let's continue to use the upper bound of the expected error provided by equation (2) above as an example to illustrate the differentiable processing process. For the differentiable transformation of outlier suppression (rounding loss), the log-sum-exponential (LSE) function can be used to replace the hard maximum value. This achieves smooth suppression of outliers. The rounding loss used to represent the vector range... Differentiable dynamic range can be obtained by performing a differentiable transformation using equation (3). .
[0126] Equation (3)
[0127] In equation (3), Temperature parameter, used to control the smoothness, where The lower, the closer Norm; The higher, the closer Norm (hard maximum); For vector dimensions; This is the second activation value. It can provide dense gradient signals while penalizing all large-value elements (outliers), thus preventing a single outlier from dominating the optimization.
[0128] For the differentiable transformation of the underflow loss, the Sigmoid function can be used instead of the discrete indicator function. ) ( (where the dead zone threshold is used), a differentiable soft underflow count is constructed using equation (4). .
[0129] Equation (4)
[0130] In equation (4), The dynamic dead zone threshold (which varies with the distribution of activation values). The temperature parameter controls the transition sharpness of the Sigmoid function. This is the second activation value. The dimension is vector. This function provides a differentiable gradient, encouraging elements to move away from the dead zone boundary and reducing underflow counts.
[0131] Using equations (3) and (4) above, we can obtain the differentiable rounding loss and the differentiable underflow loss, and thus the differentiable expected loss corresponding to the upper bound of the expected mean square loss. .
[0132] Specifically, the differentiable rounding loss and differentiable underflow loss mentioned above can be substituted into equation (2) to obtain the differentiable expected loss. .
[0133] Equation (5)
[0134] In equation (5), The second activation value after rotation transformation by the initial rotation matrix (standard rotation matrix) .
[0135] In step 232, the optimal rotation matrix is determined based on the differentiable expected loss.
[0136] For the entire model, the global CURE optimization objective is that the expected loss of all activated quantized groups satisfies equation (6).
[0137] Equation (6)
[0138] In equation (6), To activate the set of quantized groups, For the first The activation groups are rotated by the matrix The transformed vector, Orthogonal constraints must be satisfied ( The Riemann Adam optimizer is used to optimize on the Stiefel manifold, resulting in... To minimize, during the optimization process, make By moving closer to the minimum value, the orthogonality of the rotation matrix is ensured, and an optimized rotation matrix is obtained.
[0139] Those skilled in the art will understand that in this embodiment, the rounding loss and underflow loss are made differentiable in order to provide a differentiable gradient and optimize the expected loss of all activation quantization groups (all activation values). The differentiability provided by equations (3) and (4) above is only a specific implementation method, but does not limit the implementation process of differentiability. Other methods for differentiability that those skilled in the art can conceive of are within the scope of protection of this disclosure.
[0140] See Figure 6 The diagram illustrates the activation value reshaping process in this embodiment of the invention. The left-hand area of the diagram illustrates the coupling contradiction between rounding loss and underflow loss: reducing... (Suppressing outliers) reduces the quantization step size. ,lead to Increased underflow error. The middle area of the figure illustrates the differentiable optimization process, which balances outlier suppression and underflow loss through differentiable expected loss. The right area of the figure illustrates the activation value distribution after reshaping the activation values using an optimized rotation matrix. This activation value reshaping process demonstrates that the differentiable expected loss constructed in this disclosure balances outlier suppression and underflow loss, achieving geometric reshaping of activation values. This helps resolve the coupling contradiction of activation rotation transformation and achieves global optimization of quantization error.
[0141] Based on the embodiments of this disclosure, by performing differentiable processing on rounding loss and underflow loss respectively, it helps to avoid the optimization dominated by a single outlier in rounding loss, achieve smooth suppression of outliers, and reduce the dead zone underflow count of underflow loss, thereby achieving a balance between rounding loss and underflow loss and resolving the coupling contradiction between rounding loss and underflow loss.
[0142] Figure 4 This is a flowchart illustrating step 204 of the model quantization method disclosed herein. Figure 4 As shown above, in the above Figure 2 Based on the illustrated embodiment, the specific process of performing multi-stage quantization on the model may include the following steps 241-243. Each step is explained below.
[0143] In step 241, based on the optimized rotation matrix, the first activation value of each linear layer is geometrically reshaped to obtain the third activation value of each linear layer.
[0144] In a high-dimensional vector space, each activation vector corresponds to a geometric direction and length. Optimizing the rotation matrix, as an orthogonal transformation, can redistribute the components of the vector in each dimension without changing its magnitude (i.e., information energy), thereby changing its spatial orientation and reshaping the geometric structure of the activation values.
[0145] In this embodiment, the optimized rotation matrix and each first activation value can be used to perform matrix operations to obtain the third activation value of each linear layer.
[0146] In step 242, the first Hessian matrix is determined based on the optimized rotation matrix.
[0147] In this embodiment, the first Hessian matrix is the second-order partial derivative matrix of the loss function with respect to the model weights, calculated based on the calibration dataset during the forward propagation of the large language model to be quantized. It is used to characterize the local curvature relationship and mutual influence between the weight parameters. During model quantization, the first Hessian matrix is used to identify weights that are more sensitive to the model output (i.e., "important weights"), in order to guide the optimal allocation of quantization errors, prioritizing the protection of weights more sensitive to the model output during quantization.
[0148] In some implementations, the large language model to be quantized can be transformed based on the optimized rotation matrix to obtain the transformed large language model. Then, the calibration dataset is input into the transformed large language model for forward computation to obtain the first Hessian matrix.
[0149] In this embodiment, by applying the optimized rotation matrix to the large language model to be quantized and multiplying the optimized rotation matrix with the weights of the large language model to be quantized, the transformed large language model can be obtained. Then, the calibration dataset is input into the transformed large language model for forward computation. During the forward computation process, the autocorrelation matrix can be collected. Cross-correlation matrix And the first Hessian matrix (output Hessian matrix).
[0150] The autocorrelation matrix represents the distortion input. The autocorrelation matrix is given. The distorted input is the activation value (which can be called the fourth activation value) after being rotated according to the optimized rotation matrix (the first activation value) and then quantized by the quantizer. The cross-correlation matrix is the original input. (Third activation value) and distorted input The cross-correlation matrix.
[0151] In other implementations, when performing step 201, which involves inputting the calibration dataset into the large language model to be quantized for forward computation, a second Hessian matrix can be collected, and then a first Hessian matrix can be determined based on the second Hessian matrix and the optimized rotation matrix.
[0152] In this embodiment, by collecting the first activation value and the second Hessian matrix of each linear layer during the forward computation of the calibration dataset into the large language model to be quantized, the first Hessian matrix can be obtained by directly transforming the second Hessian matrix with the optimized rotation matrix when the optimized rotation matrix is obtained, without having to run the calibration dataset again.
[0153] In step 243, the large language model to be quantized is quantized based on the third activation value and the first Hessian matrix.
[0154] In this embodiment, a symmetric gradient quantizer can be used to quantize the third activation value to obtain the fourth activation value. Then, the DCR mechanism is used to solve the defect in the weight quantization of related technologies that assumes that the input activation value of each layer is a clean input. By adapting the cumulative distortion input corresponding to the fourth activation value, high-precision weight quantization is achieved.
[0155] In this embodiment, the specific implementation process of step 243 can be found in [reference needed]. Figure 5 The embodiments shown are not described in detail here.
[0156] Based on the embodiments of this disclosure, the coupling contradiction of "increased rounding loss and underflow loss" during the activation value rotation process is resolved by optimizing the rotation matrix, thereby achieving a balance between rounding loss and underflow loss and reducing global quantization error.
[0157] Figure 5 This is a flowchart illustrating step 243 of the model quantization method disclosed herein. Figure 5 As shown above, in the above Figure 4 Based on the illustrated embodiment, the quantization process of the large language model to be quantized, based on the third activation value and the first Hessian matrix, includes steps 2431-2435. Each step is explained below.
[0158] In step 2431, the third activation value is quantized to obtain the fourth activation value.
[0159] In this embodiment, the third activation value can be adjusted using a symmetric gradient quantizer. Quantization was performed to obtain the fourth activation value. .
[0160] In step 2432, the weight optimization objective is determined based on the third activation value, the fourth activation value, the first Hessian matrix, and the original weights of the large language model to be quantized.
[0161] For cumulative distortion scenarios, this embodiment constructs a weight optimization objective based on output sensitivity, which is to minimize the Hessian weighted error between the quantized output and the original output.
[0162] For example, Equation (7) can be used as the weight optimization objective.
[0163] Equation (7)
[0164] In equation (7), The first Hessian matrix in the output space is used to characterize the sensitivity of the model's task loss to the output. The larger the Hessian value, the more sensitive it is to disturbances. The weights are quantified. The quantized input activation value (fourth activation value); These are the model weights before quantization; The input activation value before quantization (third activation value); The Hessian weighted norm is defined as follows: .
[0165] The weight optimization objective illustrated by equation (7) is... Directing quantization noise to channels that are insensitive to task loss can reduce performance degradation.
[0166] In step 2433, the autocorrelation matrix of the fourth activation value and the cross-correlation matrix of the third and fourth activation values are determined.
[0167] Among them, the autocorrelation matrix of the fourth activation value is the distortion input. autocorrelation matrix The distorted input is the activation value (fourth activation value) after the input activation value (first activation value) has been rotated according to the optimized rotation matrix (third activation value), and then quantized by the quantizer. The cross-correlation matrix is the original input. (Third activation value) and distorted input cross-correlation matrix .
[0168] In step 2434, the weight offset center is determined based on the autocorrelation matrix, cross-correlation matrix, and the original model weights.
[0169] The weight offset center is used to compensate for the cumulative distortion error of the activation values of each linear layer, and can be determined by the autocorrelation matrix and the cross-correlation matrix.
[0170] For example, equation (8) can be used as the weight offset center.
[0171] Equation (8)
[0172] In equation (8), the weight offset center is the offset weight corrected by the correlation matrix (autocorrelation matrix and cross-correlation matrix), which is used to compensate for the influence of input distortion and enable the quantization weight to adapt to the distorted input. Suppressing cumulative error The cross-correlation matrix, This is the autocorrelation matrix.
[0173] After determining the weight offset center, the weight optimization objective of equation (7) can be rewritten as equation (9).
[0174] Equation (9)
[0175] In equation (9), The objective function takes into account both input distortion compensation and output sensitivity, where the trace is the matrix. The first Hessian matrix in the output space is used to characterize the sensitivity of the model's task loss to the output. The larger the Hessian value, the more sensitive it is to disturbances. These are the quantized weights.
[0176] In step 2435, the large language model to be quantized is quantized based on the weight offset center, the first Hessian matrix, and the weight optimization objective.
[0177] In this embodiment, the first Hessian matrix can be transformed into a lower triangular matrix; the transformed lower triangular matrix is normalized to obtain a strict lower triangular matrix; based on the strict lower triangular matrix, the weights of the large language model to be quantized are iteratively quantized; in response to the satisfaction of the iteration termination condition, the model quantization weights of the large language model to be quantized are obtained.
[0178] In order to project the weights, after being compensated by the weight offset center, onto the quantization grid, this embodiment uses a whitening bilateral feedback iterative rounding algorithm. Through the whitening transformation of the Hessian matrix and the input autocorrelation matrix, bidirectional correction of the cumulative error of the distorted input is achieved.
[0179] First, the Hessian matrix of the output space can be... and input autocorrelation matrix Perform Cholesky decomposition to obtain .
[0180] in, , It is a lower triangular matrix; to ensure stability, the decomposition results are normalized. , . , It is a diagonal variance matrix. , As a strictly lower triangular matrix, the diagonal scale is absorbed into the quantization step size.
[0181] Then, the quantization residual of the t-th iteration is defined as... . (The quantization weights for the t-th iteration) are obtained through bilateral feedback ( , Correct the quantization residual.
[0182] Based on the quantization weights quantized in the t-th iteration , to proceed with the first The quantization weights can be obtained through the next iteration of quantization. As shown in equation (10).
[0183] Equation (10)
[0184] In equation (10), To correct output sensitivity, the quantization residual is directed to a channel with low output sensitivity, reducing the impact of quantization noise on performance. To correct for input correlation, the quantization residuals are made to avoid the dominant feature direction of input distortion, thus suppressing error accumulation; For bilateral coupling correction, the coupling relationship between input and output is handled to improve correction accuracy in low-bit quantization scenarios; This is the center of the weight offset.
[0185] Through the above iterative quantization process, the model weights can be quantized. Once the iteration termination condition is met, the quantized weights of the large language model to be quantized can be obtained. The iteration termination condition can be that the number of iterations has reached a set number, such as 10-20, or the residual... Less than the set residual threshold.
[0186] Based on embodiments of this disclosure, the weight offset center in the cumulative distortion scenario is derived. By incorporating output sensitivity information into the weight quantization process, quantization noise is avoided in sensitive channels, further improving the performance of the quantization model.
[0187] The model quantization method provided in this disclosure is applicable to quantizing large language models deployed on electronic devices, especially for quantizing large language models on edge devices. Edge devices include various computing devices deployed on edge nodes, such as smartphones, smart home devices, and cameras. Edge devices are located closer to the data source and user, and have higher real-time requirements; quantization enables the model to meet these requirements.
[0188] The embodiments disclosed herein can improve inference performance on edge testing devices while enhancing the generalization ability of large language models.
[0189] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.
[0190] Corresponding to the aforementioned embodiments of the model quantization method, this disclosure also provides embodiments of the model quantization apparatus.
[0191] Exemplary device
[0192] Figure 7This is a schematic diagram of one embodiment of the model quantization device disclosed herein, which is applied to electronic devices (such as computer equipment, servers, etc.). Figure 7 As shown, the model quantization device may include:
[0193] The forward computation module 71 is used to input the calibration dataset into the large language model to be quantized and perform forward computation to obtain the first activation value of each linear layer.
[0194] The first determining module 72 is used to determine the expected mean square loss of the symmetric gradient quantization process based on the symmetric gradient quantizer and the first activation value. The expected mean square loss includes the rounding loss and underflow loss of the activation value.
[0195] The second determination module 73 is used to determine the optimized rotation matrix based on rounding loss and underflow loss;
[0196] The quantization module 74 is used to perform multi-stage quantization processing on the large language model to be quantized based on the first activation value and the optimized rotation matrix, so as to obtain the quantized target model, which can be deployed to the hardware device. The quantized target model can be called by the hardware device to process the corresponding task.
[0197] Figure 8 This is a schematic diagram of another embodiment of the model quantization device disclosed herein. Figure 8 As shown, in Figure 7 Based on the illustrated embodiment, in some optional implementations, the first determining module 72 may include:
[0198] The rotation transformation submodule 721 is used to perform initial rotation processing on the first activation value of each linear layer using a standard rotation matrix to obtain the second activation value of each linear layer.
[0199] The first determination submodule 722 is used to determine the expected mean square loss based on the symmetric middle ladder quantizer and the second activation value.
[0200] In some alternative implementations, the second determining module 73 may include:
[0201] The first processing submodule 731 is used to perform differentiable processing on the rounding loss and underflow loss respectively to obtain the differentiable expected loss;
[0202] The second determination submodule 732 is used to determine the optimal rotation matrix based on the differentiable expected loss.
[0203] In some alternative implementations, the quantization module 74 may include:
[0204] The second processing submodule 741 is used to perform geometric reshaping processing on the first activation value of each linear layer based on the optimized rotation matrix to obtain the third activation value of each linear layer.
[0205] The third determining submodule 742 is used to determine the first Hessian matrix based on the optimized rotation matrix;
[0206] The quantization submodule 743 is used to quantize a large language model to be quantized based on the third activation value and the first Hessian matrix.
[0207] In some optional implementations, the third determining submodule 742 is used to transform the large language model to be quantized based on the optimized rotation matrix to obtain the transformed large language model; and input the calibration dataset into the transformed large language model for forward computation to obtain the first Hessian matrix.
[0208] In some alternative implementations, the model quantization apparatus may further include:
[0209] Module 75 is used to obtain the second Hessian matrix when the calibration dataset is input into the large language model to be quantized for forward computation.
[0210] The third determining submodule 742 is used to determine the first Hessian matrix based on the second Hessian matrix and the optimized rotation matrix.
[0211] In some optional implementations, the quantization submodule 743 is used to quantize the third activation value to obtain the fourth activation value; determine the weight optimization objective based on the third activation value, the fourth activation value, the first Hessian matrix, and the original weights of the large language model to be quantized; determine the autocorrelation matrix of the fourth activation value, the cross-correlation matrix of the third activation value, and the cross-correlation matrix of the fourth activation value; determine the weight offset center based on the autocorrelation matrix, the cross-correlation matrix, and the original weights of the model, the weight offset center being used to compensate for the cumulative distortion error of the activation values of each linear layer; and quantize the large language model to be quantized based on the weight offset center, the first Hessian matrix, and the weight optimization objective.
[0212] In some optional implementations, the quantization submodule 743 is used to transform the first Hessian matrix into a lower triangular matrix; normalize the transformed lower triangular matrix to obtain a strict lower triangular matrix; perform iterative quantization on the weights of the large language model to be quantized based on the strict lower triangular matrix and the weight optimization objective; and obtain the model quantization weights of the large language model to be quantized in response to the fulfillment of the iteration termination condition.
[0213] The modules and units in this disclosed device can be further divided into finer-grained units according to actual needs, and the specific configuration can be set according to actual needs.
[0214] The apparatus of this disclosure embodiment can be used to implement the methods of the above embodiments of this disclosure. The two correspond to each other in specific implementation, and the specific implementation of related parts can be referred to each other, which will not be repeated here.
[0215] Exemplary electronic devices, computer program products, and computer-readable storage media
[0216] This disclosure also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program stored in the memory, wherein when the computer program is executed, it implements the model quantization method of any of the above embodiments of this disclosure.
[0217] Below, for reference Figure 9 This describes an electronic device according to embodiments of the present disclosure, wherein apparatus for implementing methods according to embodiments of the present disclosure may be integrated. Figure 9 This is a structural diagram of an electronic device provided in an illustrative embodiment of the present disclosure, such as... Figure 9 As shown, the electronic device includes one or more processors 91, one or more memory 92s of computer-readable storage media, and a computer program stored in the memory and executable on the processor. When the program in the memory 92 is executed, the aforementioned model quantization method can be implemented.
[0218] Specifically, in practical applications, the electronic device may also include components such as an input device 93 and an output device 94, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown). Those skilled in the art will understand that... Figure 9 The structure of the electronic device shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or certain components, or different component arrangements. Wherein:
[0219] The processor 91 may be a central processing unit (CPU) or other processing unit with model quantization capability and / or instruction execution capability. It performs various functions and processes data by running or executing software programs and / or modules stored in memory 92 and calling data stored in memory 92, thereby performing overall monitoring of the electronic device.
[0220] Memory 92 can store one or more computer program products. The memory can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products can be stored on the computer-readable storage medium, and processor 91 can run the computer program products to implement the model quantization methods of the various embodiments of this disclosure above and / or other desired functions.
[0221] The input device 93 can be used to receive input numeric or character information. The input device 93 may include a keyboard, mouse, joystick, etc., related to user settings and function control.
[0222] The output device 94 can output various information to the outside, including determined distance information, direction information, etc. The output device 94 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0223] Electronic devices may also include a power supply for powering various components, which can be logically connected to the processor 91 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply may also include one or more DC or AC power sources, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and any other components.
[0224] Of course, for the sake of simplicity, Figure 9 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0225] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the model quantization methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0226] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0227] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the model quantization methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0228] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0229] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in these embodiments are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0230] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0231] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.
[0232] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this disclosure are not limited to the order specifically described above, unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the method according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the method according to this disclosure.
[0233] The description in this disclosure is provided for illustrative and descriptive purposes only and is not intended to be exhaustive or to limit the disclosure to its forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of this disclosure and to enable those skilled in the art to understand this disclosure and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A model quantization method, characterized in that, include: The calibration dataset is input into the large language model to be quantized for forward computation to obtain the first activation value of each linear layer. Based on the symmetric ladder quantizer and the first activation value, the expected mean square loss of the symmetric ladder quantization process is determined. The expected mean square loss includes the rounding loss and underflow loss of the activation value. Based on the rounding loss and underflow loss, determine the optimal rotation matrix; Based on the first activation value and the optimized rotation matrix, the large language model to be quantized is subjected to multi-stage quantization processing to obtain the quantized target model, which is then deployed to a hardware device. The quantized target model can be called by the hardware device to process corresponding tasks.
2. The method according to claim 1, characterized in that, The step of determining the expected mean square loss of symmetric gradient quantization based on the symmetric gradient quantizer and the first activation value includes: Using a standard rotation matrix, the first activation values of each linear layer are initially rotated to obtain the second activation values of each linear layer; The expected mean square loss is determined based on the symmetric gradient quantizer and the second activation value.
3. The method according to any one of claims 1-2, characterized in that, The determination of the optimized rotation matrix based on the rounding loss and underflow loss of the activation value includes: Differentiable processing is performed on the rounding loss and the underflow loss respectively to obtain the differentiable expected loss; Based on the differentiable expected loss, the optimal rotation matrix is determined.
4. The method according to any one of claims 1-2, characterized in that, The multi-stage quantization process for the large language model to be quantized, based on the first activation value and the optimized rotation matrix, includes: Based on the optimized rotation matrix, the first activation value of each linear layer is geometrically reshaped to obtain the third activation value of each linear layer. Based on the optimized rotation matrix, the first Hessian matrix is determined; The large language model to be quantized is quantized based on the third activation value and the first Hessian matrix.
5. The method according to claim 4, characterized in that, Determining the first Hessian matrix based on the optimized rotation matrix includes: Based on the optimized rotation matrix, the large language model to be quantized is transformed to obtain the transformed large language model. The calibration dataset is input into the transformed large language model for forward computation to obtain the first Hessian matrix.
6. The method according to claim 4, characterized in that, Also includes: When the calibration dataset is input into the large language model to be quantized for forward computation, the second Hessian matrix is obtained. Determining the first Hessian matrix based on the optimized rotation matrix includes: The first Hessian matrix is determined based on the second Hessian matrix and the optimized rotation matrix.
7. The method according to claim 4, characterized in that, The quantization of the large language model to be quantized based on the third activation value and the first Hessian matrix includes: The third activation value is quantized to obtain the fourth activation value; Based on the third activation value, the fourth activation value, the first Hessian matrix, and the original weights of the large language model to be quantized, the weight optimization objective is determined. Determine the autocorrelation matrix of the fourth activation value, and the cross-correlation matrix of the third and fourth activation values; Based on the autocorrelation matrix, the cross-correlation matrix, and the original model weights, a weight offset center is determined. The weight offset center is used to compensate for the cumulative distortion error of the activation values of each linear layer. The large language model to be quantized is quantized based on the weight offset center, the first Hessian matrix, and the weight optimization objective.
8. The method according to claim 7, characterized in that, The quantization of the large language model to be quantized based on the weight offset center, the first Hessian matrix, and the weight optimization objective includes: Transform the first Hessian matrix into a lower triangular matrix; The transformed lower triangular matrix is normalized to obtain a strictly lower triangular matrix. Based on the strict lower triangular matrix and the weight optimization objective, the weights of the large language model to be quantized are iteratively quantized. In response to the fulfillment of the iteration termination condition, the model quantization weights of the large language model to be quantized are obtained.
9. A model quantization device, comprising: The forward computation module is used to input the calibration dataset into the large language model to be quantized and perform forward computation to obtain the first activation value of each linear layer. The first determining module is used to determine the expected mean square loss of the symmetric gradient quantization process based on the symmetric gradient quantizer and the first activation value. The expected mean square loss includes the rounding loss and underflow loss of the activation value. The second determining module is used to determine the optimized rotation matrix based on the rounding loss and underflow loss; The quantization module is used to perform multi-stage quantization processing on the large language model to be quantized based on the first activation value and the optimized rotation matrix to obtain the quantized target model, so as to deploy the quantized target model to the hardware device, and the quantized target model can be called by the hardware device to process corresponding tasks.
10. An electronic device, characterized in that, include: Memory, used to store computer program products; A processor for executing a computer program product stored in the memory, wherein when the computer program product is executed, it implements the method described in any one of claims 1-8.
11. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1-8.
12. A computer program product comprising computer program instructions, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1-8.