First-order error compensation quantized large-model low-resource-overhead text understanding generation method
By employing first-order error-compensated quantization and Choleski decomposition techniques, the memory and computational requirements for deploying large language models in resource-constrained environments are addressed, improving the model's accuracy and stability. This approach is suitable for low-resource-overhead deployment on edge devices.
Patent Information
- Application Number
- CN202510966948.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
AI Technical Summary
When large language models are deployed in resource-constrained environments, existing quantization techniques are insufficient to effectively reduce memory and computational requirements, and ignoring first-order gradients leads to a decrease in accuracy, affecting inference accuracy and stability.
A first-order error-compensated quantization method is adopted. By calculating the gradient of the difference between the latent weight and the full-precision weight, and combining the Cholesky decomposition technique to recover the Hessian submatrix, the quantization loss function is optimized to achieve text understanding generation with low resource overhead.
It significantly reduces model storage and computing resource consumption, improves inference accuracy and stability, is suitable for edge device deployment, is compatible with existing system architectures, and enhances the deployment flexibility and scalability of large models.
Smart Images

Figure CN120804291A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of text generation, and particularly relates to a first-order error compensation quantization large model low resource consumption text understanding generation method. BACKGROUND
[0002] Large language models (LLMs) such as Llama have shown excellent performance and wide applicability in text understanding generation tasks. With the expansion of model size and the increase of pre-training data size, its ability is constantly improving. However, a large number of parameters and higher computing requirements bring huge memory and processing burden. These requirements have caused significant limitations to actual deployment in resource-constrained environments.
[0003] Quantization is a classic model compression technique. It reduces memory usage and speeds up computation by converting high-precision floating-point parameters and activation values to low-precision fixed-point formats, while not changing the model architecture. Among various quantization methods, post-training quantization (PTQ) is known for its efficiency. It does not require gradient-based fine-tuning and can maintain nearly lossless performance at higher bit widths. Compared with quantization-aware training (QAT) which requires additional training, PTQ is usually more practical for large language models. Quantization not only reduces memory consumption by mapping full-precision weights to low-bit fixed-point formats (such as int8 or int4), but also dynamically quantizes activation values to low-bit representations. This helps to achieve efficient operations, including low-bit matrix multiplication. To alleviate the precision loss caused by the conversion from full-precision to low-bit format, reconstruction-based methods such as AdaRound, BRECQ, and QDrop have been developed. These techniques have shown strong performance on architectures such as ResNet, which can measure and minimize quantization errors within layers or blocks. However, due to the huge computational cost generated in the calibration process, these methods are difficult to directly apply to large-scale language models that usually contain billions of parameters.
[0004] Such methods usually assume that the well-trained model is close to the local optimal solution, thus proving the rationality of omitting the first-order term in the loss approximation. However, we observe that after quantizing the preceding weight column, the remaining unquantized weights will deviate significantly from their original full-precision values due to the continuous compensation of weight gaps. This cumulative deviation will very likely cause the gradient to be difficult to ignore, and in turn, the first-order term of the Taylor expansion will also have a significant impact on the construction of the quantized loss function. SUMMARY
[0005] To address the aforementioned shortcomings of the existing technology, the present invention provides a low-resource-cost method for large-model text comprehension and generation using first-order error compensation quantization. By incorporating first-order gradient terms into quantization error compensation and approximating the gradient by directly calculating the difference between latent weights and full-precision weights, this method avoids the high cost and limited generalization capability of gradient calculations based on backpropagation. It also leverages pre-computed Cholesky factors to efficiently recover the inverse of the Hessian matrix in real time. By minimizing the error in the compression process, this method enables accurate compression of large models and their application in low-resource-cost text comprehension and generation tasks.
[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: A large-model low-resource overhead text understanding generation method with first-order error compensation quantization includes the following steps: Obtain a pre-trained large language text generation model that performs a standard autoregressive text generation task and construct a calibration text dataset; Using the calibration text dataset as the model input, the weights of the neural network layer are quantized with first-order error compensation in the column quantization order when performing the standard autoregressive text generation task on the pre-trained large language text generation model; The quantized large language text generation model is used for standard autoregressive text generation.
[0007] Furthermore, using the calibration text dataset as the model input, the weights of the neural network layer are quantized according to the column quantization sequence for first-order error compensation when performing the standard autoregressive text generation task on the pre-trained large language text generation model, including: The text input matrix of the calibration text dataset is fed into a pre-trained large language text generation model to perform a standard autoregressive text generation task. Based on the text output error generated by the weight quantization of the neural network layer, an error approximation model including a first-order gradient term is established. According to the error approximation model containing the first-order gradient term, a loss optimization function is established to minimize the weight quantization loss by adjusting the latent weights of the unquantized weight columns.
[0008] Furthermore, the error approximation model including the first-order gradient term is established as:
[0009] Among them, δ E is the text output error, g is the gradient of the text output error with respect to the weight, δw is the weight quantization compensation, T is the transposed symbol, and H is the Hessian matrix of the text output error with respect to the weight.
[0010] Furthermore, the loss optimization function that minimizes the weight quantization loss is established as:
[0011] wherein min is a min function, e q is a unit vector, is the qth column element of the inverse of the Hessian matrix, q is the quantized value of the qth column weight.
[0012] Further, the Lagrange multiplier method is used to solve the loss optimization function to obtain the optimal weight quantization compensation.
[0013] Further, the optimal weight quantization compensation is:
[0014] wherein, is the qth row and qth column element of the inverse of the Hessian matrix, is the qth column element of the inverse of the Hessian matrix, is the inverse of the Hessian matrix.
[0015] Further, the Cholesky decomposition method is used to simplify the weight quantization compensation to:
[0016] wherein, is the qth row and qth column element of the upper triangular matrix of the Cholesky decomposition, is the qth column element of the upper triangular matrix of the Cholesky decomposition.
[0017] Further, the error between the full-precision weight before quantization and the quantization compensation weight is approximated as the gradient of the text output error with respect to the weight, specifically:
[0018] wherein, β is a scaling factor for mapping the weight space to the gradient space, W is the full-precision weight before quantization, is the quantization compensation weight.
[0019] Further, the text input matrix is used to reconstruct the upper triangular matrix of the Cholesky decomposition, and further recover the inverse Hessian submatrix, specifically:
[0020] wherein, is the inverse Hessian submatrix, is the matrix after removing the qth column in the text input matrix, is the submatrix of the upper triangular matrix of the Cholesky decomposition corresponding to the unquantized weight.
[0021] Further, the final optimized weight quantization compensation amount is: .
[0022] The present application has the following beneficial effects: (1) The present application proposes a precise compensation method for large language model training after quantization (PTQ) - FOEM, which is designed for resource-constrained inference deployment scenarios. This method significantly reduces model storage and runtime computing resource consumption, making it possible to deploy large models in edge devices or low-bandwidth environments, with good network resource saving and energy efficiency optimization effect.
[0023] (2) The present application innovatively introduces a first-order error compensation mechanism, breaking through the precision bottleneck caused by ignoring the first-order gradient in existing quantization technologies (such as GPTQ). By utilizing the deviation between latent weights and full-precision weights to approximate the gradient, it does not rely on high-cost backpropagation, thereby significantly improving inference accuracy and stability while ensuring model compression rate, especially suitable for scenarios that require data local inference to improve data security.
[0024] (3) To improve the compensation efficiency, the present application uses Cholesky decomposition technology to construct local Hessian sub-matrix, avoiding the repeated iteration overhead of traditional methods of Hessian inverse matrix, making the model compression process more efficient, reducing the deployment cost, saving the computing and storage resources, especially suitable for large-scale model fast low-power inference.
[0025] (4) In terms of compatibility, the present application method can seamlessly integrate existing mainstream PTQ technologies (such as GPTAQ, SpinQuant), further improving performance without changing the existing system architecture, and narrowing the accuracy gap with full-precision models, bringing greater flexibility and scalability to quantization deployment systems. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 Flowchart of large model low resource consumption text understanding generation method for first-order error compensation quantization; Figure 2 Flowchart for first-order gradient term calculation. DETAILED DESCRIPTION
[0027] The specific embodiments of the present application are described below to facilitate understanding of the present application by those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, it is obvious that various changes are within the spirit and scope of the present application defined and determined by the appended claims, and all inventions utilizing the concept of the present application are within the scope of protection.
[0028] AsFigure 1 As shown, the first-order error compensation quantization large model low resource consumption text understanding generation method provided by the embodiment of the application comprises the following steps S1 to S3: S1, obtaining a pre-trained large language text generation model for performing a standard autoregressive text generation task, and constructing a calibration text dataset; In an optional embodiment of the application, step S1 downloads a pre-trained large language model from a public channel, and constructs a calibration text dataset from multiple text data sources, which is used for error evaluation in the subsequent quantization process.
[0029] S2, performing first-order error compensation quantization on the weights of the neural network layer of the pre-trained large language text generation model when performing a standard autoregressive text generation task on the pre-trained large language text generation model with the calibration text dataset as the model input according to the column quantization order; In an optional embodiment of the application, step S2 iteratively performs weight quantization and error compensation of each column according to the column quantization order. The calculation is based on Taylor expansion as the theory, and the compensation accuracy is improved by introducing a first-order gradient term.
[0030] In this embodiment, first-order error compensation quantization is performed on the weights of the neural network layer of the pre-trained large language text generation model when performing a standard autoregressive text generation task on the pre-trained large language text generation model with the calibration text dataset as the model input, comprising: The text input matrix of the calibration text dataset is input into the pre-trained large language text generation model to perform a standard autoregressive text generation task, and an error approximation model containing a first-order gradient term is established according to the text output error generated by weight quantization of the neural network layer. Assuming that the original weight of a certain layer is W, the input is X, and it is assumed that the pruned weight is W = W+δw, then the output error δE generated thereby can be approximately obtained by Taylor series expansion:
[0031] In the loss construction, the first-order term is explicitly reserved, and the loss difference before and after quantization is:
[0032] wherein, δ E is the text output error, g is the gradient of the text output error with respect to the weight, δw is the weight quantization compensation amount, T is a transpose symbol, and H is the Hessian matrix of the text output error with respect to the weight.
[0033] According to the error approximation model containing a first-order gradient term, a loss optimization function minimizing the weight quantization loss is established by adjusting the potential weights of the unquantized weight column. When quantizing the qth column weight, the optimization target becomes to minimize the quantization-induced loss by adjusting the potential weights δw of the remaining unquantized columns:
[0034] wherein min is a minimum function, e q is a unit vector, and the value at the qth position is 1; is the qth q quantized value of the column weight.
[0035] To solve this constrained optimization problem, the present embodiment adopts the Lagrange multiplier method to solve the loss optimization function, which is expressed as:
[0036] The optimal weight quantization compensation quantity is obtained as:
[0037] wherein, is the element of the qth row and the qth column in the inverse matrix of the Hessian matrix, is the qth column element of the inverse matrix of the Hessian matrix, is the inverse matrix of the Hessian matrix.
[0038] The present embodiment adopts the Cholesky decomposition method to simplify the weight quantization compensation quantity as:
[0039] wherein, is the element of the qth row and the qth column in the upper triangular matrix of the Cholesky decomposition, is the qth column element of the upper triangular matrix of the Cholesky decomposition.
[0040] As Figure 2 shown, the present embodiment adopts the gradient of the error approximation text output error with respect to the weight between the full-precision weight before quantization and the quantization compensation weight, which is specifically:
[0041] wherein, β is a scaling factor for mapping the weight space to the gradient space, W is the full-precision weight before quantization, is the quantization compensation weight.
[0042] The present embodiment adopts the text input matrix to reconstruct the upper triangular matrix of the Cholesky decomposition, and further restores the inverse Hessian submatrix, which is specifically:
[0043] wherein, is the inverse Hessian submatrix, is the matrix after removing the qth column from the text input matrix, is the upper triangular matrix of the Cholesky decomposition corresponding to the submatrix of the unquantized weights.
[0044] The final optimized weight quantization compensation amount in the embodiment is: .
[0045] S3, using the large language text generation model after quantization to perform standard autoregressive text generation.
[0046] In an optional embodiment of the present application, step S3 uses the large language text generation model after quantization to perform standard autoregressive text generation, and the prediction target is defined as:
[0047] To verify the performance of the method proposed in the present application in large language model quantization, the present application carried out systematic experiments on multiple models such as Llama2 and Llama3. The results show that in the scenario of only 3-bit quantization of weights, the present application can reduce the perplexity of the Llama3-8B model by 89.6%, and improve the 5-shot MMLU accuracy of the Llama3-70B model from 51.7% to 74.9%, close to the full-precision model of 78.6%. At the same time, the present application is compatible with existing advanced methods such as GPTAQ and SpinQuant in the W4A4KV4 setting, and outperforms GPTQ and GPTAQ in multiple evaluation indicators, especially in large models. In addition, in terms of efficiency, although the present application introduces a first-order term, theoretically increasing the compensation calculation amount, experiments prove that the quantization time of the present application on Llama3-8B only increases by 0.6% compared with GPTQ, achieving a good balance between precision and efficiency.
[0048] In summary, the present application first explicitly introduces a first-order gradient term in post-training quantization, and modifies the traditional second-order compensation assumption. By using the difference between full-precision and compensation weights to approximate the gradient, and combining the Cholesky submatrix recovery technology, the quantization precision and model robustness are effectively improved without increasing the additional calculation overhead, providing a more efficient and accurate solution for the practical deployment of large language models.
[0049] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0050] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes A step that specifies a function in one or more boxes.
[0052] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
[0053] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A large-model low-resource-overhead text understanding and generation method with first-order error compensation quantization, characterized by: The following steps are involved: Obtain a pre-trained large language text generation model that performs a standard autoregressive text generation task and construct a calibration text dataset; Using the calibration text dataset as the model input, the weights of the neural network layer are quantized with first-order error compensation in the column quantization order when performing the standard autoregressive text generation task on the pre-trained large language text generation model; The quantized large language text generation model is used for standard autoregressive text generation.
2. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 1 is characterized in that: Using the calibration text dataset as the model input, the weights of the neural network layer are quantized according to the column quantization sequence for first-order error compensation when performing the standard autoregressive text generation task on the pre-trained large language text generation model, including: The text input matrix of the calibration text dataset is fed into a pre-trained large language text generation model to perform a standard autoregressive text generation task. Based on the text output error generated by the weight quantization of the neural network layer, an error approximation model including a first-order gradient term is established. According to the error approximation model containing the first-order gradient term, a loss optimization function is established to minimize the weight quantization loss by adjusting the latent weights of the unquantized weight columns.
3. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 2 is characterized in that: The error approximation model including the first-order gradient term is established as: Among them, δ E is the text output error, g is the gradient of the text output error with respect to the weight, δw is the weight quantization compensation, T is the transposed symbol, and H is the Hessian matrix of the text output error with respect to the weight.
4. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 3 is characterized in that: The loss optimization function that minimizes the weight quantization loss is established as: Among them, min is the minimum value function, e q is a unit vector, For the q Quantized value of the column weight.
5. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 4 is characterized in that: The Lagrange multiplier method is used to solve the loss optimization function to obtain the optimal weight quantization compensation.
6. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 5 is characterized in that: The optimal weight quantization compensation is: in, is the element in the qth row and qth column of the inverse matrix of the Hessian matrix, is the qth column element of the inverse matrix of the Hessian matrix, is the inverse matrix of the Hessian matrix.
7. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 6 is characterized in that: The Cholesky decomposition method is used to simplify the weight quantization compensation as follows: in, is the element in the qth row and qth column of the upper triangular matrix of the Cholesky decomposition, is the qth column element in the upper triangular matrix of the Cholesky decomposition.
8. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 7 is characterized in that: The error between the full-precision weight before quantization and the quantization compensation weight is used to approximate the gradient of the text output error with respect to the weight, specifically: in, β is the scaling factor that maps the weight space to the gradient space, W is the full precision weight before quantization, is the quantified compensation weight.
9. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 8 is characterized in that: The text input matrix is used to reconstruct the upper triangular matrix of the Cholesky decomposition, and the inverse Hessian matrix is further recovered, specifically: in, is the inverse Hessian matrix, is the matrix after removing the qth column from the text input matrix, The upper triangular matrix of the Cholesky decomposition corresponds to the submatrix of the unquantized weights.
10. The large-model low-resource-overhead text understanding generation method with first-order error compensation quantization according to claim 9 is characterized in that: The final optimized weight quantization compensation is: 。