Post-training quantization method and system for large language model inference deployment
Patent Information
- Application Number
- CN202610913232.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-18
AI Technical Summary
然而,当量化比特宽度降低至接近一比特的极低水平时,现有训练后量化方法面临若干尚未被妥善解决的技术问题:在极低比特离散投影过程中所产生的残差对模型前向推理输出的影响缺乏方向性区分与精准控制能力;部分方法虽然引入了额外的稠密矩阵或全精度影子权重以补偿量化误差,但这些额外参数在推理阶段仍被保留,导致存储与计算开销难以有效降低;此外,现有方法在涉及不同类型线性投影层以及不同模型规模的工程实施中,面对层间谱分布差异等数值病态工况时,缺乏系统性的稳健退化兼容机制
[0025]The specification of this application contains numerous technical features distributed across various technical solutions. Listing all possible combinations of these technical features (i.e., technical solutions) would make the specification excessively lengthy. To avoid this problem, the various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which are considered to have been described in this specification), unless such a combination of technical features is technically infeasible. For example, one example discloses feature A+B+C, and another example discloses feature A+B+D+E. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; they cannot be used simultaneously. Feature E can technically be combined with feature C. Therefore, the solution A+B+C+D should not be considered as described because it is technically infeasible, while the solution A+B+C+E should be considered as described.
Smart Images

Figure CN122779151A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural network model compression and inference deployment technology, and in particular to a post-training quantization technique for inference deployment of large language models. Background Technology
[0002] With the widespread application of large-scale pre-trained language models in tasks such as natural language understanding, text generation, and multimodal reasoning, the resource bottlenecks faced by these models in actual inference deployments are becoming increasingly prominent. Taking a typical cloud-based high-concurrency inference service scenario as an example, a Transformer architecture language model containing a large number of parameters needs to provide real-time text generation services to a large number of concurrent users on GPU clusters in data centers. The memory usage of its weight matrix and the memory access bandwidth requirements of layer-by-layer matrix multiplication operations constitute key bottlenecks restricting service throughput and deployment costs. In edge device deployment scenarios, when language models are used for local inference in mobile terminal smart assistants or IoT edge nodes, the storage capacity and computing power of the devices are extremely limited, making it difficult to directly load and run full-precision floating-point models. Furthermore, in resource-constrained calibration environments, small and medium-sized teams have limited available GPU memory and computing budgets when quantizing and compressing large language models, requiring the quantization calibration process itself to remain lightweight.
[0003] Post-training quantization, as a weight compression scheme that eliminates the need to retrain the entire model, has garnered significant attention in the aforementioned scenarios. However, when the quantization bit width is reduced to an extremely low level approaching one bit, existing post-training quantization methods face several unresolved technical issues: the residuals generated during extremely low-bit discrete projection lack directional differentiation and precise control over their impact on the model's forward inference output; while some methods introduce additional dense matrices or full-precision shadow weights to compensate for quantization errors, these additional parameters are still retained during the inference phase, making it difficult to effectively reduce storage and computational overhead; furthermore, existing methods lack a systematic robust degradation compatibility mechanism when dealing with numerically ill-conditioned conditions such as differences in inter-layer spectral distributions in engineering implementations involving different types of linear projection layers and different model sizes. Therefore, a post-training quantization method is urgently needed that can suppress directional propagation errors under extremely low-bit conditions and does not introduce additional dense parameters or additional dense matrix multiplication operations during the inference phase. Summary of the Invention
[0004] The purpose of this application is to provide a post-training quantization method and system for deploying large language model inference, so as to solve the problems mentioned in the background art.
[0005] This application provides a post-training quantization method for deploying large language model inference, targeting a neural network model containing a linear projection layer to be quantized. The method sequentially performs the following steps: Tensors with physical meaning obtained by forward and backward propagation of calibration data through the neural network model are used as input activation statistics and output gradient statistics. Based on the input activation statistics and the output gradient statistics, a proxy weight matrix containing second-order sensitivity information is constructed. Perform a bi-binary structure decomposition on the surrogate weight matrix, and represent the surrogate weight matrix as a product of two binary sign matrices and several consecutive scaling vectors; When performing the bibinary structural decomposition, the solution is obtained by alternating optimization. When fixing one side factor and updating the other side factor, the continuous surrogate variable of the other side factor is obtained, and the continuous surrogate variable is discretized projected onto the bibinary structural domain to obtain the discrete projection solution. Furthermore, in the process of processing the discrete projection residuals generated by the discrete projection of the continuous surrogate variables onto the bi-binary structural domain: an approximate null space projection operator is constructed based on the autocorrelation matrix of the fixed one-sided factor; a compensation target matrix is constructed using the continuous surrogate variables, the discrete projection residuals, and the approximate null space projection operator; the scaling ratios of the column vectors of the discrete projection solutions are solved column by column to make the column vectors of the compensation target matrix approximate each other by least squares, forming a scaling ratio vector; and the scaling ratio vector is absorbed into the continuous scaling vector corresponding to the discrete projection solution to suppress the propagation error of the discrete projection residuals under the forward mapping of the fixed one-sided factor. The two binary symbol matrices obtained from the bibinary structure decomposition are frozen, and only each continuous scaling vector is used as a learnable parameter for optimization in the reconstruction stage, so as to reduce the difference between the output distribution of the quantized model and the output distribution of the original full-precision model. In the reconstruction phase, the binary symbol matrix is not unfrozen, no additional full-precision proxy weight matrix is introduced, and the approximate null space projection operator is not explicitly retained in the inference phase; the derived quantization model is used to replace the original floating-point matrix multiplication with bit operations and channel-by-channel multiplication on the inference device.
[0006] In a preferred embodiment, the linear projection layer to be quantized includes at least one of the following layers: a query projection layer, a key projection layer, a value projection layer, and an output projection layer of the self-attention mechanism in the neural network model; and a gated projection layer, an upper projection layer, and a lower projection layer of the feedforward neural network in the neural network model.
[0007] In a preferred embodiment, constructing a proxy weight matrix containing second-order sensitivity information based on the input activation statistics and the output gradient statistics includes: Calculate the square root of the squared mean of the input activations in the input activation statistics to form the input importance vector. ,in , The first activation vector of the input activation vector of the linear projection layer to be quantized is... One element, This represents the sample expectation function on the calibration data. The input importance vector The One element; Calculate the square root of the mean of the squared output gradients in the output gradient statistics to form the output importance vector. ,in , The gradient vector of the output of the linear projection layer to be quantized with respect to the task loss is the first... One element, The output importance vector The One element; Using the input importance vector With the output importance vector For the original weight matrix The proxy weight matrix is obtained by performing a two-sided weighting. ,in This represents the element-wise multiplication function with broadcast capability; The proxy weight matrix is based on the Hessian matrix. The Kronecker factor approximation is constructed for the Hessian matrix. It is decomposed into input statistics and output gradient statistics using the Kronecker factor approximation: ,in Let be the input activation vector of the linear projection layer to be quantized. Let be the gradient vector of the output of the linear projection layer to be quantized with respect to the task loss. Represent the Kronecker product function; and retain only the diagonal energy information of the input statistics and output gradient statistics in the Kronecker factor approximation, taking the square root of each element to construct the input importance vector. With the output importance vector .
[0008] In a preferred embodiment, representing the proxy weight matrix as a product of two binary symbolic matrices and several consecutive scaling vectors is specifically expressed as follows: in, and To assign a value to an element A binary symbol matrix, , , These are continuous scaling vectors, representing the output channel scaling vector, intermediate dimension scaling vector, and input channel scaling vector, respectively. Represents the element-wise multiplication function with broadcast; the continuous scaling vector corresponding to the discrete projection solution is taken from... , , At least one of them.
[0009] In a preferred embodiment, the solution obtained through alternating optimization includes: With one side factor fixed as the right factor And update the other factor to the left factor. At that time, the continuous proxy variable The following subproblem based on the alternating direction multiplier method is obtained by solving: The subproblem has a closed-ended solution: in, For the first Discrete left factor in the next inner iteration. For the corresponding dual variable, The default non-negative penalty coefficient is... This is the intermediate dimension of the bibinary structure decomposition. for An identity matrix of order 1. This represents a function that finds the variables that minimize the objective function. The continuous proxy variable Discrete projection is performed onto the bibinarre structural domain to obtain the discrete projection solution, including performing symbol-value independent decomposition projection: taking the symbol matrix. Its element values belong to , A function that takes the sign of each element; for absolute value matrices Find the rank-1 outer product approximation by its first singular triplet. get and ,in , , They are respectively The maximum singular value, the corresponding left singular vector, and the corresponding right singular vector; denoted as the discrete projection solution is... And let the discrete projection residual be denoted as .
[0010] In a preferred embodiment, constructing the approximate null space projection operator based on the autocorrelation matrix of the fixed one-sided factor includes: Calculate the autocorrelation matrix of the fixed one-sided factor; Perform eigenvalue decomposition on the autocorrelation matrix to obtain multiple eigenvalues and their corresponding eigenvectors; Based on the multiple eigenvalues, eigenvectors corresponding to low-energy directions are selected from the eigenvectors to form an orthogonal basis matrix, and the product of the orthogonal basis matrix and its transpose is used as the approximate null space projection operator. The selection method includes at least one of the following: Method 1: Arrange the multiple feature values in descending order and determine the smallest integer that satisfies the following formula. : in, The first in descending order 1 eigenvalue, This is the intermediate dimension of the bibinary structure decomposition. For preset belonging The residual energy ratio threshold of the interval will be related to The corresponding eigenvectors form the orthogonal basis matrix; Method 2: The orthogonal basis matrix is formed by eigenvectors corresponding to eigenvalues that are not greater than a preset eigenvalue threshold among the plurality of eigenvalues; When the eigenvalue decomposition of the autocorrelation matrix exhibits an ill-conditioned condition where the condition number exceeds a preset upper limit, the eigenvalue decomposition is re-executed after adding a ridge term to the autocorrelation matrix. The ridge term is the product of preset coefficients and the identity matrix.
[0011] In a preferred embodiment, when one side factor is the right factor and the other side factor is the left factor, the continuous proxy variable of the left factor is used. Discrete projection solution with left factor Calculate the discrete projection residual According to the discrete projection residual and the approximate null space projection operator Construct the compensation target matrix ; For the discrete projection solution Each column Solve the column-by-column closed scaling ratio using the following formula. : in, This represents a function that calculates the dot product of two vectors. Represents the computation of vectors Functions of norm These are pre-defined non-negative stable terms; The column-by-column closed scaling ratio will be used to determine the scaling ratio of each column. Scaling vector ,according to The intermediate dimension scaling vector is absorbed into the continuous scaling vector through element-wise multiplication. middle; In solving the column-by-column closed scaling ratio Furthermore, it includes at least one of the following independently combinable anomaly handling mechanisms: when the energy of the corresponding column vector of the discrete projection solution... When the energy level is below the preset lower limit, the... Set to a preset backoff value; when the result of the inner product calculation is negative, set the... Set as the preset backoff value; when the absolute value of the difference... When the preset upper limit is exceeded, the... Cut off to a preset neighborhood of constant 1.
[0012] In a preferred embodiment, it further includes: Degradation processing mechanism: When the autocorrelation matrix is subjected to eigenvalue decomposition and the orthogonal basis matrix obtained by the preset rule is an empty set, the construction of the approximate null space projection operator is skipped, and the compensation target matrix is set as the continuous proxy variable. The steps of solving the scaling ratio column by column to form the scaling ratio vector and absorbing the scaling ratio vector into the continuous scaling vector corresponding to the discrete projection solution are continued. The symmetric processing steps for the other factor in the bibinary structural decomposition are as follows: With the left factor of the bibinary structural decomposition fixed... And update the right factor In the case of autocorrelation matrix based on left factor The eigenvalue decomposition constructs the corresponding approximate null space projection operator. Based on the continuous proxy variable of the right factor Discrete projection solution of right factor and the corresponding approximate null space projection operator Construct the compensation objective matrix of the right factor. ,in The scaling ratios of the row vectors of the discrete projection solutions of the right factor are calculated row by row so that the row vectors of the right factor approximate the corresponding rows of the compensation target matrix of the right factor by least squares. The corresponding scaling ratio vectors are then absorbed into the continuous scaling vectors by element-wise multiplication.
[0013] In a preferred embodiment, in the quantization model, each element of the binary symbol matrix is represented by 1 bit and stored as an integer in a bit-packed form according to a preset bit width. The continuously scaling vector is represented in floating point and stored in conjunction with the binary symbol matrix in a channel-by-channel continuous parameter form. The quantization model does not introduce any additional dense matrix or additional dense matrix multiplication operations other than the binary symbol matrix and the continuously scaling vector during the inference phase. The difference between the reduced quantization model output distribution and the original full-precision model output distribution includes using the output distribution of the original full-precision model on the calibration data as the alignment target, and using at least one of Kullback-Leibler divergence, cross-entropy, mean squared error, or contrast loss as the alignment loss function.
[0014] This application also provides a post-training quantization system for deploying large language model inference, oriented towards a neural network model containing a linear projection layer to be quantized, the system comprising: The proxy weight matrix construction module is configured to obtain a tensor with physical meaning obtained by forward and backward propagation of calibration data through the neural network model, which serves as input activation statistics and output gradient statistics, and to construct a proxy weight matrix containing second-order sensitivity information based on the input activation statistics and the output gradient statistics. The bibinary structure decomposition module is configured to perform bibinary structure decomposition on the proxy weight matrix output by the proxy weight matrix construction module, and represent the proxy weight matrix as a product of two binary symbol matrices and several continuous scaling vectors. The alternating optimization solution module is configured to solve the bibinary structure decomposition by alternating optimization when the bibinary structure decomposition module performs the bibinary structure decomposition. When fixing one side factor and updating the other side factor, the continuous surrogate variable of the other side factor is obtained, and the continuous surrogate variable is discretized onto the bibinary structure domain to obtain the discrete projection solution. The discrete projection residual compensation module is configured to perform the following operations during the processing of discrete projection residuals generated by discrete projection of the continuous surrogate variables onto the bibinary structural domain: constructing an approximate null space projection operator based on the autocorrelation matrix of the fixed one-sided factor; constructing a compensation target matrix using the continuous surrogate variables, the discrete projection residuals, and the approximate null space projection operator; solving for the scaling ratios of the column vectors of the discrete projection solution that approximate the corresponding columns of the compensation target matrix by least squares, forming a scaling ratio vector; and absorbing the scaling ratio vector into the continuous scaling vector corresponding to the discrete projection solution to suppress the propagation error of the discrete projection residuals under the forward mapping of the fixed one-sided factor. The reconstruction optimization module is configured to freeze the two binary symbol matrices obtained by the binary structure decomposition module, and only use each continuous scaling vector as a learnable parameter for optimization in the reconstruction stage, so as to reduce the difference between the output distribution of the quantized model and the output distribution of the original full-precision model. The reconstruction optimization module does not unfreeze the binary symbol matrix or introduce additional full-precision proxy weight matrix during the reconstruction phase, and the system does not explicitly retain the approximate null space projection operator during the inference phase; the quantization model derived by the system is used to replace the original floating-point matrix multiplication with bit operations and channel-by-channel multiplication on the inference device.
[0015] The post-training quantization method for large language model inference deployment described in this application involves three stages: surrogate weight matrix construction, bibinary structure decomposition (including approximate null space compensation and closed scaling ratio absorption), and lightweight global reconstruction involving only scaling vectors (see [link to relevant documentation]). Figure 1 This method achieves a synergistic balance between suppressing directional propagation errors and maintaining the inference graph under extremely low bit quantization conditions. The technical effects are explained below, focusing on the various technical methods and their collaborative approaches.
[0016] In the construction of the proxy weight matrix, the method obtains the input activation statistics and output gradient statistics from the calibration data through forward and backward propagation of the neural network model, and calculates the square root of the squared mean of the input activations to form the input importance vector. And the square root of the mean of the output gradient to form the output importance vector. And use both to adjust the original weight matrix Perform bilateral weighting to obtain the proxy weight matrix The construction of the input and output importance vectors described here is based on the approximation of the Hessian matrix using the Kronecker factor. After retaining only the diagonal energy information of each factor, the square root of each element is taken; thus, the original reconstruction objective based on Hessian weighting is equivalently transformed into a reconstruction based on the standard Frobenius norm in the surrogate space. The minimization problem is addressed. The technical effect of this equivalent transformation is that the weight regions corresponding to high-activation input channels and high-gradient output channels are proportionally enlarged in the surrogate weight matrix. When subsequent bibinarre structural decomposition is performed in this surrogate space, the limited binary representation precision can be preferentially allocated to the weight regions that have a greater impact on task loss, thus helping to alleviate quantization degradation of critical channels under extremely low bit conditions. The construction of the surrogate weight matrix is applicable to different types of linear projection layers to be quantized, such as query projection layers, key projection layers, value projection layers, and output projection layers in self-attention mechanisms, as well as gated projection layers, up projection layers, and down projection layers in feedforward neural networks.
[0017] In the bi-binary structure decomposition stage, the method represents the surrogate weight matrix as follows: The product form, where and To assign a value to an element A binary symbol matrix, , , These are the output channel scaling vector, the intermediate dimension scaling vector, and the input channel scaling vector, respectively. Through an alternating optimization strategy, with a fixed right factor... And update the left factor In this case, the method constructs a convex quadratic problem based on the alternating direction multiplier method and obtains a closed-form solution. .because Only For a matrix of this size, the subproblem can be solved efficiently using a single Cholesky decomposition, and within a fixed... The decomposition results can be reused throughout the entire inner-layer iteration, thereby reducing the computational overhead during the calibration period. For the continuous surrogate variables... When performing a sign-value independent decomposition projection, first take the sign matrix. Then, for the absolute value matrix Find the first singular triplet to obtain the output channel scaling vector components. With intermediate dimension scaling vector components This transforms the surrogate variables in the continuous domain into discrete projective solutions that satisfy the bibinary structure constraints. This decomposition method allows symbolic and amplitude information to be extracted and processed independently, which helps to preserve the directionality and scale information in continuous surrogate variables as much as possible under extremely low bit constraints.
[0018] In the process of handling discrete projection residuals—which is the core technology of this application—the method does not globally minimize the discrete projection residuals in the sense of the global Frobenius norm. Instead, it is based on the autocorrelation matrix of a fixed one-sided factor. Based on the energy distribution characteristics revealed by the eigenvalue decomposition, the directional components in the residuals are differentiated. Specifically, by... Perform eigenvalue decomposition and threshold based on residual energy ratio Alternatively, an absolute eigenvalue threshold can be used to select eigenvectors from the eigenvectors corresponding to low-energy directions to form an orthogonal basis matrix. ,by As an approximate null space projection operator, the technical advantage of this projection operator lies in its ability to handle arbitrary residual matrices. its along The projected components are Energy of forward mapping By explicitly constraining the eigenvalues of the fixed factor in the low-energy direction, a formal upper bound can be derived when using the residual energy ratio threshold selection method. This allows the "residual energy allowed to propagate along the forward mapping" to be parameterized as a threshold and a fixed-factor full energy that can be estimated a priori from calibration period statistics, enabling prior determination of the error budget for the final deployed model without relying on end-to-end measured feedback. The two optional low-energy direction selection methods—adaptive selection based on residual energy ratio and selection based on absolute eigenvalue threshold—are applicable to different inter-layer spectral distribution differences, making the method adaptable to the quantization of different types of linear projection layers to be quantized.
[0019] Furthermore, the method utilizes continuous proxy variables. Discrete projection residuals and the approximate null space projection operator Construct the compensation target matrix And solve the closed scaling ratio for each column separately. The scaling vector is composed of the closed scaling ratios of each column. By element-wise multiplication Absorb to intermediate dimension scaling vector The technical effect of this absorption process has a dual significance: on the one hand, it approximates the remaining difference. Falling on fixed factors In the direction of the approximate null space, via During forward propagation, its energy is explicitly suppressed by the aforementioned formalized upper bound; on the other hand, this absorption process does not introduce any new dense matrix or new learnable parameters, whereas it would otherwise require... The directional projection, which can be expressed by dense matrix exponentiation, is closed-formed into scalar multiplication of existing channel-wise scaling vectors, ensuring that the inference deployment computation graph strictly maintains its bi-binary structure. This technique means that the approximate null space projection operator is constructed and used only during the calibration phase and is not explicitly retained during the inference phase; its full compensation effect is fully contained within the continuous scaling vectors.
[0020] Regarding the handling of abnormal operating conditions, when the orthogonal basis matrix selected after the eigenvalue decomposition of the autocorrelation matrix is an empty set (i.e., the fixed factor is full rank overall under the current numerical precision), the method skips the construction of the approximate null space projection operator, sets the compensation target matrix as the continuous surrogate variable itself, and only performs closed-loop scaling absorption, thereby smoothly degenerating to the baseline scheme based on ADMM and SVID projections instead of introducing numerical collapse. Similarly, when the column vector energy of the discrete projection solution is too low, or the sign of the inner product calculation result is reversed, or the scaling ratio deviates significantly from the constant 1, the method performs robustness processing with a preset backoff value or truncation operation, respectively. When the eigenvalue decomposition of the autocorrelation matrix shows a numerical ill-conditioned condition with the condition number exceeding a preset upper limit, it is stabilized by adding a ridge term and re-performing the eigenvalue decomposition. The above degradation processing mechanisms can be combined independently, ensuring that the method does not produce worse results than the baseline scheme in the worst case, which helps to stably apply the method to different types of linear projection layers to be quantized and adapt to the differences in inter-layer spectral distribution under different model sizes.
[0021] For updating the right factor in a bibinarre structural decomposition, the method employs a processing approach that is completely symmetric to the left factor: based on the autocorrelation matrix of the left factor. The eigenvalue decomposition constructs the corresponding approximate null space projection operator. Construct the compensation objective matrix of the right factor. The scaling ratios are calculated row by row and then absorbed into the corresponding continuous scaling vectors. By performing the approximate null space projection and closed scaling ratio absorption symmetrically by left and right factors, the discrete projection residuals on both sides of the bibinary structural decomposition are structurally processed, which helps to improve the overall stability of the decomposition.
[0022] In the lightweight global reconstruction step of freezing the binary symbol matrix, the method freezes all binary symbol matrices determined by the aforementioned bibinarre structure decomposition, and only uses the output channel scaling vector of each linear projection layer to be quantized. , intermediate dimension scaling vector With input channel scaling vector The set As learnable parameters, the output distribution of the original full-precision model is optimized as the alignment target. Since the binary sign matrix is not unfrozen, and no full-precision shadow weights or additional surrogate weight matrices are introduced, the scale of learnable parameters per layer is significantly reduced compared to the case of unfrozen complete two-factor matrices, making it possible to perform quantized reconstruction of large-scale neural network models under memory-constrained environments. The alignment loss function can employ at least one of Kullback-Leibler divergence, cross-entropy, mean squared error, or contrastive loss. When using Kullback-Leibler divergence, a temperature coefficient can be further introduced to adjust the attention to low-probability categories. It should be noted that because the directional compensation effect has been pre-injected into the continuous scaling vector through a closed scaling ratio in the aforementioned binary structure decomposition stage, the lightweight global reconstruction can effectively complete global distribution alignment with such a limited scale of learnable parameters—there is an inseparable synergistic relationship between the two.
[0023] In terms of deployment, storage, and inference computation, each element of the binary symbol matrix in the quantization model is represented by 1 bit and stored as an integer in a bit-packed manner with a preset bit width. The continuously scaled vector is represented by floating-point numbers and stored in a channel-by-channel continuous parameter format. During the inference phase, the quantization model does not introduce any additional dense matrices or additional dense matrix multiplication operations besides the binary symbol matrix and the continuously scaled vectors. Instead, it replaces the original floating-point matrix multiplication with bitwise operations and channel-by-channel multiplication. Since the approximate null space projection operator is not retained or called as model parameters during the inference phase, the inference computation graph of the final deployed model is strictly consistent with the inference computation graph without null space compensation. Thus, while introducing a directional residual control mechanism, it strictly maintains the hardware deployment friendliness of the extremely low-bit training quantization scheme, which helps to reduce the memory usage and memory access bandwidth requirements of the inference device.
[0024] In summary, the aforementioned techniques work together under the strong constraints of "quantization after training, no unfreezing of complete weights, and no change in the inference computation graph": second-order sensitivity preconditioning allows decomposition and compensation to be performed in the proxy space aligned with task loss; directional null space compensation makes the propagation error of discrete projection residuals subject to explicit constraints of formal upper bounds; closed scaling ratio absorption folds dense projection effects into channel-wise scalar operations to keep the inference computation graph unchanged; multi-layer degradation compatibility mechanism ensures the robustness of the scheme in engineering implementation; and lightweight reconstruction of the frozen binary symbol matrix achieves global distribution alignment with a significantly reduced parameter scale. These effects are interdependent, and implementing any one of them alone is difficult to achieve a good balance among the three technical requirements that need to be considered in engineering implementation: directional propagation error suppression, zero additional inference parameters, and lightweight calibration. However, this application achieves a synergistic balance of these requirements through a specific combination of the aforementioned techniques.
[0025] The specification of this application contains numerous technical features distributed across various technical solutions. Listing all possible combinations of these technical features (i.e., technical solutions) would make the specification excessively lengthy. To avoid this problem, the various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which are considered to have been described in this specification), unless such a combination of technical features is technically infeasible. For example, one example discloses feature A+B+C, and another example discloses feature A+B+D+E. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; they cannot be used simultaneously. Feature E can technically be combined with feature C. Therefore, the solution A+B+C+D should not be considered as described because it is technically infeasible, while the solution A+B+C+E should be considered as described. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the overall process of a post-training quantization method for deploying large language model inference according to an embodiment of this application.
[0027] Figure 2 This is a schematic diagram of the structure of a computing device according to an embodiment of this application.
[0028] Figure 3 This is a schematic diagram of the structure of a post-training quantization system for deploying large language model inference according to an embodiment of this application. Detailed Implementation
[0029] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0030] Explanation of some concepts: Large language models refer to neural network models that contain a large number of parameters and are capable of performing text generation, understanding, or reasoning tasks based on input sequences, including but not limited to pre-trained language models based on Transformer architecture, hybrid expert architecture, or state space model architecture.
[0031] Post-training quantization refers to a model compression method that, after a neural network model has completed pre-training, does not re-execute the complete model training process, but instead uses only a small amount of calibration data to perform low-bit representation conversion of the model weights.
[0032] A linear projection layer to be quantized refers to a linear transformation layer in a neural network model that needs to be converted from the original floating-point matrix multiplication to a quantized computation form. This includes at least one of the query projection layer, key projection layer, value projection layer, and output projection layer in a self-attention mechanism, as well as the gated projection layer, up projection layer, and down projection layer in a feedforward neural network.
[0033] Calibration data refers to sample data used to collect input activation statistics, output gradient statistics, and for post-quantization reconstruction or output distribution alignment, and is usually sampled from the target domain corpus.
[0034] Input activation statistics refer to the input activation-related statistics collected at the input end of the linear projection layer to be quantized when calibration data is propagated forward through the neural network model.
[0035] Output gradient statistics refer to the gradient-related statistics collected at the output of the linear projection layer to be quantized when the calibration data is backpropagated through the neural network model.
[0036] The surrogate weight matrix is a weight matrix obtained by weighting the original weight matrix on both sides using the input importance vector and the output importance vector. The weighting coefficients contain second-order sensitivity information from the Hessian matrix approximation, making subsequent optimizations performed on the surrogate weight matrix equivalent to optimizations performed in a weighted space aligned with the task loss.
[0037] The input importance vector is a vector composed of the square roots of the squared mean of the input activations. ,in , is used to characterize the relative impact of different input dimensions of the linear projection layer to be quantized on the reconstruction target.
[0038] The output importance vector is a vector composed of the square roots of the squared mean of the output gradients. ,in , used to characterize the relative impact of different output dimensions of the linear projection layer to be quantized on the task loss.
[0039] Bibinarity structure decomposition refers to approximating the weight matrix as two elements whose values belong to different values. The matrix decomposition method, in the form of the product of a binary symbolic matrix and several continuously scaled vectors, is specifically expressed as follows: .
[0040] A binary symbolic matrix is a matrix in which each element takes only one value. or A matrix of two discrete values, denoted in this application as and During deployment, each element is represented by 1 bit and stored in bit-packed form.
[0041] A continuous scaling vector refers to a full-precision floating-point vector used in conjunction with a binary sign matrix in bibinarre structure decomposition, including output channel scaling vectors. , intermediate dimension scaling vector With input channel scaling vector It is used to recover the numerical scale of the weight matrix by channel-wise multiplication during inference.
[0042] Continuous surrogate variables refer to intermediate variables with continuously varying values obtained by solving a convex quadratic subproblem based on the alternating direction multiplier method during alternating optimization, where one side's factors are fixed and the other side's factors are updated. or .
[0043] Discrete projection solution refers to the matrix that satisfies the bibinarity constraint after discretizing the continuous surrogate variables onto the bibinarity structure domain, denoted as . or .
[0044] Discrete projection residuals are the difference matrix between continuous surrogate variables and their corresponding discrete projection solutions, denoted as... or , which represents the structural mismatch error generated when a continuous variable is projected onto a bi-binary structural domain.
[0045] Symbol-value independent decomposition projection refers to a discrete projection method that first extracts the element-wise symbol matrix of a continuous surrogate variable, and then performs a rank-1 outer product approximation on its absolute value matrix to obtain the scaling vector components. It is also known as SVID projection.
[0046] The approximate null space projection operator refers to the orthogonal basis matrix spanned by the eigenvalues of the autocorrelation matrix with a fixed one-sided factor, and the corresponding eigenvectors in the low-energy directions. Constructed projection operator It is used to identify and isolate the directional components in the discrete projection residuals that have a weak impact on the output after forward mapping by a fixed factor during the calibration phase.
[0047] The autocorrelation matrix is a square matrix obtained by multiplying a fixed one-sided factor by its transpose. At that time When the left factor is fixed At that time Its eigenvalue decomposition reveals the energy distribution characteristics of the fixed factor in various directions.
[0048] The compensation target matrix refers to the intermediate matrix constructed using continuous surrogate variables, discrete projected residuals, and the approximate null space projection operator, denoted as [missing information] in the left factoring case. As the target of subsequent column least squares approximation, the remaining difference after approximation falls on the direction of the approximate null space of the fixed factor.
[0049] The residual energy ratio threshold refers to a preset threshold used to adaptively determine the dimension of the approximate null space when constructing the approximate null space projection operator. Its control is the upper limit of the proportion of residual energy to the total energy of the fixed factor that is allowed to propagate along the forward mapping.
[0050] Closed scaling ratio refers to the scaling factor obtained directly through least squares calculation, used to absorb the compensation effect into the continuous scaling vector, denoted as . .
[0051] Alternating Direction Multiplier Method (ADMM) is a numerical optimization framework that decomposes the original optimization problem into subproblems that can be solved alternately, and gradually converges the solutions of the subproblems to the solution of the original problem by introducing dual variables and penalty terms.
[0052] The Kronecker factor approximation refers to representing the Hessian matrix corresponding to the linear projection layer as the Kronecker product of the input statistics and the output gradient statistics. The method is used to decompose global second-order information into two smaller independent factors.
[0053] The reconstruction phase refers to the optimization phase after the binary symbol matrix is frozen, where only the continuously scaling vectors are used as learnable parameters to make the output distribution of the quantized model approximate the output distribution of the original full-precision model.
[0054] Bit packing refers to a data organization method that packs multiple 1-bit elements in a binary symbol matrix into an integer according to a preset bit width for storage, so as to achieve efficient parallel computing using bit operation instructions during the inference stage.
[0055] Inference devices refer to computing devices used to invoke quantization models to perform forward inference calculations, including at least one of servers, graphics processing units, edge computing nodes, or dedicated neural network accelerators.
[0056] Element-wise division refers to performing division on two vectors of the same shape, corresponding to each element. .
[0057] A tensor with physical meaning refers to a statistical tensor that is collected at the input and output ends of the linear projection layer to be quantized during the forward and backward propagation of calibration data through a neural network model. It is used to characterize the input activation distribution and output gradient distribution of the layer and is different from arbitrary numerical tensors.
[0058] The following is a brief summary of some of the innovative aspects of this application: In summary, in the specific technical scenario of extremely low-bit post-training quantization for large language model inference deployment, the core technical challenge addressed by this application lies in: when the original weight matrix is represented as two binary symbol matrices through bibinarity decomposition. , With continuously scaled vectors , , After the product form (step 210), the continuous proxy variables (Step 220) The discrete projection residual generated by projecting SVID onto the bibinary structural domain (Step 230) Its right factor is fixed. The perturbation to the output after forward mapping is not the same as... The proportional relationship indicates that the conventional approach of minimizing the discrete projected residuals globally under the Frobenius norm cannot distinguish the directionality of the forward propagation error. However, this directional discrimination capability is precisely needed under extremely low bit conditions to accurately suppress error components that significantly contribute to task loss. The technical solution of this application explicitly decouples the two mathematically fundamentally different objectives of "residual minimization" and "propagation error minimization": In step 100, based on the Hessian matrix... The input importance vector obtained by approximating the diagonal energy of the Kronecker factor. With output importance vector Constructing the proxy weight matrix This places the entire decomposition and compensation process in a second-order weighted space aligned with the task loss; in the processing of discrete projection residuals, the autocorrelation matrix of the fixed factor is adjusted... Perform eigenvalue decomposition and threshold based on residual energy ratio Constructing an approximate null space projection operator by selecting low-energy directions Then construct the compensation target matrix. Solve the closed scaling ratio by column After passing This need The directional projection effect that can be expressed by the power of dense matrices folds into a vector that scales the existing intermediate dimension. Channel-by-channel multiplication—this folding depends on the discrete projection solution The column-wise structure of the target matrix and the inner product projection between each column can be scalarly scaled to approximate the non-obvious mathematical property that has not been recognized or utilized in existing post-training quantization literature. There is a deep synergy between the above steps and lightweight global reconstruction using only continuous scaling vectors as learnable parameters: because the aforementioned steps have pre-injected the directional compensation effect into the scaling vector, lightweight global reconstruction can complete global distribution alignment with a significantly lower parameter scale without unfreezing the binary sign matrix or introducing full-precision shadow weights, so that the inference computation graph strictly maintains the bibinary structure without introducing additional dense parameters. This holistic scheme, which is synergistically established under the strong constraints of "post-training quantization, not unfreezing the complete weights, and unchanged inference computation graph", constitutes the core technical contribution of this application.
[0059] Furthermore, through long-term and in-depth research, the inventors of this application have discovered that the core difficulty faced by existing ultra-low bit training post-quantization techniques stems not only from the insufficient quantization bit width itself, but also from the fundamental flaw in the implicit assumption that existing methods equate "global minimization of discrete projection residuals in the Frobenius norm sense" with "minimization of forward propagation error." Specifically, after decomposing the weight matrix into the product of two binary sign matrices and a continuous scaling vector, the perturbation of the model output by the residual matrix generated by the discrete projection of continuous surrogate variables onto the bibinarity domain, after forward mapping with a fixed one-sided factor, depends on the product of the projection magnitude of each residual component along different characteristic directions of the fixed factor and the corresponding eigenvalue of that direction—rather than being characterized by the global norm of the residual matrix. The inventors further recognized that the characteristic spectrum of the autocorrelation matrix of a fixed factor typically exhibits a phenomenon where energy is highly concentrated in a few principal directions. This means that the contribution of the residual components along low-energy characteristic directions to the output is naturally compressed by the small eigenvalues in those directions after propagation through forward mapping; conversely, the residual components along high-energy directions are amplified by large eigenvalues. Existing methods minimize the residuals as a whole without distinguishing between these two types of directions, which not only wastes resources in optimizing resource allocation but also fails to specifically suppress errors in highly sensitive directions.
[0060] The inventors also recognized another mutually restrictive technical constraint: even if the dense projection structure required for directional residual control (such as the approximate null space projection operator) is identified during the calibration phase, retaining this dense structure as an additional parameter during the inference phase would inevitably introduce additional dense matrix multiplication operations, which would be detrimental to achieving the goal of extremely low bit quantization to reduce memory access bandwidth and computational overhead during the inference phase. Through in-depth analysis of the column-by-column structure of the compensation target matrix, the inventors discovered that the least-squares approximation problem between the column vectors of the discrete projection solution and the corresponding columns of the compensation target matrix can be decoupled into independent univariate scalar optimization. Its closed-form solution can be absorbed into the existing continuous scaling vector through element-by-element multiplication, thus completely folding the dense projection effect constructed during the calibration phase into the channel-by-channel multiplication already present in the inference phase, achieving directional residual control without altering the inference computation graph. Furthermore, through practical observation of numerous different model architectures and layer types, the inventors discovered that frequent occurrences during ultra-low bit quantization, such as near-null space degradation, excessively low column vector energy, scaling ratio sign inversion, and excessively large autocorrelation matrix condition numbers, can easily lead to numerical collapse of the entire calibration pipeline if systematic degradation compatibility processing is lacking. Based on the above in-depth research, the inventors of this application propose a post-training quantization method that combines second-order sensitivity preconditioning, directional null space compensation based on fixed-factor autocorrelation spectrum with closed scaling ratio absorption, and lightweight global reconstruction involving only continuous scaling vectors. The implementation process of this application is described in detail below through specific embodiments.
[0061] The following description, in conjunction with the accompanying drawings, further illustrates several embodiments of this application; however, the scope of protection of this application is not limited to the following embodiments. It should be noted that the features involved in the following embodiments can be combined with each other without conflict, and all resulting combinations fall within the scope of protection of this application.
[0062] Example 1: Overall Quantization Process After Three-Stage Training This embodiment presents the overall flow of a post-training quantization method for deploying inference in large language models. The method is implemented for neural network models containing linear projection layers to be quantized. Specifically, the neural network model includes, but is not limited to, large language models, visual language models, and other neural network models containing large-scale linear projection layers. The linear projection layers to be quantized include, but are not limited to, query projection layers, key projection layers, value projection layers, and output projection layers in self-attention mechanisms, and at least one of gated projection layers, up projection layers, and down projection layers in feedforward neural networks. Referring to Figure 1, the method described in this embodiment executes the following steps sequentially.
[0063] Step 100: Obtain the tensors with physical meaning obtained from the forward and backward propagation of the calibration data through the neural network model, and use these tensors as input activation statistics and output gradient statistics, respectively. Construct a surrogate weight matrix containing second-order sensitivity information based on the input activation statistics and the output gradient statistics. Specifically, a small amount of text is sampled from the target domain corpus as a calibration data set. For each linear projection layer to be quantized, its forward input activation vector and backward output gradient vector are extracted, and the squared mean statistic of its corresponding components is accumulated. Based on this, a surrogate weight matrix is constructed. The specific implementation of step 100 is further described in Example 2.
[0064] Step 200: In the proxy weight matrix The process involves performing a bibinarity-based structural decomposition, where the surrogate weight matrix is represented as a product of two binary sign matrices and several continuous scaling vectors. This bibinarity-based structural decomposition is solved through alternating optimization. While fixing one side of the factors and updating the other side, a continuous surrogate variable for the other side is obtained, and this continuous surrogate variable is discretely projected onto the bibinarity-based structural domain to obtain a discrete projection solution. In processing the discrete projection residuals generated by the discrete projection of the continuous surrogate variables onto the bibinarity-based structural domain, an approximate null space projection operator is constructed based on the autocorrelation matrix of the fixed one-side factors. A compensation target matrix is constructed using the continuous surrogate variables, the discrete projection residuals, and the approximate null space projection operator. The scaling ratios of the column vectors of the discrete projection solution are calculated column-wise to approximate the corresponding columns of the compensation target matrix using least-squares scaling, forming a scaling ratio vector. This scaling ratio vector is then absorbed into the continuous scaling vector corresponding to the discrete projection solution to suppress the propagation error of the discrete projection residuals under the forward mapping of the fixed one-side factors. The specific implementation of step 200 is further described in Examples 3 to 6.
[0065] Step 300: Freeze the two binary symbol matrices obtained from the bi-binary structure decomposition, and only use each continuous scaling vector as a learnable parameter for optimization in the reconstruction stage to reduce the difference between the output distribution of the quantized model and the output distribution of the original full-precision model. The specific implementation of step 300 is further described in Example 7.
[0066] Step 400: Export the deployment model. Specifically, the binary symbol matrix is stored in bit-packed form, and the continuous scaling vector is stored in a channel-by-channel continuous parameter form in conjunction with the binary symbol matrix. The quantization model is used to replace the original floating-point matrix multiplication with bit operations and channel-by-channel multiplication on the inference device, thereby reducing the memory access bandwidth and video memory usage of the inference device. The specific implementation of step 400 is further described in Example 8.
[0067] It should be noted that the three stages described in this embodiment are not isolated from each other, but rather work together: the surrogate weight matrix obtained in step 100 explicitly amplifies the high activation channels at the input end and the high gradient channels at the output end, so that the bibinarity decomposition in the subsequent step 200 can prioritize the protection of the weight regions that are more sensitive to task loss; the approximate null space compensation in step 200 explicitly constrains the propagation energy of the extremely low bit discrete projection residual along the forward mapping without introducing inference period parameters; and step 300, under the condition that the binary structure is fixed, completes the end-to-end distribution alignment with a learnable parameter scale much lower than that of the full bifactor matrix. More specifically, the binary symbol matrix is not unfrozen in the reconstruction stage, no additional full-precision surrogate weight matrix is introduced, and the approximate null space projection operator is not explicitly retained in the inference stage; the approximate null space projection operator is only constructed and used in the calibration stage, and its compensation effect is fully carried out through the channel-by-channel multiplication of the continuously scaled vector in step 200. Therefore, the computation graph of the quantization model during the inference phase remains consistent with the computation graph without the introduction of the approximate null space projection operator.
[0068] Example 2: Construction of the surrogate weight matrix based on second-order sensitivity (step 100) This embodiment further illustrates the construction of the proxy weight matrix containing second-order sensitivity information in step 100. The specific implementation method is as follows. This embodiment no longer directly uses the unweighted Frobenius norm as the quantization target, but instead establishes a weighted reconstruction criterion based on the second-order approximation of the task loss on the weight perturbation.
[0069] Let the original weight matrix be... The compressed weight matrix is Then the increment of loss can be approximated as: in, Let Hessian be the matrix corresponding to the linear projection layer to be quantized. This function represents stacking matrices column-wise into a vector. The transpose function of a matrix or vector. This represents an approximate value for the incremental loss of the task.
[0070] Since the global Hessian matrix is enormous and difficult to store and compute directly, this embodiment approximates it by using the Kronecker factor to decompose it into input statistics and output gradient statistics: in, Let be the input activation vector of the linear projection layer to be quantized. Let be the gradient vector of the output of the linear projection layer to be quantized with respect to the task loss. This represents the Kronecker product function. In the calibration data set The sample expectation function.
[0071] Step 110: In the calibration data set For each linear projection layer to be quantized, the forward input activation vector and the backward output gradient vector are hooked together, and the squared mean of their elements is accumulated. Specifically, during the forward propagation process, the average of the squared elements is accumulated. Accumulated during backpropagation ,in The first of the input activation vectors One element, The gradient vector is the first Each element.
[0072] Step 120: Retain only the diagonal energy information of the input statistics and output gradient statistics in the Kronecker factor approximation, take the square root of each element, and define the input importance vector. With output importance vector as follows: in, The input importance vector The One element, The output importance vector The Each element.
[0073] Step 130: Based on the above definition, the reconstruction problem originally located in the Hessian weighted space is equivalently rewritten as a standard Frobenius norm minimization problem in the surrogate space: in, The function represents a diagonal matrix composed of vector elements. This represents the Frobenius norm function of a matrix.
[0074] Step 140: Based on the equivalent transformation, the proxy weight matrix is defined in this embodiment as follows: in, This represents an element-wise multiplication function with broadcasting. Since the high-activation input channel and the high-gradient output channel are proportionally amplified after being weighted on both sides, this is subsequently applied to the surrogate weight matrix. The binary structure decomposition performed on the above will prioritize protecting the weight region that has a greater impact on task loss.
[0075] It should be noted that the approximation method of the Hessian matrix described in this embodiment is not limited to the Kronecker factor approximation. In an optional implementation, the diagonal Fisher approximation, KFAC approximation, or other approximation methods that can obtain second-order sensitivity estimates for the input and output sides respectively can be used; furthermore, when Time (of which) (where the identity matrix is used), the construction degenerates into a traditional Frobenius norm minimization objective; when a diagonal Hessian approximation is used, the construction degenerates into a channel-wise weighted saliency scheme. All of the above variations fall within the reasonable variation range of this embodiment.
[0076] Example 3: Basic Representation of Bibinarity Structure Decomposition and Solution of Continuous Proxy Variables (Step 200-1) This embodiment further illustrates the basic form of the bibinarity structure decomposition in step 200 and the solution method for continuous surrogate variables. After obtaining the surrogate weight matrix... Subsequently, this embodiment employs a dual binary structure for extremely low-bit modeling.
[0077] Step 210: Transfer the proxy weight matrix It can be approximated as the product of two binary sign matrices and several continuously scaled vectors: in, and To assign a value to an element A binary symbol matrix, , , Let these be continuously scaling vectors, representing the output channel scaling vector, the intermediate dimension scaling vector, and the input channel scaling vector, respectively. For ease of describing alternating optimization, the bibinarre representation is equivalently denoteed as the product of the left and right factors. ,in , , The intermediate dimension of the bibinary structure decomposition typically satisfies , and These are the output and input dimensions of the proxy weight matrix, respectively. Optionally, the intermediate dimension... Press Take values, where The preset compression ratio is... The function is the floor function; or it is the function that makes the surrogate weight matrix... The truncated singular value decomposition retains the smallest integer selection of the preset energy ratio.
[0078] because and The optimization problem, which contains both discrete symbolic constraints and continuously scaled variables, is a highly nonconvex problem. This embodiment employs an alternating optimization strategy for solution. Steps 220 to 230 below use a "fixed right factor" approach. Update left factor "Taking an example, we fix the left factor." Update right factor The process is executed in a completely symmetrical manner, and the corresponding processing is further described in Example 6.
[0079] Before performing alternating optimization, the left factor needs to be evaluated. With right factor Initialize the proxy weight matrix. In one implementation, initialize the proxy weight matrix. Perform truncated singular value decomposition ,Pick , As the initial value, where , , Before and after retention The left singular vector matrix, the singular value diagonal matrix, and the right singular vector matrix of each singular component. To The matrix obtained by taking the square root of each element. In another optional implementation, random initialization or initialization based on K-means clustering can be used. The choice of initialization method does not affect the logical structure of the method steps.
[0080] Step 220: Fix the right factor For the current value, introduce a continuous proxy variable. Based on the Alternating Direction Method of Multipliers (ADMM), the following convex quadratic problem is constructed: in, For the first Discrete left factor in the next inner iteration. For the corresponding dual variable, The default non-negative penalty coefficient is... This represents a function of the variables that minimizes the objective function. The nonnegative penalty coefficient... Used to balance data fidelity terms and proximal terms; in one implementation, the non-negative penalty coefficient The value of is in to The range is set by those skilled in the art based on the calibration objectives during specific implementation.
[0081] The subproblem is a convex quadratic problem, and its relation to... Finding the first-order optimality condition, we can obtain a closed-form solution: in, for An identity matrix of order 1. It should be noted that... Only The matrix, whose size is much smaller than the dimension of the original weight matrix, can be efficiently solved by a single Cholesky decomposition; and in a fixed... In the entire inner-layer iteration, the results of the Cholesky decomposition can be reused to amortize the computational overhead of the calibration period.
[0082] Step 230: For the continuous proxy variables Perform a Sign-Value Independent Decomposition (SVID) projection to obtain a discrete projective solution that satisfies the bibinarity structure constraint. Specifically, the SVID projection described in this embodiment includes the following sub-steps.
[0083] Sub-step 231: Obtain the symbol matrix Its element values belong to ,in This represents a function that takes the sign of each element. Optionally, in one implementation, it is stipulated that... This is to handle the zero-element case.
[0084] Sub-step 232: For the absolute value matrix Find the rank-1 outer product approximation by its first singular triplet. get: in, , , They are respectively The maximum singular value, the corresponding left singular vector, and the corresponding right singular vector. , These are the output channel scaling vector components and the intermediate dimension scaling vector components obtained based on the SVID projection, respectively. For example, the first singular triplet can be solved using any of the following methods: truncated singular value decomposition, power iteration, Lanczos iteration, or stochastic singular value decomposition, with a preset relative error or maximum number of iterations used as the convergence criterion. Optionally, when the first singular value... With the second singular value When the numerical values are close, i.e., when the rank-1 approximation accuracy is insufficient, the rank approximation can be used. (in The singular value decomposition is truncated and multiple singular components are merged and absorbed into the scaling vector to maintain stability.
[0085] Sub-step 233: Let the discrete projection solution be: And accordingly, the discrete projection residual is defined as: The discrete projection residual This characterizes the structural mismatch error resulting from the rigid projection of the continuous surrogate variable onto a bibinarre structural domain. It is important to emphasize that this residual does not necessarily need to be minimized globally in the sense of the Frobenius norm; the core insight of this application lies in the fact that for a fixed right factor... In this regard, the component of the residual along its low-energy propagation direction has a weaker impact on the forward mapping output. Based on this, the following embodiment four will provide directional structuring processing for the discrete projection residual.
[0086] Example 4: Approximate null space projection and closed column scaling absorption based on fixed factor autocorrelation spectrum (Step 200-2) This embodiment further illustrates the step 200 regarding the discrete projection residual. The null space compensation mechanism and the solution and absorption process of the closed column scaling ratio are described. This process is one of the core aspects of this application. It should be noted that the approximate null space projection operator described in this embodiment is only constructed and used in the calibration stage, and is not used as a model parameter in the inference stage, nor does it appear in the computation graph of the final deployed model; its full effect is fully carried out through channel-by-channel multiplication of the existing continuous scaling vector.
[0087] Step 240: Construct an approximate null space projection operator.
[0088] Sub-step 241: Calculate the fixed right factor Autocorrelation matrix: And for the autocorrelation matrix Execution feature decomposition: in, The eigenvalues are arranged in descending order. , This is the corresponding orthogonal eigenvector matrix. Because... The scale is only The overhead of the eigenvalue decomposition is much lower than the scale of the original weight matrix; and at a fixed scale... In the entire inner layer iteration, the feature decomposition result only needs to be executed once.
[0089] Sub-step 242: From The orthogonal basis matrix is formed by selecting the eigenvectors corresponding to the low-energy directions. The selection method includes at least one of the following two.
[0090] The first optional implementation is an adaptive selection based on the residual energy ratio threshold. Let... To determine the preset residual energy ratio threshold, find the smallest integer that satisfies the following formula. : and take It should be noted that the residual energy is more than the threshold. Used to establish a quantifiable trade-off between compensation coverage and residual suppression strength: The smaller the value, the better the selected orthogonal basis matrix. The closer the spanned subspace is to the true null space, the more thoroughly the residual is suppressed, but the fewer directions the compensation can cover; The larger, The wider the coverage, the weaker the suppression. Optionally, the residual energy is higher than a threshold. The value of is in to The range is set by those skilled in the art based on the calibration objectives during specific implementation.
[0091] The second alternative implementation is based on the selection of an absolute threshold. Let... The preset feature value threshold will satisfy Composition of all feature vectors Optionally, the feature value threshold can be determined by... Given, among which and These are the preset relative coefficients and absolute lower bounds.
[0092] Sub-step 243: From the orthogonal basis matrix Constructing an approximate null space projection operator: The approximate null space projection operator Idempotent With symmetry The mathematical properties are used to explain the operators. The structural characteristics of [the structure / feature] do not constitute an additional limitation that must be present simultaneously in this embodiment. For any [specific structure / feature]... 3D row vector ,have: in, Represents the computation of vectors A function of the norm. Applying the above equation row by row to the residual matrix. We can obtain: Furthermore, in an implementation based on a residual energy ratio threshold, by We can obtain: In the implementation based on an absolute threshold, by Directly obtain The above inequality gives the allowable along a fixed factor. The formal absolute upper bound of the residual energy of forward propagation; the threshold and the total energy of the fixed factor in this upper bound can be estimated a priori by the calibration period statistics. By adjusting the threshold, those skilled in the art can make a priori determination of the propagation error budget of the final deployed model during the calibration phase without relying on end-to-end measured feedback.
[0093] Step 250: Construct the compensation target matrix and solve for the closed column scaling ratio.
[0094] Sub-step 251: Construct the compensation target matrix: It should be noted that the compensation target matrix The purpose is that this embodiment does not require the discrete projection solution. The residual itself falls directly into Instead of using an approximate zero space, an intermediate quantity is first constructed during the calibration phase for subsequent closed-loop absorption. Then, the discrete projection solution is approximated to the intermediate quantity by using a scaling ratio vector; thereby reducing the remaining difference after approximation. Falling In the direction of the approximate null space, via During forward propagation, its energy is explicitly suppressed according to the above inequality.
[0095] Sub-step 252: Solve the discrete projection solution column by column. The column vector least squares approximation of the compensation target matrix The scaling ratio for the corresponding column. Specifically, for each column... Solve the univariate least squares problem: in, and The discrete projection solution and the compensation target matrix are respectively the first and second parts of the solution. The above objective function is related to... Taking the derivative and setting it to zero, we obtain the column-by-column closed scaling ratio: in, This represents a function that calculates the dot product of two vectors. This is a preset non-negative stable term. Optionally, the non-negative stable term... The value of is in to The scope is set by those skilled in the art according to specific implementation needs. It should be noted that when... When zero is taken and no abnormal branch as described in Example 5 is triggered, the above column-by-column closed scaling ratio expression strictly degenerates into an equivalent closed expression in matrix diagonal form: in, This also refers to the operation of extracting the diagonal elements of a matrix to form a vector, and the division between two vectors is performed element-wise.
[0096] Sub-step 253: Will be by Scaling vector The intermediate dimension scaling vector is absorbed into the continuous scaling vector through element-wise multiplication. middle: It should be noted that the absorption process does not introduce any new dense matrix or new learnable parameters; what would normally be required in the forward mapping... The directional projection of dense matrix multiplication is closed-formed into a vector scaled over the existing intermediate dimension. The channel-wise multiplication. Therefore, the inference deployment computation graph strictly maintains the bi-binary structure decomposition expression defined in step 210. This approach achieves the synergistic effect of one-time compensation during the calibration phase and avoiding the introduction of additional dense parameters and matrix multiplications during the deployment phase. Alternatively, the scaling ratio vector can also be absorbed into the input channel scaling vector through element-wise multiplication. Or output channel scaling vector At least one of them, but this embodiment does not specifically limit this.
[0097] Step 260: Update the dual variable as follows: With this round As the initial layer for the next inner iteration The inner iteration terminates when any of the following conditions are met: the number of inner iterations reaches a preset upper limit; the number of iterations between two adjacent iterations reaches a certain threshold. The relative change is lower than the preset relative tolerance; reconstructed residual The relative tolerance increases relatively over a predetermined number of iterations. Optionally, the upper limit of the inner iteration count is in the range of 5 to 20, and the relative tolerance value is within... to Within this scope, it shall be set by those skilled in the art according to the specific implementation needs.
[0098] After completing steps 220 to 260, update the right factor in a completely symmetric manner. (See Example 6 for details), constituting a complete outer-layer alternation iteration. The outer-layer iteration terminates when any of the following conditions are met: reconstructing the residual. In the outer iterations for a predetermined number of consecutive iterations, the decrease is less than a predetermined threshold; the number of outer iterations reaches a predetermined upper limit; the flip ratio of the discrete symbol matrix relative to the previous round is less than a predetermined threshold. Optionally, the upper limit of the number of outer iterations is in the range of 3 to 50.
[0099] It is particularly important to note that the approximate null space projection operator It is constructed and used only during the calibration phase, and is not used as a model parameter in the inference phase, nor is it passed to the subsequent reconstruction phase as a learnable parameter.
[0100] Example 5: Multi-layer robust degradation compatibility mechanism (Step 200-3) This embodiment further provides a degradation handling mechanism for several abnormal operating conditions during the engineering implementation process in step 200. It should be noted that the degradation handling mechanisms described in this embodiment can be combined independently, and the triggering and processing of each item are driven by discernible conditions; through the following mechanism, this scheme strictly degrades to the "uncompensated" baseline in the worst case, rather than introducing numerical collapse.
[0101] Case 1 (approximate null space does not exist): When the orthogonal basis matrix determined in substep 242... When it is an empty set, that is, under the current numerical precision The overall rank is full, and there are no feature directions that satisfy the threshold condition. At this point, the construction of the approximate null space projection operator in sub-steps 243 and 251 is skipped, and the compensation target matrix is set to the continuous proxy variable, i.e. Only the closed scaling absorption of sub-steps 252 and 253 is performed. In this case, the process of this embodiment smoothly degenerates into a baseline scheme based on ADMM and SVID projection, at which point the forward propagation error still accounts for a certain percentage. Given the magnitude of the problem, this approach does not produce worse results than the baseline approach in this scenario.
[0102] Scenario 2 (Energy of the discrete factor column vector after scaling is too low): When the energy of the corresponding column vector of the discrete projection solution described in sub-step 252 is too low... When the energy level is below the preset column energy limit, the column-by-column closed scaling ratio is adjusted. Set it to a preset backoff value (e.g., set it to 1), and mark the column as invalid. It should be noted that, due to the discrete projection solution... The column vector elements are In the form of, when or When the overall value of the corresponding element in the column approaches zero, the energy of the scaled column may degenerate; therefore, this discrimination is based on the energy of the scaled column vector, not the pure energy. The norm of the symbol itself.
[0103] Case 3 (Inverted inner product sign): When the inner product described in sub-step 252... When the calculation result is negative, the sign of the result is... Set to the preset backoff value to avoid structural conflicts caused by the inversion of discrete symbols.
[0104] Scenario 4 (Scaling ratio significantly deviates from 1): When the absolute value of the difference When the preset upper limit is exceeded, the... The value is truncated to a preset neighborhood of constant 1 to suppress divergence caused by overcompensation.
[0105] Case 5 (ADMM Inner Iteration Divergence): When the reconstruction residual of the subproblem monotonically increases in a predetermined number of consecutive inner iterations, reset the dual variable. Increase the non-negative penalty coefficient by a preset ratio. The inner iteration is restarted; if the number of restarts exceeds a preset limit, the intermediate dimension is reduced by a preset ratio. And reinitialize the bi-binary structure decomposition of this layer.
[0106] Case 6 (Numerically ill-conditioned autocorrelation matrix eigenvalue decomposition): When When the preset upper limit is exceeded, a ridge term is added to the autocorrelation matrix. The feature decomposition is then re-executed, wherein For the preset floating-point safety items, The coefficients are the preset non-negative ridge terms; at the same time, this layer is marked as an ill-conditioned layer, and the learning rate of the intermediate dimension scaling vector of this layer is scaled separately according to the preset ratio in the subsequent vector-only scaling reconstruction stage.
[0107] Scenario 7 (Alignment Loss Divergence in Reconstruction Phase): When the alignment loss in the subsequent vector-only reconstruction phase monotonically increases within a preset number of steps, according to... Write in sub-step 253 The scaling vector absorption effect is addressed by retaining only the baseline results based on ADMM and SVID projections for subsequent reconstruction, where... This represents the element-wise division function.
[0108] Additionally, optionally, the residual energy ratio threshold In each outer iteration, the residual energy ratio can be used as a basis: Perform feedback adjustment: when If the value remains above a preset validity threshold during a preset number of consecutive outer iterations, the value is reduced by a preset proportion. ;when When the value remains below another preset threshold, the value increases by a preset ratio. The feedback adjustment is performed within a preset value range and does not affect the logical structure of the method steps.
[0109] Through the above seven types of anomaly handling mechanisms and optional feedback adjustment, this embodiment enables the aforementioned scheme to be stably applied to different types of linear projection layers such as query projection layer, key projection layer, value projection layer, output projection layer in self-attention mechanisms, and gated projection layer, up projection layer, and down projection layer in feedforward neural networks, and can adapt to the differences in inter-layer spectral distribution under different model scales.
[0110] Example 6: Symmetry treatment of left and right factors (step 200-4) This embodiment further illustrates the fixing of the left factor in step 200. Update right factor The symmetrical processing flow described in this embodiment is completely symmetrical to the flow of fixing the right factor and updating the left factor described in Embodiments 3 and 4.
[0111] Specifically, this involves introducing a continuous proxy variable with a right factor. The solution is obtained through ADMM subproblems. And obtain the discrete projected solution of the right factor by projecting according to SVID. With the corresponding discrete projection residual Its solution form is completely symmetrical to that of Example 3, and will not be repeated here.
[0112] Furthermore, the autocorrelation matrix described in sub-step 241 of Embodiment 4... Replace with Regarding the above After performing eigenvalue decomposition, construct the corresponding approximate null space projection operator according to any of the selection methods described in sub-step 242. .
[0113] Construct the compensation objective matrix for the right factor: Solve for the discrete projection solution of the right factor row by row. The scaling ratio of the corresponding row of the compensation target matrix of the row vector least squares approximation of the right factor. The closed-form solution is: in, and The discrete projection solution of the right factor and the compensation target matrix of the right factor are respectively the first and second solutions of the right factor. OK, This is a pre-defined non-negative stable term. It will be determined by... The resulting scaling vector is absorbed into the intermediate dimension scaling vector contained in the continuous scaling vector through element-wise multiplication. Or input channel scaling vector One of them, the specific target of absorption is set by those skilled in the art in the specific implementation according to the actual deployment constraints.
[0114] By symmetrically performing the approximate null space projection and the closed scaling ratio absorption using left and right factors, this embodiment can structure the discrete projection residuals on both sides of the bibinary structural decomposition, thereby improving the overall stability of the bibinary structural decomposition.
[0115] Example 7: Lightweight Global Reconstruction of a Frozen Binary Symbol Matrix (Step 300) This embodiment further illustrates the lightweight global reconstruction process in step 300, which uses only continuously scaling vectors as learnable parameters.
[0116] After step 200 is completed, the two binary symbol matrices in each linear projection layer to be quantized and This has been determined. At this point, each linear projection layer retains only three full-precision continuous scaling vectors for subsequent optimization, namely the output channel scaling vectors. , intermediate dimension scaling vector and input channel scaling vector ,in This represents the layer index of the linear projection layer. Let the first layer be... The equivalent weights of each linear projection layer in the reconstruction stage are: Let the set of all learnable, continuously scaled vectors be denoted as . In the calibration data set Above, the alignment target is generated using the original full-precision model (i.e., the unquantized pre-trained model), all binary sign matrices in the quantized model are frozen, and only the set is optimized. This makes the logits of the quantized model approximate the logits of the original full-precision model. It should be noted that the alignment target is generated by the original full-precision model (i.e., the unquantized pre-trained model) and does not require a separate, separately trained teacher network.
[0117] Specifically, the difference between the output distribution of the reduced quantization model and the original full-precision model output distribution can be achieved using at least one of Kullback-Leibler divergence (hereinafter referred to as KL divergence), cross-entropy, mean squared error, or contrast loss as the alignment loss function. Optionally, when using KL divergence as the alignment loss function, the alignment loss is: in, For the calibration data set The sample, For the original full-precision model in the sample The output logits, For the quantization model in the sample The above, with the set The output logits when the continuously scaled vector in the vector is used as a learnable parameter. Represents the normalized exponential function, Denotes the KL divergence function. This represents the desired function over the calibration dataset.
[0118] It should be noted that the alignment loss function is not limited to KL divergence. In an optional implementation, the KL divergence can be replaced by any one of cross-entropy, mean squared error, contrastive loss, Hinge loss, inverse KL divergence, symmetric KL divergence, or Jensen-Shannon divergence, or a weighted combination thereof; in another optional implementation, a temperature coefficient can be introduced into the output logits before applying the softmax function to adjust the alignment process's attention to low-probability classes.
[0119] More specifically, since this stage does not require unfreezing any binary sign matrices, nor does it require introducing full-precision shadow weights or additional surrogate weight matrices, the number of trainable parameters per layer is only [amount missing]. Relative to the thawed complete two-factor matrix The magnitude is significantly reduced. This design makes it possible to perform quantization and reconstruction of large-scale neural network models in memory-constrained environments.
[0120] Example 8: Deployment Model Export and Inference Computation Graph (Step 400) This embodiment further illustrates the specific implementation of the deployment model export stage and the inference stage in step 400. After completing the lightweight global reconstruction in step 300, only the binary symbol matrix and the continuous scaling vector are retained when exporting the deployment model.
[0121] Specifically, each element of the binary symbol matrix is represented by 1 bit and stored as an integer in bit-packed form with a preset bit width. In one embodiment, the binary symbol matrix is stored as an unsigned integer in row-major order with a preset bit width, and the elements... and These are mapped to 0 and 1 of the corresponding bits, respectively. Those skilled in the art can select an appropriate packing bit width based on the word width and alignment requirements of the specific hardware platform; this application does not impose any restrictions on this. The continuously scaled vector is represented in floating-point and stored in conjunction with the binary symbol matrix in the form of channel-by-channel continuous parameters. Optionally, the continuously scaled vector can adopt single-precision floating-point, half-precision floating-point, or half-precision floating-point formats, which those skilled in the art can select based on the data type support of the specific inference hardware.
[0122] Furthermore, the quantization model does not introduce any additional dense matrix or additional dense matrix multiplication operations besides the binary sign matrix and the continuous scaling vector during the inference phase. In other words, the linear projection layer calculation during the inference phase is strictly performed according to the bi-binary structure form shown in the equivalent weight expression defined in Embodiment 7. Specifically, it replaces the dense floating-point matrix multiplication used by the original linear projection layer by combining bitwise operations (including but not limited to bitwise XOR, bit counting, etc.) with channel-by-channel floating-point multiplication.
[0123] It should be particularly noted that the approximate null space projection operator or It is not retained or called as a model parameter during the inference phase, nor does it participate in any dense matrix multiplication operations during inference; its entire role in the calibration phase has been completely folded into the intermediate dimension scaling vector through the above-described column-by-column closed scaling ratio solution and intermediate dimension scaling vector absorption process. Therefore, the inference computation graph of the final deployed model is strictly consistent with the inference computation graph without zero-space compensation. Thus, while introducing a directional residual control mechanism, this application strictly maintains the hardware deployment friendliness of the quantization scheme after extremely low bit training. The quantization model is used to replace the original floating-point matrix multiplication with bit operations and channel-by-channel multiplication on the inference device, which can reduce the memory usage and memory access bandwidth requirements of the inference device.
[0124] It should be further noted that the post-training quantization method only performs quantization processing on the linear projection layers in the neural network model, and does not perform post-training quantization processing on layer normalization layers, root mean square normalization layers, activation function layers, etc. in the neural network model. This method is applicable to linear projection layers in various neural network model architectures, including but not limited to Transformer architecture, hybrid expert architecture, state-space model architecture, and visual Transformer architecture.
[0125] Example 9: Computing Device Example This embodiment provides a computing device for implementing the aforementioned method. The computing device includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores the binary symbol matrix and the continuously scaled vector, and also stores program instructions executable by the at least one processor. When executed by the at least one processor, the program instructions cause the computing device to perform the post-training quantization method described in Embodiments 1 to 8 during the calibration phase.
[0126] Specifically, the computing device may further include a function for accessing the calibration data set during the calibration phase. The system includes a data interface, a model interface for communicating with the original full-precision model, and an inference interface for providing quantized model inference services during the inference phase. During the calibration phase, the at least one processor performs the following operations: completes forward and backward propagation based on the calibration data, thereby constructing the surrogate weight matrix; performs the bibinaristic structure decomposition on the surrogate weight matrix; constructs the approximate null space projection operator based on the autocorrelation matrix with a fixed one-sided factor, and completes the construction of the compensation target matrix and the solution and absorption of the closed column (or row) scaling ratio; freezes the binary sign matrix, and completes alignment reconstruction using only the continuous scaling vector as a learnable parameter.
[0127] More specifically, after the calibration phase is completed, the quantization model in the memory for the inference phase only retains the binary sign matrix and the channel-by-channel continuous scaling vector, and does not include the full-precision surrogate weight matrix and the explicit approximate null space projection operator. During the inference phase, the at least one processor replaces the original floating-point matrix multiplication with bitwise operations and channel-by-channel multiplication to perform the forward computation of the quantization model. Optionally, the computing device can be a single server, a distributed server cluster, a single graphics processing unit, a collaborative node of multiple graphics processing units, an edge computing node, a dedicated neural network accelerator, or other computing devices with sufficient computing power. This embodiment does not limit the specific hardware implementation of the computing device.
[0128] Example 10: Computer-readable storage medium and computer program product examples This embodiment provides a computer-readable storage medium storing computer program instructions. When executed by a processor, the computer program instructions implement the post-training quantization method described in Embodiments 1 to 8 during the calibration phase; and ensure that after executing the method, the derived quantization model retains only the binary symbol matrix and the channel-by-channel continuous scaling vector for inference phase calls, without explicitly retaining the approximate null space projection operator.
[0129] This embodiment also provides a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, they implement the post-training quantization method described in Embodiments 1 to 8.
[0130] It should be noted that the computer-readable storage medium includes, but is not limited to, any one or more combinations of volatile storage media, non-volatile storage media, magnetic storage media, optical storage media, and semiconductor storage media, such as hard disks, solid-state drives, optical disks, flash memory, random access memory, etc.; the computer program product may be distributed in the form of source code, object code, executable files, installation packages, container images, software development kits, or other forms.
[0131] Example 11: Description of some steps of the method that can be performed independently It should be further noted that some steps of the method described in this application can be implemented independently, and it is not necessarily required that all three stages be executed. Specifically, the surrogate weight matrix construction based on second-order sensitivity described in Example 2 can be used independently for the preconditioning stage in other two-factor decomposition methods or other post-training quantization schemes; the approximate null space projection and closed column scaling absorption based on fixed factor autocorrelation spectrum described in Example 4 can be used independently for residual suppression in other discrete projection scenarios; the lightweight global reconstruction of the frozen binary sign matrix and using only the continuous scaling vector as a learnable parameter described in Example 7 can be used independently for post-calibration of other quantization models where the sign matrix has been determined.
[0132] Furthermore, in some optional implementations, only steps 100 and 200 may be performed, omitting the lightweight global reconstruction in step 300, and the deployment model may be directly derived from the binary symbolic matrix and continuously scaling vector obtained in step 200; in other optional implementations, only steps 200 and 300 may be performed. The various phased implementation methods described in this embodiment do not constitute a limitation on the scope of protection of this application.
[0133] Example of Supplementary Details for Implementing the Project (Example XII) To facilitate complete reproduction of the method described in this application by those skilled in the art, this embodiment provides several engineering implementation details by way of example. It should be noted that the numerical ranges and specific configurations described below are merely illustrative examples and do not constitute a limitation on the scope of protection of this application. Those skilled in the art can make adjustments in specific implementations according to the calibration target, target hardware platform, and specific application scenario.
[0134] Regarding the size and sequence length of the calibration data set: Optional, the calibration data set The number of samples can be selected within a reasonable range according to the calibration target, and the sequence length of each sample can also be selected within a reasonable range according to the maximum context length of the neural network model. Several hundred samples (e.g., 128 to 1024 samples) are randomly sampled from the target domain general corpus. Each sample is truncated or concatenated to a sequence length of several thousand tokens (e.g., 512 to 2048 tokens) as the calibration data set. During the sampling process, deduplication is performed at the document level to avoid duplicate samples from biasing the statistics. The data is then divided into an input activation statistics collection subset, an output gradient statistics collection subset, and a lightweight global reconstruction subset according to a preset ratio.
[0135] Regarding ADMM penalty coefficient With iteration termination condition: Optional, the non-negative penalty coefficient The number of iterations in the inner layer, the number of alternating iterations in the outer layer, and the convergence tolerance of the ADMM are selected within a reasonable range; these are set by those skilled in the art based on the specific implementation objectives. exist to The value is taken within the order of magnitude, with the initial value set to approximately 1.0. During inner-layer iterative divergence, the value is adaptively increased proportionally from 1.5 to 2.0. For the query projection layer and key projection layer in the self-attention mechanism, which have larger dimensions... The initial value can be chosen to be relatively small, between 0.1 and 0.5, to avoid the proximal term excessively suppressing the data fidelity term; the upper limit for the number of inner iterations is set to approximately 5 to 20, and the upper limit for the number of outer alternating iterations is set to approximately 3 to 50. The relative change in the reconstructed residual between two adjacent outer iterations is less than approximately It may be terminated early.
[0136] Regarding the residual energy ratio threshold With non-negative stable terms Optionally, the residual energy is higher than the threshold. exist A reasonable range within the interval is selected; the non-negative stable term Select within a reasonable small range. exist to The value is taken within a range, and the initial value is set to approximately And based on the residual energy ratio obtained from the calibration period for each layer. Perform feedback adjustment, when consistently higher proportionally Decrease ,when Persistently below proportionally Increase ; Take contract Non-negative stable term exist to Values are taken within a range, and when continuously scaled vectors are stored using single-precision floating-point, they are set to approximately [value range missing]. When using half-precision floating-point storage, the value should be appropriately increased to approximately [value missing]. To avoid underflow on small orders of magnitude; eigenvalue threshold Depend on Given, among which exist to Range of values.
[0137] Regarding the intermediate dimension Selection: Optional, the intermediate dimension Press Take values, where The compression ratio is set to a preset value; or it is set according to the proxy weight matrix. The singular value decomposition is truncated, preserving the smallest integer selection of the preset energy ratio. For the query, key, value, and output projection layers in the self-attention mechanism, the compression ratio coefficient is... Take approximately 1 / 8 to 1 / 4; for the gating, upper, and lower projection layers in a feedforward neural network, the compression ratio coefficient is... Take approximately 1 / 16 to 1 / 8; if the energy ratio preserved by truncated singular value decomposition within the same layer is difficult to achieve the preset target (e.g., approximately 95%), then increase accordingly. .
[0138] Regarding the algorithm for solving the first singular triplet in SVID projection: Optionally, the first singular triplet can be obtained through truncated singular value decomposition, power iteration, Lanczos iteration, or random singular value decomposition. When using the power iteration method to solve the first singular triplet, the initial vector is initialized according to a Gaussian random distribution, and the upper limit of the number of iterations is approximately 50. When the relative change of singular values between two adjacent iterations is less than approximately... The process terminates early; to improve numerical stability, Gram-Schmidt orthogonalization is performed on the intermediate vector after each iteration.
[0139] Regarding the optimizer and hyperparameters in step 300: Optionally, an optimizer based on adaptive moment estimation, such as AdamW, can be used in step 300, with the learning rate, number of training steps, and temperature coefficient selected within a reasonable range. When using the AdamW optimizer, the learning rate is within... to Select within the range, with the initial learning rate set to approximately The cosine annealing learning rate scheduling strategy is adopted; the gradient clipping norm threshold is set to approximately 1.0; the total number of training steps is several hundred to several thousand steps (e.g., 200 to 2000 steps), with the first few dozen steps being a linear warm-up stage; the temperature coefficient is set to approximately 1.0 to 2.0; the output channel scaling vector, the intermediate dimension scaling vector, and the input channel scaling vector all use the same learning rate.
[0140] Regarding the bit-packing format of the binary symbol matrix: Optionally, the binary symbol matrix is packaged into an unsigned integer for each preset bit width in row-major order; when the number of columns in the matrix is not an integer multiple of the packaged bit width, zeros can be padded at the end and the valid column number can be recorded. On the target GPU platform, the binary symbol matrix is packaged into a uint32 integer for each 32 bits in row-major order, and the elements... Mapped to bit 0 Mapped to bit 1; in another exemplary implementation, it is packed in 64 bits to fit the bit operation unit of a 64-bit register; by combining the popcount bit counting instruction with the bit XOR instruction, multi-path parallel binary multiplication and addition operations are performed within each word, and then fused with the channel-by-channel floating-point scaling vector to complete the output calculation of the linear projection layer.
[0141] Regarding the anomaly handling threshold: optionally, the column energy lower limit, scaling ratio truncation neighborhood, ADMM restart count upper limit, and ridge term coefficient. The threshold values for anomaly handling, such as the upper limit of the condition number, shall be set by those skilled in the art based on the robustness requirements of the specific implementation. The lower limit of column energy is set to... ,For example The scaling factor truncates the neighborhood to a symmetrical interval around 1, for example... The maximum number of ADMM restarts can be set to a certain number, for example, 3 times; Ridge term coefficient. Take values within a small range, for example Interval; the upper limit of the condition number takes a larger order of magnitude, for example... .
[0142] Regarding the expansion of applicable models and layer types: This method can be adapted to large language models with Transformer architecture, large language models with hybrid expert architecture, and linear projection layers of visual Transformer models; for query, key, value, and output projection layers in self-attention mechanisms, and gating, up, and down projection layers in feedforward neural networks, the same three-stage process is used for quantization; the original floating-point precision is maintained for all normalization layers, embedding layers, output logits projection layers (i.e., language model heads), and residual connections to avoid excessive perturbation to the model output distribution.
[0143] Regarding the ablation experiment design: For example, the following four schemes are constructed for ablation comparison: Scheme A, fully executing the method described in this application; Scheme B, removing the approximate null space projection operator. (i.e., directly place) The rest of the process remains unchanged; Option C retains the approximate null space projection operator. However, Scheme D removes the closed column scaling absorption while keeping the rest of the process unchanged; Scheme D removes the degradation handling mechanism for seven types of abnormal operating conditions while keeping the rest of the process unchanged. The four schemes are evaluated on the calibration dataset, and the propagation error norm is recorded. Reconstructing residuals By comparing the accuracy of downstream tasks, the actual impact of each module can be assessed.
[0144] Regarding end-to-end measured performance data: On a large language model with a certain parameter scale, after quantizing the weights of all linear projection layers to an equivalent bit width close to 1 bit using the method described in this application, the perplexity of the quantized model on the general language modeling benchmark is reduced compared to the baseline scheme without the approximate null space projection and closed scaling absorption, and the zero-sample accuracy on downstream tasks of commonsense reasoning is improved compared to the baseline scheme; at the same time, the deployment storage footprint of the quantized model is significantly reduced compared to the full-precision model, the learnable parameter scale in the calibration phase is significantly reduced compared to the control scheme of unfrozen complete two-factor matrices, and no additional dense matrix multiplication operations are introduced in the inference phase. The specific numerical data were obtained by the inventors through actual measurement during implementation, and this specification does not make any specific commitments regarding the numerical values.
[0145] The second embodiment of this application relates to a post-training quantization system for deploying large language model inference, oriented towards a neural network model containing a linear projection layer to be quantized, the system comprising: The proxy weight matrix construction module is configured to obtain a tensor with physical meaning obtained by forward and backward propagation of calibration data through the neural network model, which serves as input activation statistics and output gradient statistics, and to construct a proxy weight matrix containing second-order sensitivity information based on the input activation statistics and the output gradient statistics. The bibinary structure decomposition module is configured to perform bibinary structure decomposition on the proxy weight matrix output by the proxy weight matrix construction module, and represent the proxy weight matrix as a product of two binary symbol matrices and several continuous scaling vectors. The alternating optimization solution module is configured to solve the bibinary structure decomposition by alternating optimization when the bibinary structure decomposition module performs the bibinary structure decomposition. When fixing one side factor and updating the other side factor, the continuous surrogate variable of the other side factor is obtained, and the continuous surrogate variable is discretized onto the bibinary structure domain to obtain the discrete projection solution. The discrete projection residual compensation module is configured to perform the following operations during the processing of discrete projection residuals generated by discrete projection of the continuous surrogate variables onto the bibinary structural domain: constructing an approximate null space projection operator based on the autocorrelation matrix of the fixed one-sided factor; constructing a compensation target matrix using the continuous surrogate variables, the discrete projection residuals, and the approximate null space projection operator; solving for the scaling ratios of the column vectors of the discrete projection solution that approximate the corresponding columns of the compensation target matrix by least squares, forming a scaling ratio vector; and absorbing the scaling ratio vector into the continuous scaling vector corresponding to the discrete projection solution to suppress the propagation error of the discrete projection residuals under the forward mapping of the fixed one-sided factor. The reconstruction optimization module is configured to freeze the two binary symbol matrices obtained by the binary structure decomposition module, and only use each continuous scaling vector as a learnable parameter for optimization in the reconstruction stage, so as to reduce the difference between the output distribution of the quantized model and the output distribution of the original full-precision model. The reconstruction optimization module does not unfreeze the binary symbol matrix or introduce additional full-precision proxy weight matrix during the reconstruction phase, and the system does not explicitly retain the approximate null space projection operator during the inference phase; the quantization model derived by the system is used to replace the original floating-point matrix multiplication with bit operations and channel-by-channel multiplication on the inference device.
[0146] The embodiments described above are method implementations corresponding to this implementation. The technical details in the embodiments described above can be applied to this implementation, and the technical details in this implementation can also be applied to the embodiments described above.
[0147] It should be noted that those skilled in the art should understand that the implementation functions of each module shown in the above-described embodiments of the post-training quantization system for large language model inference deployment can be understood with reference to the relevant descriptions of the post-training quantization method for large language model inference deployment. The functions of each module shown in the above-described embodiments of the post-training quantization system for large language model inference deployment can be implemented by a program (executable instructions) running on a processor, or by specific logic circuits. If the post-training quantization system for large language model inference deployment described in this application is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application embodiment, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any particular combination of hardware and software.
[0148] Accordingly, this application also provides a computer storage medium storing computer-executable instructions, which, when executed by a processor, implement the various method implementations of this application.
[0149] Furthermore, this application also provides a post-training quantization system for deploying large language model inference, including a memory for storing computer-executable instructions and a processor; the processor is used to implement the steps in the above-described method embodiments when executing the computer-executable instructions in the memory. The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The aforementioned memory can be read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or solid-state drive, etc. The steps of the methods disclosed in the various embodiments of this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0150] The above embodiments have the following technical effects: First, by constructing the surrogate weight matrix based on second-order sensitivity as described in Example 2, the second-order approximation of the task loss on the weight perturbation is transformed into a standard Frobenius norm minimization problem in the surrogate space through the Kronecker factor approximation of the Hessian matrix and diagonal energy preservation. This allows the subsequent bibinary structure decomposition to be performed in the surrogate space aligned with the task loss, which helps to prioritize the reconstruction accuracy of the key weight region. Compared with the control scheme that does not introduce second-order sensitivity information, this helps to alleviate the quantization degradation of key channels under extremely low bit conditions.
[0151] Second, by constructing the approximate null space projection operator based on the eigenvalue decomposition of the autocorrelation matrix of a fixed one-sided factor as described in Examples 3 and 4, the two objectives of "minimizing residuals" and "minimizing propagation errors" are explicitly distinguished. Furthermore, by constructing the compensation target matrix and solving the closed column scaling ratio, the directional projection that originally required dense matrix multiplication is folded into a channel-by-channel multiplication of the existing continuous scaling vector. This allows the approximate null space projection operator to be constructed and used only during the calibration phase. The computational graph of the quantization model during the inference phase remains consistent with that without the approximate null space projection operator, thus avoiding the introduction of additional dense parameter matrices and additional dense matrix multiplication operations during the inference phase.
[0152] Third, the two optional rank selection methods—adaptive rank selection based on the residual energy ratio threshold and rank selection based on the absolute threshold—ensure that the discrete projective residual energy allowed to propagate along the forward mapping is subject to a formalized absolute upper bound. Explicit constraints. The threshold and the total energy of the fixed factor in this upper bound can both be estimated a priori from the calibration statistics, enabling prior determination of the error budget of the final deployed model during the calibration phase without relying on end-to-end measured feedback.
[0153] Fourth, by using the left and right factor symmetry processing described in Example 6, the discrete projection residuals on both sides of the bibinary decomposition can be structured, thereby improving the overall stability of the bibinary structural decomposition.
[0154] Fifth, through the degradation processing mechanism for the seven types of abnormal working conditions described in Example 5, including degradation when the near null space does not exist, regression when the energy of the scaled discrete factor column vector is too low, regression when the inner product sign is reversed, truncation when the scaling ratio deviates significantly from 1, restart when the inner layer of ADMM diverges, ridge term stabilization when the autocorrelation matrix eigenvalue decomposition is numerically ill-conditioned, and regression of absorption effect when the alignment loss diverges during the reconstruction stage, this scheme strictly degrades to an "uncompensated" baseline in the worst case rather than introducing numerical collapse. This helps to stably apply this scheme to different types of linear projection layers such as query projection layer, key projection layer, value projection layer, output projection layer in self-attention mechanisms, and gated projection layer, upper projection layer, and lower projection layer in feedforward neural networks.
[0155] Sixth, through the lightweight global reconstruction described in Example 7, which freezes the binary symbol matrix and uses only continuously scaled vectors as learnable parameters, the scale of trainable parameters for each layer is reduced from that corresponding to a complete two-factor matrix. The magnitude dropped significantly to Compared to the approach that requires unfreezing the complete weight matrix or introducing full-precision shadow weights, this scale helps reduce memory usage and training complexity during the reconstruction phase, making it possible to perform quantization and reconstruction on large language models in memory-constrained environments.
[0156] Seventh, the quantization model derived from the bit-packed storage and channel-by-channel multiplication inference computation graph described in Example 8 only contains a binary symbol matrix in bit-packed form and a continuous scaling vector for each channel. It does not contain a full-precision surrogate weight matrix and an explicit approximate null space projection operator, which helps to reduce the deployment storage and memory access bandwidth requirements. It is suitable for scenarios with limited video memory, limited bandwidth, or strict requirements on model throughput and deployment costs, including but not limited to high-concurrency inference in the cloud, edge device deployment, and inference deployment in resource-constrained environments.
[0157] It should be noted that among the above effects, "formal error upper bound" and "parameter scale" are different. relatively The "decline" is an effect that can be self-derived from the formula, while the end-to-end performance of the final deployed model (such as downstream task accuracy, perplexity, inference latency, etc.) can be verified through actual tests during specific implementation. This specification does not make any specific commitments to the measured values.
[0158] Furthermore, the aforementioned technical effects are achieved synergistically under the strong constraints of "quantization after training, not unfreezing the complete weights, and keeping the inference computation graph unchanged." The implementation of any single point can be regarded as a local contribution of this application, but the seven effects are presented synergistically as a whole, constituting the overall technical contribution of this application relative to existing post-training quantization schemes.
[0159] It should be noted that in this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this application, if a reference is made to performing an action based on an element, it means performing the action at least based on that element, including two cases: performing the action only based on that element, and performing the action based on that element and other elements. Expressions such as "multiple," "repeatedly," and "various" include two, two times, two kinds, and more than two, more than two times, and more than two kinds.
[0160] Furthermore, it should be understood that after reading the above disclosure of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope of protection claimed in this application.
Claims
1. A post-training quantization method for large language model inference deployment, characterized in that, For a neural network model containing a linear projection layer to be quantized, the method performs the following steps sequentially: Tensors with physical meaning obtained by forward and backward propagation of calibration data through the neural network model are used as input activation statistics and output gradient statistics. Based on the input activation statistics and the output gradient statistics, a proxy weight matrix containing second-order sensitivity information is constructed. Perform a bi-binary structure decomposition on the surrogate weight matrix, and represent the surrogate weight matrix as a product of two binary sign matrices and several consecutive scaling vectors; When performing the bibinary structural decomposition, the solution is obtained by alternating optimization. When fixing one side factor and updating the other side factor, the continuous surrogate variable of the other side factor is obtained, and the continuous surrogate variable is discretized projected onto the bibinary structural domain to obtain the discrete projection solution. Furthermore, in the process of processing the discrete projection residuals generated by the discrete projection of the continuous surrogate variables onto the bi-binary structural domain: an approximate null space projection operator is constructed based on the autocorrelation matrix of the fixed one-sided factor; a compensation target matrix is constructed using the continuous surrogate variables, the discrete projection residuals, and the approximate null space projection operator; the scaling ratios of the column vectors of the discrete projection solutions are solved column by column to make the column vectors of the compensation target matrix approximate each other by least squares, forming a scaling ratio vector; and the scaling ratio vector is absorbed into the continuous scaling vector corresponding to the discrete projection solution to suppress the propagation error of the discrete projection residuals under the forward mapping of the fixed one-sided factor. The two binary symbol matrices obtained from the bibinary structure decomposition are frozen, and only each continuous scaling vector is used as a learnable parameter for optimization in the reconstruction stage, so as to reduce the difference between the output distribution of the quantized model and the output distribution of the original full-precision model. In the reconstruction phase, the binary symbol matrix is not unfrozen, no additional full-precision proxy weight matrix is introduced, and the approximate null space projection operator is not explicitly retained in the inference phase; the derived quantization model is used to replace the original floating-point matrix multiplication with bit operations and channel-by-channel multiplication on the inference device.
2. The method as described in claim 1, characterized in that, The linear projection layer to be quantized includes at least one of the following layers: the query projection layer, key projection layer, value projection layer, and output projection layer of the self-attention mechanism in the neural network model; and the gated projection layer, upper projection layer, and lower projection layer of the feedforward neural network in the neural network model.
3. The method as described in claim 1, characterized in that, The construction of the proxy weight matrix containing second-order sensitivity information based on the input activation statistics and the output gradient statistics includes: Calculate the square root of the squared mean of the input activations in the input activation statistics to form the input importance vector. ,in , The first activation vector of the input activation vector of the linear projection layer to be quantized is... One element, This represents the sample expectation function on the calibration data. The input importance vector The One element; Calculate the square root of the mean of the squared output gradients in the output gradient statistics to form the output importance vector. ,in , The gradient vector of the output of the linear projection layer to be quantized with respect to the task loss is the first... One element, The output importance vector The One element; Using the input importance vector With the output importance vector For the original weight matrix The proxy weight matrix is obtained by performing a two-sided weighting. ,in This represents the element-wise multiplication function with broadcast capability; The proxy weight matrix is based on the Hessian matrix. The Kronecker factor approximation is constructed for the Hessian matrix. It is decomposed into input statistics and output gradient statistics using the Kronecker factor approximation: ,in Let be the input activation vector of the linear projection layer to be quantized. Let be the gradient vector of the output of the linear projection layer to be quantized with respect to the task loss. Represent the Kronecker product function; and retain only the diagonal energy information of the input statistics and output gradient statistics in the Kronecker factor approximation, taking the square root of each element to construct the input importance vector. With the output importance vector .
4. The method as described in claim 1, characterized in that, The proxy weight matrix is represented as a product of two binary symbolic matrices and several consecutive scaling vectors, specifically as follows: in, and To assign a value to an element A binary symbol matrix, , , These are continuous scaling vectors, representing the output channel scaling vector, intermediate dimension scaling vector, and input channel scaling vector, respectively. Represents the element-wise multiplication function with broadcast; the continuous scaling vector corresponding to the discrete projection solution is taken from... , , At least one of them.
5. The method as described in claim 1, characterized in that, The method of solving by alternating optimization includes: With one side factor fixed as the right factor And update the other factor to the left factor. At that time, the continuous proxy variable The following subproblem based on the alternating direction multiplier method is obtained by solving: The subproblem has a closed-ended solution: in, For the first Discrete left factor in the next inner iteration. For the corresponding dual variable, The default non-negative penalty coefficient is... This is the intermediate dimension of the bibinary structure decomposition. for An identity matrix of order 1. This represents a function that finds the variables that minimize the objective function. The continuous proxy variable Discrete projection is performed onto the bibinarre structural domain to obtain the discrete projection solution, including performing symbol-value independent decomposition projection: taking the symbol matrix. Its element values belong to , A function that takes the sign of each element; for absolute value matrices Find the rank-1 outer product approximation by its first singular triplet. get and ,in , , They are respectively The maximum singular value, the corresponding left singular vector, and the corresponding right singular vector; denoted as the discrete projection solution is... And let the discrete projection residual be denoted as .
6. The method as described in claim 1, characterized in that, The construction of the approximate null space projection operator based on the autocorrelation matrix of the fixed one-sided factor includes: Calculate the autocorrelation matrix of the fixed one-sided factor; Perform eigenvalue decomposition on the autocorrelation matrix to obtain multiple eigenvalues and their corresponding eigenvectors; Based on the multiple eigenvalues, eigenvectors corresponding to low-energy directions are selected from the eigenvectors to form an orthogonal basis matrix, and the product of the orthogonal basis matrix and its transpose is used as the approximate null space projection operator. The selection method includes at least one of the following: Method 1: Arrange the multiple feature values in descending order and determine the smallest integer that satisfies the following formula. : in, The first in descending order 1 eigenvalue, This is the intermediate dimension of the bibinary structure decomposition. For preset belonging The residual energy ratio threshold of the interval will be related to The corresponding eigenvectors form the orthogonal basis matrix; Method 2: The orthogonal basis matrix is formed by eigenvectors corresponding to eigenvalues that are not greater than a preset eigenvalue threshold among the plurality of eigenvalues; When the eigenvalue decomposition of the autocorrelation matrix exhibits an ill-conditioned condition where the condition number exceeds a preset upper limit, the eigenvalue decomposition is re-executed after adding a ridge term to the autocorrelation matrix. The ridge term is the product of preset coefficients and the identity matrix.
7. The method as described in claim 6, characterized in that: When one side of the fixed factor is the right factor and the other side of the fixed factor is the left factor, the continuous proxy variable of the left factor is used. Discrete projection solution with left factor Calculate the discrete projection residual According to the discrete projection residual and the approximate null space projection operator Construct the compensation target matrix ; For the discrete projection solution Each column Solve the column-by-column closed scaling ratio using the following formula. : in, This represents a function that calculates the dot product of two vectors. Represents the computation of vectors Functions of norm These are pre-defined non-negative stable terms; The column-by-column closed scaling ratio will be used to determine the scaling ratio of each column. Scaling vector ,according to The intermediate dimension scaling vector is absorbed into the continuous scaling vector through element-wise multiplication. middle; In solving the column-by-column closed scaling ratio Furthermore, it includes at least one of the following independently combinable anomaly handling mechanisms: when the energy of the corresponding column vector of the discrete projection solution... When the energy level is below the preset lower limit, the... Set to a preset backoff value; when the result of the inner product calculation is negative, set the... Set as the preset backoff value; when the absolute value of the difference... When the preset upper limit is exceeded, the... Cut off to a preset neighborhood of constant 1.
8. The method as described in claim 6, characterized in that, Further includes: Degradation processing mechanism: When the autocorrelation matrix is subjected to eigenvalue decomposition and the orthogonal basis matrix obtained by the preset rule is an empty set, the construction of the approximate null space projection operator is skipped, and the compensation target matrix is set as the continuous proxy variable. The steps of solving the scaling ratio column by column to form the scaling ratio vector and absorbing the scaling ratio vector into the continuous scaling vector corresponding to the discrete projection solution are continued. The symmetric processing steps for the other factor in the bibinary structural decomposition are as follows: With the left factor of the bibinary structural decomposition fixed... And update the right factor In the case of autocorrelation matrix based on left factor The eigenvalue decomposition constructs the corresponding approximate null space projection operator. Based on the continuous proxy variable of the right factor Discrete projection solution of right factor and the corresponding approximate null space projection operator Construct the compensation objective matrix of the right factor. ,in The scaling ratios of the row vectors of the discrete projection solutions of the right factor are calculated row by row so that the row vectors of the right factor approximate the corresponding rows of the compensation target matrix of the right factor by least squares. The corresponding scaling ratio vectors are then absorbed into the continuous scaling vectors by element-wise multiplication.
9. The method as described in claim 1, characterized in that: In the quantization model, each element of the binary symbol matrix is represented by 1 bit and stored as an integer in a bit-packed manner according to a preset bit width. The continuous scaling vector is represented in floating point and stored in conjunction with the binary symbol matrix in a channel-by-channel continuous parameter form. The quantization model does not introduce any additional dense matrix or additional dense matrix multiplication operations other than the binary symbol matrix and the continuous scaling vector during the inference phase. The difference between the reduced quantization model output distribution and the original full-precision model output distribution includes using the output distribution of the original full-precision model on the calibration data as the alignment target, and using at least one of Kullback-Leibler divergence, cross-entropy, mean squared error, or contrast loss as the alignment loss function.
10. A computing device, characterized in that, include: At least one processor; A memory communicatively connected to the at least one processor, the memory storing the binary symbol matrix and the continuously scaled vector, and storing program instructions executable by the at least one processor; When the program instructions are executed by the at least one processor, the computing device performs the method as described in any one of claims 1 to 9 during the calibration phase; and, after the calibration phase is completed, the quantization model in the memory for the inference phase retains only the binary symbol matrix and the channel-by-channel continuous scaling vector, and does not contain the full-precision surrogate weight matrix and the explicit approximate null space projection operator.
11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they enable the execution of the method described in any one of claims 1 to 9 during the calibration phase; and ensure that, after the execution of the method, the derived quantization model retains only the binary symbol matrix and the channel-by-channel continuous scaling vector for inference phase calls, without explicitly retaining the approximate null space projection operator.
12. A post-training quantization system for deploying large language model inference, characterized in that, For a neural network model containing a linear projection layer to be quantized, the system includes: The proxy weight matrix construction module is configured to obtain a tensor with physical meaning obtained by forward and backward propagation of calibration data through the neural network model, which serves as input activation statistics and output gradient statistics, and to construct a proxy weight matrix containing second-order sensitivity information based on the input activation statistics and the output gradient statistics. The bibinary structure decomposition module is configured to perform bibinary structure decomposition on the proxy weight matrix output by the proxy weight matrix construction module, and represent the proxy weight matrix as a product of two binary symbol matrices and several continuous scaling vectors. The alternating optimization solution module is configured to solve the bibinary structure decomposition by alternating optimization when the bibinary structure decomposition module performs the bibinary structure decomposition. When fixing one side factor and updating the other side factor, the continuous surrogate variable of the other side factor is obtained, and the continuous surrogate variable is discretized onto the bibinary structure domain to obtain the discrete projection solution. The discrete projection residual compensation module is configured to perform the following operations during the processing of discrete projection residuals generated by discrete projection of the continuous surrogate variables onto the bibinary structural domain: constructing an approximate null space projection operator based on the autocorrelation matrix of the fixed one-sided factor; constructing a compensation target matrix using the continuous surrogate variables, the discrete projection residuals, and the approximate null space projection operator; solving for the scaling ratios of the column vectors of the discrete projection solution that approximate the corresponding columns of the compensation target matrix by least squares, forming a scaling ratio vector; and absorbing the scaling ratio vector into the continuous scaling vector corresponding to the discrete projection solution to suppress the propagation error of the discrete projection residuals under the forward mapping of the fixed one-sided factor. The reconstruction optimization module is configured to freeze the two binary symbol matrices obtained by the binary structure decomposition module, and only use each continuous scaling vector as a learnable parameter for optimization in the reconstruction stage, so as to reduce the difference between the output distribution of the quantized model and the output distribution of the original full-precision model. The reconstruction optimization module does not unfreeze the binary symbol matrix or introduce additional full-precision proxy weight matrix during the reconstruction phase, and the system does not explicitly retain the approximate null space projection operator during the inference phase; the quantization model derived by the system is used to replace the original floating-point matrix multiplication with bit operations and channel-by-channel multiplication on the inference device.