Model quantification method and device, electronic equipment, storage medium and program product

By constructing a variance equalization rotation matrix and an adaptive variance equalization transformation strategy, the problem of inconsistent variance between channels in large language models is solved, achieving higher quantization accuracy and model efficiency.

CN120706481APending Publication Date: 2025-09-26NANJING HOUMO TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510915033.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing quantization techniques cannot effectively balance the variance between channels in different large language models, resulting in uneven distribution of quantized data values ​​and reduced model accuracy.

Method used

The variance equalization rotation matrix is ​​constructed through feature decomposition, and the variance equalization transformation is performed on the quantized data to achieve precise alignment of the variances between channels. The adaptive variance equalization transformation strategy and smoothing parameter fusion are adopted to improve the consistency of the variances between channels.

Benefits of technology

It achieves higher quantization accuracy and more uniform data distribution, improving the operating efficiency and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706481A_ABST
    Figure CN120706481A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a model quantification method and apparatus, an electronic device, a storage medium and a program product. The method comprises the steps of obtaining to-be-quantized data in a target model; performing characteristic decomposition on the covariance matrix of the to-be-quantized data to obtain a characteristic vector matrix; constructing a variance equilibrium rotation matrix based on the feature vector matrix; based on the variance equilibrium rotation matrix, performing variance equilibrium transformation on the to-be-quantized data; and based on the to-be-quantized data after variance equilibrium transformation, performing quantization processing on the target model so as to deploy the quantized target model to hardware equipment, the quantized target model being capable of being called by the hardware equipment to perform corresponding task processing, thereby performing principal component analysis on the to-be-quantized data based on eigendecomposition, and obtaining the to-be-quantized data. According to the technical scheme, accurate alignment of the variances between the channels can be achieved, the variances between the channels of the to-be-quantized data tend to be consistent, numerical distribution of the quantized data is more uniform, higher quantization precision is achieved, and then the execution efficiency of hardware equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a model quantization method, device, electronic device, storage medium, and program product. Background Art

[0002] As large language models (LLMs) continue to grow in size and complexity, achieving efficient model deployment within limited computing resources has become a key issue in LLM applications. Quantization technology, as an important tool, significantly reduces memory usage, computational burden, and energy consumption by converting neural network weights and activation data from high-precision (such as 32-bit floating-point numbers) to low-precision (such as 8-bit integers), thereby enabling the effective application of LLMs in resource-constrained environments.

[0003] While quantization techniques have been applied to various neural network architectures, the varying characteristics of different LLM models present numerous challenges. For example, some model weights and activation data contain significant outliers. The computational operations of different models affect the data distribution, making quantization more difficult. These differences lead to uneven distribution of quantized data, which in turn reduces model accuracy. Summary of the Invention

[0004] In view of the above problems in the prior art, embodiments of the present disclosure provide a model quantization method, apparatus, electronic device, storage medium, and program product.

[0005] A first aspect of the present disclosure provides a model quantization method, including:

[0006] Obtaining data to be quantized in a target model, where the target model includes a large language model;

[0007] Perform eigendecomposition on the covariance matrix of the quantized data to obtain the eigenvector matrix;

[0008] Construct a variance equalization rotation matrix based on the eigenvector matrix;

[0009] Based on the variance equalization rotation matrix, the quantized data is transformed into a variance equalization matrix;

[0010] Based on the data to be quantized after the variance equalization transformation, the target model is quantized to deploy the quantized target model to the hardware device. The quantized target model can be called by the hardware device to process the corresponding task.

[0011] As a possible implementation of the first aspect, the data to be quantized includes weight data of a matrix multiplication module in a target model; and based on a variance equalization rotation matrix, performing a variance equalization transformation on the data to be quantized includes:

[0012] Based on the variance equalization rotation matrix, the value matrix in the weight data is subjected to variance equalization transformation to obtain the transformed value matrix;

[0013] Performing a first matrix multiplication calculation on the query matrix and the key matrix in the weight data, and performing a normalized exponential operation on the result of the first matrix multiplication calculation;

[0014] Perform a second matrix multiplication calculation on the result of the normalized exponential operation and the transformed value matrix;

[0015] Based on the transposed matrix of the variance equalization rotation matrix, the output transformation matrix of the matrix multiplication module is subjected to variance equalization transformation;

[0016] The result of the second matrix multiplication calculation is processed based on the transformed output transformation matrix to obtain an output result of the matrix multiplication module.

[0017] As a possible implementation manner of the first aspect, the data to be quantized includes inter-block connection data in the target model, and the inter-block connection data includes at least one of output projection data, gate projection data, and state projection data.

[0018] As a possible implementation manner of the first aspect, the data to be quantized includes at least one of weight data and activation data of a low-rank adaptive module in the target model.

[0019] As a possible implementation of the first aspect, performing a variance equalization transformation on the quantized data based on a variance equalization rotation matrix includes:

[0020] The preset smoothing parameters are fused into the linear layer of the target model to obtain the smoothed fused model to be quantized;

[0021] Based on the variance equalization rotation matrix, the variance equalization transformation is performed on the model to be quantized after smooth fusion.

[0022] As a possible implementation of the first aspect, a preset smoothing parameter is integrated into the linear layer of the target model, including:

[0023] Add a smoothing parameter to the weight matrix of the projection layer of the matrix multiplication module of the target model.

[0024] As a possible implementation of the first aspect, the method further includes:

[0025] Based on the modified multiplication operator, the projection operation result of the projection layer based on the data processing unit is modified.

[0026] A second aspect of the present disclosure provides a model quantization device, including:

[0027] An acquisition unit, configured to: acquire data to be quantized in a target model, where the target model includes a large language model;

[0028] A decomposition unit is used to perform eigendecomposition on the covariance matrix of the quantized data to obtain an eigenvector matrix;

[0029] A construction unit, used for: constructing a variance equalization rotation matrix based on the eigenvector matrix;

[0030] A transformation unit, configured to perform a variance equalization transformation on the quantized data based on a variance equalization rotation matrix;

[0031] The quantization unit is used to: quantize the target model based on the data to be quantized after the variance equalization transformation, so as to deploy the quantized target model to the hardware device. The quantized target model can be called by the hardware device to process the corresponding task.

[0032] As a possible implementation of the second aspect, the data to be quantized includes weight data of a matrix multiplication module in a target model; and the transform unit is configured to:

[0033] Based on the variance equalization rotation matrix, the value matrix in the weight data is subjected to variance equalization transformation to obtain the transformed value matrix;

[0034] Performing a first matrix multiplication calculation on the query matrix and the key matrix in the weight data, and performing a normalized exponential operation on the result of the first matrix multiplication calculation;

[0035] Perform a second matrix multiplication calculation on the result of the normalized exponential operation and the transformed value matrix;

[0036] Based on the transposed matrix of the variance equalization rotation matrix, the output transformation matrix of the matrix multiplication module is subjected to variance equalization transformation;

[0037] The result of the second matrix multiplication calculation is processed based on the transformed output transformation matrix to obtain an output result of the matrix multiplication module.

[0038] As a possible implementation manner of the second aspect, the data to be quantized includes inter-block connection data in the target model, and the inter-block connection data includes at least one of output projection data, gate projection data, and state projection data.

[0039] As a possible implementation manner of the second aspect, the data to be quantized includes at least one of weight data and activation data of a low-rank adaptive module in the target model.

[0040] As a possible implementation of the second aspect, the transform unit includes:

[0041] The smoothing subunit is used to fuse the preset smoothing parameters into the linear layer of the target model to obtain a smoothed fused model to be quantized;

[0042] The transformation subunit is used to perform variance equalization transformation on the model to be quantized after smooth fusion based on the variance equalization rotation matrix.

[0043] As a possible implementation manner of the second aspect, the smoothing subunit is configured to:

[0044] Add a smoothing parameter to the weight matrix of the projection layer of the matrix multiplication module of the target model.

[0045] As a possible implementation manner of the second aspect, the smoothing subunit is further configured to:

[0046] Based on the modified multiplication operator, the projection operation result of the projection layer based on the data processing unit is modified.

[0047] A third aspect of the present disclosure provides an electronic device, including:

[0048] a memory for storing a computer program product;

[0049] The processor is configured to execute a computer program product stored in the memory, and when the computer program product is executed, implements any one of the methods of the first aspect.

[0050] A fourth aspect of an embodiment of the present disclosure provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the method of any one of the above-mentioned first aspects is implemented.

[0051] A fifth aspect of the present disclosure provides a computer program product, including computer program instructions, which, when executed by a processor, implement any of the methods described in the first aspect.

[0052] Based on the disclosed embodiments, principal component analysis is performed on the data to be quantized in the target model based on eigendecomposition to construct a variance equalization rotation matrix. A variance equalization transformation is then performed on the data to be quantized based on the variance equalization rotation matrix. The target model is then quantized based on the data to be quantized after the variance equalization transformation. This allows for precise alignment of inter-channel variances, making the inter-channel variances of the data to be quantized consistent. By improving the consistency of inter-channel variances, the numerical distribution of the quantized data becomes more uniform, thereby achieving higher quantization accuracy.

[0053] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0055] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:

[0056] Figure 1 Schematic diagram of model data distribution differences in related technologies;

[0057] Figure 2 A flowchart of an embodiment of the model quantification method disclosed herein;

[0058] Figure 3 A flowchart of an embodiment of the model quantization method disclosed herein;

[0059] Figure 4 This is a schematic diagram of an application of an embodiment of the disclosed model quantization method;

[0060] Figure 5 This is a schematic diagram of an application of an embodiment of the disclosed model quantization method;

[0061] Figure 6 This is a schematic diagram of an application of an embodiment of the disclosed model quantization method;

[0062] Figure 7 A flowchart of an embodiment of the model quantization method disclosed herein;

[0063] Figure 8 This is a schematic diagram of an application of an embodiment of the disclosed model quantization method;

[0064] Figure 9 This is a schematic diagram of an application of an embodiment of the disclosed model quantization method;

[0065] Figure 10 This is a schematic structural diagram of an embodiment of the model quantization device disclosed herein;

[0066] Figure 11 This is a schematic structural diagram of an embodiment of the model quantization device disclosed herein;

[0067] Figure 12 The figure is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0068] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention.

[0069] The following description of at least one exemplary embodiment is merely illustrative in nature and is not intended to limit the invention, its application, or uses.

[0070] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0071] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0072] Embodiments of the present invention may be applied to electronic devices such as computer systems / servers, which may operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with electronic devices such as computer systems / servers include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above.

[0073] Computer systems / servers and other electronic devices may be described in the general context of computer system-executable instructions (such as program modules) executed by the computer system. Generally, program modules may include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers may be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules may be located on local or remote computer system storage media, including storage devices.

[0074] In order to accurately describe the technical content of the present disclosure and to accurately understand the present disclosure, the following explanations or definitions of the terms used in this specification are given before describing the specific embodiments:

[0075] 1) Token: This refers to the basic unit of input into the model, typically a word, character, or subword. When working with natural language processing (NLP) tasks, tokens are the fundamental building blocks of text data. Each token represents a word, punctuation mark, or other linguistic unit. Samples in large models can often include multiple tokens. For example, in large language models, a token is the smallest unit for segmenting and encoding input text, and can be a word, subword, character, or other form of text fragment.

[0076] 2) Mamba: A deep learning architecture based on the State Space Model (SSM), primarily designed for processing sequence data. Its goal is to capture long-range dependencies through efficient structured matrix operations while avoiding the computational bottlenecks of traditional sequence models.

[0077] 3) GPTQ (Generative Pretrained Quantization): This is a quantization technique for generative pretrained Transformer models. The main purpose of GPTQ is to significantly reduce storage requirements and computational costs while maintaining model performance.

[0078] 4) Transformer: Essentially an encoder-decoder architecture, the Transformer can be divided into two parts: the encoder and the decoder. The Transformer model architecture uses a self-attention mechanism, replacing the recurrent neural network (RNN) architecture commonly used in NLP (natural language processing) tasks. Its biggest advantage over the RNN architecture is its parallel computing capabilities.

[0079] 5) RMSNorm (Root Mean Square Layer Normalization): This is a normalization technique used in deep learning models, particularly well-suited for architectures like the Transformer. A variant of LayerNorm, it aims to simplify the normalization process and reduce computational complexity while maintaining or improving model performance. Normalization is performed by calculating the root mean square (RMS) of the input vector.

[0080] 6) SiLU (Sigmoid Linear Unit) activation function: also known as the Swish activation function, it is defined as follows:

[0081] SiLU(x)=x·σ(x)

[0082] Among them, σ(x) is the standard sigmoid function, and its value is between 0 and 1. The characteristics of the Swish function include nonlinearity, continuous differentiability, and are defined in the range from negative infinity to positive infinity.

[0083] 7) DIMC (Digital In-Memory Compute) is an innovative computing architecture that eliminates the latency and energy consumption of data transmission by integrating computing units into memory, making it particularly suitable for AI inference operations. Compared with traditional GPUs and other alternatives, DIMC provides higher performance, energy efficiency and cost savings. The core advantage of DIMC technology lies in its ability to provide ultra-high memory bandwidth and fast interaction speeds. This technology significantly improves the performance of AI inference by reducing the latency and energy consumption of data transmission. In addition, DIMC technology is also suitable for other compute-intensive applications such as machine learning and big data processing. By integrating computing units into memory, DIMC can significantly improve processing speed and efficiency, and is particularly suitable for application scenarios that require frequent access to large static data sets.

[0084] 8) SIMD (Single Instruction Multiple Data) is a parallel computing technology that uses a single instruction to operate on multiple data elements simultaneously, achieving data-level parallelism. It is widely used in fields such as image processing and scientific computing. Its core feature is that multiple processing units execute the same instruction synchronously under a single controller, significantly improving computing throughput. By managing multiple processing units with a single controller, the same operation is performed on each element in a set of data (such as a vector), achieving spatial parallelism. Typical applications include CPU instruction sets and parallel computing architectures.

[0085] The following first introduces the relevant technologies, and then introduces the technical solutions of the embodiments of the present disclosure in detail.

[0086] In related technologies, taking the Mamba model as an example, there are obvious outliers in its weight data and activation data, especially in the gate projection layer weights for language tasks and the output projection layer activations for vision tasks. In addition, in the Mamba structure, the output of the parallel scan (PScan) operation is the input of the matrix multiplication, and the parallel scan (PScan) operation further exacerbates the problem of outliers in activations. Taking the Transformer model as an example, similar distribution differences are also observed in the feedforward network (FFN) and multi-head self-attention (MHSA) modules. See Figure 1 Figure (a) and Figure 1 As shown in Figure (b), in the example of the Transformer model, the variance distribution between the channels of the feedforward network input data (FFN Input) and the input data of the multi-head self-attention module (MHSA Input) is quite different.

[0087] Since both Mamba and Transformer are sequence models containing fully connected layers, these layers need to be quantized. In the existing quantization method based on Hadamard rotation transform, the linear part of RMSNorm (Root Mean Square Layer Normalization) is first absorbed into the adjacent weight matrix, and then the model is rotated using the Hadamard matrix Q, and finally the quantization operation is performed. This quantization method has shortcomings in balancing the variance between channels. Specifically, given a centered data matrix X (with dimensions of n×m) and a variance equalization transformation matrix H (with dimensions of m×m), applying this method cannot ensure the consistency of variance between channels, resulting in the problem of uneven numerical distribution during the quantization process, thereby reducing the accuracy of the model.

[0088] In summary, the relevant technologies have at least the following problems: the Hadamard transform cannot achieve variance alignment between channels. The above inconsistent variances will inevitably cause uneven numerical distribution of quantized data, thereby reducing model accuracy.

[0089] Based on the technical problems existing in the above-mentioned related technologies, the embodiment of the present disclosure provides a model quantization method. First, based on the feature decomposition, principal component analysis is performed on the data to be quantized in the target model to construct a variance equalization rotation matrix, and the variance equalization transformation is performed on the data to be quantized based on the variance equalization rotation matrix; then, based on the data to be quantized after the variance equalization transformation, the target model is quantized, which can achieve accurate alignment of the variances between channels and make the variances between the channels of the data to be quantized tend to be consistent. By improving the consistency of the variances between channels, the numerical distribution of the quantized data is made more uniform, thereby achieving higher quantization accuracy, thereby solving the technical problems in the prior art that variance alignment cannot be achieved between channels and the numerical distribution of the quantized data is uneven.

[0090] In the related art, there is currently no unified, universal quantization method for Mamba and Transformer models. The model quantization method provided by the disclosed embodiments can achieve unprecedented accuracy preservation on Transformer models, and for the first time, effectively applies quantization to Mamba models. It is applicable to various LLM models, including Mamba and Transformer. By improving the consistency of inter-channel variance, the quantization difficulty is reduced, thereby achieving higher quantization accuracy.

[0091] Figure 2 FIG. 1 is a flow chart of an embodiment of the model quantization method disclosed herein. Figure 2 As shown, the method may specifically include:

[0092] Step S110 , obtaining data to be quantized in a target model, where the target model includes a large language model.

[0093] Model quantization is a technique for compressing models by reducing the precision of model parameters, aiming to reduce storage requirements and computational costs while maintaining high inference accuracy. For example, model quantization can be the process of mapping model parameters from high-precision floating-point numbers (such as FP32) to low-precision integers (such as 8 bits or lower). The core idea is to improve the model's operational efficiency by reducing the data representation precision, reducing the model's storage requirements and computational complexity, especially on resource-constrained devices (such as mobile devices or embedded systems).

[0094] The target model is a model that needs to be quantized, for example, various LLM models including but not limited to Mamba and Transformer. The data to be quantized in the target model may include various parameters of the target model, such as weight data or activation data of the target model. In one example, the data to be quantized can be a centered data matrix (the columns of the matrix are zero mean) X; wherein X represents a weight matrix or an activation matrix, and its dimension is (n, m). The matrix X can come from calibration data. Calibration data is data used to adjust the model's predicted probability distribution so that it matches the true probability distribution, which is equivalent to sampling the actual distribution of the data.

[0095] Step S120 , performing eigendecomposition on the covariance matrix of the data to be quantized to obtain an eigenvector matrix.

[0096] Among them, C X Represents the covariance matrix of the data to be quantized X. The centered covariance matrix C of the matrix X from the calibration data X Perform eigenvalue decomposition to obtain the eigenvector matrix K and the eigenvalue matrix Λ. The eigenvector matrix K and the eigenvalue matrix Λ are obtained according to the following formula (1). T Represents the transposed matrix of matrix X.

[0097]

[0098] Step S130: constructing a variance equalization rotation matrix based on the eigenvector matrix.

[0099] Randomly construct the Hadamard matrix H, as shown in the following formula (2). In the constructed Hadamard matrix H, the matrix element values ​​are required to satisfy That is, the value of each element of the matrix is One of these two values, and satisfying the orthogonality of the row vector and column vector of the matrix.

[0100]

[0101] Furthermore, based on the eigenvector matrix K and the constructed Hadamard matrix H, the variance equalization rotation matrix H is constructed. k , use the following formula (3) to construct the variance equalization rotation matrix H k .

[0102] H k =KH Formula (3)

[0103] Step S140 : performing variance equalization transformation on the data to be quantized based on the variance equalization rotation matrix.

[0104] Apply the variance equalization rotation matrix H k Rotate the weight matrix or activation matrix X. The specific method of variance equalization transformation can be designed according to the network structure of the target model. For example, the variance equalization rotation matrix H can be k Perform left or right multiplication on the weight matrix or activation matrix.

[0105] In step S150, the target model is quantized based on the data to be quantized after the variance equalization transformation, and the quantized target model is deployed to a hardware device. The quantized target model can be called by the hardware device to process the corresponding task. The target model can include various LLM models, including Mamba and Transformer. After the quantization process, a quantized LLM is obtained, which can be called by the hardware device to process the corresponding task.

[0106] After performing variance equalization on the data to be quantized, the target model is quantized. For example, quantization can be performed using the same weight reconstruction method as in GPTQ. Specifically, this involves quantizing each column of weights, calculating the quantization error, and applying error compensation to subsequent columns.

[0107] The above embodiment can be applied to an offline transformation scenario, and in step S120, principal component analysis is performed on the data to be quantified in the target model based on eigendecomposition. The matrix is ​​decomposed into eigenvectors and eigenvalues ​​by eigendecomposition, and these eigenvectors and eigenvalues ​​reveal the intrinsic properties of the matrix. The eigenvectors represent the main directions of the matrix transformation, while the eigenvalues ​​represent the amplitude of the transformation in these directions. Based on principal component analysis, the data can be transformed into a new coordinate system so that the first largest variance of any data projection is on the first coordinate (called the first principal component), the second largest variance is on the second coordinate (called the second principal component), and so on. Principal component analysis is often used to reduce the dimensionality of a data set while maintaining the features that contribute most to the variance of the data set.

[0108] Therefore, the embodiment of the present disclosure can perform principal component analysis on the data to be quantized in the target model based on eigendecomposition, and then introduce the principal component analysis method into the adaptive variance equalization transformation strategy (AVETS). Specifically, the covariance matrix of the data to be quantized is subjected to eigendecomposition to obtain an eigenvector matrix, which is used as the analysis result of the principal component analysis on the data to be quantized. The variance equalization rotation matrix constructed based on the eigenvector matrix is ​​used to perform variance equalization transformation on the data to be quantized. After the transformation, the target model is quantized again, and the result of the principal component analysis is applied to the variance equalization transformation. Through the variance equalization transformation based on principal component analysis, the variance difference between channels can be reduced and the data distribution can be balanced. The above adaptive variance equalization transformation strategy can make the variance of the variables to be quantized consistent between channels, thereby forming a more uniform distribution.

[0109] Based on the disclosed embodiment, principal component analysis is performed on the data to be quantized in the target model based on eigendecomposition, and then a variance equalization rotation matrix is ​​constructed. A variance equalization transformation is performed on the data to be quantized based on the variance equalization rotation matrix. The target model is then quantized based on the data to be quantized after the variance equalization transformation, which can achieve precise alignment of the variances between channels and make the variances between the channels of the data to be quantized tend to be consistent. By improving the consistency of the variances between channels, the numerical distribution of the quantized data is made more uniform, thereby reducing the difficulty of quantization and achieving higher quantization accuracy, which can accelerate calculations, reduce memory access, and improve the execution efficiency of hardware devices.

[0110] The quantized LLM obtained based on the embodiments of this disclosure can be deployed in electronic devices, such as terminals and servers, for processing tasks in areas such as natural language processing, conversational interaction, and computer vision. When used for natural language processing tasks, it can generate human-like text or answer natural language questions based on input text information. When used for computer vision tasks, it can accurately classify and recognize input images.

[0111] When applying the LLM model to perform corresponding tasks, data processing operations such as matrix multiplication and multi-data parallel processing are typically involved. In a specific application example, matrix multiplication can be performed using a digital in-memory compute unit (DIMC), or multi-data parallel processing can be performed using a single instruction, multiple data unit (SIMD). For example, building on the aforementioned model quantization, the DIMC engine architecture can be further employed to move computation closer to RAM (memory), merging the memory with the multiplication and accumulation units in the compute unit to achieve greater computational bandwidth and efficiency, reduce latency, and reduce energy consumption. Another example is that building on the aforementioned model quantization, the single instruction, multiple data unit (SIMD) can be further employed to accelerate matrix operations and the forward and backward propagation processes of neural networks, improving training and inference speeds. When used for computer vision tasks, the single instruction, multiple data unit (SIMD) can be further employed to parallelize graphics data processing, significantly improving graphics processing efficiency, reducing rendering time, and enhancing the performance of gaming and virtual reality applications.

[0112] like Figure 3 As shown, in one embodiment, the data to be quantized includes weight data of a matrix multiplication module in a target model; based on a variance equalization rotation matrix, performing a variance equalization transformation on the data to be quantized includes:

[0113] Step S210, performing a variance equalization transformation on the value matrix in the weight data based on the variance equalization rotation matrix to obtain a transformed value matrix;

[0114] Step S220, performing a first matrix multiplication calculation on the query matrix and the key matrix in the weight data, and performing a normalized exponential operation on the result of the first matrix multiplication calculation;

[0115] Step S230 , performing a second matrix multiplication calculation on the result of the normalized exponential operation and the transformed value matrix;

[0116] Step S240, performing variance equalization transformation on the output transformation matrix of the matrix multiplication module based on the transposed matrix of the variance equalization rotation matrix;

[0117] Step S250 , processing the result of the second matrix multiplication calculation based on the transformed output transformation matrix to obtain an output result of the matrix multiplication module.

[0118] In the example of a Transformer model, the above adaptive variance equalization transformation strategy can also be used in the matrix multiplication module (ovmatmul) of the multiheadattention. In the Transformer model, by using three trainable parameter matrices Wq (query matrix), W k (key matrix), W v (Value matrix), transform the input matrix to obtain three matrices Q, K, and V respectively. Among them, the Q (Query) matrix is ​​used to find the correlation between the current input and other inputs; the K (Key) matrix represents the characteristics of the input and is used to compare with the query vector to calculate the similarity; the V matrix (Value) represents the actual information of the input, which is weighted with the similarity weight to generate the final output. The three matrices Q, K, and V are obtained from the input sequence through linear transformation. By using the above three trainable parameter matrices W q 、W k 、W v , which can enhance the model's fitting ability.

[0119] like Figure 4 As shown, in step S210, based on the variance equalization rotation matrix H k , for the value matrix W in the weight data v Perform variance equalization transformation to obtain the transformed value matrix. Specifically, the variance equalization rotation matrix H can be used k Logarithm W v Perform right multiplication to obtain the transformed value matrix.

[0120] like Figure 4 As shown, in step S220, the query matrix W in the weight data is q and bond matrix W k Perform the first matrix multiplication calculation (matmul), and then perform normalized exponential operation on the result of the first matrix multiplication calculation, that is, Figure 4 The softmax operation in . Softmax is an operation that converts a real number vector into a probability distribution.

[0121] like Figure 4 As shown, in step S230, a second matrix multiplication calculation is performed on the result of the normalized exponential operation in step S220 and the value matrix transformed in step S210.

[0122] like Figure 4 As shown, in step S240, the transposed matrix H of the variance equalization rotation matrix is k .T, transform the output matrix W of the matrix multiplication module o Perform variance equalization transformation. Among them, the output transformation matrix W o is the weight matrix of the output layer of the matrix multiplication module. Specifically, the transposed matrix H of the variance equalization rotation matrix can be used k .T output transformation matrix W o Perform left multiplication to obtain the transformed output transformation matrix.

[0123] like Figure 4 As shown, in step S250, the result of the second matrix multiplication calculation is processed based on the output transformation matrix transformed in step S240 to obtain the output result of the matrix multiplication module.

[0124] In large models, matrix multiplication can be viewed as the multiplication of two matrices, where one matrix represents the input data and the other represents the model parameters. Large models usually use matrix multiplication to establish the mapping relationship between input data and model parameters. Figure 3 and Figure 4 In the embodiment shown, the variance equalization rotation matrix H constructed based on the eigenvector matrix is ​​used. k By performing a variance equalization transformation on the data to be quantized in the matrix multiplication module, we can reduce the variance differences between channels after the target model is quantized, making the variance of the data to be quantized more consistent across channels. By improving the consistency of the variance between channels, the numerical distribution of the quantized data becomes more uniform, thereby achieving higher quantization accuracy.

[0125] In one embodiment, the data to be quantified includes inter-block connection data in the target model, and the inter-block connection data includes at least one of output projection data, gate projection data, and state projection data.

[0126] In this embodiment, the above adaptive variance equalization transformation strategy AVETS is applied to the offline transformation of Transformer and Mamba models. Specifically, AVETS can be applied between blocks of Transformer or Mamba models. Figure 5 As shown, the adaptive variance equalization rotation matrix H can be fused in the rotation between blocks k , the specific fusion method can be matrix multiplication fusion, such as the variance equalization rotation matrix H k Perform matrix multiplication with the weight matrix in the block, and obtain the transformed new weight matrix through fusion, that is, the new model parameters, to achieve the effect of data variance equalization. Specifically, the gated weight matrix, state weight matrix and variance equalization rotation matrix H in the block can be k Fusion, the output transformation matrix and the transposed matrix H of the variance equalization rotation matrix k .T fusion, making H k and H k .T is absorbed into the corresponding weight matrix.

[0127] like Figure 5As shown in the figure, the "output head" in the model is the top part at the end of the entire network structure, usually the last layer of the model, usually composed of a fully connected layer or a convolutional layer, which is used to map the extracted features to the final output space and generate the final prediction results of the network. The "embedding layer" in the model is the front part of the model (encoding layer), in which high-dimensional data is mapped to a low-dimensional space. It is usually used to convert discrete, non-continuous data into a continuous vector representation for easy computer processing. In the process of rotation transformation, the weight matrix of the embedding layer can be compared with the transposed matrix H of the variance equalization rotation matrix. k .T, and the weight matrix of the output head is combined with the variance equalization rotation matrix H k Combined, H k and H k .T is absorbed into the corresponding weight matrix.

[0128] In one embodiment, the above method also includes: the data to be quantized includes at least one of weight data and activation data of a low-rank adaptive module in the target model.

[0129] For example, the above adaptive variance equalization transformation strategy AVETS can be used in the LoRA (Low-Rank Adaptation) module of the Mamba model. Figure 6 As shown, Wx_dt and Wdt represent the weight matrices of the input layer and output layer of the LoRA module respectively. During the rotation transformation, the weight matrix of the input layer can be combined with the variance equalization rotation matrix H k Fusion, the weight matrix of the output layer and the transposed matrix H of the variance equalization rotation matrix k .T fusion, making H k and H k .T is absorbed into the corresponding weight matrix, and the transformed new weight matrix is ​​obtained through fusion to achieve the purpose of data variance balancing and effectively improve the quantization effect.

[0130] like Figure 7 As shown, in one embodiment, Figure 1 Step S140 in the embodiment of the present invention performs variance equalization transformation on the quantized data based on the variance equalization rotation matrix, including:

[0131] Step S310, fusing the preset smoothing parameters into the linear layer of the target model to obtain a smoothed fused model to be quantized;

[0132] Step S320 , performing variance equalization transformation on the smoothed fused model to be quantized based on the variance equalization rotation matrix.

[0133] For the Transformer model and the Mamba model, since there is a SiLU activation function, the nonlinear SiLU activation function cannot transfer linear parameters, thus hindering the fusion of smoothing parameters. In the embodiment of the present disclosure, a static smooth fusion rotation strategy (SSFRS) is adopted to introduce smoothing parameters before the online variance equalization rotation, and the preset smoothing parameters are fused into the linear layer of the target model to obtain a smoothed fused model to be quantized. These smoothing parameters can equalize the channel variance between the variables to be quantized, that is, to reduce the difference between the numerical values. Then, based on the variance equalization rotation matrix H k The variance equalization transformation is performed on the model to be quantized after smooth fusion to balance the channel variance of the variables to be quantized and make up for the defects of the traditional Hadamard rotation transformation.

[0134] Specifically, in the embodiment of the present disclosure, the S-SiLU (smoothed gated linear unit) activation function is used to replace the traditional SiLU activation function, which is defined as shown in formula (4):

[0135] S-SiLU(x,s)=x☉σ(s☉x), Formula (4)

[0136] Where s represents the preset smoothing parameter. Based on the smoothed gated linear activation function shown in formula (4), the final output of the output projection layer has the form shown in formula (5).

[0137] y out =[y ssm ☉SiLU(x g W g )]W o =[y ssm ☉S-SiLU(x g W′ g ,s out )]W′ o Formula (5)

[0138] Among them, x g Represents the input data of the gate projection layer; W g and W′ g Respectively represent the gate weight matrices before and after the fusion transformation; W o and W′ o Respectively represent the output transformation matrix before and after the fusion transformation; s out Represents the smoothing parameter preset in the output projection layer; y ssm It represents the result obtained by the output projection layer processing the input data before smoothing; out Indicates yssm After smoothing, the final output result of the projection layer is output.

[0139] In this way, the smoothing parameters are integrated into the calculation of the output projection layer without changing the output result. The conversion process is as follows Figure 8 shown. Figure 8 W in gate Represents the gate weight matrix, that is, W g ;W down Represents the weight matrix of the last linear layer, namely W o ; Mul represents multiplication operation. Figure 8 The figure above shows the conversion process based on the traditional SiLU activation function. Figure 8 The following figure shows the conversion process of the S-SiLU (smooth gated linear unit) activation function based on the embodiment of the present disclosure. Figure 8 In the following figure, W gate Divide by the smoothing parameter s (i.e. Figure 8 Lieutenant General W gate Multiply by the inverse of the smoothing parameter 1 / s), and W down Multiply by the smoothing parameter s. In other words, during the transformation process, smaller values ​​can be multiplied by a value and larger values ​​can be divided by a value. This processing method is mathematically equivalent to the transformation and can achieve the purpose of keeping the output unchanged.

[0140] In the embodiment of the present disclosure, smoothing parameters are introduced before the online variance equalization rotation. These smoothing parameters can be used to equalize the channel variances between the variables to be quantized, so that the inter-channel variances of the data to be quantized tend to be consistent, making the numerical distribution of the quantized data more uniform, thereby achieving higher quantization accuracy.

[0141] In one embodiment, the smoothing parameter of is integrated into the linear layer of the target model, including:

[0142] Add a smoothing parameter to the weight matrix of the projection layer of the matrix multiplication module of the target model.

[0143] Specifically, the fusion of smoothing parameters can be applied to the matrix multiplication module of the Mamba model. The fusion structure of smoothing parameters is as follows: Figure 9 See Figure 9 , Mul represents the multiplication instruction; EXP represents the natural exponential function; matmul represents the matrix multiplication calculation. Figure 9 The left figure in the figure shows the network structure of the projection layer of the traditional matrix multiplication module (Origin). Figure 9 The right figure in shows the network structure of the projection layer of the matrix multiplication module (Ours) that integrates the smoothing parameter in an embodiment of the present disclosure.

[0144] See also Figure 9 In the right figure, one of the data sources of the matrix multiplication module is the C projection layer in the Mamba block, which enables the model to filter out irrelevant information. Figure 9 Wc in the C-projection layer represents the weight matrix. For the C-projection layer in the Mamba block, its weight matrix can be directly incorporated into the smoothing parameters. For example, the weight matrix of the C-projection layer can be directly multiplied by the smoothing parameters to directly incorporate the weight matrix into the smoothing parameters.

[0145] See also Figure 9 In the right figure, another data source of the matrix multiplication module is the output of the PScan operator, where the B projection layer is used to perform linear transformation and feature extraction on the input data and project the input data into different feature spaces. Figure 9 W in B Represents the weight matrix of the B-projection layer. The B-projection layer can directly incorporate the smoothing parameter. For example, the weight matrix of the B-projection layer can be directly multiplied by the inverse of the smoothing parameter 1 / s to fuse the weight matrix with the smoothing parameter.

[0146] After the above smoothing process, the variance of the activation channels of the matrix multiplication module becomes relatively uniform, which makes the numerical distribution of the quantized data more uniform and achieves higher quantization accuracy.

[0147] In one embodiment, the above method further comprises:

[0148] Based on the modified multiplication operator, the projection operation result of the projection layer based on the data processing unit is modified.

[0149] See also Figure 9 In the right figure, the A projection layer in the Mamba block refers to the process of mapping the input data into the state space model (SSM), that is, Figure 9 The "A" in . Figure 9 "Δ(1), Δ(2), ..., Δ(t)" in represents a sequence of processing units or processing objects. Figure 9 As shown, for the A projection layer in the Mamba block, it needs to be continuously applied to multiple tokens over time, making it difficult to directly perform an absorption transformation. Therefore, in the embodiment of the present disclosure, a correction multiplication operator (addcmul operator) is introduced to correct the projection operation result of the projection layer based on the data processing unit based on the correction multiplication operator, wherein the data processing unit may include the first token in the input data of the projection layer. The calculation of its first token is modified, as shown in formula (6):

[0150] addcmul(-ln(s mm ),Δ(1),A)=AΔ(1)-ln(smm ) Formula (6)

[0151] Among them, Δ(1) represents the data of the first token; A represents the projection operation; s mm represents the smoothing parameter, that is Figure 9 The s in.

[0152] Through the above correction processing based on the modified multiplication operator, the influence of time passage on the absorption transformation is overcome, so that the weight matrix can directly smooth the parameters better and achieve a better smoothing effect.

[0153] In summary, in the embodiments of the present disclosure, the adaptive variance equalization transformation strategy AVETS based on principal component analysis is introduced in the offline transformation scenario, and the static smooth fusion rotation strategy SSFRS is adopted for the online transformation scenario. The above strategies achieve the following beneficial effects:

[0154] 1) Inter-channel variance consistency: Through AVETS and SSFRS, the inter-channel variance of the quantified variables is made consistent, forming a more uniform distribution.

[0155] 2) High-precision quantization: This approach improves quantization accuracy across multiple LLM models (including Mamba and Transformer) and various tasks (such as vision and natural language processing). In one example, applying the model quantization method of an embodiment of this disclosure can achieve a W4A8 quantization loss accuracy of less than 1%.

[0156] 3) Wide applicability: The disclosed embodiments are not only applicable to Transformer and Mamba models, but can also be extended to other LLM models, and have strong versatility.

[0157] like Figure 10 As shown, the present disclosure also provides an embodiment of a corresponding model quantization device. For the beneficial effects or technical problems solved by the device, please refer to the description of the method corresponding to each device, or refer to the description in the invention content, and will not be repeated here.

[0158] In an embodiment of the model quantization device, the device includes:

[0159] The acquisition unit 100 is used to: acquire the data to be quantized in the target model, where the target model includes a large language model;

[0160] The decomposition unit 200 is used to perform eigendecomposition on the covariance matrix of the quantized data to obtain an eigenvector matrix;

[0161] The construction unit 300 is used to: construct a variance equalization rotation matrix based on the eigenvector matrix;

[0162] The transformation unit 400 is configured to perform a variance equalization transformation on the quantized data based on the variance equalization rotation matrix;

[0163] The quantization unit 500 is used to: quantize the target model based on the data to be quantized after the variance equalization transformation, so as to deploy the quantized target model to the hardware device. The quantized target model can be called by the hardware device to process the corresponding task.

[0164] In one embodiment, the data to be quantized includes weight data of a matrix multiplication module in a target model; the transform unit 400 is configured to:

[0165] Based on the variance equalization rotation matrix, the value matrix in the weight data is subjected to variance equalization transformation to obtain the transformed value matrix;

[0166] Performing a first matrix multiplication calculation on the query matrix and the key matrix in the weight data, and performing a normalized exponential operation on the result of the first matrix multiplication calculation;

[0167] Perform a second matrix multiplication calculation on the result of the normalized exponential operation and the transformed value matrix;

[0168] Based on the transposed matrix of the variance equalization rotation matrix, the output transformation matrix of the matrix multiplication module is subjected to variance equalization transformation;

[0169] The result of the second matrix multiplication calculation is processed based on the transformed output transformation matrix to obtain an output result of the matrix multiplication module.

[0170] In one embodiment, the data to be quantified includes inter-block connection data in the target model, and the inter-block connection data includes at least one of output projection data, gate projection data, and state projection data.

[0171] In one embodiment, the data to be quantized includes at least one of weight data and activation data of a low-rank adaptive module in the target model.

[0172] like Figure 11 As shown, in one embodiment, the transformation unit 400 includes:

[0173] The smoothing subunit 410 is used to: fuse the preset smoothing parameters into the linear layer of the target model to obtain a smoothed and fused model to be quantized;

[0174] The transformation subunit 420 is configured to perform a variance equalization transformation on the smoothed fused model to be quantized based on the variance equalization rotation matrix.

[0175] In one embodiment, the smoothing subunit 410 is configured to:

[0176] Add a smoothing parameter to the weight matrix of the projection layer of the matrix multiplication module of the target model.

[0177] In one embodiment, the smoothing subunit 410 is further configured to:

[0178] Based on the modified multiplication operator, the projection operation result of the projection layer based on the data processing unit is modified.

[0179] The model quantization device of the embodiment of the present disclosure corresponds to the above-mentioned model quantization method embodiments of the present disclosure in terms of specific implementation and beneficial technical effects. The relevant contents can be referenced to each other and will not be repeated here.

[0180] Below, reference Figure 12 The electronic device according to the embodiment of the present disclosure is described. The electronic device may be either or both of the first device and the second device, or a standalone device independent of them, and the standalone device may communicate with the first device and the second device to receive collected input signals from them.

[0181] Figure 12 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.

[0182] like Figure 12 As shown, the electronic device includes one or more processors and memory.

[0183] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0184] The memory may store one or more computer program products, and the memory may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program products may be stored on a computer-readable storage medium, and the processor may execute the computer program products to implement the model quantization method of each embodiment of the present disclosure and / or other desired functions.

[0185] In one example, the electronic device may further include an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0186] In addition, the input device may also include, for example, a keyboard, a mouse, and the like.

[0187] The output device can output various information to the outside, including determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.

[0188] Of course, to simplify, Figure 12 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.

[0189] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps in the model quantization method according to various embodiments of the present disclosure described in the above part of this specification.

[0190] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of the disclosed embodiments, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0191] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps in the model quantization method according to various embodiments of the present disclosure described in the above part of this specification.

[0192] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0193] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0194] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.

[0195] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0196] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.

[0197] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0198] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0199] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A model quantization method, characterized in that: include: Acquiring data to be quantized in a target model, wherein the target model includes a large language model; Performing eigendecomposition on the covariance matrix of the data to be quantized to obtain an eigenvector matrix; Constructing a variance equalization rotation matrix based on the eigenvector matrix; Based on the variance equalization rotation matrix, performing variance equalization transformation on the data to be quantized; Based on the data to be quantized after the variance equalization transformation, the target model is quantized to deploy the quantized target model to a hardware device. The quantized target model can be called by the hardware device to process the corresponding task.

2. The method according to claim 1, characterized in that The data to be quantized includes weight data of a matrix multiplication module in the target model; performing a variance equalization transformation on the data to be quantized based on the variance equalization rotation matrix includes: Based on the variance equalization rotation matrix, performing a variance equalization transformation on the value matrix in the weight data to obtain a transformed value matrix; Performing a first matrix multiplication calculation on the query matrix and the key matrix in the weight data, and performing a normalized exponential operation on a result of the first matrix multiplication calculation; Performing a second matrix multiplication calculation on the result of the normalized exponential operation and the transformed value matrix; Based on the transposed matrix of the variance equalization rotation matrix, performing a variance equalization transformation on the output transformation matrix of the matrix multiplication module; The result of the second matrix multiplication calculation is processed based on the transformed output transformation matrix to obtain the output result of the matrix multiplication module.

3. The method according to claim 1, characterized in that The data to be quantified includes inter-block connection data in the target model, and the inter-block connection data includes at least one of output projection data, gate projection data, and state projection data.

4. The method according to claim 1, wherein The data to be quantized includes at least one of weight data and activation data of a low-rank adaptive module in the target model.

5. The method according to any one of claims 1 to 4, characterized in that The performing variance equalization transformation on the data to be quantized based on the variance equalization rotation matrix includes: Fusing the preset smoothing parameters into the linear layer of the target model to obtain a smoothed fused model to be quantized; Based on the variance equalization rotation matrix, a variance equalization transformation is performed on the model to be quantized after smooth fusion.

6. The method according to claim 5, characterized in that The step of fusing the preset smoothing parameters into the linear layer of the target model comprises: The smoothing parameter is added to the weight matrix of the projection layer of the matrix multiplication module of the target model.

7. The method according to claim 6, characterized in that The method further comprises: Based on the modified multiplication operator, the projection operation result of the projection layer based on the data processing unit is modified.

8. A model quantization device, characterized in that: include: An acquisition unit, configured to: acquire data to be quantized in a target model, wherein the target model includes a large language model; A decomposition unit, configured to perform eigendecomposition on the covariance matrix of the data to be quantized to obtain an eigenvector matrix; A construction unit, configured to: construct a variance equalization rotation matrix based on the eigenvector matrix; A transform unit, configured to: perform a variance equalization transform on the data to be quantized based on the variance equalization rotation matrix; The quantization unit is used to: quantize the target model based on the data to be quantized after the variance equalization transformation, so as to deploy the quantized target model to the hardware device, and the quantized target model can be called by the hardware device to process the corresponding task.

9. An electronic device, characterized in that: include: a memory for storing a computer program product; A processor is configured to execute the computer program product stored in the memory, and when the computer program product is executed, implements the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.

11. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Quantification method and device, electronic device, storage medium and electronic equipment

    CN121235002A