Model quantification method and device, electronic equipment and storage medium
By determining the weight matrix to be optimized in the large model and using weight transfer parameter optimization, the problem of increased storage and computing power in edge-side deployment is solved, the model accuracy and performance are maintained, and efficient quantization is achieved.
Patent Information
- Application Number
- CN202410251890.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-05
AI Technical Summary
When deploying large models on the client side, existing quantitative technologies lead to increased storage and computing power, as well as loss of accuracy, which affects the model's inference performance.
By determining the weight matrix to be optimized in the model to be quantified, obtaining the weight transfer parameters based on the target optimization function, optimizing and quantizing the weight matrix, reducing outliers and lowering precision loss.
Maintain model accuracy and performance in end-to-end scenarios, reduce precision loss after quantization, and achieve efficient deployment.
Smart Images

Figure CN120597958A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of model quantization, and in particular to a model quantization method, device, electronic device, and storage medium. Background Art
[0002] With the continuous development of large models, deploying large models on the end side (such as mobile terminals) is a future trend, and model quantization technology is one of the main ways to reduce deployment costs.
[0003] To ensure the quality of large model training, they are typically trained in the cloud. The resulting large models are also in the FP32 data type. While this offers high precision, it also consumes a large amount of storage. If large models are deployed on the client side for inference, the storage and computational requirements of the large models will place a significant burden on the client side.
[0004] In related technologies, quantizing large models usually causes significant accuracy loss, which affects the performance of the large models on the end. Summary of the Invention
[0005] The present application proposes a model quantization method, device, electronic device and storage medium, aiming to solve one of the technical problems in the related art at least to a certain extent.
[0006] The first embodiment of the present application proposes a model quantification method, including:
[0007] Determine each weight matrix to be optimized in the model to be quantified;
[0008] Obtaining a weight transfer parameter corresponding to each weight matrix based on a target optimization function associated with the weight matrix;
[0009] Optimizing the model to be quantized based on the weight transfer parameters corresponding to each weight matrix;
[0010] The optimized model to be quantized is quantized to obtain a target model.
[0011] Optionally, determining each weight matrix to be optimized in the model to be quantized includes:
[0012] The weight matrices to be optimized are determined according to the model structure of the model to be quantized.
[0013] Optionally, determining the weight matrices to be optimized according to the model structure of the model to be quantized includes:
[0014] In the case where the model to be quantized includes a root mean square normalization module RMSNorm, the weight matrices to be optimized include:
[0015] The first weight matrices connected to the RMSNorm module and the second weight matrices not connected to the RMSNorm module.
[0016] Optionally, the model to be quantized contains multiple decoder modules, each of which contains an attention layer and a feedforward layer.
[0017] The first weight matrices include: a query weight matrix, a key weight matrix, and a value weight matrix contained in the attention layer, and a third weight matrix in the feedforward layer;
[0018] The second weight matrices include: the output weight matrix contained in the attention layer, and the fourth weight matrix in the feedforward layer, wherein the third weight matrix and the fourth weight matrix have different corresponding positions in the feedforward layer.
[0019] Optionally, obtaining a weight transfer parameter corresponding to each weight matrix based on a target optimization function associated with the weight matrix includes:
[0020] determining a first objective optimization function associated with the first weight matrix and a second objective optimization function associated with the second weight matrix;
[0021] Determining a first weight transfer parameter corresponding to each decoder module based on the first objective optimization function and a query weight matrix, a key weight matrix, and a value weight matrix corresponding to each decoder module;
[0022] Determining a second weight transfer parameter corresponding to each decoder module based on the first objective optimization function and a third weight matrix corresponding to each decoder module;
[0023] Determining a third weight transfer parameter corresponding to each decoder module based on the second objective optimization function and the output weight matrix corresponding to each decoder module;
[0024] Based on the second objective optimization function and the fourth weight matrix corresponding to each decoder module, a fourth weight transfer parameter corresponding to each decoder module is determined.
[0025] Optionally, the optimizing the model to be quantized based on the weight transfer parameters corresponding to each weight matrix includes:
[0026] Multiplying the weight matrix by the corresponding weight transfer parameter respectively to perform weight transfer processing on the weight matrix;
[0027] Performing element-wise multiplication of a first learnable parameter in the RMSNorm module and an inverse vector of the first weight transfer parameter;
[0028] A second learnable parameter in the RMSNorm module is element-wise multiplied by an inverse vector of the second weight transfer parameter, where the first learnable parameter is different from the second learnable parameter.
[0029] The second embodiment of the present application provides a model quantization device, including:
[0030] A first determination module is used to determine each weight matrix to be optimized in the model to be quantized;
[0031] An acquisition module, configured to acquire a weight transfer parameter corresponding to each weight matrix based on a target optimization function associated with the weight matrix;
[0032] An optimization module, configured to optimize the model to be quantized based on the weight transfer parameters corresponding to each weight matrix;
[0033] The quantization module is used to perform quantization processing on the optimized model to be quantized to obtain a target model.
[0034] Optionally, the first determining module includes:
[0035] The first determining unit is configured to determine the weight matrices to be optimized according to the model structure of the model to be quantized.
[0036] Optionally, the determining unit is specifically configured to:
[0037] In the case where the model to be quantized includes a root mean square normalization module RMSNorm, the weight matrices to be optimized include:
[0038] The first weight matrices connected to the RMSNorm module and the second weight matrices not connected to the RMSNorm module.
[0039] Optionally, the model to be quantized contains multiple decoder modules, each of which contains an attention layer and a feedforward layer.
[0040] The first weight matrices include: a query weight matrix, a key weight matrix, and a value weight matrix contained in the attention layer, and a third weight matrix contained in the feedforward layer;
[0041] The second weight matrices include: the output weight matrix contained in the attention layer, and the fourth weight matrix contained in the feedforward layer, wherein the third weight matrix and the fourth weight matrix have different corresponding positions in the feedforward layer.
[0042] Optionally, the acquisition module includes:
[0043] a second determining unit, configured to determine a first objective optimization function associated with the first weight matrix and a second objective optimization function associated with the second weight matrix;
[0044] a third determining unit, configured to determine a first weight transfer parameter corresponding to each of the decoder modules based on the first objective optimization function and the query weight matrix, the key weight matrix, and the value weight matrix corresponding to each of the decoder modules;
[0045] a fourth determining unit, configured to determine a second weight transfer parameter corresponding to each of the decoder modules based on the first objective optimization function and a third weight matrix corresponding to each of the decoder modules;
[0046] a fifth determining unit, configured to determine a third weight transfer parameter corresponding to each of the decoder modules based on the second objective optimization function and the output weight matrix corresponding to each of the decoder modules;
[0047] A sixth determining unit is configured to determine a fourth weight transfer parameter corresponding to each of the decoder modules based on the second objective optimization function and a fourth weight matrix corresponding to each of the decoder modules.
[0048] Optionally, the optimization module includes:
[0049] a first processing unit, configured to multiply the weight matrix by corresponding weight transfer parameters respectively, so as to perform weight transfer processing on the weight matrix;
[0050] A second processing unit is configured to perform element-wise multiplication of a first learnable parameter in the RMSNorm module and an inverse vector of the first weight transfer parameter;
[0051] A third processing unit is configured to perform element-wise multiplication of a second learnable parameter in the RMSNorm module and an inverse vector of the second weight transfer parameter, where the first learnable parameter is different from the second learnable parameter.
[0052] Optionally, the first processing unit is specifically configured to:
[0053] Multiplying the query weight matrix, the key weight matrix, and the value weight matrix by the first weight transfer parameter element-by-element, column-by-column;
[0054] Multiplying the third weight matrix by the second weight transfer parameter element by element;
[0055] Multiplying the output weight matrix column by column by element by the third weight transfer parameter;
[0056] The fourth weight matrix is multiplied element-wise by the fourth weight transfer parameter column by column.
[0057] The third aspect embodiment of the present application proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the quantization method of the model of the embodiment of the present application.
[0058] The fourth aspect of the present application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the quantization method of the model disclosed in the embodiment of the present application.
[0059] In the disclosed embodiment, each weight matrix to be optimized in the model to be quantized is first determined, and then based on the target optimization function associated with the weight matrix, the weight transfer parameters corresponding to each weight matrix are obtained. Then, based on the weight transfer parameters corresponding to each weight matrix, the model to be quantized is optimized, and finally, the optimized model to be quantized is quantized to obtain the target model. Thus, by selecting appropriate weight transfer parameters and performing weight transfer on the weight matrix of the model to be quantized, outliers in the weights can be reduced, thereby reducing the accuracy loss after quantization, improving the model performance in the end-side scenario, and allowing the model to maintain its original accuracy and performance as much as possible after quantization.
[0060] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure, in which:
[0062] Figure 1 is a flowchart of a quantization method of a model provided according to the first embodiment of the present disclosure;
[0063] Figure 2 This is an application scenario diagram of the quantization method of the model proposed in this disclosure;
[0064] Figure 3 is a flowchart of a quantization method of a model provided according to the second embodiment of the present disclosure;
[0065] Figure 4 is a schematic diagram of performing weight transfer on a first weight matrix according to an embodiment of the present disclosure;
[0066] Figure 5 is a schematic diagram of performing weight transfer on a second weight matrix according to an embodiment of the present disclosure;
[0067] Figure 6 is a schematic diagram of a quantization device of a model according to an embodiment of the present disclosure;
[0068] Figure 7 It is a block diagram of an electronic device used to implement the quantization method of the model according to the embodiment of the present disclosure. DETAILED DESCRIPTION
[0069] Some embodiments of the present disclosure will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as becomes apparent after understanding the present disclosure, except for operations that must be performed in a specific order. In addition, for the sake of clarity and brevity, descriptions of features known in the art may be omitted.
[0070] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0071] It's important to note that while current quantization schemes can reduce storage and accelerate computation, they also suffer from precision loss. This precision loss is primarily due to the error in converting floating-point numbers to integers. This error is unavoidable, but outliers can amplify it. For example, consider the vector [0.2, 0.3, 0.8]. If this vector is mapped to a 1-bit integer (0 or 1), the quantization result is [0, 0, 1]. If one outlier is added: [0.2, 0.3, 0.8, 4.6], the mapping result becomes [0, 0, 0, 1], significantly increasing the error. The activation value refers to the output value of a neuron in a neural network layer after weighted summation and activation function processing.
[0072] Activation values are the output values of the activation function in a neural network. Outliers in activation values are extremely or abnormally large or small values relative to other activation values. Related techniques eliminate outliers in activation values by scaling them to the weights. However, outliers in weights can also reduce the model's inference accuracy.
[0073] Related solutions usually target situations where weights and activation values use the same data type, such as W8A8 (indicates that the data type of weights is int8 and the data type of activation values is int8). In this case, obvious outliers in the activation values are the main cause of accuracy degradation. However, when deploying large models on the client side, a quantization scheme of W4A16 (indicates that the data type of weights is int4 and the data type of activation values is int16) is generally adopted. This has the advantage of small model storage space and better inference results. However, in this configuration, the int4 weights become the main cause of accuracy loss. These technical solutions transfer outliers to the weights, which in turn amplifies the accuracy loss.
[0074] Among them, int4, int8, and int16 refer to 4-byte integer types, 8-byte integer types, and 2-byte integer types, respectively.
[0075] The embodiments of the present disclosure propose a model quantization method, device, electronic device, and storage medium, aiming to solve one of the technical problems in the related art at least to a certain extent.
[0076] It should be noted that the executor of the model quantization method of this embodiment can be a model quantization device, which can be implemented by software and / or hardware. The device can be configured in a cloud device, or it can also be a server, which is not limited here.
[0077] In the embodiments of the present disclosure, the quantization method of the model will be described by taking the “quantization device of the model” as the execution subject, and no limitation is given here.
[0078] Figure 1 FIG. 1 is a flow chart of a quantization method of a model according to the first embodiment of the present disclosure. Figure 1 As shown, the method includes:
[0079] S101: Determine each weight matrix to be optimized in the model to be quantized.
[0080] The model to be quantized may be a large model that requires significant storage space. In the disclosed embodiments, the model to be quantized may be a large model that needs to be deployed on the device. To reduce the storage space occupied by the large model on the device and to ensure its accuracy, it needs to be quantized.
[0081] It should be noted that the model to be quantized can be a large language model in the direction of natural language processing, or a large model for computer vision (for image processing), or a large model for speech processing, etc., and there is no limitation here.
[0082] As an implementation method, in the embodiment of the present disclosure, a large model with a transformer architecture can be used as the model to be quantized, such as the LLaMA model (Large Language Model Meta AI), which is not limited here.
[0083] The weight matrices to be optimized may be parameters in the quantized model that require outlier transfer. As an example, the weight matrices to be optimized may include the weight parameter matrix in the attention layer and the weight parameter matrix in the feedforward layer, which are not limited here.
[0084] S102: Based on the target optimization function associated with the weight matrix, obtain the weight transfer parameters corresponding to each weight matrix.
[0085] Among them, the weight transfer parameter is used to transfer the parameters of the model to remove outliers in the weight matrix, thereby reducing the accuracy loss after quantization.
[0086] It can be understood that during the model optimization process, the weight transfer parameters can maintain the original function of the model to be quantized, maintain similar prediction results after quantization, and greatly reduce the error.
[0087] The objective optimization function is the objective function used to select the weight transfer parameters. It should be noted that the objective optimization functions associated with different weight matrices can be the same or different.
[0088] For example, if each weight matrix includes A1, A2, A3, B1, and B2, the target optimization functions corresponding to A1, A2, and A3 can be the same target optimization function M1, and the target optimization functions corresponding to B1 and B2 can be the same target optimization function M2, which is not limited here.
[0089] It can be understood that by selecting the optimal value of the weight transfer parameter through the objective optimization function corresponding to each weight matrix, the weight transfer parameter finally selected can minimize the error in the model prediction accuracy before and after quantization.
[0090] Optionally, the weight transfer parameter ranges from 0 to 1.
[0091] S103: Optimizing the model to be quantized based on the weight transfer parameters corresponding to each weight matrix.
[0092] Specifically, each weight matrix in the model to be quantized may be multiplied element-by-element by the weight transfer parameter column by column to achieve weight transfer of the weight matrix, thereby optimizing the model parameters of the model to be quantized.
[0093] Alternatively, you can refer to the following formula: Among them, S is the weight transfer parameter, W x As the weight matrix, W x Multiplying S element by element column by column can get the parameter W after weight transfer y .
[0094] S104: quantizing the optimized model to be quantized to obtain a target model.
[0095] The target model can be the quantized model to be quantized. The target model can be deployed on the client side (such as mobile devices such as mobile phones and tablets, or IoT devices) or other low-precision hardware platforms, and can maintain the model prediction accuracy of the original quantized model.
[0096] As a possible implementation method, after the optimization of the quantization model is completed, the parameters and activation values of the model are quantized using post-training quantization technology (PTQ). The quantization error can be reduced by determining appropriate quantization parameters, such as quantization step size and bias calibration.
[0097] As a possible implementation method, the target model can be a quantized model in the "W4A16" format, that is, a model quantized using the "W4A16" quantization scheme. This target model can run efficiently on a low-precision hardware platform while trying to maintain the performance of a high-precision model to meet the needs of efficiently deploying deep learning models in a resource-constrained environment.
[0098] Among them, "W4" means that the weights are quantized into 4-bit fixed-point numbers, that is, the weight parameters in the model will be stored in the form of 4-bit binary numbers. Compared with 32-bit floating-point numbers (fp32), this can greatly compress the size of the model, which is conducive to deploying the model on devices with limited memory and storage space (such as mobile devices or IoT devices), while also speeding up the calculation speed.
[0099] Among them, "A16" means that the activation value (Activations) is quantized into a 16-bit fixed-point number. The activation value is the output result of each layer in the neural network and is generated during the forward propagation process. In many application scenarios, it can effectively reduce the computational and storage burden while meeting the accuracy requirements.
[0100] like Figure 2 As shown, Figure 2 This is an application scenario diagram proposed in the present disclosure. First, the large model is trained in the cloud. The data type of the model parameters of the large model is the fp32 (32-bit floating point) data type. After the large model is quantized, the data type of the model parameters is converted from fp32 to int4 (4-bit integer) and can be deployed on the end side for end-side reasoning.
[0101] In the disclosed embodiment, each weight matrix to be optimized in the model to be quantized is first determined, and then based on the target optimization function associated with the weight matrix, the weight transfer parameters corresponding to each weight matrix are obtained. Then, based on the weight transfer parameters corresponding to each weight matrix, the model to be quantized is optimized, and finally, the optimized model to be quantized is quantized to obtain the target model. Thus, by selecting appropriate weight transfer parameters and performing weight transfer on the weight matrix of the model to be quantized, outliers in the weights can be reduced, thereby reducing the accuracy loss after quantization, improving the model performance in the end-side scenario, and allowing the model to maintain its original accuracy and performance as much as possible after quantization.
[0102] Figure 3 FIG. 1 is a flow chart of a quantization method of a model according to the second embodiment of the present disclosure. Figure 3 As shown, the method includes:
[0103] S201: Determine each weight matrix to be optimized according to the model structure of the model to be quantized.
[0104] It should be noted that the model structure can be the model architecture of the model to be quantized. For large models with different architectures, the corresponding weight matrices to be optimized may be different, or they may be the same. Therefore, when determining the weight matrix to be optimized in the large model, it is necessary to determine it based on the specific architecture of the large model.
[0105] Optionally, when the model to be quantized includes a root mean square normalization module RMSNorm, the weight matrices to be optimized include: the first weight matrices connected after the RMSNorm module, and the second weight matrices not connected to the RMSNorm module.
[0106] Among them, the RMSNorm (Root Mean Square Normalization) module is a normalization layer in deep learning, which is used to standardize the features within the neural network. It adjusts the data distribution by calculating the root mean square of the input data, helping the network to converge more stably and quickly during training.
[0107] The first weight matrix may be a weight matrix connected after the RMSNorm module, and the second weight matrix may be a weight matrix not connected to the RMSNorm module.
[0108] In the disclosed embodiments, the LLaMA model, a large-scale model used in fields such as natural language processing and computer vision, is used as an example of a model to be quantized. The LLaMA model includes an RMSNorm module. The model to be quantized includes multiple decoder modules, each of which includes an attention layer and a feedforward layer.
[0109] Optionally, the feedforward layer in the decoder module uses a gated linear unit with a Swish activation function, as follows:
[0110]
[0111] Where x is the input vector, W1 is the weight coefficient in the first linear transformation layer in the feedforward layer, W2 is the weight coefficient in the second linear transformation layer in the feedforward layer, and W3 is the additional weight coefficient in the feedforward layer, which is used to adjust the output y. In this formula, the input vector x first undergoes a linear transformation using the first weight matrix W1, then the Swish activation function is applied, followed by a linear transformation using the second weight matrix W2, and finally multiplied by W3 to obtain the output vector y.
[0112] In the embodiment of the present disclosure, W1 and W2 in the above formula can be used as the third weight matrix. W1 in the above formula is the weight matrix from the input layer to the hidden layer, and W2 is the weight matrix from the hidden layer to the output layer. Optionally, W3 in the above formula can be used as the fourth weight matrix.
[0113] The fourth weight matrix may be an additional parameter in the feedforward neural network, used for the final linear transformation or output adjustment. The third weight matrix and the fourth weight matrix may have different corresponding positions in the feedforward layer.
[0114] Among them, each first weight matrix includes: a query weight matrix, a key weight matrix and a value weight matrix contained in the attention layer, and a third weight matrix contained in the feedforward layer.
[0115] Among them, each second weight matrix includes: the output weight matrix contained in the attention layer, and the fourth weight matrix contained in the feedforward layer.
[0116] Among them, the attention layer can be a multi-head self-attention layer.
[0117] Among them, the formula corresponding to the multi-head self-attention layer is as follows:
[0118]
[0119] MultiHead(Q,K,V)=Concat(head1,head2,...,headn)W O
[0120] Among them, W O is the output weight matrix.
[0121] It should be noted that in the attention layer of the Transformer architecture, the query matrix (Q), key matrix (K), and value matrix (V) can be obtained by multiplying the input data with the weight matrix learned by the model. Specifically, for the word embedding vector at each position in the input sequence, the model uses a different weight matrix (W Q 、W K 、W V ) is converted into the corresponding query vector, key vector and value vector.
[0122] Among them, the query matrix is the word embedding of the input sequence through the query weight matrix W Q The key matrix is the word embedding of the input sequence through the key weight matrix W K The value matrix is obtained by transforming the word embedding of the input sequence through the value weight matrix W V Obtained by transformation.
[0123] In the embodiment of the present disclosure, the query weight matrix W in the attention layer connected after RMSNorm can be Q , key weight matrix W K Sum value weight matrix W V as the first weight matrix.
[0124] In the multi-head self-attention layer, the outputs of each attention head can be concatenated and then linearly transformed through an additional output weight matrix (also called a merge matrix or projection matrix) to generate the final attention output. In the disclosed embodiment, the output weight matrix in the multi-head self-attention layer can be used as the second weight matrix.
[0125] S202: Determine a first objective optimization function associated with the first weight matrix and a second objective optimization function associated with the second weight matrix.
[0126] The first objective optimization function is the objective optimization function corresponding to the first weight matrix, and is used to determine the weight transfer parameters that match the first weight matrix.
[0127] As an example, the formula of the first objective optimization function may be:
[0128]
[0129] Where X represents the input vector, W m represents the first weight matrix, s m represents the weight transfer parameter corresponding to the first weight matrix (range between 0 and 1). int16 Indicates that the result of quantizing the input to int16 is then dequantized to a value of type fp, Q int4 Indicates that the result of quantizing the input to int4 is then dequantized to a value of type fp.
[0130] The second objective optimization function is the objective optimization function corresponding to the second weight matrix, and is used to determine the weight transfer parameters that match the second weight matrix.
[0131] As an example, the formula of the second objective optimization function may be:
[0132]
[0133] Where X represents the input vector, W n represents the second weight matrix, s n Represents the weight transfer parameter corresponding to the second weight matrix (ranges between 0 and 1).
[0134] S203: Determine a first weight transfer parameter corresponding to each decoder module based on the first objective optimization function and the query weight matrix, key weight matrix, and value weight matrix corresponding to each decoder module.
[0135] It should be noted that the model to be quantized consists of multiple layers of decoder modules, and the weight matrix in each layer of decoder modules has corresponding weight transfer parameters.
[0136] Optionally, the first weight transfer parameter corresponding to each decoder module can be determined by using the first objective optimization function and the query weight matrix, key weight matrix and value weight matrix corresponding to each decoder module.
[0137] The first weight transfer parameter is a weight transfer parameter associated with the query weight matrix, the key weight matrix, and the value weight matrix.
[0138] Optionally, the query weight matrix, key weight matrix and value weight matrix can be substituted into the first objective optimization function, and the weight transfer parameters in the first objective optimization function can be iterated until the weight transfer parameters that can make the first objective optimization function achieve the minimum value are determined and used as the first weight transfer parameters.
[0139] S204: Determine a second weight transfer parameter corresponding to each decoder module based on the first objective optimization function and the third weight matrix corresponding to each decoder module.
[0140] The second weight transfer parameter is a weight transfer parameter associated with the third weight matrix.
[0141] Optionally, the third weight matrix can be substituted into the first objective optimization function, and the weight transfer parameters in the first objective optimization function can be iterated until the weight transfer parameters that can make the first objective optimization function achieve the minimum value are determined and used as the second weight transfer parameters.
[0142] S205: Determine a third weight transfer parameter corresponding to each decoder module based on the second objective optimization function and the output weight matrix corresponding to each decoder module.
[0143] The third weight transfer parameter is a weight transfer parameter associated with the output weight matrix.
[0144] Optionally, the output weight matrix can be substituted into the second objective optimization function, and the weight transfer parameters in the second objective optimization function can be iterated until the weight transfer parameters that can make the second objective optimization function achieve the minimum value are determined and used as the third weight transfer parameters.
[0145] S206: Determine a fourth weight transfer parameter corresponding to each decoder module based on the second objective optimization function and the fourth weight matrix corresponding to each decoder module.
[0146] The fourth weight transfer parameter is a weight transfer parameter associated with the fourth weight matrix.
[0147] Optionally, the fourth weight matrix can be substituted into the second objective optimization function, and the weight transfer parameters in the second objective optimization function can be iterated until a weight transfer parameter that can make the second objective optimization function achieve the minimum value is determined and used as the fourth weight transfer parameter.
[0148] S207: Multiply the weight matrix by the corresponding weight transfer parameter respectively to perform weight transfer processing on the weight matrix.
[0149] Optionally, the query weight matrix, key weight matrix and value weight matrix can be multiplied element-by-element with the first weight transfer parameter column by column, the third weight matrix can be multiplied element-by-element with the second weight transfer parameter column by column, the output weight matrix can be multiplied element-by-element with the third weight transfer parameter column by column, and the fourth weight matrix can be multiplied element-by-element with the fourth weight transfer parameter column by column.
[0150] Optional, the specific formula is:
[0151]
[0152]
[0153]
[0154]
[0155] in, Indicates element-by-element multiplication, the matrix is multiplied column by column and element by element by s, s1 is the first weight transfer parameter, s2 is the second weight transfer parameter, s3 is the third weight transfer parameter, s4 is the fourth weight transfer parameter, W Q ,W K and W V are the query weight matrix, key weight matrix, and value weight matrix respectively. W1 and W2 are the third weight matrix, W3 is the fourth weight matrix, and W O is the output weight matrix.
[0156] like Figure 4 As shown, Figure 4 A schematic diagram of weight transfer for a first weight matrix.
[0157] Where X is the input vector, which enters the RMSNorm module for normalization to obtain vector Y, then vector Y is multiplied by the learnable parameter to obtain vector Z, and then vector Z is multiplied by the first weight matrix to obtain the output. Figure 4 The W in can be W Q 、W K 、W V , W1 and W2, etc. When optimizing the quantized model, it is necessary to multiply the first weight matrix column by column and the corresponding weight transfer parameter s element by element.
[0158] like Figure 5 As shown, Figure 5 A schematic diagram of weight transfer for a second weight matrix.
[0159] Where X is the input vector, and the output Y can be obtained by multiplying the input vector X with the second weight matrix W. Figure 5 The W in can be W O And W3 and other second weight matrices. When optimizing the quantized model, it is necessary to multiply the second weight matrix column by column and the corresponding weight transfer parameter s element by element, that is, Figure 5 shown In addition, the output vector Y and the inverse vector s of the weight transfer parameter s need to be -1 Multiply.
[0160] S208: Multiply the first learnable parameter in the RMSNorm module and the inverse vector of the first weight transfer parameter element by element.
[0161] Among them, before the input vector is input into each decoder module for calculation, it needs to be normalized using RMSNorm. The formula is as follows:
[0162]
[0163] Among them, g i is a learnable parameter. In the embodiment of the present disclosure, g i can be a scaling parameter, Rms() is the root mean square function, n is the dimension of the input vector, a i is any input vector i.
[0164] Among them, the first learnable parameter can be a learnable parameter corresponding to the attention layer.
[0165] Specifically, it can be calculated according to the following formula:
[0166]
[0167] Among them, s1 is the first weight transfer parameter, is the inverse vector of the first weight transfer parameter, g1 is the first learnable parameter, and g1' is the processed first learnable parameter.
[0168] S209: performing element-by-element multiplication of a second learnable parameter in the RMSNorm module and an inverse vector of a second weight transfer parameter, where the first learnable parameter and the second learnable parameter are different.
[0169] The second learnable parameter may be a learnable parameter corresponding to the feedforward layer.
[0170] Specifically, it can be calculated according to the following formula:
[0171]
[0172] Among them, s2 is the second weight transfer parameter, is the inverse vector of the second weight transfer parameter, g2 is the second learnable parameter, and g2' is the processed second learnable parameter.
[0173] like Figure 4 As shown in Figure 2, since the learnable parameters are then quantized to int16 and are not affected by outliers, the outliers in the first weight matrix can be transferred to the learnable parameters g in RMSNorm. i middle.
[0174] S210: performing quantization processing on the optimized model to be quantized to obtain a target model.
[0175] It should be noted that the specific implementation of step S210 can refer to the above embodiment and will not be described in detail here.
[0176] In an embodiment of the present disclosure, first, according to the model structure of the model to be quantized, each weight matrix to be optimized is determined, then a first objective optimization function associated with the first weight matrix and a second objective optimization function associated with the second weight matrix are determined, then based on the first objective optimization function and the query weight matrix, key weight matrix and value weight matrix corresponding to each decoder module, a first weight transfer parameter corresponding to each decoder module is determined, then based on the first objective optimization function and the third weight matrix corresponding to each decoder module, a second weight transfer parameter corresponding to each decoder module is determined, then based on the second objective optimization function and the output weight matrix corresponding to each decoder module, a third weight transfer parameter corresponding to each decoder module is determined, then based on the second objective optimization function and the fourth weight matrix corresponding to each decoder module, a fourth weight transfer parameter corresponding to each decoder module is determined, then the weight matrices are multiplied by the corresponding weight transfer parameters respectively to perform weight transfer processing on the weight matrices, then the first learnable parameter in the RMSNorm module is multiplied by the inverse vector of the first weight transfer parameter element-by-element, then the second learnable parameter in the RMSNorm module is multiplied by the inverse vector of the second weight transfer parameter element-by-element, and the first learnable parameter and the second learnable parameter in the RMSNorm module are multiplied by the inverse vector of the second weight transfer parameter element-by-element, where the first learnable parameter and the second learnable parameter are different. Therefore, the weight transfer coefficients associated with each weight matrix are first accurately determined. Since they are determined by the corresponding target optimization function, the model parameters are quantization-friendly. The outliers in the first weight matrix are transferred to the learnable parameters in the RMSNorm module. This ensures that the model has lower quantization accuracy loss after quantization, and has better model performance in the end-side scenario.
[0177] Figure 6 FIG. 1 is a schematic diagram of a quantization device of a model according to another embodiment of the present disclosure. Figure 6 As shown, the quantization device 600 of the model includes:
[0178] A first determining module 610 is used to determine each weight matrix to be optimized in the model to be quantized;
[0179] An acquisition module 620 is configured to acquire a weight transfer parameter corresponding to each weight matrix based on a target optimization function associated with the weight matrix;
[0180] An optimization module 630, configured to optimize the model to be quantized based on the weight transfer parameters corresponding to each weight matrix;
[0181] The quantization module 640 is configured to perform quantization processing on the optimized model to be quantized to obtain a target model.
[0182] Optionally, the first determining module includes:
[0183] The first determining unit is configured to determine the weight matrices to be optimized according to the model structure of the model to be quantized.
[0184] Optionally, the determining unit is specifically configured to:
[0185] In the case where the model to be quantized includes a root mean square normalization module RMSNorm, the weight matrices to be optimized include:
[0186] The first weight matrices connected to the RMSNorm module and the second weight matrices not connected to the RMSNorm module.
[0187] Optionally, the model to be quantized contains multiple decoder modules, each of which contains an attention layer and a feedforward layer.
[0188] The first weight matrices include: a query weight matrix, a key weight matrix, and a value weight matrix contained in the attention layer, and a third weight matrix contained in the feedforward layer;
[0189] The second weight matrices include: the output weight matrix contained in the attention layer, and the fourth weight matrix contained in the feedforward layer, wherein the third weight matrix and the fourth weight matrix have different corresponding positions in the feedforward layer.
[0190] Optionally, the acquisition module includes:
[0191] a second determining unit, configured to determine a first objective optimization function associated with the first weight matrix and a second objective optimization function associated with the second weight matrix;
[0192] a third determining unit, configured to determine a first weight transfer parameter corresponding to each of the decoder modules based on the first objective optimization function and the query weight matrix, the key weight matrix, and the value weight matrix corresponding to each of the decoder modules;
[0193] a fourth determining unit, configured to determine a second weight transfer parameter corresponding to each of the decoder modules based on the first objective optimization function and a third weight matrix corresponding to each of the decoder modules;
[0194] a fifth determining unit, configured to determine a third weight transfer parameter corresponding to each of the decoder modules based on the second objective optimization function and the output weight matrix corresponding to each of the decoder modules;
[0195] A sixth determining unit is configured to determine a fourth weight transfer parameter corresponding to each of the decoder modules based on the second objective optimization function and a fourth weight matrix corresponding to each of the decoder modules.
[0196] Optionally, the optimization module includes:
[0197] a first processing unit, configured to multiply the weight matrix by corresponding weight transfer parameters respectively, so as to perform weight transfer processing on the weight matrix;
[0198] A second processing unit is configured to perform element-wise multiplication of a first learnable parameter in the RMSNorm module and an inverse vector of the first weight transfer parameter;
[0199] A third processing unit is configured to perform element-wise multiplication of a second learnable parameter in the RMSNorm module and an inverse vector of the second weight transfer parameter, where the first learnable parameter is different from the second learnable parameter.
[0200] Optionally, the first processing unit is specifically configured to:
[0201] Multiplying the query weight matrix, the key weight matrix, and the value weight matrix by the first weight transfer parameter element-by-element, column-by-column;
[0202] Multiplying the third weight matrix by the second weight transfer parameter element by element;
[0203] Multiplying the output weight matrix column by column by element by the third weight transfer parameter;
[0204] The fourth weight matrix is multiplied element-wise by the fourth weight transfer parameter column by column.
[0205] In the disclosed embodiment, each weight matrix to be optimized in the model to be quantized is first determined. Then, based on the target optimization function associated with the weight matrix, the weight transfer parameters corresponding to each weight matrix are obtained. Then, based on the weight transfer parameters corresponding to each weight matrix, the model to be quantized is optimized. Finally, the optimized model to be quantized is quantized to obtain the target model. Thus, by selecting appropriate weight transfer parameters and performing weight transfer on the weight matrix of the model to be quantized, outliers in the weights can be reduced, thereby reducing the accuracy loss after quantization and improving the model performance in the end-side scenario.
[0206] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0207] Figure 7 A block diagram of an exemplary computer device suitable for implementing embodiments of the present application is shown. Figure 7 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0208] like Figure 7 As shown, computer device 12 is implemented as a general-purpose computing device. Components of computer device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).
[0209] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of such architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnection (PCI) bus.
[0210] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0211] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 7 Not shown, often called a "hard drive").
[0212] although Figure 7 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a Compact Disc Read Only Memory (hereinafter referred to as: CD-ROM), a Digital Video Disc Read Only Memory (hereinafter referred to as: DVD-ROM), or other optical media) may be provided. In these cases, each drive can be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present application.
[0213] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally implement the functions and / or methods of the embodiments described herein.
[0214] The computer device 12 can also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). This communication can occur via an input / output (I / O) interface 22. Furthermore, the computer device 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the computer device 12 via a bus 18. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the computer device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0215] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28 , such as implementing the quantization method of the model mentioned in the above embodiment.
[0216] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0217] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
[0218] It should be noted that, in the description of this application, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.
[0219] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0220] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0221] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0222] In addition, the functional units in the various embodiments of the present application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disk, etc.
[0223] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0224] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A quantization method for a model, characterized in that: include: Determine each weight matrix to be optimized in the model to be quantified; Obtaining a weight transfer parameter corresponding to each weight matrix based on a target optimization function associated with the weight matrix; Optimizing the model to be quantized based on the weight transfer parameters corresponding to each weight matrix; The optimized model to be quantized is quantized to obtain a target model.
2. The method according to claim 1, characterized in that The step of determining each weight matrix to be optimized in the model to be quantized includes: The weight matrices to be optimized are determined according to the model structure of the model to be quantized.
3. The method according to claim 2, characterized in that The step of determining the weight matrices to be optimized according to the model structure of the model to be quantized includes: In the case where the model to be quantized includes a root mean square normalization module RMSNorm, the weight matrices to be optimized include: The first weight matrices connected to the RMSNorm module and the second weight matrices not connected to the RMSNorm module.
4. The method according to claim 3, characterized in that in, The model to be quantized contains multiple decoder modules, each of which contains an attention layer and a feedforward layer. The first weight matrices include: a query weight matrix, a key weight matrix, and a value weight matrix contained in the attention layer, and a third weight matrix contained in the feedforward layer; The second weight matrices include: the output weight matrix contained in the attention layer, and the fourth weight matrix contained in the feedforward layer, wherein the third weight matrix and the fourth weight matrix have different corresponding positions in the feedforward layer.
5. The method according to claim 4, characterized in that The obtaining of the weight transfer parameters corresponding to each weight matrix based on the objective optimization function associated with the weight matrix includes: determining a first objective optimization function associated with the first weight matrix and a second objective optimization function associated with the second weight matrix; Determining a first weight transfer parameter corresponding to each decoder module based on the first objective optimization function and a query weight matrix, a key weight matrix, and a value weight matrix corresponding to each decoder module; Determining a second weight transfer parameter corresponding to each decoder module based on the first objective optimization function and a third weight matrix corresponding to each decoder module; Determining a third weight transfer parameter corresponding to each decoder module based on the second objective optimization function and the output weight matrix corresponding to each decoder module; Based on the second objective optimization function and the fourth weight matrix corresponding to each decoder module, a fourth weight transfer parameter corresponding to each decoder module is determined.
6. The method according to claim 5, characterized in that The optimizing the model to be quantized based on the weight transfer parameters corresponding to each weight matrix includes: Multiplying the weight matrix by the corresponding weight transfer parameter respectively to perform weight transfer processing on the weight matrix; Performing element-wise multiplication of a first learnable parameter in the RMSNorm module and an inverse vector of the first weight transfer parameter; A second learnable parameter in the RMSNorm module is element-wise multiplied by an inverse vector of the second weight transfer parameter, where the first learnable parameter is different from the second learnable parameter.
7. The method according to claim 6, characterized in that The multiplying the weight matrix by the corresponding weight transfer parameter to perform weight transfer processing on the weight matrix includes: Multiplying the query weight matrix, the key weight matrix, and the value weight matrix by the first weight transfer parameter element-by-element, column-by-column; Multiplying the third weight matrix by the second weight transfer parameter element by element; Multiplying the output weight matrix column by column by element by the third weight transfer parameter; The fourth weight matrix is multiplied element-wise by the fourth weight transfer parameter column-by-column.
8. A quantization device for a model, characterized in that: include: A first determination module is used to determine each weight matrix to be optimized in the model to be quantized; An acquisition module, configured to acquire a weight transfer parameter corresponding to each weight matrix based on a target optimization function associated with the weight matrix; An optimization module, configured to optimize the model to be quantized based on the weight transfer parameters corresponding to each weight matrix; The quantization module is used to perform quantization processing on the optimized model to be quantized to obtain a target model.
9. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.