Model quantization method and apparatus, device, and medium
By filtering and quantifying the weight parameters in the model to be quantized whose importance is less than the preset threshold, the problem of accuracy loss caused by model quantization is solved, and the effect of maintaining model accuracy while reducing resource usage is achieved.
Patent Information
- Application Number
- PCT/CN2024/130524
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-11-07
- Publication Date
- 2025-06-26
AI Technical Summary
When existing model quantization methods reduce model complexity and resource usage, they often lead to loss of model accuracy and even make the model unavailable. How to reduce the impact on model accuracy while meeting the needs of quantitative use has become an urgent problem.
By obtaining each weight parameter in the model to be quantified, calculating its importance, and filtering out the weight parameters whose importance is less than the preset threshold as the target weights, quantifying these target weights to obtain the quantized model.
While meeting the needs of quantitative use, the accuracy loss caused by model quantization is reduced, so that the quantized model maintains high accuracy while reducing resource usage.
Smart Images

Figure CN2024130524_26062025_PF_FP_ABST
Abstract
Description
Model quantification method, device, equipment and medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311762021.2 and invention name “A model quantization method, device, equipment and medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of artificial intelligence technology, and in particular to a model quantization method, device, equipment and medium. Background Art
[0003] Currently, with the continuous development of deep learning technology, deep neural network models are widely used in many fields such as image classification and natural language processing, and have achieved very good results. However, in actual deployment, there are still problems such as large models, high computational complexity, and high hardware cost requirements. To address these problems, model quantization has achieved good results. It quantizes the model from floating-point type to fixed-point type. Although traditional neural network model quantization methods can reduce model complexity and model resource usage, because quantization modifies the accuracy of model weights, biases, or activation values, it often causes a certain degree of accuracy loss and may even make the model unusable. Therefore, how to reduce the accuracy loss caused by model quantization so that the model meets the requirements of quantization while having a smaller impact on model accuracy has become an urgent problem to be solved.
[0004] Summary of the Invention
[0005] Based on this, a model quantization method, device, equipment and medium are provided to solve the problem of how to reduce the accuracy loss caused by model quantization so that the model can meet the needs of quantitative use while having less impact on model accuracy.
[0006] In a first aspect, an embodiment of the present invention provides a model quantization method, the method comprising the following steps:
[0007] Obtain each weight parameter in any network layer to be quantized in the model to be quantized;
[0008] Calculating the importance of each weight parameter, and selecting a weight parameter whose importance is less than a preset threshold from all weight parameters as a target weight;
[0009] The target weight in the network layer to be quantized is quantized to obtain a quantized network layer, and all the network layers to be quantized are traversed to obtain a quantized model.
[0010] In a second aspect, an embodiment of the present invention provides a model quantization device, the device comprising:
[0011] The first acquisition module is used to obtain each weight parameter in any network layer to be quantized in the model to be quantized;
[0012] A screening module is used to calculate the importance of each weight parameter, and screen out the weight parameters whose importance is less than a preset threshold from all weight parameters as target weights;
[0013] The first quantization module is used to quantize the target weight in the network layer to be quantized to obtain a quantized network layer, and traverse all the network layers to be quantized to obtain a quantized model.
[0014] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned model quantization method when executing the computer program.
[0015] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned model quantization method are implemented.
[0016] The present invention achieves a technical effect that is different from existing solutions: the present invention obtains each weight parameter in any network layer to be quantized in the model to be quantized, and based on the importance of each weight parameter, selects weight parameters with an importance less than a preset threshold as target weights, quantizes the target weights, and obtains a quantized model. According to the importance of each weight parameter, only weight parameters with an importance less than a preset threshold are quantized, while weights with higher importance are not quantized. This reduces the accuracy loss caused by model quantization to a certain extent, allowing the model to meet the requirements of quantization while having a smaller impact on model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] FIG1 is a schematic diagram of an application environment of a model quantization method provided in Example 1 of the present invention;
[0018] FIG2 is a schematic diagram of a flow chart of a model quantization method provided in Example 1 of the present invention;
[0019] FIG3 is a flow chart of a model quantization method provided in a second embodiment of the present invention;
[0020] FIG4 is a schematic diagram of a flow chart of a model quantization method provided in Example 3 of the present invention;
[0021] FIG5 is a schematic diagram of a flow chart of a model quantization method provided in a fourth embodiment of the present invention;
[0022] FIG6 is a flow chart of a model quantization method provided in a fifth embodiment of the present invention;
[0023] FIG7 is a flow chart of a model quantization method provided in accordance with a sixth embodiment of the present invention;
[0024] FIG8 is a flow chart of a model quantization method provided in Embodiment 8 of the present invention;
[0025] FIG9 is a schematic flow chart of a model quantization method provided in Embodiment 9 of the present invention;
[0026] FIG10 is a schematic flow chart of a model quantization method provided in Example 10 of the present invention;
[0027] FIG11 is a schematic structural diagram of a model quantization device provided in Example 11 of the present invention;
[0028] FIG12 is a schematic structural diagram of a computer device provided in accordance with a twelfth embodiment of the present invention. DETAILED DESCRIPTION
[0029] Referring to FIG1 , which is a schematic diagram of an application environment of a model quantization method provided in a first embodiment of the present invention, the model quantization method is to quantize the model to obtain a quantized model, so that the quantized model can meet the needs of quantitative use while having little impact on the model accuracy. The quantized model can be configured in the client or the server, and both the client and the server can provide model services to users. Of course, when the server is configured with the quantized model, the client communicating with the server can apply for model services from the server through the network. The client includes but is not limited to PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, server-side computer devices, personal digital assistants (PDAs), and other devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers, which can be specifically embodied as a computer device.
[0030] 2 is a flow chart of a model quantization method provided in Embodiment 1 of the present invention. As shown in FIG2 , the model quantization method includes the following steps:
[0031] Step S201: Obtain each weight parameter in any network layer to be quantized in the model to be quantized.
[0032] In this embodiment, the model to be quantized is a model configured on the server side of the application environment shown in FIG1 . The model to be quantized can be a trained neural network model or a neural network model in the training process. The model to be quantized can be quantized and compressed to convert high-bit numerical calculations into low-bit numerical calculations, thereby enabling the model to be used for lightweight requirements.
[0033] According to the network architecture of the model to be quantized, each network layer in the architecture can be decomposed. For different network architectures, the number of network layers in the model and the settings of weight parameters may be different. For example, for a general convolutional neural network (CNN), its network layers may include an input layer, a convolution layer, a pooling layer, an activation function layer, and a fully connected layer. Among them, the present invention can quantize any network layer in the model as a layer to be quantized. Of course, in one embodiment, only a specific network layer in the model to be quantized can be quantized without quantizing other network layers, for example, only the above-mentioned activation function layer can be quantized.
[0034] The weight parameter may refer to a parameter used to calculate the relationship between the input and output of the model. The role of the weight parameter is to adjust the input of the model to obtain the corresponding output. In the process of obtaining the weight parameters of the network layer to be quantized, the corresponding network layer to be quantized for which the weight parameters need to be obtained can be located according to the model structure of the model to be quantized, the network layer to be quantized can be accessed, and the weight parameters in the network layer to be quantized can be extracted and saved in a preset storage file or storage device in the form of a matrix or binary. In other embodiments, after locating the network layer to be quantized for which the weight parameters need to be obtained, the weight parameters in the network layer to be quantized can be directly accessed, and the weight parameters can be quantized in the network layer to be quantized.
[0035] Step S202: Calculate the importance of each weight parameter, and select the weight parameter whose importance is less than a preset threshold from all weight parameters as the target weight.
[0036] In this embodiment, the importance of the weight parameter may refer to the size of the role played by the weight parameter in the network layer of the model, or the degree of influence of the change of the weight parameter on the output of the network layer of the model. For example, for a convolution layer of a trained model, the weight parameter in its convolution kernel is more important and can affect the comprehensiveness of the model's feature analysis of the input data. Therefore, the weight parameter in the convolution kernel can be regarded as an important parameter.
[0037] The importance of the weight parameters can be calculated according to a preset importance analysis strategy, wherein the importance analysis strategy can be a mapping table, which stores the mapping relationship between weight parameters and importance, so that the importance of each weight parameter can be calculated through the mapping table. Currently, the mapping table is configured based on modeling requirements when establishing the above-mentioned model; in addition, the importance analysis strategy can be a numerical calculation process based on logical analysis, that is, analyzing the usage results of each weight parameter in the network layer of the model to obtain the size of the role played by the weight parameter; in other embodiments, each weight parameter can be changed separately through experimental means, and the change in output caused by the change in the corresponding weight parameter when the input remains unchanged is calculated. If the change in output is greater, the change in the corresponding weight parameter has a greater impact, which means that the weight parameter is more important. Of course, the experimental means can be used in combination with the above-mentioned mapping table, that is, experimental means are used to determine the importance of each weight parameter in the mapping table.
[0038] [Corrected 12.12.2024 according to Rule 91] A screen is performed for each of the above-mentioned weight parameters, wherein the screening is performed based on the degree of importance. The greater the degree of importance, the greater the impact of the corresponding weight parameter on the model after being quantized, and the smaller the degree of importance, the smaller the impact of the corresponding weight parameter on the model after being quantized. Therefore, this embodiment provides a preset threshold value, which can be set according to demand. That is, the larger the preset threshold value is set, the lower the requirement for the degree of quantization of the model (that is, the fewer weight parameters that can be quantized), and the higher the requirement for high precision of the model. The lower the preset threshold value is set, the greater the requirement for the degree of quantization of the model (that is, the more weights that can be quantized), and the lower the requirement for high precision of the model. In addition, the preset threshold value is a value that can be compared with the degree of importance. If the degree of importance of the weight parameter is in the interval of [1,100], the preset threshold value can be a value within the interval. If the degree of importance of the weight parameter is relatively dispersed, the degree of importance can also be normalized so that it is between [0,1]. The preset threshold value is a value within the interval of [0,1].
[0039] Step S203: quantize the target weights in the network layer to be quantized to obtain a quantized network layer, and traverse all the network layers to be quantized to obtain a quantized model.
[0040] In this embodiment, quantization may refer to converting high-bit values into low-bit values to reduce the complexity of model calculations. For example, quantization converts a floating-point model into a fixed-point model. For example, if the weights in the original model are float32 floating-point, quantization converts the model weights into fixed-point int8.
[0041] The target weight obtained in step S202 is quantized. The quantization can adopt methods such as symmetric quantization, asymmetric quantization, channel quantization, group-by-group quantization, layer-by-layer quantization, combined quantization, and fixed-point quantization. For example, the process of quantizing the target weight using the symmetric quantization method can be as follows: first, determining the target number of quantization bits and the quantization range of the target weight, wherein the quantization range is a symmetric quantization range, that is, the quantization range is divided into two parts, one part is a positive range, and the other part is a negative range. Then, according to the maximum and minimum values of the target weight and the quantization range, a quantization parameter and a zero point are calculated, wherein the quantization parameter = (maximum value of the quantization range - minimum value of the quantization range) / (maximum value of the target weight - minimum value of the target weight), and the zero point can be the average value of the maximum and minimum values of the quantization range. Finally, the target weight is quantized according to the quantization parameter and the zero point to obtain the quantized target weight, wherein the quantized target weight = (target weight - zero point) / quantization parameter. For example, the target weight range is [-10, 10], represented by floating-point type, and the quantization range is [-128, 127], represented by 8 as an integer. If a target weight value is 5.2, then according to the maximum and minimum values of the target weight and the quantization range, the quantization parameter is calculated to be 1.25, and the zero point is -128. The target weight is quantized according to the quantization parameter and the zero point, and the quantized target weight is 100.48.
[0042] If all network layers in the model to be quantized are quantized, then according to the network architecture of the model and the input-output relationship between each network layer, each network layer is traversed in turn, and the target weights in the network layer are quantized to obtain the quantized model; if only individual network layers in the model to be quantized are quantized, then the position of the network layer in the model can be determined according to the name of the network layer or the weight parameter information in the network layer, and the target weights in each determined network layer are quantized in turn to obtain the quantized model.
[0043] In this embodiment, after obtaining the quantized model, the process of using the quantized model for inference is as follows: first, multiply the input sample of the model by the quantized target weight to obtain the output sample corresponding to the target weight, and multiply the input sample by the weight parameter whose importance is greater than the preset threshold to obtain the output sample of the weight parameter whose importance is greater than the preset threshold. Then, the output sample corresponding to the target weight and the output sample of the weight parameter whose importance is greater than the preset threshold are added together to obtain the inference output model of the model. After the input sample is multiplied by the quantized target weight, the quantized target weight also needs to be dequantized. The dequantization process is the inverse process of the quantization process, so as to restore the output sample corresponding to the target weight to the original accuracy category.
[0044] In this embodiment, each weight parameter in any network layer to be quantized in the model to be quantized is obtained. Based on the importance of each weight parameter, weight parameters with an importance less than a preset threshold are screened as target weights, and the target weights are quantized to obtain a quantized model. Based on the importance of each weight parameter, only weight parameters with an importance less than the preset threshold are quantized, while weights with higher importance are not quantized. This reduces the accuracy loss caused by model quantization to a certain extent, allowing the model to meet the requirements of quantization while having a smaller impact on model accuracy.
[0045] 3 , which is a flow chart of a model quantization method according to a second embodiment of the present invention, the calculation of the importance of each weight parameter in step S202 in the first embodiment may include the following steps:
[0046] Step S301: obtaining input samples of the network layer to be quantized, and activating the input samples using a weight matrix formed by all weight parameters of the network layer to be quantized to obtain output samples.
[0047] In this embodiment, the input sample can be data input into the model to be quantized, and enters the network layer to be quantized after being processed by other network layers before the network layer to be quantized, wherein the data can be data in the form of text, image, audio, etc.
[0048] After passing through the input layer of the model to be quantized, the input sample can be processed into vector data, where the specific embodiment of the vector data can be in matrix form. For example, a two-dimensional image can be represented by a two-dimensional matrix, and text can be represented by a one-dimensional matrix.
[0049] Of course, for the network layer to be quantized, the weight parameters therein can be expressed in the form of a weight matrix, and the weight matrix can reflect the physical meaning and function of the network layer to be quantized when it is constructed.
[0050] For example, for an activation function layer, after the input sample enters the activation function layer, the input sample is nonlinearly mapped under the combined action of the weight matrix and the activation function to obtain the output sample, where the output sample = activation function (weight matrix * input sample), where the activation function can be a Sigmoid function and a ReLU function, etc.
[0051] Step S302: Calculate the contribution of each element in the weight matrix based on the input sample and the output sample.
[0052] Step S303: Determine the contribution of each element in the weight matrix as the importance of the weight parameter of the corresponding element.
[0053] In this embodiment, the output sample may be output data of the model to be quantized, and is input into other network layers following the network layer to be quantized, wherein the data may be data in the form of vectors or the like.
[0054] After passing through the output layer of the model to be quantized, the output samples can be processed into data in the form of text, images, and audio, such as the classification labels of text, the predicted probabilities of images, and the conversion results of text to speech.
[0055] Since the weight parameters are expressed in the form of a weight matrix, each element in the weight matrix corresponds to a weight parameter. It can be seen from the description of step S301 that each element will participate in the effect on the input sample, thereby obtaining the output sample.
[0056] For example, by multiplying the input sample by the weight matrix, it can be seen that all the weight parameters in the first column of the weight matrix act on all the elements in the first row of the matrix of the input sample, and the elements in the first column and the first row of the matrix of the output sample are obtained, and so on.
[0057] Among them, due to the effect of the weight parameters on the input samples, there are certain differences between the output samples and the input samples. Among them, under the same input sample conditions, the output samples obtained after the action of different weight parameters, the difference between each output sample and the input sample, the greater the difference, the greater the impact of the corresponding weight parameter change on the output of the network layer of the model, that is, the greater the contribution, conversely, the smaller the corresponding contribution.
[0058] In this embodiment, the input samples of the network layer to be quantized are activated using a weight matrix to obtain output samples. Based on the difference between the input and output samples, the contribution of each element in the weight matrix to the difference is calculated, and the contribution is determined as the importance of the weight parameter at the corresponding position in the weight matrix. Through the above steps, the importance of the elements at the corresponding position in the weight matrix is calculated based on the real-time input and output of the network layer to be quantized, making the calculation of the importance of each weight parameter more real-time and accurate.
[0059] 4 is a flow chart of a model quantization method provided in a third embodiment of the present invention. In step S302 of the second embodiment, the contribution of each element in the weight matrix is calculated based on the input sample and the output sample, which may include the following steps:
[0060] Step S401: Calculate the root mean square of each column element in the matrix expression corresponding to the input sample to obtain the root mean square value of the corresponding column, use the root mean square value of each column as the element of the corresponding column of a row matrix to obtain the first row matrix, and transpose the first row matrix to obtain the first column matrix.
[0061] Specifically, the calculation formula for the root mean square of each column element in the matrix expression corresponding to the input sample is:
[0062] Among them, bsz is the number of input samples, seq is the length of the input sample, X in_out is the root mean square value of each column element, xi is the element in each column of the matrix expression corresponding to the input sample, x i =[x1,x2,…x bsz*seq ].
[0063] For example, if the size of the matrix representation corresponding to the input matrix is m×n, the size of the calculated first row matrix is 1×n, and the size of the first column matrix after transposition is n×1.
[0064] Step S402: Calculate the mean square error of each column element in the matrix expression corresponding to the output sample to obtain the mean square error value of the corresponding column, and use the mean square error value of each column as the element of the corresponding column of a row matrix to obtain the second row matrix.
[0065] Specifically, the calculation formula for the mean square error of each column element in the matrix expression corresponding to the output sample is:
[0066] Among them, bsz is the number of input samples, seq is the length of the input sample, X ou_out is the mean square error of each column element, yi is the element in each column of the matrix expression corresponding to the output sample, y i =[y1,y2,…y bsz*seq ].
[0067] For example, if the size of the matrix expression corresponding to the output sample is m×n, the size of the calculated second row matrix is 1×n.
[0068] Step S403: multiply the elements of each row in the first column matrix by the absolute value of each element of the corresponding row in the weight matrix to obtain an intermediate matrix.
[0069] Step S404: Multiply each column element in the second row matrix with each element in the corresponding column of the middle matrix to obtain an influence matrix, where the element value of each element in the influence matrix represents the contribution degree of the element at the corresponding position in the weight matrix.
[0070] For example, if the first column of the matrix is The weight matrix is Then multiply the elements of each row in the first column matrix with the absolute value of each element in the corresponding row in the weight matrix, and the intermediate matrix is obtained as If the second row matrix is (y1 y2 y3), then multiply each column element in the second row matrix with each element in the corresponding column of the middle matrix to obtain the influence matrix: The element value of each element in the influence degree matrix represents the contribution degree of the element at the corresponding position in the weight matrix.
[0071] In this embodiment, the first column matrix is obtained by calculating the root mean square of each column element in the matrix expression corresponding to the input sample, the mean square error of each column element in the matrix expression corresponding to the output sample is calculated to obtain the second row matrix, the elements of each row in the first column matrix are multiplied by the absolute value of each element in the corresponding row in the weight matrix to obtain an intermediate matrix, and the elements of each column in the second row matrix are multiplied by each element in the corresponding column in the intermediate matrix to obtain an influence matrix. By analyzing the size of the input sample, output sample, and weight value, a corresponding weight importance that can be expressed in data volume is quickly and accurately obtained, so that the target weight can be subsequently screened based on the importance threshold.
[0072] 5 is a flow chart of a model quantization method provided in a fourth embodiment of the present invention. In step S203 of the first embodiment, the target weights in the network layer to be quantized are quantized to obtain the quantized network layer, which may include the following steps:
[0073] Step S501: scaling the target weights in the network layer to be quantized to obtain scaled target weights.
[0074] In this embodiment, the scaling process may refer to scaling the target weights in the network layer to be quantized to a smaller range according to a certain ratio. The target weight quantization may refer to converting the target weights from a floating point type to a fixed point type. For example, if the scaled target weights are of 32-bit floating point type, then after quantization, the scaled quantized target weights may be of 4-bit integer type. The inverse scaling process may refer to restoring the scaled quantized target weights to the original unscaled range according to a certain ratio, which is the inverse process of the scaling process.
[0075] Step S502: quantize the scaled target weight to obtain a scaled and quantized target weight, and perform inverse scaling on the scaled and quantized target weight to obtain a quantized target weight.
[0076] In this embodiment, the forward calculation process before the target weight is scaled and quantized is: multiply the input sample and the target weight to obtain the output sample. The forward calculation process after the target weight is scaled and quantized is: perform inverse scaling processing on the scaled and quantized target weight to obtain the quantized target weight, multiply the quantized target weight by the input sample to obtain the output sample, wherein the inverse scaling processing process is: divide the input sample of the network layer where the scaled and quantized target weight is located by the scaling factor of the network layer, and at the same time, since the weight parameters with an importance greater than the preset threshold in the network layer are not scaled, the weight parameters with an importance greater than the preset threshold are multiplied by the scaling factor of the network layer to maintain the original high-precision level, wherein the scaling factor in the inverse scaling processing process is placed between the network layer where the scaled and quantized target weight is located and the upper network layer.
[0077] For example, the network layer where the target weight after scaling and quantization is located is L1, and the scaling factor of the target weight in the network layer L1 is 0.5. The input sample corresponding to the network layer L1 is output by the network layer L2. When the target weight in the network layer L1 is inversely scaled, the weight in the network layer L2 is divided by the scaling factor 0.5 in the network layer L1, that is, the input sample corresponding to the network layer L1 is divided by the scaling factor 0.5, and at the same time, the weight parameter in the network layer L1 whose importance is greater than the preset threshold is multiplied by the scaling factor 0.5, wherein the scaling factor 0.5 is placed between the network layer L2 and the network layer L1.
[0078] The forward calculation process of the above-mentioned scaled and quantized target weight corresponds to the following formula: ou_act = in_act / scale * Q (w * scale)
[0079] Among them, ou_act is the output sample, in_act is the input sample, scale is the scaling factor, Q() is the quantization method, w is the target weight, w*scale means scaling the target weight to obtain the scaled target weight, Q(w*scale) means quantizing the scaled target weight to obtain the scaled and quantized target weight.
[0080] Step S503: replacing the corresponding target weights in the network to be quantized with the quantized target weights to obtain a quantized network layer.
[0081] In this embodiment, after obtaining the quantized target weights, the quantized target weights are used to replace the corresponding target weights in the network to be quantized. After completing the replacement of all target weights in the network layer to be quantized, the quantized network layer can be obtained.
[0082] In this embodiment, after obtaining the quantized network layer, all the network layers to be quantized are traversed to obtain the quantized model. After obtaining the quantized model, the inference process using the quantized model is as follows: first, the input sample is multiplied by the target weight parameter after scaling and quantization to obtain the output sample corresponding to the target weight, and the input sample is multiplied by the weight parameter whose importance is greater than the preset threshold to obtain the output sample of the weight parameter whose importance is greater than the preset threshold. Then, the output sample corresponding to the target weight and the output sample corresponding to the weight parameter whose importance is greater than the preset threshold are summed to obtain the output sample of the model inference. Among them, the input sample is the input sample that has been divided by the scaling factor during the inverse scaling process, and the weight parameter whose importance is greater than the preset threshold is the weight parameter that has been multiplied by the scaling factor during the inverse scaling process. After the input sample is multiplied by the scaled and quantized target weight, the quantized target weight needs to be inversely quantized to restore the output sample corresponding to the target weight to the accuracy level corresponding to the target weight before quantization.
[0083] The corresponding reasoning calculation formula for the above reasoning process is: ou_act = in_act*w_important+in_act*w_low*revert_scale+bais
[0084] Among them, ou_act is the output sample, in_act is the input sample that has been divided by the scaling factor, w_important is the weight parameter whose importance is greater than the preset threshold, where the weight parameter has been multiplied by the scaling factor, w_low is the target weight after scaling quantization, that is, w_low = Q (w * scale), revert_scale and bais are inverse quantization coefficients.
[0085] In this embodiment, the target weights are scaled to obtain scaled target weights, the scaled target weights are quantized to obtain scaled and quantized target weights, the scaled and quantized target weights are inversely scaled to obtain quantized target weights, and the quantized target weights are replaced with the corresponding target weights in the network layer to be quantized to obtain the quantized network layer. By only quantizing target weights whose importance is less than a preset threshold and not quantizing weight parameters whose importance is greater than the preset threshold, the accuracy of weight parameters whose importance is greater than the preset threshold is preserved, reducing the accuracy loss of the model caused by quantization. At the same time, by quantizing the target weights, the resources occupied by the model are also reduced.
[0086] 6 is a flow chart of a model quantization method provided in a fifth embodiment of the present invention. In step S501 of the fourth embodiment, the target weights in the network layer to be quantized are scaled to obtain the scaled target weights. The process may include the following steps:
[0087] Step S601: Obtain an initial scaling factor.
[0088] Step S602: scaling all original weights in the network layer to be quantized based on the initial scaling factor to obtain scaled original weights, and quantizing the scaled original weights to obtain quantized original weights.
[0089] In this embodiment, the scaling factor may refer to a proportional factor that scales the target weight to a smaller range, wherein the scaling factor may be in the form of a scalar value or a matrix. The original weight may refer to a weight that has not been scaled or quantized.
[0090] For example, if the initial scaling factor is a scalar value, all the original weights in the weight matrix are multiplied by the scalar value to obtain the scaled original weights. If the initial scaling factor is a matrix, the elements at each position in the initial scaling factor matrix are multiplied by the original weights at the corresponding positions in the weight matrix to obtain the scaled original weights. The scaled original weights are quantized to obtain the quantized original weights.
[0091] Step S603: Calculate the quantization loss before and after quantization based on all original weights and quantized original weights, and adjust the initial scaling factor based on the minimum quantization loss to obtain the adjusted scaling factor.
[0092] Step S604: Use the adjusted scaling factor as the initial scaling factor, return to execute the step of scaling all the original weights in the network layer to be quantized based on the initial scaling factor until the quantization loss is minimized, and obtain the initial scaling factor when the quantization loss is minimized as the target scaling factor.
[0093] Step S605: Using the target scaling factor, the target weight in the network layer to be quantized is scaled to obtain the scaled target weight.
[0094] In this embodiment, the quantization loss may refer to the precision loss caused by converting the original weight into the quantized original weight through the quantization operation.
[0095] Specifically, the quantization loss before and after quantization is calculated as follows: quant_loss = || inact / scale*Q(w*scale)-inact*w||
[0096] Among them, quant_loss is the quantization loss value, inact is the input sample, scale is the initial scaling factor, Q() is the quantization method, w is the original weight, alpha is the initial weighting coefficient, and argmin represents the value of the independent variable alpha required to minimize the formula quant_loss(scale(alpha)).
[0097] For example, in the process of obtaining the target scaling factor corresponding to the minimum quant_loss, the initial alpha value is between 0 and 1, and 100 values can be evenly taken from 0 to 1, that is, Calculate the quant_loss values corresponding to these 100 values respectively, compare the calculated quant_loss values based on the minimum quant_loss, screen out the minimum quant_loss value, and adjust the initial scaling factor according to the alpha value corresponding to the minimum quant_loss value to obtain the target scaling factor. Scale the target weight in the network layer to be quantized according to the obtained target scaling factor to obtain the scaled target weight.
[0098] In this embodiment, the original weight is scaled according to the initial scaling factor to obtain the scaled original weight, the scaled original weight is quantized to obtain the quantized original weight, and the quantization loss before and after quantization is calculated based on the quantized original weight and the original weight. Based on the minimum quantization loss, the initial scaling factor is adjusted to obtain the target scaling factor corresponding to the minimum quantization loss. Through the above steps, the target weight is scaled according to the target scaling factor to obtain the scaled target weight, and the target weight is scaled to a smaller range for subsequent quantization operations. At the same time, the accuracy of the model is maintained to a certain extent, which helps retain the accuracy of the model and reduces the error caused by quantization.
[0099] 7 , which is a flow chart of a model quantization method according to a sixth embodiment of the present invention, obtaining the initial scaling factor in step S601 in the fifth embodiment may include the following steps:
[0100] Step S701: Obtain an input scaling factor based on an input sample, and obtain a weight scaling factor based on a weight matrix.
[0101] Step S702: using the initial weighting coefficient, perform weighting processing on the input scaling coefficient and the weight scaling coefficient, and obtain the weighted result as the initial scaling factor.
[0102] The input scaling factor and the weight scaling factor are weighted, and the weighting process can be performed by weighted summation, weighted averaging, weighted harmonic mean, weighted geometric mean, etc. For example, if the input scaling factor and the weight scaling factor are weighted by weighted averaging, the input scaling factor and the weight scaling factor are multiplied by the initial weighting factor respectively, and the two obtained values are added and divided by the sum of the initial weighting coefficients to obtain a weighted result. For example, if the input scaling factor is 0.5, the weight scaling factor is 0.3, and the initial weighting coefficient is 0.2, then the weighted result obtained by weighted averaging the input scaling factor and the weight scaling factor is 0.4.
[0103] Specifically, according to the input scaling factor, weight scaling factor and weighting factor, the calculation formula for the initial scaling factor is:
[0104] Among them, scale_mid is the middle value of the scaling factor, in_act_scale is the input scaling factor, w_scale is the weight scaling factor, alpha is the initial weighting factor, scale_mid max is the maximum value of the middle value of the scaling factor, scale_mid min is the minimum of the intermediate values of the scaling factors, and scale is the initial scaling factor.
[0105] In this embodiment, when the initial weighting coefficient is 1, the corresponding intermediate value of the scaling factor is the same as the input scaling coefficient, and the input scaling coefficient can be calculated to obtain the initial scaling factor; when the initial weighting coefficient is 0, the corresponding intermediate value of the scaling factor is the same as the weight scaling coefficient, and the weight scaling coefficient can be calculated to obtain the initial scaling factor; when the initial weighting coefficient is any value other than 0 and 1, the input scaling coefficient and the weight scaling coefficient are weighted according to the weighting coefficient to obtain the initial scaling factor.
[0106] In this embodiment, the initial scaling factor is calculated based on the input scaling coefficient, the weight scaling coefficient and the initial weighting coefficient. Through the above steps, an accurate scaling factor is calculated. The scaling factor takes into account the influence of the input sample and the weight matrix on the scaling, thereby reducing the error loss caused by the scaling process. The scaling factor can be dynamically adjusted by the weighting coefficient according to different situations, thereby improving the adaptability of the scaling.
[0107] In one embodiment, adjusting the initial scaling factor in step S701 in the sixth embodiment to obtain the adjusted scaling factor may include:
[0108] The initial weighting coefficient is adjusted to obtain an adjusted weighting coefficient, and the input scaling coefficient and the weight scaling coefficient are weighted using the adjusted weighting coefficient to obtain a weighted result as the adjusted scaling factor.
[0109] Specifically, the quantization loss before and after quantization is continuously iteratively calculated based on the initial weighting coefficient, and the initial weighting coefficient is adjusted according to the calculated quantization loss. Based on the minimum quantization loss, the weighting coefficient corresponding to the minimum quantization loss is selected, and the input scaling coefficient and the weight scaling coefficient are weighted according to the weighting coefficient to obtain the adjusted scaling factor, that is, the target scaling factor.
[0110] In this embodiment, the initial weighting coefficient is adjusted, and the input scaling coefficient and the weight scaling coefficient are weighted according to the adjusted weighting coefficient to obtain the adjusted scaling factor. By adjusting the scaling factor only by adjusting the weighting coefficient, fewer parameters need to be adjusted, and the adjustment result corresponding to the minimum quantization loss can be quickly iterated.
[0111] 8 is a flow chart of a model quantization method according to an eighth embodiment of the present invention. In step S701 of the sixth embodiment, the input scaling factor is obtained according to the input sample, which may include the following steps:
[0112] Step S801: Calculate the average value of the norm of each column element in the matrix expression corresponding to the input sample to obtain the first average value of the corresponding column.
[0113] Step S802: taking the first average value of each column as an element of a corresponding column of a row of a matrix, and obtaining the third row of the matrix as the input scaling factor.
[0114] Specifically, the calculation formula for the first average value is:
[0115] Among them, in_act_scale is the first average, bsz is the number of input samples, seq is the length of the input sample, x i For each column of the matrix expression corresponding to the input sample, x i =[x1,x2,…,x bsz*seq ].
[0116] For example, if the size of the matrix expression corresponding to the input sample is m×n, after calculating the first average value of all columns in the input sample, the size of the third row matrix obtained is 1×n, where each element in the third row matrix is the input scaling factor.
[0117] In this embodiment, the first average value of the corresponding column is obtained by calculating the average value of the norm of each column element in the matrix expression corresponding to the input sample. After calculating the first average values of all columns in the input sample, the third row matrix representing the input scaling factor is obtained. By calculating the average value of the norm of the input sample, the input sample is normalized to the same scale, thereby avoiding the situation where certain features in the input sample have too large or too small an impact. The scaling process is performed using the average value of the norm as the input scaling factor, taking into account the impact of the input sample on the scaling, reducing the error loss caused by the scaling process to a certain extent, and avoiding the overall feature distortion or instability after scaling.
[0118] 9 is a flow chart of a model quantization method according to a ninth embodiment of the present invention. In step S701 of the sixth embodiment, obtaining a weight scaling coefficient according to a weight matrix may include the following steps:
[0119] Step S901: Divide the weight matrix equally to obtain N sub-matrices, and compare the absolute value of each element in each sub-matrix with the maximum absolute value of all elements in the corresponding sub-matrix to obtain a first ratio of the corresponding element.
[0120] Step S902: Calculate the average of the first ratios of the elements in each column to obtain the second average of the corresponding column, use the second average of each column as the element of the corresponding column of a row of the matrix, and obtain the fourth row of the matrix as the weight scaling coefficient.
[0121] For example, if the weight matrix is The weight matrix is evenly divided to obtain 4 sub-matrices, which are If the maximum absolute value of all elements in submatrix w1 is |w 11 |, then the first ratio of the corresponding element in submatrix w1 is If the maximum absolute value of all elements in submatrix w2 is |w 14 |, then the first ratio of the corresponding element in submatrix w2 is If the maximum absolute value of all elements in submatrix w3 is |w 41 |, then the first ratio of the corresponding element in submatrix w3 is , if the maximum absolute value of all elements in submatrix w4 is |w 44 |, then the first ratio of the corresponding element in submatrix w4 is Then the first ratio of the corresponding element in the weight matrix is Calculate the average of the first ratios of the elements in each column to obtain the second average, and use the second average of each column as the element of the corresponding column of a row of the matrix. The size of the obtained fourth row matrix is 1×4, where each element in the fourth row matrix is a weight scaling coefficient.
[0122] In this example, the weight matrix is divided equally to obtain N sub-matrices, and the absolute value of each element in each sub-matrix is compared with the maximum absolute value of all elements in the corresponding sub-matrix to obtain the first ratio of the corresponding element, and the average value of the first ratio of the elements in each column is calculated to obtain the second average value of the corresponding column. The second average value of each column is used as the element of the corresponding column of a row matrix to obtain the fourth row matrix as the weight scaling coefficient. The weight matrix is divided into multiple sub-matrices, and the first ratio of the elements in each sub-matrix is calculated. The calculated first ratio can characterize the importance of each element in the corresponding sub-matrix. The first ratio is averaged, and the obtained second average value can better reflect the overall importance distribution of each element in the weight matrix. The second average value is used as the weight scaling coefficient for scaling. The influence of the weight matrix on scaling is taken into account, and the error loss caused by the scaling process is reduced to a certain extent, and the overall feature distortion or instability after scaling is avoided.
[0123] 10 is a flow chart of a model quantization method according to a tenth embodiment of the present invention. In step S502 of the fourth embodiment, the scaled target weight is quantized to obtain the scaled quantized target weight, which may include the following steps:
[0124] Step S1001: Obtain the maximum value and minimum value of the scaled target weight, as well as the target number of bits.
[0125] In this embodiment, the target number of bits may refer to the number of bits occupied by the quantized target weight obtained after quantization.
[0126] Step S1002: For any weight in the scaled target weights, compare the difference between the weight and the minimum value with the difference between the maximum value and the minimum value to obtain a second ratio.
[0127] Step S1003: Convert the second ratio to the magnitude corresponding to the target number of bits to obtain an intermediate value, round the intermediate value to obtain the rounded value as the scaled quantization result of the weight, traverse all weights in the scaled target weight to obtain the scaled quantized target weight.
[0128] Specifically, the weight scaling quantization calculation formula is:
[0129] Among them, w_low is the target weight after scaling and quantization, Q() is the quantization method, round() is the rounding function, w_mid is the target weight after scaling, w_mid min is the minimum value of the scaled target weight, w_mid max is the maximum value of the scaled target weight, and bit_width is the target number of bits.
[0130] in, is the second ratio, The round() function converts the calculated high-bit value into the target bit value according to the principle of rounding.
[0131] In this embodiment, the scaled quantized target weight is calculated based on the scaled target weight, the maximum and minimum values of the scaled target weight, and the target number of bits. The calculation is simple and the accurate quantization of each target weight is achieved quickly. At the same time, by adjusting the target number of bits, the quantization level of the weight can be automatically adjusted according to different task conditions, thereby improving the adaptability of the model quantization.
[0132] The models quantized using the above-mentioned model quantization method may include neural network models such as image classification models, natural language processing models, and speech recognition models. The corresponding image classification model can be applied to the scene of sorting different objects in the image in object detection, the corresponding natural language processing model can be applied to the scene of classifying text in text processing, and the corresponding speech recognition model can be applied to the scene of identifying different information in the speech in speech recognition.
[0133] For example, in a scenario where the above-mentioned image classification model is used to sort different items in an image during object detection, the specific image classification model may be an AlexNet model, a VGGNet model, a ResNet model, etc. The image classification model may be composed of network layers such as an input layer, a convolutional layer, a pooling layer, and a fully connected layer. The above-mentioned model quantization method may quantize all network layers in the image classification model, or may quantize any network layer in the image classification model. When the above-mentioned image classification model is used in a scenario where different items in an image are sorted, the image classification model may be a model that has been quantized using the above-mentioned model quantization method, or the image classification model may be quantized using the above-mentioned model quantization method during the image classification process.
[0134] For another example, during the image classification process, the convolutional layer in the image classification model is quantized. The operational relationship between the input data, weight parameter data, and output data during the model quantization process can be referred to the following example.
[0135] The input of the model is an object image, which is encoded through the input layer to obtain the matrix expression of N channels corresponding to the object image. One of the matrix expressions X is input as the input data of the convolution layer. The matrix expression corresponding to the convolution kernel parameters of this convolution layer is After convolving the input data X input according to the convolution kernel parameter W1, the output data
[0136] The root mean square of each column of the X input is calculated to obtain the first row matrix, and then the first column matrix is obtained after transposition. The mean square error of each column of the Y output is calculated to obtain the second row matrix. Based on the first column matrix, the second row matrix and W1, the influence matrix representing the importance of the corresponding position elements in W1 is calculated.
[0137] W1 is screened based on the importance of each weight parameter in W1 represented by the influence matrix and a preset threshold, and a target weight W2 having an importance less than the preset threshold is obtained. That is, if x1, x3, x7, and x9 in the X input represent image features of the portion of the object image containing the object, and x2, x4, x5, x6, and x8 represent image features of the portion of the object image without the object, then based on the relationship between w1, the X input, and the Y output, a target weight W2 having a smaller influence on the consistency between the X input and the Y output is screened.
[0138] In the process of inference of the image classification model, the above-mentioned quantization method is used for quantization, and some features of the input data can be used to screen the weights. The important weights are obtained by screening the important features of the part of the input data with objects, and the general weights are obtained by screening the general features of the part of the input data without objects. The general weights are quantized as target weights, so that the model has better targeting of the data and the quantized model is more consistent with the input data.
[0139] FIG11 shows a model quantization device according to an eleventh embodiment of the present invention. The model quantization device corresponds to the model quantization method according to the above embodiment. The model quantization device includes a first acquisition module 111, a screening module 112, and a first quantization module 113. The functional modules are described in detail as follows:
[0140] The first acquisition module 111 is used to obtain each weight parameter in any network layer to be quantized in the model to be quantized;
[0141] A screening module 112 is configured to calculate the importance of each weight parameter and select a weight parameter whose importance is less than a preset threshold from all weight parameters as a target weight;
[0142] The first quantization module 113 is configured to quantize the target weight in the network layer to be quantized to obtain a quantized network layer, and traverse all the network layers to be quantized to obtain a quantized model.
[0143] Optionally, the screening module 112 includes:
[0144] An activation submodule, configured to obtain an input sample of the network layer to be quantized, activate the input sample using a weight matrix formed by all weight parameters of the network layer to be quantized, and obtain an output sample;
[0145] A first calculation submodule is configured to calculate a contribution degree of each element in the weight matrix based on the input sample and the output sample;
[0146] The determination submodule is used to determine the contribution degree of each element in the weight matrix as the importance degree of the weight parameter of the corresponding element.
[0147] Optionally, the first calculation submodule includes:
[0148] a second calculation unit, configured to calculate the root mean square of each column element in the matrix expression corresponding to the input sample to obtain a root mean square value of the corresponding column, use the root mean square value of each column as an element of a corresponding column of a row matrix to obtain a first row matrix, and transpose the first row matrix to obtain a first column matrix;
[0149] A third calculation unit is used to calculate the mean square error of each column element in the matrix expression corresponding to the output sample to obtain the mean square error value of the corresponding column, and use the mean square error value of each column as the element of the corresponding column of a row matrix to obtain a second row matrix;
[0150] a fourth calculation unit, configured to perform a multiplication operation on the elements of each row in the first column matrix by the absolute value of each element in the corresponding row in the weight matrix to obtain an intermediate matrix;
[0151] The fifth calculation unit is used to multiply each column element in the second row matrix with each element in the corresponding column of the intermediate matrix to obtain an influence degree matrix, wherein the element value of each element in the influence degree matrix represents the degree of contribution of the element at the corresponding position in the weight matrix.
[0152] Optionally, the first quantization module 113 includes:
[0153] A first scaling submodule is configured to scale the target weight in the network layer to be quantized to obtain a scaled target weight;
[0154] A second quantization submodule is configured to quantize the scaled target weight to obtain a scaled and quantized target weight, and perform an inverse scaling process on the scaled and quantized target weight to obtain a quantized target weight;
[0155] The replacement submodule is used to replace the corresponding target weight in the network to be quantized with the quantized target weight to obtain a quantized network layer.
[0156] Optionally, the first scaling submodule includes:
[0157] A second obtaining unit, configured to obtain an initial scaling factor;
[0158] A third quantization unit is configured to scale all original weights in the network layer to be quantized based on the initial scaling factor to obtain scaled original weights, and quantize the scaled original weights to obtain quantized original weights;
[0159] a sixth calculating unit, configured to calculate quantization losses before and after quantization based on all original weights and the quantized original weights, and adjust the initial scaling factor based on minimizing the quantization loss to obtain an adjusted scaling factor;
[0160] a seventh calculation unit, configured to use the adjusted scaling factor as the initial scaling factor, return to the step of scaling all original weights in the network layer to be quantized based on the initial scaling factor, until the quantization loss is minimized, and obtain the initial scaling factor when the quantization loss is minimized as the target scaling factor;
[0161] The second scaling unit is configured to scale the target weight in the network layer to be quantized using the target scaling factor to obtain the scaled target weight.
[0162] Optionally, the second obtaining unit includes:
[0163] a third acquisition subunit, configured to obtain an input scaling coefficient according to the input sample, and obtain a weight scaling coefficient according to the weight matrix;
[0164] The weighting subunit is used to use the initial weighting coefficient to perform weighted processing on the input scaling coefficient and the weight scaling coefficient, and obtain a weighted result as the initial scaling factor.
[0165] Optionally, the sixth calculation unit includes:
[0166] The adjustment subunit is used to adjust the initial weighting coefficient to obtain an adjusted weighting coefficient, and use the adjusted weighting coefficient to weight the input scaling coefficient and the weight scaling coefficient to obtain a weighted result as the adjusted scaling factor.
[0167] Optionally, the third obtaining subunit includes:
[0168] Calculating the average of the norms of the elements in each column of the matrix expression corresponding to the input sample to obtain a first average value of the corresponding column;
[0169] The first average value of each column is used as the element of the corresponding column of a row of the matrix, and the third row of the matrix is the input scaling factor.
[0170] Optionally, the third obtaining subunit includes:
[0171] The weight matrix is equally divided to obtain N sub-matrices, and the absolute value of each element in each sub-matrix is compared with the maximum absolute value of all elements in the corresponding sub-matrix to obtain a first ratio of the corresponding element;
[0172] Calculate the average of the first ratios of the elements in each column to obtain the second average of the corresponding column. Use the second average of each column as the element of the corresponding column of a row of the matrix to obtain the fourth row of the matrix as the weight scaling coefficient.
[0173] Optionally, the second quantization submodule includes:
[0174] a fourth acquiring unit, configured to acquire a maximum value and a minimum value of the scaled target weights, and a target number of bits;
[0175] an eighth calculating unit, configured to compare, for any weight among the scaled target weights, a difference between the weight and the minimum value with a difference between the maximum value and the minimum value to obtain a second ratio;
[0176] The ninth calculation unit is used to convert the second ratio to the magnitude corresponding to the target number of bits, obtain an intermediate value, round the intermediate value, and obtain the rounded value as the scaled quantization result of the weight, traverse all weights in the scaled target weight, and obtain the scaled quantized target weight.
[0177] For the specific definition of the model quantization device, please refer to the definition of the model quantization method above and will not be repeated here. Each module in the above-mentioned model quantization device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0178] FIG12 is a schematic diagram of the structure of a terminal device provided in accordance with Embodiment 12 of the present invention. As shown in FIG12 , the terminal device of this embodiment includes: at least one processor (only one is shown in FIG12 ), a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps of any of the aforementioned model quantization method embodiments are implemented.
[0179] The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that FIG12 is merely an example of a terminal device and does not limit the terminal device. The terminal device may include more or fewer components than shown in the figure, or may combine certain components or different components, for example, it may also include a network interface, a display screen, and an input device.
Claims
1. A model quantization method, characterized in that: The method comprises the following steps: Obtain each weight parameter in any network layer to be quantized in the model to be quantized; Calculate the importance of each weight parameter, and select the weight parameter whose importance is less than a preset threshold from all weight parameters as the target weight; The target weight in the network layer to be quantized is quantized to obtain a quantized network layer, and all network layers to be quantized are traversed to obtain a quantized model.
2. The model quantization method according to claim 1, characterized in that: The calculating the importance of each weight parameter comprises: Obtaining input samples of the network layer to be quantized, activating the input samples using a weight matrix formed by all weight parameters of the network layer to be quantized, and obtaining output samples; Calculating the contribution of each element in the weight matrix according to the input sample and the output sample; The contribution degree of each element in the weight matrix is determined as the importance degree of the weight parameter of the corresponding element.
3. The model quantization method according to claim 2, characterized in that: The step of calculating, based on the input sample and the output sample, a contribution degree of each element in the weight matrix, comprises: Calculate the root mean square of each column element in the matrix expression corresponding to the input sample to obtain the root mean square value of the corresponding column, use the root mean square value of each column as the element of the corresponding column of a row of the matrix to obtain the first row matrix, and transpose the first row matrix to obtain the first column matrix; Calculate the mean square error of each column element in the matrix expression corresponding to the output sample to obtain the mean square error value of the corresponding column, and use the mean square error value of each column as the element of the corresponding column of a row of the matrix to obtain a second row of the matrix; Multiplying the elements of each row in the first column matrix by the absolute value of each element of the corresponding row in the weight matrix to obtain an intermediate matrix; Multiply each column element in the second row matrix with each element in the corresponding column of the intermediate matrix to obtain an influence degree matrix, wherein the element value of each element in the influence degree matrix represents the contribution degree of the element at the corresponding position in the weight matrix.
4. The model quantization method according to claim 1, characterized in that: The step of quantizing the target weight in the network layer to be quantized to obtain a quantized network layer includes: Scaling the target weight in the network layer to be quantized to obtain the scaled target weight; quantizing the scaled target weight to obtain a scaled and quantized target weight, and performing a reverse scaling process on the scaled and quantized target weight to obtain a quantized target weight; The quantized target weights are used to replace the corresponding target weights in the network to be quantized to obtain a quantized network layer.
5. The model quantization method according to claim 4, characterized in that: The scaling process of the target weight in the network layer to be quantized to obtain the scaled target weight includes: Get the initial scaling factor; Scaling all original weights in the network layer to be quantized based on the initial scaling factor to obtain scaled original weights, and quantizing the scaled original weights to obtain quantized original weights; Calculating the quantization loss before and after quantization according to all the original weights and the quantized original weights, and adjusting the initial scaling factor based on the minimum quantization loss to obtain an adjusted scaling factor; The adjusted scaling factor is used as the initial scaling factor, and the step of scaling all the original weights in the network layer to be quantized based on the initial scaling factor is returned to be executed until the quantization loss is minimized, and the initial scaling factor when the quantization loss is minimized is obtained as the target scaling factor; The target weight in the network layer to be quantized is scaled using the target scaling factor to obtain the scaled target weight.
6. The model quantization method according to claim 5, characterized in that: The obtaining of the initial scaling factor comprises: According to the input sample, an input scaling factor is obtained, and according to the weight matrix, a weight scaling factor is obtained; The input scaling factor and the weight scaling factor are weighted using the initial weighting coefficient, and the weighted result is obtained as the initial scaling factor.
7. The model quantization method according to claim 6, characterized in that: The adjusting the initial scaling factor to obtain an adjusted scaling factor includes: The initial weighting coefficient is adjusted to obtain an adjusted weighting coefficient, and the input scaling coefficient and the weight scaling coefficient are weighted using the adjusted weighting coefficient to obtain a weighted result as the adjusted scaling factor.
8. The model quantization method according to claim 6, characterized in that: The step of obtaining an input scaling factor according to the input sample comprises: Calculate the average value of the norm of each column element in the matrix expression corresponding to the input sample to obtain a first average value of the corresponding column; The first average value of each column is used as the element of the corresponding column of a row of the matrix, and the third row of the matrix is obtained as the input scaling factor.
9. [Corrected 12.12.2024 in accordance with Rule 91] The model quantization method according to claim 6, characterized in that: The step of obtaining a weight scaling coefficient according to the weight matrix includes: The weight matrix is equally divided to obtain N sub-matrices, and the absolute value of each element in each sub-matrix is compared with the maximum absolute value of all elements in the corresponding sub-matrix to obtain a first ratio of the corresponding element; Calculate the average of the first ratios of the elements in each column to obtain the second average of the corresponding column, use the second average of each column as the element of the corresponding column of a row of the matrix, and obtain the fourth row of the matrix as the weight scaling coefficient.
10. [Corrected 12.12.2024 in accordance with Rule 91] The model quantization method according to claim 4, characterized in that: The step of quantizing the scaled target weight to obtain the scaled quantized target weight includes: Obtaining the maximum value and the minimum value of the scaled target weights, and the target number of bits; For any weight in the scaled target weights, compare the difference between the weight and the minimum value with the difference between the maximum value and the minimum value to obtain a second ratio; Convert the second ratio to the magnitude corresponding to the target number of bits to obtain an intermediate value, round the intermediate value to obtain a rounded value as the scaled quantization result of the weight, traverse all weights in the scaled target weight to obtain the scaled quantized target weight.
11. A model quantization device, characterized in that: The device comprises: The first acquisition module is used to obtain each weight parameter in any network layer to be quantized in the model to be quantized; A screening module is used to calculate the importance of each weight parameter, and screen out the weight parameters whose importance is less than a preset threshold from all weight parameters as target weights; The first quantization module is used to quantize the target weight in the network layer to be quantized to obtain a quantized network layer, and traverse all the network layers to be quantized to obtain a quantized model.
12. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the model quantization method according to any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the model quantization method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Neural network model quantification method, system and device and computer medium
CN114970822A
Quantitative model generation method and device, electronic equipment and storage medium
CN115759238A
Model quantification method, model quantification device, electronic equipment and medium
CN116822594A
Positive electrode active material
KR1020250052187A
Method and apparatus with neural network parameter quantization
US20200380360A1