A neural network model quantization method, system, device and computer medium

The Transformer-based neural network quantization method efficiently compresses models by generating layer-specific bit widths, addressing inefficiencies in existing methods and preserving accuracy.

CN114970822BActive Publication Date: 2025-07-15LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210609520.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-15
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The existing neural network model quantization methods have problems such as high limitations, long time periods, requiring multiple GPU resources to support, and not being able to reflect cross-layer impact.

Method used

The Transformer model is used to linearly embed the weight values, hyperparameters and position numbers of each network layer in the neural network model to generate a target embedding matrix, and the pre-trained Transformer model is processed to determine the quantization bit number, so as to achieve mixed precision quantization.

Benefits of technology

It reduces the model size and memory footprint, while retaining the original network's accuracy loss, reducing the amount of computing and reducing limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114970822B_ABST
    Figure CN114970822B_ABST
Patent Text Reader

Abstract

The present application discloses a neural network model quantization method, system, device and computer medium for quantizing a neural network model, including obtaining the weight values, hyperparameters and position serial numbers of each network layer in the target neural network model to be quantized; performing linear embedding on the weight values, hyperparameters and position serial numbers to generate a target embedding matrix; processing the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model; and quantizing the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model. In the present application, by processing the target embedding matrix with a Transformer model to obtain the quantization bit numbers of each layer in the target neural network model, the model size and memory occupancy can be reduced, while the precision loss of the original network is small. In addition, the amount of computation can be greatly reduced and the limitation is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of neural network models, and more specifically, to a method, system, device and computer medium for quantifying neural network models. Background Art

[0002] With the development of artificial intelligence technology, while the accuracy of deep neural network models is continuously improving, the number of model parameters has increased sharply. This makes the model have problems such as a large number of model parameters and high computational complexity during actual deployment, especially for some edge devices, these problems are particularly obvious. To solve these problems, the neural network model can be quantified. Model quantization is a mature and commonly used means in the field of model compression, which can convert a floating-point model into an integer model with a smaller number of bits (binary digits, bits). 8-bit quantization is already relatively mature in the industrial field, but the compression effect on the model is limited. Extremely low 1 or 2-bit quantization has also been studied a lot in the academic field, but it often comes with a decrease in model accuracy. Since the importance of each layer of the model is different, if different quantization bit numbers are used for each layer, that is, mixed-precision quantization, the model accuracy can be guaranteed while reducing the number of parameters.

[0003] Existing mixed-precision quantization methods include using reinforcement learning methods to obtain the quantization bit positions of each layer under the constraints of hardware latency and accuracy loss requirements, or using neural network architecture search methods. By setting multiple edges representing different bit positions on the nodes of each layer, and then finding the optimal bit distribution through random search. However, the disadvantages of these two methods are long time periods and the need for multiple GPU resources to provide computing power support. Another method is to calculate the relative sensitivity of each layer of the network using the hessian matrix to provide a reference for the bit distribution of each layer. The disadvantage is that it can only calculate the importance of the weights of each layer in the loss unidirectionally from back to front and cannot reflect the cross-layer influence. This makes the limitations of existing methods for quantifying neural network models high.

[0004] In summary, how to reduce the limitations of neural network model quantization methods is an urgent problem for those skilled in the art currently. Summary of the Invention

[0005] The purpose of this application is to provide a method for quantifying neural network models, which can, to a certain extent, solve the technical problem of how to reduce the limitations of neural network model quantization methods. This application also provides a neural network model quantization system, an electronic device and a computer-readable storage medium.

[0006] To achieve the above purpose, this application provides the following technical solutions:

[0007] A neural network model quantization method, comprising:

[0008] Obtain the weight values, hyperparameters and position numbers of each network layer in the target neural network model to be quantized;

[0009] Perform linear embedding on the weight values, the hyperparameters and the position numbers to generate a target embedding matrix;

[0010] Process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model;

[0011] Quantize the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model.

[0012] Preferably, the performing linear embedding on the weight values, the hyperparameters and the position numbers to generate a target embedding matrix includes:

[0013] Flatten the weight values of the network layer into corresponding first vectors;

[0014] Serialize the hyperparameters of the network layer into corresponding second vectors;

[0015] Perform position encoding on the network layer based on the position number to obtain a position encoding matrix;

[0016] Generate the target embedding matrix based on the first vector, the second vector and the position encoding matrix.

[0017] Preferably, the serializing the hyperparameters of the network layer into corresponding second vectors includes:

[0018] If the type of the network layer is a convolutional layer, serialize the hyperparameters of the network layer into the corresponding second vectors based on a first serialization formula;

[0019] If the type of the network layer is a fully connected layer, serialize the hyperparameters of the network layer into the corresponding second vectors based on a second serialization formula

[0020] The first serialization formula includes:

[0021] H p =(c in ,c out ,s kernel ,s stride ,s feat ,n params ,i dw ,i w / a );

[0022] The second serialization formula includes:

[0023] H P =(h in ,h out ,1,0,s feat ,n params ,0,i w / a )

[0024] Wherein, H p represents the second vector; c in represents the number of input channels; c out represents the number of output channels; s kernel represents the convolution kernel size; s stride represents the size of the sliding step; s feat represents the size of the input feature map; n params represents the number of parameters; i dw represents the binary indicator symbol of the separable convolution; i w / a represents the binary indicator symbol of the weight w or the activation a; h in represents the number of input hidden units; h out represents the number of output hidden units; s feat represents the size of the input feature vector; n params represents the number of parameters.

[0025] Preferably, generating the target embedding matrix based on the first vector, the second vector, and the position encoding matrix includes:

[0026] Combining the first vector and the second vector of the network layer to obtain a third vector of the network layer;

[0027] Concatenating all the third vectors to obtain a vector matrix;

[0028] Generating the target embedding matrix based on the vector matrix and the position encoding matrix.

[0029] Preferably, generating the target embedding matrix based on the vector matrix and the position encoding matrix includes:

[0030] Obtaining a pre-learned target matrix;

[0031] Generating the target embedding matrix according to the matrix generation formula based on the target matrix, the vector matrix, and the position encoding matrix;

[0032] The matrix generation formula includes:

[0033] Z0 = E * X + PE;

[0034] Among them, Z0 represents the target embedding matrix; E represents the target matrix; X represents the vector matrix; and PE represents the position encoding matrix.

[0035] Preferably, the position encoding of the network layer based on the position sequence number includes:

[0036] Performing sinusoidal position encoding on the network layer based on the position sequence number.

[0037] Preferably, the Transformer model includes a preset number of encoder layers for processing the target embedding matrix; a first LayerNorm layer connected to the encoder layer; a first fully connected layer connected to the first LayerNorm layer; a softmax layer connected to the fully connected layer; and a discrete mapping layer connected to the softmax layer;

[0038] Among them, the encoder layer is used to calculate the attention of each row in the target embedding matrix; the discrete mapping layer is used to output the quantization bit number.

[0039] Preferably, each encoder layer includes a second LayerNorm layer connected to the input layer of the encoder layer; a multi-head attention mechanism layer connected to the second LayerNorm layer; a first residual layer connected to the multi-head attention mechanism layer and the input layer; a third LayerNorm layer connected to the residual layer; a feed-forward neural network layer connected to the third LayerNorm layer; and a second residual layer connected to the feed-forward neural network layer and the first residual layer.

[0040] Preferably, the feed-forward neural network layer includes: a second fully connected layer connected to the third LayerNorm layer; a ReLU activation layer connected to the second fully connected layer; and a third fully connected layer connected to the ReLU activation layer.

[0041] Preferably, the operation formula of the quantization bit number includes:

[0042] b p = round(b min - 0.5 + y p × (b max - b min + 1));

[0043] y = softmax(LN(Z L )W o + b o )

[0044] Among them, b pRepresents the quantization bit number of the p-th network layer, where p = [1, 2, … N], and N represents the total number of network layers; round represents the rounding algorithm; b min Represents the minimum value of quantization; b max Represents the maximum value of quantization; L represents the total number of consecutive encoder layers; Z L Represents the processing result of the L-th encoder layer; W o , b o Represents a preset value.

[0045] Preferably, quantizing the target neural network model based on the quantization bit number to obtain a target quantized neural network model includes:

[0046] Statistical weight distribution of the network layer to obtain a weight value distribution result;

[0047] Discard the preset number of weight values at the front and back in the weight value distribution result to obtain the remaining weight values;

[0048] Statistical weight maximum and weight minimum in the remaining weight values;

[0049] Use the maximum value among the weight maximum and the weight minimum as the truncation range;

[0050] Truncate the network layer based on the truncation range to obtain a truncated value;

[0051] Based on the layer-by-layer symmetric quantization algorithm, quantize and dequantize the weight values of the network layer to obtain corresponding quantization results and dequantization results;

[0052] Determine the target quantized neural network model based on the quantization result and the dequantization result;

[0053] Among them, the layer-by-layer symmetric quantization algorithm includes:

[0054] w q = round(clamp(w, c) / s p ); w′ = w q s p ;

[0055] Among them, w q Represents the quantization result of the q-th weight value in the p-th network layer; clamp(w, c) represents truncating the weight value w to [-c, c], where c represents the truncated value; w′ represents the dequantization result of the q-th weight value in the p-th network layer.

[0056] Preferably, the loss function of the Transformer model includes:

[0057] L(w,w') = λ(Y F (x,w) - Y Q (x,w'));

[0058]

[0059] where Loss(w,w') represents the value of the loss function; λ represents a hyperparameter that adjusts the initial value around 1; x represents the test set images; Y F (x,w) represents the accuracy of the floating-point model F; Y Q (x,w') represents the accuracy of the quantized model Q; log represents the logarithmic function; γ represents the weight that adjusts the model Size and the loss term.

[0060] A neural network model quantization system, comprising:

[0061] A first acquisition module, configured to acquire the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized;

[0062] A first generation module, configured to perform linear embedding on the weight values, the hyperparameters, and the position serial numbers to generate a target embedding matrix;

[0063] A first processing module, configured to process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model;

[0064] A first quantization module, configured to quantize the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model.

[0065] An electronic device, comprising:

[0066] A memory, configured to store a computer program;

[0067] A processor, configured to implement the steps of any one of the above neural network model quantization methods when executing the computer program.

[0068] A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above neural network model quantization methods are implemented.

[0069] A neural network model quantization method provided by the present application obtains the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized; performs linear embedding on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix; processes the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model; and quantizes the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model. In the present application, by using the Transformer model to process the target embedding matrix generated by the weight values, hyperparameters, and position serial numbers to obtain the quantization bit numbers of each layer in the target neural network model, the model size and memory occupancy can be reduced, and at the same time, the precision loss of the original network is small. In addition, since it is not necessary to consider each bit number layer by layer, the amount of computation can be greatly reduced, and the limitation is low. A neural network model quantization system, an electronic device, and a computer-readable storage medium provided by the present application also solve the corresponding technical problems. Description of the Drawings

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0071] Figure 1 It is the first flowchart of a neural network model quantization method provided by an embodiment of the present application;

[0072] Figure 2 It is the second flowchart of a neural network model quantization method provided by an embodiment of the present application;

[0073] Figure 3 It is the structural schematic diagram of the Transformer model in a neural network model quantization method provided by an embodiment of the present application;

[0074] Figure 4 It is the structural schematic diagram of the encoder layer;

[0075] Figure 5 It is the third flowchart of a neural network model quantization method provided by an embodiment of the present application;

[0076] Figure 6 It is the structural schematic diagram of a neural network model quantization system provided by an embodiment of the present application;

[0077] Figure 7 It is the structural schematic diagram of an electronic device provided by an embodiment of the present application;

[0078] Figure 8 Another structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0079] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0080] With the development of artificial intelligence technology, while the accuracy of deep neural network models is continuously improving, the number of model parameters is also increasing rapidly. This makes the model have problems such as a large number of model parameters and high computational complexity during actual deployment, especially for some edge devices, these problems are particularly obvious. To solve these problems, the neural network model can be quantized. Model quantization belongs to a mature and commonly used method in the field of model compression, and can convert a floating-point model into an integer model that occupies a smaller number of bits (binary digits, bits). 8-bit quantization is already relatively mature in the industrial field, but the compression effect on the model is limited. Extremely low 1-bit or 2-bit quantization has also been studied a lot in the academic field, but it is often accompanied by a decrease in model accuracy. Since the importance of each layer of the model is different, if different quantization bit numbers are used for each layer, that is, mixed-precision quantization, the model accuracy can be guaranteed while reducing the number of parameters.

[0081] Existing mixed-precision quantization methods include using reinforcement learning methods to obtain the quantization bit positions of each layer under the constraints of hardware latency and accuracy loss requirements, or using neural network architecture search methods. By setting multiple edges representing different bit positions on the nodes of each layer, and then finding the optimal bit distribution through random search methods. However, the disadvantages of these two methods are long time periods and the need for multiple GPU resources to provide computing power support. Another method is to use the hessian matrix to calculate the relative sensitivity of each layer of the network to provide a reference for the bit distribution of each layer. However, the disadvantage is that it can only calculate the importance of the weights of each layer in the loss unidirectionally from back to front and cannot reflect the cross-layer influence. This makes the limitations of existing methods for quantizing neural network models high. To solve this technical problem, the present application provides a neural network model quantization method, system, electronic device, and computer-readable storage medium.

[0082] Please refer to Figure 1 , Figure 1 The first flowchart of a neural network model quantization method provided by an embodiment of the present application.

[0083] A neural network model quantization method provided by an embodiment of the present application may include the following steps:

[0084] Step S101: Obtain the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized.

[0085] In practical applications, the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized can be obtained first. Among them, the type of the target neural network model can be determined according to actual needs, and the number, type of network layers, and the weight values, hyperparameters, and position serial numbers of each network layer can also be determined according to actual needs. Among them, the position serial number can be the serial number of the network layer in the target neural network model, etc. The present application does not make specific limitations here.

[0086] Step S102: Perform linear embedding on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix.

[0087] In practical applications, after obtaining the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized, linear embedding can be performed on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix, so as to process the weight values, hyperparameters, and position serial numbers with the help of the target embedding matrix.

[0088] Step S103: Process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model.

[0089] In practical applications, after performing linear embedding on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix, the target embedding matrix can be processed based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model. In other words, the present application uses the Transformer model to process the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to obtain the quantization bit numbers of each network layer in the target neural network model.

[0090] Step S104: Quantize the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model.

[0091] In practical applications, after processing the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model, the target neural network model can be quantized based on the quantization bit numbers to obtain the target quantized neural network model. It should be noted that the process of quantizing the target neural network model based on the quantization bit numbers can be determined according to actual needs, and this application does not make specific limitations here.

[0092] A neural network model quantization method provided by this application includes obtaining the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized; performing linear embedding on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix; processing the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each layer in the target neural network model; and quantizing the target neural network model based on the quantization bit numbers to obtain the target quantized neural network model. In this application, by using the Transformer model to process the target embedding matrix generated by the weight values, hyperparameters, and position serial numbers to obtain the quantization bit numbers of each layer in the target neural network model, the model size and memory occupancy can be reduced, and at the same time, the accuracy loss of the original network is relatively small. In addition, since there is no need to consider each bit number layer by layer, the amount of computation can be greatly reduced, and the limitations are low.

[0093] Please refer to Figure 2 , Figure 2 which is the second flowchart of a neural network model quantization method provided by an embodiment of this application.

[0094] A neural network model quantization method provided by an embodiment of this application may include the following steps:

[0095] Step S201: Obtain the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized.

[0096] Step S202: Flatten the weight values of the network layer into corresponding first vectors.

[0097] Step S203: Serialize the hyperparameters of the network layer into corresponding second vectors.

[0098] Step S204: Perform position encoding on the network layer based on the position serial number to obtain a position encoding matrix.

[0099] Step S205: Generate a target embedding matrix based on the first vector, the second vector, and the position encoding matrix.

[0100] In practical applications, in the process of linearly embedding weight values, hyperparameters, and position serial numbers to generate a target embedding matrix, a vectorization method can be used to generate the target embedding matrix. That is, for each network layer, the weight values of the network layer can be flattened into a corresponding first vector, the hyperparameter sequence of the network layer can be serialized into a corresponding second vector, the network layer is position-encoded based on the position serial number to obtain a position encoding matrix, and then the corresponding target embedding matrix is generated based on the first vector, the second vector, and the position encoding matrix.

[0101] In a specific application scenario, in the process of flattening the weight values of the network layer into a corresponding first vector, the number of parameters of all network layers that need to be quantized in the target neural network model can be first counted, the maximum value of the number of parameters can be determined, and finally the weight values of each network layer can be flattened into a corresponding one-dimensional vector w p ∈R d , p ∈ [1, 2, 3, … N], where N represents the total number of network layers. Those with fewer parameters than d can be supplemented with 0, that is, the first vector is sufficient, so as to unify the lengths of the first vectors of all network layers to d, facilitating subsequent batch processing of the first vectors.

[0102] In a specific application scenario, in the process of serializing the hyperparameters of the network layer into a corresponding second vector, the hyperparameters can be accurately serialized into a corresponding second vector according to the type of the network layer. Specifically, if the type of the network layer is a convolutional layer, the hyperparameters of the network layer are serialized into a corresponding second vector based on the first serialization formula; if the type of the network layer is a fully connected layer, the hyperparameters of the network layer are serialized into a corresponding second vector based on the second serialization formula

[0103] The first serialization formula includes:

[0104] H p =(c in , c out , s kernel , s stride , s feat , n params , i dw , i w / a );

[0105] The second serialization formula includes:

[0106] H P =(h in , h out , 1, 0, s feat , n params , 0, i w / a )

[0107] Among them, H p represents the second vector; cin Indicates the number of input channels; c out Indicates the number of output channels; s kernel Indicates the size of the convolutional kernel; s stride Indicates the size of the sliding step; s feat Indicates the size of the input feature map; n params Indicates the number of parameters; i dw Indicates the binary indicator symbol of the depthwise separable convolution; i w / a Indicates the binary indicator symbol of the weight w or the activation a; h in Indicates the number of input hidden units; h out Indicates the number of output hidden units; s feat Indicates the size of the input feature vector; n params Indicates the number of parameters.

[0108] In a specific application scenario, during the process of generating the target embedding matrix based on the first vector, the second vector, and the position encoding matrix, for each network layer, the first vector and the second vector of the network layer can be merged to obtain the third vector of the network layer. Assume the third vector is represented by x p Indicates, then x p = Concat(w p , H p ), Concatenate all the third vectors to obtain a vector matrix. Assume the vector matrix is represented by X, then X = [x1; x2;...; x N ; Generate the target embedding matrix based on the vector matrix and the position encoding matrix.

[0109] In a specific application scenario, during the process of generating the target embedding matrix based on the vector matrix and the position encoding matrix, a pre-learned target matrix can be obtained; according to the matrix generation formula, generate the target embedding matrix based on the target matrix, the vector matrix, and the position encoding matrix; the matrix generation formula includes:

[0110] Z0 = E * X + PE;

[0111] where, Z0 represents the target embedding matrix; E represents the target matrix; X represents the vector matrix; PE represents the position encoding matrix.

[0112] In a specific application scenario, during the process of performing position encoding on the network layer based on the position index, sinusoidal position encoding can be performed on the network layer based on the position index, and its encoding process can be as follows:

[0113] PE (p,2j) = sin(p / 10000 2j / D ); PE (p,2j+1) = cos(p / 10000 2j / D);

[0114] Among them, D = d + 8; j ∈ [0, 1,..., (D - 1) / 2], representing the dimension serial number; of course, there can also be other position encoding methods, which are not specifically limited in this application.

[0115] Step S206: Process the target embedding matrix based on the pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model.

[0116] Step S207: Quantize the target neural network model based on the quantization bit numbers to obtain the target quantized neural network model.

[0117] Please refer to Figure 3 and Figure 4 , Figure 3 , which is a schematic structural diagram of the Transformer model in a neural network model quantization method provided by an embodiment of this application, Figure 4 is a schematic structural diagram of the encoder layer.

[0118] In a neural network model quantization method provided by an embodiment of this application, the Transformer model may include a preset number of encoder layers for processing the target embedding matrix; a first LayerNorm (layer normalization) layer connected to the encoder layer; a first fully connected layer connected to the first LayerNorm layer; a softmax (normalization) layer connected to the fully connected layer; a discrete mapping layer connected to the softmax layer; among them, the encoder layer is used to calculate the attention of each row in the target embedding matrix; the discrete mapping layer is used to output the quantization bit numbers. It should be noted that the value of the preset number can be determined according to the specific application scenario. For example, the value of the preset number can be 6, etc., which is not specifically limited in this application.

[0119] In practical applications, as Figure 4 shown, each encoder layer in this application may include a second LayerNorm layer connected to the input layer of the encoder layer; a multi-head attention mechanism (Mualti-HeadAttention, MHA) layer connected to the second LayerNorm, etc.; a first residual layer connected to the multi-head attention mechanism layer and the input layer; a third LayerNorm layer connected to the residual layer; a feed-forward neural network (Feed-Forward Networks, FFN) layer connected to the third LayerNorm layer; a second residual layer connected to the feed-forward neural network layer and the first residual layer.

[0120] In a specific application scenario, the feedforward neural network layer may include: a second fully connected layer connected to the third LayerNorm layer; a ReLU activation layer connected to the second fully connected layer; and a third fully connected layer connected to the ReLU activation layer.

[0121] It should be noted that in the encoder layer, assuming that the input of the l-th encoder layer is Z l-1 , where l ∈ [1, 2... L] and L represents the total number of encoder layers. Here, Z0 represents the target embedding matrix. Then, we can first calculate the query vector query, key vector key, and value vector value corresponding to each Z l-1 , which are simply referred to as q, k, and v vectors respectively. Then

[0122]

[0123] where LN() represents LayerNorm; a ∈ [1, 2... A] represents each head in the multi-head attention mechanism, and A represents the total number of heads; represents the learnable mapping parameter matrix, and the dimension D of each head h = D / A, and both q, k, and v

[0124] Then, we calculate the attention output of each head

[0125]

[0126] Then, the outputs of each head are concatenated and projected to obtain the output of the multi-head attention mechanism:

[0127] FFN(x) = max(0, xW1 + b1)W2 + b2 ∈ R N×D , and FFN(x) ∈ R N×D ;

[0128] where W1, W2 ∈ R D×D , and b2 ∈ R N ;

[0129] After passing through L encoder blocks, the output Z L still has the dimension of R N×D , and finally, after passing through the LayerNorm layer, the fully connected layer, and the softmax layer, the final output of the model is a one-dimensional vector y ∈ R N , and the range of y is [0, 1]:

[0130] y = softmax(MLP(LN(Z L))) = softmax(LN(Z L )W o +b o ));

[0131] Finally, y is discretized to the corresponding bits:

[0132] b p = round(b min - 0.5 + y p × (b max - b min + 1));

[0133] That is, in this application, the operation formula for the quantization bit number can include:

[0134] b p = round(b min - 0.5 + y p × (b max - b min + 1));

[0135] y = softmax(LN(Z L )W o +b o );

[0136] Wherein, b p represents the quantization bit number of the p-th network layer, p = [1, 2,... N], N represents the total number of network layers; round represents the rounding algorithm; b min represents the minimum value of quantization; b max represents the maximum value of quantization; L represents the total number of sequentially connected encoder layers; Z L represents the processing result of the L-th encoder layer; W o , b o represent preset values.

[0137] It should be noted that the values of b min , b max can be determined according to actual needs. For example, the value of b min can be 2, and the value of b max can be 8, etc. This application does not make specific limitations here.

[0138] Please refer to Figure 5 , Figure 5 which is the third flowchart of a neural network model quantization method provided by an embodiment of this application.

[0139] A neural network model quantization method provided by an embodiment of this application may include the following steps:

[0140] Step S301: Obtain the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized.

[0141] Step S302: Perform linear embedding on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix.

[0142] Step S303: Process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model.

[0143] Step S304: Statistically analyze the weight distribution of the network layer to obtain a weight value distribution result.

[0144] Step S305: Discard a preset number of weight values at the front and back of the weight value distribution result to obtain the remaining weight values.

[0145] Step S306: Statistically analyze the maximum and minimum weight values among the remaining weight values.

[0146] Step S307: Use the maximum value among the maximum and minimum weight values as the truncation range.

[0147] Step S308: Truncate the network layer based on the truncation range to obtain truncation values.

[0148] Step S309: Quantize and de-quantize the weight values of the network layer based on the layer-by-layer symmetric quantization algorithm to obtain corresponding quantization results and de-quantization results.

[0149] In practical applications, during the process of quantizing the target neural network model based on the quantization bit numbers to obtain the target quantized neural network model, the weight values of the target neural network model can be quantized based on the quantization bit numbers. Specifically, the weight distribution of the network layer can be statistically analyzed to obtain a weight value distribution result; a preset number of weight values at the front and back of the weight value distribution result can be discarded, for example, 1% of the weight values at the front and back of the weight value distribution result are discarded, to obtain the remaining weight values; the maximum weight value w max and the minimum weight value w min among the remaining weight values are statistically analyzed; the maximum value among the maximum and minimum weight values is used as the truncation range, c = max(|w min |, |w max |); the network layer is truncated based on the truncation range to obtain truncation values [-c, c]; the weight values of the network layer are quantized and de-quantized based on the layer-by-layer symmetric quantization algorithm to obtain corresponding quantization results and de-quantization results; the target quantized neural network model is determined based on the quantization results and de-quantization results;

[0150] Among them, the layer-by-layer symmetric quantization algorithm includes:

[0151] w q = round(clamp(w, c) / s p )); w' = w q s p ;

[0152] Among them, w q represents the quantization result of the q-th weight value in the p-th network layer; clamp(w, c) represents truncating the weight value w to [-c, c], where c represents the truncation value; w' represents the dequantization result of the q-th weight value in the p-th network layer.

[0153] In a specific application scenario, the loss function of the Transformer model may include:

[0154] L(w, w') = λ(Y F (x, w) - Y Q (x, w'));

[0155]

[0156] Among them, Loss(w, w') represents the loss function value; λ represents a hyperparameter that adjusts the initial value near 1; x represents the test set image; Y F (x, w) represents the accuracy of the floating-point model F; Y Q (x, w') represents the accuracy of the quantization model Q; log represents the logarithmic function; γ represents the weight that adjusts the model Size and the loss term. Specifically, during the training process, a threshold can be set before training. When the loss during training is greater than the threshold, backpropagation is performed, and the gradient is calculated by minimizing the loss function, that is, min(Loss(w, w')), to update the weights of the Transformer model; when the loss during training is lower than the threshold, the output of the Transformer model is that the optimal quantization bit of the target model can minimize the loss function.

[0157] Please refer to Figure 6 , Figure 6 , which is the structural schematic diagram of a neural network model quantization system provided by an embodiment of this application.

[0158] A neural network model quantization system provided by an embodiment of this application may include:

[0159] The first acquisition module 101 is used to acquire the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized;

[0160] The first generation module 102 is used to perform linear embedding on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix;

[0161] The first processing module 103 is configured to process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each of the network layers in the target neural network model;

[0162] The first quantization module 104 is configured to quantize the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model.

[0163] For a neural network model quantization system provided by an embodiment of the present application, the first generation module may include:

[0164] The first flattening sub-module is configured to flatten the weight values of the network layer into corresponding first vectors;

[0165] The first serialization sub-module is configured to serialize the hyperparameters of the network layer into corresponding second vectors;

[0166] The first encoding sub-module is configured to perform position encoding on the network layer based on the position serial number to obtain a position encoding matrix;

[0167] The first generation sub-module is configured to generate a target embedding matrix based on the first vector, the second vector, and the position encoding matrix.

[0168] For a neural network model quantization system provided by an embodiment of the present application, the first serialization sub-module may include:

[0169] The first serialization unit is configured to, if the type of the network layer is a convolutional layer, serialize the hyperparameters of the network layer into corresponding second vectors based on the first serialization formula;

[0170] The second serialization unit is configured to, if the type of the network layer is a fully connected layer, serialize the hyperparameters of the network layer into corresponding second vectors based on the second serialization formula

[0171] The first serialization formula includes:

[0172] H p =(c in ,c out ,s kernel ,s stride ,s feat ,n params ,i dw ,i w / a );

[0173] The second serialization formula includes:

[0174] H P =(h in ,h out ,1,0,sfeat ,n params ,0,i w / a )

[0175] Among them, H p represents the second vector; c in represents the number of input channels; c out represents the number of output channels; s kernel represents the size of the convolutional kernel; s stride represents the size of the sliding step; s feat represents the size of the input feature map; n params represents the number of parameters; i dw represents the binary indicator symbol of the separable convolution; i w / a represents the binary indicator symbol of the weight w or the activation a; h in represents the number of input hidden units; h out represents the number of output hidden units; s feat represents the size of the input feature vector; n params represents the number of parameters.

[0176] A neural network model quantization system provided by an embodiment of the present application, the first generation sub-module may include:

[0177] The first merging unit is used to merge the first vector and the second vector of the network layer to obtain the third vector of the network layer;

[0178] The first splicing unit is used to splice all the third vectors to obtain a vector matrix;

[0179] The first generation unit is used to generate a target embedding matrix based on the vector matrix and the position encoding matrix.

[0180] A neural network model quantization system provided by an embodiment of the present application, the first generation unit may specifically be used to: obtain a pre-learned target matrix; generate a target embedding matrix based on the target matrix, the vector matrix, and the position encoding matrix according to the matrix generation formula; the matrix generation formula includes:

[0181] Z0 = E * X + PE;

[0182] Among them, Z0 represents the target embedding matrix; E represents the target matrix; X represents the vector matrix; PE represents the position encoding matrix.

[0183] A neural network model quantization system provided by an embodiment of the present application, the first encoding sub-module may include:

[0184] The first encoding unit is used to perform sinusoidal position encoding on the network layer based on the position sequence number.

[0185] A neural network model quantization system provided by an embodiment of the present application. The Transformer model includes a preset number of encoder layers for processing a target embedding matrix; a first LayerNorm layer connected to the encoder layer; a first fully connected layer connected to the first LayerNorm layer; a softmax layer connected to the fully connected layer; a discrete mapping layer connected to the softmax layer;

[0186] Among them, the encoder layer is used to calculate the attention of each row in the target embedding matrix; the discrete mapping layer is used to output the quantization bit number.

[0187] A neural network model quantization system provided by an embodiment of the present application. Each encoder layer includes a second LayerNorm layer connected to the input layer of the encoder layer; a multi-head attention mechanism layer connected to the second LayerNorm layer; a first residual layer connected to the multi-head attention mechanism layer and the input layer; a third LayerNorm layer connected to the residual layer; a feed-forward neural network layer connected to the third LayerNorm layer; a second residual layer connected to the feed-forward neural network layer and the first residual layer.

[0188] A neural network model quantization system provided by an embodiment of the present application. The feed-forward neural network layer includes: a second fully connected layer connected to the third LayerNorm layer; a ReLU activation layer connected to the second fully connected layer; a third fully connected layer connected to the ReLU activation layer.

[0189] A neural network model quantization system provided by an embodiment of the present application. The operation formula for the quantization bit number includes:

[0190] b p = round(b min - 0.5 + y p × (b max - b min + 1));

[0191] y = softmax(LN(Z L )W o + b o )

[0192] Among them, b p represents the quantization bit number of the p-th network layer, p = [1, 2,... N], N represents the total number of network layers; round represents the rounding algorithm; b min represents the minimum value of quantization; b max represents the maximum value of quantization; L represents the total number of sequentially connected encoder layers; Z L represents the processing result of the L-th encoder layer; Wo and b o represent preset values.

[0193] A neural network model quantization system provided by an embodiment of the present application. The first quantization module may include:

[0194] A first statistical unit for statistically analyzing the weight distribution of a network layer to obtain a weight value distribution result;

[0195] A first discarding unit for discarding a preset number of weight values at the front and back of the weight value distribution result to obtain remaining weight values;

[0196] A second statistical unit for statistically analyzing the maximum weight value and the minimum weight value among the remaining weight values;

[0197] A first setting unit for using the maximum value among the maximum weight value and the minimum weight value as a truncation range;

[0198] A first truncation unit for truncating the network layer based on the truncation range to obtain a truncation value;

[0199] A first quantization unit for quantizing and de - quantizing the weight values of the network layer based on a layer - by - layer symmetric quantization algorithm to obtain corresponding quantization results and de - quantization results;

[0200] A first determination unit for determining a target quantized neural network model based on the quantization results and the de - quantization results;

[0201] Among them, the layer - by - layer symmetric quantization algorithm includes:

[0202] w q = round(clamp(w,c) / s p ); w' = w q s p ;

[0203] Among them, w q represents the quantization result of the q - th weight value in the p - th network layer; clamp(w,c) represents truncating the weight value w to [-c,c], where c represents the truncation value; w' represents the de - quantization result of the q - th weight value in the p - th network layer.

[0204] A loss function of the Transformer model in a neural network model quantization system provided by an embodiment of the present application includes:

[0205] L(w,w') = λ(Y F (x,w)-Y Q (x,w'));

[0206]

[0207] Among them, Loss(w, w') represents the loss function value; λ represents a hyperparameter for adjusting the initial value around 1; x represents the test set images; Y F (x, w) represents the accuracy of the floating-point model F; Y Q (x, w') represents the accuracy of the quantized model Q; log represents the logarithmic function; γ represents the weight for adjusting the model Size and the loss term.

[0208] This application also provides an electronic device and a computer-readable storage medium, both of which have the corresponding effects of a neural network model quantization method provided by the embodiments of this application. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device provided by the embodiments of this application.

[0209] An electronic device provided by the embodiments of this application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented:

[0210] Obtain the weight values, hyperparameters, and position sequence numbers of each network layer in the target neural network model to be quantized;

[0211] Perform linear embedding on the weight values, hyperparameters, and position sequence numbers to generate a target embedding matrix;

[0212] Process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model;

[0213] Quantize the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model.

[0214] An electronic device provided by the embodiments of this application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: Flatten the weight values of the network layer into corresponding first vectors; Serialize the hyperparameter sequences of the network layer into corresponding second vectors; Perform position encoding on the network layer based on the position sequence numbers to obtain a position encoding matrix; Generate a target embedding matrix based on the first vector, the second vector, and the position encoding matrix.

[0215] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: If the type of the network layer is a convolutional layer, then based on the first serialization formula, serialize the hyperparameters of the network layer into a corresponding second vector; if the type of the network layer is a fully connected layer, then based on the second serialization formula, serialize the hyperparameters of the network layer into a corresponding second vector

[0216] The first serialization formula includes:

[0217] H p =(c in ,c out ,s kernel ,s stride ,s feat ,n params ,i dw ,i w / a );

[0218] The second serialization formula includes:

[0219] H P =(h in ,h out ,1,0,s feat ,n params ,0,i w / a )

[0220] Wherein, H p represents the second vector; c in represents the number of input channels; c out represents the number of output channels; s kernel represents the kernel size; s stride represents the size of the sliding step; s feat represents the size of the input feature map; n params represents the number of parameters; i dw represents the binary indication symbol of depthwise separable convolution; i w / a represents the binary indication symbol of weight w or activation a; h in represents the number of input hidden units; h out represents the number of output hidden units; s feat represents the size of the input feature vector; n params represents the number of parameters.

[0221] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: combining a first vector and a second vector in a network layer to obtain a third vector in the network layer; concatenating all the third vectors to obtain a vector matrix; generating a target embedding matrix based on the vector matrix and a position encoding matrix.

[0222] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: obtaining a pre-learned target matrix; generating a target embedding matrix based on the target matrix, the vector matrix, and the position encoding matrix according to a matrix generation formula.

[0223] The matrix generation formula includes:

[0224] Z0 = E * X + PE;

[0225] Wherein, Z0 represents the target embedding matrix; E represents the target matrix; X represents the vector matrix; PE represents the position encoding matrix.

[0226] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: performing sine position encoding on the network layer based on a position serial number.

[0227] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: The Transformer model includes a preset number of encoder layers for processing the target embedding matrix; a first LayerNorm layer connected to the encoder layer; a first fully connected layer connected to the first LayerNorm layer; a softmax layer connected to the fully connected layer; a discrete mapping layer connected to the softmax layer; wherein, the encoder layer is used to calculate the attention of each row in the target embedding matrix; the discrete mapping layer is used to output the quantization bit number.

[0228] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: Each encoder layer includes a second LayerNorm layer connected to the input layer of the encoder layer; a multi-head attention mechanism layer connected to the second LayerNorm layer; a first residual layer connected to the multi-head attention mechanism layer and the input layer; a third LayerNorm layer connected to the residual layer; a feed-forward neural network layer connected to the third LayerNorm layer; a second residual layer connected to the feed-forward neural network layer and the first residual layer.

[0229] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: The feed-forward neural network layer includes: a second fully-connected layer connected to the third LayerNorm layer; a ReLU activation layer connected to the second fully-connected layer; a third fully-connected layer connected to the ReLU activation layer.

[0230] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: The operation formula for the quantization bit number includes:

[0231] b p = round(b min - 0.5 + y p × (b max - b min + 1));

[0232] y = softmax(LN(Z L )W o + b o )

[0233] Wherein, b p represents the quantization bit number of the p-th network layer, p = [1, 2,... N], N represents the total number of network layers; round represents the rounding algorithm; b min represents the minimum value of quantization; b max represents the maximum value of quantization; L represents the total number of sequentially connected encoder layers; Z L represents the processing result of the L-th encoder layer; W o , b o represent preset values.

[0234] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: statistically analyze the weight distribution of the network layer to obtain a weight value distribution result; discard the preset number of weight values at the front and back of the weight value distribution result to obtain the remaining weight values; statistically analyze the maximum weight value and the minimum weight value among the remaining weight values; use the maximum value of the maximum weight value and the minimum weight value as the truncation range; truncate the network layer based on the truncation range to obtain a truncation value; based on the layer-by-layer symmetric quantization algorithm, quantize and dequantize the weight values of the network layer to obtain corresponding quantization results and dequantization results; determine a target quantization neural network model based on the quantization results and dequantization results; where the layer-by-layer symmetric quantization algorithm includes:

[0235] w q = round(clamp(w, c) / s p ); w' = w q s p ;

[0236] where w q represents the quantization result of the q-th weight value in the p-th network layer; clamp(w, c) represents truncating the weight value w to [-c, c], and c represents the truncation value; w' represents the dequantization result of the q-th weight value in the p-th network layer.

[0237] An electronic device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: The loss function of the Transformer model includes:

[0238] L(w, w') = λ(Y F (x, w) - Y Q (x, w'));

[0239]

[0240] where Loss(w, w') represents the loss function value; λ represents a hyperparameter for adjusting the initial value near 1; x represents a test set image; Y F (x, w) represents the accuracy of the floating-point model F; Y Q (x, w') represents the accuracy of the quantization model Q; log represents the logarithmic function; γ represents the weight for adjusting the model Size and the loss term.

[0241] Please refer to Figure 8, another electronic device provided by an embodiment of the present application may further include: an input port 203 connected to the processor 202, configured to transmit an externally input command to the processor 202; a display unit 204 connected to the processor 202, configured to display the processing result of the processor 202 to the outside; a communication module 205 connected to the processor 202, configured to implement communication between the electronic device and the outside. The display unit 204 may be a display panel, a laser scanning display, etc.; the communication methods adopted by the communication module 205 include but are not limited to Mobile High-Definition Link technology (HML), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connections: Wi-Fi technology, Bluetooth communication technology, Low Energy Bluetooth communication technology, communication technology based on IEEE802.11s.

[0242] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0243] Obtain the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized;

[0244] Perform linear embedding on the weight values, hyperparameters, and position serial numbers to generate a target embedding matrix;

[0245] Process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model;

[0246] Quantize the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model.

[0247] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: flatten the weight value of the network layer into a corresponding first vector; serialize the hyperparameter sequence of the network layer into a corresponding second vector; perform position encoding on the network layer based on the position serial number to obtain a position encoding matrix; generate a target embedding matrix based on the first vector, the second vector, and the position encoding matrix.

[0248] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: if the type of the network layer is a convolutional layer, serialize the hyperparameters of the network layer into a corresponding second vector based on a first serialization formula; if the type of the network layer is a fully connected layer, serialize the hyperparameters of the network layer into a corresponding second vector based on a second serialization formula

[0249] The first serialization formula includes:

[0250] H p =(c in , c out , s kernel , s stride , s feat , n params , i dw , i w / a );

[0251] The second serialization formula includes:

[0252] H P =(h in , h out , 1, 0, s feat , n params , 0, i w / a )

[0253] Among them, H p represents the second vector; c in represents the number of input channels; c out represents the number of output channels; s kernel represents the size of the convolutional kernel; s stride represents the size of the sliding step; s feat represents the size of the input feature map; n params represents the number of parameters; i dw represents the binary indicator symbol of the depthwise separable convolution; i w / a represents the binary indicator symbol of the weight w or the activation a; h in represents the number of input hidden units; h out represents the number of output hidden units; s feat represents the size of the input feature vector; n params represents the number of parameters.

[0254] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: combining the first vector and the second vector of the network layer to obtain a third vector of the network layer; concatenating all the third vectors to obtain a vector matrix; generating a target embedding matrix based on the vector matrix and the position encoding matrix.

[0255] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: obtaining a pre-learned target matrix; generating a target embedding matrix based on the target matrix, the vector matrix, and the position encoding matrix according to the matrix generation formula;

[0256] The matrix generation formula includes:

[0257] Z0 = E * X + PE;

[0258] Among them, Z0 represents the target embedding matrix; E represents the target matrix; X represents the vector matrix; PE represents the position encoding matrix.

[0259] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: performing sinusoidal position encoding on the network layer based on the position sequence number.

[0260] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: The Transformer model includes a preset number of encoder layers for processing the target embedding matrix; a first LayerNorm layer connected to the encoder layer; a first fully connected layer connected to the first LayerNorm layer; a softmax layer connected to the fully connected layer; a discrete mapping layer connected to the softmax layer; among them, the encoder layer is used to calculate the attention of each row in the target embedding matrix; the discrete mapping layer is used to output the quantization bit number.

[0261] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: Each encoder layer includes a second LayerNorm layer connected to the input layer of the encoder layer; a multi-head attention mechanism layer connected to the second LayerNorm layer; a first residual layer connected to the multi-head attention mechanism layer and the input layer; a third LayerNorm layer connected to the residual layer; a feed-forward neural network layer connected to the third LayerNorm layer; a second residual layer connected to the feed-forward neural network layer and the first residual layer.

[0262] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: The feed-forward neural network layer includes: a second fully connected layer connected to the third LayerNorm layer; a ReLU activation layer connected to the second fully connected layer; a third fully connected layer connected to the ReLU activation layer.

[0263] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: The operation formula of the quantization bit number includes:

[0264] b p = round(b min-0.5 + y p ×(b max -b min + 1));

[0265] y = softmax(LN(Z L )W o + b o ));

[0266] wherein, b p represents the quantization bit number of the p-th network layer, p = [1, 2,... N], N represents the total number of network layers; round represents the rounding algorithm; b min represents the minimum value of quantization; b max represents the maximum value of quantization; L represents the total number of sequentially connected encoder layers; Z L represents the processing result of the L-th encoder layer; W o , b o represent preset values.

[0267] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: statistically analyze the weight distribution of the network layer to obtain a weight value distribution result; discard a preset number of weight values at the front and back of the weight value distribution result to obtain remaining weight values; statistically analyze the maximum weight value and the minimum weight value in the remaining weight values; use the maximum value of the maximum weight value and the minimum weight value as the truncation range; perform truncation on the network layer based on the truncation range to obtain truncation values; perform quantization and inverse quantization on the weight values of the network layer based on the layer-by-layer symmetric quantization algorithm to obtain corresponding quantization results and inverse quantization results; determine a target quantization neural network model based on the quantization results and the inverse quantization results; wherein, the layer-by-layer symmetric quantization algorithm includes:

[0268] w q = round(clamp(w, c) / s p )); w' = w q s p ;

[0269] wherein, w q represents the quantization result of the q-th weight value in the p-th network layer; clamp(w, c) represents truncating the weight value w to [-c, c], c represents the truncation value; w' represents the inverse quantization result of the q-th weight value in the p-th network layer.

[0270] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: The loss function of the Transformer model includes:

[0271] L(w,w') = λ(Y F (x,w) - Y Q (x,w'));

[0272]

[0273] where Loss(w,w') represents the loss function value; λ represents a hyperparameter that adjusts the initial value around 1; x represents the test set images; Y F (x,w) represents the accuracy of the floating-point model F; Y Q (x,w') represents the accuracy of the quantization model Q; log represents the logarithmic function; γ represents the weight that adjusts the model Size and the loss term.

[0274] The computer-readable storage medium involved in the present application includes random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.

[0275] For the description of the relevant parts in a neural network model quantization system, an electronic device, and a computer-readable storage medium provided by an embodiment of the present application, please refer to the corresponding detailed description in a neural network model quantization method provided by an embodiment of the present application, which will not be elaborated here. In addition, the parts in the above technical solutions provided by the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.

[0276] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0277] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A neural network model quantization method, characterized in that, Including: Obtain the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized; Perform linear embedding on the weight values, the hyperparameters, and the position serial numbers to generate a target embedding matrix; Process the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model; Quantize the target neural network model based on the quantization bit numbers to obtain a target quantized neural network model, so as to reduce the model size and memory occupancy; Deploy the target quantized neural network model to a device; Wherein, the target neural network model can be used to process pictures; and the Transformer model includes a model trained by test set pictures; Wherein, the operation formula of the quantization bit numbers includes: ; ; in, Indicates The number of quantization bits of the network layer, , Indicates the total number of layers of the network; Indicates the rounding algorithm; Indicates the minimum value of quantization; Indicates the maximum value of quantization; Indicates the total number of encoder layers connected in sequence; Indicates The processing result of the encoder layer; , Indicates the preset value.

2. The method according to claim 1, wherein The performing linear embedding on the weight values, the hyperparameters, and the position serial numbers to generate a target embedding matrix includes: Flatten the weight values of the network layer into corresponding first vectors; Serialize the hyperparameters of the network layer into corresponding second vectors; Perform position encoding on the network layer based on the position serial number to obtain a position encoding matrix; Generate the target embedding matrix based on the first vector, the second vector, and the position encoding matrix.

3. The method according to claim 2, wherein The serializing the hyperparameters of the network layer into corresponding second vectors includes: If the type of the network layer is a convolutional layer, serialize the hyperparameters of the network layer into the corresponding second vectors based on a first serialization formula; If the type of the network layer is a fully connected layer, serialize the hyperparameters of the network layer into the corresponding second vectors based on a second serialization formula; The first serialization formula includes: ; The second serialization formula includes: ; Among them, represents the second vector; represents the number of input channels; represents the number of output channels; represents the convolutional kernel size; represents the size of the sliding step; represents the size of the input feature map; represents the number of parameters; represents the binary indicator symbol of the separable convolution; represents the weight or activation binary indicator symbol; represents the number of input hidden units; represents the number of output hidden units; represents the size of the input feature vector; represents the number of parameters.

4. The method according to claim 2, wherein The generating the target embedding matrix based on the first vector, the second vector, and the position encoding matrix includes: Merge the first vector and the second vector of the network layer to obtain a third vector of the network layer; Concatenate all the third vectors to obtain a vector matrix; Generate the target embedding matrix based on the vector matrix and the position encoding matrix.

5. The method according to claim 4, wherein The generating the target embedding matrix based on the vector matrix and the position encoding matrix includes: Obtain a pre-learned target matrix; Generate the target embedding matrix based on the target matrix, the vector matrix, and the position encoding matrix according to a matrix generation formula; The target embedding matrix generation formula includes: ; Among them, represents the target embedding matrix; represents the target matrix; represents the vector matrix; represents the position encoding matrix.

6. The method according to claim 2, wherein The performing position encoding on the network layer based on the position serial number includes: Perform sine position encoding on the network layer based on the position serial number.

7. The method according to any one of claims 1 to 6, characterized in that The Transformer model includes a preset number of encoder layers that process the target embedding matrix; a first LayerNorm layer connected to the encoder layers; a first fully connected layer connected to the first LayerNorm layer; a softmax layer connected to the fully connected layer; a discrete mapping layer connected to the softmax layer; wherein, the encoder layer is used to calculate the attention of each row in the target embedding matrix; the discrete mapping layer is used to output the quantization bit number.

8. The method according to claim 7, wherein Each encoder layer includes a second LayerNorm layer connected to the input layer of the encoder layer; a multi-head attention mechanism layer connected to the second LayerNorm layer; a first residual layer connected to the multi-head attention mechanism layer and the input layer; a third LayerNorm layer connected to the residual layer; a feed-forward neural network layer connected to the third LayerNorm layer; a second residual layer connected to the feed-forward neural network layer and the first residual layer.

9. The method according to claim 8, wherein The feed-forward neural network layer includes: a second fully connected layer connected to the third LayerNorm layer; a ReLU activation layer connected to the second fully connected layer; a third fully connected layer connected to the ReLU activation layer.

10. The method according to claim 9, wherein Quantizing the target neural network model based on the quantization bit number to obtain a target quantized neural network model includes: Statistical weight distribution of the network layer to obtain a weight value distribution result; Discarding a preset number of weight values at the front and back in the weight value distribution result to obtain remaining weight values; Statistical maximum and minimum weight values in the remaining weight values; Taking the maximum value of the maximum weight value and the minimum weight value as the truncation range; Truncating the network layer based on the truncation range to obtain truncation values; Based on the layer-by-layer symmetric quantization algorithm, quantizing and dequantizing the weight values of the network layer to obtain corresponding quantization results and dequantization results; Determining the target quantized neural network model based on the quantization result and the dequantization result; wherein, the layer-by-layer symmetric quantization algorithm includes: ; ; ; Among them, represents the th quantization result of the th weight value in the th network layer; represents truncating the weight value to where represents the truncation value; represents the th dequantization result of the th weight value in the th network layer.

11. The method according to claim 10, wherein The loss function of the Transformer model includes: ; ; ; ; Among them, represents the loss function value; represents the hyperparameter with the adjusted initial value around 1; represents the test set images; represents the floating-point model accuracy; represents the quantized model accuracy; represents the logarithmic function; represents the adjusted model and the weights of the loss terms.

12. A neural network model quantization system, characterized in that, Including: A first acquisition module for acquiring the weight values, hyperparameters, and position serial numbers of each network layer in the target neural network model to be quantized; A first generation module for linearly embedding the weight values, the hyperparameters, and the position serial numbers to generate a target embedding matrix; A first processing module for processing the target embedding matrix based on a pre-trained Transformer model to obtain the quantization bit numbers of each network layer in the target neural network model; A first quantization module for quantizing the target neural network model based on the quantization bit number to obtain a target quantized neural network model to reduce the model size and memory occupancy; deploying the target quantized neural network model to a device; Among them, the target neural network model can be used to process pictures; and the Transformer model includes a model trained by test set pictures; Among them, the operation formula of the quantization bit number includes: ; ; in, Indicates The number of quantization bits of the network layer, , Indicates the total number of layers of the network; Indicates the rounding algorithm; Indicates the minimum value of quantization; Indicates the maximum value of quantization; Indicates the total number of encoder layers connected in sequence; Indicates The processing result of the encoder layer; , Indicates the preset value.

13. An electronic device, characterized in that Including: A memory for storing a computer program; A processor for implementing the steps of the neural network model quantization method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the neural network model quantization method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Antibacterial peptide prediction method and device based on protein pre-training representation learning

    CN112614538A

  • Automatic quantification method for target detection model Faster R-CNN

    CN113627593A