Quantization operation method of neural network model, electronic device and program product

By determining the common quantization coefficient in the neural network model and reducing the number of quantization and inverse quantization processes, the problem of decreasing model accuracy and operating speed in the prior art is solved, and a more efficient quantization operation effect is achieved.

CN114819147BActive Publication Date: 2025-06-24ARM TECH CHINA CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210519436.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-06-24
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

When performing quantization operations on neural network models, the prior art requires frequent quantization and inverse quantization processes, resulting in a decrease in model accuracy and operating speed. Especially in complex models such as time series models, there are more quantitative inverse quantization steps and more obvious impact.

Method used

By determining the two operators that need to perform correlation operations, obtain the common quantization coefficient based on the quantization coefficients corresponding to their output data, reduce the number of inverse quantization and quantization processes, and perform correlation operations directly to obtain the quantization operation results.

Benefits of technology

It effectively improves the running speed and accuracy of the model, reduces the number of quantitative inverse quantization processes, especially in complex models, and improves the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114819147B_ABST
    Figure CN114819147B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a quantization operation method, an electronic device, and a program product for a neural network model. The quantization operation method includes: determining a first operator and a second operator that need to perform an association operation, where the first operator has first output data and the second operator has second output data; obtaining a common quantization coefficient based on a first quantization coefficient corresponding to the first output data and a second quantization coefficient corresponding to the second output data; obtaining first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient, and obtaining second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient; performing an association operation on the first quantization data and the second quantization data corresponding to the first operator and the second operator to obtain a quantization operation result of the association operation of the first operator and the second operator. Based on the above solution, the model operation speed and accuracy can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a quantization operation method, an electronic device, and a program product for a neural network model. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, neural network models (e.g., Long Short-Term Memory (LSTM) models) have achieved very good results in the fields of computer vision, speech, natural language, reinforcement learning, etc. in recent years. After a neural network model is trained, it has hundreds or even tens of millions of parameters, such as weight parameters and bias parameters in each network layer, and these parameters are stored based on 32-bit floating-point or higher-bit storage when stored. Due to the large amount of data of these parameters, a large amount of storage space and computing resources will be consumed during the entire convolution calculation process. Therefore, a quantization operation method is needed to compress the memory occupied by the neural network model, so as to facilitate the deployment of the neural network model to terminals with limited computing resources such as mobile phones.

[0003] Currently, for the quantization of neural network models, generally a quantization parameter is configured for each data in the model. For example, a quantization parameter is configured for the output data of each operator, so that the operator can directly output the operation result with a fixed-point data type (i.e., the quantization operation result). However, the quantization parameters for the output data of each operator in the model are mostly inconsistent. Therefore, when performing corresponding operations on the output data of operators with associated operations, generally, the output data needs to be dequantized to obtain floating-point data, then the floating-point data is subjected to the corresponding operation, and the operation result is quantized again to obtain the corresponding quantization operation result. This method involves multiple quantization and dequantization processes. When all operations in the neural network model are quantized in this way, a large number of quantization and dequantization operations will be involved, so the model accuracy and running speed will both decrease. And for more complex models, generally more quantization parameters need to be set, that is, more quantization and dequantization steps are introduced. For example, for a time series model, such as an LSTM model, since the state of the data is related to the time t input to the model, when the time series model is quantized, it requires t times the quantization parameters compared to a general neural network model. Therefore, when performing quantization operations on complex models, such as time series models, there will be more quantization and dequantization steps. In this way, the model accuracy and running speed will be greatly reduced. Summary of the Invention

[0004] To solve the above problems, the embodiments of the present application provide a quantization operation method, an electronic device, and a program product for a neural network model.

[0005] In a first aspect, an embodiment of the present application provides a quantization operation method for a neural network model, which is used in an electronic device and includes:

[0006] Determine a first operator and a second operator that need to perform an association operation, where the first operator has first output data and the second operator has second output data;

[0007] Obtain a common quantization coefficient based on a first quantization coefficient corresponding to the first output data and a second quantization coefficient corresponding to the second output data;

[0008] Obtain first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient, and obtain second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient;

[0009] Perform an association operation on the first quantization data and the second quantization data corresponding to the first operator and the second operator to obtain a quantization operation result of the association operation between the first operator and the second operator.

[0010] It can be understood that based on the above solution, when the common quantization coefficient is one of the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data, only the output data corresponding to the other quantization coefficient needs to be dequantized to obtain floating-point data, and the floating-point data is quantized through the common quantization coefficient to obtain quantization data, which can then be directly operated with the other output data that has not been dequantized to obtain the corresponding quantization operation result. Compared with the prior art method of dequantizing both output data to obtain the corresponding floating-point data, adding the floating-point data to obtain the floating-point operation result, and then quantizing the floating-point operation result to obtain the quantization operation result, the dequantization and quantization processes of one of the data and the quantization process of the final operation result are reduced. In this way, the model running speed can be effectively improved, and the model accuracy can be improved.

[0011] In a possible implementation of the above first aspect, obtaining a common quantization coefficient based on a first quantization coefficient corresponding to the first output data and a second quantization coefficient corresponding to the second output data includes:

[0012] When it is determined that the difference between the first quantization coefficient corresponding to the first output data and the second quantization coefficient corresponding to the second output data is greater than a set difference, the largest quantization coefficient among the first quantization coefficient and the second quantization coefficient is used as the common quantization coefficient.

[0013] It can be understood that when the floating-point data ranges are the same, the larger the quantization coefficient, the larger the range corresponding to the quantized data. When the range corresponding to the quantized data is larger, it means that the number of discrete values that the continuous values (or a large number of possible discrete values) of the signal can be approximated to is more, that is, the higher the quantization accuracy. Therefore, using the largest quantization coefficient among the quantization coefficients corresponding to the first output data and the quantization coefficient corresponding to the second output data as the common quantization coefficient can further improve the model accuracy.

[0014] In a possible implementation of the first aspect above, obtaining a common quantization coefficient based on the first quantization coefficient corresponding to the first output data and the second quantization coefficient corresponding to the second output data includes:

[0015] When it is determined that the difference between the first quantization coefficient corresponding to the first output data and the second quantization coefficient corresponding to the second output data is less than or equal to a set difference, obtain a common quantization coefficient based on the first quantization coefficient, the second quantization coefficient, and a preset rule.

[0016] In a possible implementation of the first aspect above, obtaining a common quantization coefficient based on the first quantization coefficient, the second quantization coefficient, and a preset rule includes:

[0017] Equally spaced divide the numerical range between the first quantization coefficient and the second quantization coefficient to obtain a set number of dividing points;

[0018] Using each dividing point as a preset common quantization coefficient, obtain the quantization operation results of the correlation operations of the first operator and the second operator corresponding to each preset common quantization coefficient;

[0019] Obtain the floating-point operation result of the correlation operation of the first operator and the second operator;

[0020] Using the preset common quantization coefficient corresponding to the quantization operation result with the smallest accuracy loss relative to the floating-point operation result among each quantization operation result as the common quantization coefficient.

[0021] It can be understood that when the difference between the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data is small, multiple preset common quantization coefficients can be obtained from the numerical range between the first quantization coefficient and the second quantization coefficient, and the preset common quantization coefficient corresponding to the quantization operation result with the smallest accuracy loss finally obtained is used as the common quantization coefficient. In this way, the operation accuracy of the model can be effectively guaranteed.

[0022] In a possible implementation of the first aspect above, using the preset common quantization coefficient corresponding to the quantization operation result with the smallest accuracy loss relative to the floating-point operation result among each quantization operation result as the common quantization coefficient includes:

[0023] Among the results of various quantization operations, the preset common quantization coefficient corresponding to the quantization operation result with the highest similarity to the floating-point operation result is used as the common quantization coefficient.

[0024] It can be understood that the magnitude of the precision loss corresponding to the above quantization operation results can be judged based on the root mean square error, similarity, signal-to-noise ratio, KL divergence (Kullback–Leibler divergence, abbreviated as KLD) dimension, etc. between the quantization operation results and the floating-point operation results. For example, the greater the similarity, the smaller the root mean square error, the smaller the signal-to-noise ratio, and the smaller the KL divergence, the smaller the quantization precision loss.

[0025] Therefore, among the results of various quantization operations, the preset common quantization coefficient corresponding to the quantization operation result with the smallest precision loss relative to the floating-point operation result is used as the common quantization coefficient; it may also include:

[0026] The preset common quantization coefficient corresponding to the quantization operation result with the smallest root mean square error, the smallest signal-to-noise ratio, or the smallest KL divergence among the results of various quantization operations relative to the floating-point operation result is used as the common quantization coefficient.

[0027] In a possible implementation of the above first aspect, when the common quantization coefficient is the first quantization coefficient, the first quantization data is the first output data, and

[0028] Obtaining the second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient; includes:

[0029] Obtaining the floating-point data corresponding to the second output data according to the second output data and the second quantization data, and obtaining the quantization conversion data of the second output data according to the floating-point data corresponding to the second output data and the common quantization coefficient;

[0030] Taking the quantization conversion data of the second output data as the second quantization data.

[0031] In a possible implementation of the above first aspect, when the common quantization coefficient is the second quantization coefficient, the second quantization data is the second output data, and

[0032] Obtaining the first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient, including:

[0033] Obtaining the floating-point data corresponding to the first output data according to the first output data and the first quantization data, and obtaining the quantization conversion data of the first output data according to the floating-point data corresponding to the first output data and the common quantization coefficient;

[0034] Taking the quantization conversion data of the first output data as the first quantization data.

[0035] In a possible implementation of the foregoing first aspect, when the common quantization coefficient is a value other than the first quantization coefficient and the second quantization coefficient,

[0036] Obtaining first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient, and obtaining second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient, includes:

[0037] Obtaining floating-point data corresponding to the first output data according to the first output data and the first quantization data, and obtaining quantization conversion data of the first output data according to the floating-point data corresponding to the first output data and the common quantization coefficient;

[0038] Obtaining floating-point data corresponding to the second output data according to the second output data and the second quantization data, and obtaining quantization conversion data of the second output data according to the floating-point data corresponding to the second output data and the common quantization coefficient;

[0039] Taking the quantization conversion data of the second output data as the second quantization data.

[0040] In a second aspect, an embodiment of the present application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the foregoing quantization operation method.

[0041] In a third aspect, an embodiment of the present application provides a computer program product, including instructions for implementing the foregoing quantization operation method.

[0042] In a fourth aspect, an embodiment of the present application provides a storage medium, on which instructions are stored, and when the instructions are executed on a computer, the computer executes the foregoing quantization operation method.

[0043] Based on the foregoing solutions, the present application has the following beneficial effects:

[0044] In the foregoing solution provided by the embodiment of the present application, when the common quantization coefficient is one of the quantization coefficients corresponding to the first output data and the second output data, then only the output data corresponding to the other quantization coefficient can be dequantized to obtain floating-point data, and the floating-point data is quantized through the common quantization coefficient to obtain quantization data, which can be directly operated with the other output data that has not been dequantized to obtain the corresponding quantization operation result. Compared with the prior art in which both output data need to be dequantized to obtain the corresponding floating-point data, and the floating-point operation result obtained after adding the floating-point data is quantized again to obtain the quantization operation result, the dequantization and quantization processes of one of the data and the quantization process of the final operation result are reduced. In this way, the model operation speed can be effectively improved, and the model accuracy can be improved.

[0045] Moreover, by using the largest quantization coefficient among the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data as the common quantization coefficient, the model accuracy can be further improved. Secondly, when the difference between the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data is small, multiple preset common quantization coefficients can be obtained from the numerical range between the first quantization coefficient and the second quantization coefficient, and the preset common quantization coefficient corresponding to the quantization operation result with the smallest accuracy loss finally obtained is used as the common quantization coefficient. In this way, the operation accuracy of the model can be effectively guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 The figure shows a schematic structural diagram of an LSTM model provided by an embodiment of the present application;

[0047] Figure 2 The figure shows a schematic structural diagram of a single neural network layer of an LSTM model provided by an embodiment of the present application;

[0048] Figure 3 The figure shows a schematic diagram of the operation steps of a part of the operations in an LSTM model provided by an embodiment of the present application;

[0049] Figure 4 The figure shows a schematic diagram of the quantization operation steps of a part of the operations in an LSTM model provided by an embodiment of the present application;

[0050] Figure 5 The figure shows a schematic diagram of the quantization operation steps of a part of the operations in an LSTM model provided by an embodiment of the present application;

[0051] Figure 6 The figure shows a schematic diagram of the quantization operation steps of a part of the operations in an LSTM model provided by an embodiment of the present application;

[0052] Figure 7a The figure shows a schematic hardware structure diagram of an electronic device applying a neural network model provided by an embodiment of the present application;

[0053] Figure 7b The figure shows a schematic software structure diagram of an electronic device applying a neural network model provided by an embodiment of the present application;

[0054] Figure 8 The figure shows a schematic flowchart of a quantization operation method of a neural network model provided by an embodiment of the present application;

[0055] Figure 9 The figure shows a schematic diagram of a quantization operation device of a neural network model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] Exemplary embodiments of the present application include, but are not limited to, a quantization operation method for a neural network model, an electronic device, and a program product. The embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0057] In the following description, many technical details are provided to help the reader better understand the present invention. However, those of ordinary skill in the art can understand that the technical solutions claimed in the claims of the present invention can be implemented even without these technical details and various changes and modifications based on the following embodiments.

[0058] To better understand the solutions in the embodiments of the present application, the following briefly describes the terms related to the present application:

[0059] A recurrent neural network (RNN) is a multi-layer neural network that includes an input layer, a hidden layer, and an output layer. Among them, there are generally multiple hidden layers. The output of each hidden layer is related not only to the input from the input layer to this hidden layer but also to the output of the previous hidden layer of this hidden layer, that is, the hidden layers are associated with each other.

[0060] The LSTM model is a special RNN model. In the LSTM model, the neural network layers in the hidden layer can transfer the state and output of the neural network layer. That is, the output of each neural network layer in the hidden layer is related to the input from the input layer, the output of the previous neural network layer, and the state of the previous neural network layer. The neural network layers in the hidden layer can filter the input data through a gate structure, that is, forget the unimportant and select the important ones. Correspondingly, the gate structure includes a forget gate, a select gate, and an output gate. Therefore, compared with ordinary RNNs, LSTMs can perform better in longer sequences.

[0061] Quantization refers to the process of approximating the continuous values (or a large number of possible discrete values) of a signal to a finite number (or fewer) of discrete values. Specifically, it can be the process of converting floating-point data into fixed-point data.

[0062] Among them, the floating-point to fixed-point range conversion coefficient (i.e., the quantization coefficient) can be calculated through the following formula:

[0063] S = (Q_max - Q_min) / (F_max - F_min) Formula (1)

[0064] Among them, F_max represents the maximum value of a floating-point number. F_min represents the minimum value of a floating-point number. Q_max represents the maximum value of the quantization data range corresponding to the corresponding quantization bit number. Q_min represents the minimum value of the quantization data range corresponding to the corresponding quantization bit number. S represents the range conversion coefficient (i.e., quantization coefficient) of floating-point to fixed-point.

[0065] The conversion of floating-point data to fixed-point data (quantized data) is achieved through the following formula:

[0066] Q = F * S Formula (2)

[0067] Among them, Q represents the quantized data. S represents the range conversion coefficient (i.e., quantization coefficient) of floating-point quantization, and F represents the fixed-point data.

[0068] Dequantization refers to the process of converting fixed-point data into floating-point data through the quantization coefficient.

[0069] Before introducing the quantization method in the embodiments of the present application, the LSTM model in the embodiments of the present application will be further introduced below with reference to the accompanying drawings.

[0070] As Figure 1 shown, it is a schematic structural diagram of an LSTM model in an embodiment of the present application. The LSTM model 100 includes an input layer 110, hidden layers 120-1 to 120-t (hereinafter referred to as neural network layers 120-1 to 120-t), and an output layer 130. Among them, the data input by the input layer 110 includes x1, x2, x3,..., xt, etc., and the data output by the output layer 130 includes h1, h2, h3,..., ht, etc. The state c and output h of the neural network layer 120 can be transmitted between the neural network layers 120, such as state c1, c2 and output h1, h2, etc.

[0071] It can be understood that the LSTM model can be applied to fields such as speech recognition, translation, and can also be applied to various fields of time series forecasting.

[0072] For example, when the LSTM model is used for speech recognition, as Figure 1 shown, x1, x2, x3,..., xt can be the speech data to be recognized. Among them, x1 is the speech input data input to the hidden layer 120-1 at the first time step, x2 is the speech input data input to the hidden layer 120-2 at the second time step, and xt is the speech input data input to the hidden layer 120-t at the t-th time. The data output by the output layer 130 includes h1, h2, h3,..., h t can be the speech recognition results corresponding to different time steps respectively.

[0073] Among them, during the operation of the neural network layer of the LSTM model 200, such as the neural network layer 120-t, the input data x of the input layer 110 can be used t , the data h output by the output layer 130 at the previous time step t-1 , the state c of the neural network layer 120-(t-1) at the previous time step t-1 , the weight coefficient matrix, the error coefficient matrix, etc., to determine the data h output by the output layer 130 at the current time step t . Figure 2 FIG. shows a schematic structural diagram of a neural network layer in an embodiment of the present application. Specifically, the operation of the neural network layer can be divided into the following three stages. Taking the neural network layer 120-t as an example, the description is as follows:

[0074] 1. Forget stage

[0075] The forget stage can selectively forget the state output by the previous neural network layer 120-(t-1), that is, forget the unimportant ones. Specifically, assuming that the current neural network layer is the neural network layer of the t-th node, the forget stage can be based on the output h of the previous neural network layer 120-(t-1) t-1 and the input x of the current neural network layer t , calculate the forget gate f t , and selectively forget the state c of the previous neural network layer 120-(t-1) through the forget gate f t . The forget gate f t-1 can be understood as t the output of the σ layer 1211 in Figure 2 .

[0076] Exemplarily, the forget gate f t can be calculated by the following formula:

[0077] f t =σ(W xf x t +W hf h t-1 +b f ) Formula (3)

[0078] where σ represents the sigmoid function, that is, the calculation result of W xf x t +W hf h t-1 +b f is converted to between (0,1) through the sigmoid function. W xf represents the weight coefficient of the input x of the neural network layer 120-t t , W hfDenote the output \(h\) of the previous neural network layer \(120-(t - 1)\) t-1 as the weight coefficient, \(b\) f is the bias term.

[0079] Exemplarily, through the forget gate \(f\) t selectively forgets the state \(c\) of the previous neural network layer \(120-(t - 1)\) t-1 It can be understood that multiplying the forget gate \(f\) t by the state \(c\) of the previous neural network layer \(120-(t - 1)\) t-1 i.e., \(f\) t * \(c\) t-1 . \(f\) t * \(c\) t-1 can be understood as Figure 2 the dot product operation 1215 in

[0080] 2. Selection and Memory Stage

[0081] The selection and memory stage can selectively remember the input of the current neural network layer \(120 - t\), that is, select the important ones. Specifically, assume that the current neural network layer is the neural network layer of the \(t\)-th node. The selection and memory stage can calculate the selection gate \(i\) t-1 and the state candidate value \(g\) t based on the output \(h\) of the previous neural network layer \(120-(t - 1)\) t and the input \(x\) of the current neural network layer t . Then use the selection gate \(i\) t to achieve selective memory of the state candidate value \(g\) t . The selection gate \(i\) t can be understood as Figure 2 the output of the \(\sigma\) layer 1212 in t and the state candidate value \(g\) Figure 2 can be understood as the output of the \(\psi\) layer 1213 in

[0082] Exemplarily, the selection gate \(i\) t and the state candidate value \(g\) t can be calculated by the following formula:

[0083]

[0084] where \(\sigma\) represents the sigmoid function, that is, the calculation result of \(W\) xi \(x\) t +\(W\) hi \(h\) t-1 +\(b\) i is transformed to the range \((0, 1)\) through the sigmoid function. \(W\) xi represents the input \(x\) of the neural network layer \(120 - t\) tThe weight coefficient, W hi represents the output h of the previous neural network layer 120-t t-1 The weight coefficient, b i is the bias term.

[0085] where ψ represents the tanh function, that is, for W xg x t +W hg h t-1 +b i The calculation result of is transformed to between (-1, 1) through the tanh function. W xg represents the input x of the neural network layer 120-t t The weight coefficient, W hg represents the output h of the previous neural network layer 120-(t-1) t-1 The weight coefficient, b g is the bias term.

[0086] Exemplarily, using the selective gate i t to achieve selective memory of the state candidate value g t can be understood as multiplying the selective gate i t with the state candidate value g t i.e., i t *g t . i t *g t can be understood as Figure 2 the dot product operation 1214 in

[0087] Exemplarily, f obtained in the forgetting stage can be t *c t-1 and i obtained in the selective memory stage t *g t are added to obtain the state c of the current neural network layer 120-t t , that is, the state c of the current neural network layer 120-t t can be calculated by the following formula:

[0088] c t = f t *c t-1 + i t *g t Formula (5)

[0089] 3. Output stage

[0090] The output stage can determine the output and state of the previous neural network layer 120-t and perform the output. Specifically, assuming that the current neural network layer is the neural network layer of the t-th node, the output stage can be based on the output h of the previous neural network layer 120-(t-1)t-1 and the input x of the current neural network layer t , the output gate o is calculated t . And the state c of the current neural network layer 120 - t t is scaled, for example, it can be achieved through the tanh function. Then, the output gate o t is used to control the scaled state c of the current neural network layer 120 - t t . The output gate o t can be understood as Figure 2 the output of the σ layer 1217 in t and the scaled state c of the current neural network layer 120 - t Figure 2 can be understood as the output of the ψ layer 1218 in

[0091] Exemplarily, the output gate o t and the output h t can be calculated by the following formula:

[0092]

[0093] where σ represents the sigmoid function, that is, the calculation result of W xo x t +W ho h t-1 +b o W xi x t +W hi h t-1 +b i is transformed to the range (0, 1) through the sigmoid function. W xo represents the weight coefficient of the input x of the neural network layer 120 - t t , and W ho represents the weight coefficient of the output h of the previous neural network layer 120 - t t-1 , and b o is the bias. Where ψ represents the tanh function, that is, the calculation result of c t is transformed to the range (-1, 1) through the tanh function.

[0094] It can be understood that the LSTM model includes multiple neural network layers, and each neural network layer needs to perform the above - mentioned complex calculations, and in each step, multiple quantization and de - quantization processes are required. For example, as Figure 3 shown, when performing the quantization operation of W xf X t +W xh h t-1 in the above formula (1), the operation steps required include: for Wxf , X t , W xh , h t-1 Quantize them respectively to obtain the corresponding quantization data W1, X1, W2, h1; then calculate the product Y1 of W1 and X1 and the product Y2 of W2 and h1;

[0095] Since the quantization coefficient corresponding to Y1 is the product of the quantization coefficients of W1 and X1, and the quantization coefficient corresponding to Y2 is the product of the quantization coefficients of W2 and h1; therefore, the quantization coefficients corresponding to Y1 and Y2 are generally different, that is, the quantization data ranges corresponding to Y1 and Y2 are different. And numbers with different quantization ranges are difficult to perform direct operations. Therefore, as Figure 4 shown, the operation steps required for the above quantization operation also include dequantizing Y1 through the corresponding first quantization coefficient to obtain the floating-point data Y11 corresponding to Y1, and dequantizing Y2 through the corresponding second quantization coefficient to obtain the floating-point data Y21 corresponding to Y2. And add the floating-point data Y11 corresponding to Y1 and the floating-point data Y21 corresponding to Y2 to obtain W xf X t +W xh h t-1 The floating-point operation result Y3 of, and then quantize Y3 through the corresponding third quantization coefficient to obtain the quantization data Y31, that is, the fixed-point operation result of W xf X t +W xh h t-1 .

[0096] As described above, any associated operation requires a relatively large number of quantization steps to obtain the corresponding quantization result. Thus, the model running speed is slowed down. Secondly, since a certain precision loss will be caused during each quantization process, and the above quantization operation method includes a large number of quantization processes, it will cause a large precision loss of the model.

[0097] To solve the above problems, an embodiment of the present application provides a quantization operation method. The method includes: determining a first operator and a second operator having an operation association, where the first operator has a first output data and the second operator has a second output data. For example, when performing the quantization operation in the above formula (3) of W xf X t +W xh h t-1 This quantization operation, the first operator is the operator that quantizes W xf and X t respectively to obtain the corresponding quantization data W1 and X1, and multiplies W1 and X1. Then the first output data can be the product of W1 and X1, for example, Y1. The second operator is the operator that quantizes W xh and h t-1An operator that quantifies respectively to obtain corresponding quantization data W2 and h1, and multiplies W2 and h1. Then the second output data can be the product of W2 and h1, for example, Y2.

[0098] Then, based on the quantization coefficients corresponding to the first output data and the second output data respectively, a common quantization coefficient is obtained. Based on the common quantization coefficient, the first quantization data corresponding to the first output data and the second quantization data corresponding to the second output data are obtained. The first quantization data and the second quantization data are subjected to corresponding correlation operations to obtain an operation result. It can be understood that this operation result is the quantization operation result of the corresponding operations of the first operator and the second operator.

[0099] Among them, the way to obtain the common quantization coefficient can include: when it is determined that the difference between the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data is greater than a set value, then the largest quantization coefficient among the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data is used as the common quantization coefficient.

[0100] When it is determined that the difference between the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data is less than the set value, then based on the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data, the corresponding quantization coefficient is calculated according to the set rule.

[0101] It can be obtained from the foregoing formula (1) that when the floating-point data ranges are the same, the larger the quantization coefficient, the larger the range corresponding to the quantization data. When the range corresponding to the quantization data is larger, it means that the number of discrete values that the continuous values (or a large number of possible discrete values) of the signal can be approximated to is more, that is, the quantization accuracy is higher.

[0102] For example, assume that the quantization coefficient corresponding to the first output data is A1, and the quantization coefficient corresponding to the second output data is A2, and A1 is greater than A2. If the difference between A1 and A2 is greater than the set difference, then A1 is used as the common quantization coefficient. Assume that the difference between A1 and A2 is less than or equal to the set value, then based on A1 and A2, the corresponding quantization coefficient is calculated according to the set rule.

[0103] Among them, the above set rule can be to equally divide the numerical range between the first quantization coefficient and the second quantization coefficient to obtain a set number of dividing points. Each dividing point is used as a preset common quantization coefficient respectively, the quantization operation results of the corresponding operations of the first operator and the second operator are obtained, and the accuracy loss of the corresponding quantization operation result is calculated. The preset common quantization coefficient corresponding to the quantization operation result with the smallest accuracy loss is used as the finally determined common quantization coefficient.

[0104] It can be understood that the magnitude of the precision loss corresponding to the above quantization operation result can be judged according to the root mean square error, similarity, signal-to-noise ratio, KL divergence (Kullback–Leibler divergence, abbreviated as KLD) dimension, etc. between the quantization operation result and the floating-point operation result. For example, the greater the similarity, the smaller the root mean square error, the smaller the signal-to-noise ratio, and the smaller the KL divergence, the smaller the quantization precision loss.

[0105] It can be understood that both the above first output data and the second data can be fixed-point data, that is, quantized data.

[0106] It can be understood that based on the above solution, when the common quantization coefficient is one of the quantization coefficients corresponding to the first output data and the quantization coefficient corresponding to the second output data, only the output data corresponding to the other quantization coefficient needs to be dequantized to obtain floating-point data, and the floating-point data is quantized through the common quantization coefficient to obtain quantized data, which can then be directly operated with the other output data that has not been dequantized to obtain the corresponding quantization operation result. Compared with the prior art where both output data need to be dequantized to obtain the corresponding floating-point data, and the floating-point operation result obtained by adding the floating-point data is quantized again to obtain the quantization operation result, the dequantization and quantization processes of one of the data and the quantization process of the final operation result are reduced. In this way, the model operation speed can be effectively improved, and the model precision can be improved.

[0107] Moreover, using the largest quantization coefficient among the quantization coefficients corresponding to the first output data and the quantization coefficient corresponding to the second output data as the common quantization coefficient can further improve the model precision.

[0108] Secondly, when the difference between the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data is small, multiple preset common quantization coefficients can be obtained from the numerical range between the first quantization coefficient and the second quantization coefficient, and the preset common quantization coefficient corresponding to the quantization operation result with the smallest precision loss finally obtained is used as the common quantization coefficient. In this way, the operation precision of the model can be effectively guaranteed.

[0109] Next, taking the W in the above formula (3) xf X t +W xh h t-1 This quantization operation is used to illustrate the model quantization method in the embodiments of the present application.

[0110] It can be understood that when performing the quantization operation of W in the above formula (3) xf X t +W xh h t-1 This quantization operation, for example Figure 3As shown, the first operator mentioned above is to quantize W xf and X t respectively to obtain the corresponding quantized data W1 and X1, and multiply W1 and X1. The first output data can be the product of W1 and X1, for example, Y1. The second operator mentioned above is to quantize W xh and h t-1 respectively to obtain the corresponding quantized data W2 and h1, and multiply W2 and h1. Then the second output data can be the product of W2 and h1, for example, Y2.

[0111] First, the quantization coefficients corresponding to Y1 and Y2 can be obtained respectively. Among them, the quantization coefficient corresponding to Y1 is the product of the quantization coefficients corresponding to W xf and X t respectively, assumed to be 500; the quantization coefficient corresponding to Y2 is the product of the quantization coefficients corresponding to W xh and h t-1 respectively, assumed to be 10; if the difference is set to 9 times the minimum quantization coefficient, that is, 90. It can be judged that the difference between the quantization coefficient corresponding to Y1 and the quantization coefficient corresponding to Y2 is greater than the set difference. At this time, the quantization coefficient 500 corresponding to Y1 can be used as the common quantization coefficient.

[0112] As Figure 5 shown, at this time, Y2 can be dequantized, that is, multiply Y2 by the initial quantization coefficient 10 to obtain the floating-point data Y21 corresponding to Y2, and then quantize the floating-point data Y21 corresponding to Y2 with the common quantization coefficient 500 to obtain the quantized data Y22 corresponding to W xh h t-1 . Then directly add Y1 and Y22 to obtain the quantized operation result Y31 of W xf X t +W xh h t-1 .

[0113] Assume that the quantization coefficient corresponding to Y1 is 500 and the quantization coefficient corresponding to Y2 is 300. If the difference is set to 9 times the minimum quantization coefficient, that is, 2700. It can be judged that the difference between the quantization coefficient corresponding to Y1 and the quantization coefficient corresponding to Y2 is less than the set difference. At this time, the common quantization coefficient can be calculated according to the set rules.

[0114] For example, the way to calculate the common quantization coefficient can be to divide the numerical range between the quantization coefficient 500 corresponding to Y1 and the quantization coefficient 300 corresponding to Y2 at an interval of 20 to obtain 11 separation points, namely 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, 500. These 11 separation points are respectively used as the preset common quantization coefficients to obtain the corresponding W xf X t +W xh h t-1 The quantization operation results of, and calculate the similarity between the corresponding quantization operation results and the floating-point operation results of W xf X t +W xh h t-1 respectively. The preset common quantization coefficient corresponding to the quantization operation result with the maximum similarity is used as the finally determined common quantization coefficient, for example, S0.

[0115] At this time, as Figure 6 shown, Y1 and Y2 can be dequantized first, that is, Y1 is multiplied by the initial quantization coefficient 500 to obtain the floating-point data Y11 corresponding to Y1, and Y2 is multiplied by the initial quantization coefficient 300 to obtain the floating-point data Y21 corresponding to Y2. Then Y1 is quantized with the common quantization coefficient S0 to obtain the quantization data Y12 corresponding to W xf X t and Y2 is quantized with the common quantization coefficient S0 to obtain the quantization data Y22 corresponding to W xh h t-1 . Then Y12 and Y22 are directly added to obtain the quantization operation result of W xf X t +W xh h t-1 .

[0116] It can be understood that the above schemes are all based on performing the quantization operation of W in the above formula (3) xf X t +W xh h t-1 as an example to illustrate the quantization operation method in the embodiments of the present application. The quantization operation method in the embodiments of the present application can be used in any operator with associated operations.

[0117] For example, in the above formula (3) in the LSTM model, when performing the operation of W xf X t +W xh h t-1 +b f the first operator can refer to performing the execution of W xf X tThe operator of the quantization operation, and the second operator can refer to W xh h t-1 +b f The operator of the quantization operation.

[0118] In the above formula (4) of the LSTM model, when performing the operation of W xi X t +W hi h t-1 +b i In the operation, the first operator can refer to the operator that performs the quantization operation of W xi X t The operator of the quantization operation, and the second operator can refer to the operator that performs W hi h t-1 +b i The operator of the quantization operation.

[0119] In the above formula (5) of the LSTM model, when performing the operation of f t c t-1 +i t g t In the operation, the first operator can refer to the operator that performs the quantization operation of f t c t-1 The operator of the quantization operation, and the second operator can refer to the operator that performs i t g t The operator of the quantization operation.

[0120] It can be understood that when any two operators in the neural network model perform an associated operation, the above operation method can be adopted. The neural network model mentioned in the technical solution of this application can be a recurrent neural network model, for example, a GRU model, an LSTM model, a Bi-RNN model, etc. It can also be a convolutional neural network model, a deep learning model, etc. It can be understood that according to the actual application, this application does not make specific limitations on the neural network models applicable to model quantization.

[0121] It can be understood that the quantization method of the neural network model provided by this application can be implemented on various electronic devices, including but not limited to servers, a distributed server cluster composed of multiple servers, mobile phones, tablet computers, face recognition access control, laptop computers, desktop computers, wearable devices, head-mounted displays, mobile email devices, portable game consoles, portable music players, reader devices, personal digital assistants, virtual reality or augmented reality devices, televisions and other electronic devices in which one or more processors are embedded or coupled.

[0122] Next, the electronic device provided by the embodiments of this application will be briefly introduced by taking a mobile phone as an example. As Figure 7aAs shown, the hardware structure of the mobile phone 20 may include a processor 220, a power module 240, a memory 280, a mobile communication module 230, a wireless communication module 220, a sensor module 290, an audio module 250, a camera 270, an interface module 260, a button 202, and a display screen 202, etc.

[0123] It can be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the mobile phone 20. In other embodiments of the present application, the mobile phone 20 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0124] The processor 220 may include one or more processing units. For example, it may include a central processing unit CPU (Central Processing Unit), a graphics processing unit GPU (Graphics Processing Unit), a digital signal processor DSP, a microprocessor MCU (Micro-programmed Control Unit), an AI (Artificial Intelligence) processor, or a processing module or processing circuit such as a field programmable gate array FPGA (Field Programmable Gate Array). Among them, different processing units may be independent devices or integrated in one or more processors. A storage unit may be provided in the processor 220 for storing instructions and data. In some embodiments, the storage unit in the processor 220 is a cache memory 280. Among them, the processor can be used to execute the quantization operation method in the embodiments of the present application.

[0125] The power module 240 may include a power source, a power management component, etc. The power source may be a battery. The power management component is used to manage the charging of the power source and the power supply from the power source to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module is used to receive charging input from a charger; the power management module is used to connect the power source, the charging management module, and the processor 220. The power management module receives the input of the power source and / or the charging management module and supplies power to the processor 220, the display screen 202, the camera 270, the wireless communication module 220, etc.

[0126] The mobile communication module 230 may include, but is not limited to, an antenna, a power amplifier, a filter, an LNA (Low noise amplify), etc. The mobile communication module 230 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the mobile phone 20.

[0127] The wireless communication module 220 may include an antenna and realize the transceiver of electromagnetic waves via the antenna. The wireless communication module 220 may provide solutions for wireless communications applied to the mobile phone 20, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The mobile phone 20 may communicate with the network and other devices through wireless communication technologies.

[0128] In some embodiments, the mobile communication module 230 and the wireless communication module 220 of the mobile phone 20 may also be located in the same module.

[0129] The display screen 202 is used to display the human-machine interaction interface, images, videos, etc. The display screen 202 includes a display panel. The display panel may adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc.

[0130] The sensor module 290 may include a proximity light sensor, a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0131] The audio module 250 is used to convert digital audio information into an analog audio signal for output, or convert an analog audio input into a digital audio signal. The audio module 250 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 250 can be disposed in the processor 220, or some functional modules of the audio module 250 can be disposed in the processor 220. In some embodiments, the audio module 250 can include a speaker, a receiver, a microphone, and a headphone jack.

[0132] The camera 270 is used to capture still images or videos. An object forms an optical image through a lens and projects it onto a photosensitive element. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP (Image Signal Processing) to convert it into a digital image signal. The mobile phone 20 can implement the shooting function through the ISP, the camera 270, a video codec, a GPU (Graphic Processing Unit), the display screen 202, and an application processor, etc.

[0133] The interface module 260 includes an external memory interface, a universal serial bus (USB) interface, a subscriber identification module (SIM) card interface, etc. The external memory interface can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the mobile phone 20. The external memory card communicates with the processor 220 through the external memory interface to implement the data storage function. The universal serial bus interface is used for the mobile phone 20 to communicate with other electronic devices. The subscriber identification module card interface is used to communicate with the SIM card installed in the mobile phone 20, such as reading the phone number stored in the SIM card, or writing the phone number into the SIM card.

[0134] In some embodiments, the mobile phone 20 further includes keys 202, a motor, and an indicator, etc. Among them, the keys 202 can include volume keys, a power on / off key, etc. The motor is used to make the mobile phone 20 generate a vibration effect, for example, generating a vibration when the mobile phone 20 of the user is called to prompt the user to answer the incoming call of the mobile phone 20. The indicator can include a laser indicator, a radio frequency indicator, an LED indicator, etc.

[0135] Figure 7b The software structure block diagram of a mobile phone 20 according to an embodiment of the present application is shown. The mobile phone 20 can include an application layer, an application framework layer, an Android runtime, a system library, and a kernel layer.

[0136] The application layer can include a series of application packages.

[0137] As Figure 7b shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0138] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0139] As Figure 7b shown, the application framework layer may include window manager, content provider, view system, telephone manager, resource manager, notification manager, etc.

[0140] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0141] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include video, image, audio, dialed and answered calls, browsing history and bookmarks, phone book, etc.

[0142] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures.

[0143] The telephone manager is used to provide the communication function of the mobile phone 20. For example, the management of call status (including connection, disconnection, etc.).

[0144] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.

[0145] The notification manager enables applications to display notification information in the status bar, can be used to convey notification-type messages, and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, and can also be a notification that appears on the screen in the form of a dialogue window. For example, prompt text information in the status bar, emit a prompt tone, the electronic device vibrates, the indicator light flashes, etc.

[0146] The Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0147] The core libraries consist of two parts: one is the functional functions that the Java language needs to call, and the other is the core libraries of Android.

[0148] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0149] The system libraries can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing libraries (such as OpenGL ES), 2D graphics engines (such as SGL), etc.

[0150] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0151] The media libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0152] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0153] The 2D graphics engine is the drawing engine for 2D drawing.

[0154] The kernel layer is the layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.

[0155] The model quantization operation method in the embodiments of the present application will be briefly described below in combination with the above-mentioned electronic device. Figure 8 The flowchart of a model quantization operation method according to an embodiment of the present application is shown. The model quantization operation method provided by the embodiment of the present application can be executed by the processor 220. As Figure 8 shown, the model quantization method includes:

[0156] 401: Determine a first operator and a second operator that need to perform an association operation, where the first operator has first output data and the second operator has second output data.

[0157] It can be understood that the first operator and the second operator can refer to a node or neuron in a neural network model that performs the corresponding operation, or can be a combination of multiple nodes or multiple neurons, or can also refer to a unit that performs a certain operation.

[0158] For example, in the operation of xf W t × xh X t-1 + xf W t × xh h t-1 the first operator can refer to the operator that performs the quantization operation of

[0159] For example, in the operation of xf W t × xh X t-1 + f W xf × t h xh + t-1 b f the first operator can refer to the operator that performs the quantization operation of

[0160] For example, when performing the quantization operation of xf W t × xh X t-1 + xf W t × xh h t-1 in formula (1) above, the first operator is the operator that quantizes

[0161] W

[0162] 402: Determine whether the difference between the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data is greater than a set difference.

[0163] If so, go to 4031 and use the largest quantization coefficient among the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data as the common quantization coefficient.

[0164] If not, go to 4032, and obtain the corresponding quantization coefficient according to the set rules based on the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data.

[0165] It can be understood that the quantization coefficient corresponding to the first output data is the ratio of the floating-point output data of the first operator to the first output data. Among them, the floating-point output data of the first operator is the floating-point operation result obtained when all parameters in the first operator are not quantized. When the first operator is the product of multiple parameters, the quantization coefficient corresponding to the first output data is the product of the quantization coefficients of all parameters (including all parameters such as input data and weights) in the first operator.

[0166] The method for obtaining the quantization coefficient corresponding to the second output data is the same as that for the first output data, and will not be elaborated here.

[0167] For example, W in the above formula (1) xf X t +W xh h t-1 During this quantization operation, the quantization coefficient corresponding to the first output data Y1 is the product of the quantization coefficients corresponding to W xf and X t respectively; the quantization coefficient corresponding to the second output data Y2 is the product of the quantization coefficients corresponding to W xh and h t-1 respectively.

[0168] It can be understood that the above set difference can be 9 times the smaller quantization coefficient among the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data.

[0169] For example, the quantization coefficient corresponding to Y1 is the product A1 of the quantization coefficients corresponding to W xf and X t respectively, assumed to be 500; the quantization coefficient corresponding to Y2 is the product A2 of the quantization coefficients corresponding to W xh and h t-1 respectively, assumed to be 10; if the set difference is set to 9 times the minimum quantization coefficient, that is, 90. It can be judged that the difference between the quantization coefficient corresponding to Y1 and the quantization coefficient corresponding to Y2 is greater than the set difference. At this time, the quantization coefficient 500 corresponding to Y1 can be used as the common quantization coefficient.

[0170] It can be obtained from the foregoing formula (1) that when the floating-point data ranges are the same, the larger the quantization coefficient, the larger the range corresponding to the quantized data. When the range corresponding to the quantized data is larger, it means that the number of discrete values that the continuous values (or a large number of possible discrete values) of the signal can be approximated to is more, and to a certain extent, the quantization accuracy is higher.

[0171] Therefore, when the difference between the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data is large, for example, greater than the set difference, it can be clearly determined that the quantization operation result obtained using the quantization coefficients corresponding to the output data with the largest quantization coefficient has higher accuracy.

[0172] When the difference between the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data is small, it indicates that the quantization precisions of the first output data and the second output data are not much different. Therefore, the set rules can be used to calculate and obtain the corresponding quantization coefficients to ensure the model accuracy.

[0173] For example, the quantization coefficient corresponding to Y1 is the product A1 of the quantization coefficients corresponding to W xf and X t respectively, assumed to be 500; the quantization coefficient corresponding to Y2 is the product A2 of the quantization coefficients corresponding to W xh and h t-1 respectively, assumed to be 300; if the set difference is set to 9 times the minimum quantization coefficient, that is, 2700. It can be judged that the difference between the quantization coefficients corresponding to Y1 and the quantization coefficients corresponding to Y2 is less than the set difference. At this time, the common quantization coefficient can be calculated according to the set rules.

[0174] 4031: Use the largest quantization coefficient among the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data as the common quantization coefficient.

[0175] When the difference between the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data is large, for example, greater than the set difference, and the quantization coefficient corresponding to the first output data is large, then use the quantization coefficient corresponding to the first output data as the common quantization coefficient.

[0176] When the difference between the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data is large, for example, greater than the set difference, and the quantization coefficient corresponding to the second output data is large, then use the quantization coefficient corresponding to the second output data as the common quantization coefficient.

[0177] 4032: Obtain the corresponding quantization coefficients according to the set rules based on the quantization coefficients corresponding to the first output data and the quantization coefficients corresponding to the second output data.

[0178] Among them, the above set rules can be to equally divide the numerical range between the first quantization coefficient and the second quantization coefficient to obtain a set number of separation points, use each separation point as the preset common quantization coefficient respectively, obtain the quantization operation results of the corresponding first operator and second operator, and calculate the accuracy loss of the corresponding quantization operation result. Use the preset common quantization coefficient corresponding to the quantization operation result with the smallest accuracy loss as the finally determined common quantization coefficient.

[0179] Assume that the quantization coefficient corresponding to Y1 is 500 and the quantization coefficient corresponding to Y2 is 300. If the set difference is set to 9 times the minimum quantization coefficient, i.e., 2700. It can be determined that the difference between the quantization coefficient corresponding to Y1 and the quantization coefficient corresponding to Y2 is less than the set difference. At this time, the common quantization coefficient can be calculated according to the set rules.

[0180] For example, the way to calculate the common quantization coefficient can be to divide the numerical range between the quantization coefficient 500 corresponding to Y1 and the quantization coefficient 300 corresponding to Y2 at an interval of 20 to obtain 11 separation points, which are 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, 500 respectively. Use these 11 separation points as the preset common quantization coefficients respectively to obtain the corresponding W xf X t +W xh h t-1 of the quantization operation results, and calculate the similarity between the corresponding quantization operation results and the floating-point operation results of W xf X t +W xh h t-1 respectively. Take the preset common quantization coefficient corresponding to the quantization operation result with the largest similarity as the finally determined common quantization coefficient, for example, S0.

[0181] It can be understood that the magnitude of the accuracy loss of the above corresponding quantization operation results can be judged according to the root mean square error, similarity, signal-to-noise ratio, KL divergence (Kullback–Leibler divergence, abbreviated as KLD) dimension, etc. of the quantization operation results and the floating-point operation results. For example, the greater the similarity, the smaller the root mean square error, the smaller the signal-to-noise ratio, and the smaller the KL divergence, the smaller the quantization accuracy loss.

[0182] 404: Obtain the first quantization data corresponding to the first output data and the second quantization data corresponding to the second output data based on the common quantization coefficient.

[0183] Assume that the common quantization coefficient is the quantization coefficient corresponding to one of the output data. For example, if the quantization coefficient A1 corresponding to the first output data above is used as the common quantization coefficient, then the quantization coefficient A1 of the first output data itself is the common quantization coefficient. Therefore, no conversion is required. And the quantization coefficient corresponding to the second output data is A2, then the second output data needs to be converted to data with a quantization coefficient of A1. Among them, the specific conversion method is: perform inverse quantization on the second quantization data, for example, Y2, that is, multiply Y2 by the quantization coefficient A2 to obtain the floating-point data Y21 corresponding to Y2, and then quantize the floating-point data Y21 corresponding to Y2 with the common quantization coefficient A1 to obtain the quantization conversion data Y22 corresponding to the second output data.

[0184] Assume that the common quantization coefficient is the common quantization coefficient calculated according to the preset rule. At this time, it is necessary to perform inverse quantization on both Y1 and Y2, that is, multiply the first output data Y1 by the initial quantization coefficient A1 to obtain the floating-point data Y11 corresponding to Y1, multiply the second output data Y2 by the initial quantization coefficient A2 to obtain the floating-point data Y21 corresponding to Y2, and then quantize Y1 with the common quantization coefficient S0 to obtain W xf X t The corresponding quantization conversion data Y11 of Y1, and quantize Y2 with the common quantization coefficient S0 to obtain W xh h t-1 The corresponding quantization conversion data Y21

[0185] 405: Perform corresponding correlation operations on the first quantization data and the second quantization data to obtain the operation result

[0186] It can be understood that this operation result is the quantization operation result of the corresponding operations of the first operator and the second operator

[0187] When the quantization coefficient corresponding to the output data is the common quantization coefficient, the corresponding quantization data is the output data

[0188] When the quantization coefficient corresponding to the output data is not the common quantization coefficient, the corresponding quantization data is the quantization conversion data corresponding to the output data

[0189] For example, when the quantization coefficient corresponding to the first output data is the common quantization coefficient, the first quantization data is the first output data

[0190] When the quantization coefficient corresponding to the first output data is not the common quantization coefficient, the corresponding quantization data is the quantization conversion data corresponding to the first output data

[0191] For example, when the quantization coefficient A1 corresponding to the first output data above is used as the common quantization coefficient, the first output data Y1 above can be directly added to the quantization conversion data Y21 of the second output data to obtain W xf X t +W xh h t-1 The quantization operation result

[0192] When the common quantization coefficient above is the common quantization coefficient calculated according to the preset rule, the quantization conversion data Y11 of Y1 of the first output data can be directly added to the quantization conversion data Y21 of Y1 of the first output data to obtain W xf X t +W xh h t-1 The quantization operation result

[0193] It can be understood that the above schemes are all for performing W in the above formula (3). xf X t +W xh h t-1 Taking this quantization operation as an example, the quantization operation method in the embodiments of the present application will be described. The quantization operation method in the embodiments of the present application can be used in any operator with associated operations.

[0194] For example, in the above formula (3) in the LSTM model, when performing the operation of W xf X t +W xh h t-1 +b f the first operator can refer to the operator for performing the quantization operation of executing W xf X t and the second operator can refer to the operator for performing the quantization operation of W xh h t-1 +b f .

[0195] In the above formula (4) in the LSTM model, when performing the operation of W xi X t +W hi h t-1 +b i the first operator can refer to the operator for performing the quantization operation of executing W xi X t and the second operator can refer to the operator for performing the quantization operation of executing W hi h t-1 +b i .

[0196] In the above formula (5) in the LSTM model, when performing the operation of f t c t-1 +i t g t the first operator can refer to the operator for performing the quantization operation of executing f t c t-1 and the second operator can refer to the operator for performing the quantization operation of executing i t g t .

[0197] It can be understood that based on the above solution, when the common quantization coefficient is one of the quantization coefficients corresponding to the first output data and the quantization coefficient corresponding to the second output data, only the output data corresponding to the other quantization coefficient needs to be dequantized to obtain floating-point data, and the floating-point data is quantized through the common quantization coefficient to obtain quantization data, which can then be directly operated on with the other output data that has not been dequantized to obtain the corresponding quantization operation result. Compared with the prior art where both output data need to be dequantized to obtain the corresponding floating-point data, and the floating-point operation result obtained by adding the floating-point data is quantized again to obtain the quantization operation result, the dequantization process of one of the data and the quantization process of the final operation result are reduced. In this way, the model running speed can be effectively improved, and the model accuracy can be improved.

[0198] Moreover, by using the largest quantization coefficient among the quantization coefficients corresponding to the first output data and the quantization coefficient corresponding to the second output data as the common quantization coefficient, the model accuracy can be further improved.

[0199] Secondly, when the difference between the quantization coefficient corresponding to the first output data and the quantization coefficient corresponding to the second output data is small, multiple preset common quantization coefficients can be obtained from the numerical range between the first quantization coefficient and the second quantization coefficient, and the preset common quantization coefficient corresponding to the operation result with the smallest accuracy loss finally obtained is used as the common quantization coefficient. In this way, the operation accuracy of the model can be effectively guaranteed.

[0200] The embodiment of the present application also provides a quantization operation device for a neural network model, which is used to execute the above quantization operation method. Specifically, as Figure 9 shown, the quantization operation device may include:

[0201] A determination module, configured to determine a first operator and a second operator that need to perform an associated operation, where the first operator has first output data and the second operator has second output data;

[0202] A first acquisition module, configured to obtain a common quantization coefficient based on a first quantization coefficient corresponding to the first output data and a second quantization coefficient corresponding to the second output data;

[0203] A second acquisition module, configured to obtain first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient, and obtain second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient;

[0204] An operation module, configured to perform an associated operation on the first quantization data and the second quantization data corresponding to the first operator and the second operator to obtain a quantization operation result of the associated operation of the first operator and the second operator.

[0205] An embodiment of the present application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the above quantization operation method.

[0206] An embodiment of the present application provides a computer program product including instructions for implementing the above quantization operation method.

[0207] An embodiment of the present application provides a storage medium having instructions stored thereon, and when the instructions are executed on a computer, the computer is caused to execute the above quantization operation method.

[0208] Embodiments of the mechanisms disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.

[0209] The program code can be applied to the input instructions to execute the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, a processing system includes any system having a processor such as, for example, a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.

[0210] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. This includes but is not limited to OpenCL, C language, C++, Java, etc. For languages such as C++ and Java, since they perform storage conversions, there will be some differences in the application of the data processing method in the embodiments of the present application. Those skilled in the art can make transformations based on specific high-level languages, all without departing from the scope of the embodiments of the present application.

[0211] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or via other computer-readable media. Thus, machine-readable media may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, magneto-optical disks, read-only memory (ROM), CD-ROMs, random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms using the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0212] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0213] It should be noted that the units / modules mentioned in the device embodiments of this application are all logical units / modules. Physically, a logical unit / module may be a physical unit / module, a part of a physical unit / module, or may be implemented as a combination of multiple physical units / modules. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. In addition, in order to highlight the innovative part of this application, the above device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that there are no other units / modules in the above device embodiments.

[0214] It should be noted that in the examples and the description of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0215] Although this application has been illustrated and described by reference to certain preferred embodiments thereof, those of ordinary skill in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.

Claims

1. A quantization operation method for a neural network model, used in an electronic device, characterized in that Including: Determine a first operator and a second operator that need to perform an associative operation. Among them, the first operator has first output data, and the second operator has second output data. Correspondingly, when the neural network model is a speech recognition model, the input data of the first operator includes first speech data, and the first operator is used to calculate the first speech data to obtain the first output data; Obtain a common quantization coefficient based on a first quantization coefficient corresponding to the first output data and a second quantization coefficient corresponding to the second output data; Obtain first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient, and obtain second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient; Perform the associative operation corresponding to the first operator and the second operator on the first quantization data and the second quantization data to obtain a quantization operation result of the associative operation of the first operator and the second operator; The obtaining of the common quantization coefficient based on the first quantization coefficient corresponding to the first output data and the second quantization coefficient corresponding to the second output data includes: When it is determined that the difference between the first quantization coefficient corresponding to the first output data and the second quantization coefficient corresponding to the second output data is greater than a set difference, the largest quantization coefficient among the first quantization coefficient and the second quantization coefficient is used as the common quantization coefficient.

2. The quantization operation method according to claim 1, wherein The obtaining of the common quantization coefficient based on the first quantization coefficient corresponding to the first output data and the second quantization coefficient corresponding to the second output data; further includes: When it is determined that the difference between the first quantization coefficient corresponding to the first output data and the second quantization coefficient corresponding to the second output data is less than or equal to the set difference, obtain the common quantization coefficient based on the first quantization coefficient, the second quantization coefficient and a preset rule.

3. The quantization operation method according to claim 2, wherein The obtaining of the common quantization coefficient based on the first quantization coefficient, the second quantization coefficient and the preset rule; includes: Equally divide the numerical range between the first quantization coefficient and the second quantization coefficient to obtain a set number of separation points; Use each of the separation points as a preset common quantization coefficient, and obtain quantization operation results of the associative operation of the first operator and the second operator corresponding to each of the preset common quantization coefficients; Obtain a floating-point operation result of the associative operation of the first operator and the second operator; Use the preset common quantization coefficient corresponding to the quantization operation result with the smallest precision loss relative to the floating-point operation result among each of the quantization operation results as the common quantization coefficient.

4. The quantization operation method according to claim 3, wherein The using of the preset common quantization coefficient corresponding to the quantization operation result with the smallest precision loss relative to the floating-point operation result among each of the quantization operation results as the common quantization coefficient includes: Use the preset common quantization coefficient corresponding to the quantization operation result with the highest similarity relative to the floating-point operation result among each of the quantization operation results as the common quantization coefficient.

5. The quantization operation method according to claim 1, wherein When the common quantization coefficient is the first quantization coefficient, the first quantization data is the first output data, and Obtaining the second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient includes: Obtaining the floating-point data corresponding to the second output data according to the second output data and the second quantization data, and obtaining the quantization conversion data of the second output data according to the floating-point data corresponding to the second output data and the common quantization coefficient; Taking the quantization conversion data of the second output data as the second quantization data.

6. The quantization operation method according to claim 1, characterized in that When the common quantization coefficient is the second quantization coefficient, the second quantization data is the second output data, and Obtaining the first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient includes: Obtaining the floating-point data corresponding to the first output data according to the first output data and the first quantization data, and obtaining the quantization conversion data of the first output data according to the floating-point data corresponding to the first output data and the common quantization coefficient; Taking the quantization conversion data of the first output data as the first quantization data.

7. The quantization operation method according to claim 1, wherein When the common quantization coefficient is a value other than the first quantization coefficient and the second quantization coefficient, Obtaining the first quantization data corresponding to the first output data according to the first output data and the common quantization coefficient, and obtaining the second quantization data corresponding to the second output data according to the second output data and the common quantization coefficient include: Obtaining the floating-point data corresponding to the first output data according to the first output data and the first quantization data, and obtaining the quantization conversion data of the first output data according to the floating-point data corresponding to the first output data and the common quantization coefficient; Obtaining the floating-point data corresponding to the second output data according to the second output data and the second quantization data, and obtaining the quantization conversion data of the second output data according to the floating-point data corresponding to the second output data and the common quantization coefficient; Taking the quantization conversion data of the second output data as the second quantization data.

8. An electronic device, characterized in that, Includes: A memory for storing instructions executed by one or more processors of an electronic device, and a processor, which is one of the processors of the electronic device, for executing the quantization operation method according to any one of claims 1 to 7.

9. A computer program product, characterized in that, Includes instructions for implementing the quantization operation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Quantification method, computing device and computer readable storage medium

    CN114118341A