Quantization Processing Method, Device, Equipment and Storage Medium of Neural Network Model
By predicting the arithmetic overflow frequency that occurs after quantization of each layer in the neural network model and adjusting the calculation data based on the prediction results, the problems of poor quantization operation flexibility and high overflow frequency in the prior art are solved, and a more flexible and efficient quantization training process is achieved.
Patent Information
- Application Number
- CN202010088221.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-02-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-02-12
AI Technical Summary
In the prior art, the quantization method of neural network models is poor in flexibility, which leads to the arithmetic overflow problem that the quantized model is prone to arithmetic overflow during the calculation process.
During the quantitative training of the neural network model, the frequency of arithmetic overflow occurs after quantization of each neural network layer is predicted, and the target calculation data is determined based on the predicted overflow frequency to perform quantitative training.
By predicting and adjusting the computational data, the overflow of the quantized neural network model is effectively controlled, which improves the flexibility of quantization operations and reduces the probability of arithmetic overflow during use of the model after training.
Smart Images

Figure CN113255877B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technologies, and in particular, to a method, apparatus, device, and storage medium for quantifying a neural network model. Background Art
[0002] A neural network model contains a large number of model parameters, most of which are floating-point types and occupy a relatively large amount of storage space. At the same time, the neural network model performs floating-point operations based on the floating-point model parameters, which also occupies a relatively large amount of computing resources.
[0003] Generally, the neural network model can be quantified to reduce the data volume of the neural network model and improve the computing speed of the neural network model.
[0004] However, the quantization method of the neural network model provided by the prior art has the defect of poor flexibility. Therefore, there is a need to propose a new solution. Summary of the Invention
[0005] Multiple aspects of this application provide a method, apparatus, device, and storage medium for quantifying a neural network model to improve the flexibility of the quantization operation of the neural network model.
[0006] An embodiment of this application provides a method for quantifying a neural network model, including: obtaining the calculation data required by a first neural network layer in the neural network model; predicting the frequency of arithmetic overflow of the first neural network layer after quantization according to the calculation data; determining target calculation data for quantizing and training the first neural network layer according to the frequency of arithmetic overflow and the calculation data, so as to perform quantization training on the first neural network layer.
[0007] An embodiment of this application further provides a device for quantifying a neural network model, including: an input module, configured to: obtain the calculation data required by a first neural network layer in the neural network model; an overflow prediction module, configured to: predict the frequency of arithmetic overflow of the first neural network layer after quantization according to the calculation data; a quantization training module, configured to: determine target calculation data for quantizing and training the first neural network layer according to the frequency of arithmetic overflow and the calculation data, so as to perform quantization training on the first neural network layer.
[0008] An embodiment of this application further provides an electronic device, including: a memory and a processor; the memory is configured to store one or more computer instructions; the processor is configured to execute the one or more computer instructions to: execute the method for quantifying a neural network model provided by an embodiment of this application.
[0009] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the quantization processing method of the neural network model provided by the embodiment of the present application.
[0010] In the embodiment of the present application, during the quantization training of the neural network model, the frequency of arithmetic overflow after quantization of any neural network layer in the prediction neural network model is predicted, and the target calculation data for the quantization training of the first neural network layer is determined according to the predicted frequency of arithmetic overflow. Based on this method, it is convenient to control the overflow situation of the quantized neural network model, which is beneficial to improving the flexibility of the quantization operation of the neural network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0012] Figure 1 It is a schematic flowchart of the quantization method of the neural network model provided by an exemplary embodiment of the present application;
[0013] Figure 2a It is a schematic flowchart of the quantization method of the neural network model provided by another exemplary embodiment of the present application;
[0014] Figure 2b It is a schematic diagram of the data quantization process provided by an exemplary embodiment of the present application;
[0015] Figure 3 It is a schematic diagram of the step size provided by an exemplary embodiment of the present application;
[0016] Figure 4 It is a schematic diagram of the step size provided by another exemplary embodiment of the present application;
[0017] Figure 5 It is a schematic diagram of controlling the mapping range through a scaling operation provided by an exemplary embodiment of the present application;
[0018] Figure 6 It is a schematic structural diagram of the data quantization device of the neural network model provided by an exemplary embodiment of the present application;
[0019] Figure 7 It is a schematic structural diagram of a convolutional layer including an overflow prediction module provided by an exemplary embodiment of the present application;
[0020] Figure 8 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0022] A neural network model refers to a deep learning model obtained by training an artificial neural network. A neural network model contains various model parameters. Generally, the model parameters and their operation processes are numerically represented using floating-point numbers (float data type). For example, single-precision floating-point numbers of 32 bits (the smallest unit for measuring information) or double-precision floating-point numbers of 64 bits are often used for representation.
[0023] Model quantization usually refers to numerically representing the floating-point numbers in a neural network model with integers. For example, fixed-point integers of 8 bits or 16 bits are used for numerical representation.
[0024] The quantization operation of a neural network model includes quantization during training and quantization after training. Among them, quantization after training means training a neural network model with floating-point numbers for numerical representation. Then, a batch of sample inputs are used to input into the neural network model, the numerical range of the output data of the neural network model is statistically analyzed, and then the neural network model is quantized according to the numerical range of the output data. Among them, quantization during training means simulating the quantization behavior during the process of training a neural network model, using floating-point numbers to save fixed-point parameters, and directly using the saved fixed-point parameters when the neural network model is put into use after training is completed.
[0025] In some scenarios, there is an arithmetic overflow problem in the quantized neural network model. Arithmetic overflow refers to the result of an arithmetic operation by a computer exceeding the range that the machine can represent. That is, the numerical value representing the output data exceeds the range that a fixed-point number of a certain length can express.
[0026] To address the above technical problem of arithmetic overflow, in some embodiments of this application, a solution is provided. In this solution, based on quantization during training, further improvements are made. An overflow prediction module is added to the neural network model. This overflow prediction module can predict the arithmetic overflow situation of the quantized neural network model, which is beneficial to reducing the occurrence of arithmetic overflow in the quantized neural network model. The technical solutions provided by each embodiment of this application will be described in detail below in conjunction with the drawings.
[0027] Figure 1The flowchart of the quantization processing method for the neural network model provided by an exemplary embodiment of the present application is shown as follows Figure 1 The method includes:
[0028] Step 101: Obtain the calculation data required for the first neural network layer in the neural network model.
[0029] Step 102: Predict the frequency of arithmetic overflow occurring in the quantized first neural network layer according to the calculation data.
[0030] Step 103: Determine the target calculation data for quantizing and training the first neural network layer according to the frequency of arithmetic overflow and the calculation data, so as to perform quantized training on the first neural network layer.
[0031] The first neural network layer refers to any neural network layer in the neural network model. Here, "first" is used to limit the neural network layer, which is only for convenient description and distinction, and does not impose any restrictions on the order or position of the neural network layer. For example, for a Convolutional Neural Networks (CNN) model, the first neural network layer may include: the input layer, any convolutional layer, pooling layer, fully connected layer, etc. in the convolutional neural network. This embodiment does not make any restrictions.
[0032] The calculation data refers to one or more data required for the first neural network layer to complete the calculation of this layer, including the data passed from the previous layer and may also include the model parameters within this layer. For example, when the first neural network layer is implemented as any convolutional layer in the CNN, the calculation data required for this convolutional layer may include the image data input to the convolutional layer, the convolutional kernel for image feature extraction, and the activation function for passing the output data of this layer to the next layer, etc., which will not be elaborated here.
[0033] The operation of predicting the frequency of arithmetic overflow occurring in the output data of the quantized first neural network layer according to the calculation data can be implemented based on the overflow prediction module in the neural network model. This overflow prediction module can sense whether there is an arithmetic overflow phenomenon in the calculation process of the quantized neural network model and count the overflow frequency, which is beneficial to further optimizing the accuracy of the quantized neural network model.
[0034] Among them, the target calculation data refers to the data used for quantized training of the first neural network layer, that is, the data actually participating in the quantized training process of the first neural network layer. The target calculation data can be adjusted according to the overflow situation of the output data of the first neural network layer. Based on this method, according to the predicted overflow situation of the quantized neural network model, the calculation data participating in the actual quantized training process of the first neural network layer can be adjusted backward, which is beneficial to reducing the probability of arithmetic overflow of the neural network model after training is completed.
[0035] In this embodiment, during the process of quantized training of the neural network model, the frequency of arithmetic overflow after quantization of any neural network layer in the neural network model is predicted, and the target calculation data used for quantized training of the first neural network layer is determined according to the predicted frequency of arithmetic overflow. Based on this method, it is convenient to control the overflow situation of the quantized neural network model, which is beneficial to improving the flexibility of the quantization operation of the neural network model.
[0036] Figure 2a For the flow diagram of the quantization processing method of the neural network model provided by another exemplary embodiment of this application, as Figure 2a shown, the method includes:
[0037] Step 201, obtain the calculation data required for the first neural network layer in the neural network model.
[0038] Step 202, according to the calculation data, predict the frequency of arithmetic overflow of the quantized first neural network layer.
[0039] Step 203, determine whether the frequency of arithmetic overflow is less than or equal to the set overflow frequency threshold; if yes, execute Step 204, if no, execute Step 205.
[0040] Step 204, if the frequency of arithmetic overflow is less than or equal to the set overflow frequency threshold, then use the calculation data as the target calculation data, and perform quantized training on the first neural network layer according to the target data.
[0041] Step 205, if the frequency of arithmetic overflow is greater than the set overflow frequency threshold, then adjust the calculation data, and execute Step 202 according to the adjusted calculation data.
[0042] In Step 201, the calculation data required for the first neural network layer may include the model parameters of the first neural network layer and the input data of the first neural network layer. When the first neural network layer is not the first neural network layer in the neural network model, the input data of the first neural network layer may be the output data of the previous neural network layer of the first neural network layer.
[0043] Among them, when the neural network model is different, the model parameters are also different. For example, for the convolutional layer in a CNN, the model parameters may include the weights and activation functions of this layer, etc.
[0044] In step 202, based on the calculation data, the frequency of arithmetic overflow occurring in the first quantized neural network layer can be predicted.
[0045] Optionally, in this embodiment, the calculation process of the first quantized neural network layer can be inferred, and the frequency of arithmetic overflow can be obtained based on the inference result. In deep learning, inference refers to putting the trained neural network model into the usage link so that the neural network model can apply the performance it has learned.
[0046] Optionally, in this step, the quantization operation on the calculation data can be inferred to obtain the quantized calculation data, and based on the quantized calculation data, the calculation process of the first quantized neural network layer can be inferred to obtain the intermediate calculation result and output data of the first neural network layer.
[0047] Among them, the output data refers to the calculation result that the first neural network layer needs to pass to the next layer. Usually, the calculation process of the first neural network calculation layer includes multiple calculation operations. Before obtaining the output data, the result obtained by each calculation operation can be called the intermediate calculation result.
[0048] Among them, the quantization operation on the calculation data refers to the operation of converting the calculation data into a fixed-point integer. Optionally, the fixed-point integer can be implemented as an integer data of a set number of bits (int data type), for example, 8-bit integer data, 7-bit integer data, etc.
[0049] Optionally, first, the numerical range to which the calculation data belongs can be obtained. For example, the numerical range to which the input data of the first neural network layer belongs and the numerical range to which the model parameters of the first neural network layer belong can be obtained. For the convenience of description below, the numerical range to which the input data belongs is described as [Imin, Imax], and the numerical range to which the model parameters belong is described as [Pmin, Pmax].
[0050] Next, based on the numerical range to which the calculation data belongs and the numerical range to which the integer data of the set number of bits belongs, the quantization mapping relationship can be determined. Taking the model parameters as an example, the quantization mapping relationship can be described as:
[0051]
[0052] Among them, P refers to the floating-point model parameter before quantization, P` refers to the integer model parameter after quantization, 2 B The number of bits of the integer data, int() represents the function of obtaining an integer.
[0053] Next, according to the quantization mapping relationship, the calculation data is mapped to integer calculation data with a set number of bits. That is, the floating-point input data and the floating-point model parameters can be mapped to integer input data and integer model parameters with a set number of bits. The following will be described with specific examples.
[0054] For example, the floating-point input data is: Imin = 0, Imax = 1.0, B = 3. Based on the mapping relationship expressed by the above formula, float:0 can be quantized to int:0, float:0.5 can be quantized to int:128, float:0.0039 (1. / 255) can be quantized to int:1, and values within the range of float:[0,0.0039] can be quantized to int:1.
[0055] Based on the above steps, the integer model parameters with a set number of bits and the integer input data with a set number of bits can be quantized. Next, the calculation process of the first neural network layer after quantization can be inferred.
[0056] Optionally, when inferring the calculation process of the first neural network layer after quantization, integer operations can be performed on the calculation process of the first neural network layer after quantization according to the integer model parameters and the integer input data, in accordance with the calculation logic of the first neural network layer. For example, when the first neural network layer is implemented as a convolutional layer in a CNN, the convolutional calculation process within this layer can be executed according to the quantized integer image data and weight parameters, in accordance with the integer operation rules, and the intermediate calculation results generated during the convolutional calculation process and the finally calculated feature vector can be obtained.
[0057] After obtaining the intermediate calculation results and the output data of the first neural network layer, it can be determined whether an arithmetic overflow occurs in the intermediate calculation results and whether an arithmetic overflow occurs in the output data. Based on the judgment results, the frequency of arithmetic overflow in the first neural network layer can be determined. In this embodiment, the frequency of this arithmetic overflow can be calculated according to the intermediate calculation results, the output data, and the set data storage range. For example, if a certain intermediate calculation result exceeds the set data storage range, one arithmetic overflow can be recorded; if the output data exceeds the set data storage range, one arithmetic overflow can be recorded.
[0058] Optionally, the set data storage range can include the numerical ranges corresponding to any integer data from 1 to 64 bits. For example, it can be the numerical range corresponding to 1-bit integer data, the numerical range corresponding to 8-bit integer data, the numerical range corresponding to 16-bit integer data, or the numerical range corresponding to 32-bit integer data. This embodiment includes but is not limited to this.
[0059] Among them, the data storage ranges corresponding to the intermediate calculation results and the data storage ranges corresponding to the output data may be the same or different, and this embodiment does not make any restrictions. For example, in some scenarios, the storage range corresponding to the intermediate calculation results may be the numerical range corresponding to 8-bit integer data, and the storage range corresponding to the output data may be the numerical range corresponding to 16-bit integer data. In some other scenarios, the storage ranges corresponding to the intermediate calculation results and the output data may both be the numerical ranges corresponding to 16-bit integer data, which will not be elaborated here.
[0060] For example, in some typical cases, the calculation process of the first neural network layer includes multiplication and addition operations, as shown in the following formula:
[0061] R = x1 * w1 + x2 * w2 + … + x n * w n
[0062] Among them, x n represents the nth input data, w n represents the nth weight parameter of the first neural network layer, x n * w n represents the intermediate calculation result obtained by performing the nth multiplication operation, R represents the output data of the first neural network layer, and n is a positive integer. It should be understood that based on the above formula, the result obtained after each multiplication operation is called an intermediate calculation result, and the result obtained after each addition operation can also be called an intermediate calculation result. For example, the intermediate calculation result can also be: x1 * w1 + x2 * w2, x1 * w1 + x2 * w2 + x3 * w3, x1 * w1 + x2 * w2 + … + x n-1 * w n-1 and so on, which will not be listed one by one.
[0063] Assume that the quantized model parameters w1, w2 … w n and the quantized input data x1, x2 … x n are all represented in 8-bit integers. The numerical range corresponding to 8-bit integer data is [-127, 128]. Assume that the values of x1, x2 … x n are all 127, and the values of w1, w2 … w n are all 127.
[0064] If n = 100, then after performing multiplication and addition operations based on the above formula, 100 results of 127 * 127 added together can be obtained, R = 1,612,900. When storing this output data as 16-bit integer data, this output data R exceeds the numerical range [-32768, 32767] corresponding to 16-bit integer data, so it can be considered that arithmetic overflow has occurred in the output data R.
[0065] When storing the output data as 32-bit integer data, if n is large enough, the output data R will exceed the numerical range [-214783648, +2147483647] corresponding to int32 with a certain probability, and thus arithmetic overflow will occur.
[0066] If the intermediate calculation result is stored as 10-bit integer data, then, if the intermediate calculation result x n *w n exceeds the numerical range [-1024, 1023] corresponding to 10-bit integer data, it can be considered that arithmetic overflow has occurred in the intermediate calculation result.
[0067] In step 203, optionally, after obtaining the frequency of arithmetic overflow, it can be determined whether the frequency of arithmetic overflow is less than or equal to a set overflow frequency threshold. Among them, the set overflow frequency threshold can be set according to the actual situation. The set overflow frequency threshold can represent the tolerance for arithmetic overflow. If the tolerance for arithmetic overflow is high, a larger overflow frequency threshold can be set; if the tolerance for arithmetic overflow is low, a smaller overflow frequency threshold can be set.
[0068] Optionally, in some embodiments, the overflow frequency threshold can be set to zero to ensure the accuracy of the neural network model obtained by quantization training.
[0069] In step 204, if the frequency of arithmetic overflow is less than or equal to the set overflow frequency threshold, the calculation data can be used as the target calculation data, and the first neural network layer can be quantized and trained according to the target data.
[0070] In step 205, if the frequency of the arithmetic overflow is greater than the set overflow frequency threshold, the calculation data can be adjusted, and the adjusted calculation data can be used to return to step 202 to perform the prediction operation until the arithmetic overflow no longer occurs in the predicted output data.
[0071] In some alternative embodiments, adjusting the calculation data may include: replacing the existing calculation data with new calculation data with a smaller value. For example, the value range of the input data can be reduced, or the value range of the model parameters can be reduced.
[0072] In some other alternative embodiments, adjusting the calculation data may include: scaling the calculation data according to a set scaling factor. The principle of how scaling the calculation data affects arithmetic overflow will be described in detail below with reference to the accompanying drawings and specific examples.
[0073] It should be understood that quantifying data means changing continuous floating-point numbers within a certain range into discrete stepped integers. The step size between adjacent integers is determined according to the upper and lower limit ranges of the floating-point numbers and the mapping range of the corresponding integers for the quantization operation. The following will be described by way of specific examples.
[0074] Suppose the model parameter Pmin = 0.0 and Pmax = 255.0. The numerical range of 8-bit unsigned integer data is [0, 255]. When mapping the model parameter to an 8-bit unsigned integer, the step size (stride) = 1. As Figure 2b shown, 253.0, 253.1, 253.2, and 253.3 can all be mapped to the integer 253.
[0075] The numerical range of 6-bit unsigned integer data is [0, 63]. If the above model parameter is mapped to a 6-bit unsigned integer, the step size = 4. Then Pmax = 255.0 will be mapped to 63, and all floating-point values between [252.0, 255.0] will also be mapped to 63.
[0076] When the upper and lower limit ranges of the floating-point numbers are the same, the step size for quantifying the floating-point numbers into 8-bit integer data can refer to the illustration in Figure 3 and the step size for quantifying the floating-point numbers into 6-bit integer data can refer to the illustration in Figure 4 Based on Figure 3 and Figure 4 it can be seen that when the numerical ranges to which the data to be quantified belong are the same, the larger the mapping range corresponding to the integer data, the smaller the step size of the discrete integers, and the larger the value range of the integer data obtained after quantifying the data to be quantified.
[0077] To prevent calculation results from overflowing, the step size can be appropriately increased to narrow the value range of the calculated data after quantization. Based on the above analysis, an optional way to increase the step size is to increase the numerical range to which the calculated floating-point data to be quantified belongs while keeping the mapping range corresponding to the integer data unchanged.
[0078] Optionally, in this embodiment, the numerical range to which the calculated data belongs can be scaled.
[0079] First, the calculated data can be statistically analyzed to obtain the upper limit value and the lower limit value of the calculated data, and the numerical range to which the calculated data belongs can be determined according to the upper limit value and the lower limit value. For example, the numerical range [Pmin, Pmax] to which the model parameters of the first neural network layer belong and the numerical range [Imin, Imax] to which the input data belongs can be statistically analyzed respectively.
[0080] Next, the numerical range to which the calculated data belongs can be scaled.
[0081] As Figure 5 shown, assume Pmin = -1, Pmax = 1. If it is mapped to 8-bit integer data, the step size of the ladder is 1 / 128. After scaling it by 4 times, [4*Pmin, 4*Pmax] = [-4, 4]. After mapping the scaled range to 8-bit integer data, the step size of the ladder is 1 / 32. Then, the actual mapping range of [-1, 1] will be reduced to [-32, 31].
[0082] For the convenience of description, the scaling multiple can be represented by α. Among them, the value of the scaling factor α can be an integer or can include a decimal, and this embodiment does not make a limitation. When the value range of α is flexible enough, the calculated data can be mapped to any integer range, greatly improving the flexibility of quantization.
[0083] Optionally, in this embodiment, the model parameters of the first neural network layer can be scaled according to the first scaling multiple; or, the input data of the first neural network layer can be scaled according to the second scaling multiple. Or, the above scaling processing can be performed on the model parameters and the input data at the same time, and this embodiment does not make a limitation.
[0084] Optionally, the first scaling multiple and the second scaling multiple can be the same or different, and the two can be determined according to the frequency and / or the amount of arithmetic overflow of the output data. When the frequency of arithmetic overflow is high, a larger scaling multiple can be set. When the frequency of arithmetic overflow is low, a smaller scaling multiple can be set; or, when the amount of arithmetic overflow is large, a larger scaling multiple can be set. When the amount of arithmetic overflow is small, a smaller scaling multiple can be set, and this embodiment does not make a limitation.
[0085] It should be noted that in a scenario, if it is predicted that no arithmetic overflow occurs in the first neural network layer, and when inferring the calculation process of the first neural network layer, the numerical ranges of the intermediate calculation results and the output data of the first neural network layer are both much smaller than the set data storage range, then the model parameters and / or the input data corresponding to the first neural network layer can be scaled. Among them, the scaling multiple can be set between 0 and 1 to reduce the model parameters and / or the input data. Based on this method, the set data storage range can be fully utilized and reasonably used, which is beneficial to improving the accuracy of the trained model.
[0086] In this embodiment, by changing the numerical range to which the calculated data belongs to control the quantization mapping range, the overflow situation of the neural network model can be flexibly controlled, which is beneficial to improving the flexibility of the quantization operation of the neural network model.
[0087] In addition, in the quantization method of the neural network model provided in the embodiments of the present application, after predicting the frequency of arithmetic overflow of the calculation result of the first neural network layer, the target calculation data can be determined according to the frequency of arithmetic overflow. The target calculation data is more suitable for quantizing and training the neural network model. On the one hand, it reduces the probability of arithmetic overflow occurring during the use of the neural network model after training is completed, and improves the accuracy of the quantized model; on the other hand, the process of predicting overflow is related to the actual training task of the model. Without overflow occurring, the floating-point calculation data can be mapped to the largest possible integer range, which can avoid over-compression of the model to a certain extent.
[0088] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 201 to 204 can be device A; for another example, the execution subject of steps 201 and 202 can be device A, and the execution subject of step 203 can be device B; and so on.
[0089] In addition, in some processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 201 and 202 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in order or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are different types.
[0090] It is worth noting that the quantization processing method of the neural network model provided in the above embodiments of the present application can be encapsulated into a tool available for third parties to use, such as a SaaS (Software-as-a-Service) tool. Optionally, the SaaS tool can be implemented as a plug-in or an application program. Based on this SaaS tool, third-party users can conveniently use the quantization processing service of the neural network model.
[0091] For example, in some scenarios, the SaaS tool can be deployed on a cloud platform, and third-party users can access the cloud platform to use the SaaS tool online. Based on this SaaS tool, third-party users can quantize and train the neural network model online, predict the arithmetic overflow situation of the neural network model in real time through this SaaS tool, and can adjust the quantization training process according to the arithmetic overflow situation.
[0092] It is also worth noting that the quantization processing method of the neural network model provided in the above embodiments of the present application can be applied to a variety of different neural network models. For example, a Convolutional Neural Networks (CNN) model, a Deep Neural Network (DNN) model, a Graph Convolutional Networks (GCN) model, a Recurrent Neural Network (RNN) model, and a Long Short-Term Memory (LSTM) model, one or more of them, or it can also be applied to other neural network models obtained by transforming one or more of the above neural networks.
[0093] When applied to different neural network models, the parameters required for the SaaS tool to execute the quantization processing method of the neural network model are also different. Among them, the parameters may include: the overflow frequency threshold and the scaling factor required for adjusting the calculation data, etc.
[0094] Optionally, the above parameters can be provided by the SaaS tool or provided by the user actively. For example, in some scenarios, the SaaS tool can provide different default parameters for different types of neural networks. In other scenarios, the SaaS tool can interact with third-party users to obtain personalized customized parameters provided by the third-party users, and execute the quantization processing method according to the parameters provided by the third-party users, which will not be elaborated here.
[0095] Figure 6 For the quantization processing device of the neural network model provided in an exemplary embodiment of the present application, as Figure 6 shown, the device includes:
[0096] An input module 61, configured to: obtain the calculation data required for the first neural network layer in the neural network model.
[0097] An overflow prediction module 62, configured to: predict the frequency of arithmetic overflow of the quantized first neural network layer according to the calculation data.
[0098] A quantization training module 63, configured to: determine the target calculation data for quantizing and training the first neural network layer according to the frequency of the arithmetic overflow and the calculation data, so as to perform quantization training on the first neural network layer.
[0099] Further optionally, the overflow prediction module 62 includes: a quantization sub-module 621. When predicting the frequency of arithmetic overflow in the first neural network layer after quantization according to the calculation data, the overflow prediction module 62 is specifically configured to: through the quantization sub-module 621, infer the quantization operation on the calculation data to obtain the quantized calculation data; through the calculation sub-module 622, infer the calculation process of the first neural network layer after quantization according to the quantized calculation data to obtain the intermediate calculation result and the output data of the first neural network layer; calculate the frequency of the arithmetic overflow according to the intermediate calculation result, the output data, and the set data storage range.
[0100] Further optionally, when inferring the quantization operation on the calculation data, the quantization sub-module 621 is specifically configured to: determine the quantization mapping relationship according to the numerical range to which the calculation data belongs and the numerical range of the integer data with the set number of bits; map the calculation data to the integer calculation data with the set number of bits according to the quantization mapping relationship.
[0101] Further optionally, the integer calculation data with the set number of bits includes: integer model parameters and integer input data; when inferring the calculation process of the first neural network layer after quantization according to the quantized calculation data, the overflow prediction module 62 is specifically configured to: through the calculation sub-module 622, perform integer calculation on the calculation process of the first neural network layer after quantization according to the integer model parameters and the integer input data according to the calculation logic of the first neural network layer.
[0102] Further optionally, the set data storage range includes: the numerical range corresponding to any integer data from 1 to 64 bits.
[0103] Further optionally, when determining the target calculation data for quantizing and training the first neural network layer according to the frequency of the arithmetic overflow and the calculation data, the quantization training module 63 is specifically configured to: if the frequency of the arithmetic overflow is less than or equal to the set overflow frequency threshold, use the calculation data as the target calculation data.
[0104] Further optionally, the overflow prediction module 62 further includes: an adjustment sub-module 623; the adjustment sub-module 623 is configured to: when the frequency of the arithmetic overflow is greater than the set overflow frequency threshold, adjust the calculation data so that the overflow prediction module 62 performs the prediction step according to the adjusted calculation data.
[0105] Further optionally, when adjusting the calculation data, the adjustment sub-module 623 is specifically configured to: perform a scaling process on the numerical range to which the calculation data belongs according to the set scaling factor.
[0106] Further optionally, when the adjustment sub-module 623 scales the numerical range to which the calculation data belongs according to a set scaling factor, it performs at least one of the following operations: scales the numerical range to which the model parameters of the first neural network layer belong according to a first scaling factor; scales the numerical range to which the input data of the first neural network layer belongs according to a second scaling factor.
[0107] Further optionally, the first scaling factor and / or the second scaling factor are determined according to the frequency of the arithmetic overflow.
[0108] The quantization processing device will be further described below in conjunction with the structure of the neural network model.
[0109] Figure 7 Illustrates the typical operation steps when quantizing the convolutional layer in the neural network model, such as Figure 7 shown, the overflow prediction module 62 can execute Figure 7 the following steps: scaling, quantization, quantized convolution, and calculating the frequency N of arithmetic overflow. Among them, the quantized convolution operation refers to performing a convolution operation with a set number of integer bits on the quantized input data and storing the calculation result as an integer with a specified number of bits. For example, the input data is 8-bit integer data, and the calculation result is stored as a 16-bit integer. In addition to the steps executed by the overflow prediction module 62, other steps such as fake quantization, multiplication, addition, and convolution are the steps required for model quantization training and will not be elaborated here.
[0110] During the model quantization training process, the overflow prediction module 62 can sense whether an overflow occurs. Among them, in the quantization step, the overflow prediction module 62 can quantize the float-type input data into 8-bit integers according to the upper and lower limit values of the input data. And, it can quantize the float-type weights into 8-bit integers according to the upper and lower limit values of the weights. Then, the overflow prediction module 62 can perform a convolution calculation based on the quantized 8-bit input data and weights in the convolution-int16 step, and compare the result of the convolution calculation with the value range of int16 to predict the frequency N of arithmetic overflow in the convolutional layer. If the frequency N of arithmetic overflow is greater than the set overflow frequency threshold, the overflow prediction module 62 can, in the scaling step, correspondingly scale the upper and lower limit ranges of the input data and the upper and lower limit ranges of the weights, and re-execute the operations in the quantization step until the frequency N of arithmetic overflow is less than or equal to the set overflow frequency threshold.
[0111] The advantages of the above structure are as follows. Firstly, by scaling the upper and lower limit ranges of the input data and the upper and lower limit ranges of the weights, the purpose of controlling the quantization mapping range is achieved, overcoming the shortcomings of using a fixed number of bits for quantization operations. It is very flexible to map the data to be quantized to any range. Secondly, the control of overflow is related to the training task. Without overflow, the data to be quantized can be mapped to the largest possible integer range, thus ensuring that the capacity of the model is not overly compressed. Thirdly, the overflow prediction module 62 can work simultaneously with the quantization link during training, ensuring the accuracy of the quantized model.
[0112] Figure 8 FIG. shows a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application, such as Figure 8 shown, the electronic device includes: a memory 801, a processor 802, and a communication component 803.
[0113] The memory 801 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0114] Among them, the memory 801 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0115] The processor 802 is coupled to the memory 801 and is used to execute the computer program in the memory 801 for: obtaining the calculation data required for the first neural network layer in the neural network model; predicting the frequency of arithmetic overflow of the quantized first neural network layer according to the calculation data; determining the target calculation data for quantizing and training the first neural network layer according to the frequency of the arithmetic overflow and the calculation data, so as to perform quantization training on the first neural network layer.
[0116] Further optionally, when predicting the frequency of arithmetic overflow of the quantized first neural network layer according to the calculation data, the processor 802 is specifically used for: inferring the quantization operation on the calculation data to obtain the quantized calculation data; inferring the calculation process of the quantized first neural network layer according to the quantized calculation data to obtain the intermediate calculation result and the output data of the first neural network layer; calculating the frequency of the arithmetic overflow according to the intermediate calculation result, the output data, and the set data storage range.
[0117] Further optionally, when the processor 802 infers the quantization operation on the calculation data, it is specifically configured to: determine a quantization mapping relationship according to the numerical range to which the calculation data belongs and the numerical range to which the integer data with a set number of bits belongs; and map the calculation data to the integer calculation data with the set number of bits according to the quantization mapping relationship.
[0118] Further optionally, the integer calculation data with the set number of bits includes: integer model parameters and integer input data; when the processor 802 infers the calculation process of the first neural network layer after quantization according to the quantized calculation data, it is specifically configured to: perform integer calculation on the calculation process of the first neural network layer after quantization according to the integer model parameters and the integer input data according to the calculation logic of the first neural network layer.
[0119] Further optionally, the set data storage range includes: any integer from 1 to 64 bits.
[0120] Further optionally, when the processor 802 determines the target calculation data for quantizing and training the first neural network layer according to the frequency of arithmetic overflow and the calculation data, it is specifically configured to: if the frequency of arithmetic overflow is less than or equal to a set overflow frequency threshold, use the calculation data as the target calculation data; if the frequency of arithmetic overflow is greater than the set overflow frequency threshold, adjust the calculation data and re - execute the prediction step according to the adjusted calculation data.
[0121] Further optionally, when the processor 802 adjusts the calculation data, it is specifically configured to: perform a scaling process on the numerical range to which the calculation data belongs according to a set scaling factor.
[0122] Further optionally, when the processor 802 performs a scaling process on the numerical range to which the calculation data belongs according to a set scaling factor, it performs at least one of the following operations: perform a scaling process on the numerical range to which the model parameters of the first neural network layer belong according to a first scaling factor; perform a scaling process on the numerical range to which the input data of the first neural network layer belongs according to a second scaling factor.
[0123] Further optionally, the first scaling factor and / or the second scaling factor are determined according to the frequency of arithmetic overflow.
[0124] Further, as Figure 8 shown, the electronic device further includes: a display 804, a power supply component 805 and other components. Figure 8 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 8The components shown.
[0125] Among them, the communication component 803 is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as WiFi, 2G, 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can be implemented based on near-field communication (NFC) technology, radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0126] Among them, the display 804 includes a screen, and the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from users. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations.
[0127] Among them, the power supply component 805 provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device where the power supply component is located.
[0128] In this embodiment, during the process of quantizing and training the neural network model, the frequency of arithmetic overflow after quantization of any neural network layer in the neural network model is predicted, and the target calculation data for quantizing and training the first neural network layer is determined according to the predicted frequency of arithmetic overflow. Based on this method, it is convenient to control the overflow situation of the quantized neural network model, which is beneficial to improving the flexibility of the quantization operation of the neural network model.
[0129] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executable by an electronic device in the above method embodiment.
[0130] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0131] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0132] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0134] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0135] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0136] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0137] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0138] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A quantization processing method for a neural network model, characterized in that, Including: Obtain the calculation data required for the first neural network layer in the neural network model. The calculation data refers to one or more types of data required for the first neural network layer to complete the calculation of this layer. When using the neural network model to process image data or video, the calculation data includes image data or video. The calculation data is a floating-point number, and the calculation data has a numerical range to which it belongs and the numerical range to which the calculation data belongs can be enlarged; According to the numerical range to which the calculation data belongs, determine the quantization mapping relationship for converting the calculation data into a fixed-point number with a set number of digits. Among them, the largest floating-point number in the numerical range to which the calculation data belongs is used to map to the maximum value of the fixed-point number with the set number of digits, and the smallest floating-point number in the numerical range to which the calculation data belongs is used to map to the minimum value of the fixed-point number with the set number of digits; According to the calculation data and the quantization mapping relationship, predict the frequency of arithmetic overflow occurring in the quantized first neural network layer; According to the frequency of arithmetic overflow and the calculation data, determine the target calculation data for quantizing and training the first neural network layer, so as to perform quantization training on the first neural network layer; Among them, determining the target calculation data for quantizing and training the first neural network layer according to the frequency of arithmetic overflow and the calculation data includes: if the frequency of arithmetic overflow is greater than the set overflow frequency threshold, magnify the numerical range to which the calculation data belongs according to a set multiple, and re-execute the steps of determining the quantization mapping relationship for converting the calculation data into a fixed-point number with a set number of digits and prediction according to the adjusted numerical range.
2. The method according to claim 1, characterized in that Predicting the frequency of arithmetic overflow occurring in the quantized first neural network layer according to the calculation data and the quantization mapping relationship includes: According to the quantization mapping relationship, map the calculation data to the fixed-point number with the set number of digits to obtain the quantized calculation data; According to the quantized calculation data, infer the calculation process of the quantized first neural network layer to obtain the intermediate calculation result and output data of the first neural network layer; According to the intermediate calculation result, the output data and the set data storage range, calculate the frequency of arithmetic overflow.
3. The method according to claim 2, characterized in that, The calculation data includes: model parameters and input data; Inferring the calculation process of the quantized first neural network layer according to the quantized calculation data includes: According to the quantized model parameters and the quantized input data, perform integer calculation on the quantized calculation process of the first neural network layer according to the calculation logic of the first neural network layer.
4. The method according to claim 2 or 3, characterized in that, The set data storage range includes the numerical ranges corresponding to any integer data from 1 to 64 bits.
5. The method according to any one of claims 1 to 3, characterized in that Determining the target calculation data for quantizing and training the first neural network layer according to the frequency of arithmetic overflow and the calculation data further includes: If the frequency of arithmetic overflow is less than or equal to the set overflow frequency threshold, use the calculation data as the target calculation data.
6. The method according to claim 1, wherein Perform a magnification process on the numerical range to which the calculation data belongs according to a set magnification factor, including at least one of the following: Perform a magnification process on the numerical range to which the model parameters of the first neural network layer belong according to a first magnification factor; Perform a magnification process on the numerical range to which the input data of the first neural network layer belongs according to a second magnification factor.
7. The method according to claim 6, wherein The first magnification factor and / or the second magnification factor are determined according to the frequency of the arithmetic overflow.
8. A quantization processing device for a neural network model, characterized in that, It includes: An input module, configured to: obtain the calculation data required for the first neural network layer in the neural network model, where the calculation data refers to one or more types of data required for the first neural network layer to complete the calculation of this layer. When using the neural network model to process image data or video, the calculation data includes image data or video. The calculation data is a floating-point number, the calculation data has a numerical range to which it belongs, and the numerical range to which the calculation data belongs can be magnified; An overflow prediction module, configured to: determine a quantization mapping relationship for converting the calculation data into a fixed-point number with a set number of digits according to the numerical range to which the calculation data belongs, where the largest floating-point number in the numerical range to which the calculation data belongs is used to map to the maximum value of the fixed-point number with the set number of digits, and the smallest floating-point number in the numerical range to which the calculation data belongs is used to map to the minimum value of the fixed-point number with the set number of digits; predict the frequency of arithmetic overflow in the quantized first neural network layer according to the calculation data and the quantization mapping relationship, where the numerical range to which the calculation data belongs is used to determine the quantization mapping relationship, and the quantization mapping relationship is used to convert the calculation data into an integer data with a set number of digits; A quantization training module, configured to: determine target calculation data for quantizing and training the first neural network layer according to the frequency of the arithmetic overflow and the calculation data, so as to perform quantization training on the first neural network layer; where determining the target calculation data for quantizing and training the first neural network layer according to the frequency of the arithmetic overflow and the calculation data includes: if the frequency of the arithmetic overflow is greater than a set overflow frequency threshold, perform a magnification process on the numerical range to which the calculation data belongs according to a set magnification factor, and re-execute the steps of determining the quantization mapping relationship for converting the calculation data into a fixed-point number with a set number of digits and predicting according to the quantization mapping relationship according to the adjusted numerical range.
9. The device according to claim 8, wherein The overflow prediction module includes: A quantization sub-module, configured to: map the calculation data into the fixed-point number with the set number of digits according to the quantization mapping relationship to obtain the quantized calculation data; A calculation sub-module, configured to infer the calculation process of the quantized first neural network layer according to the quantized calculation data to obtain the intermediate calculation result and the output data of the first neural network layer; and calculate the frequency of the arithmetic overflow according to the intermediate calculation result, the output data, and a set data storage range.
10. An electronic device, characterized in that, It includes: A memory and a processor; The memory is used to store one or more computer instructions; The processor is used to execute the one or more computer instructions for: executing the method according to any one of claims 1-7.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed, it can implement the method according to any one of claims 1-7.
Citation Information
Patent Citations
Neural network based on fixed-point operation
CN108345939A