A quantitative calibration method, computing device and computer-readable storage medium

By using a new quantization difference metric to evaluate quantization performance and optimize the truncation threshold, the problem of accuracy degradation caused by quantization processing is solved, and quantization inference with high accuracy is achieved while reducing the amount of computation and resource consumption.

CN113947177BActive Publication Date: 2025-09-23ANHUI CAMBRICON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010682877.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-15
Publication Date
2025-09-23
Estimated Expiration
2040-12-20

AI Technical Summary

Technical Problem

While quantization processing reduces the amount of neural network computation and saves resources, it also leads to a decrease in inference accuracy. Therefore, it is necessary to optimize the quantization parameters while maintaining a certain level of accuracy.

Method used

A new quantization difference metric is used to evaluate the quantization performance. By dividing the input data into a quantized part and a truncated part, the truncation threshold is optimized to achieve higher quantization inference accuracy.

Benefits of technology

While reducing the amount of computation and saving resources, it maintains the accuracy of quantitative reasoning and improves the computational efficiency and accuracy of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113947177B_ABST
    Figure CN113947177B_ABST
Patent Text Reader

Abstract

This disclosure discloses a quantization calibration method, a computing device, and a computer-readable storage medium. The computing device may be included in a combined processing device, which may also include an interface device and other processing devices. The computing device interacts with the other processing devices to jointly complete user-specified computing operations. The combined processing device may also include a storage device, which is connected to the computing device and the other processing devices, respectively, and is used to store data from the computing device and the other processing devices. The disclosed solution uses a new quantization difference metric to optimize quantization parameters, thereby maintaining a certain level of quantization inference accuracy while achieving various advantages through quantization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of data processing and, more particularly, to a quantitative calibration method, a computing device, and a computer-readable storage medium. Background Art

[0002] With the development of artificial intelligence technology, the amount of computation required for neural network operations is increasing, and the computing resources required are also increasing. Quantizing the data of neural network operations is a good way to reduce the amount of computation and save computing resources.

[0003] However, quantization will reduce the inference accuracy, so quantization calibration is needed to solve the technical problem of reducing the amount of calculation and saving computing resources while still achieving a certain level of quantitative inference accuracy. Summary of the Invention

[0004] In order to at least solve the technical problems mentioned above, the present disclosure proposes a solution in multiple aspects to use a new quantization difference metric to optimize the quantization parameters, so as to achieve the advantages of reducing the amount of calculation, saving computing resources, saving storage resources, and speeding up the processing cycle through quantization while maintaining a certain level of quantization inference accuracy.

[0005] In a first aspect, the present disclosure provides a method for calibration quantization in a neural network, performed by a processor, comprising: receiving a calibration data set; performing quantization processing on the calibration data set using a truncation threshold; determining a quantization total difference metric of the quantization processing; and determining an optimized truncation threshold based on the quantization total difference metric, wherein the optimized truncation threshold is used by the processor to perform quantization processing on data during a neural network operation; wherein the calibration data set is divided into quantized partial data and truncated partial data according to the truncation threshold, and the quantization total difference metric is determined based on the quantization difference metric of the quantized partial data and the quantization difference metric of the truncated partial data.

[0006] In a second aspect, the present disclosure provides a computing device for calibrated quantization in a neural network, comprising: at least one processor; and at least one memory in communication with the at least one processor, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the at least one processor, the at least one processor executes the method described in any embodiment of the first aspect of the present disclosure.

[0007] In a third aspect, the present disclosure provides a computer-readable storage medium having program instructions stored therein. When the program instructions are loaded and executed by a processor, the processor executes the method described in any embodiment of the first aspect of the present disclosure.

[0008] Through the quantization calibration method, computing device and computer-readable storage medium provided above, the disclosed solution uses a new quantization difference metric to evaluate the performance of quantization, thereby optimizing the quantization parameters to achieve the various advantages brought by quantization (such as reducing the amount of calculation, saving computing resources, saving storage resources, speeding up the processing cycle, etc.) while maintaining a certain quantization reasoning accuracy. According to the quantization calibration scheme disclosed in the present invention, the total quantization difference metric can be divided into: a metric of the quantized part data DQ of the input data and a metric of the truncated part data DC of the input data. By dividing the input data into two categories according to the quantization operation to evaluate the quantization difference, the impact of quantization on the effective information of the data can be more accurately characterized, which is conducive to the optimization of the quantization parameters to provide higher quantization reasoning accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0010] Figure 1 An exemplary structural block diagram of a neural network to which the embodiments of the present disclosure can be applied is shown;

[0011] Figure 2 A schematic diagram of a forward propagation process of a hidden layer of a neural network including a quantization operation to which the disclosed embodiments can be applied is shown;

[0012] Figure 3 A schematic diagram of a back-propagation process of a hidden layer of a neural network including a quantization operation to which the disclosed embodiments can be applied is shown;

[0013] Figure 4 A schematic diagram showing a quantization operation to which the disclosed embodiments can be applied;

[0014] Figure 5 A schematic diagram exemplarily shows a quantization error of a quantized portion of data and a truncation error of a truncated portion of data;

[0015] Figure 6 An exemplary flow chart of a quantitative calibration method according to an embodiment of the present disclosure is shown;

[0016] Figure 7 An exemplary logic flow for implementing the quantitative calibration method according to an embodiment of the present disclosure is shown;

[0017] Figure 8 A hardware configuration block diagram of a computing device that can implement the quantitative calibration scheme of the disclosed embodiment is shown;

[0018] Figure 9 A schematic diagram showing the application of the computing device of an embodiment of the present disclosure to an artificial intelligence processor chip;

[0019] Figure 10 A structural diagram showing a combined processing device according to an embodiment of the present disclosure; and

[0020] Figure 11 2 is a schematic diagram showing the structure of a board according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.

[0022] It should be understood that the terms "first," "second," and "third," etc., which may be used in the claims, specification, and drawings of this disclosure, are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0023] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0024] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0025] First, explanations of technical terms that may be used in this disclosure are given.

[0026] Floating-point numbers: The IEEE floating-point standard represents a number in the form of V = (-1)∧sign*mantissa*2∧E. Sign is the sign bit, with 0 representing a positive number and 1 representing a negative number. E represents the exponent, which weights the floating-point number to the power of 2 (which may be a negative power). Mantissa represents the mantissa, a binary fraction in the range of 1 to 2-ε, or 0-ε. The representation of a floating-point number in a computer is divided into three fields, which are used to encode these fields:

[0027] (1) A single sign bit s directly encodes the sign s.

[0028] (2) The k-bit exponent field encodes the exponent, exp = e(k-1)...e(1)e(0).

[0029] (3) The n-bit decimal field mantissa encodes the mantissa. However, the encoding result depends on whether the exponent stage is all zeros.

[0030] Fixed-point numbers are composed of three parts: a shared exponent, a sign, and a mantissa. The shared exponent means that the exponent is shared within a set of real numbers to be quantized; the sign indicates the positive or negative value of the fixed-point number. The mantissa determines the number of significant digits in the fixed-point number, i.e., its precision. Taking an 8-bit fixed-point number as an example, its numerical calculation method is:

[0031] value=(-1) sign ×(mantissa)×2 (exponent-127)

[0032] KL (Kullback–Leibler) divergence: also known as relative entropy, information divergence, or information gain. KL divergence measures the asymmetry between two probability distributions P and Q. It measures the average number of extra bits required to encode samples from P using a code based on Q. Typically, P represents the true distribution of the data, and Q represents the theoretical distribution, model distribution, or an approximation of P.

[0033] Data bit width: how many bits are used to represent the data.

[0034] Quantization: The process of converting high-precision numbers previously expressed in 32-bit or 64-bit into fixed-point numbers that occupy less memory space, generally 16-bit or 8-bit. The process of converting high-precision numbers to fixed-point numbers will cause a certain loss in accuracy.

[0035] The following briefly introduces a neural network environment to which the disclosed embodiments can be applied.

[0036] A neural network (NN) is a mathematical model that mimics the structure and function of biological neural networks. Neural networks perform computations by connecting a large number of neurons. Therefore, a neural network is a computational model composed of a large number of interconnected nodes (or "neurons"). Each node represents a specific output function, called an activation function. Each connection between two neurons represents a weighted value for the signal passing through that connection, called a weight, which acts as the neural network's memory. The output of a neural network varies depending on the connections between neurons, the weights, and the activation function. In a neural network, a neuron is the basic unit. It receives a certain number of inputs and a bias, and when a signal (value) arrives, it is multiplied by a weight. A connection connects a neuron to another neuron in another layer or within the same layer, and the connection is associated with a weight. Furthermore, a bias is an additional input to a neuron that is always 1 and has its own connection weight.

[0037] In practice, if a nonlinear function isn't applied to the neurons in a neural network, the neural network is simply a linear function and is no more powerful than a single neuron. If the output of a neural network is set between 0 and 1, for example, in the example of identifying a cat or dog, an output close to 0 can be considered a cat, and an output close to 1 can be considered a dog. To achieve this goal, activation functions, such as the sigmoid activation function, are introduced into the neural network. All you need to know about this activation function is that its return value is a number between 0 and 1. Therefore, the activation function is used to introduce nonlinearity into the neural network, narrowing the results of the neural network's calculations to a smaller range. The choice of activation function affects the expressive power of the resulting network. Activation functions can take many forms, all of which parameterize a nonlinear function through weights. These weights can be changed to modify the nonlinear function.

[0038] Figure 1 FIG. 1 is an exemplary block diagram showing a neural network 100 to which the embodiments of the present disclosure can be applied. Figure 1 The neural network shown in the figure includes three layers: input layer, hidden layer and output layer. Figure 1 The number of hidden layers shown is 5.

[0039] The leftmost layer of a neural network is called the input layer, and the neurons in the input layer are called input neurons. The input layer is the first layer in a neural network, which receives the required input signals (values) and passes them to the next layer. It generally does not operate on the input signals (values) and has no associated weights and biases. Figure 1 In the neural network shown, there are 4 input signals x1, x2, x3, x4.

[0040] The hidden layer contains neurons (nodes) that are used to apply different transformations to the input data. Figure 1 The neural network shown has five hidden layers. The first hidden layer has four neurons (nodes), the second layer has five neurons, the third layer has six neurons, the fourth layer has four neurons, and the fifth layer has three neurons. Finally, the hidden layers pass the calculated values ​​of the neurons to the output layer. Figure 1 The neural network shown fully connects each neuron in the five hidden layers, meaning that every neuron in each hidden layer is connected to every neuron in the next layer. It should be noted that not every hidden layer in a neural network is fully connected.

[0041] Figure 1 The rightmost layer of the neural network is called the output layer, and the neurons in the output layer are called output neurons. The output layer receives the output from the last hidden layer. Figure 1 In the neural network shown, the output layer has 3 neurons and 3 output signals y1, y2, and y3.

[0042] In practical applications, a large amount of sample data (including input and output) is given in advance to train the initial neural network. After the training is completed, a trained neural network is obtained. This neural network can give a correct output for the input of the future real environment.

[0043] Before we begin discussing the training of a neural network, we need to define the loss function. The loss function is a function that measures the performance of a neural network in performing a specific task. In some embodiments, the loss function can be obtained as follows: during the training of a neural network, for each sample data, it is passed along the neural network to obtain an output value, and then the output value is subtracted from the expected value and then squared. The loss function calculated in this way is the distance between the predicted value and the true value, and the purpose of training the neural network is to reduce the value of this distance or loss function. In some embodiments, the loss function can be expressed as:

[0044]

[0045] In the above formula, y represents the expected value, Refers to the actual result obtained by the neural network for each sample data in the sample data set, and i is the index of each sample data in the sample data set. Represents the expected value y and the actual result The error value between them. m is the number of sample data in the sample data set.

[0046] Take the actual application scenario of cat and dog identification as an example. Assume that a data set consists of pictures of cats and dogs. If the picture is of a dog, the corresponding label is 1, and if the picture is of a cat, the corresponding label is 0. This label corresponds to the expected value y in the above formula. When passing each sample picture to the neural network, the actual goal is to obtain the recognition result through the neural network, that is, whether the animal in the picture is a cat or a dog. In order to calculate the loss function, it is necessary to traverse each sample picture in the sample data set and obtain the actual result corresponding to each sample picture. The loss function is then calculated according to the above definition. If the loss function is large, for example, exceeding a predetermined threshold, it means that the neural network has not been trained well, and the weights need to be further adjusted.

[0047] When training a neural network, the weights are randomly initialized. In most cases, this initialization does not provide good training results. However, during training, it is possible to start with a poorly trained neural network and then train it to produce a highly accurate one.

[0048] The training process of a neural network is divided into two stages. The first stage is the forward signal processing operation (referred to as the forward propagation process in this disclosure), which trains the signal from the input layer through the hidden layer to the output layer. The second stage is the backward gradient operation (referred to as the backward propagation process in this disclosure), which trains the signal from the output layer to the hidden layer and finally to the input layer. The weights and biases of each layer in the neural network are adjusted in sequence according to the gradient.

[0049] During the forward propagation process, input values ​​are fed into the input layer of a neural network. After the corresponding operations are performed by the relevant operators in multiple hidden layers, the output layer of the neural network produces the so-called predicted value. When the input value is provided to the input layer of the neural network, it may not be processed at all, or some necessary preprocessing may be performed depending on the application scenario. Within the hidden layer, the second hidden layer obtains the intermediate predicted value from the first hidden layer, performs calculations and activation operations, and then passes the resulting intermediate predicted value to the next hidden layer. The following layers perform the same operations, and finally the output value is obtained at the output layer of the neural network. After the forward propagation process, an output value, called a predicted value, is typically obtained. To calculate the error, the predicted value can be compared with the actual output value to obtain the corresponding error value.

[0050] During the back propagation process, the chain rule of differential calculus can be used to update the weights of each layer in order to obtain a lower error value relative to the previous one in the next forward propagation process. In the chain rule, the derivative of the error value of the last layer of weights of the corresponding neural network is first calculated. These derivatives are called gradients, and then these gradients are used to calculate the gradient of the penultimate layer in the neural network. Repeat this process until the gradient corresponding to each weight in the neural network is obtained. Finally, the corresponding gradient is subtracted from each weight in the neural network, thereby updating the weight once to achieve the purpose of reducing the error value. Similar to the use of various operators (referred to as forward operators in this disclosure) in the forward propagation process, there are also reverse operators corresponding to the forward operators in the forward propagation process in the corresponding back propagation process. For example, for the convolution operator in the convolution layer, it includes the forward convolution operator in the forward propagation process and the deconvolution operator in the back propagation process.

[0051] For neural networks, fine-tuning involves loading a trained neural network. The fine-tuning process is identical to the training process and consists of two phases: the first is the forward signal processing (referred to in this disclosure as the forward propagation process), and the second is the backward propagation of gradients (referred to in this disclosure as the backward propagation process), which updates the weights of the trained neural network. The difference between training and fine-tuning is that training randomly processes an initialized neural network, training the neural network from scratch, while fine-tuning does not.

[0052] During the training or fine-tuning of a neural network, each time the network undergoes a forward propagation process (forward processing of signals) and a corresponding backpropagation process (backward propagation of errors), the weights in the neural network are updated using the gradient. This is called an iteration. To achieve a neural network with the desired accuracy, the training process requires a very large sample dataset, but it is almost impossible to input all of the sample datasets into a computing device (such as a computer) all at once. Therefore, to solve this problem, the sample dataset is divided into multiple batches and fed into the computer. After each batch of data undergoes the forward propagation process, it undergoes a corresponding backpropagation process (updating the neural network weights). When a complete sample dataset passes through the neural network once and returns a corresponding weight update, this process is called an epoch. In practice, passing a complete dataset through the neural network once is not sufficient; it needs to be passed through the same neural network multiple times, i.e., multiple epochs, to ultimately achieve a neural network with the desired accuracy.

[0053] During the training or fine-tuning of a neural network, users usually hope that the training or fine-tuning speed will be as fast as possible and the accuracy will be as high as possible, but such expectations are usually affected by the data type of the neural network data. In many application scenarios, the data of the neural network is represented by high-precision data formats (such as floating-point numbers). Taking the convolution operation in the forward propagation process and the reverse convolution operation in the backward propagation process as examples, when these two operations are performed on the central processing unit ("CPU") and graphics processing unit ("GPU") of the computing device, in order to ensure data accuracy, almost all inputs, weights and gradients are floating-point type data.

[0054] Taking the floating-point type format as an example of a high-precision data format, according to computer architecture, based on the arithmetic representation rules of floating-point numbers and fixed-point numbers, for floating-point and fixed-point operations of the same length, the floating-point operation calculation mode is more complex, requiring more logic devices to construct the floating-point arithmetic unit. Therefore, in terms of volume, the floating-point arithmetic unit is larger than the fixed-point arithmetic unit. Furthermore, the floating-point arithmetic unit requires more resources to process, resulting in a power consumption difference of orders of magnitude between fixed-point and floating-point operations, which results in a significant difference in computing cost. However, according to experimental findings, fixed-point operations are faster than floating-point operations and the loss of precision is not significant. Therefore, using fixed-point operations to process a large number of neural network operations (such as convolution and fully connected operations) in artificial intelligence chips is a feasible solution. For example, the floating-point data involved in the input, weights, and gradients of the forward convolution, forward fully connected, reverse convolution, and reverse fully connected operators can be quantized and then fixed-point operations can be performed. After the operator operation is completed, the low-precision data can be converted to high-precision data.

[0055] Figure 2 A schematic diagram of the forward propagation process of the hidden layer of a neural network including a quantization operation to which the disclosed embodiments can be applied is shown.

[0056] like Figure 2 As shown, the hidden layers (e.g., convolutional layers and fully connected layers) of the neural network are represented by a fixed-point computing device 250. The activation values ​​210 and weights 220 associated with the fixed-point computing device 250 are typically floating-point data. The activation values ​​210 and weights 220 are quantized to obtain fixed-point activation values ​​230 and weights 240, which are then provided to the fixed-point computing device 250 for fixed-point computation, resulting in a computation result 260 of fixed-point data.

[0057] Depending on the structure of the neural network, the calculation result 260 of the fixed-point calculation device 250 can be provided to the next hidden layer of the neural network as its activation value, or provided to the output layer as an output result. Therefore, the calculation result can be dequantized as needed to obtain a calculation result of floating-point data.

[0058] Figure 3 A schematic diagram of the backpropagation process of a hidden layer of a neural network including quantization operations to which the disclosed embodiments can be applied is shown. As previously described, the forward propagation process forwardly transmits information until an error is generated in the output, and the backpropagation process backpropagates the error information to update the weights.

[0059] like Figure 3 As shown, the floating-point data gradient 310 used in the back-propagation calculation is quantized to obtain the fixed-point data gradient 320. The fixed-point gradient 320 is provided to the fixed-point calculation device 330 of the previous hidden layer of the neural network. Similarly, the calculation of the fixed-point calculation device 330 also requires the corresponding weights and activation values. Figure 3 3 shows the weight 340 and activation value 360 ​​of floating point data, which are quantized into the weight 350 and activation value 370 of fixed point data respectively. It will be understood by those skilled in the art that although Figure 3 3 shows the quantization of the weights 340 and activation values ​​360, but when the fixed-point weights and activation values ​​have been obtained during the forward propagation process, there is no need to re-quantize them.

[0060] The fixed-point computing device 330 performs fixed-point calculations based on the fixed-point gradient 320 provided by the subsequent layer, the current corresponding fixed-point weight 350, and the activation value 370 to calculate the gradient of the corresponding weight and activation value. Next, the fixed-point weight gradient 380 calculated by the fixed-point computing device 330 is dequantized into a floating-point weight gradient 390. Finally, the floating-point weight gradient 390 is used to update the floating-point weight 340 corresponding to the fixed-point computing device 330. For example, the corresponding gradient 390 can be subtracted from the weight 340 to update the weight to achieve the purpose of reducing the error value. The fixed-point computing device 330 can continue to propagate the gradient of the current layer to the previous layer to adjust the parameters of the previous layer.

[0061] Quantization operations are involved in both the forward and back propagation processes mentioned above.

[0062] Figure 4 Schematic diagram showing the quantization operation to which the disclosed embodiment can be applied. Figure 4 In the example shown, 32-bit floating-point data is quantized into n-bit fixed-point data, where n is the fixed-point number bit width. Figure 4 The points on the upper horizontal line represent the floating-point data to be quantized, and the points on the lower horizontal line represent the fixed-point data after quantization.

[0063] Figure 4 The number domain of the data to be quantized is asymmetrically distributed relative to "0". In this quantization operation, there is a threshold T that maps ±T to ±(2 n-1-1). From Figure 4 As can be seen, the floating-point data outside the threshold ±T is directly mapped to the fixed-point number ±(2 n-1 -1). For example, Figure 4 The three points on the horizontal line above that are less than -T are directly mapped to -(2 n-1 -1). Floating point data within the ±T threshold range can be proportionally mapped to ±(2 n-1 -1). This mapping relationship is saturated and asymmetric.

[0064] While quantization can reduce computational complexity and conserve computing resources, it also reduces inference accuracy. Therefore, the technical challenges addressed by the disclosed embodiments are how to replace floating-point units with fixed-point units to achieve the speed of fixed-point operations, thereby increasing the peak computing power of AI processor chips while maintaining the required floating-point accuracy.

[0065] Based on the description of the above technical problems, one of the characteristics of neural networks is that they have a high tolerance for input noise. If you consider identifying objects in photos, the neural network can ignore the main noise and focus on important similarities. This function means that the neural network can use low-precision calculations as a noise source and still produce accurate prediction results in a numerical format that accommodates less information. In the description below, the error caused by quantization is understood from the perspective of noise, that is, the quantization error can be understood as noise that is correlated with the original signal. In this sense, the quantization error is sometimes also called quantization noise, and the two can be used interchangeably. However, those skilled in the art should understand that the quantization noise in this article is different from white noise that is independent of the signal, such as Gaussian noise. For Figure 4 For the quantization operation shown in FIG, the above technical problem is converted into the need to find the optimal threshold T so as to minimize the loss of accuracy after quantization.

[0066] In the noise-based quantization calibration scheme of the disclosed embodiment, it is proposed to use a new quantization difference metric to evaluate the performance of quantization, thereby optimizing the quantization parameters so as to achieve the various advantages brought by quantization (such as reducing the amount of calculation, saving computing resources, saving storage resources, speeding up the processing cycle, etc.) while still maintaining the required quantization inference accuracy.

[0067] According to the disclosed noise-based quantization calibration scheme, the total quantization difference metric can be divided into two types: a metric for the quantized portion of the input data and a metric for the truncated portion of the input data. By classifying the input data into two categories based on the quantization operation to evaluate the quantization difference, the impact of quantization on the effective information of the data can be more accurately characterized, thereby facilitating the optimization of quantization parameters and providing higher quantization inference accuracy.

[0068] To facilitate understanding of the embodiments of the present disclosure, the quantitative total difference metric used in the embodiments of the present disclosure is first explained below.

[0069] In some embodiments, the input data (eg, calibration data) may be represented as:

[0070] D=[x1,x2,…,x N ],D∈R N (2)

[0071] Wherein, N is the number of data in data D, and R represents the real number domain.

[0072] When Figure 4 When the quantization operation shown in the figure quantizes the input data, the data exceeding the threshold ±T is directly mapped to the fixed-point number ±(2 n-1 -1). Therefore, in the disclosed embodiment, the input data D is divided into the quantized portion data DQ and the truncated portion data DC according to the truncation threshold T. Accordingly, the quantized total difference metric is also divided into: a metric for the quantized portion data DQ of the input data D and a metric for the truncated portion data DC of the input data D.

[0073] Figure 5 A schematic diagram exemplarily shows a quantization error of a quantized portion of data and a truncation error of a truncation portion of data. Figure 5 The horizontal axis is the value x of the input data, and the vertical axis is the frequency y of the corresponding value. Figure 5 It can be seen that when the quantized data DQ is within the threshold T, each data is quantized into a fixed-point data, so the quantization error is small. In contrast, when the truncated data DC is outside the threshold T, no matter how large the truncated data DC is, it is uniformly quantized into the fixed-point data corresponding to the threshold T. For example, 2 n-1 -1. Therefore, the truncation error is large and widely distributed. Thus, the quantization errors of the quantized data and the truncated data have different manifestations. It should be noted that in the KL divergence calibration method, the histogram of the input data is usually used to evaluate the quantization error. In the embodiments disclosed herein, the input data is directly used without employing any form of histogram.

[0074] In the embodiments of the present disclosure, by performing quantization difference evaluation on the quantized partial data DQ and the truncated partial data DC respectively, the impact of quantization on the effective information of the data can be more accurately characterized, which is beneficial to the optimization of the quantization parameters to provide higher quantization inference accuracy.

[0075] In some embodiments, the quantized partial data DQ and the truncated partial data DC may be expressed as:

[0076]

[0077] DC=[x|Abs(x)≥T,x∈D] (4)

[0078] Among them, Abs() means taking the absolute value, and n is the bit width of the quantized fixed-point number.

[0079] In this embodiment, the Because this part of the data has little impact on quantization, but through experimental analysis, it has a greater impact on the quantitative difference measurement of the embodiment of the disclosure, so this part of the data is removed.

[0080] In the embodiments of the present disclosure, corresponding quantization difference metrics are constructed for the quantized partial data DQ and the truncated partial data DC, respectively, such as the quantization difference metric DistQ of the quantized partial data DQ and the quantization difference metric DistC of the truncated partial data DC. Subsequently, the quantization total difference metric Dist(D, T) can be expressed as a function of the quantization difference metrics DistQ and DistC. Various functions can be constructed to characterize the relationship between the quantization total difference metric Dist(D, T) and the quantization difference metrics DistQ and DistC.

[0081] In some embodiments, the quantitative total difference metric Dist(D, T) may be calculated as follows:

[0082] Dist(D,T)=DistQ+DistC (5)

[0083] In some embodiments, when constructing a quantization difference metric for the quantized partial data DQ and the truncated partial data DC, two aspects can be considered: the magnitude of the quantization noise and the correlation of the quantization noise with the input data. On the one hand, the magnitude of the quantization noise reflects the difference in the absolute value of the quantization error; on the other hand, the correlation of the quantization noise with the input data considers the relationship between the different manifestations of the quantization error between the quantized partial data and the truncated partial data and the distribution of the input data relative to the optimal truncation threshold T.

[0084] Specifically, the quantization difference measure DistQ of the quantized partial data DQ can be expressed as a function of the amplitude of the quantization noise of the quantized partial data DQ and the correlation coefficient between the quantization noise and the input data; and / or the quantization difference measure DistC of the truncated partial data DC can be expressed as a function of the amplitude of the quantization noise of the truncated partial data DC and the correlation coefficient between the quantization noise and the input data. Various functions can be constructed to characterize the relationship between the quantization difference measure and the amplitude of the quantization noise and the correlation coefficient between the quantization noise and the input data.

[0085] In some embodiments, the amplitude of the quantization noise may be weighted using a correlation coefficient, for example, by calculating the quantization difference metric DistQ and the quantization difference metric DistC according to the following formulas:

[0086] DistQ=(1+EQ)×AQ (6)

[0087] DistC=(1+EC)×AC (7)

[0088] The quantization noise amplitude AQ of the quantized partial data DQ and the quantization noise amplitude AC of the truncated partial data DC in the above formulas (6) and (7) can be calculated as follows:

[0089]

[0090]

[0091] Wherein, Quantize(x, T) is a function that quantizes the data x with T as the maximum value. Those skilled in the art will appreciate that the disclosed embodiments can be applied to various quantization methods. The purpose of the disclosed embodiments is to find the optimal quantization parameter that conforms to the currently used quantization method, that is, the optimal truncation threshold. Depending on the quantization method used, Quantize(x, T) can have different expressions. In one example, the data can be quantized according to the following formula:

[0092]

[0093] Where s is the point position parameter, round is the rounding operation, ceil is the ceiling operation, Ix is the n-bit binary representation of the data x after quantization, and Fx is the floating-point value of the data x before quantization.

[0094] The correlation coefficient EQ of the quantization noise of the quantized partial data DQ and the input data and the correlation coefficient EC of the quantization noise of the truncated partial data DC and the input data in the above formulas (6) and (7) can be calculated as follows:

[0095]

[0096]

[0097] The above describes the quantization total difference metric used in the embodiments of the present disclosure. From the above description, it can be seen that by dividing the input data into two categories (quantized partial data and truncated partial data) according to the quantization operation to evaluate the quantization total difference metric, the impact of quantization on the effective information of the data can be more accurately characterized, which is beneficial to the optimization of the quantization parameters to provide higher quantization reasoning accuracy. Furthermore, in some embodiments, the quantization difference metric of each portion of data takes into account two aspects: the amplitude of the quantization noise and the correlation between the quantization noise and the input data. Thus, the impact of quantization on the effective information of the data can be further accurately characterized. The quantization total difference metric Dist(D,T) described above can be used to calibrate the quantization noise of the computational data in the neural network.

[0098] Figure 6 FIG. 6 is an exemplary flow chart of a quantization noise calibration method 600 according to an embodiment of the present disclosure. The quantization noise calibration method 600 may be executed by a processor, for example. Figure 6 The illustrated technical solution determines calibrated / optimized quantization parameters (e.g., a cutoff threshold T), which are used by an artificial intelligence processor to quantize data (e.g., activation values, weights, gradients, etc.) during neural network operations, thereby determining quantized fixed-point data. The quantized fixed-point data can be used by the artificial intelligence processor for training, fine-tuning, or inference of the neural network.

[0099] like Figure 6 As shown, in step S610, the processor receives input data D. The input data D is, for example, a calibration data set or a sample data set prepared for calibrating quantization noise. The input data D can be received from a collaborative processing circuit in a neural network environment applied in the embodiments of the present disclosure.

[0100] If the input data is large, the calibration data set can be provided to the processor in batches.

[0101] For example, in some examples, the calibration dataset can be represented as:

[0102] D=[D1,D2,…,D B ],D i ∈R N×S ,i∈[1…B] (13)

[0103] Among them, B is the number of data batches; N is the data batch size, that is, the number of data samples in each data batch; S is the number of data in a single data sample; R represents the real number domain.

[0104] Next, in step S620, the processor quantizes the input data D using the cutoff threshold. The input data can be quantized using a variety of quantization methods. For example, the aforementioned formula (10) can be used for quantization, which will not be described in detail here.

[0105] Then, in step S630, the processor determines the quantization total difference metric of the quantization processing performed in step S620, wherein the input data is divided into quantized partial data and truncated partial data according to the truncation threshold, and the quantization total difference metric is determined based on the quantization difference metric of the quantized partial data and the quantization difference metric of the truncated partial data.

[0106] Furthermore, in some embodiments, the quantization difference metric of the quantized portion of data and / or the quantization difference metric of the truncated portion of data may be determined based on at least two factors: the amplitude of the quantization noise; and the correlation coefficient between the quantization noise and the corresponding quantized data.

[0107] Specifically, in some embodiments, the input data may be divided into a quantized data portion DQ and a truncated data portion DC, for example, with reference to the aforementioned formulas (3) and (4). Then, for example, with reference to the aforementioned formulas (8) and (9), the quantization noise amplitudes AQ and AC of the quantized data portion DQ and the truncated data portion DC may be calculated, respectively; and, for example, with reference to the aforementioned formulas (11) and (12), the correlation coefficients EQ and EC between the quantization noise of the quantized data portion DQ and the truncated data portion DC and the corresponding quantized data may be calculated, respectively.

[0108] Next, for example, referring to the aforementioned formulas (6) and (7), the quantization difference metrics DistQ and DistC of the quantized data portion DQ and the truncated data portion DC are calculated respectively. Finally, for example, referring to the aforementioned formula (5), the quantization total difference metric is calculated.

[0109] continue Figure 6 , the method 600 may proceed to step S640, where the processor determines an optimized truncation threshold based on the quantized total difference metric determined in step S630. In this step, the processor may select the truncation threshold that minimizes the quantized total difference metric as the calibrated / optimized truncation threshold.

[0110] In some embodiments, when the input data or calibration dataset includes multiple data batches, the processor may determine a corresponding total quantitative difference metric for each batch of data. The total quantitative difference metric for the entire calibration dataset may then be considered as a whole to determine the total quantitative difference metric for the entire calibration dataset, thereby determining a calibration / optimization cutoff threshold. In one example, the total quantitative difference metric for the calibration dataset may be the sum of the total quantitative difference metrics for each batch.

[0111] Reference above Figure 6 An exemplary process of the quantization noise calibration method of the disclosed embodiment is described. In actual operation, a search method can be used to determine the calibration / optimization truncation threshold. Specifically, by searching and comparing the corresponding quantization total difference metric Dist(D, Tc) of each candidate truncation threshold Tc within the possible range of the truncation threshold (herein referred to as the search space) for a given calibration data set D, the candidate truncation threshold Tc that optimizes the quantization total difference metric is determined as the calibration / optimization truncation threshold.

[0112] Figure 7 An exemplary logic flow 700 for implementing the quantization noise calibration method according to an embodiment of the present disclosure is shown. The flow 700 may be executed by a processor for a calibration data set, for example.

[0113] like Figure 7 As shown, in step S710, the calibration data set is quantized using multiple candidate truncation thresholds Tc in the truncation threshold search space.

[0114] In some embodiments, the search space for the truncation threshold can be determined based at least on the maximum value of the calibration data set. The search space can be set to, for example, (0, max], where max is the maximum value of the calibration data set. When calibration is performed using the calibration data set in batches, max can be initialized to max = max(D1), where max(D1) is the maximum value of the first calibration data batch.

[0115] The number of candidate truncation thresholds Tc in the search space can be referred to as the search precision M. The search precision M can be preset. In some examples, the search precision M can be set to 2048. In other examples, the search precision M can be set to 64. The search precision determines the search interval. Thus, the jth candidate truncation threshold Tc in the search space can be determined at least in part based on the preset search precision M as follows:

[0116]

[0117] After determining the candidate truncation threshold Tc, the input data can be quantized using various quantization methods. For example, the aforementioned formula (10) can be used for quantization.

[0118] Next, in step S720, for each candidate truncation threshold Tc, the corresponding quantization total difference metric Dist(D, Tc) of the quantization process is determined. Specifically, the following sub-steps may be included:

[0119] In sub-step S721, the calibration data set D is divided into the quantized data portion DQ and the truncated data portion DC according to the candidate truncation threshold Tc, with reference to the above formulas (3) and (4). In this embodiment, formulas (3) and (4) can be adjusted to:

[0120]

[0121] DC=[x|Abs(x)≥Tc,x∈D],

[0122] Where n is the bit width of the quantized data after quantization processing.

[0123] Sub-step S722, respectively determine the quantization difference metric DistQ of the quantized partial data DQ and the quantization difference metric DistC of the truncated partial data DC. For example, the quantization difference metric DistQ and the quantization difference metric DistC can be determined with reference to the aforementioned formulas (6) and (7):

[0124] DistQ=(1+EQ)×AQ,

[0125] DistC=(1+EC)×AC,

[0126] Wherein AQ represents the amplitude of the quantization noise of the quantized partial data DQ, EQ represents the correlation coefficient between the quantization noise of the quantized partial data DQ and the quantized partial data DQ, AC represents the amplitude of the quantization noise of the truncated partial data DC, and EC represents the correlation coefficient between the quantization noise of the truncated partial data DC and the truncated partial data DC.

[0127] Furthermore, the quantization noise amplitudes AQ and AC of the quantized data portion DQ and the truncated data portion DC can be calculated respectively with reference to the aforementioned formulas (8) and (9); and the correlation coefficients EQ and EC of the quantization noise of the quantized data portion DQ and the truncated data portion DC with the corresponding quantized data can be calculated respectively with reference to the aforementioned formulas (11) and (12). In this embodiment, the aforementioned formulas can be adjusted to:

[0128]

[0129]

[0130]

[0131]

[0132] Where N is the number of data in the current calibration data set D, and Quantize(x,Tc) is a function that quantizes the data x with Tc as the maximum value.

[0133] Sub-step S723: Determine the corresponding quantized total difference metric Dist(D, Tc) based on the quantized difference metrics DistQ and DistC calculated in sub-step S722. In some embodiments, for example, the corresponding quantized total difference metric Dist(D, Tc) can be determined according to the following formula:

[0134] Dist(D,Tc)=DistQ+DistC.

[0135] Finally, in step S730 , the candidate truncation threshold Tc that minimizes the quantized total difference measure Dist(D, Tc) is selected from the plurality of candidate truncation thresholds Tc as the calibrated / optimized truncation threshold T.

[0136] In some embodiments, when the calibration dataset includes multiple data batches, the processor may determine a corresponding total quantitative difference metric for each batch of data. The total quantitative difference metric for the entire calibration dataset may then be considered as a whole to determine the total quantitative difference metric for the entire calibration dataset, thereby determining a calibration / optimization cutoff threshold. In one example, the total quantitative difference metric for the calibration dataset may be the sum of the total quantitative difference metrics for each batch. The above calculation may be expressed as:

[0137]

[0138] Where B is the number of data batches.

[0139] The quantization noise calibration scheme according to the embodiment of the present disclosure has been described above with reference to the flowchart.

[0140] The inventors conducted experiments comparing the aforementioned KL divergence calibration method with the quantization noise calibration method of the present disclosure on the classification models MobileNet V1, MobileNet V2, ResNet 50 V1.5, and DenseNet 121, as well as the translation model GNMT. The experiments employed different batch sizes B, batch numbers N, and search accuracies M.

[0141] Experimental results show that the quantization noise calibration method of the disclosed embodiment achieves performance similar to KL on MobileNet V1; exceeds KL on MobileNet V2 and GNMT; and is lower than KL on ResNet 50 and DenseNet 121. In summary, the disclosed embodiment provides a new quantization noise calibration scheme that can calibrate quantization parameters (e.g., truncation thresholds) to achieve various advantages brought by quantization (such as reducing the amount of computation, saving computing resources, saving storage resources, speeding up processing cycles, etc.) while maintaining a certain degree of quantization inference accuracy. The quantization noise calibration scheme of the disclosed embodiment is particularly suitable for neural networks with more concentrated distribution of quantized data and more difficult to quantize, such as the MobileNet series models and the GNMT model.

[0142] Figure 8 FIG. 8 is a block diagram showing a hardware configuration of a computing device 800 that can implement the quantization noise calibration scheme of the disclosed embodiment. Figure 8 As shown, the computing device 800 may include a processor 810 and a memory 820. Figure 8 In the computing device 800, only the components related to this embodiment are shown. Therefore, it is obvious to those skilled in the art that the computing device 800 may also include Figure 8 The components shown in the figure are different from the common components. For example: fixed-point arithmetic units.

[0143] The computing device 800 may correspond to a computing device having various processing functions, such as functions for generating a neural network, training or learning a neural network, quantizing a floating-point neural network to a fixed-point neural network, or retraining a neural network. For example, the computing device 800 may be implemented as various types of devices, such as a personal computer (PC), a server device, a mobile device, etc.

[0144] The processor 810 controls all functions of the computing device 800. For example, the processor 810 controls all functions of the computing device 800 by executing a program stored in the memory 820 on the computing device 800. The processor 810 may be implemented by a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), an artificial intelligence processor chip (IPU), etc. provided in the computing device 800. However, the present disclosure is not limited thereto.

[0145] In some embodiments, the processor 810 may include an input / output (I / O) unit 811 and a computing unit 812. The I / O unit 811 may be configured to receive various data, such as a calibration data set. The computing unit 812 may be configured to perform quantization processing on the calibration data set received via the I / O unit 811 using a truncation threshold, determine a quantized total difference metric for the quantization processing, and determine an optimized truncation threshold based on the quantized total difference metric. The optimized truncation threshold may be output by the I / O unit 811, for example. The output data may be provided to the memory 820 for reading and use by other devices (not shown), or may be directly provided to other devices for use.

[0146] The memory 820 is hardware for storing various data processed in the computing device 800. For example, the memory 820 can store processed data and data to be processed in the computing device 800. The memory 820 can store data sets involved in the neural network operation process that has been processed or is to be processed by the processor 810, such as data of an untrained initial neural network, intermediate data of a neural network generated during the training process, data of a neural network that has completed all training, data of a quantized neural network, etc. In addition, the memory 820 can store applications, drivers, etc. to be driven by the computing device 800. For example, the memory 820 can store various programs related to the training algorithm, quantization algorithm, calibration algorithm, etc. of the neural network to be executed by the processor 810. The memory 820 can be a DRAM, but the present disclosure is not limited thereto. The memory 820 can include at least one of a volatile memory and a non-volatile memory. The non-volatile memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a phase-change RAM (PRAM), a magnetic RAM (MRAM), a resistive RAM (RRAM), a ferroelectric RAM (FRAM), etc. The volatile memory may include a dynamic RAM (DRAM), a static RAM (SRAM), a synchronous DRAM (SDRAM), a PRAM, an MRAM, an RRAM, a ferroelectric RAM (FeRAM), etc. In an embodiment, the memory 820 may include at least one of a hard disk drive (HDD), a solid-state drive (SSD), a high-density flash memory (CF), a secure digital (SD) card, a micro secure digital (Micro-SD) card, a mini secure digital (Mini-SD) card, an extreme digital (xD) card, caches, or a memory stick.

[0147] Processor 810 can generate a trained neural network by repeatedly training (learning) a given initial neural network. In this state, to ensure the processing accuracy of the neural network, the parameters of the initial neural network are in a high-precision data representation format, such as a data representation format with 32-bit floating-point precision. The parameters can include various types of data input to / output from the neural network, such as the neural network's input / output neurons, weights, biases, etc. Compared to fixed-point operations, floating-point operations require a relatively large number of operations and relatively frequent memory accesses. Specifically, the majority of operations required for neural network processing are known to be various convolution operations. Therefore, in mobile devices with relatively low processing power (such as smartphones, tablets, wearable devices, embedded devices, etc.), high-precision neural network data operations can lead to underutilization of mobile device resources. Consequently, in order to drive neural network operations within an acceptable precision loss range and minimize the amount of operations in such devices, the high-precision data involved in the neural network operations can be quantized and converted into low-precision fixed-point numbers.

[0148] Taking into account the processing performance of devices such as mobile devices and embedded devices that deploy neural networks, the computing device 800 performs quantization to convert the parameters of the trained neural network into fixed-point type with a specific number of bits, and the computing device 800 sends the corresponding quantization parameters (e.g., truncation threshold) to the device that deploys the neural network, so that when the artificial intelligence processor chip performs training, fine-tuning, and other operations, fixed-point operations are performed. The device that deploys the neural network can be an autonomous vehicle, a robot, a smart phone, a tablet device, an augmented reality (AR) device, an Internet of Things (IoT) device, etc. that performs speech recognition, image recognition, etc. by using the neural network, but the present disclosure is not limited thereto.

[0149] The processor 810 obtains data from the memory 820 during the neural network operation process. The data includes at least one of neurons, weights, biases and gradients. Figure 6-Figure 7 The technical solution shown determines a corresponding truncation threshold and uses it to quantize the target data during the neural network operation. The quantized data is then used to perform neural network operations. These operations include, but are not limited to, training, fine-tuning, and inference.

[0150] In summary, the specific functions implemented by the memory 820 and processor 810 of the computing device 800 provided in the embodiments of this specification can be interpreted in comparison with the aforementioned embodiments in this specification, and can achieve the technical effects of the aforementioned embodiments, so they will not be repeated here.

[0151] In this embodiment, the processor 810 can be implemented in any suitable manner. For example, the processor 810 can take the form of a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, an embedded microcontroller, etc.

[0152] Figure 9 The following is a schematic diagram showing the application of the computing device for quantization noise calibration of a neural network according to an embodiment of the present disclosure to an artificial intelligence processor chip. Figure 9 As described above, in a computing device 800 such as a PC or server, the processor 810 performs a quantization operation to quantize the floating-point data involved in the neural network operation process into fixed-point numbers. The fixed-point operator 922 on the artificial intelligence processor chip 920 uses the fixed-point numbers obtained by quantization to perform training, fine-tuning, or inference. The artificial intelligence processor chip is a dedicated hardware for driving neural networks. Since the artificial intelligence processor chip is implemented with relatively low power or performance, the present technical solution uses low-precision fixed-point numbers to implement neural network operations. Compared with high-precision data, the memory bandwidth required to read low-precision fixed-point numbers is smaller, and the caches of the artificial intelligence processor chip can be better used to avoid memory access bottlenecks. At the same time, when executing SIMD instructions on the artificial intelligence processor chip, more calculations are achieved within one clock cycle, achieving faster execution of neural network operations.

[0153] Furthermore, when comparing fixed-point and high-precision data operations of the same length, especially when comparing fixed-point and floating-point operations, it is clear that floating-point operations have a more complex calculation model and require more logic devices to construct a floating-point unit. Therefore, floating-point units are physically larger than fixed-point units. Furthermore, floating-point units consume more processing resources, resulting in a power consumption difference of orders of magnitude between fixed-point and floating-point operations.

[0154] In summary, the disclosed embodiments enable the replacement of floating-point arithmetic units (FPUs) on AI processor chips with fixed-point units (FPUs), resulting in lower power consumption for AI processor chips. This is particularly important for mobile devices.

[0155] In the embodiments of the present disclosure, the artificial intelligence processor chip may correspond to, for example, a neural processing unit (NPU), a tensor processing unit (TPU), a neural engine, etc., which are dedicated chips for driving neural networks, but the present disclosure is not limited thereto.

[0156] In the embodiments of the present disclosure, the artificial intelligence processor chip can be implemented in a separate device independent of the computing device 800, or the computing device 800 can be implemented as a functional module of the artificial intelligence processor chip. However, the present disclosure is not limited thereto.

[0157] In the embodiment of the present disclosure, the operating system of a general-purpose processor (such as a CPU) generates instructions based on the embodiment of the present disclosure, and sends the generated instructions to an artificial intelligence processor chip (such as a GPU), which executes the instruction operations to implement the quantization noise calibration process and the quantization process of the neural network. There is also an application scenario in which the general-purpose processor directly determines the corresponding truncation threshold based on the embodiment of the present disclosure, and the general-purpose processor directly quantizes the corresponding target data according to the truncation threshold, and the artificial intelligence processor chip uses the quantized data to perform fixed-point operations. What's more, the general-purpose processor (such as a CPU) and the artificial intelligence processor chip (such as a GPU) are pipelined, and the operating system of the general-purpose processor (such as a CPU) generates instructions based on the embodiment of the present disclosure, and while copying the target data, the artificial intelligence processor chip (such as a GPU) performs neural network operations, so that certain time consumption can be hidden. However, the present disclosure is not limited to this.

[0158] In an embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the above-mentioned quantization noise calibration method in the neural network.

[0159] As can be seen from the above, during the neural network operation process, the disclosed embodiment is used to determine the truncation threshold during quantization. The truncation threshold is used by the artificial intelligence processor to quantize the data in the neural network operation process, converting high-precision data into low-precision fixed-point numbers, which can reduce the size of all data storage spaces involved in the neural network operation process. For example: converting float32 to fix8 can reduce the model parameters by 4 times. Since the data storage space becomes smaller, the neural network uses a smaller space when deployed, so that the on-chip memory on the artificial intelligence processor chip can accommodate more data, reducing the artificial intelligence processor chip's memory access data and improving computing performance.

[0160] Figure 10 FIG. 1 is a structural diagram showing a combined processing device 1000 according to an embodiment of the present disclosure. Figure 10 As shown in FIG, the combined processing device 1000 includes a computing processing device 1002, an interface device 1004, other processing devices 1006, and a storage device 1008. According to different application scenarios, the computing processing device may include one or more computing devices 1010, which may be configured as Figure 8 The computing device 800 shown is used to execute the Figure 6-7 The described operation.

[0161] In various embodiments, the computing and processing device of the present disclosure may be configured to perform user-specified operations. In exemplary applications, the computing and processing device may be implemented as a single-core artificial intelligence processor or a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the computing and processing device may be implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core. When multiple computing devices are implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core, the computing and processing device of the present disclosure may be considered to have a single-core structure or a homogeneous multi-core structure.

[0162] In exemplary operation, the computing processing device of the present disclosure can interact with other processing devices through interface means, to jointly complete the operation specified by the user. Depending on the difference in implementation, the other processing devices of the present disclosure may include one or more types of processors in general and / or special processors such as central processing unit (Central Processing Unit, CPU), graphics processing unit (Graphics Processing Unit, GPU), artificial intelligence processor. These processors may include but are not limited to digital signal processor (Digital Signal Processor, DSP), application specific integrated circuit (Application Specific Integrated Circuit, ASIC), field programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As previously mentioned, only with respect to the computing processing device of the present disclosure, it can be regarded as having a single-core structure or a homogeneous multi-core structure. However, when the computing processing device and other processing devices are considered together, the two can be regarded as forming a heterogeneous multi-core structure.

[0163] In one or more embodiments, the other processing device may serve as an interface between the computing device disclosed herein (which may be embodied as an artificial intelligence computing device such as a neural network computing device) and external data and control, performing basic control including but not limited to data transfer, starting and / or stopping the computing device, and so on. In other embodiments, the other processing device may also collaborate with the computing device to jointly complete computing tasks.

[0164] In one or more embodiments, the interface device can be used to transmit data and control instructions between the computing and processing device and other processing devices. For example, the computing and processing device can obtain input data from other processing devices via the interface device and write it to the storage device (or memory) on the computing and processing device chip. Furthermore, the computing and processing device can obtain control instructions from other processing devices via the interface device and write them to the control cache on the computing and processing device chip. Alternatively or optionally, the interface device can also read data from the storage device of the computing and processing device and transmit it to other processing devices.

[0165] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is connected to the computing processing device and the other processing device, respectively. In one or more embodiments, the storage device may be used to store data of the computing processing device and / or the other processing device. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the computing processing device or other processing device.

[0166] In some embodiments, the present disclosure further discloses a chip (e.g. Figure 11 In one implementation, the chip is a system on chip (SoC) and integrates one or more components such as Figure 7 The chip can be connected to the external interface device (such as Figure 11 The external interface device 1106 shown in the figure is connected to other related components. The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card or a wifi interface. In some application scenarios, other processing units (such as video codecs) and / or interface modules (such as DRAM interfaces) can be integrated on the chip. In some embodiments, the present disclosure also discloses a chip packaging structure, which includes the above-mentioned chip. In some embodiments, the present disclosure also discloses a board card, which includes the above-mentioned chip packaging structure. The following will be combined with Figure 11 The board is described in detail.

[0167] Figure 11 FIG. 1 is a schematic diagram showing the structure of a board 1100 according to an embodiment of the present disclosure. Figure 11As shown in , the board includes a storage device 1104 for storing data, which includes one or more storage units 1110. The storage device can be connected and data can be transmitted with the control device 1108 and the chip 1102 described above by means of, for example, a bus. Furthermore, the board also includes an external interface device 1106, which is configured for data relay or transfer function between the chip (or the chip in the chip packaging structure) and the external device 1112 (such as a server or computer, etc.). For example, the data to be processed can be passed from the external device to the chip through the external interface device. For another example, the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms, for example, it can adopt a standard PCIE interface, etc.

[0168] In one or more embodiments, the control device in the disclosed board can be configured to regulate the state of the chip. To this end, in one application scenario, the control device can include a microcontroller unit (MCU) for regulating the working state of the chip.

[0169] According to the above combination Figure 10 and Figure 11 Based on the description, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which may include one or more of the above-mentioned boards, one or more of the above-mentioned chips and / or one or more of the above-mentioned combined processing devices.

[0170] According to different application scenarios, the electronic devices or devices disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, driving recorders, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, automatic driving terminals, vehicles, household appliances, and / or medical equipment. The vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs. The electronic devices or devices disclosed herein may also be applied to the Internet, Internet of Things, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, medical care and other fields. Furthermore, the electronic devices or devices disclosed herein may also be used in cloud, edge, terminal and other application scenarios related to artificial intelligence, big data and / or cloud computing. In one or more embodiments, electronic devices or apparatuses with high computing power according to the disclosed solution can be applied to cloud devices (such as cloud servers), while electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or edge device are compatible with each other, so that according to the hardware information of the terminal device and / or edge device, appropriate hardware resources can be matched from the hardware resources of the cloud device to simulate the hardware resources of the terminal device and / or edge device, so as to complete the unified management, scheduling and collaborative work of end-to-end or cloud-edge-to-end.

[0171] It should be noted that, for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art will understand that the scheme of the present disclosure is not limited by the order of the actions described. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art will understand that some of the steps therein can be performed in other orders or simultaneously. Further, those skilled in the art will understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved therein are not necessarily necessary for the implementation of one or more schemes of the present disclosure. In addition, depending on the different schemes, the description of some embodiments of the present disclosure also has different emphases. In view of this, those skilled in the art will understand that the parts that are not described in detail in a certain embodiment of the present disclosure may also refer to the relevant descriptions of other embodiments.

[0172] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented in other ways not disclosed herein. For example, with respect to the various units in the electronic device or device embodiments described above, this document divides them based on the consideration of logical functions, and there may be other ways of division in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. With respect to the connection relationship between different units or components, the connection discussed above in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.

[0173] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network elements. In addition, according to actual needs, some or all of the units may be selected to achieve the purpose of the solution described in the embodiments of this disclosure. In addition, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically separately.

[0174] In some implementation scenarios, the above-mentioned integrated unit can be implemented in the form of a software program module. If implemented in the form of a software program module and sold or used as an independent product, the integrated unit can be stored in a computer-readable memory. Based on this, when the scheme of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in a memory, which may include several instructions to enable a computer device (such as a personal computer, a server or a network device, etc.) to perform some or all of the steps of the method described in the embodiment of the present disclosure. The aforementioned memory may include, but is not limited to, various media that can store program code, such as a USB flash drive, a flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0175] In some other implementation scenarios, the above-mentioned integrated unit can also be implemented in the form of hardware, that is, a specific hardware circuit, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include but is not limited to physical devices, and the physical devices may include but are not limited to devices such as transistors or memristors. In view of this, the various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM and RAM, etc.

[0176] The foregoing content can be better understood in accordance with the following terms:

[0177] Clause 1. A method, performed by a processor, for calibrating quantization noise in a neural network, comprising:

[0178] receiving a calibration data set;

[0179] quantizing the calibration data set using a cutoff threshold;

[0180] determining a quantitative total difference measure for the quantization process; and

[0181] Determining an optimized truncation threshold based on the quantized total difference metric, wherein the optimized truncation threshold is used for quantizing data in a neural network operation process by an artificial intelligence processor;

[0182] The calibration dataset is divided into quantized partial data and truncated partial data according to the truncation threshold, and the quantized total difference metric is determined based on the quantized difference metric of the quantized partial data and the quantized difference metric of the truncated partial data.

[0183] Clause 2. The method of clause 1, wherein the quantized difference metric of the quantized portion of data and / or the quantized difference metric of the truncated portion of data is determined based on at least two of the following factors:

[0184] the magnitude of the quantization noise; and

[0185] The correlation coefficient between the quantization noise and the corresponding quantized data.

[0186] Clause 3. The method according to any one of clauses 1-2, wherein quantizing the calibration data set using a truncation threshold comprises:

[0187] The calibration data set is quantized using a plurality of candidate truncation thresholds in a truncation threshold search space.

[0188] Clause 4. The method of clause 3, wherein determining a quantized total difference metric for the quantization process comprises:

[0189] For each candidate truncation threshold Tc, the calibration data set D is divided into quantized data DQ and truncated data DC according to the following formula:

[0190]

[0191] DC=[x|Abs(x)≥Tc,x∈D],

[0192] Wherein n is the bit width of the quantized data after the quantization process;

[0193] respectively determining a quantization difference metric DistQ of the quantized partial data DQ and a quantization difference metric DistC of the truncated partial data DC; and

[0194] A corresponding quantized total difference metric Dist(D, Tc) is determined based on the quantized difference metric DistQ and the quantized difference metric DistC.

[0195] Clause 5. The method according to clause 4, wherein the quantization difference metric DistQ of the quantized partial data DQ and the quantization difference metric DistC of the truncated partial data DC are determined according to the following formula:

[0196] DistQ=(1+EQ)×AQ,

[0197] DistC=(1+EC)×AC,

[0198] Wherein AQ represents the amplitude of the quantization noise of the quantized partial data DQ, EQ represents the correlation coefficient between the quantization noise of the quantized partial data DQ and the quantized partial data DQ, AC represents the amplitude of the quantization noise of the truncated partial data DC, and EC represents the correlation coefficient between the quantization noise of the truncated partial data DC and the truncated partial data DC.

[0199] Clause 6. The method according to clause 5, wherein:

[0200] The amplitudes AQ and AC of the quantization noise are determined as follows:

[0201]

[0202] and / or

[0203] The correlation coefficients EQ and EC are determined according to the following formula:

[0204]

[0205]

[0206] Wherein N is the number of data in the calibration data set D, and Quantize(x, Tc) is a function that quantizes the data x with Tc as the maximum value.

[0207] Clause 7. A method according to any one of clauses 4 to 6, wherein the corresponding quantitative total difference measure Dist(D, Tc) is determined according to the following formula:

[0208] Dist(D,Tc)=DistQ+DistC.

[0209] Clause 8. The method of any one of clauses 4-7, wherein determining the optimized truncation threshold based on the quantitative total difference metric comprises:

[0210] From the multiple candidate truncation thresholds Tc, a candidate truncation threshold that minimizes the quantized total difference measure Dist(D, Tc) is selected as the optimized truncation threshold.

[0211] Clause 9. A method according to any one of clauses 3-8, wherein the search space for the truncation threshold is determined based at least on a maximum value of the calibration data set, and the candidate truncation threshold is determined at least in part based on a preset search accuracy.

[0212] Clause 10. The method of any one of clauses 1-9, wherein the calibration dataset comprises data from a plurality of batches, and the quantitative total difference metric is based on the quantitative total difference metric for each batch of data.

[0213] Clause 11. A computing device for calibrating quantization noise in a neural network, comprising:

[0214] at least one processor; and

[0215] At least one memory in communication with the at least one processor, having computer-readable instructions stored thereon, which, when loaded and executed by the at least one processor, cause the at least one processor to perform the method described in any one of clauses 1-10.

[0216] Clause 12. A computer-readable storage medium having program instructions stored therein, which, when loaded and executed by a processor, causes the processor to perform the method according to any one of clauses 1 to 10.

[0217] Although a plurality of embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may conceive of many modifications, changes, and alternatives without departing from the ideas and spirit of the present disclosure. It should be understood that in practicing the present disclosure, various alternatives to the embodiments of the present disclosure described herein may be adopted. The appended claims are intended to define the scope of protection of the present disclosure and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A method for calibrating quantization in a neural network, executed by a processor, wherein the neural network is configured to perform at least one of speech recognition and image recognition, the method comprising: receiving a calibration data set; quantizing the calibration data set using a cutoff threshold; determining a quantitative total difference measure for said quantization process; as well as Determining an optimized truncation threshold based on the quantized total difference metric, the optimized truncation threshold being used for performing quantization processing on data in a neural network operation process by a processor so as to minimize precision loss of the neural network in the processor after the quantization processing, wherein the data in the neural network operation process includes at least one of neurons, weights, biases, and gradients; The calibration dataset is divided into quantized partial data and truncated partial data according to the truncation threshold, and the quantized total difference metric is determined based on the quantized difference metric of the quantized partial data and the quantized difference metric of the truncated partial data.

2. The method according to claim 1, wherein The quantization difference metric of the quantized portion of data and / or the quantization difference metric of the truncated portion of data is determined based on at least the following two factors: the magnitude of the quantization noise; and The correlation coefficient between the quantization noise and the corresponding quantized data.

3. The method according to any one of claims 1-2, wherein: The quantization processing of the calibration data set using a truncation threshold comprises: The calibration data set is quantized using a plurality of candidate truncation thresholds in a truncation threshold search space.

4. The method according to claim 3, wherein: Determining a quantized total difference metric for the quantization process includes: For each candidate truncation threshold Tc, the calibration data set D is divided into quantized data DQ and truncated data DC according to the following formula: DC=[x|Abs(x)≥Tc,x∈D], Wherein n is the bit width of the quantized data after the quantization process; respectively determining a quantization difference metric DistQ of the quantized partial data DQ and a quantization difference metric DistC of the truncated partial data DC; and A corresponding quantized total difference metric Dist(D, Tc) is determined based on the quantized difference metric DistQ and the quantized difference metric DistC.

5. The method according to claim 4, wherein The quantization difference metric DistQ of the quantized partial data DQ and the quantization difference metric DistC of the truncated partial data DC are determined according to the following formula: DistQ=(1+EQ)×AQ, DistC=(1+EC)×AC, Wherein AQ represents the amplitude of the quantization noise of the quantized partial data DQ, EQ represents the correlation coefficient between the quantization noise of the quantized partial data DQ and the quantized partial data DQ, AC represents the amplitude of the quantization noise of the truncated partial data DC, and EC represents the correlation coefficient between the quantization noise of the truncated partial data DC and the truncated partial data DC.

6. The method according to claim 5, wherein: The amplitudes AQ and AC of the quantization noise are determined as follows: and / or The correlation coefficients EQ and EC are determined according to the following formula: Wherein N is the number of data in the calibration data set D, and Quantize(x, Tc) is a function that quantizes the data x with Tc as the maximum value.

7. The method according to any one of claims 4 to 6, wherein: The corresponding quantitative total difference metric Dist(D, Tc) is determined according to the following formula: Dist(D,Tc)=DistQ+DistC.

8. The method according to claim 4, wherein Determining an optimized truncation threshold based on the quantitative total difference metric includes: From the multiple candidate truncation thresholds Tc, a candidate truncation threshold that minimizes the quantized total difference measure Dist(D, Tc) is selected as the optimized truncation threshold.

9. The method according to claim 3, wherein: The search space for the truncation threshold is determined based at least on a maximum value of the calibration data set, and the candidate truncation threshold is determined at least in part based on a preset search accuracy. 10 . The method of claim 1 , wherein the calibration dataset comprises a plurality of batches of data, and the quantitative total difference metric is based on the quantitative total difference metric for each batch of data.

11. A computing device for calibration quantization in a neural network, comprising: at least one processor; as well as At least one memory in communication with the at least one processor, having computer-readable instructions stored thereon, which, when loaded and executed by the at least one processor, causes the at least one processor to execute the method according to any one of claims 1 to 10.

12. A computer-readable storage medium storing program instructions, wherein when the program instructions are loaded and executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Quantification realization method and related product

    CN109993296A

  • Convolutional neural network quantification method and device, computer and storage medium

    CN110363281A