Neural network calculation method, device, board and computer readable storage medium
By introducing multiple quantization intervals and non-uniform fixed-point data formats into neural networks, the problem that existing quantization methods cannot effectively reduce energy consumption is solved, and more efficient training and reasoning in neural networks are achieved.
Patent Information
- Application Number
- CN202010457247.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-26
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2040-09-28
AI Technical Summary
The existing quantization methods cannot effectively reduce energy consumption during the inference and training of neural networks, especially the quantization bit width of the backpropagation gradient is difficult to reduce, limiting the efficiency of edge-end fixed-point network training.
A neural network calculation method is proposed, including a control unit, a quantization unit and a calculation unit. By dividing the numerical distribution of floating-point data into multiple intervals, and using different precision quantization methods according to different intervals, fixed-point data is generated for calculation.
It realizes the reduction of model storage capacity and computing resource consumption in neural networks, improves training speed and inference efficiency, and avoids training crashes and non-convergence problems.
Smart Images

Figure CN113723600B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of neural networks. More specifically, the present disclosure relates to a method, device, board, and computer-readable storage medium for neural network computing. Background Art
[0002] In order to process massive amounts of data, modern neural networks continue to increase the depth, width, and resolution of the network, resulting in an increasing amount of model storage capacity and computational complexity. This causes the neural network to be extremely resource-intensive during training and reasoning. Therefore, technicians in this field are all trying their best to reduce energy consumption.
[0003] Quantized neural networks can compress models, reduce computing power consumption, and accelerate model reasoning and training without losing accuracy. Quantization is done by using low-bit fixed-point numbers instead of floating-point numbers in the network, and the data size can be reduced from the original 32 bits (single-precision floating-point numbers) to 8 bits or 4 bits. Quantized networks have such advantages and have long been a common technical means in this field.
[0004] In training quantization, in addition to quantizing weights and activation values, the back-propagation gradient must also be quantized. The amount of calculation for the back-propagation gradient is about twice that of the forward propagation. A large number of experiments have shown that in the fixed-point quantization training process, the quantization accuracy requirement for the back-propagation gradient is higher than the quantization accuracy requirement for the forward propagation weights and activation values. In other words, the quantization bit width of the back-propagation gradient must be higher than the quantization bit width of the forward propagation weights and activation values. Therefore, the quantization bit width of the back-propagation gradient cannot be low enough at present, and is generally maintained at 16 bits. The benefit of the compression model is limited, which has become the main factor limiting the quantization of the edge fixed-point network training process. If the training cannot be completed effectively, the neural network cannot be applied.
[0005] Therefore, there is a technical problem at this stage that the existing quantitative methods cannot effectively reduce energy consumption in both the inference and training processes. Summary of the invention
[0006] In order to at least partially solve the technical problems mentioned in the background technology, the solution of the present disclosure provides a method, device, board and computer-readable storage medium for neural network calculation.
[0007] In one aspect, the present disclosure discloses a neural network computing device, comprising a control unit, a quantization unit, and a computing unit. The control unit is used to provide a segmentation parameter; the quantization unit is used to quantize floating-point data according to the segmentation parameter to generate fixed-point data; and the computing unit is used to calculate the neural network using the fixed-point data.
[0008] In another aspect, the present disclosure discloses a multiplication unit, including a multiplier, an addition module, and a fixed-point to floating-point converter. The multiplier is used to multiply first fixed-point data and second fixed-point data to generate a fixed-point product; the addition module is used to add a plurality of quantization offset coefficients corresponding to the first fixed-point data and the second fixed-point data to generate a quantization offset coefficient sum; the fixed-point to floating-point converter is used to convert the fixed-point product into floating-point data according to the quantization offset coefficient sum.
[0009] In another aspect, the present disclosure discloses an integrated circuit device, including the aforementioned neural network computing device or multiplication unit, and discloses a board card, including the aforementioned integrated circuit device.
[0010] In another aspect, the present disclosure discloses a method for calculating a neural network, comprising: receiving floating-point data for calculating the neural network, the floating-point data falling within a numerical distribution; based on a segmentation parameter, segmenting the numerical distribution into a first interval and a second interval; determining whether the floating-point data falls within the first interval; if so, quantizing the floating-point data according to the segmentation parameter to generate fixed-point data; and calculating the neural network using the fixed-point data.
[0011] In another aspect, the present disclosure discloses a method for calculating a neural network based on floating-point data, including: presetting multiple quantization intervals, each quantization interval corresponds to a quantization formula, and each quantization formula exhibits a different gradient of quantization; determining whether the floating-point data falls within a specific interval of the multiple quantization intervals; and quantizing the floating-point data according to the quantization formula corresponding to the specific interval to generate fixed-point data; and calculating the neural network using the fixed-point data.
[0012] In another aspect, the present disclosure discloses a method for calculating a neural network based on floating-point data, comprising: presetting multiple quantization intervals; determining whether the floating-point data falls within a specific interval of the multiple quantization intervals; and quantizing the floating-point data according to the specific interval to generate fixed-point data; setting an N-bit flag in a data structure of the fixed-point data, the flag recording the specific interval, wherein N is a positive integer; and calculating the neural network using the fixed-point data.
[0013] In another aspect, the present disclosure discloses a method for calculating a neural network, comprising: providing a first segmentation parameter and a second segmentation parameter; quantizing first floating-point data according to the first segmentation parameter to generate first fixed-point data; quantizing second floating-point data according to the second segmentation parameter to generate second fixed-point data; performing multiplication operation on the first fixed-point data and the second fixed-point data to generate intermediate data; and calculating the neural network according to the intermediate data.
[0014] In another aspect, the present disclosure discloses a method for forward propagation in a neural network, comprising: receiving activation values and weights required for calculating the current layer; providing a first segmentation parameter, a first offset value and a first segmentation value corresponding to the activation value; providing a second segmentation parameter, a second offset value and a second segmentation value corresponding to the weight; quantizing the activation value according to the first segmentation parameter, the first offset value and the first segmentation value to generate a first fixed-point data; quantizing the weight according to the second segmentation parameter, the second offset value and the second segmentation value to generate a second fixed-point data; performing multiplication operation on the first fixed-point data and the second fixed-point data to generate intermediate data; performing floating-point calculation on the intermediate data to generate the activation value of the next layer; and repeating the above steps to perform calculations on each layer to complete the neural network.
[0015] In another aspect, the present disclosure discloses a method for backpropagation in a neural network, wherein the neural network includes weights and fixed-point data of the weights after the weights are fixed-point processed, including: receiving a next-layer error value; providing a segmentation parameter, an offset value, and a segmentation value corresponding to the next-layer error value; quantizing the next-layer error value according to the segmentation parameter, the offset value, and the segmentation value to generate error value fixed-point data; performing multiplication operation on the error value fixed-point data and the weight fixed-point data to generate a gradient of the weight; performing fixed-point calculation on the gradient to generate a current-layer error value; and adjusting the weight according to the current-layer error value.
[0016] In another aspect, the present disclosure discloses a method for training a neural network, comprising: performing forward propagation; calculating the next layer error value according to the next layer activation value; performing back propagation; and adjusting the weight according to the current layer error value. The step of performing forward propagation includes: receiving the activation value and weight required for calculating the current layer; quantizing the activation value according to the first segmentation parameter, the first offset value and the first segmentation value corresponding to the activation value to generate the first fixed-point data; quantizing the weight according to the second segmentation parameter, the second offset value and the second segmentation value corresponding to the weight to generate the second fixed-point data; multiplying the first fixed-point data and the second fixed-point data to generate intermediate data; and performing floating-point calculation on the intermediate data to generate the next layer activation value. The step of performing back propagation includes: quantizing the next layer error value according to the third segmentation parameter, the third offset value and the third segmentation value corresponding to the next layer error value to generate the third fixed-point data; multiplying the second fixed-point data and the third fixed-point data to generate the gradient of the weight; and performing fixed-point calculation on the gradient to generate the current layer error value.
[0017] In another aspect, the present disclosure discloses an electronic device, comprising one or more processors and a memory. The memory stores computer executable instructions, and when the computer executable instructions are executed by the one or more processors, the electronic device executes any one of the methods described above.
[0018] In another aspect, the present disclosure discloses a computer-readable storage medium, comprising computer-executable instructions. When the computer-executable instructions are executed by one or more processors, any one of the methods described above is performed.
[0019] This disclosure proposes a technical solution for adaptive non-uniform low-bit fixed-point training. By introducing multiple quantization intervals, non-uniform fixed-point data formats and corresponding hardware, the technical problem that the existing quantization method cannot effectively reduce energy consumption in both the reasoning and training processes is solved, achieving the technical effects of reducing the memory usage of the network model, reducing the resource consumption of model training and improving the training speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0021] Figure 1 is a graph showing the input and output of the sigmoid function;
[0022] Figure 2 is a schematic diagram showing the four-layer structure of a neural network;
[0023] Figure 3 is a schematic diagram showing a neural network computing device according to an embodiment of the present disclosure;
[0024] Figure 4 is a schematic diagram showing a control unit of an embodiment of the present disclosure;
[0025] Figure 5 is a graph showing possible numerical distribution of input data of an embodiment of the present disclosure;
[0026] Figure 6 is a schematic diagram showing an 8-bit non-uniform fixed-point quantization data structure according to an embodiment of the present disclosure;
[0027] Figure 7 is a flowchart showing a method for calculating a neural network according to another embodiment of the present disclosure;
[0028] Figure 8 is a schematic diagram showing a multiplication unit according to another embodiment of the present disclosure;
[0029] Fig. 9 is a flow chart showing a method for executing a multiplication calculation neural network according to another embodiment of the present disclosure;
[0030] Fig.10 is a schematic diagram showing a non-uniform fixed-point quantization data structure according to another embodiment of the present disclosure;
[0031] Fig.11 is a graph showing possible numerical distribution of input data according to another embodiment of the present disclosure;
[0032] Fig.12 is a schematic diagram showing a neural network computing device according to another embodiment of the present disclosure;
[0033] Fig.13 is a flow chart showing a forward propagation method according to another embodiment of the present disclosure;
[0034] Fig.14 is a flow chart showing a method of back propagation according to another embodiment of the present disclosure;
[0035] Fig.15 is a structural diagram showing an integrated circuit device according to another embodiment of the present disclosure; and
[0036] Fig.16 is a structural diagram of a board according to another embodiment of the present disclosure. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0038] It should be understood that the terms "first", "second", "third", "fourth", etc. in the claims, specifications and drawings of the present disclosure are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprise" used in the specifications and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their collections.
[0039] It should also be understood that the terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used in this disclosure and claims, the singular forms of "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in this disclosure and claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations.
[0040] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context.
[0041] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0042] Neural networks are built based on the concept of neurons. Neurons are like perceptrons. Generally, the activation function of a perceptron is a step function, while the activation function of a neuron is commonly a sigmoid function, which is defined as follows:
[0043]
[0044] The input and output diagram of the sigmoid function is as follows Figure 1 As shown, it can map a real number to the interval between 0 and 1, which is suitable for binary classification.
[0045] A neural network is a system of multiple neurons connected according to certain rules. Taking a convolutional neural network as an example, it is generally composed of the following four layer structures: input layer, convolution layer, pooling layer, and fully connected layer. Figure 2 Schematic diagram showing the four-layer structure of neural network 200.
[0046] The input layer 201 extracts part of the information from the input data and converts it into a feature matrix, which contains the features corresponding to the part of the information. The input data here may include but is not limited to image data, voice data or text data.
[0047] The convolution layer 202 is configured to receive the feature matrix from the input layer 201, and extract features from the input data through convolution operations. In actual use, the convolution layer 202 can be constructed with multiple convolution layers. Taking image data as an example, the convolution layer in the first half is used to capture local and detailed information of the image. For example, each pixel of the output image only senses the result of calculation of a very small range of values of the input image, and the subsequent convolution layers sense a larger range of values layer by layer to capture more complex and abstract information of the image. After the operation of multiple convolution layers, an abstract representation of the image at different scales is finally obtained. Although the feature extraction of the input image is completed through the convolution operation, the amount of information of the feature image is too large and the dimension is too high, which is not only time-consuming to calculate, but also easy to cause overfitting, and further dimensionality reduction is required.
[0048] The pooling layer 203 is configured to replace a certain area of the data with a value, which is usually the maximum or average value of all the values in the area. If the maximum value is used, it is called maximum pooling; if the average value is used, it is called mean pooling. Through pooling, the model size can be reduced and the calculation speed can be increased without losing too much information.
[0049] The fully connected layer 204 acts as a classifier in the entire convolutional neural network 200, which is equivalent to feature space transformation. It extracts and integrates all the previous useful information, and adds the nonlinear mapping of the aforementioned activation function. Multi-layer fully connected layers can theoretically simulate any nonlinear transformation to compare information based on different classifications, so as to determine whether the input data is similar to the comparison target.
[0050] Neural networks require weights as model parameters, and the best solution for the weights is obtained during the training of the neural network. In addition, some parameters such as the connection method of the neural network, the number of layers of the network, the number of nodes in each layer, etc. are not obtained through learning, but are set in advance. These artificially set parameters are called hyperparameters.
[0051] Neural networks are divided into forward propagation and back propagation. The so-called forward propagation is like Figure 2 The calculation is forward in the direction shown, from the input layer to the output layer, and the state value and activation value of each neuron are calculated in turn. Back propagation is to calculate from the output end to the input end in order to obtain the gradient value.
[0052] When input data is calculated in a neural network, it is generally hoped that a solution with a minimum loss function can be found, indicating that the result closest to the actual situation is obtained. However, the loss function of a neural network is very complex, and it is difficult to find the most suitable analytical expression. The general practice is to calculate the negative direction of the gradient, because the maximum value of the negative gradient is the direction in which the loss function decreases the most, and the back-propagation algorithm is to find this gradient value, so as to update the model parameters based on gradient descent. Therefore, the back-propagation algorithm starts from the output layer of the neural network model, and uses the chain rule of function derivation to find the model gradient layer by layer, hoping to obtain a solution with the minimum loss function.
[0053] By calculating the gradient in this way, each neural unit will only be calculated once, and will not be calculated repeatedly. The fundamental reason why this calculation direction is efficient is that when calculating the gradient, the previous level unit depends on the calculation of the next level unit. By first calculating the gradient value of the next level unit and then calculating the previous level unit, the calculated result can be fully utilized, avoiding repeated calculations.
[0054] Regularization penalty and ReLU activation function are often used in neural network training. Regularization penalty is because when overfitting, in order to match all the data in the test set, the poorly generalized high-order function will produce a lot of jitter, which causes the derivative to become very large, and requires a large parameter to fit all the data. Therefore, adding a penalty term can penalize the situation where the parameters are very large to avoid these parameters with very large jitter.
[0055] The ReLU activation function is when the model has N layers, theoretically the activation rate of the neuron will be reduced by 2 N times, ReLU can better realize the sparse model to mine relevant features and fit the training data. In addition, compared with other activation functions, ReLU has the following advantages: for linear functions, ReLU has stronger expressive power, especially in deep networks; for nonlinear functions, the gradient of ReLU in the non-negative interval is a constant, so there is no gradient vanishing problem, so that the convergence speed of the model is maintained in a stable state. The gradient vanishing problem means that when the gradient is less than 1, the error between the predicted value and the true value will decay once per layer. If the aforementioned sigmoid function is used as the activation function in a deep model, this phenomenon is particularly obvious, which will cause the model convergence to stagnate.
[0056] Due to the effects of regularization penalty and ReLU activation function, the weights, activation values and gradients of each layer are not uniform in numerical distribution, and the numerical distribution of weights, activation values and gradients tends to be close to 0 during the entire training process, like a Gaussian distribution. Considering that the values will be concentrated near 0, uniform quantization will quantize all these values close to 0 to 0, which seriously affects the direction of network training. For example, if 7-bit uniform quantization is used, the absolute value is less than The value of will be quantized to 0. This quantization is unbearable for training. Under the prior art uniform quantization scheme, the factor limiting the accuracy of network training mainly comes from the gradient quantization error, especially the quantization error of the gradient near 0 in the value distribution.
[0057] The present disclosure provides an adaptive non-uniform low-bit quantization scheme, which is applicable to the training tasks and reasoning tasks of various neural networks (such as convolutional neural networks, recurrent neural networks, graph neural networks, etc.). As mentioned above, in the process of improving the neural network, two types of parameters are involved. One is the general parameter, that is, the parameter data obtained through training, such as the weight w and offset b in y=wx+b; the other is called a hyperparameter, which is a parameter set before starting learning and cannot be obtained through training.
[0058] The apparatus described in the embodiments of the present disclosure and the various devices, units, modules, etc. described below can be implemented in the form of hardware circuits, such as digital circuits or analog circuits. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor mentioned in the embodiments can be any appropriate hardware processor, such as CPU, GPU, FPGA, DSP and ASIC, etc. Unless otherwise specified, the memory, storage device, and storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM (Resistive Random Access Memory), dynamic random access memory DRAM (Dynamic Random Access Memory), static random access memory SRAM (Static Random-Access Memory), enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), high-bandwidth memory HBM (High-Bandwidth Memory), hybrid memory cube HMC (Hybrid Memory Cube), etc.
[0059] Optionally, when the following devices, various devices, units, and modules are implemented as ASIC, their advantages over other hardware implementations are in terms of power consumption, reliability, area, etc., especially when used in high-performance, low-power mobile terminals.
[0060] One embodiment of the present disclosure is a neural network computing device, which divides the numerical distribution of floating-point input data into a first interval and a second interval, and adopts different quantization methods for floating-point data falling in different intervals. Figure 3 shown.
[0061] The neural network computing device of this embodiment includes a control unit 31, a quantization unit 32 and a computing unit 33. The control unit 31 is used to provide various parameters and hyperparameters required in the quantization process; the quantization unit 32 is used to quantize floating-point data according to the parameters and hyperparameters to generate fixed-point data; the computing unit 33 is used to calculate the neural network using the fixed-point data. Figure 3 Each unit can be implemented in the form of hardware circuit.
[0062] The control unit 31 is used to provide the segmentation parameter α, the offset value shift and the segmentation value x required in the quantization process. p Etc. Figure 4 As shown, the control unit 31 includes an input / output module 311 , a segmentation parameter generator 312 , an offset value generator 313 and a segmentation value generator 314 .
[0063] The input / output module 311 serves as a channel for signal transmission between the control unit 31 and the external unit. It can output signals when a control requirement occurs, and can also receive control signals from the external unit and send the control signals to the segmentation parameter generator 312, the offset value generator 313 or the segmentation value generator 314.
[0064] The segmentation parameter generator 312 is used to generate a hyperparameter: a segmentation parameter α. The segmentation parameter α is used to define the first interval and the second interval. If the input data x (floating point number) is distributed in the numerical distribution D, the relationship between the segmentation parameter α and the first interval A and the second interval B is:
[0065] A={x|x<2 α max{abs(D)},x∈D} (1)
[0066] B={x|x≥2 α max{abs(D)},x∈D} (2)
[0067] See also Figure 5, the Gaussian distribution curve in the figure is the possible value distribution of the input data x. The input data x will basically fall completely within the range of the value distribution D. The first interval A is a positive and negative symmetrical interval including 0, that is, the range between the two dotted lines in the figure, and the range outside the dotted lines is the second interval B. When the input data x is less than 2 α max{abs(D)} indicates that the input data x falls within the first interval A; when the input data x is greater than or equal to 2 α When max{abs(D)}, it means that the input data x falls within the second interval B.
[0068] The segmentation parameter α determines Figure 5 The position of the dotted line in the middle determines the range of the first interval A. The smaller the absolute value of the segmentation parameter α is, the larger the range of the first interval A is. The segmentation parameter α is related to the quantization bit width b, which represents the number of bits of a fixed-point number. The quantization bit width b is a positive integer, and the segmentation parameter α is a negative integer. In this embodiment, the relationship between the two is as follows:
[0069] 50%×b≤abs(α)≤90%×b
[0070] Taking the quantization bit width b as 8 (ie, an 8-bit fixed-point number) as an example, the segmentation parameter α can be optionally one of -4, -5, -6, and -7.
[0071] The offset value generator 313 is used to generate an offset value shift, where the offset value shift represents a quantization offset, and the offset value generator 313 generates the offset value shift according to the following expression:
[0072]
[0073] The ceil function returns the smallest integer greater than or equal to the expression.
[0074] The segmentation value generator 314 is used to generate the segmentation value x p , split value x p Indicates that the first interval A and the second interval B are Figure 5 The dividing line on the horizontal axis, that is, Figure 4 The segmentation value generator 314 generates the segmentation value x according to the following expression: p :
[0075] x p =2 α max {abs(D)} (4)
[0076] The above segmentation parameter α, offset value shift and segmentation value x p It is a parameter required in the quantization process. After being generated by the control unit 31 , it will be sent to the quantization unit 32 through the input / output module 311 .
[0077] Optionally, in one embodiment, the control unit 31 may also have other hardware circuit implementations, which will not be described in detail here.
[0078] Back to Figure 3 The quantization unit 32 includes: an absolute value operator 321, a comparator 322, a two-way selector 323, an adder 324, a quantizer 325 and a bus converter 326.
[0079] The absolute value operator 321 receives input data x, takes the absolute value of the input data x in floating point format, and outputs the absolute value abs(x). The comparator 322 receives the absolute value abs(x) and the division value x p , the absolute value abs(x) and the segmentation value x p If the split value x p is greater than the absolute value abs(x), referring to formula (1), indicating that the input data x falls within the first interval A, and the comparator 322 sets the value of the flag bit flag to 1. If the segmentation value x p is less than or equal to the absolute value abs(x), referring to formula (2), indicating that the input data x falls within the second interval B, and the value of the flag is set to 0. The value of the flag reflects whether the input data x falls within the first interval A or the second interval B.
[0080] The two-way selector 323 receives the segmentation parameter α from the control unit 31, and determines whether the output is the segmentation parameter α or 0 based on the value of the flag. When the value of the flag is 1, it means that the input data x falls within the first interval A, and the values within the first interval A including 0 are easily quantized to 0. This embodiment adjusts the quantization accuracy for the values within the first interval A, so the two-way selector 323 sets the output value to the segmentation parameter α for the subsequent stage to adjust the accuracy. When the value of the flag is 0, it means that the input data x falls within the second interval B, and the values within this interval do not need to adjust the accuracy, so the two-way selector 323 sets the output value to 0.
[0081] The adder 324 receives the output of the two-way selector 323 and the offset value shift from the control unit 31, and adds the offset value shift to the output value of the two-way selector 323 to generate a quantized offset value s. That is, when the value of the flag flag is 1, the quantized offset value s is the segmentation parameter α plus the offset value shift; when the value of the flag flag is 0, the quantized offset value s is the offset value shift.
[0082] The quantizer 325 quantizes the input data x into n-bit fixed-point data, and the quantization formula is:
[0083]
[0084] Among them, 2 s It is called the quantization interval. The round function is the result of rounding. x[n-1:0] is the fixed-point data after the input data x is quantized. However, the output x[n-1:0] of the quantizer 325 is only intermediate data, not the final quantization result.
[0085] The bus converter 326 is used to combine the value of the flag flag with the intermediate data x[n-1:0] to generate the fixed-point data x[n:0]. More specifically, the value of the flag flag is added to the intermediate data x[n-1:0] so that the final fixed-point data x[n:0] is n+1 bits.
[0086] In another scenario, if the value of the flag flag is not important to the computing unit 33 , this embodiment may not include the bus converter 326 , and the output x[n-1:0] of the quantizer 325 is the final quantization result, which is directly transmitted to the computing unit 33 .
[0087] Optionally, in one embodiment, the quantization unit 32 may also have other hardware circuit implementations, which will not be described in detail here.
[0088] This embodiment is equipped with a new non-uniform fixed-point quantization data structure. The data format of the non-uniform fixed-point data x[n:0] will be described below. Figure 6 An 8-bit non-uniform fixed-point quantization data structure is shown, which includes a sign bit 61, a value bit 62 and a flag bit 63. The sign bit 61 is 1 bit, which is the most significant bit (MSB), and is used to record the positive and negative signs of the fixed-point data. The flag bit 63 is 1 bit, which is the least significant bit (LSB), and is used to record the value of the flag bit flag. The middle 6 bits are the value bit 62, which is used to record the value of the fixed-point data. The output x[n-1:0] of the quantizer 325 corresponds to the values of the sign bit 61 and the value bit 62 of the non-uniform fixed-point quantization data structure. The bus converter 326 then records the value of the flag bit flag in the flag bit 63 to generate the complete fixed-point data x[n:0]. Based on Figure 6 The relationship between the corresponding floating-point input data x and the sign bit 61, the value bit 62 and the flag bit 63 is as follows:
[0089] x=(-1) sign ×value×2 α·flag ×interval (6)
[0090] Among them, sign is the value of the sign bit 61, and value is the value of the value bit 62.
[0091] If the number of bits of the fixed-point number is not 8 but n, then under this fixed-point data structure, its most significant bit and least significant bit are also the sign bit 61 and the flag bit 63, and the middle n-2 bits are the value bits 62.
[0092] Back to Figure 3 In the quantization unit 32, the floating point input data x is converted into fixed point data x[n:0], and the fixed point data x[n:0] is transmitted to the calculation unit 33 for calculation. The calculation unit 33 can perform specific calculations according to actual needs, such as Figure 2 The convolution calculation of the fully connected layer 204 generates an intermediate result y[n:0] after the calculation is completed. The intermediate result y[n:0] is also fixed-point data. The calculation unit 33 then restores the intermediate result y[n:0] to floating-point data y to complete the entire calculation process.
[0093] The above description takes the input data x as an activation value as an example, but the present disclosure is not limited to this. That is, in a neural network, any data that needs to be quantized (such as weights and gradients, etc.) can be converted into fixed-point data using the quantization unit 32.
[0094] In summary, this embodiment implements a method for calculating a neural network. Figure 7 A flow chart of this method is shown. The process of this embodiment is described based on the aforementioned hardware design. It is understandable that it is not limited to the hardware design of the present disclosure, such as digital circuits or analog circuits. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. The artificial intelligence processor mentioned in the embodiment can be any appropriate hardware processor, such as CPU, GPU, FPGA, DSP and ASIC, etc. Unless otherwise specified, the memory, storage device, storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM, dynamic random access memory DRAM, static random access memory SRAM, enhanced dynamic random access memory EDRAM, high bandwidth memory HBM, hybrid storage cube HMC, etc.
[0095] In step 71, the absolute value operator 321 receives floating point data for computing a neural network, the floating point data falling within a numerical distribution. Figure 3 , the absolute value operator 321 receives input data x, takes the absolute value of the input data x in floating point format, and outputs the absolute value abs(x). It should be particularly emphasized that in this embodiment, a series of floating point data are received, each of which is based on Figure 7 process to process.
[0096] In step 72, the control unit 31 divides the numerical distribution into a first interval and a second interval based on the segmentation parameter. Figure 4, the segmentation parameter generator 312 of the control unit 31 generates a segmentation parameter α. The segmentation parameter α is used to define the first interval A and the second interval B according to equations (1) and (2).
[0097] In step 73, the comparator 322 determines whether the floating point data falls within the first interval. Figure 3 The comparator 322 receives the absolute value abs(x) and the segmentation value x p , the absolute value abs(x) and the segmentation value x p If the split value x p is greater than the absolute value abs(x), indicating that the input data x falls within the first interval A. The comparator 322 sets the value of the flag bit flag to 1. If the segmentation value x p Less than or equal to the absolute value abs(x), indicating that the input data x falls within the second interval B, and the value of the flag is set to 0.
[0098] If the floating point data falls within the first interval, step 74 is executed, and the two-way selector 323, the adder 324, the quantizer 325 and the bus converter 326 quantize the floating point data according to the segmentation parameter to generate fixed point data. Figure 3 , the two-way selector 323 receives the segmentation parameter α from the control unit 31, and determines to output the segmentation parameter α or 0 based on the value of the flag. Since the floating-point data falls within the first interval, the value of the flag is 1, and the two-way selector 323 sets the output value to be the segmentation parameter α. The quantization offset value s output by the adder 324 is the segmentation parameter α plus the offset value shift. The quantizer 325 quantizes the input data x into the intermediate data x based on equation (5): q The bus converter 326 then adds the value of the flag bit flag to the intermediate data x q to generate fixed-point data x[n:0].
[0099] If the floating-point data does not fall within the first interval, step 75 is executed, and the two-way selector 323, the adder 324, the quantizer 325 and the bus converter 326 do not quantize the floating-point data according to the segmentation parameter, but quantize the floating-point data according to the following expression to generate the fixed-point data x q :
[0100]
[0101] interval=2 shift (8)
[0102]
[0103] These expressions are not substantially different from equations (3) and (5). The difference is that since the floating point data falls within the second interval B, the value of the flag is 0, and the two-way selector 323 sets the output value to 0. The quantization offset value s output by the adder 324 is only the offset value shift. The quantizer 325 quantizes the input data x into the intermediate data x q , and the bus converter 326 then adds the value of the flag bit flag to the intermediate data x q to generate fixed-point data x[n:0].
[0104] In step 76, the calculation unit 33 calculates the neural network using the fixed-point data x[n:0]. Figure 3 The calculation unit 33 can perform specific calculations according to actual needs, and generate an intermediate result y[n:0] after the calculation is completed. The intermediate result y[n:0] is also fixed-point data. The calculation unit 33 then restores the intermediate result y[n:0] to floating-point data y to complete the entire calculation process.
[0105] Since this embodiment can generate a flag through the comparator 322 to record whether the input data x falls within the first interval A or the second interval B, so that the quantization unit 32 can select the appropriate precision for quantization, this embodiment is an "adaptive" system. Furthermore, the precision of the first interval A and the second interval B are different. The first interval A including 0 uses 2 α The precision is divided into finer parts to avoid a large amount of data close to 0 from being quantized to 0, so this embodiment is still a "non-uniform" system.
[0106] For a neural network, the computing unit 33 in the above-mentioned embodiment can perform an important calculation, namely, matrix multiplication. Matrix multiplication involves a large number of fixed-point multiplications. Figure 6 A fixed-point data structure with a special structure of multiplication unit is proposed.
[0107] For Figure 6 When performing multiplication calculation on two fixed-point data of a fixed-point data structure, it is assumed that the first fixed-point data x1 falls within the first numerical distribution, and the second fixed-point data x2 falls within the second numerical distribution. The first segmentation parameter α1 is used to segment the first numerical distribution into the first interval and the second interval, and the second segmentation parameter α2 is used to segment the second numerical distribution into the third interval and the fourth interval. The first flag flag1 is used to reflect that the first fixed-point data x1 falls within the first interval or the second interval, and the second flag flag2 is used to reflect that the second fixed-point data x2 falls within the third interval or the fourth interval. Based on formula (6), its floating-point product x1×x2 is:
[0108]
[0109] Another embodiment of the present disclosure is a multiplication unit, a schematic diagram of which is shown in FIG. Figure 8 As shown, the multiplication unit 80 of this embodiment includes a multiplier 81, an addition module 82 and a fixed-point to floating-point converter 83.
[0110] The multiplier 81 is used to multiply the first fixed-point data x1[n-1:0] and the second fixed-point data x2[n-1:0] to generate a fixed-point product y[2n-1:0], that is, to implement equation (10) The first fixed-point data x1[n-1:0] and the second fixed-point data x2[n-1:0] can come from Figure 3 The output of the quantizer 325 or the output from the bus converter 326 is then removed, leaving the sign bit and the value bit. Since the first fixed-point data x1[n-1:0] and the second fixed-point data x2[n-1:0] are both n bits, the fixed-point product y[2n-1:0] is 2n bits.
[0111] The adding module 82 is used to add a plurality of quantization offset coefficients corresponding to the first fixed-point data x1[n-1:0] and the second fixed-point data x2[n-1:0] to generate a quantization offset coefficient sum, that is, to implement the formula (10) As shown in formula (8), formula (10) is equal to Therefore, the quantization offset coefficient involves a segmentation parameter α, a flag and an offset value shift.
[0112] The addition module 82 includes a first selector 821, a second selector 822 and an adder 823. The first selector 821 sets the first output value to the first segmentation parameter α1 or 0 according to the first flag bit flag1. When the first flag bit flag1 is 1, that is, the value of the flag bit x1[n] is 1, it means that the first fixed-point data x1[n-1:0] falls within the first interval, and the first segmentation parameter α1 needs to participate in the calculation, so the first segmentation parameter α1 is output; when the first flag bit flag1 is 0, that is, the value of the flag bit x1[n] is 0, it means that the first fixed-point data x1[n-1:0] falls within the second interval, and the first segmentation parameter α1 does not participate in the calculation, so 0 is output. Similarly, the second selector 822 sets the second output value to the second segmentation parameter α2 or 0 according to the second flag bit flag2. Its specific operation is the same as that of the first selector 821 and will not be repeated.
[0113] The adder 823 adds the first offset value shift1, the second offset value shift2, the first output value and the second output value to generate a quantized offset coefficient sum s, that is, α1·flag1+α2·flag2+shift1+shift2. The first offset value shift1 and the second offset value shift2 are calculated according to formula (3).
[0114] The fixed-point to floating-point converter 83 converts the fixed-point product y[2n-1:0] into floating-point data y according to the quantization offset coefficient and s. The fixed-point to floating-point converter 83 includes a 2-power calculator 831 and a multiplier 832. The 2-power calculator 831 generates a quantization offset value 2 based on the quantization offset coefficient and s. s , to achieve Multiplier 832 adds the fixed-point product y[2n-1:0] to the quantization offset value 2 s Multiplying them, we get the floating-point product x1×x2.
[0115] Optionally, in one embodiment, the multiplication unit 80 may also have other hardware circuit implementations, which will not be described in detail here.
[0116] In summary, this embodiment implements a method for executing a multiplication calculation neural network. The flowchart of this method is as follows: Fig. 9 As shown. The process of this embodiment is described based on the aforementioned hardware design. It is understandable that it is not limited to the hardware design of the present disclosure, such as digital circuits or analog circuits. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. The artificial intelligence processor mentioned in the embodiment can be any appropriate hardware processor, such as CPU, GPU, FPGA, DSP and ASIC, etc. Unless otherwise specified, the memory, storage device, storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM, dynamic random access memory DRAM, static random access memory SRAM, enhanced dynamic random access memory EDRAM, high bandwidth memory HBM, hybrid memory cube HMC, etc.
[0117] In step 91, the control unit 31 provides a first segmentation parameter α1 and a second segmentation parameter α2. More specifically, the segmentation parameter generator 312 of the control unit 31 generates the first segmentation parameter α1 and the second segmentation parameter α2.
[0118] In step 92, the quantization unit 32 quantizes the first floating point data x1 according to the first segmentation parameter α1 to generate the first fixed point data x1[n:0]. The comparator 322 generates the segmentation value x1 based on the first segmentation parameter α1 and according to formula (4): p, dividing the first numerical distribution into a first interval A and a second interval B; then determining whether the first floating-point data x1 falls within the first interval A, where the first interval A is defined by equation (1). If the first floating-point data x1 falls within the first interval A, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x1[n-1:0] according to the following expression:
[0119]
[0120] If the first floating-point data x1 does not fall within the first interval A, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x1[n-1:0] according to the following expression:
[0121]
[0122] Finally, the bus converter 326 fills in the flag bit flag to generate the complete first fixed-point data x1[n:0].
[0123] In step 93, the quantization unit 32 quantizes the second floating point data x2 according to the second segmentation parameter α2 to generate the second fixed point data x2[n:0]. Similarly, the comparator 322 generates the segmentation value x based on the second segmentation parameter α2 and according to formula (4): p , dividing the second numerical distribution into a third interval and a fourth interval; then determining whether the second floating-point data x2 falls within the third interval, the third interval being defined by formula (1). If the second floating-point data x2 falls within the third interval, the two-way selector 323, the adder 324 and the quantizer 325 generate the second fixed-point data x2[n-1:0] according to the following expression:
[0124]
[0125]
[0126] If the second floating-point data x2 does not fall within the third interval, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x2[n-1:0] according to the following expression:
[0127]
[0128] Finally, the bus converter 326 fills in the flag bit flag to generate the complete second fixed-point data x2[n:0].
[0129] In step 94, the multiplication unit 80 performs a multiplication operation on the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0] to generate the intermediate data y. As mentioned above, the flag bit of the data structure of the fixed-point data disclosed in the present invention is used to record the interval of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0], the sign bit is used to record the positive and negative signs of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0], and the value bit is used to record the values V1 and V2 of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0]. The multiplier 81 adds the values of the sign bits of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0] to generate the signed sum value sign t ; The task of the first selector 821 is equivalent to multiplying the first segmentation parameter α1 and the value of the flag bit x1[n] of the first fixed-point data to generate the first parameter multiplication value pm1; The task of the second selector 822 is equivalent to multiplying the second segmentation parameter α2 and the value of the flag bit x2[n] of the second fixed-point data to generate the second parameter multiplication value pm2; Finally, the fixed-point to floating-point converter 83 executes the following expression to generate the intermediate data:
[0130]
[0131] In step 95, a neural network is calculated based on the intermediate data. Figure 2 In the neural network architecture shown, the input data can be diverse, such as images, voices, text data, etc. These data will undergo a large number of multiplication operations in the input layer 201, the convolution layer 202, the pooling layer 203, and the fully connected layer 204. These multiplication operations can all be implemented using the aforementioned steps until the reasoning process is completed and the image, voice, and text data are finally recognized. In addition to the input data, the parameters in the neural network can also be converted into fixed-point data using the aforementioned steps to perform multiplication operations with the input data.
[0132] Since data close to 0 will be quantized to 0 during the quantization process, affecting the calculation, the aforementioned embodiments divide the numerical distribution of the input data into two: one is a positive and negative symmetrical interval including 0, and the other is an interval other than the aforementioned interval. However, the present disclosure does not limit the number of intervals, and as long as the numerical distribution within different ranges needs to be quantized with different precisions, it can be appropriately divided into multiple intervals.
[0133] When the interval exceeds two, Figure 6 The data structure in the example only needs to adjust the size of the flag bit 63. Taking 3 or 4 intervals as an example, in order to fully record that the data falls into one of the 3 or 4 intervals, the flag bit 63 needs 2 bits, such as Fig.10As shown in the flag bit 10. In other words, if the value distribution is divided into N intervals, the flag bit requires ceil[log2N] bits.
[0134] Another embodiment of the present disclosure is a neural network computing device, which is different from the above embodiment in that the neural network computing device of this embodiment divides the numerical distribution of floating-point data into a first interval, a second interval, and a third interval, and different quantization methods are used for floating-point data falling in different intervals. Fig.11 As shown, the numerical distribution D is divided into a first interval A, a second interval B and a third interval C according to different quantization accuracies.
[0135] This embodiment uses two segmentation parameters to define the three intervals. The segmentation parameter α1 is used to define the first interval A and the second interval B, and the segmentation parameter α2 is used to define the second interval B and the third interval C. The relationship between the segmentation parameters α1, α2 and the first interval A, the second interval B and the third interval C is:
[0136]
[0137] When calculating the fixed-point data x that falls in the first interval A q , the following expressions can be used:
[0138]
[0139]
[0140] When calculating the fixed-point data x that falls in the second interval B q , the following expressions can be used:
[0141]
[0142] When calculating the fixed-point data x that falls in the third interval C q , the following expressions can be used:
[0143]
[0144] interval3=2 shift
[0145]
[0146] The schematic diagram of this embodiment is as follows Fig.12 As shown, it is Figure 3 There is no significant difference in the framework of FIG. 1 , and the only difference lies in the control unit 121 , the comparator 122 and the three-way selector 123 .
[0147] Compared to Figure 3The control unit 31 outputs the segmentation parameters α1 and α2 to the three-way selector 123, and outputs the first segmentation value x corresponding to the segmentation parameter α1. p1 and the second segmentation value x corresponding to the segmentation parameter α2 p2 To the comparator 122, its expression is as follows:
[0148]
[0149] Compared to Figure 3 The comparator 322 is implemented as a two-stage comparison circuit. The first stage compares the absolute value abs(x) of the input data x with the first segmentation value x. p1 , if the absolute value abs(x) is less than the first segmentation value x p1 , indicating that the input data x falls within the first interval A, so it does not need to be compared with the second segmentation value x p2 Compare and output flag flag is 00 if the absolute value abs(x) is not less than the first segmentation value x p1 , indicating that the input data x falls in the second interval B or the third interval C, then it enters the second stage circuit to compare the absolute value abs(x) of the input data x with the second segmentation value x p2 If the absolute value abs(x) is less than the second segmentation value x p2 , indicating that the input data x falls within the second interval B, so the output flag flag is 01. If the absolute value abs(x) is not less than the second segmentation value x p2 , indicating that the input data x falls within the third interval C, so the output flag is 10.
[0150] Compared to Figure 3 The two-way selector 323 of the three-way selector 123 receives the segmentation parameters α1 and α2, and determines the output as the segmentation parameters α1, α2 or 0 based on the value of the flag. When the value of the flag is 00, it means that the input data x falls within the first interval A, so the three-way selector 123 sets the output value to the segmentation parameter α1. When the value of the flag is 01, it means that the input data x falls within the second interval B, so the three-way selector 123 sets the output value to the segmentation parameter α1. 2。 When the value of the flag is 10, it indicates that the input data x falls within the third interval C, so the three-way selector 123 sets the output value to 0.
[0151] The remaining components operate similar to Figure 3 The corresponding components are the same, so they are not described in detail. Optionally, in one embodiment, Fig.12 The embodiments may also be implemented in other hardware circuits, which will not be described in detail here.
[0152] Fig.12The embodiment is illustrated with three intervals as an example, and the present disclosure is not limited to the number of intervals, and those skilled in the art can easily extend to multiple intervals without creative investment.
[0153] Another embodiment of the present disclosure is a method for forward propagation in a neural network, that is, the reasoning process of the neural network, which can be used Figure 3 or Fig.12 For the convenience of explanation, the following will be combined with Figure 3 The embodiments are described below. Fig.13 A flow chart of the method of this embodiment is shown. The process of this embodiment is described based on the aforementioned hardware design. It is understandable that it is not limited to the hardware design of the present disclosure, such as digital circuits or analog circuits. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. The artificial intelligence processor mentioned in the embodiment can be any appropriate hardware processor, such as CPU, GPU, FPGA, DSP and ASIC, etc. Unless otherwise specified, the memory, storage device, storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM, dynamic random access memory DRAM, static random access memory SRAM, enhanced dynamic random access memory EDRAM, high bandwidth memory HBM, hybrid memory cube HMC, etc.
[0154] In step 1301, the absolute value operator 321 receives the activation value x1 and weight x2 required for calculating the current layer. The activation value x1 and weight x2 in floating point format are input to the absolute value operator 321 as input data, and the absolute value abs(x) is output.
[0155] In step 1302, the control unit 31 provides a first segmentation parameter α1, a first offset value shift1 and a first segmentation value x1 corresponding to the activation value x1. p1 The first segmentation parameter α1, the first offset value shift1 and the first segmentation value x p1 All of them have been explained in the above embodiments and will not be described again. p1 It can be obtained by calculating formula (4).
[0156] In step 1303, the control unit 31 provides a second segmentation parameter α2, a second offset value shift2 and a second segmentation value x corresponding to the weight x2. p2 . The second split value x p2 It can be obtained by calculating formula (4).
[0157] In step 1304, the quantization unit 32 calculates the first segmentation parameter α1, the first offset value shift1 and the first segmentation value x p1, quantize the activation value x1 to generate the first fixed-point data. The comparator 322 is based on the first segmentation value x p1 , it is determined whether the activation value x1 falls within the first interval A, and the first interval A is defined by formula (1). If the activation value x1 falls within the first interval A, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x1[n-1:0] according to the following expression:
[0158]
[0159] If the activation value x1 does not fall within the first interval A, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x1[n-1:0] according to the following expression:
[0160]
[0161] Finally, the bus converter 326 fills in the flag bit flag to generate the complete first fixed-point data x1[n:0].
[0162] In step 1305, the quantization unit 32 calculates the second segmentation parameter α2, the second offset value shift2 and the second segmentation value x p2 , quantize the weight x2 to generate the second fixed-point data. More specifically, the weight x2 falls within the second numerical distribution, and the second numerical distribution is divided into a third interval and a fourth interval. The comparator 322 is based on the second segmentation value x p2 , determine whether the weight value x2 falls within the third interval, and the third interval is also defined by formula (1). If the weight value x2 falls within the third interval, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x2[n-1:0] according to the following expression:
[0163]
[0164] If the weight value x2 does not fall within the third interval, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x2[n-1:0] according to the following expression:
[0165]
[0166] Finally, the bus converter 326 fills in the flag bit flag to generate the complete second fixed-point data x2[n:0].
[0167] In step 1306, the calculation unit 33 performs a multiplication operation on the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0] to generate intermediate data. In this embodiment, the calculation unit 33 has Figure 8The structure of the multiplication unit 80 performs multiplication operation on the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0] to generate the intermediate data y. As mentioned above, the flag bit of the data structure of the fixed-point data disclosed in the present invention is used to record the interval of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0], the sign bit is used to record the positive and negative signs of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0], and the value bit is used to record the values V1 and V2 of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0]. The multiplier 81 adds the values of the sign bits of the first fixed-point data x1[n:0] and the second fixed-point data x2[n:0] to generate the signed sum value sign t ; The task of the first selector 821 is equivalent to multiplying the first segmentation parameter α1 and the value of the flag bit x1[n] of the first fixed-point data to generate the first parameter multiplication value pm1; The task of the second selector 822 is equivalent to multiplying the second segmentation parameter α2 and the value of the flag bit x2[n] of the second fixed-point data to generate the second parameter multiplication value pm2; Finally, the fixed-point to floating-point converter 83 executes the following expression to generate intermediate data:
[0168]
[0169] This intermediate data is the output result of this layer and also the input data of the next layer, namely the activation value. The method returns to step 1301, and the activation value of the next layer obtained in this step is input to the next layer, and the process is repeated until all layers are fully calculated.
[0170] In step 1307, the neural network inference is completed. Figure 2 In the neural network architecture shown, the input data can be diverse, such as images, speech, text data, etc. These data are repeatedly multiplied in the input layer 201, the convolution layer 202, the pooling layer 203, and the fully connected layer 204 for multiple times until the reasoning process is completed, and finally the image, speech, and text data are recognized.
[0171] Another embodiment of the present disclosure is a method for back propagation in a neural network, which quantifies the back propagation gradient by propagating error values. The method can also be used Figure 3 or Fig.12 For the convenience of explanation, the following will be combined with Figure 3 The embodiments are described below. Fig.14A flow chart of the method of this embodiment is shown. The process of this embodiment is described based on the aforementioned hardware design. It is understandable that it is not limited to the hardware design of the present disclosure, such as digital circuits or analog circuits. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. The artificial intelligence processor mentioned in the embodiment can be any appropriate hardware processor, such as CPU, GPU, FPGA, DSP and ASIC, etc. Unless otherwise specified, the memory, storage device, storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM, dynamic random access memory DRAM, static random access memory SRAM, enhanced dynamic random access memory EDRAM, high bandwidth memory HBM, hybrid memory cube HMC, etc.
[0172] In step 1401, the quantization unit 32 receives the next layer error value x R . Reference Figure 2 , the back propagation is in the direction of the fully connected layer 204, the pooling layer 203, the convolution layer 202, and the input layer 201, and the error value of the output value is transmitted back in reverse. For the final output node, the difference between the activation value generated by the network and the actual value is taken as the error value x of the next layer R , the error value x of the next layer R That is Figure 3 The input data x of the device.
[0173] In step 1402, the control unit 31 provides the error value x corresponding to the next layer. R The segmentation parameter α R , offset value shift R and the split value x pR It should be noted that the forward segmentation parameter α, the offset value shift and the segmentation value x p With the reverse segmentation parameter α R , offset value shift R and the split value x pR Different, where the split value x pR The same can be obtained by calculating the formula (4).
[0174] In step 1403, the quantization unit 32 calculates the segmentation parameter α according to the segmentation parameter α. R , offset value shift R and the split value x pR , quantize the error value x of the next layer R , to generate error value fixed point data. In more detail, the comparator 322 is based on the segmentation value x pR , determine the next layer error value x R Whether it falls within the first interval A, the first interval A is defined by formula (1).R If it falls within the first interval A, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x according to the following expression: R [n-1:0]:
[0175]
[0176] The next layer error value x R If it does not fall within the first interval A, the two-way selector 323, the adder 324 and the quantizer 325 generate the intermediate fixed-point data x according to the following expression: R [n-1:0]:
[0177]
[0178] Finally, the bus converter 326 fills in the flag bit to generate the complete fixed-point data x R [n:0].
[0179] In step 1404, the calculation unit 33 calculates the error value fixed point data x R [n:0] and the weight fixed-point number x2[n:0] perform multiplication operation to generate the gradient of weight x2. The weight fixed-point number x2[n:0] can be quantized in the forward propagation process (step 1305). In this embodiment, the calculation unit 33 has Figure 8 The structure of the multiplication unit 80 is to calculate the error value fixed-point data x R [n:0] and the weight fixed-point data x2[n:0] perform multiplication operation to generate the gradient of the weight x2. As mentioned above, the flag bit of the fixed-point data data structure of the present disclosure is used to record the error value fixed-point data x R [n:0] and weight fixed-point data x2[n:0], the sign bit is used to record the error value fixed-point data x R [n:0] and the positive and negative signs of the weighted fixed-point data x2[n:0], and the numerical bit is used to record the error value fixed-point data x R [n:0] and the value V of the weighted fixed-point data x2[n:0] R , V2. Multiplier 81 converts the error value fixed-point data x R [n:0] and the sign bit of the weighted fixed-point data x2[n:0] are added to generate the sign sum value sign t The task of the first selector 821 is equivalent to dividing the parameter α R and the error value fixed point data flag x RThe task of the second selector 822 is equivalent to multiplying the second segmentation parameter α2 and the value of the flag bit x2[n] of the weight fixed-point data to generate the second parameter multiplication value pm2; finally, the fixed-point to floating-point converter 83 executes the following expression to generate the gradient of the weight x2:
[0180]
[0181] In step 1405, the quantization unit 32 performs fixed-point calculation on the gradient to generate the error value of this layer. The detailed process of performing the fixed-point calculation is as described above and will not be repeated.
[0182] In step 1406, the control unit 31 adjusts the weight x2 according to the error value of the current layer. The back propagation algorithm starts from the output layer of the neural network model, and uses the chain rule of function derivation to find the model gradient layer by layer to adjust the weight x2.
[0183] Another embodiment of the present disclosure is a method for training a neural network. The training of a neural network generally uses the output value obtained by forward propagation to calculate its error value, and then back-propagates the error value back to the input end, and adjusts the weight of each layer according to the error value of the layer, so that the reasoning structure of the neural network model is closer to the actual situation. In other words, the method for training a neural network of this embodiment includes Fig.13 The forward propagation process is to obtain the activation value of the next layer, and then the error value of the next layer is calculated according to the activation value of the next layer, and then execute Fig.13 The back propagation process is used to generate the error value of this layer, and then the appropriate weights are adjusted according to the error value of this layer. Push back all the way to obtain the appropriate weights for each layer.
[0184] Fig.15 1 is a structural diagram showing an integrated circuit device 1500 according to an embodiment of the present disclosure. Fig.15 As shown, the integrated circuit device 1500 includes a computing device 1502 , which is a neural network computing device in the aforementioned embodiments. In addition, the integrated circuit device 1500 also includes a universal interconnection interface 1504 and other processing devices 1506 .
[0185] Other processing devices 1506 may be one or more types of processors such as a central processing unit, a graphics processing unit, an artificial intelligence processor, and general and / or special processors, and their number is not limited but determined according to actual needs. Other processing devices 1506 serve as an interface between computing device 1502 and external data and control, and perform basic control including but not limited to data transfer and starting and stopping computing device 1502. Other processing devices 1506 may also cooperate with computing device 1502 to jointly complete computing tasks.
[0186] The universal interconnection interface 1504 may be used to transmit data and control instructions between the computing device 1502 and other processing devices 1506. For example, the computing device 1502 may obtain required input data from other processing devices 1506 via the universal interconnection interface 1504 and write the data into the storage unit on the computing device 1502. Further, the computing device 1502 may obtain control instructions from other processing devices 1506 via the universal interconnection interface 1504 and write the control instructions into the control cache on the computing device 1502. Alternatively or optionally, the universal interconnection interface 1504 may also read the data in the storage module of the computing device 1502 and transmit the data to the other processing devices 1506.
[0187] The integrated circuit device 1500 further includes a storage device 1508, which can be connected to the computing device 1502 and other processing devices 1506. The storage device 1508 is used to store data of the computing device 1502 and other processing devices 1506, and is particularly suitable for data that cannot be fully stored in the internal storage of the computing device 1502 or other processing devices 1506.
[0188] According to different application scenarios, the integrated circuit device 1500 can be used as a system on chip (SOC) for mobile phones, robots, drones, video acquisition devices, etc., thereby effectively reducing the core area of the control part, improving the processing speed and reducing the overall power consumption. In this case, the universal interconnection interface 1504 of the integrated circuit device 1500 is connected to certain components of the device. The certain components here can be, for example, a camera, a display, a mouse, a keyboard, a network card or a wifi interface.
[0189] The present disclosure also discloses a chip or an integrated circuit chip, which includes the integrated circuit device 1500. The present disclosure also discloses a chip packaging structure, which includes the above chip.
[0190] Another embodiment of the present disclosure is a board card, which includes the above chip packaging structure. Fig.16 In addition to the plurality of chips 1602 described above, the board 1600 may also include other supporting components, including a storage device 1604 , an interface device 1606 and a control device 1608 .
[0191] The memory device 1604 is connected to the chip 1602 in the chip package structure via a bus 1616 for storing data. The memory device 1604 may include a plurality of memory cells 1610 .
[0192] The interface device 1606 is electrically connected to the chip 1602 in the chip package structure. The interface device 1606 is used to realize data transmission between the chip 1602 and an external device 1612 (such as a server or a computer). In this embodiment, the interface device 1606 is a standard PCIe interface, and the data to be processed is transmitted from the server to the chip 1602 through the standard PCIe interface to realize data transfer. The calculation result of the chip 1602 is also transmitted back to the external device 1612 by the interface device 1606.
[0193] The control device 1608 is electrically connected to the chip 1602 so as to monitor the state of the chip 1602. Specifically, the chip 1602 and the control device 1608 may be electrically connected via an SPI interface. The control device 1608 may include a single chip microcomputer ("MCU", Micro Controller Unit).
[0194] Another embodiment of the present disclosure is an electronic device or apparatus, which includes the above-mentioned board 1600. According to different application scenarios, the electronic device or apparatus may include a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a driving recorder, a navigator, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, a headset, a mobile storage, a wearable device, a means of transportation, a household appliance, and / or a medical device. The means of transportation include an airplane, a ship and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, an electric rice cooker, a humidifier, a washing machine, an electric lamp, a gas stove, and a range hood; the medical device includes an MRI, an ultrasound machine and / or an electrocardiograph.
[0195] Another embodiment of the present disclosure is an electronic device, including one or more processors and a memory, wherein the memory stores computer executable instructions, and when the computer executable instructions are executed by the one or more processors, the electronic device executes the method as described above, in particular, executes the method as described above. Figure 7 , Fig. 9 , Fig.13 and Fig.14 The method described.
[0196] Another embodiment of the present disclosure is a computer-readable storage medium having stored thereon computer-executable instructions for calculating data in a computing device. When the computer-executable instructions are executed by one or more processors, the method as described above is executed, in particular, the method as described above is executed. Figure 7 , Fig. 9 , Fig.13 and Fig.14 The method described.
[0197] This disclosure proposes a technical solution for adaptive non-uniform low-bit fixed-point training. By introducing multiple quantization intervals, non-uniform fixed-point data formats and corresponding hardware technical means, the technical problem that the quantization bit width is not low enough to effectively reduce energy consumption is solved, thereby achieving the following technical effects:
[0198] 1. Reduce the memory usage of network models, reduce the resource consumption of model training, and improve the training speed.
[0199] 2. When using the neural network model disclosed in the present invention for reasoning, the reasoning speed can be accelerated and the accuracy degradation can be reduced.
[0200] 3. When training the neural network model disclosed in the present invention, the problems of training collapse and non-convergence can be avoided.
[0201] The foregoing content can be better understood in accordance with the following terms:
[0202] Item A1. A method for forward propagation in a neural network, comprising: receiving activation values and weights required for calculating the current layer; providing a first segmentation parameter, a first offset value and a first segmentation value corresponding to the activation value; providing a second segmentation parameter, a second offset value and a second segmentation value corresponding to the weight; quantizing the activation value according to the first segmentation parameter, the first offset value and the first segmentation value to generate a first fixed-point data; quantizing the weight according to the second segmentation parameter, the second offset value and the second segmentation value to generate a second fixed-point data; multiplying the first fixed-point data and the second fixed-point data to generate an activation value for the next layer; and repeating the above steps to perform calculations for each layer to complete the neural network.
[0203] Clause A2. A method according to clause A1, wherein the activation value falls within a first numerical distribution, and the step of quantifying the activation value comprises: based on the first segmentation parameter, segmenting the first numerical distribution into a first interval and a second interval; and determining whether the activation value falls within the first interval.
[0204] Clause A3. The method of clause A2, wherein the first segmentation value x p1 for:
[0205]
[0206] Among them, α1 is the first segmentation parameter, and D1 is the first numerical distribution.
[0207] Clause A4. The method according to clause A3, wherein if the activation value falls within the first interval, the step of quantizing the activation value generates the first fixed-point data according to the following expression:
[0208]
[0209] in, is the first fixed-point data, shift1 is the first offset value, the ceil function returns the smallest integer greater than or equal to the expression, and b is the quantization bit width.
[0210] Clause A5. A method according to clause A4, wherein the weight falls within a second numerical distribution, and the step of quantizing the weight includes: based on the second segmentation parameter, segmenting the second numerical distribution into a third interval and a fourth interval; and determining whether the weight falls within the third interval.
[0211] Clause A6. The method of clause A5, wherein the second segmentation value x p2 for:
[0212]
[0213] Among them, α2 is the second segmentation parameter, and D2 is the second numerical distribution.
[0214] Clause A7. The method according to clause A6, wherein if the weight falls within the third interval, the quantization weight step generates the second fixed-point data according to the following expression:
[0215]
[0216] in, is the second fixed-point data, and shift2 is the second offset value.
[0217] Item A8. A method according to Item A7, wherein the data structure of the first fixed-point data and the second fixed-point data includes: a flag bit for recording the interval of the first fixed-point data and the second fixed-point data; a sign bit for recording the positive and negative signs of the first fixed-point data and the second fixed-point data; and a value bit for recording the values V1, V2 of the first fixed-point data and the second fixed-point data.
[0218] Item A9. The method according to Item A8, wherein the step of performing a multiplication operation comprises: adding the numerical values of the sign bits of the first fixed-point data and the second fixed-point data to generate a sign sum value sign t ; multiplying the first segmentation parameter and the value of the flag bit of the first fixed-point data to generate a first parameter multiplication value pm1; multiplying the second segmentation parameter and the value of the flag bit of the second fixed-point data to generate a second parameter multiplication value pm2; and executing the following expression to generate the activation value of the next layer:
[0219]
[0220] Item A10. An electronic device, comprising: one or more processors; and a memory, wherein the memory stores computer executable instructions, and when the computer executable instructions are executed by the one or more processors, the electronic device executes a method as described in any one of items A1-9.
[0221] Item A11. A computer-readable storage medium comprising computer-executable instructions, which, when executed by one or more processors, execute the method described in any one of Items A1-9.
[0222] Item A12. A method of backpropagation in a neural network, wherein the neural network includes weights and fixed-point data of the weights after the weights are fixed-point processed, comprising: receiving a next-layer error value; providing a segmentation parameter, an offset value and a segmentation value corresponding to the next-layer error value; quantizing the next-layer error value according to the segmentation parameter, the offset value and the segmentation value to generate error value fixed-point data; performing multiplication operation on the error value fixed-point data and the weight fixed-point data to generate the gradient of the weight; and performing fixed-point calculation on the gradient to generate the error value of this layer; and adjusting the weight according to the error value of this layer.
[0223] Clause A13. The method according to clause A12, wherein the next layer of error values falls within a numerical distribution, and the step of quantizing the next layer of error values comprises: based on the segmentation parameter, segmenting the numerical distribution into a first interval and a second interval; and determining whether the next layer of error values falls within the first interval.
[0224] Clause A14. The method of clause A13, wherein the split value x p1 for:
[0225]
[0226] Wherein, α1 is the segmentation parameter, and D1 is the numerical distribution.
[0227] Clause A15. The method according to clause A14, wherein if the next layer error value falls within the first interval, the step of quantizing the next layer error value generates the fixed-point data according to the following expression:
[0228]
[0229] in, is the fixed-point data, shift1 is the offset value, the ceil function returns the smallest integer greater than or equal to the expression, and b is the quantization bit width.
[0230] Item A16. An electronic device, comprising: one or more processors; and a memory, wherein the memory stores computer executable instructions, and when the computer executable instructions are executed by the one or more processors, the electronic device executes a method as described in any one of Items A12-15.
[0231] Item A17. A computer-readable storage medium comprising computer-executable instructions, which, when executed by one or more processors, execute the method described in any one of Items A12-15.
[0232] Clause A18. A method of training a neural network, comprising: performing forward propagation, comprising:
[0233] Receive activation values and weights required for calculating this layer; quantize the activation values according to a first segmentation parameter, a first offset value and a first segmentation value corresponding to the activation values to generate first fixed-point data; quantize the weights according to a second segmentation parameter, a second offset value and a second segmentation value corresponding to the weights to generate second fixed-point data; multiply the first fixed-point data and the second fixed-point data to generate activation values for a next layer; calculate error values for a next layer according to the activation values for the next layer; perform back propagation, including: quantize the error values for the next layer according to a third segmentation parameter, a third offset value and a third segmentation value corresponding to the error values for the next layer to generate third fixed-point data; multiply the second fixed-point data and the third fixed-point data to generate a gradient of the weight; and perform fixed-point calculation on the gradient to generate an error value for this layer; and adjust the weights according to the error value for this layer.
[0234] Clause A19. A method according to clause A18, wherein the activation value falls within a first numerical distribution, and the step of quantifying the activation value comprises: based on the first segmentation parameter, segmenting the first numerical distribution into a first interval and a second interval; and determining whether the activation value falls within the first interval.
[0235] Clause A20. The method of clause A19, wherein the first segmentation value x p1 for:
[0236]
[0237] Among them, α1 is the first segmentation parameter, and D1 is the first numerical distribution.
[0238] Clause A21. The method according to clause A20, wherein if the activation value falls within the first interval, the step of quantizing the activation value generates the first fixed-point data according to the following expression:
[0239]
[0240] Wherein, x1 is the activation value, is the first fixed-point data, shift1 is the first offset value, the ceil function returns the smallest integer greater than or equal to the expression, and b is the quantization bit width.
[0241] Clause A22. A method according to clause A21, wherein the weight falls within a second numerical distribution, and the step of quantizing the weight includes: based on the second segmentation parameter, segmenting the second numerical distribution into a third interval and a fourth interval; and determining whether the weight falls within the third interval.
[0242] Clause A23. The method of clause A22, wherein the second segmentation value x p2 for:
[0243]
[0244] Among them, α2 is the second segmentation parameter, and D2 is the second numerical distribution.
[0245] Clause A24. The method according to clause A23, wherein if the weight falls within the third interval, the quantization weight step generates the second fixed-point data according to the following expression:
[0246]
[0247] Among them, x2 is the weight, is the second fixed-point data, and shift2 is the second offset value.
[0248] Clause A25. A method according to clause A24, wherein the next layer of error values falls within a third numerical distribution, and the step of quantizing the next layer of error values includes: based on the third segmentation parameter, segmenting the third numerical distribution into a fifth interval and a sixth interval; and determining whether the next layer of error values falls within the fifth interval.
[0249] Clause A26. The method of clause A25, wherein the third segmentation value x p3 for:
[0250]
[0251] Wherein, α3 is the third segmentation parameter, and D3 is the third numerical distribution.
[0252] Clause A27. The method according to clause A26, wherein if the next layer error value falls within the fifth interval, the step of quantizing the next layer error value generates the fixed-point data according to the following expression:
[0253]
[0254] Among them, x3 is the error value of the next layer, is the third fixed-point data, shift3 is the third offset value, and b is the quantization bit width.
[0255] Item A28. A method according to Item A27, wherein the data structure of the first fixed-point data, the second fixed-point data and the third fixed-point data includes: a flag bit for recording the interval of the first fixed-point data, the second fixed-point data and the third fixed-point data; a sign bit for recording the positive and negative signs of the first fixed-point data, the second fixed-point data and the third fixed-point data; and a value bit for recording the values V1, V2, V3 of the first fixed-point data, the second fixed-point data and the third fixed-point data.
[0256] Item A29. A method according to Item A28, wherein the step of performing a multiplication operation on the first fixed-point data and the second fixed-point data comprises: adding the numerical values of the sign bits of the first fixed-point data and the second fixed-point data to generate a sign sum value sign t ; multiplying the first segmentation parameter and the value of the flag bit of the first fixed-point data to generate a first parameter multiplication value pm1; multiplying the second segmentation parameter and the value of the flag bit of the second fixed-point data to generate a second parameter multiplication value pm2; and executing the following expression to generate the activation value of the next layer:
[0257]
[0258] Item A30. A method according to Item A28, wherein the step of performing a multiplication operation on the second fixed-point data and the third fixed-point data comprises: adding the numerical values of the sign bits of the second fixed-point data and the third fixed-point data to generate a sign sum value sign t ; multiply the second segmentation parameter and the value of the flag bit of the second fixed-point data to generate a first parameter multiplication value pm1; multiply the third segmentation parameter and the value of the flag bit of the third fixed-point data to generate a second parameter multiplication value pm2; and execute the following expression to generate the activation value of the next layer:
[0259]
[0260] Item A31. A method according to Item A25, wherein the first numerical distribution, the second numerical distribution and the third numerical distribution are Gaussian distributions, and the first interval, the third interval and the fifth interval are positive and negative symmetric intervals including 0.
[0261] Item A32. An electronic device comprising: one or more processors; and a memory, wherein the memory stores computer executable instructions, and when the computer executable instructions are executed by the one or more processors, the electronic device executes a method as described in any one of items A18-31.
[0262] Item A33. A computer-readable storage medium comprising computer-executable instructions, which, when executed by one or more processors, execute a method as described in any one of Items A18-31.
[0263] The embodiments of the present disclosure are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core idea. At the same time, for those skilled in the art, according to the ideas of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.
Claims
1. A method for forward propagation in a neural network, comprising: Receive the activation values and weights required to calculate this layer; providing a first segmentation parameter, a first offset value, and a first segmentation value corresponding to the activation value; providing a second segmentation parameter, a second offset value, and a second segmentation value corresponding to the weight; quantizing the activation value according to the first segmentation parameter, the first offset value and the first segmentation value to generate first fixed-point data; quantizing the weight according to the second segmentation parameter, the second offset value and the second segmentation value to generate second fixed-point data; Performing a multiplication operation on the first fixed-point data and the second fixed-point data to generate an activation value of a next layer; as well as Repeat the above steps to perform calculations at each layer to complete the neural network; Wherein, the activation value falls within a first numerical distribution, and the step of quantizing the activation value comprises: Based on the first segmentation parameter, segment the first value distribution into a first interval and a second interval; and Determining whether the activation value falls within the first interval; The first segmentation value for: in, is the first segmentation parameter, is the first numerical distribution; The neural network includes an input layer, a convolution layer, a pooling layer and a fully connected layer, wherein the input data of the input layer includes image data, voice data or text data.
2. The method according to claim 1, wherein if the activation value falls within the first interval, the step of quantizing the activation value generates the first fixed-point data according to the following expression: in, is the activation value, is the first fixed-point data, is the first offset value, the ceil function returns the smallest integer greater than or equal to the expression, and b is the quantization bit width.
3. The method according to claim 2, wherein the weight falls within a second numerical distribution, and the step of quantizing the weight comprises: Based on the second segmentation parameter, segment the second value distribution into a third interval and a fourth interval; as well as It is determined whether the weight falls within the third interval.
4. The method according to claim 3, wherein the second segmentation value for: in, is the second segmentation parameter, is the second numerical distribution.
5. The method according to claim 4, wherein if the weight falls within the third interval, the step of quantizing the weight generates the second fixed-point data according to the following expression: in, is the weight, For the second fixed-point data, is the second offset value.
6. The method according to claim 5, wherein the data structures of the first fixed-point data and the second fixed-point data include: A flag bit, used to record the interval of the first fixed-point data and the second fixed-point data; A sign bit, used to record the positive and negative signs of the first fixed-point data and the second fixed-point data; as well as A value bit, used to record the value of the first fixed-point data and the second fixed-point data .
7. The method according to claim 6, wherein the performing multiplication step comprises: Add the values of the sign bits of the first fixed-point data and the second fixed-point data to generate a signed sum value ; Multiply the first segmentation parameter and the value of the flag bit of the first fixed-point data to generate a first parameter multiplication value ; Multiply the second segmentation parameter and the value of the flag bit of the second fixed-point data to generate a second parameter multiplication value ; as well as The following expression is executed to generate the activation value of the next layer: 。 8. An electronic device, comprising: one or more processors; as well as A memory, wherein computer executable instructions are stored in the memory, and when the computer executable instructions are executed by the one or more processors, the electronic device executes the method according to any one of claims 1 to 7.
9. A computer-readable storage medium comprising computer-executable instructions, which, when executed by one or more processors, execute the method according to any one of claims 1 to 7.
10. A method for back propagation in a neural network, the neural network comprising weights and fixed-point weight data after the weights are fixed-point processed, comprising: Receive the next layer error value; Providing segmentation parameters, offset values and segmentation values corresponding to the next layer error value; quantizing the next layer error value according to the segmentation parameter, the offset value and the segmentation value to generate error value fixed-point data; Performing a multiplication operation on the error value fixed-point data and the weight fixed-point data to generate a gradient of the weight; as well as Performing fixed-point calculation on the gradient to generate a layer error value; as well as Adjusting the weight according to the error value of the current layer; Wherein, the next layer error value falls within the numerical distribution, and the step of quantizing the next layer error value comprises: Based on the segmentation parameter, segment the value distribution into a first interval and a second interval; and Determining whether the next layer error value falls within the first interval; The split value for: in, is the segmentation parameter, is the numerical distribution; The neural network includes an input layer, a convolution layer, a pooling layer and a fully connected layer, wherein the input data of the input layer includes image data, voice data or text data.
11. The method according to claim 10, wherein if the next layer error value falls within the first interval, the step of quantizing the next layer error value generates the fixed-point data according to the following expression: in, is the fixed-point data, is the offset value, the ceil function returns the smallest integer greater than or equal to the expression, and b is the quantization bit width.
12. An electronic device, comprising: one or more processors; as well as A memory, wherein computer executable instructions are stored in the memory, and when the computer executable instructions are executed by the one or more processors, the electronic device executes the method as claimed in claim 10 or 11.
13. A computer-readable storage medium comprising computer-executable instructions, which, when executed by one or more processors, perform the method according to claim 10 or 11.
14. A method for training a neural network, comprising: Perform forward propagation, including: Receive the activation values and weights required to calculate this layer; quantizing the activation value according to a first segmentation parameter corresponding to the activation value, a first offset value, and a first segmentation value to generate first fixed-point data; quantizing the weight value according to a second segmentation parameter corresponding to the weight value, a second offset value, and a second segmentation value to generate second fixed-point data; Performing a multiplication operation on the first fixed-point data and the second fixed-point data to generate an activation value for a next layer; Calculating the error value of the next layer according to the activation value of the next layer; Perform backpropagation, including: quantizing the next layer error value according to a third segmentation parameter, a third offset value and a third segmentation value corresponding to the next layer error value to generate third fixed-point data; performing a multiplication operation on the second fixed-point data and the third fixed-point data to generate a gradient of the weight; and performing fixed-point calculation on the gradient to generate a layer error value; and Adjusting the weight according to the error value of the current layer; Wherein, the activation value falls within a first numerical distribution, and the step of quantizing the activation value comprises: Based on the first segmentation parameter, segment the first value distribution into a first interval and a second interval; and Determining whether the activation value falls within the first interval; The first segmentation value for: in, is the first segmentation parameter, is the first numerical distribution; The neural network includes an input layer, a convolution layer, a pooling layer and a fully connected layer, wherein the input data of the input layer includes image data, voice data or text data.
15. The method according to claim 14, wherein if the activation value falls within the first interval, the step of quantizing the activation value generates the first fixed-point data according to the following expression: in, is the activation value, is the first fixed-point data, is the first offset value, the ceil function returns the smallest integer greater than or equal to the expression, and b is the quantization bit width.
16. The method according to claim 15, wherein the weight falls within a second numerical distribution, and the step of quantizing the weight comprises: Based on the second segmentation parameter, segment the second value distribution into a third interval and a fourth interval; as well as It is determined whether the weight falls within the third interval.
17. The method according to claim 16, wherein the second segmentation value for: in, is the second segmentation parameter, is the second numerical distribution.
18. The method according to claim 17, wherein if the weight falls within the third interval, the step of quantizing the weight generates the second fixed-point data according to the following expression: in, is the weight, For the second fixed-point data, is the second offset value.
19. The method according to claim 18, wherein the next layer error value falls within a third numerical distribution, and the step of quantizing the next layer error value comprises: Based on the third segmentation parameter, segment the third value distribution into a fifth interval and a sixth interval; It is determined whether the next layer error value falls within the fifth interval.
20. The method according to claim 19, wherein the third segmentation value for: in, is the third segmentation parameter, is the third numerical distribution.
21. The method according to claim 20, wherein if the next layer error value falls within the fifth interval, the step of quantizing the next layer error value generates the fixed-point data according to the following expression: in, is the error value of the next layer, is the third fixed-point data, is the third offset value, and b is the quantization bit width.
22. The method according to claim 21, wherein the data structures of the first fixed-point data, the second fixed-point data and the third fixed-point data include: A flag bit, used to record the interval of the first fixed-point data, the second fixed-point data and the third fixed-point data; A sign bit, used to record the positive and negative signs of the first fixed-point data, the second fixed-point data, and the third fixed-point data; and A value bit is used to record the values of the first fixed-point data, the second fixed-point data and the third fixed-point data. .
23. The method according to claim 22, wherein the step of performing a multiplication operation on the first fixed-point data and the second fixed-point data comprises: Add the values of the sign bits of the first fixed-point data and the second fixed-point data to generate a signed sum value ; Multiply the first segmentation parameter and the value of the flag bit of the first fixed-point data to generate a first parameter multiplication value ; Multiply the second segmentation parameter and the value of the flag bit of the second fixed-point data to generate a second parameter multiplication value ; as well as The following expression is executed to generate the activation value of the next layer: 。 24. The method according to claim 22, wherein the step of performing a multiplication operation on the second fixed-point data and the third fixed-point data comprises: Add the values of the sign bits of the second fixed-point data and the third fixed-point data to generate a sign sum value ; Multiply the second segmentation parameter and the value of the flag bit of the second fixed-point data to generate the first parameter multiplication value ; The third segmentation parameter and the value of the flag bit of the third fixed-point data are multiplied to generate a second parameter multiplication value. ; as well as The following expression is executed to generate the activation value of the next layer: 。 25 . The method according to claim 19 , wherein the first numerical distribution, the second numerical distribution, and the third numerical distribution are Gaussian distributions, and the first interval, the third interval, and the fifth interval are positive and negative symmetric intervals including 0.
26. An electronic device comprising: one or more processors; as well as A memory, wherein computer executable instructions are stored in the memory, and when the computer executable instructions are executed by the one or more processors, the electronic device executes the method as described in any one of claims 14-25.
27. A computer-readable storage medium comprising computer-executable instructions, which, when executed by one or more processors, perform the method according to any one of claims 14 to 25.
Citation Information
Patent Citations
Quantification realization method and related product
CN109993296A
Data processing method and device, computer equipment and storage medium
CN110889503A
Computing device for neural network operation and integrated circuit board card thereof
CN111027691A