A Low-Bit Neural Network Training Method and Accelerator
Through the low-bit neural network training method and accelerator designed in the feature channel dimension expansion and full addition convolution design, the existing low-bit convolution neural network has been solved, and the computing unit is complex is achieved, efficient and low-power computing capabilities are achieved, and the deployment capabilities on small resource platforms are increased.
Patent Information
- Application Number
- CN202411282693.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-09-13
AI Technical Summary
The existing low-bit convolutional neural network has low accuracy and complex computing unit design, making it difficult to achieve high computing efficiency while maintaining low power consumption.
A low-bit neural network training method is proposed, and the flexibility of convolution calculation is realized by unfolding in the characteristic channel dimension, and a fully additive convolution accelerator is designed, which avoids the demand for multiplier by traditional convolution, and realizes a low-bit convolution accelerator with simple computing logic, low power consumption, high efficiency and high flexibility.
It effectively avoids the loss of neural network accuracy, improves training accuracy and computing efficiency, reduces the demand for hardware platforms, and increases the deployment capabilities of neural networks on small resource platforms.
Smart Images

Figure CN119227746B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network acceleration, and proposes a low-bit neural network training method and accelerator. Background Art
[0002] With the rapid development of artificial intelligence technology, neural network models are increasingly widely used in various applications. These models usually require a large amount of computing resources. In mobile devices and embedded systems, computing resources have become a key factor restricting the application of models. To solve this problem, researchers have proposed a variety of neural network compression and acceleration technologies.
[0003] In the existing technologies, there are mainly two methods to reduce the computational complexity of neural networks: one is network pruning, which reduces the amount of computation by removing unimportant connections in the network; the other is quantization, that is, reducing the number of bits of the weights in the network, thereby reducing the storage requirements and computational complexity.
[0004] Although the existing quantization technologies can reduce the computational complexity of neural networks to a certain extent, these methods generally have the following problems: the existing quantization technologies often lead to a significant decline in the performance of neural networks, especially in low-bit quantization (such as below four bits); the design of existing neural network computing units is complex, and it is difficult to achieve high computational efficiency while maintaining low power consumption; the existing technologies have limitations in hardware implementation, such as poor flexibility and insufficient scalability.
[0005] In view of the above problems, this application is committed to proposing a low-bit neural network training method and accelerator. Summary of the Invention
[0006] The purpose of the present invention is to solve the problems that the accuracy of existing low-bit convolutional neural networks is relatively low and the design of convolutional neural network computing units is complex, and it is difficult to achieve high computational efficiency while maintaining low power consumption, and proposes a low-bit neural network training method and accelerator. The method avoids excessive loss of the accuracy of the neural network while quantizing the neural network to a lower bit width; the accelerator provides hardware computing support for the method, realizes the flexibility of convolutional computing by expanding in the feature channel dimension, and can adapt to convolutional computing of different sizes; a configurable adder tree realizes configurable computing parallelism at runtime, improving the versatility and flexibility of the accelerator; the accelerator also proposes a design of full adder convolution, avoiding the need for multipliers in traditional convolution, reducing the requirements of the accelerator for the hardware platform, and realizing a low-bit convolutional accelerator with simple computing logic, low power consumption, high efficiency, and high flexibility, partially solving the problem that it is difficult to deploy neural networks on small resource platforms.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] As a first aspect of the present invention, a method for training a low-bit neural network is proposed, including the following steps:
[0009] S1. Initialize the parameters of the full-precision neural network model embedded with quantization blocks, including: initializing the full-precision neural network model of the quantization block with the parameters of the full-precision neural network model; initializing the trainable parameters of the quantization block with random numbers;
[0010] The quantization block includes a first quantization block and a second quantization block, and both include division, rounding, and multiplication;
[0011] S2. Perform forward inference on the full-precision neural network model embedded with quantization blocks. The specific operation process for a certain layer is as follows: the weights of each convolutional layer are quantized by the first quantization block and then convolved with the input feature map; the convolution result is quantized by the second quantization block to obtain the output feature map;
[0012] S3. Backpropagate and calculate the gradients of all layer parameters in the full-precision neural network model embedded with quantization blocks; S3, specifically in implementation, the operation process for a certain layer is as follows:
[0013] S31. Obtain the gradient of the loss function and backpropagate it to obtain the gradient of the output feature map of this convolutional layer;
[0014] S32. Calculate the partial gradients of the quantization step size and the gradient of the result after rounding by taking the gradient of the multiplication in the second quantization block;
[0015] S33. Use the gradient approximation function of the rounding function to replace the gradient of the rounding function in the second quantization block, and obtain the gradient of the result after division;
[0016] S34. Calculate the complete gradient of the quantization step size and the gradient of the input of the second quantization block by taking the gradient of the division in the second quantization block;
[0017] S35. Calculate the gradients of the quantization bit width and quantization range of the second quantization block;
[0018] S36. Replace the second quantization block in S32 to S35 with the first quantization block, and repeat S32 to S35 to obtain the gradient of the input of the first quantization block and the gradients of the quantization bit width and quantization range of the first quantization block;
[0019] S4. Update all the trainable parameters according to the calculated gradients.
[0020] The method can quantize the neural network to a lower bit width; avoid excessive loss of the accuracy of the neural network;
[0021] The full-precision neural network model with embedded quantization blocks includes a network structure and the weights of each network layer; the trainable parameters of the quantization blocks include the quantization bit width and the quantization range.
[0022] Each convolutional layer in the full-precision neural network model with embedded quantization blocks described in S1 embeds a first quantization block and a second quantization block.
[0023] The first quantization block quantizes the input weights to obtain quantized input weights, and then these weights are convolved with the input feature map to obtain an input quantized convolution result; the second quantization block quantizes the input quantized convolution result to obtain an output feature map, which is the output of this convolutional layer.
[0024] The quantization described in S2 includes the following sub-processes:
[0025] S21. Calculate the quantization step size from the trainable parameters;
[0026] S22. Divide the input of the quantization block by the quantization step size, round the obtained result, and then multiply by the quantization step size to obtain the quantized parameter.
[0027] The input of the quantization block corresponds to the weight of the convolutional layer for the first quantization block; for the second quantization block, it is the convolution result obtained by convolving the quantized weight of this convolutional layer with the input feature map.
[0028] The input of the first quantization block described in S36 is the weight of this convolutional layer; all the trainable parameters described in S4 include the parameters of the full-precision neural network model with embedded quantization blocks and the training parameters of the quantization blocks.
[0029] A full adder low-bit neural network accelerator is connected to unified memory and a general-purpose processor, and includes a hardware data rearrangement module, a cache scheduling module, an input feature map cache, a parallel computing array, and an output feature map cache; the hardware data rearrangement module is connected to the unified memory and the cache scheduling module; the cache scheduling module is connected to the hardware data rearrangement module, the general-purpose processor, the input feature map cache, and the output feature map cache; the input feature map cache is connected to the unified memory, the cache scheduling module, and the parallel computing array; the output feature map cache is connected to the unified memory, the cache scheduling module, and the parallel computing array; the parallel computing array is connected to the cache scheduling module, the input feature map cache, and the output feature map cache.
[0030] The hardware data rearrangement module receives the size and data address information of the input feature map of each layer from the cache scheduling module, reads the data information of the current layer input feature map from the unified memory according to the size and data address information, and stores the data back in the unified memory after converting the data arrangement method into the order of height, width, and number of feature channels;
[0031] The cache scheduling module receives the input feature map size, input feature map data address, weight size, weight data address, output feature map size, output feature map data address, and current layer calculation type information from the general - purpose processor. At the beginning of each layer's calculation, it sends the input feature map size and input feature map data address to the hardware data rearrangement module; calculates the parallelism division information of the parallel computing array according to the input feature map size and weight size and sends it to the parallel computing array; calculates different input cache read - in addresses and input cache read lengths according to the parallelism division information, input feature map size, input feature map data address, weight size, and weight data address, and sends them to the input feature map cache; calculates different output cache write addresses and output cache write lengths according to the parallelism division information, output feature map size, and output feature map data address, and sends them to the output feature map cache;
[0032] The input feature map cache receives information such as the input cache read - in address and input cache read length from the cache scheduling module, reads the block - divided input feature map data and weight data from the unified memory, and sends them to the parallel computing array;
[0033] The output feature map cache receives information such as the output cache write address and output cache write length from the cache scheduling module, and sends the block - divided output feature map data from the parallel computing array to the unified memory;
[0034] The parallel computing array receives the parallelism division information from the cache scheduling module, configures the parallelism of parallel computing in the input channel and output channel according to the parallelism division information; reads the block - divided input feature map data and weight data from the input feature map cache for convolution calculation, and after the calculation is completed, stores the output feature map in the output feature map cache;
[0035] The parallel computing array includes a weight cache array, a first - level adder tree array, a second - level adder tree unit, an accumulator array, and a quantization unit array; the weight cache array is connected to the input feature map cache and the first - level adder tree array, the first - level adder tree array is connected to the input feature map cache and the weight cache array; the second - level adder tree unit is connected to the first - level adder tree array and the accumulator array; the accumulator array is connected to the second - level adder tree unit and the quantization unit array; the quantization unit array is connected to the accumulator array and the output feature map cache.
[0036] The weight cache array includes a number of weight caches; each first-level adder tree array includes a number of first-level adder trees; each first-level adder tree further includes three adders, which can implement low-bit multiplication calculation; the second-level adder tree unit includes an adder stage controller and a number of stages of adders; the adder stage controller is used to enable the number of stages of the adder; the accumulator array includes a number of accumulators; the quantization unit array includes a number of quantization units; the accumulator array receives the addition result from the second-level adder tree unit, and after accumulating for a specified number of cycles, passes the accumulated result to the quantization unit array; the quantization unit array receives the accumulated result from the accumulator array, performs quantization operations, obtains the output feature map result, and sends it to the output feature map cache.
[0037] The weight cache array receives the weight data passed by the input feature map cache, temporarily stores the weight data in the weight cache, and passes it to the first-level adder tree array when needed; the first-level adder tree array reads the input feature map data and weight data from the input feature map cache and the weight cache, calculates the multiplication result, and then passes the multiplication result to the second-level adder tree unit; the second-level adder tree unit receives the multiplication result from the first-level adder tree array, performs addition operations for a specified number of stages according to the configuration information, and then passes the addition result to the accumulator array;
[0038] The accelerator provides hardware computing support for the method, realizes the flexibility of convolution calculation by expanding in the feature channel dimension, and can adapt to different sizes of convolution calculations; the accelerator also proposes a design of full adder convolution, avoiding the need for multipliers in traditional convolutions, reducing the requirements of the accelerator for the hardware platform, and realizing a low-bit convolution accelerator with simple calculation logic, low power consumption, high efficiency, and high flexibility, partially solving the problem that it is difficult to deploy neural networks on small resource platforms. A configurable adder tree design realizes configurable computing parallelism at runtime, improving the versatility and flexibility of the accelerator.
[0039] Advantages:
[0040] The present invention proposes a low-bit neural network training method and an accelerator, which have the following advantages compared with the existing low-bit neural networks:
[0041] 1. The low-bit neural network training method proposes a rounding function gradient approximation function, which improves the gradient mismatch problem caused by using the approximate straight-through estimator method in the traditional low-bit neural network training method, and improves the training accuracy of the low-bit neural network;
[0042] 2. The low-bit neural network accelerator can configure the parallelism of convolutional calculations in the input channel dimension and the output channel dimension during runtime, improving the problem of low generality of traditional neural network accelerators. This enables the low-bit neural network accelerator to simultaneously calculate convolutions with different numbers of feature map channels and different groupings, enhancing the generality and computational speed of the neural network accelerator.
[0043] 3. The bit widths of the input feature map data, weight data, and output feature map data for the convolutional calculations of the low-bit neural network accelerator are all four bits. Compared with traditional neural network accelerators that use eight bits, the low-bit neural network accelerator can consume fewer resources and achieve higher parallelism.
[0044] 4. The low-bit neural network accelerator uses full addition operations to calculate convolutions, reducing the demand for computational power of the neural network accelerator. This allows the low-bit neural network accelerator to be deployed on a wider range of hardware devices, expanding the scope of use of the neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 Schematic diagram of the low-bit convolutional calculation method during the training process of the present invention;
[0046] Figure 2 Schematic diagram of the quantization block calculation method of the present invention;
[0047] Figure 3 Schematic diagram of the rounding function gradient approximation function of the present invention;
[0048] Figure 4 Schematic diagram of the accelerator structure of the present invention;
[0049] Figure 5 Schematic diagram of the parallel computing array structure of the present invention;
[0050] Figure 6 Schematic diagram of the structure of the secondary adder tree unit of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following elaborates in detail on the deployment, actual application process, and advantages (beneficial effects) of a low-bit convolutional neural network computing unit of the present invention in conjunction with the accompanying drawings and embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the present invention, various equivalent forms of modifications made by those skilled in the art fall within the scope defined by the appended claims of this application.
[0052] Example 1
[0053] This example elaborates on the specific implementation steps of the described low-bit neural network training method.
[0054] A low-bit neural network training method, relying on the existing low-bit neural network training methods, proposes a rounding function gradient approximation function, including the following steps:
[0055] Step1. Initialize the parameters of the low-bit neural network training model using the full-precision neural network model parameters. At the same time, initialize the trainable parameters of the quantization block in the low-bit neural network training model with random numbers. The trainable parameters of the quantization block include the quantization bit width and the quantization range.
[0056] Step2. Perform forward inference of the neural network. As shown in the appendix Figure 1 As shown, in each layer of convolution, the weights are quantized by the quantization block and then convolved with the input feature map. The result of the convolution is then quantized by the quantization block to obtain the output feature map. As shown in the appendix Figure 2 As shown, the quantization process includes the following sub-processes:
[0057] Step21. Calculate the quantization step size from the trainable parameter quantization bit width and the quantization range.
[0058] Step22. Divide the input parameter by the quantization step size, round the obtained result, and then multiply by the quantization step size to obtain the quantized parameter.
[0059] Step3. Perform backward gradient calculation and parameter update of the neural network. The process includes the following sub-processes:
[0060] Step31. According to the chain rule, calculate the gradient of the loss function to obtain the gradient of the output feature map.
[0061] Step32. According to the chain rule, calculate the gradient of the multiplication in the quantization block to obtain the partial gradient of the quantization step size and the gradient of the rounded result.
[0062] Step33. Use the rounding function gradient approximation function to replace the gradient of the rounding function in the quantization block, and calculate the gradient of the result after division according to the chain rule.
[0063] Step34. According to the chain rule, calculate the gradient of the division in the quantization block to obtain the gradient of the quantization step size and the gradient of the input parameter.
[0064] Step35. According to the chain rule, calculate the gradients of the quantization bit width and the quantization range.
[0065] Step36. Update all the trainable parameters according to the calculated gradients.
[0066] A full adder low-bit neural network accelerator, which proposes a parallel computing array with a runtime-configurable computing parallelism and a full adder convolution. It is connected to a unified memory and a general-purpose processor, and includes a hardware data rearrangement module, a cache scheduling module, an input feature map cache, an output feature map cache, and a parallel computing array.
[0067] The parallel computing array includes a weight cache array, a first-level adder tree array, a second-level adder tree unit, an accumulator array, and a quantization unit array.
[0068] The first-level adder tree is composed of three adders and can implement low-bit multiplication calculation.
[0069] The second-level adder tree unit is composed of several levels of adders, and the number of enabled adder levels can be configured during runtime to implement grouped convolution operations.
[0070] The low-bit quantization module includes a shift controller and a shifter.
[0071] Regarding Advantage 1, the advantages of the low-bit neural network training method are as follows:
[0072] In the inference calculation of a low-bit neural network, it is necessary to quantize the weights and convolution results, and one of the steps is rounding. The operation of the rounding function is to round the input data. The problem is that the rounding function is not a continuous function, so its derivative function cannot be obtained, resulting in the inability to calculate the gradient of the trainable parameters during the training process.
[0073] In the traditional inference calculation of a low-bit neural network, the way to solve the above problem is to use the approximate straight-through estimation method, and replace the derivative function of the rounding function with the constant function g(x)=1. That is, use the gradient of the rounded result as the gradient of the value before rounding. In this way, the gradient of each trainable parameter can be obtained.
[0074] However, this method has a disadvantage that there is a large error between the true derivative function and the constant function used in the approximate straight-through estimation method. By introducing the concept of the impulse function, the derivative function of the rounding function is the sum of the impulse functions at all integers plus 0.5, and there is a large error between the constant function g(x)=1 used in the approximate straight-through estimation method and this derivative function. As the number of layers of the neural network increases, this error will gradually accumulate, resulting in a large gradient mismatch problem for the parameters during the training process. This leads to unsatisfactory training effects for traditional low-bit neural networks.
[0075] In the low-bit convolutional neural network training method proposed in the present invention, the approximate straight-through estimation method is no longer used to approximate the gradient of the rounding function. Instead, use Figure 3The gradient approximation function of the rounding function is used to replace the gradient of the rounding function. Compared with the constant function g(x)=1, the gradient approximation function of the rounding function is closer to the derivative function of the true rounding function, brings less error in the gradient calculation process of backpropagation, and the gradient of each training parameter calculated is closer to the true gradient, thus obtaining a better training effect for the low-bit neural network.
[0076] Embodiment 2
[0077] This embodiment elaborates on the specific implementation steps of the low-bit neural network accelerator.
[0078] The full adder low-bit neural network accelerator is connected to unified memory and a general-purpose processor, and includes a hardware data rearrangement module, a cache scheduling module, an input feature map cache, an output feature map cache, and a parallel computing array, as shown in the appendix Figure 4 shown.
[0079] The hardware data rearrangement module is connected to the unified memory and the cache scheduling module. The cache scheduling module is connected to the hardware data rearrangement module, the general-purpose processor, the input feature map cache, and the output feature map cache. The input feature map cache is connected to the unified memory, the cache scheduling module, and the parallel computing array. The output feature map cache is connected to the unified memory, the cache scheduling module, and the parallel computing array. The parallel computing array is connected to the cache scheduling module, the input feature map cache, and the output feature map cache.
[0080] The hardware data rearrangement module receives the size and data address information of the input feature map of each layer from the cache scheduling module, reads the data information of the current layer input feature map from the unified memory according to this information, and stores it back in the unified memory after converting the data arrangement method into the order of height, width, and number of feature channels.
[0081] The cache scheduling module receives information such as the input feature map size, input feature map data address, weight size, weight data address, output feature map size, output feature map data address, and current layer calculation type from the general-purpose processor. At the beginning of each layer of calculation, it sends the input feature map size and input feature map data address to the hardware data rearrangement module. It sends the parallelism division information of the parallel computing array calculated according to the input feature map size and weight size to the parallel computing array. It calculates different input cache read addresses and input cache read lengths according to the parallelism division information, input feature map size, input feature map data address, weight size, and weight data address, and sends them to the input feature map cache. It calculates different output cache write addresses and output cache write lengths according to the parallelism division information, output feature map size, and output feature map data address, and sends them to the output feature map cache.
[0082] The input feature map cache receives information such as the input cache read address and the input cache read length from the cache scheduling module, reads the block-based input feature map data and weight data from the unified memory, and sends them to the parallel computing array.
[0083] The output feature map cache receives information such as the output cache write address and the output cache write length from the cache scheduling module, and sends the block-based output feature map data from the parallel computing array to the unified memory.
[0084] The parallel computing array receives the parallelism division information from the cache scheduling module, configures the parallelism for parallel computing in the input channel and the output channel according to the parallelism division information. It reads the block-based input feature map data and weight data from the input feature map cache for convolution calculation, and stores the obtained output feature map in the output feature map cache after the calculation is completed.
[0085] The parallel computing array includes a weight cache array, a first-level adder tree array, a second-level adder tree unit, an accumulator array, and a quantization unit array, as shown in the appendix. Figure 5 as shown.
[0086] The weight cache array is connected to the input feature map cache and the first-level adder tree. It includes several weight caches.
[0087] The weight cache array receives the weight data passed from the input feature map cache, temporarily stores the weight data in the weight cache, and passes it to the first-level adder tree array when needed.
[0088] The first-level adder tree array is connected to the input feature map cache and the weight cache array, and is composed of several first-level adder trees; each first-level adder tree includes three adders and can perform low-bit multiplication calculations.
[0089] The first-level adder tree array reads the input feature map data and weight data from the input feature map cache and the weight cache, and passes the multiplication result to the second-level adder tree unit after calculating the multiplication result.
[0090] The second-level adder tree unit is connected to the first-level adder tree array and the accumulator array, and is composed of configuration information and several levels of adders, as shown in the appendix. Figure 6 The configuration information includes the enabled number of levels of the adder.
[0091] The second-level adder tree unit receives the multiplication result from the first-level adder tree array, performs the addition operation for the specified number of levels according to the configuration information, and then passes the addition result to the accumulator array.
[0092] The accumulator array, connected to the second-level adder tree unit and the quantization unit array, consists of a number of accumulators.
[0093] The accumulator array receives the addition results from the second-level adder tree unit. After accumulating for a specified number of cycles, it passes the accumulated results to the quantization unit array.
[0094] The quantization unit array, connected to the accumulator array and the output feature map cache, consists of a number of quantization units.
[0095] The quantization unit array receives the accumulated results from the accumulator array. After performing quantization operations, it obtains the output feature map results and sends them to the output feature map cache.
[0096] During the process of calculating the low-bit neural network, parameters such as the network structure, weights, and quantization scale of each layer of the low-bit neural network are pre-stored in the unified memory.
[0097] At the beginning of each layer's calculation, the general-purpose processor passes data such as the input feature map address, input feature map size, weight address, weight size, output feature map address, and output feature map size to the cache scheduling module. The cache scheduling module divides the feature map data into blocks according to the input feature map size and controls the parallelism in the input channel dimension and output channel dimension of the parallel computing array.
[0098] The cache scheduling module controls the hardware data rearrangement module to read the input feature map data from the unified memory according to the input feature map address and input feature map size, and changes the arrangement form of the input feature map data from the organization form of channel number, height, and width to the organization form of height, width, and channel number.
[0099] At the same time, the cache scheduling module reads the weight data according to the weight address, weight size, and output channel dimension parallelism and stores it in the weight cache array in the parallel computing array.
[0100] After the weight data is read, the cache scheduling module controls the input feature map cache to read the input feature map data from the unified memory according to the input feature map address and input feature map size. After reading in enough data volume, the parallel computing array starts to calculate. After reading in the input feature map data once, a part of the output feature map data is calculated.
[0101] After the parallel computing array finishes the calculation, it stores the calculation results in the output feature map cache. After storing enough data volume, the output feature map cache stores the output feature map data in the unified memory according to the output feature map address and output feature map size.
[0102] Regarding Advantage 2, the advantages of the low-bit neural network accelerator are as follows:
[0103] In a traditional neural network accelerator, if the overall parallelism of the parallel computing array is 1024, the parallelism of the input channel dimension is set to 16, and the parallelism of the output channel dimension is set to 64. When calculating the first layer convolution of the Resnet-18 neural network, after reading in the weight data once and reading through the neural network input feature map data once, the result can be obtained. The output data requires a total of 50176 computing clock cycles and two data accesses. When calculating the third layer convolution of the Resnet-18 neural network, it is necessary to read in the weight data twice and read in the input feature map data twice to calculate the result. The output data requires a total of 18432 computing clock cycles and four data accesses. When calculating the fifth layer convolution of the Resnet-18 neural network, it is necessary to read in the weight data four times and read in the input feature map data four times to calculate the result. The output data requires a total of 18432 computing clock cycles and eight data accesses.
[0104] The low-bit neural network accelerator performs parallel computing in the input channel and output channel dimensions. The parallelism in each channel dimension can be configured at runtime. If the overall parallelism of the parallel computing array is 1024, when calculating the first layer convolution of the Resnet-18 neural network, the parallelism of the input channel dimension can be set to 16, and the parallelism of the output channel dimension can be set to 64. After reading in the weight data once and reading through the neural network input feature map data once, the result can be obtained. The output data requires a total of 50176 computing clock cycles and two data accesses, which is the same as that of the traditional neural network accelerator. When calculating the third layer convolution of the Resnet-18 neural network, the parallelism of the input channel dimension can be set to 8, and the parallelism of the output channel dimension can be set to 128. After reading in the weight data once and reading through the neural network input feature map data once, the result can be obtained. The output data requires a total of 18432 computing clock cycles and two data accesses, which is two data access times less than that of the traditional neural network accelerator. When calculating the fifth layer convolution of the Resnet-18 neural network, the parallelism of the input channel dimension can be set to 4, and the parallelism of the output channel dimension can be set to 256. After reading in the weight data once and reading through the neural network input feature map data once, the result can be obtained. The output data requires a total of 18432 computing clock cycles and two data accesses, which is six data access times less than that of the traditional neural network accelerator.
[0105] If the input channel dimension parallelism of a traditional neural network accelerator is set to 4 and the output channel dimension parallelism is set to 256, when calculating the first-layer convolution of the Resnet-18 neural network, the result can be obtained after reading the weight data once and reading the neural network input feature map data once. However, it will require 200,704 computing clock cycles, which is 150,528 more computing clock cycles than the low-bit neural network accelerator described in the present invention. When calculating the third-layer convolution of the Resnet-18 neural network, the result can be obtained after reading the weight data once and reading the neural network input feature map data once. However, it will require 36,864 computing clock cycles, which is 18,432 more computing clock cycles than the low-bit neural network accelerator described in the present invention.
[0106] For the grouped convolution and depth convolution commonly used in the mobileNet neural network, a traditional neural network accelerator cannot calculate these two types of convolutions simultaneously. When using the traditional neural network accelerator solution, accelerators for grouped convolution and depth convolution need to be deployed simultaneously.
[0107] When using the low-bit neural network accelerator described in the present invention, a lower input channel parallelism can be configured during the calculation of grouped convolution at runtime, and a higher input channel parallelism can be configured during the calculation of depth convolution. Only one low-bit neural network accelerator described in the present invention needs to be deployed.
[0108] The low-bit neural network accelerator can configure the parallelism of convolution calculation in the input channel dimension and output channel dimension at runtime, improving the problem of low generality of traditional neural network accelerators, enabling the low-bit neural network accelerator to calculate convolutions with different feature map channel numbers and different groups simultaneously, and improving the generality and computing speed of the neural network accelerator.
[0109] Regarding beneficial effect 3, the advantages of the low-bit neural network accelerator are as follows:
[0110] In the neural network accelerators with parallel computing in the input channel and output channel dimensions in the past, the bit widths of the input feature map data and output feature map data were generally eight bits, and each convolution layer's feature map data and weight data had a floating-point scaling factor and a floating-point zero offset respectively. The bit widths of the two parts of the input data of the multiplication calculation unit were both eight bits, and the output data bit width was sixteen bits.
[0111] The advantage of this implementation method is that the neural network has a relatively high accuracy calculation and good performance. However, this implementation method has two disadvantages: First, using eight bits to represent each feature map data and each weight data, the overall data volume is relatively large, which requires relatively high storage resources and bandwidth resources of the platform. Second, the bit widths of the two parts of the input data of the multiplication calculation unit are both eight bits, and the required calculation amount is relatively large, which requires relatively high calculation resources of the platform.
[0112] After changing the bit widths of the feature map data and the weight data of the convolutional layer to four bits in the low-bit neural network accelerator according to the present invention, relying on the low-bit neural network training method according to the present invention, weights with a four-bit width can be obtained, which can better maintain the accuracy of the neural network. The size of the network model can be reduced to half of the original, and the storage spaces required for the feature map data and the weight data of the entire neural network are both reduced to half of the original, greatly reducing the storage resource requirements and bandwidth resource requirements of the traditional convolutional neural network calculation unit for the platform, and enabling the same neural network to be deployed on a platform with more limited resources.
[0113] At the same time, after changing the bit widths of the feature map data and the weight data of the convolutional layer to four bits, the bit widths of the two parts of the input data of the multiplication calculation unit are both reduced to four bits. Compared with the multiplication calculation unit with an eight-bit width, the calculation resources required by the multiplication calculation unit with a four-bit width are only one-fourth of it. And if the same calculation resources are used, the parallelism can be achieved four times that of the original.
[0114] Therefore, after changing the bit widths of the feature map data and the weight data of the convolutional layer to four bits, the consumption of storage resources, bandwidth resources, and calculation resources of the deployment platform is greatly reduced, solving the problem that the neural network accelerators for parallel calculation of input channel and output channel dimensions in the past have excessive requirements for storage resources, bandwidth resources, and digital signal processors. On a platform with the same amount of resources, compared with the neural network accelerators for parallel calculation of input channel and output channel dimensions in the past, a higher parallelism can be achieved.
[0115] Regarding beneficial effect 4, the advantages of the low-bit neural network accelerator are as follows:
[0116] In traditional neural network accelerators, multiplication calculations are all implemented using multipliers. The advantage of this method is that the multiplier has good timing and a relatively large operation bit width. However, the disadvantage is that it consumes calculation resources, has relatively high requirements for the calculation platform, and requires the calculation platform to have the ability to carry a multiplier.
[0117] The low-bit neural network accelerator according to the present invention designs a full adder convolutional calculation unit. As shown in the appendix Figure 5As shown, the parallel computing array in the low-bit neural network accelerator includes a weight cache array, a first-level adder tree array, a second-level adder tree unit, an accumulator array, and a quantization unit array. The first-level adder tree array implements a four-bit multiplication function, replacing each multiplier with three adders. Only addition operations are used to implement the convolution calculation function.
[0118] The calculation method eliminates the dependence of traditional neural network accelerators on multipliers, enabling the low-bit neural network accelerator to only implement addition operations and be deployed on hardware platforms without multipliers, improving the versatility of the neural network accelerator.
[0119] In summary, the low-bit neural network accelerator of the present invention reduces the computing power requirements of the neural network accelerator, enabling the low-bit neural network accelerator to be deployed on a wider range of hardware devices and increasing the scope of use of neural networks.
Claims
1. A full-additive low-bit neural network accelerator connected to a unified memory and a general-purpose processor, characterized in that: It includes a hardware data reordering module, a cache scheduling module, an input feature map cache, a parallel computing array and an output feature map cache; The hardware data rearrangement module is connected to the unified memory and the cache scheduling module; the cache scheduling module is connected to the hardware data rearrangement module, the general processor, the input feature map cache and the output feature map cache; the input feature map cache is connected to the unified memory, the cache scheduling module and the parallel computing array; the output feature map cache is connected to the unified memory, the cache scheduling module and the parallel computing array; the parallel computing array is connected to the cache scheduling module, the input feature map cache and the output feature map cache; The hardware data rearrangement module receives the size and data address information of each layer of input feature map from the cache scheduling module, reads the data information of the current layer input feature map from the unified memory according to the size and data address information, converts the arrangement of the data information into the order of height, width, and number of feature channels, and stores it back into the unified memory; The cache scheduling module receives the input feature map size, input feature map data address, weight size, weight data address, output feature map size, output feature map data address, and current layer calculation type from the general processor; at the beginning of each layer calculation, the input feature map size and the input feature map data address are sent to the hardware data rearrangement module; the parallelism division information of the parallel computing array is calculated according to the input feature map size and the weight size and sent to the parallel computing array; different input cache read addresses and input cache read lengths are calculated according to the parallelism division information, the input feature map size, the input feature map data address, the weight size and the weight data address, and sent to the input feature map cache; different output cache write addresses and output cache write lengths are calculated according to the parallelism division information, the output feature map size, and the output feature map data address, and sent to the output feature map cache; The input feature map cache receives the input cache read address and input cache read length information from the cache scheduling module, reads the block input feature map data and weight data from the unified memory, and sends them to the parallel computing array; The output feature map cache receives the output cache write address and output cache write length information from the cache scheduling module, and sends the block output feature map data from the parallel computing array to the unified memory; The parallel computing array receives the parallelism division information from the cache scheduling module, and configures the parallelism of the parallel computing of the input channel and the output channel according to the parallelism division information; reads the block input feature map data and weight data from the input feature map cache for convolution calculation, and after the calculation is completed, obtains the output feature map and stores it in the output feature map cache.
2. The full-addition low-bit neural network accelerator according to claim 1, characterized in that: The parallel computing array comprises a weight cache array, a first-level adder tree array, a second-level adder tree unit, an accumulator array and a quantization unit array; the weight cache array is connected to the input feature map cache and the first-level adder tree array, and the first-level adder tree array is connected to the input feature map cache and the weight cache array; The secondary addition tree unit is connected to the primary addition tree array and the accumulator array; the accumulator array is connected to the secondary addition tree unit and the quantization unit array; the quantization unit array is connected to the accumulator array and the output feature map buffer.
3. The full-addition low-bit neural network accelerator according to claim 1, characterized in that: The weight cache array includes a plurality of weight caches; each first-level adder tree array includes a plurality of first-level adder trees; each first-level adder tree includes three adders, which can realize low-bit multiplication calculation; the second-level adder tree unit includes an adder level controller and a plurality of adders; the adder level controller is used to enable the adder level; the accumulator array includes a plurality of accumulators; the quantization unit array includes a plurality of quantization units; the accumulator array receives the addition result from the second-level adder tree unit, and after performing the accumulation of the specified number of cycles, passes the accumulation result to the quantization unit array; The quantization unit array receives the accumulation result from the accumulator array, performs a quantization operation, obtains an output feature map result, and sends it to the output feature map cache.
4. The full-addition low-bit neural network accelerator according to claim 1, characterized in that: A low-bit neural network training method is applied to obtain a full-additive low-bit neural network accelerator. The low-bit neural network training method includes the following steps: S1. Initializing a full-precision neural network model embedded in a quantization block, including: initializing the full-precision neural network model embedded in a quantization block using parameters of the full-precision neural network model; initializing trainable parameters of the quantization block using random numbers; The quantization blocks include a first quantization block and a second quantization block, and both include division, rounding and multiplication; S2. Forward inference is performed on the full-precision neural network model embedded with the quantization block. The specific operation process for a certain layer is as follows: the weight of each convolutional layer is quantized by the first quantization block and then convolved with the input feature map; the convolution result is then quantized by the second quantization block to obtain the output feature map; S3, back-propagate and calculate the gradients of all layer parameters in the full-precision neural network model embedded in the quantization block; S3, when implemented, the operation process for a certain layer is as follows: S31. Calculate the gradient of the loss function and back-propagate to obtain the gradient of the output feature map of the convolutional layer; S32, calculating the gradient of the multiplication in the second quantization block to obtain a partial gradient of the quantization step length and a gradient of a rounded result; S33, using the rounding function gradient approximation function to replace the gradient of the rounding function in the second quantization block, to obtain the gradient of the result after the division; S34, calculating the gradient of the division in the second quantization block to obtain the complete gradient of the quantization step length and the gradient of the input of the second quantization block; S35, calculating the quantization bit width and the gradient of the quantization range of the second quantization block; S36, replacing the second quantization block in S32 to S35 with the first quantization block, repeating S32 to S35, and obtaining the gradient of the first quantization block input and the gradient of the quantization bit width and quantization range of the first quantization block; S4. Update all trainable parameters according to the calculated gradients.
5. The full-addition low-bit neural network accelerator according to claim 4, characterized in that: S1: The full-precision neural network model of the embedded quantization block includes the network structure and the weights of each network layer; the trainable parameters of the quantization block include the quantization bit width and the quantization range.
6. The full-addition low-bit neural network accelerator according to claim 4, characterized in that: Each convolution layer in the full-precision neural network model embedded with quantization blocks described in S1 is embedded with a first quantization block and a second quantization block.
7. The full-addition low-bit neural network accelerator according to claim 6, characterized in that: The first quantization block quantizes the input weight to obtain a quantized input weight, and the weight is then convolved with the input feature map to obtain an input quantized convolution result; The second quantization block quantizes the input quantization convolution result to obtain an output feature map, which is the output of the convolution layer.
8. The full-addition low-bit neural network accelerator according to claim 4, characterized in that: The quantification described in S2 includes the following sub-processes: S21, calculating the quantization step size based on the trainable parameters; S22. Divide the input of the quantization block by the quantization step size, round the result, and then multiply it by the quantization step size to obtain a quantized parameter.
9. The full-addition low-bit neural network accelerator according to claim 8, characterized in that: The input of the quantization block is the weight of the convolution layer for the first quantization block; and the input of the quantization block is the convolution result of the weight of the convolution layer after quantization and convolution with the input feature map for the second quantization block.
10. The full-addition low-bit neural network accelerator according to claim 4, characterized in that: S36: the input of the first quantization block, i.e., the weight of the convolutional layer; S4: all the trainable parameters, including the parameters of the full-precision neural network model embedded in the quantization block and the training parameters of the quantization block.
Citation Information
Patent Citations
Low-bit quantization neural network accelerator implementation method and system
CN114757347A
Ultra-low power keyword spotting neural network circuit
US20210089874A1