A General Neural Network Acceleration Method Based on 3D Loop Unrolling
By using a three-dimensional circular expansion method in the neural network accelerator, parallel computing is performed in three dimensions of the input channel, the output channel and the output feature map, the problems of low computing efficiency of the convolutional computing unit and high memory access energy consumption are solved, and high energy-efficient neural network acceleration is achieved.
Patent Information
- Application Number
- CN202411288900.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-09-13
AI Technical Summary
In existing neural network accelerators, the convolutional computing unit has low computing efficiency and high memory access energy consumption, making it difficult to achieve high-efficiency hardware acceleration.
A general neural network acceleration method based on three-dimensional circular expansion is proposed. By performing parallel calculations in three dimensions of expansion, input channel, output channel and output feature map, combining multi-level buffers to achieve efficient data reuse and resource utilization.
Improves computing performance and data reuse, reduces power consumption caused by memory access, and enhances the flexibility and configurability of neural network accelerators.
Smart Images

Figure CN119227760B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network acceleration and processing unit design, and proposes a general neural network acceleration method based on three-dimensional loop unrolling. Background Art
[0002] The excellent performance of deep convolutional neural networks in many fields such as computer vision and speech recognition has made them the mainstream direction of current machine learning. However, with the increase in the complexity and computational requirements of deep networks, traditional central processing units (CPUs) and graphics processing units (GPUs) face efficiency and performance bottlenecks when processing these computationally intensive tasks. Therefore, heterogeneous computing platforms such as FPGAs and ASICs are widely used in the hardware accelerator design of deep networks.
[0003] Convolution, as the main operation in neural networks, is a key factor that significantly affects the efficiency and performance of accelerators. The limited computational resources and storage capacity on the hardware platform make the acceleration of convolution calculation a complex multi-dimensional optimization problem. In addition, the energy consumption cost associated with the large amount of data movement and memory access involved in convolution operations even exceeds the energy consumption of computing. Based on the above problems, achieving high-energy-efficiency hardware acceleration of convolutional neural networks requires maximizing data reuse and resource utilization while minimizing data communication, which poses high requirements for the optimized acceleration of convolution.
[0004] In view of the above problems, this application is committed to proposing a general neural network acceleration method based on three-dimensional loop unrolling. Summary of the Invention
[0005] The purpose of the present invention is to address the problems of low computational efficiency and high memory access energy consumption in the convolution calculation unit of current neural network accelerators, and proposes a general neural network acceleration method based on three-dimensional loop unrolling.
[0006] To achieve the above object, the present invention adopts the following technical solutions.
[0007] As a first aspect of the present invention, a general neural network computing unit connected to an external memory and an external controller is proposed. The general neural network computing unit includes a central control module, a core computing unit, a first-level buffer, and a first-level buffer read / write controller;
[0008] After the central control module configures each buffer and control module, the external controller sends a start signal;
[0009] The central control module controls the first-level reading and writing of input data, weight data, output data, partial sum data, and quantized weight data, and the zero-level reading and writing of input data and weight data through sending, receiving, and decoding instructions, so as to configure each buffer and control module; and enables each control module when the start signal arrives.
[0010] The first-level buffer includes a quantized weight data buffer, a partial sum data buffer, an input data first-level buffer, a weight data first-level buffer, and an output data buffer; the general neural network computing unit performs first-level buffering on the input feature map data, weight data, and quantized weight data required before neural network computing, the partial sum data generated during neural network computing, and the output data generated after neural network computing, and adjusts the read-write control logic in real time through instructions.
[0011] The general neural network computing unit further includes a zero-level buffer and a zero-level buffer read-write controller; the zero-level buffer includes an input data zero-level buffer and a weight data zero-level buffer; the zero-level buffer read-write controller includes an input data zero-level read-write control module and a weight data zero-level read-write control module; the zero-level buffer is respectively connected to the first-level buffer read-write controller and the zero-level buffer read-write controller.
[0012] The first-level buffer read-write controller includes a quantized weight data read-write control module, a partial sum data read-write control module, an input data first-level read-write control module, a weight data first-level read-write control module, and an output data read-write control module; the input data first-level read-write control module and the weight data first-level read-write control module respectively detect the states of the input data and weight data first-level buffers and the zero-level buffer, and enter the blocking state when the buffer states do not meet the read and write requirements, and wait in the blocking state until the read and write requirements are met before performing data read and write operations.
[0013] The input data zero-level read-write control module and the weight data zero-level read-write control module read data from the zero-level buffer and write it into the core computing unit for convolution calculation, and the output data obtained by convolution is sent to the output data buffer.
[0014] After the output data buffer receives the output data obtained by convolution, the core computing unit sends a calculation end signal to the central control module; after receiving the calculation end signal, the central control module sends a clearing instruction to each control module and the core computing unit to complete the current calculation.
[0015] As the second aspect of the present invention, a general neural network acceleration method with three-dimensional loop unrolling is proposed, including the following steps:
[0016] S1. The central control module controls the first-level reading and writing of input data, the first-level reading and writing of weight data, the reading and writing of output data, the reading and writing of partial sum data, and the reading and writing of quantized weight data, as well as the zero-level reading and writing of input data and the zero-level reading and writing of weight data through the sending, receiving, and decoding of instructions, so as to configure each buffer and control module. After the configuration is completed, the external controller sends a start signal.
[0017] S2. When the start signal arrives, the central control module enables each control module.
[0018] S3. The first-level reading and writing control module of input data and the first-level reading and writing control module of weight data respectively detect the status of the first-level buffer of input data and weight data and the status of the zero-level buffer. When the buffer status does not meet the requirements of reading and writing, it enters the blocking state and waits until the requirements of reading and writing are met before performing data reading and writing operations.
[0019] S4. The zero-level reading and writing control module of input data and the zero-level reading and writing control module of weight data read data from the zero-level buffer and write it into the core computing unit for convolution calculation. The output data obtained by the convolution is sent to the output data buffer.
[0020] S5. When all the output data obtained by the convolution are sent to the output data buffer, the core computing unit sends a calculation end signal.
[0021] S6. After receiving the calculation end signal, the central control module sends a clearing instruction to each control module and the core computing unit to complete this convolution calculation.
[0022] The configuration of each buffer and control module in S1 is specifically as follows: The central control module sends instructions to each first-level buffer to inform the operator type of this operation, the storage location, form, and size of each data in the buffer, and the outer calculation order; sends instructions to each zero-level buffer to inform its inner calculation order of this calculation; sends instructions to the zero-level buffer and the core computing unit to configure the registers related to the calculation.
[0023] In S2, when the start signal arrives, the central control module enables each control module. Specifically, after receiving the start signal, the central control module samples the rising edge of the start signal and pulls up the enable signal of each control module when the rising edge of the start signal arrives.
[0024] In S3, the buffer status does not meet the requirements of reading and writing. Specifically, when the buffer status is empty, it indicates that the buffer status does not meet the reading requirements; when the buffer status is full, it indicates that the buffer status does not meet the writing requirements.
[0025] Under the S3 blocking state, wait until the read and write requirements are met before performing data read and write operations. Specifically: In the blocking state, when the buffer state is empty, wait until the buffer state is not empty before performing data read. Specifically: Read the data from the first-level buffer, perform data rearrangement and partial zero-padding operations according to the instruction information, and then send the data to the zero-level buffer; when the buffer state is full, wait until the buffer state is not full before performing data write. Specifically: Write the data after data rearrangement and partial zero-padding operations to the zero-level buffer.
[0026] The convolution calculation described in S4 specifically relies on the core calculation unit to continuously generate pre-read enable signals to complete the pre-read operations of partial sum data, quantized weight data, and output data.
[0027] Reading data from the zero-level buffer and writing it to the core calculation unit described in S4 is specifically as follows: The input data zero-level read / write control module and the weight data zero-level read / write control module respectively detect the status of the zero-level buffer corresponding to the input data and the weight data and the status of the core calculation unit. If the corresponding read and write requirements are not met, enter the blocking state and wait until the requirements are met before performing data read and write operations; if the corresponding read and write requirements are met, read the data from the zero-level buffer, and after completing the remaining zero-padding operation of the data according to the instruction information, send the input data and the weight data to the core calculation unit.
[0028] The convolution calculation described in S4 is implemented relying on a core calculation unit including a broadcast control module, a multiplier array, an adder tree array, an accumulator array, and a quantization array. The specific calculation process includes the following sub-steps:
[0029] S41: The core calculation unit receives the input data and the weight data and completes the broadcast, multiplication, and reduction operations of the data according to the instruction information;
[0030] S42: The calculation in the multiplier array of the core calculation unit starts and generates an enable signal for reading and writing partial sum data;
[0031] S43: After receiving the enable signal for reading and writing partial sum data, the partial sum data read / write control module reads the corresponding partial sum data in the first-level buffer required for calculation, and then sends the partial sum data to the core calculation unit;
[0032] S44: After receiving the partial sum data, the core calculation unit completes the accumulation operation between the data after reduction and the partial sum data;
[0033] S45: The calculation in the accumulator array of the core calculation unit starts and generates an enable signal for reading and writing quantized weight data;
[0034] S46. After the quantization weight data reading and writing control module receives the enable signal for quantization weight data reading and writing, it reads the corresponding quantization weight data in the first-level buffer required for calculation, and then sends the quantization weight data to the core calculation unit;
[0035] S47. After the core calculation unit receives the quantization weight data, it completes the quantization operation;
[0036] S48. The quantization array calculation in the core calculation unit starts and generates an enable signal for output data reading and writing;
[0037] S49. After the output data reading and writing control module receives the enable signal for output data reading and writing, it sends the calculated output data to the output buffer.
[0038] Advantageous Effects
[0039] The present invention proposes a general neural network acceleration method based on three-dimensional loop unrolling. Compared with the existing neural network acceleration methods, it has the following advantageous effects:
[0040] 1. The method proposes to perform parallel calculations by unrolling in the input channel, output channel, and output feature Figure 3 dimensions. Compared with the conventional loop unrolling method, it has higher computing performance and data reuse rate, and reduces the power consumption caused by memory access;
[0041] 2. The method proposes an address signal generation mechanism based on multiple loop unrollings. By real-time configuring the cascaded counter, it realizes that one module can simultaneously adapt to various address signal generation methods based on loop unrolling, greatly improving the flexibility and configurability of the neural network accelerator;
[0042] 3. The method proposes an efficient zero-padding and data rearrangement mechanism for neural networks. With the help of multi-level buffers, it realizes phased zero-padding of the input feature map, realizes partial real-time data rearrangement and zero-padding, simplifies the zero-padding and data rearrangement logic, reduces the external bandwidth requirement, and improves the data utilization rate of the on-chip buffer. Description of the Drawings
[0043] Figure 1 It is a schematic diagram of the composition and connection relationship of the acceleration unit relied on by a general neural network acceleration method based on three-dimensional loop unrolling of the present invention;
[0044] Figure 2 It is a schematic diagram of the convolution calculation principle of three-dimensional loop unrolling described in the present invention;
[0045] Figure 3 It is a flow control mechanism diagram of a general neural network acceleration method based on three-dimensional loop unrolling described in the present invention;
[0046] Figure 4 Flowchart of the buffer controller for a general neural network acceleration method with three-dimensional loop unfolding according to the present invention;
[0047] Figure 5 Principle diagram of data rearrangement for a general neural network acceleration method with three-dimensional loop unfolding according to the present invention. Detailed implementation manners
[0048] The following combines the accompanying drawings and embodiments to elaborate in detail on the actual application process and advantages (beneficial effects) of a general neural network acceleration method based on three-dimensional loop unfolding according to the present invention.
[0049] Embodiment 1
[0050] This embodiment describes a general neural network acceleration method with three-dimensional loop unfolding, belonging to the technical field of neural network acceleration and processing unit design. The method relies on a general neural network computing unit connected to an external memory and an external controller, such as Figure 1 shown, the general neural network computing unit includes a central control module, a first-level buffer, a first-level buffer controller, and a core computing unit; the first-level buffer includes an input data first-level buffer, a weight data first-level buffer, a quantized weight data buffer, a partial sum data buffer, and an output data buffer; the first-level buffer controller, the first-level buffer is connected to the external controller and the first-level buffer controller, and the computing unit further includes a zero-level buffer and a zero-level buffer controller; the zero-level buffer includes an input data zero-level buffer and a weight data zero-level buffer; the zero-level buffer controller, the zero-level buffer is respectively connected to the first-level buffer controller and the zero-level buffer controller. The method includes: receiving an instruction and sending the decoded instruction information to configure each control module and the core computing unit; enabling after receiving a start signal; reading input data and weight data; reading input data and weights from the zero-level buffer and writing them into the core computing unit; the core computing unit continuously generates pre-read enable signals to respectively control the reading and writing of partial sum data, quantized weights, and output data; after the output data is written into the zero-level buffer, the output data reading and writing control module reads the output data from the zero-level buffer and writes it into the first-level buffer, and after completing all the writing work, sends a calculation end signal. The method performs parallel calculations in the input channels, output channels, and output feature Figure 3 dimensions, has higher computing performance and data reuse rate, and reduces the power consumption caused by memory access.
[0051] A general neural network acceleration method with three-dimensional loop unfolding, specifically in implementation, in the input channels, output channels, and output features Figure 3Expand in multiple dimensions and store them in the buffer in sequence. The expansion dimensions are set as PC, PK, and PE, all of which are configurable parameters. For three-dimensional loop expansion, each loop calculates the matrix multiplication of PC×PK×PE input pixels and the weights at corresponding positions. After completing the matrix multiplication, input an adder tree with a fan-in of PC to sum the multiplication results; each loop generates PK×PE partial sums for loop accumulation. Specifically, the implementation includes the following steps:
[0052] Step1. The central control module receives instructions, generates configuration information for each control module and the core computing unit according to the received instructions, and sends the decoded instruction information to each control module and the core computing unit.
[0053] Among them, the specific content is as follows:
[0054] The central control module sends instructions to each first-level buffer controller, informing it of the operator type of this operation, the storage location, form, and size of each data in the buffer, and the outer-layer calculation order; sends instructions to each zero-level buffer, informing it of the inner-layer calculation order of this calculation; sends instructions to the zero-level buffer and the computing unit to configure their registers related to the calculation.
[0055] Step2. The central control module receives the start signal and samples the rising edge of the start signal. When the rising edge of the start signal arrives, it raises the enable signals of each control module and the core computing unit.
[0056] Step3. Read the input data and weight data. The reading process includes the following sub-steps:
[0057] Step31. The input data reading and writing control module and the weight data reading and writing control module in the first-level buffer controller first detect the status of the first-level buffer and the zero-level buffer. If the corresponding reading and writing requirements are not met, they enter the blocking state and wait until the requirements are met before performing data reading and writing operations.
[0058] Step32. If the corresponding reading and writing requirements are met, read the data from the first-level buffer, perform data rearrangement and partial zero-padding operations according to the instruction information, and then send the data to the zero-level buffer.
[0059] Step4. The input data reading and writing control module and the weight data reading and writing control module in the zero-level buffer controller read the data from the zero-level buffer and write it into the core computing unit. The reading process includes the following sub-steps:
[0060] Step41. The two modules in Step 4 respectively detect the status of their corresponding zero-level buffers and the status of the computing unit. If the corresponding reading and writing requirements are not met, they enter the blocking state and wait until the requirements are met before performing data reading and writing operations.
[0061] Step42. If the corresponding reading and writing requirements are met, read the data from the zero-level buffer, and after completing the remaining zero-padding operation of the data according to the instruction information, send the data to the core computing unit.
[0062] Step5. During the process of the data flowing in the core computing unit, pre-read enable signals are continuously generated to separately control the partial sum data read / write control module, the quantized weight data read / write control module, and the output data read / write control module to generate enable signals in advance, so as to realize the pre-reading operations of the partial sum data, the quantized weight data, and the output data, and complete the calculation process of the core computing unit. The calculation process includes the following sub-steps:
[0063] Step51. After the core computing unit receives the input data and the weight data, according to the instruction information, complete the data broadcast, multiplication, and reduction operations.
[0064] Step52. When the multiplier array in the core computing unit starts to calculate, generate the enable signal of the partial sum data read / write control module.
[0065] Step53. After the partial sum data read / write control module receives the enable signal, according to the previously received instruction information, read the corresponding partial sum and quantized weight in the first-level buffer required for the calculation in real time, and then send the partial sum data to the core computing unit.
[0066] Step54. After the core computing unit receives the partial sum data, according to the instruction information, complete the accumulation operation between the reduced data and the partial sum data.
[0067] Step55. When the accumulator array in the core computing unit starts to calculate, generate the enable signal of the quantized weight data read / write control module.
[0068] Step56. After the quantized weight data read / write control module receives the enable signal, according to the previously received instruction information, read the corresponding quantized weight data in the first-level buffer required for the calculation in real time, and then send the quantized weight data to the core computing unit.
[0069] Step57. After the core computing unit receives the quantized weight data, according to the instruction information, complete the quantization operation.
[0070] Step58. When the quantization array in the core computing unit starts to calculate, generate the enable signal of the output data read / write control module.
[0071] Step59. After the output data read / write control module receives the enable signal, according to the previously received instruction information, send the calculated output data to the zero-level buffer in real time.
[0072] Step 6. After the output data is written into the zero-level buffer, the output data read-write control module in the first-level buffer controller reads the output data from the zero-level buffer and writes it into the first-level buffer. After completing all the writing work, it sends a calculation end signal to the central control module.
[0073] Step 7. After receiving the calculation end signal, the central control module sends a clearing instruction to each control module and the core calculation unit to complete the clearing operation for this calculation.
[0074] During specific implementation, the characteristics of the convolutional neural network are analyzed to determine three unrolled loops. According to the selected three-dimensional loops, the specific calculation order is determined:
[0075] 1. Analysis of the convolutional kernel dimension (loop1). The commonly used convolutional kernel sizes in convolutional neural networks range from 1×1 to 13×13, with a relatively small range. In modern neural network accelerators, except for very few specific scenarios, parallel computing on the convolutional kernel is rarely seen because this will cause a significant reduction in the types of convolutional kernel sizes and strides that the accelerator can adapt to, thus limiting the applicable range of the accelerator. Therefore, the loop unrolling method will not perform parallel computing on the convolutional kernel (loop 1);
[0076] 2. Analysis of the input channel dimension (loop 2) and the output channel dimension (loop 4); The common input channels and output channels in convolutional neural networks range from 1 to 1024. Since the coupling degree between the input channels and output channels is low, the control logic is simple, and it can be compatible with various sizes and strides of convolutions, modern neural network accelerators will choose to perform parallel computing on the input channels or output channels. The loop unrolling method chooses to perform parallel computing on both the input channels and output channels;
[0077] 3. Analysis of the output feature map dimension (loop 3): The size of the output feature map ranges from 20×20 to 512×512, which can provide rich parallelism without worrying about low hardware utilization. However, since the output feature map dimension is a double loop, if both loops are unrolled simultaneously, problems such as complex data paths and poor adaptability will be faced; Therefore, the loop unrolling method chooses to perform parallel computing on one of the dimensions of the output feature map.
[0078] In summary, the three-dimensional unrolling method chooses to perform parallel computing on the input channels, output channels, and output feature Figure 3 map dimensions to improve the performance ceiling and simplify the data path.
[0079] Figure 2This is the schematic diagram of a three-dimensional cyclic expansion described in the present invention; based on the three-dimensional cyclic expansion method proposed in this embodiment, the output channels, the vertical direction of the output feature map, and the input channels are expanded in parallel in three dimensions, and the parallelism degrees are denoted as PK, PE, and PC respectively. At this time, the convolution operation can be mapped to the following cyclic expansion:
[0080] for(nos = 0; nos < Nof; nos++){ loop4
[0081] for(ys = 0; ys < Noy; ys++){ loop3-h
[0082] for(x = 0; x < Nox; x++){ loop3-w
[0083] for(nis = 0; nis < Nif; nis++){ loop2
[0084] for(ky = 0; ky < Nky; ky++){ loop1-h
[0085] for(kx = 0; kx < Nkx; kx++){ loop1-w
[0086] #parallel
[0087] for(nop = 0; nop < PK; nop++){
[0088] for(yp = 0; yp < PE; yp++){
[0089] for(nip = 0; nip < PC; nip++){
[0090] #ComputeOneCycle
[0091] }}}}}}}}}
[0092] In the above formula, the loop can be divided into two major parts. loop4, loop3-h, loop3-w, loop2, loop1-h, and loop1-w are the outer loop expansion and the inner loop expansion, corresponding to the process of data flowing from each primary buffer into the computing unit; the innermost three-layer loop is the parallel loop expansion, corresponding to the process of data being computed in the computing unit. At this time, if the loop in the outer loop expansion and the inner loop expansion is closer to the parallel loop expansion, it is called that the computation is closer to the inside; if it is farther from the parallel loop expansion, it is called that the computation is closer to the outside.
[0093] On this basis, we split the outer loop unrolling and the inner loop unrolling and swap the loop order, and determine the data formats corresponding to the outer loop and the inner loop, so as to map the input data and the weight data to the zero-level buffer respectively.
[0094] To determine the calculation order of the six-fold remaining loops in the outer loop unrolling and the inner loop unrolling parts, we first analyze the influence of different loop layers on the neural network calculation process:
[0095] 1. Output channels (loop 4): The data of different output channels are generated by convolving with different convolution kernels respectively, and the calculations between different convolution kernels are independent of each other. Therefore, in modern neural network accelerators, the output channels are usually placed in the outermost layer of the loop. That is, the output channels are calculated last.
[0096] 2. Input channels (loop 2): When performing convolution calculations, if the input channel dimension is switched, both the input feature map and the convolution kernel will change simultaneously. Since the coupling degree of the input channels with the other three dimensions is relatively low, adjusting their calculation order in the loop unrolling will not increase the complexity of the data path and control signals. However, the calculation order of the input channels will significantly affect the scale of the partial sums. The specific impacts are as follows:
[0097] When the calculation of the input channels is closer to the outside, the scale of the partial sums is larger, but it is easier to obtain a simpler data path because, without involving the unrolling of the input channels and output channels in the outer layer, the data of the convolution kernel will remain unchanged for a long time, thus reducing the memory access bandwidth requirement for the memory subsystem; when the calculation of the input channels is closer to the inside, the scale of the partial sums is smaller, but this also means that the feature map and the convolution kernel of the input calculation unit need to be switched more frequently. At this time, in order to achieve data reuse and reduce the power consumption overhead of accessing the memory subsystem, a larger on-chip memory access bandwidth is required.
[0098] Therefore, to meet the bandwidth requirements of the memory subsystem in different application scenarios, the calculation unit can freely adjust the order of the input channels in the loop unrolling, so as to balance the memory access bandwidth and the size of the partial sum buffer.
[0099] 3. Output feature maps (loop 3) and convolution kernels (loop 1): Convolution operations have locality, and there is a large amount of data reuse during the sliding window process. This kind of reuse can be divided into data reuse between rows and data reuse within a row. Data reuse between rows means that the same data may be repeatedly used during the sliding window process of the convolution kernel between different rows. Data reuse within a row means that the same data may be repeatedly used during the sliding window process of the convolution kernel within the same row.
[0100] Take the case where the convolution kernel size is 3×3 and the input feature map is the same as the output feature map. Ideally, if the data reuse generated during the sliding window process can be fully utilized, the on-chip buffer size will be reduced by 9 times compared to directly converting the convolution to matrix multiplication; if the on-chip register file can be used to fully cache the data that needs to be reused, the number of times of reading data from the global buffer will also be reduced by 9 times.
[0101] Both the output feature map and the convolution kernel have two layers of loops in the horizontal and vertical directions. When the calculation order of the output feature map and the convolution kernel is closer to the inside, the data reusability is higher, but the complexity of the data path will also increase significantly; when the calculation order of the output feature map and the convolution kernel is closer to the outside, although the complexity of the data path is greatly reduced, due to the reduction of data reusability, it brings greater bandwidth and area pressure to the storage subsystem. Therefore, it is necessary to appropriately arrange these two layers of loops of the four layers of loops to balance the contradiction between data reusability and data path complexity.
[0102] When the two layers of loops in the horizontal and vertical directions are close together, it means that the computing unit needs to frequently read data in different rows in a short time to complete a convolution operation, which will make the design of the data path very difficult. Therefore, it is necessary to separate the horizontal and vertical directions as much as possible. Inserting other loops between these two layers of loops can effectively utilize the data reusability within the row without sacrificing too much hardware resources and control logic for realizing the data reusability between rows.
[0103] In summary, the calculation order corresponding to the serial loop unrolling of the computing unit is as follows:
[0104] for(nos=0;nos<Nof;nos++){ loop4
[0105] for(ys=0;ys<Noy;ys++){ loop3-h
[0106] for(ky=0;ky<Nky;ky++){ loop1-h
[0107] for(x=0;x<Nox;x++){ loop3-w
[0108] for(kx=0;kx<Nkx;kx++){ loop1-w
[0109] for(nis=0;nis<Nif;nis++){ loop2
[0110] #parallel
[0111] }}}}}}
[0112] Among them, loop unrolling can be divided into four parts. The position of the output channel (loop 4) is absolutely fixed and located on the outermost side of the loop; the positions of the output feature map (loop 3) and the convolution kernel (loop 1) can be swapped, corresponding to the convolution operations unfolded by rows and columns respectively; the position of the input channel (loop 2) is absolutely flexible. By adjusting the position of the loop at this layer, different calculation orders can be obtained. A total of 2×5 = 10 different calculation orders can be obtained, and thus different power consumptions and on-chip cache overheads can be obtained. It can be seen from this that the flexibility and configurability of the computing unit are high.
[0113] Assume that the input channel is on the innermost side. When using the above-mentioned serial loop unrolling for convolution calculation, it can not only ensure a certain degree of data reuse but also greatly reduce the complexity of the data path. Taking the convolution size of 3×3 and the input feature map and the output feature map having the same size as an example, directly using this loop calculation and using the zero-level buffer in the computing unit for data buffering can save 3 times the on-chip buffer size and reduce the number of times of reading data from the global buffer by 3 times.
[0114] Step 2: The method reads data, performs single-stage data rearrangement, and two-stage zero padding according to the three dimensions of loop unrolling. Specifically: read the feature map data from the first-level buffer, perform zero padding between rows and data rearrangement on it according to the three-dimensional loop unrolling, and then write it into the zero-level buffer; read the feature map data from the zero-level buffer, perform zero padding between columns on it according to the three-dimensional loop unrolling, and then write it into the convolution computing unit.
[0115] Analyze the process of mapping the method to the convolution computing unit.
[0116] When loading data from the global buffer into the convolution computing unit, there are the following two mismatches: the mismatch between the read / write bandwidth of the global buffer and the read / write bandwidth of the computing unit; the mismatch between the data storage format in the global buffer and the data storage format required by the computing unit.
[0117] The former mismatch is caused by the uncontrollable computing intensity of the operator. For example, the convolution operator has a high computing intensity and the data has a high data reuse rate, so the requirement for the read / write bandwidth of the global buffer is low. To address this mismatch, an additional zero-level buffer is added and implemented through a data flow control and transmission mechanism.
[0118] Figure 3 This is the flow control mechanism diagram of a general neural network acceleration method with three-dimensional loop unrolling according to the present invention. There are three parts in this flow control mechanism: an upstream control module, a buffer, and a downstream control module. Among them, the upstream control module sends a write enable and a write pointer to the buffer, and the downstream control module sends a read enable and a read pointer to the buffer.
[0119] Take the case where the convolution kernel size is 3*3 and the input channel (loop 2) is in the innermost layer. At this time, the upstream control module is a first-level buffer controller, the buffer is a zero-level buffer, and the downstream control module is a zero-level buffer controller. At this time, when the buffer is not full, the upstream will write the data at one address every clock cycle and increment the write pointer by 1; when the buffer is not empty, the downstream will read the data at one address every clock cycle, and after an interval of 3 clock cycles, increment the read pointer by 3. When the input channel moves outward, the downstream will increment by 3× the number of rows after an interval of 3× the number of rows of clock cycles.
[0120] When the convolution loop unrolling calculation order is adjusted, this mechanism can automatically receive instruction information and adjust the behavior logic to start the calculation immediately without excessive waiting, while the traditional ping-pong structure has to wait until the data fills at least one buffer before starting, reaching the speed limit of data reading in this process.
[0121] There are two reasons for the latter mismatch: on the one hand, due to loop unrolling, the data that the computing unit can receive needs to be an integer multiple of the parallelism. Taking the loop unrolling of 4×4×4 as an example, when the input data is 6×6, it has to be padded with zeros to 8×8 before being sent to the computing unit; on the other hand, due to the specific zero-padding logic of the convolution operator, additional data rearrangement and zero-padding logic need to be added to achieve it. The computing unit realizes this mismatch problem through the collaborative work of the first-level buffer read-write controller and the zero-level buffer read-write controller.
[0122] Figure 4 This is the flowchart of the buffer controller for the general neural network acceleration method of three-dimensional loop unrolling described in the present invention.
[0123] The functions of the first-level buffer read-write controller are as follows: on the one hand, it completes the data rearrangement operation, converting the data format in the global buffer into the data format required for convolution operation; on the other hand, it completes the zero-padding operation between rows.
[0124] The functions of the zero-level buffer read-write controller are as follows: sequentially read the data in the zero-level buffer and complete the zero-padding operation between columns.
[0125] For example, when the convolution kernel size is 3×3 and the stride is 1, the input data first-level read-write control module pads the input feature map in the global buffer from 60×60 with zeros to 64×60, padding one row of zeros at the top and three rows of zeros at the bottom, and sends it to the zero-level buffer; the input data zero-level read-write control module pads 64*60 in the zero-level buffer with zeros to 64×62 and sends it to the computing unit for calculation.
[0126] By means of single-stage data rearrangement and two-stage zero-padding operations, the original complex zero-padding logic is decoupled, the number of combinational logic levels required for zero-padding operations is reduced by 50%, and the upper limit of the system's operating frequency is increased. Through the above structure, a total of 9 times the on-chip buffer size can be saved, and at the same time, the number of times of reading data from the global buffer is reduced by 9 times, reaching the theoretical limit.
[0127] Step 3: The convolution calculation unit reads the feature map data from the zero-level buffer and performs convolution calculations after three-dimensional loop unrolling.
[0128] Embodiment 2
[0129] This embodiment specifically elaborates on the data rearrangement and zero-padding processes of the method.
[0130] In convolution calculations, the input feature map is usually a multi-channel multi-dimensional data. To more efficiently utilize memory and cache, the input data can be rearranged into a continuously stored form. When rearranging, the input channel dimension is usually placed in the innermost layer, so that data can be accessed sequentially during the calculation process, reducing the randomness of memory access, reducing the number of times data moves from memory to the processing unit, and improving calculation efficiency. Similarly, the weights can also be rearranged, merging the size and channels of the weights into a larger continuous block, so that they can be accessed and calculated more quickly during convolution operations.
[0131] In order to match the read-write bandwidth of the global buffer with that of the calculation unit, the present invention performs ping-pong caching and data rearrangement in the first-level buffer control unit. Figure 5 The schematic diagram of data rearrangement is shown.
[0132] For example, when the read-write bandwidths of both the global buffer and the calculation unit are PE×PC, due to the locality of data storage, the data read from the first-level buffer is PE×PC data along the input channels, while the calculation unit requires PE×PC data expanded in the vertical direction of the feature map and the input channel direction. The present invention opens up two buffers in the first-level buffer controller for ping-pong caching. When writing PE×PC data row by row to one buffer, at the same time, reading PE×PC data column by column from the other buffer. Through this method, the matching of the read bandwidth and the efficient reading of data are achieved.
[0133] Zero-padding can keep the spatial dimensions of the input and output feature maps relatively unchanged, avoiding the problem that the size of the output feature map shrinks due to the sliding of the convolution kernel. The size of zero-padding is specified when configuring the convolutional layer and is dynamically implemented during the loading and calculation of the input feature map to adapt to different sizes of convolution kernels and strides.
[0134] The present invention realizes row zero-padding and column zero-padding of the feature map through two-stage zero-padding: after the primary buffer controller reads data from the primary buffer, it performs row zero-padding and then sends it to the zero-level buffer; after the zero-level buffer controller reads data from the zero-level buffer, it performs column zero-padding and then sends it to the convolution calculation unit.
[0135] For example, when the size of the input feature map is 60×60, the size of the convolution kernel is 3×3, and the stride is 1, the input data primary buffer controller pads the feature map to 64×60, with one row of zeros padded above and three rows of zeros padded below. In the primary buffer controller, the feature map vertical expansion counter, input channel counter, feature map row counter, convolution kernel column counter, feature map column counter, and output channel counter work in cascade in sequence, generating address signals and zero-padding signals in real time. When the sum of the feature map vertical expansion counter, convolution kernel column counter, and feature map column counter is less than 1, the zero-padding signal is valid, and the first row is padded with zeros; when the sum of the feature map vertical expansion counter, convolution kernel column counter, and feature map column counter is greater than 60, the zero-padding signal is valid, and the last three rows are padded with zeros.
[0136] The input data buffer controller pads the feature map in the zero-level buffer from 64*60 to 64×62 and sends it to the calculation unit for calculation. In the zero-level buffer controller, the input channel counter, convolution kernel row counter, feature map row counter, convolution kernel column counter, feature map column counter, and output channel counter work in cascade in sequence, generating address signals and zero-padding signals in real time. When the sum of the convolution kernel row counter and feature map row counter is less than 1, the zero-padding signal is valid, and the first column is padded with zeros; when the sum of the convolution kernel row counter and feature map row counter is greater than 60, the zero-padding signal is valid, and the last column is padded with zeros.
[0137] Step1. The method reads, rearranges, and pads zeros for data according to three dimensions of loop unrolling. Specifically: read the feature map data from the primary buffer, perform row zero-padding and data rearrangement on it according to three-dimensional loop unrolling, and then write it into the zero-level buffer.
[0138] Step11. Configure convolution-related parameters, specifically including: convolution kernel size, feature map size, number of input channels, number of output channels, zero-padding size, configuration of cascade counters, buffer address offset.
[0139] Step12. Cascade counters in sequence: feature map vertical expansion counter, input channel counter, feature map row counter, convolution kernel column counter, feature map column counter, output channel counter, and increment and clear according to the configured parameters.
[0140] Step13. Synchronously generate the data read address of the primary buffer according to the counter value.
[0141] Step14. Determine whether to pad with zeros according to the configuration parameters and the counter value. If not, read the data at the corresponding address from the primary buffer; otherwise, replace it with zeros.
[0142] Step15. Rearrange the read data or the zero-padded data; the specific process is as follows: cache the read data and the zero-padded data into the register in the form of a matrix, and then transpose and output.
[0143] Step16. Write the rearranged data into the L0 buffer.
[0144] Step2. The method reads and pads data according to three dimensions of loop unrolling. Specifically: read the feature map data from the L0 buffer, pad zeros between columns according to three-dimensional loop unrolling, and then write it into the convolution calculation unit.
[0145] Step21. Configure the convolution-related parameters, specifically including: convolution kernel size, feature map size, number of input channels, number of output channels, zero-padding size, configuration of the cascaded counter, buffer address offset.
[0146] Step22. Cascade the counters in sequence: input channel counter, convolution kernel row counter, feature map row counter, convolution kernel column counter, feature map column counter, output channel counter, and increment and clear them according to the configured parameters.
[0147] Step23. Synchronously generate the data read address of the L0 buffer according to the counter value.
[0148] Step24. Determine whether to pad with zeros according to the configuration parameters and the counter value. If not, read the data at the corresponding address from the L0_cache; otherwise, replace it with zeros;
[0149] Step25. Write the zero-padded data into the calculation unit;
[0150] The above is the preferred embodiment of the present invention, and the present invention should not be limited to the content disclosed in this embodiment and the drawings. All equivalent or modified implementations completed without departing from the spirit disclosed by the present invention fall within the protection scope of the present invention.
Claims
1. A general neural network acceleration method based on three-dimensional loop unrolling, relying on a general neural network computing unit connected to an external memory and an external controller, including a central control module, a core computing unit, a primary buffer, and a primary buffer read-write controller; the primary buffer includes a quantized weight data buffer, a partial sum data buffer, an input data primary buffer, a weight data primary buffer, and an output data buffer; the general neural network computing unit performs primary buffering on the input feature map data, weight data, and quantized weight data required before the neural network calculation, the partial sum data generated during the neural network calculation, and the output data generated after the neural network calculation, and adjusts the read-write control logic in real time through instructions; It is characterized in that The general neural network computing unit further includes a zero-level buffer and a zero-level buffer read-write controller; the zero-level buffer includes an input data zero-level buffer and a weight data zero-level buffer; the zero-level buffer read-write controller includes an input data zero-level read-write control module and a weight data zero-level read-write control module; the zero-level buffer is connected to the first-level buffer read-write controller and the zero-level buffer read-write controller respectively, and the general neural network acceleration method includes the following steps: S1. The central control module controls the first-level reading and writing of input data, the first-level reading and writing of weight data, the reading and writing of output data, the reading and writing of partial sum data and the reading and writing of quantized weight data through the sending, receiving and decoding of instructions, and controls the zero-level reading and writing of input data and the zero-level reading and writing of weight data to realize the configuration of each buffer and control module. After the configuration is completed, the external controller sends a start signal; S2, the central control module enables each control module when the start signal arrives; S3, the input data level read-write control module and the weight data level read-write control module detect the state of the input data and weight data level buffer and the state of the zero-level buffer respectively, and enter the blocking state when the buffer state does not meet the reading and writing requirements, and wait in the blocking state until the reading and writing requirements are met before performing data reading and writing operations; S4, the input data zero-level read-write control module and the weight data zero-level read-write control module read data from the zero-level buffer and write it into the core computing unit for convolution calculation, and the output data obtained by the convolution is sent to the output data buffer; S5. When all the output data obtained by the convolution are sent to the output data buffer, the core computing unit sends a computing end signal; S6. After receiving the calculation end signal, the central control module sends a clear instruction to each control module and core computing unit to complete the convolution calculation.
2. The general neural network acceleration method based on three-dimensional loop unrolling according to claim 1, characterized in that: The configuration of each buffer and control module in S1 is as follows: the central control module sends instructions to each first-level buffer to inform the operator type of this operation and the storage location, form, size and outer calculation order of each data in the buffer; sends instructions to each zero-level buffer to inform it of the inner calculation order of this calculation; sends instructions to the zero-level buffer and core computing unit to configure registers related to the calculation.
3. The general neural network acceleration method based on three-dimensional loop unrolling according to claim 1 is characterized in that: The S2 central control module enables each control module when the start signal arrives. Specifically, after receiving the start signal, the central control module samples the rising edge of the start signal and pulls up the enable signal of each control module when the rising edge of the start signal arrives.
4. The general neural network acceleration method based on three-dimensional loop unrolling according to claim 1, characterized in that: S3 The buffer state does not meet the reading and writing requirements, specifically: when the buffer state is empty, it indicates that the buffer state does not meet the reading requirements; when the buffer state is full, it indicates that the buffer state does not meet the writing requirements.
5. The general neural network acceleration method based on three-dimensional loop unrolling according to claim 4 is characterized in that: In the S3 blocking state, the read and write requirements are met before performing data reading and writing operations. Specifically, in the blocking state, when the buffer state is empty, wait until the buffer state is not empty before reading data. Specifically, read the data from the first-level buffer and rearrange the data and partially fill in zeros according to the instruction information, and then send the data to the zero-level buffer; when the buffer state is full, wait until the buffer state is not full before writing data. Specifically, write the data after data rearrangement and partial zero filling operations into the zero-level buffer.
6. The general neural network acceleration method based on three-dimensional loop unrolling according to claim 1, characterized in that: The convolution calculation described in S4 specifically relies on the core computing unit to continuously generate a pre-read enable signal to complete the pre-read operation of partial sum data, quantized weight data and output data.
7. A general neural network acceleration method based on three-dimensional loop unrolling according to claim 1, characterized in that: S4 describes reading data from the zero-level buffer and writing it into the core computing unit, specifically: the input data zero-level read-write control module and the weight data zero-level read-write control module respectively detect the status of the zero-level buffer and the status of the core computing unit corresponding to the input data and the weight data. If the corresponding read and write requirements are not met, a blocking state is entered, and data reading and writing operations are performed after the requirements are met; if the corresponding read and write requirements are met, the data is read out from the zero-level buffer, and after the remaining zero-filling operation of the data is completed according to the instruction information, the input data and weight data are sent to the core computing unit.
8. A general neural network acceleration method based on three-dimensional loop unrolling according to claim 1, characterized in that: The convolution calculation in S4 is implemented by a core computing unit including a broadcast control module, a multiplier array, an adder tree array, an accumulator array, and a quantization array. The specific calculation process includes the following sub-steps: S41, the core computing unit receives input data and weight data, and completes data broadcasting, multiplication and reduction operations according to the instruction information; S42, the multiplier array in the core computing unit starts computing and generates an enable signal for partial and data reading and writing; S43, after receiving the enable signal of partial sum data reading and writing, the partial sum data reading and writing control module reads the corresponding partial sum data in the primary buffer required for calculation, and then sends the partial sum data to the core computing unit; S44, after receiving the partial sum data, the core computing unit completes the accumulation operation between the reduced data and the partial sum data; S45, the accumulator array in the core computing unit starts calculation and generates an enable signal for reading and writing quantized weight data; S46, after receiving the enable signal for reading and writing the quantized weight data, the quantized weight data read and write control module reads the corresponding quantized weight data in the primary buffer required for calculation, and then sends the quantized weight data to the core calculation unit; S47, after receiving the quantization weight data, the core computing unit completes the quantization operation; S48, the quantization array calculation in the core calculation unit starts and generates an enable signal for reading and writing output data; S49, after receiving the output data read and write enable signal, the output data read and write control module sends the calculated output data to the output buffer.
Citation Information
Patent Citations
Convolutional neural network acceleration device and method based on SIMD technology
CN112418417A
Neural network acceleration device and method and communication equipment
CN113807509A