High-Throughput Object Detection Accelerator Based on Systolic Array

Through the computing architecture and pipeline technology based on pulsating arrays, combined with the roofline model to optimize the pulsating array scale and data multiplexing method of the YOLOv2-tiny object detection algorithm, the high throughput and high resource utilization problems of the YOLOv2-tiny object detection algorithm under limited resources are solved, and an efficient object detection accelerator design is achieved.

CN116136798BActive Publication Date: 2025-07-18RES INST OF SOUTHEAST UNIV IN SUZHOU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310260292.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-17
Publication Date
2025-07-18
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently realize the high throughput and high resource utilization of the YOLOv2-tiny object detection algorithm under limited resource constraints, especially in the design of pulsating arrays, which have low resource utilization or storage wall problems.

Method used

Using a computing architecture based on pulsating arrays, pipeline technology and hybrid data multiplexing strategies are introduced, and the optimal scale is dynamically explored through design space, and combining the roofline model to optimize the scale and data multiplexing method of pulsating arrays to improve computing parallelism and throughput.

Benefits of technology

Without affecting network accuracy, optimize memory operations, reduce external storage and read bandwidth, improve data utilization, enhance computing parallelism, and realize high throughput and high resource utilization target detection accelerator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116136798B_ABST
    Figure CN116136798B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-throughput object detection accelerator based on a systolic array. The accelerator includes: an input feature map storage unit (1), an input feature map read / write cache unit (2), a systolic array calculation part (3), a weight read / write cache unit (4), a pooling unit (5), an output result read / write cache unit (6), and a global configuration unit (7); it not only simplifies the convolutional calculation and the complex data flow design of the hardware, but also introduces a step of design space exploration, ensuring high resource utilization and high throughput of the hardware. The pipeline technology is introduced into the calculation units of the systolic array, enabling this architecture to improve the calculation parallelism and using a hybrid data reuse strategy to reduce the bandwidth pressure. The core of this accelerator is to dynamically find the optimal scale of the systolic array through design space exploration, enabling the object detection algorithm to fully utilize the hardware resources of the given FPGA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the design of a high-throughput object detection accelerator based on a systolic array, belonging to the field of computing architectures. Background Art

[0002] In recent years, object detection technology has broad applications in fields such as road monitoring, driverless, medical image analysis, and national defense security. Neural networks are the core algorithms for object detection, featuring large amounts of data, large amounts of computation, and large amounts of memory access. However, the performance growth rate of general-purpose processors has slowed down and it is difficult to keep up with the rapid application development pace of neural networks. Therefore, optimizing the computing system is becoming an increasingly concerned issue.

[0003] The mainstream optimization strategies for computing systems can be divided into: multiplication optimization at the circuit level, approximate computing, quantization at the parameter level, pruning at the network structure level, bandwidth optimization at the memory access level, and computing parallelism and data reuse at the architecture level. Approximate computing, quantization, and pruning, etc. will affect the original accuracy of the network. While the computing parallelism and data reuse strategies based on systolic arrays can optimize memory operations, reduce external storage and read bandwidth, and improve data utilization without affecting the original accuracy of the network.

[0004] Since the amount of computation varies greatly among different layers of YOLOv2-tiny, if the scale of the systolic array is designed too large, it will lead to low resource utilization; if the scale of the systolic array is designed too small, it may be affected by the memory wall problem. Therefore, it is necessary to reasonably explore the design space of the systolic array. Thus, the YOLOv2-tiny accelerator based on the systolic array needs to use the roofline modeling method to explore the design space of the systolic array under limited resource constraints, balance the problems of low resource utilization and computing latency, so as to find a solution with high throughput and high utilization. Summary of the Invention

[0005] Technical Problem: Aiming at the above problems, the present invention discloses a design of a high-throughput object detection accelerator based on a systolic array, which not only simplifies the convolutional calculation and the complex data flow design of the hardware, but also introduces the step of design space exploration, ensuring high resource utilization and high throughput of the hardware. Pipeline technology is introduced into the computing units of the systolic array, enabling this architecture to improve computing parallelism and using a hybrid data reuse strategy to reduce bandwidth pressure. The core of this accelerator is to dynamically find the optimal scale of the systolic array through design space exploration, enabling the object detection algorithm to make full use of the given FPGA hardware resources.

[0006] Technical solution: A high-throughput object detection accelerator based on a systolic array according to the present invention includes: an input feature map storage unit 1, an input feature map read / write cache unit 2, a systolic array calculation part 3, a weight read / write cache unit 4, a pooling unit 5, an output result read / write cache unit 6, and a global configuration unit 7; the input feature map storage unit 1, the input feature map read / write cache unit 2, the systolic array, the weight read / write cache unit 4, the pooling unit 5, the output result read / write cache unit 6, and the global configuration unit 7; wherein the output of the global configuration unit 7 is respectively connected to the input feature map storage unit 1, the input feature map read / write cache unit 2, the weight read / write cache unit 4, the pooling unit 5, and the output result read / write cache unit 6; the output of the input feature map storage unit 1 is connected to the input feature map read / write cache unit 2, and the output of the input feature map read / write cache unit 2 is connected to the systolic array calculation part 3; the output of the weight read / write cache unit 4 is respectively connected to the systolic array calculation part 3; the input of the pooling unit 5 is connected to the systolic array calculation part 3; the output of the pooling unit 5 is connected to the output result read / write cache unit 6; YOLOv2-Tiny object detection is implemented in the high-throughput object detection accelerator based on the systolic array.

[0007] The systolic array includes a plurality of single computing units, and each single computing unit includes: a feature map input port 31, an input feature map register unit 32, a weight input port 33, a weight register unit 34, and then convolution calculation is performed through a multiplier 35 and an adder 36. Whether the complete convolution calculation is completed is judged by a judge 37. If not, the calculation part and the result are stored in a register 38 for continued accumulation. If the calculation is completed, the result is subjected to pooling processing 39; the input feature map is transmitted to the right computing unit 314 through a register 312, the weight is transmitted to the lower computing unit 315 through a register 313, and the output result 316 of the lower computing unit is temporarily stored in a register 317. The data selector 310 is used to select whether to output the current calculation result or the calculation result transmitted from the next systolic unit.

[0008] For the single computing unit, a register unit is added at the output result systolic transmission place to form a pipeline design. According to the parameters of different layers of the YOLOv2-Tiny object detection algorithm, the delay time of the pipeline can be flexibly designed to increase the parallelism of the calculation and improve the throughput of the accelerator.

[0009] The described high-throughput object detection accelerator operates in the following mode: The input feature map stored in the external memory of the accelerator is transmitted to the input feature map register unit 1 via the AXI4-Stream bus. Then, the global configuration unit control unit configures the values in the input feature map register unit 1 into the form required by the systolic array and stores them in the input feature map read / write cache unit 2. Finally, they are sequentially output to each row of the systolic array for calculation. The weight data is transmitted to the weight read / write cache unit 4 via the AXI4-Stream bus, and then the weight data is sequentially output to each column of the systolic array for convolution calculation with the data input to the rows of the systolic array. The results of each calculation unit are passed to the previous calculation unit. The calculation results of the first layer will be stored in the output result read / write cache unit 6 in the designed data arrangement manner after the pooling operation in the pooling unit 5, and finally, the output results are transmitted to the external memory of the accelerator via the AXI4-Stream bus.

[0010] For the described systolic array calculation part, for the given YOLOv2-Tiny object detection algorithm, the systolic scales that can be designed on different hardware platforms are different. To achieve higher resource utilization and throughput of the YOLOv2-Tiny object detection algorithm on the hardware platform, the steps of design space exploration are introduced. The input data, weight data, and output data of each layer of the YOLOv2-Tiny object detection algorithm are sequentially selected. Denote the three-dimensional scale of the output data as R rows, C columns, and M channels, and the four-dimensional scale of the weight data as K rows, K columns, N channels, and M blocks. Therefore, the number of calculation operations Ops for this layer of the YOLOv2-Tiny object detection algorithm is given by Formula 1.

[0011] Ops = 2 × R × C × M × N × K × K Formula 1

[0012] The output data is divided into blocks. Denote Tm as the block size of the channels, Tr as the row block size, and Tc as the column block size. The total number of blocks Blocks and the time required to calculate each block Timeperblock are given by Formula 2 and Formula 3.

[0013]

[0014] Timeperblock = K × K × N + Tr + Tc - 1 Formula 3

[0015] Denote B in 、B w 、B out as the sizes of the input feature map register unit, weight read / write cache unit, and output result read / write cache unit required when calculating one block of output data respectively. S is the sliding window step size of the convolution calculation, as given by Formulas 4-1, 4-2, and 4-3 respectively.

[0016] Bin = N × (S × Tr + K - S) × (S × Tc + K - S) Formula 4-1

[0017] B w = K × K × N × Tm Formula 4-2

[0018] B out = Tr × Tc × Tm Formula 4-3

[0019] Denote α in as the number of times the input data needs to be transferred from outside the accelerator to the input feature map register unit inside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated. α w is the number of times the weight data needs to be transferred from outside the accelerator to the weight read / write cache unit inside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated. α out is the number of times the output data needs to be transferred from the output result read / write cache unit inside the accelerator to outside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated, as shown in Formulas 5-1, 5-2, and 5-3 respectively.

[0020]

[0021]

[0022]

[0023] The computing performance Performance and the operation intensity CTC of YOLOv2-Tiny can be calculated using the roofline model as shown in Formulas 6 and 7 respectively;

[0024]

[0025]

[0026] By enumerating Tr, Tc, and Tm in Formula 6 and under the constraint of Formula 7, the most suitable systolic array scale for the YOLOv2-Tiny object detection algorithm can be obtained.

[0027] For the systolic array computing part, a hybrid data reuse strategy is adopted; for the YOLOv2-Tiny object detection network, the data volume of the convolution kernels in the first 4 layers is small, so weight reuse is adopted, and the data volume of the convolution kernels in the last 5 layers is much larger than that of the input feature map data, so the input feature map reuse method is adopted.

[0028] The steps to implement YOLOv2-Tiny object detection in the accelerator based on the systolic array are as follows:

[0029] Step 1: Find a feasible hardware circuit mapping method in the systolic array to ensure that correct data is available at specific positions in the computing units in each cycle. Considering the significant difference in the computational workload between layers of the YOLOv2-Tiny object detection network, based on the comparison of the input feature map and weight data volume of each layer, it is determined that weight reuse is adopted for the first four layers in the systolic array, and input feature map reuse is adopted for the last five layers.

[0030] Step 2: Use the roofline model to find the optimal scale of the systolic array to achieve the best computing power and appropriate bandwidth of the YOLOv2-tiny model on the FPGA platform. First, calculate the total number of operations required for one layer in the YOLOv2-tiny network model, as shown in Equation 1. The total memory access volume only considers the data transfer from outside the accelerator to inside the accelerator, and the data transfer inside the accelerator is not included in the total memory access volume. The total number of times that need to be calculated for this layer in the systolic array is as shown in Equation 2. The sizes of the input feature map read / write cache unit, weight read / write cache unit, and output result read / write cache unit are as shown in Equations 4-1, 4-2, and 4-3. The number of times that the input feature map, weight, and output result need to be transmitted through AXI4-Stream is as shown in Equations 5-1, 5-2, and 5-3. Using the roofline formula, the computing performance and operation intensity of the model can be calculated as shown in Equations 6 and 7 respectively. Taking the ninth layer in the YOLOv2-Tiny network as an example, by enumerating Tr, Tc, and Tm in Equation 6 and under the constraint of Equation 7, the roofline model of a certain layer of YOLOv2-Tiny, the relationship between the systolic array scale and the calculation intensity in Equation 6, and the relationship between the systolic array scale and the operation intensity in Equation 7 can be obtained. Sequentially count the systolic array scales corresponding to the nine layers in the YOLOv2-Tiny network, and finally select the most appropriate scale for the accelerator design.

[0031] Step 3: After determining the data reuse method and the scale of the systolic array through the first two steps, next, partition the input data and design the sizes of the input feature map register unit 1, input feature map read / write cache unit 2, weight read / write cache unit 4, output result read / write cache unit 6, as well as the parameters and control logic of the global configuration unit 7.

[0032] Step 4: After completing the above work, the systolic array can be executed. First, the input feature map stored in the external memory of the accelerator is transmitted to the input feature map register unit 1 through the AXI4-Stream bus, and the weight data is transmitted to the weight read / write cache unit 4 through the AXI4-Stream bus. Then, the global configuration unit 7 configures the values in the input feature map register unit 1 into the form required by the systolic array and stores them in the input feature map read / write cache unit 2. Finally, the input feature map is output to each row of the systolic array in sequence, and the weight data is output to each column of the systolic array in sequence.

[0033] For the described single computing unit, its output data is passed to the previous computing unit in a pipelined manner. The calculation results of the first row of the systolic array calculation part will be pooled in the pooling unit and then stored in the output result read / write cache unit according to the designed data arrangement method. Finally, the output result is transmitted to the external memory of the accelerator through the AXI4-Stream bus.

[0034] Beneficial effects: Since the present invention adopts the above technical solutions, the present invention has the following advantages:

[0035] 1. Based on the systolic array computing architecture, without affecting the original accuracy of the network, it can optimize memory operations, reduce external storage and read bandwidth, and improve data utilization.

[0036] 2. Based on the pipelined method to transfer calculation results, the pipeline delay time can be flexibly designed according to the parameters of different layers, so as to increase the computing parallelism and improve the throughput of the target detection accelerator.

[0037] 3. The systolic array utilizes FPGA resources to achieve high throughput. As the capacity of the hardware resources in the FPGA continues to increase, the systolic array-based YOLOv2-Tiny target detection network can be applied to other devices using design space exploration. This optimization method helps to balance the relationship between high throughput and clock frequency.

[0038] 4. For the YOLOv2-Tiny target detection network, the data volume of the convolutional kernels in the first four layers is small, and the data volume of the convolutional kernels in the last five layers is much larger than that of the input feature map. The corresponding output bandwidth varies greatly under different data reuse modes. Using different data reuse modes for different layers of the network can relieve the huge pressure on the data transfer bandwidth. The convolutional kernel reuse is used in the first four layers, and the input feature map reuse is used in the last five layers, thereby reducing the area and power consumption caused by bandwidth.

[0039] From the perspective of improving the throughput of the accelerator, the present invention explores the optimal scale of the systolic array in the field of YOLOv2-Tiny object detection, and combines the strategies of pipelining and hybrid data reuse to improve the utilization rate of hardware resources while ensuring the throughput. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a schematic diagram of the overall structure of a high-throughput object detection accelerator based on a systolic array implemented in the present invention;

[0041] Figure 2 is a schematic diagram of the structure of a single computing unit within a systolic array implemented in the present invention;

[0042] Figure 3 is a schematic diagram of roofline analysis for YOLOv2-Tiny implemented in the present invention;

[0043] Figure 4 is a schematic diagram of the relationship between the scale of the systolic array and the computational intensity implemented in the present invention;

[0044] Figure 5 is a schematic diagram of the relationship between the scale of the systolic array and the operation intensity implemented in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] A high-throughput object detection accelerator based on a systolic array of the present invention, the accelerator includes: an input feature map storage unit 1, an input feature map read / write cache unit 2, a systolic array computing part 3, a weight read / write cache unit 4, a pooling unit 5, an output result read / write cache unit 6, and a global configuration unit 7; wherein the output of the global configuration unit 7 is respectively connected to the input feature map storage unit 1, the input feature map read / write cache unit 2, the weight read / write cache unit 4, the pooling unit 5, and the output result read / write cache unit 6; the output of the input feature map storage unit 1 is connected to the input feature map read / write cache unit 2, and the output of the input feature map read / write cache unit 2 is connected to the systolic array computing part 3; the output of the weight read / write cache unit 4 is respectively connected to the systolic array computing part 3; the input of the pooling unit 5 is connected to the systolic array computing part 3; the output of the pooling unit 5 is connected to the output result read / write cache unit 6; YOLOv2-Tiny object detection is implemented in the high-throughput object detection accelerator based on the systolic array.

[0046] The pulsating array calculation part 3 includes multiple individual calculation units, and each individual calculation unit includes: a feature map input port 31, an input feature map register unit 32, a weight input port 33, a weight register unit 34, and then convolution calculation is performed through a multiplier 35 and an adder 36. It is judged by a judge 37 whether the complete convolution calculation is completed. If not, the calculation part and sum are stored in a register 38 for continued accumulation. If the calculation is completed, the result is subjected to pooling processing 39; the input feature map is transmitted to the right calculation unit 314 through a register 312, the weight is transmitted to the lower calculation unit 315 through a register 313, and the output result 316 of the lower calculation unit is temporarily stored in a register 317. The data selector 310 is used to select whether to output the current calculation result or the calculation result transmitted from the next pulsating unit.

[0047] In the described individual calculation unit, a register unit is added at the pulsating transmission of the output result to form a pipeline design. According to the parameters of different layers of the YOLOv2-Tiny object detection algorithm, the delay time of the pipeline can be flexibly designed to increase the parallelism of calculation and improve the throughput of the accelerator.

[0048] The high-throughput object detection accelerator operates in the following mode: The input feature map existing in the external memory of the accelerator is transmitted to the input feature map register unit 1 through the AXI4-Stream bus. Then, the global configuration unit 7 controls the unit to configure the value in the input feature map register unit 1 into the form required by the pulsating array and store it in the input feature map read-write cache unit 2, and finally output it to each row of the pulsating array for calculation in turn; the weight data is transmitted to the weight read-write cache unit 4 through the AXI4-Stream bus, and then the weight data is output to each column of the pulsating array in turn to perform convolution calculation with the data input to the rows of the pulsating array; the result of each calculation unit is transmitted to the previous calculation unit. The calculation result of the first layer will be subjected to pooling operation in the pooling unit 5 and then stored in the output result read-write cache unit 6 according to the designed data arrangement method. Finally, the output result is transmitted to the external memory of the accelerator through the AXI4-Stream bus.

[0049] For the described pulsating array calculation part, for the given YOLOv2-Tiny object detection algorithm, the pulsating scale that can be designed on different hardware platforms is different. In order to enable the YOLOv2-Tiny object detection algorithm to obtain higher resource utilization and throughput on the hardware platform, the step of design space exploration is introduced. The input data, weight data, and output data of each layer of the YOLOv2-Tiny object detection algorithm are selected in turn. The three-dimensional scale of the output data is recorded as R rows, C columns, and M channels, and the four-dimensional scale of the weight data is K rows, K columns, N channels, and M blocks. Therefore, the number of calculation operations Ops of this layer of the YOLOv2-Tiny object detection algorithm is formula 1.

[0050] Ops = 2 × R × C × M × N × K × K Formula 1

[0051] Chunk the output data. Denote Tm as the chunk size of the channel, Tr as the row chunk size, and Tc as the column chunk size. The total number of chunks Blocks and the time Timeperblock required to calculate each chunk are as shown in Formula 2 and Formula 3

[0052]

[0053] Timeperblock = K × K × N + Tr + Tc - 1 Formula 3

[0054] Denote B in 、B w 、B out as the sizes of the input feature map register unit 1, the weight read / write cache unit 4, and the output result read / write cache unit 6 required respectively when calculating the output data of one chunk. S is the sliding window step size of the convolution calculation. They are as shown in Formula 4-1, 4-2, and 4-3 respectively

[0055] B in = N × (S × Tr + K - S) × (S × Tc + K - S) Formula 4-1

[0056] B w = K × K × N × Tm Formula 4-2

[0057] B out = Tr × Tc × Tm Formula 4-3

[0058] Denote α in as the number of times the input data needs to be transferred from outside the accelerator to the input feature map register unit 1 inside the accelerator via the AXI4-Stream bus when calculating the output data of all chunks in one layer. α w as the number of times the weight data needs to be transferred from outside the accelerator to the weight read / write cache unit 4 inside the accelerator via the AXI4-Stream bus when calculating the output data of all chunks in one layer. α out as the number of times the output data needs to be transferred from the output result read / write cache unit 6 inside the accelerator to outside the accelerator via the AXI4-Stream bus when calculating the output data of all chunks in one layer. They are as shown in Formula 5-1, 5-2, and 5-3 respectively

[0059]

[0060]

[0061]

[0062] The computational performance and operation intensity of YOLOv2-Tiny can be calculated using the roofline model as shown in Equations 6 and 7 respectively;

[0063]

[0064]

[0065] By enumerating Tr, Tc, and Tm in Equation 6 and under the constraint of Equation 7, the most suitable systolic array scale for the YOLOv2-Tiny object detection algorithm is obtained.

[0066] For the systolic array calculation part, a hybrid data reuse strategy is adopted; for the YOLOv2-Tiny object detection network, the data volume of the first 4 convolutional kernels is small, so weight reuse is adopted, and the data volume of the last five convolutional kernels is much larger than that of the input feature map, so the input feature map reuse method is adopted.

[0067] As Figure 1 shown, the high-throughput YOLOv2-Tiny object detection accelerator based on systolic array of the present invention includes: an input feature map register unit 1, an input feature map read / write cache unit 2, a systolic array calculation part 3, a weight read / write cache unit 4, a pooling unit 5, an output result read / write cache unit 6, and a global configuration unit 7. The systolic array accelerator has the characteristics of low global data transmission and high clock frequency, and is suitable for large-scale parallel design on FPGA. The invention can be extended to the design of other network models based on systolic array calculation. The steps to implement YOLOv2-Tiny object detection in a systolic array-based accelerator are as follows:

[0068] Step 1: Find a feasible hardware circuit mapping method in the systolic array to ensure that correct data is available at specific positions in the PE array in each cycle. Considering the large difference in the computational amount between layers of the YOLOv2-Tiny object detection network, according to the comparison of the data volume of the input feature map and weights of each layer, it is determined that weight reuse is adopted for the first four layers in the systolic array, and input feature map reuse is adopted for the last five layers.

[0069] Step 2: Use the roofline model to find the optimal scale of the systolic array to achieve the best computing power and appropriate bandwidth of the YOLOv2-Tiny model on the FPGA platform. The specific calculation method is as Figure 3 , first calculate one layer in the YOLOv2-Tiny network model, denote Tm as the channel block size, Tr as the row block size, Tc as the column block size,, α inWhen calculating the number of times the input data needs to be transferred from outside the accelerator to the input feature map storage unit 1 inside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated, α w When calculating the number of times the weight data needs to be transferred from outside the accelerator to the weight read / write cache unit 4 inside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated, α out When calculating the number of times the output data needs to be transferred from the output result read / write cache unit 6 inside the accelerator to outside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated, B in 、B w 、B out are the sizes of the input feature map storage unit 1, the weight read / write cache unit 4, and the output result read / write cache unit 6 required when calculating the output data of one block respectively. S is the sliding step. The total number of operations required for a certain layer is as shown in Formula 1. Considering only the data transfer from outside the accelerator to inside the accelerator for the total memory access volume, the total number of sub-blocks in the systolic array for this layer is as shown in Formula 2. The sizes of the input feature map read / write cache unit, the weight read / write cache unit, and the output result read / write cache unit are as shown in Formulas 4-1, 4-2, and 4-3. The number of times the input feature map, weight, and output result need to be transferred via AXI4-Stream is as shown in Formulas 5-1, 5-2, and 5-3. Using the roofline formula, the computing performance and operation intensity of the model can be calculated as shown in Formulas 6 and 7 respectively. By enumerating Tr, Tc, and Tm in Formula 6 and under the constraint of Formula 7, the most suitable systolic array scale for a certain layer of YOLOv2-Tiny can be obtained as Figure 4 . Sequentially count the systolic array scales corresponding to the nine layers in the YOLOv2-Tiny network, and finally select the most suitable scale for the accelerator design.

[0070] Step 3: After determining the data reuse method and the scale of the systolic array through the first two steps, the next step is to divide the input data into blocks, and design the sizes of the input feature map storage unit 1, the input feature map read / write cache unit 2, the weight read / write cache unit 4, the output result read / write cache unit 6, as well as the parameters and control logic of the global configuration unit 7.

[0071] Step 4: After completing the above work, the systolic array can be executed. First, the input feature map stored in the external memory of the accelerator is transmitted to the input feature map register unit 1 through AXI4-Stream, and the weight data is transmitted to the weight read / write cache unit 4 through AXI4-Stream. Then, the global configuration unit 7 configures the values in the input feature map register unit 1 into the form required by the systolic array and stores them in the input feature map read / write cache unit 2. Finally, the input feature map is sequentially output to each row of the systolic array, and the weight data is sequentially output to each column of the systolic array.

[0072] Further, as Figure 2 shown, the computing unit passes the input data to the computing unit on the right, the weight data to the computing unit below, and the output data is passed to the previous computing unit in a pipelined manner. The calculation result of the Tm column in the first row of the systolic array will be pooled in the pooling unit 5 and stored in the output result read / write cache unit 6 according to the designed data arrangement method. Finally, the output result is transmitted to the outside of the accelerator through AXI4-Stream.

Claims

1. A high-throughput object detection accelerator based on a systolic array, characterized in that: The accelerator includes: an input feature map storage unit (1), an input feature map read / write cache unit (2), a systolic array computing part (3), a weight read / write cache unit (4), a pooling unit (5), an output result read / write cache unit (6), and a global configuration unit (7); the output of the global configuration unit (7) is respectively connected to the input feature map storage unit (1), the input feature map read / write cache unit (2), the weight read / write cache unit (4), the pooling unit (5), and the output result read / write cache unit (6); the output of the input feature map storage unit (1) is connected to the input feature map read / write cache unit (2), and the output of the input feature map read / write cache unit (2) is connected to the systolic array computing part (3); the output of the weight read / write cache unit (4) is respectively connected to the systolic array computing part (3); the input of the pooling unit (5) is connected to the systolic array computing part (3); the output of the pooling unit (5) is connected to the output result read / write cache unit (6); YOLOv2-Tiny object detection is implemented in the high-throughput object detection accelerator based on the systolic array; The high-throughput object detection accelerator operates in the following mode: the input feature map existing in the external memory of the accelerator is transmitted to the input feature map storage unit (1) through the AXI4-Stream bus, and then the control unit of the global configuration unit (7) configures the value in the input feature map storage unit (1) into the form required by the systolic array and stores it in the input feature map read / write cache unit (2), and finally outputs it to each row of the systolic array for calculation in sequence; the weight data is transmitted to the weight read / write cache unit (4) through the AXI4-Stream bus, and then the weight data is output to the systolic array column by column in sequence to perform convolution calculation with the data input to the row of the systolic array; the result of each computing unit is passed to the previous computing unit, and the result of the first layer will be stored in the output result read / write cache unit (6) after the pooling operation in the pooling unit (5) according to the designed data arrangement method, and finally the output result is transmitted to the external memory of the accelerator through the AXI4-Stream bus; The systolic array computing part (3) sequentially selects the input data, weight data, and output data of each layer of the YOLOv2-Tiny object detection algorithm. Denote the three-dimensional scale of the output data as R rows, C columns, and M channels, and the four-dimensional scale of the weight data as K rows, K columns, N channels, and M blocks. Therefore, the number of computing operations Ops of this layer of the YOLOv2-Tiny object detection algorithm is given by Formula 1, Ops = 2 × R × C × M × N × K × K Formula 1 The output data is divided into blocks. Denote Tm as the block size of the channel, Tr as the row block size, and Tc as the column block size. The total number of blocks Blocks and the time required to calculate each block Time per block are given by Formula 2 and Formula 3, Time perblock = K × K × N + Tr + Tc - 1 Formula 3 Denote B in , B w , B out are respectively the sizes of the input feature map register unit, weight read / write cache unit, and output result read / write cache unit required for calculating a piece of output data. S is the sliding window step size of the convolution calculation, as shown in formulas 4-1, 4-2, and 4-3 respectively, B in = N×(S×Tr + K - S)×(S×Tc + K - S) Formula 4-1 B w = K × K × N × Tm Formula 4-2 B out = Tr × Tc × Tm Equation 4-3 Denote α in as the number of times the input data needs to be transferred from the outside of the accelerator to the input feature map storage unit (1) inside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated. α w is the number of times the weight data needs to be transferred from the outside of the accelerator to the weight read / write cache unit (4) inside the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated. α out is the number of times the output data needs to be transferred from the output result read / write cache unit (6) inside the accelerator to the outside of the accelerator via the AXI4-Stream bus when the output data of all blocks in one layer is calculated, as shown in formulas 5-1, 5-2, 5-3, The computing performance Performance and the operation intensity CTC of YOLOv2-Tiny can be calculated using the roofline model as shown in Formulas 6 and 7 respectively; By enumerating Tr, Tc, and Tm in Formula 6 and under the constraint of Formula 7, the most suitable systolic array scale for the YOLOv2-Tiny object detection algorithm is obtained.

2. The high-throughput target detection accelerator based on a systolic array according to claim 1, wherein: The systolic array computing part (3) includes multiple single computing units, and each single computing unit includes: a feature map input port (31), an input feature map register unit (32), a weight input port (33), a weight register unit (34), and then convolution calculations are performed through a multiplier (35) and an adder (36). The judge (37) determines whether the complete convolution calculation is completed. If not, the calculation part and the sum are stored in the register (38) and continue to be accumulated. If the calculation is completed, the result is pooled (39); the input feature map is passed to the right computing unit (314) through the register (312), and the weight is passed to the lower computing unit (315) through the register (313). The output result (316) of the lower computing unit is temporarily stored in the register (317), and the data selector (310) is used to select whether to output the current calculation result or the calculation result passed from the next systolic unit.

3. The high-throughput target detection accelerator based on systolic array according to claim 2, characterized in that: For the described single computing unit, a register unit is added at the output result systolic transfer point to form a pipelined design. According to the parameters of different layers of the YOLOv2-Tiny object detection algorithm, the pipeline delay time can be flexibly designed to increase the computing parallelism and improve the throughput of the accelerator.

4. The high-throughput target detection accelerator based on systolic array according to claim 1, characterized in that: The described systolic array computing part (3) adopts a hybrid data reuse strategy; for the YOLOv2-Tiny object detection network, the data volume of the convolution kernels in the first 4 layers is small, and weight reuse is adopted. The data volume of the convolution kernels in the last 5 layers is much larger than the data volume of the input feature map, and the input feature map reuse method is adopted.

5. The high-throughput target detection accelerator based on a systolic array according to claim 1, characterized in that: The steps to implement YOLOv2-Tiny object detection in the accelerator based on the systolic array are as follows: Step 1: Find a feasible hardware circuit mapping method in the systolic array to ensure that correct data is available at specific positions in the computing unit in each cycle; considering the large difference in the computing volume between different layers of the YOLOv2-Tiny object detection network, according to the comparison of the data volume of the input feature map and the weight of each layer, it is determined that weight reuse is adopted in the systolic array for the first 4 layers, and input feature map reuse is adopted for the last 5 layers; Step 2: Use the roofline model to find the optimal size of the systolic array to achieve the best computing power and appropriate bandwidth of the YOLOv2-tiny model on the FPGA platform. First, calculate one layer in the YOLOv2-tiny network model, and calculate the total number of operations required for this layer as shown in Equation 1. The total memory access volume only considers the data transfer from outside the accelerator to inside the accelerator, and the data transfer inside the accelerator is not included in the total memory access volume. The total number of times this layer needs to be calculated in the systolic array is as shown in Equation 2. The sizes of the weight read / write cache unit and the output result read / write cache unit are as shown in Equations 4-1, 4-2, and 4-3. The number of times the input feature map, weights, and output results need to be transmitted through the AXI4-Stream is as shown in Equations 5-1, 5-2, and 5-3. For the input feature map read / write cache unit, the computing performance and operation intensity of the model can be calculated using the roofline formula as shown in Equations 6 and 7 respectively; Step 3: After determining the data reuse method and the size of the systolic array through the first two steps, the next step is to block the input data and design the sizes of the input feature map register unit (1), the input feature map read / write cache unit (2), the weight read / write cache unit (4), the output result read / write cache unit (6), as well as the parameters and control logic of the global configuration unit (7); Step 4: After completing the above work, the calculation of the systolic array (3) can be executed; first, the input feature map stored in the external memory of the accelerator is transmitted to the input feature map register unit (1) through the AXI4-Stream bus, and the weight data is transmitted to the weight read / write cache unit (4) through the AXI4-Stream bus. Then, the global configuration unit (7) configures the values in the input feature map register unit (1) into the form required by the systolic array and stores them in the input feature map read / write cache unit (2). Finally, the input feature map is sequentially output to each row of the systolic array, and the weight data is sequentially output to each column of the systolic array.

6. The high-throughput target detection accelerator based on a systolic array according to claim 2, characterized in that: For the described single computing unit, its output data is passed to the previous computing unit in a pipelined manner. The calculation results of the first row of the systolic array calculation part (3) will be pooled in the pooling unit (5) and then stored in the output result read / write cache unit (6) according to the designed data arrangement method. Finally, the output result is transmitted to the memory outside the accelerator through the AXI4-Stream bus.

Citation Information

Patent Citations

  • A universal convolutional neural network accelerator based on a one-dimensional pulsation array

    CN109934339A

  • Convolutional neural network hardware acceleration architecture based on FPGA

    CN110135554A