Sparse neural network accelerated calculation system and method based on FPGA (Field Programmable Gate Array)
By designing a sparse neural network accelerated computing system based on FPGA, the sparse data decoding unit is used to decode non-zero weighted data and its location, efficient sparse neural network computing is realized, solving the problem that the existing FPGA architecture cannot handle sparse coded data, and improving computing efficiency and resource utilization.
Patent Information
- Application Number
- CN202510389899.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing FPGA neural network computing architecture cannot process sparsely encoded data, resulting in wasted computing resources and inefficient, making it difficult to directly use it for sparse neural network inference calculations.
A sparse neural network accelerated computing system based on FPGA is designed, including feature map cache unit, sparse weight cache unit, sparse data decoding unit and calculation engine unit. The sparse data decoding unit realizes hardware real-time decoding of sparse encoded data, and only non-zero weight-related convolutional calculations are performed to make full use of network sparseness.
It realizes efficient sparse neural network inference calculation, reduces the number of memory accesses, improves computing efficiency, and has strong system architecture versatility, and is suitable for a variety of FPGA scenarios.
Smart Images

Figure CN120338006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sparse neural network acceleration computing system and method based on FPGA, belonging to the technical field of neural networks. Background Art
[0002] Due to its internally array-distributed computing resources and hardware reconfigurability, FPGA can fully fit the data dependence relationship and hierarchical parallel relationship in the neural network computing process in the design of the computing architecture, realizing high parallelism and pipelining, with low memory access bandwidth requirements and small hardware overhead, which can improve computing throughput, reduce computing latency and improve energy efficiency, and is widely used in dense neural network acceleration computing. However, the neural network sparsification process introduces compressed encoded data to mark the positions of non-zero elements. Facing the sparse neural network computing tasks with irregular network connections and uncertain computing amounts between layers and channels, the FPGA computing architecture for accelerating dense neural networks cannot process the encoded data and cannot utilize the sparsity of the network, resulting in a waste of a large amount of computing resources and low computing efficiency. The existence of the above problems makes it difficult to directly use the existing FPGA neural network computing architecture for sparse neural network inference computing. Summary of the Invention
[0003] The present invention aims to solve the technical problems in the prior art that the existing FPGA neural network computing architecture cannot process the encoded data and cannot utilize the sparsity of the network, resulting in a waste of a large amount of computing resources and low computing efficiency, and further proposes a sparse neural network acceleration computing system and method based on FPGA.
[0004] The technical solution adopted by the present invention to solve the above problems is: The present invention includes a sparse neural network acceleration computing system based on FPGA, and the system includes:
[0005] A feature map cache unit, configured to receive and cache input data from a memory access unit DMA; receive the final calculation result sent by a calculation engine unit, and send the final calculation result to the memory access unit DMA;
[0006] A sparse weight cache unit, configured to read and cache the weight parameters after sparse encoding from the memory access unit DMA;
[0007] A sparse data decoding unit, configured to read the weight parameters after sparse encoding from the sparse weight cache unit and perform decoding to obtain non-zero weight data and their corresponding position coordinates;
[0008] The computing engine unit is used to read the non-zero weight data and its corresponding position coordinates from the sparse data decoding unit, and read the input data from the feature map cache unit, and perform convolution calculations to obtain the final calculation result.
[0009] In some embodiments, the computing engine unit includes:
[0010] A computing array unit, an output cache unit, an activation unit, and a pooling unit;
[0011] Among them, the computing array unit is composed of a plurality of processing elements (PEs) interconnected to form a streaming structure for performing convolution calculations.
[0012] In some embodiments, the output cache unit is used to cache the intermediate calculation results of the computing array unit, the activation calculation results of the activation unit, and the pooling calculation results of the pooling unit, and output the intermediate calculation results, the activation calculation results, and the pooling calculation results to the computing array unit when needed.
[0013] In some embodiments, the processing element (PE) includes:
[0014] A multiply-accumulate unit, an internal cache unit, and an adder;
[0015] Among them, the adder is used to take the partial sum output result of the previous row of processing elements (PEs) as the partial sum input, accumulate it with the partial sum result of the current row of processing elements (PEs), obtain the partial sum output result of the current row of processing elements (PEs), and output it to the next row of processing elements (PEs) as the partial sum input.
[0016] In some embodiments, the multiply-accumulate unit includes:
[0017] A sub-multiplier and a sub-adder;
[0018] Among them, the sub-multiplier is used to multiply the input data read from the feature map cache unit by the non-zero weight data to obtain a multiplication result; the sub-adder is used to add the multiplication result to the partial sum data stored in the internal cache unit and then store it back into the internal cache unit.
[0019] In some embodiments, the internal cache unit includes:
[0020] An address generation unit and a sub-cache unit;
[0021] Among them, the address generation unit is used to determine the validity of the input data based on the position coordinates corresponding to the non-zero weight data and the position coordinates in the input data, and calculate the cache address of the calculation result of the multiply-accumulate unit in the sub-cache unit; the sub-cache unit is used to output partial sum data and partial sum results.
[0022] In some embodiments, the sparse data decoding unit includes:
[0023] A shift register unit, a control logic unit, a row-column counter unit, and a row-column coordinate cache unit.
[0024] The present invention also includes an FPGA-based accelerated calculation method for sparse neural networks, which is applied to the FPGA-based accelerated calculation system for sparse neural networks described above. The method includes:
[0025] The feature map cache unit receives and caches the input data from the memory access unit DMA;
[0026] The sparse weight cache unit reads and caches the weight parameters after sparse coding from the memory access unit DMA;
[0027] The sparse data decoding unit reads the weight parameters after sparse coding from the sparse weight cache unit and decodes them to obtain non-zero weight data and their corresponding position coordinates;
[0028] The calculation engine unit reads the non-zero weight data and their corresponding position coordinates from the sparse data decoding unit, and reads the input data from the feature map cache unit, and performs convolution calculation to obtain the final calculation result;
[0029] The feature map cache unit receives the final calculation result and sends the final calculation result to the memory access unit DMA.
[0030] The beneficial effects of the present invention are:
[0031] 1. The present invention realizes the hardware real-time decoding of sparse coding data through the design of the sparse data decoding unit.
[0032] 2. The present invention decodes the weight parameters after sparse coding through the sparse data decoding unit to obtain non-zero weight data and their corresponding position coordinates, and then only performs convolution calculations related to non-zero weights, making full use of the network sparsity to realize efficient sparse neural network inference calculations;
[0033] 3. The system architecture of the present invention is based on the FPGA architecture, so it can be used in any scenario where the FPGA architecture can be used, and has a wide range of applications. Description of the Drawings
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0035] Figure 1 It is a schematic structural diagram of a sparse neural network acceleration computing system based on FPGA provided by this application;
[0036] Figure 2 It is a schematic structural diagram of a computing engine unit provided by this application;
[0037] Figure 3 It is a schematic structural diagram of a computing array unit provided by this application;
[0038] Figure 4 It is a schematic diagram of the ReLU activation function image provided by this application;
[0039] Figure 5 It is a schematic diagram of the process of performing max-pooling calculation on a 2×2 area provided by this application;
[0040] Figure 6 It is a schematic structural diagram of the max-pooling unit design provided by this application;
[0041] Figure 7 It is a schematic diagram of the structure of the processing unit PE provided by this application;
[0042] Figure 8 It is a schematic diagram of the results of using COO, CSR, and CSC compression encoding methods for a two-dimensional sparse matrix respectively provided by this application;
[0043] Figure 9 It is a schematic diagram of the Bitmap compression format encoding and decoding process provided by this application;
[0044] Figure 10 It is a schematic structural diagram of the sparse data decoding unit provided by this application;
[0045] Figure 11 It is a schematic flowchart of a sparse neural network acceleration computing method based on FPGA provided by this application;
[0046] Figure 12 It is a schematic structural diagram of a sparse neural network acceleration computing module based on Zynq deployed with a sparse neural network acceleration computing system based on FPGA provided by this application;
[0047] Figure 13Schematic diagram of the sparse neural network acceleration computing test system built according to the Zynq-based sparse neural network acceleration computing module provided by this application. Detailed implementation manners Detailed implementation manner one:
[0049] Combined with Figure 1 Describe this implementation manner. This implementation manner provides a sparse neural network acceleration computing system based on FPGA. The system includes:
[0050] The feature map cache unit 101 is used to receive and cache the input data from the memory access unit (Direct Memory Access, DMA); receive the final calculation result sent by the computing engine unit, and send the final calculation result to the memory access unit DMA.
[0051] The sparse weight cache unit 102 is used to read and cache the weight parameters after sparse coding from the memory access unit DMA.
[0052] The sparse data decoding unit 103 is used to read the weight parameters after sparse coding from the sparse weight cache unit and perform decoding to obtain the non-zero weight data and their corresponding position coordinates.
[0053] The computing engine unit 104 is used to read the non-zero weight data and their corresponding position coordinates from the sparse data decoding unit, and read the input data from the feature map cache unit, and perform convolution calculation to obtain the final calculation result.
[0054] In some embodiments, the computing engine unit includes:
[0055] A computing array unit, an output cache unit, an activation unit, and a pooling unit;
[0056] Among them, the computing array unit is composed of multiple processing elements (PEs) interconnected to form a streaming structure for performing convolution calculation.
[0057] As Figure 2 shown, it is the schematic diagram of the computing engine unit, consisting of Figure 2It can be seen that the computing engine unit includes a computing array unit, an output buffer unit, an activation unit, and a pooling unit. Among them, the computing array reads data from the feature map caches of the decoding units respectively and performs convolution calculations in the computing array according to requirements. The computing array unit is formed by interconnecting multiple processing elements (PEs) into a streaming structure. Each processing element PE internally includes a multiply-accumulate unit and an internal buffer unit. Since the resource scale of the computing array unit is limited, for relatively large weights, block calculations need to be performed. Therefore, an output buffer unit is required to temporarily store the intermediate calculation results. The activation unit and the pooling unit perform activation and pooling calculations on the output results after accumulation and return the final results to the output buffer for subsequent calculations or directly output the calculation results.
[0058] Specifically, for the sparse convolution calculation process and characteristics, the specific structure of the computing array unit adopted by the present invention is as Figure 3 shown, which is a schematic structural diagram of the computing array unit.
[0059] From Figure 3 it can be seen that the computing array unit is responsible for performing convolution calculations in the sparse neural network and is composed of PEs arranged in N rows and N columns interconnected. Among them, the PEs in the same row share the input data of the feature map cache unit of the same input channel. The leftmost PE in each row is connected to the feature map cache unit and the decoding unit (i.e., the sparse data decoding unit, Figure 3 and the weight cache in front of the decoding unit in the figure refers to the sparse weight cache unit), receives the input data, the decoded non-zero weight data, and the corresponding position coordinates, and passes the input data to all the PEs in the same row to the right. In this way, the full reuse of the input data can be realized, the number of memory accesses can be reduced, and the computing efficiency can be improved.
[0060] The PEs in the same column are responsible for processing the input data of the feature map cache unit of the same output channel. After the PEs in a column complete the calculation, the PEs in the same column pass and accumulate the calculation results of each input channel from top to bottom, and finally obtain the final output data at the bottom and store it in the output buffer unit for subsequent calculation processes.
[0061] It should be further noted that the activation function in the activation unit is a non-linear function, and its main function is to increase non-linear calculations in the neural network model. Theoretically, it can make the network approximate any function and greatly increase the expression ability of the neural network. Otherwise, the relationship between the network output and the input is a strictly linear relationship, and the fitting and approximation ability is poor. Considering that the ReLU activation function is widely used in common object detection networks, this project designs a hardware structure to implement the ReLU activation function calculation, as Figure 4As shown, it is a schematic diagram of the ReLU activation function image, where the horizontal axis represents the input data and the vertical axis represents the output of the activation function.
[0062] When the input of the ReLU activation function is less than 0, the output is 0; when the output is greater than 0, the output value is equal to the input value. This calculation mode can be directly implemented by hardware logic through conditional judgment and numerical comparison. When comparing the input value with 0, if it is true, the numerical value is directly output; otherwise, 0 is output.
[0063] In the pooling unit, pooling is also known as downsampling. Its main function is to reduce the dimension of the feature map while keeping the original features unchanged, reduce the network complexity, remove redundant information, and prevent the model from overlearning the features, resulting in poor network generality. Max pooling is a commonly used pooling method in object detection networks. Therefore, the architecture of this topic uses hardware to implement the max pooling operation. For each pooling region, the maximum value in the region is selected as the output of the pooling result. As Figure 5 shown, it is a schematic diagram of the process of performing max pooling calculation on a 2×2 region.
[0064] From Figure 5 it can be seen that for each 2×2 region, the maximum value in the region is selected as the output result of the max pooling operation. Correspondingly, as Figure 6 shown, it is a schematic diagram of the design structure of the max pooling unit.
[0065] When the max pooling unit performs pooling calculation, it first calculates the addresses of each corresponding element in the sliding window in the output cache according to the pooling region size and the current pooling sliding window position. According to the calculated addresses, the output cache outputs the numerical values of the corresponding positions of the elements, compares the read numerical values with the values in the maximum value register (initially 0). If the read numerical value is greater than the current maximum value, it controls the gating of the current data and updates the maximum value register to the current numerical value. Repeating the above process for each element in the sliding window can obtain the maximum element value in the sliding window and store it back in the output cache. By continuously changing the sliding window position and repeating the above operations, the complete pooling calculation result is finally obtained.
[0066] In some embodiments, the output cache unit is used to cache the intermediate calculation results of the calculation array unit, the activation calculation results of the activation unit, and the pooling calculation results of the pooling unit, and output the intermediate calculation results, the activation calculation results, and the pooling calculation results to the calculation array unit when needed.
[0067] In some embodiments, the processing unit PE includes:
[0068] a multiply-accumulate unit, an internal cache unit, and an adder;
[0069] Among them, the adder is used to take the partial sum output result of the previous row processing unit PE as the partial sum input, accumulate it with the partial sum result of the current row processing unit PE, obtain the partial sum output result of the current row processing unit PE, and output it to the next row processing unit PE as the partial sum input.
[0070] In some embodiments, the multiply-accumulate unit includes:
[0071] A sub-multiplier and a sub-adder;
[0072] Among them, the sub-multiplier is used to multiply the input data read from the feature map cache unit by the non-zero weight data to obtain a multiplication result; the sub-adder is used to add the multiplication result to the partial sum data stored in the internal cache unit and then store it back into the internal cache unit.
[0073] In some embodiments, the internal cache unit includes:
[0074] An address generation unit and a sub-cache unit;
[0075] Among them, the address generation unit is used to judge the validity of the input data based on the position coordinates corresponding to the non-zero weight data and the position coordinates in the input data, and calculate the cache address of the multiply-accumulate unit calculation result in the sub-cache unit; the sub-cache unit is used to output partial sum data and partial sum results.
[0076] It should be noted that the schematic diagram of the structure of the processing unit PE is as Figure 7 shown. Each processing unit PE internally includes two parts: a multiply-accumulate unit and an internal cache unit, as well as an adder. The adder is used to accumulate the partial sum output result of the previous row processing unit PE with the partial sum result of the current processing unit PE and output it to the next row processing unit PE. The multiply-accumulate unit consists of a sub-multiplier and a sub-adder. The sub-multiplier multiplies the input data of the input feature map unit by the non-zero weight data, and then the sub-adder adds the multiplication result to the partial sum stored in the sub-cache unit and stores it back into the sub-cache unit. The internal cache unit includes an address generation unit and a sub-cache unit. The address generation unit judges the validity of the input data based on the position coordinates of the input data and the position coordinates corresponding to the non-zero weight, and calculates the corresponding cache address of the multiply-accumulate unit calculation result in the sub-cache unit.
[0077] When performing convolution calculations, first, each row of processing elements (PEs) reads the input data, non-zero weight data, and corresponding position coordinates from the corresponding feature map cache unit and sparse data decoding unit. Subsequently, the processing element PE starts the convolution calculation, successively multiplying each non-zero element in the convolution kernel with the corresponding feature map element in the input data and accumulating the result with the partial sum already stored in the sub-cache unit, and then storing the result back in the sub-cache unit. After all the processing elements PE in the computing array have completed the calculation, the inter-row accumulation process is started. The calculation results of each row of processing elements PE are passed and accumulated from top to bottom. Finally, the calculated output data feature map is stored in the output cache unit, and the calculation results of each output channel in each column are separately cached and stored.
[0078] In some embodiments, the sparse data decoding unit includes:
[0079] a shift register unit, a control logic unit, a row-column counter unit, and a row-column coordinate cache unit
[0080] It should be noted that compressing the weight data and only retaining the non-zero weight data can effectively reduce the storage space occupation and bandwidth requirements of the weight data. However, since the weights are irregular after sparsification, encoding is required to record the position information of the weights. Adopting a reasonable sparse weight encoding and making full use of the weight sparsity can bring improvements in performance and energy efficiency. Common compression encoding methods include the Coordinate Format (COO), Compressed Sparse Row Format (CSR), Compressed Sparse Column Format (CSC), etc. Among them, COO is the simplest compression method, which stores the data value, row coordinate, and column coordinate respectively. CSR also requires three types of data to represent, storing the data value, column number, and row offset. The row offset indicates the position of the first non-zero element in a row among all non-zero elements. CSC is a compression method similar to CSR. CSC compresses by column and stores the data value, row number, and column offset. As Figure 8 shown, it is a schematic diagram of the results of using the COO, CSR, and CSC compression encoding methods for a two-dimensional sparse matrix.
[0081] The weight parameters in the neural network are trained offline and are static data. Therefore, the trained weight data can be pre-compressed according to a specific encoding method, but the decoding process still needs to be completed in real time in the hardware architecture. To facilitate the implementation of the weight decoding process using the hardware architecture, this project uses the Bitmap compression format to encode the sparse weights and designs a corresponding hardware decoding unit to decode the position coordinates of each non-zero weight value in the convolution kernel according to the Bitmap value before convolution calculation and output them to the processing element PE of the computing array.
[0082] As Figure 9 shown, it is a schematic diagram of the encoding and decoding process of the Bitmap compression format. The Bitmap compression format consists of three components, namely the data length, the Bitmap encoding, and the non-zero element values, where the Bitmap encoding is used to record the position information of the non-zero elements. The encoding process compresses the original weight matrix into the Bitmap format. According to the row and column sequence of the original weight matrix, each element is represented by a 1-bit binary number in turn to indicate whether it is non-zero. If the element is non-zero, it is recorded as 1, otherwise it is recorded as 0. In this way, the position of the non-zero elements can be recorded by the Bitmap. For a sparse matrix, this can avoid recording a large number of zero-valued elements, save the storage space occupied by the matrix, and reduce the memory access times during calculation. Taking the 4×4 weights in Figure 5 as an example, there are 4 non-zero elements in the original weight matrix, so the data length is recorded as 4. Taking the first row of the matrix as an example for the Bitmap part, it is (10,0,0,0) in turn. Therefore, the corresponding bitmap is (1,0,0,0) from the low bit to the high bit in turn. The remaining rows are (0,1,0,0), (0,0,1,0), and (0,0,0,1) in turn, and the values of the non-zero elements are directly recorded in sequence.
[0083] The weights in the Bitmap compression format cannot be directly output to the processing element PE for calculation. Therefore, it is also necessary to decode the compressed format data before calculation to obtain the position coordinates of each non-zero element in the original matrix. Different from the encoding process, decoding needs to be completed inside the computing architecture. Therefore, it is necessary to design a specific decoding unit to implement the weight data decoding. As Figure 10 shown, it is a schematic diagram of the structure of the sparse data decoding unit.
[0084] The sparse data decoding unit consists of a shift register unit, a control logic unit, a row-column counter unit, and a row-column coordinate cache unit. When performing decoding, first, 8-bit Bitmap data is input into the shift register unit in parallel, and then a shift operation is performed to serially output the Bitmap data from the low bit to the high bit to the control logic unit. At the same time, each time a shift is performed, the row-column counter unit counts once. The control logic unit judges the input value. If the input is 1, a write operation is performed to write the corresponding row-column coordinates into the row-column coordinate cache unit; otherwise, the next input data is read until all the current Bitmap data is decoded and the next Bitmap data is read in.
[0085] The beneficial effects of the present invention are as follows:
[0086] 1. The present invention realizes the hardware real-time decoding of sparse-coded data through the design of the sparse data decoding unit.
[0087] 2. The present invention decodes the weight parameters after sparse coding through the sparse data decoding unit to obtain non-zero weight data and its corresponding position coordinates, and then only performs convolution calculations related to non-zero weights, making full use of the network sparsity to realize efficient sparse neural network inference calculations;
[0088] 3. The system architecture of the present invention is based on the FPGA architecture, so it can be used in all scenarios where the FPGA architecture can be used, and the application range is wide. Specific Embodiment 2:
[0090] Based on the above-mentioned FPGA-based sparse neural network acceleration calculation system disclosed in the embodiments of the present invention, Figure 11 a specific FPGA-based sparse neural network acceleration calculation method applied to the FPGA-based sparse neural network acceleration calculation system is specifically disclosed.
[0091] As Figure 11 shown, Embodiment 2 of the present invention discloses an FPGA-based sparse neural network acceleration calculation method, and the method includes:
[0092] S1101. The feature map cache unit receives and caches the input data from the memory access unit DMA;
[0093] S1102. The sparse weight cache unit reads and caches the weight parameters after sparse coding from the memory access unit DMA;
[0094] S1103. The sparse data decoding unit reads the weight parameters after sparse coding from the sparse weight cache unit and decodes them to obtain non-zero weight data and its corresponding position coordinates;
[0095] S1104. The computing engine unit reads the non-zero weight data and its corresponding position coordinates from the sparse data decoding unit, and reads the input data from the feature map caching unit, and performs convolution calculation to obtain the final calculation result;
[0096] S1105. The feature map caching unit receives the final calculation result and sends the final calculation result to the memory access unit DMA.
[0097] For the further specific working processes of the feature map caching unit, sparse weight caching unit, sparse data decoding unit, and computing engine unit in the FPGA-based sparse neural network acceleration calculation method disclosed in the embodiments of the present invention, reference may be made to the corresponding content in the FPGA-based sparse neural network acceleration calculation system disclosed in the above embodiments of the present invention, and details will not be elaborated here. Specific Embodiment 3:
[0099] To facilitate understanding of the technical solution of Embodiment 1, this application is illustrated by specific examples.
[0100] As Figure 12 shown, it is a schematic structural diagram of a Zynq-based sparse neural network acceleration calculation module deploying an FPGA-based sparse neural network acceleration calculation system provided in Specific Embodiment 3 of this application.
[0101] From Figure 12 it can be seen that the sparse calculation architecture of the present invention is deployed on the FPGA side (PL side) of the Zynq device, and corresponding data paths and data interaction interfaces are designed around the calculation architecture, and the ARM processor is used for software design to control the software and hardware collaborative calculation scheduling and perform data interaction with the host computer.
[0102] As Figure 13 shown, it is a schematic structural diagram of a sparse neural network acceleration calculation test system built according to the Zynq-based sparse neural network acceleration calculation module.
[0103] Among them, the test system consists of two parts: a host computer and a ZCU102 development board. The two are connected through a gigabit Ethernet interface and can transmit remote sensing image datasets, model parameters, and calculation results. The host computer is mainly responsible for human-computer interaction and data source simulation and is the interaction medium between the user and the development board. The development board can be controlled through the host computer. The ZCU102 development board is equipped with a Zynq UltraScale+ MPSoC device based on the FPGA architecture, which also includes a quad-core ARM Cortex-A53 and a dual-core Cortex-R5F real-time processor. The FPGA and the ARM together constitute a heterogeneous computing architecture. Inside the SoC, the PS side runs software programs responsible for control and resource scheduling, and the PL side deploys a sparse streaming computing architecture responsible for the accelerated calculation of sparse neural networks. The two are interconnected through the AXI bus for data interaction.
[0104] From the above content, it can be determined that the sparse neural network acceleration computing system based on FPGA of the present invention can solve the technical problems in the prior art that the traditional FPGA computing architecture cannot process encoded data and cannot utilize the sparsity of the network, resulting in a waste of a large amount of computing resources and low computing efficiency, and realize the accelerated calculation of sparse neural networks.
[0105] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement of the above embodiments still fall within the protection scope of the technical solution of the present invention.
Claims
1. A sparse neural network acceleration computing system based on FPGA, characterized in that, The system includes: A feature map cache unit, configured to receive and cache input data from a memory access unit DMA; receive the final calculation result sent by a calculation engine unit, and send the final calculation result to the memory access unit DMA; A sparse weight cache unit, configured to read and cache the weight parameters after sparse coding from the memory access unit DMA; A sparse data decoding unit, configured to read the weight parameters after sparse coding from the sparse weight cache unit and perform decoding to obtain non-zero weight data and their corresponding position coordinates; The calculation engine unit, configured to read the non-zero weight data and their corresponding position coordinates from the sparse data decoding unit, and read the input data from the feature map cache unit, and perform convolution calculation to obtain the final calculation result.
2. The system according to claim 1, characterized in that, The calculation engine unit includes: A calculation array unit, an output cache unit, an activation unit, and a pooling unit; Wherein, the calculation array unit is composed of a plurality of processing elements PE interconnected to form a streaming structure for performing convolution calculation.
3. The system according to claim 2, wherein The output cache unit is configured to cache the intermediate calculation result of the calculation array unit, the activation calculation result of the activation unit, and the pooling calculation result of the pooling unit, and output the intermediate calculation result, the activation calculation result, and the pooling calculation result to the calculation array unit when needed.
4. The system according to claim 2, wherein The processing element PE includes: A multiply-accumulate unit, an internal cache unit, and an adder; Wherein, the adder is configured to use the partial sum output result of the previous row of processing elements PE as a partial sum input, accumulate it with the partial sum result of the current row of processing elements PE, obtain the partial sum output result of the current row of processing elements PE and output it to the next row of processing elements PE as a partial sum input.
5. The system according to claim 4, wherein The multiply-accumulate unit includes: A sub-multiplier and a sub-adder; Wherein, the sub-multiplier is configured to multiply the input data read from the feature map cache unit by the non-zero weight data to obtain a multiplication result; the sub-adder is configured to add the multiplication result to the partial sum data stored in the internal cache unit and then store it back in the internal cache unit.
6. The system according to claim 4, wherein The internal cache unit includes: An address generation unit and a sub-cache unit; Wherein, the address generation unit is configured to determine the validity of the input data based on the position coordinates corresponding to the non-zero weight data and the position coordinates in the input data, and calculate the cache address of the calculation result of the multiply-accumulate unit in the sub-cache unit; the sub-cache unit is configured to output partial sum data and partial sum results.
7. The system according to claim 1, characterized in that, The sparse data decoding unit includes: A shift register unit, a control logic unit, a row-column counter unit, and a row-column coordinate cache unit.
8. A sparse neural network acceleration calculation method based on FPGA, characterized in that, Applied to the FPGA-based sparse neural network acceleration calculation system according to claim 1, the method includes: The feature map cache unit receives and caches input data from the memory access unit DMA; The sparse weight cache unit reads and caches the weight parameters after sparse coding from the memory access unit DMA; The sparse data decoding unit reads the sparsely encoded weight parameters from the sparse weight cache unit and decodes them to obtain the non-zero weight data and their corresponding position coordinates; The computing engine unit reads the non-zero weight data and their corresponding position coordinates from the sparse data decoding unit, and reads the input data from the feature map cache unit, and performs convolution calculations to obtain the final calculation result; The feature map cache unit receives the final calculation result and sends the final calculation result to the memory access unit DMA.