Packet high-parallel high-throughput CNN hardware acceleration architecture and acceleration method

By designing a hardware-accelerated architecture for high-parallelism and high-throughput CNNs, optimizing convolution computation and parallel processing, the problem of complex object detection with limited resources at the edge is solved, achieving high-throughput CNN object detection and meeting the requirements of high-time-efficiency tasks.

CN120952075APending Publication Date: 2025-11-14BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511050451.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

With limited edge resources on satellite platforms, existing lightweight CNN algorithms struggle to perform complex target detection tasks, exhibiting issues of insufficient accuracy and limited processing power.

Method used

Design a hardware-accelerated architecture for grouped high-parallelism and high-throughput CNNs, including an input unit, a grouped multi-level convolution parallel processing engine, a hidden layer high-parallel processing module, and an output unit. Through caching, parallel processing, and further operations, the convolution computation is optimized to achieve parallel processing of large-size multi-channel data streams.

Benefits of technology

With limited edge computing power, a high throughput of complex CNN object detection algorithm was achieved, meeting the requirements of time-sensitive tasks, with high resource utilization and a throughput of 770 GOPS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952075A_ABST
    Figure CN120952075A_ABST
Patent Text Reader

Abstract

The invention discloses a grouped high-parallel high-throughput CNN (Convolutional Neural Network) hardware acceleration architecture and acceleration method, and belongs to the technical field of hardware acceleration. The architecture comprises an input unit for caching an input data stream, a packet multi-level convolution parallel processing engine for realizing parallel processing of a large-size multi-channel data stream, and a hidden layer high parallel processing module for further processing the data stream after parallel processing, the input unit is used for caching data streams from the hidden layer high-parallel processing module and outputting the data streams, the output unit is used for caching the data streams from the hidden layer high-parallel processing module and outputting the data streams, and the configuration unit is used for regulating and controlling the input unit, the grouped multistage convolution parallel processing engine, the hidden layer high-parallel processing module and the output unit. Large-size multi-channel data parallel processing is realized, and a complex CNN target detection algorithm can be realized under limited edge end computing power through data packet processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hardware acceleration technology, and in particular to a hardware acceleration architecture and method for grouped high-parallelism, high-throughput CNNs. Background Technology

[0002] In recent years, with the rapid development of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) have shown great application potential in time-sensitive tasks such as emergency response and disaster relief, and military reconnaissance. However, due to the strict limitations of satellite platforms on processor size, weight, and power consumption, and the limited computing and storage resources at the satellite edge, it is difficult to deploy computationally intensive and complex CNN object detection algorithms.

[0003] Currently, the main implementations for edge computing are lightweight CNN algorithms, such as YOLOv2-Tiny, YOLOv3-Tiny, YOLOv4-Tiny, and MobileNetv2, all with fewer than 100 GOPs. While these algorithms meet the basic operational requirements in resource-constrained environments, they still have significant shortcomings in terms of accuracy, robustness, and handling more complex tasks. To address these issues, there is an urgent need to construct novel CNN acceleration architectures to overcome the dual constraints of computation and storage under limited hardware resources, enabling the implementation of more complex CNN detection algorithms. Based on this, this invention proposes a grouped, high-parallelism, high-throughput CNN hardware acceleration architecture and acceleration method. Summary of the Invention

[0004] The purpose of this invention is to provide a hardware acceleration architecture and method for high-parallelism and high-throughput grouped CNNs, which can realize complex CNN object detection algorithms with limited edge computing power and achieve higher throughput, thus meeting the requirements of high-time-efficiency tasks.

[0005] To achieve the above objectives, the present invention provides a hardware acceleration architecture and method for grouped high-parallelism and high-throughput CNNs, including an input unit for buffering the input data stream, a grouped multi-level convolutional parallel processing engine for implementing parallel processing of large-size multi-channel data streams, a hidden layer high-parallel processing module for further processing of the data stream after parallel processing, an output unit for buffering and outputting the data stream from the hidden layer high-parallel processing module, and a configuration unit for adjusting the input unit, the grouped multi-level convolutional parallel processing engine, the hidden layer high-parallel processing module, and the output unit.

[0006] Preferably, the input unit includes an input weight cache unit for caching the weight parameter data of all neurons required for a single parallel computation and an input feature map cache unit for caching feature maps. Both the input weight cache unit and the input feature map cache unit include several RAM cache units and a multiplexer.

[0007] Preferably, the grouped multi-level convolutional parallel processing engine includes a multi-level parallel processing engine array that receives and processes data from input weight cache units and input feature map cache units in parallel, and an adder group that adds the data processed by the multi-level parallel processing engine array.

[0008] Preferably, the multi-level parallel processing engine array includes the same number of convolutional layers as the input feature maps, each convolutional layer consists of a convolutional processing engine consisting of half the number of output feature maps and a weight cache, and the addition group includes several adders.

[0009] Preferably, the hidden layer high-parallel processing module includes a number of hidden processing groups equal to the number of output feature maps.

[0010] This invention also provides a hardware acceleration method for grouped high-parallelism, high-throughput CNNs, employing the aforementioned hardware acceleration architecture for grouped high-parallelism, high-throughput CNNs, and comprising the following steps:

[0011] S1. Use the input unit to cache the weights and feature maps of the input data stream, and send the data stream and cached content to the grouped multi-level convolution parallel processing engine;

[0012] S2. Optimize standard convolution in the grouped multi-level convolution parallel processing engine;

[0013] S3. After optimization, a grouped multi-level convolutional parallel processing engine is used to perform convolutional layer calculations on the data stream to obtain N. out One calculation result;

[0014] S4, N out Each calculation result is input into the hidden layer height parallel processing module for further processing, and the result is then input into the output unit.

[0015] S5, The output unit simultaneously uses N to process the received result. out The RAM is used for caching, and the final result is output after caching is complete.

[0016] Preferably, the process of S1 is as follows:

[0017] S11. Use the input feature map caching unit to cache N. in Pixel data from ×K×K input feature maps;

[0018] S12, using N in One RAM cache simultaneously feeds the pixel data of the input feature map to the multi-level convolution parallel processing engine;

[0019] S13. Use the input weight cache unit to cache N in the data stream. in ×N outThe weight parameter data for ×K×K neurons is obtained through the following process:

[0020] First send N out The weights and convolution parameters corresponding to the first convolutional layer with ×K×K output channels are then sent to N. out The weights and convolution parameters corresponding to the second convolutional layer with ×K×K output channels, up to N. in All the weights and convolution parameters required for each convolutional layer have been sent.

[0021] S14. After obtaining the weight parameter data from the data stream, send it to the multi-level convolution parallel processing engine, and simultaneously send N... out The parameters are batch normalized in each hidden processing group.

[0022] Preferably, the process of optimizing the standard convolution in S2 is as follows:

[0023] S21. Perform parallel computation of the K×K convolution kernel in the first and second layers of the convolutional layer.

[0024] S22. After cyclic parallel computation, the input channels are grouped and traversed in the third layer of the convolutional layer, and the output channels are grouped and traversed in the fourth layer. The parallel computation unit for the input channels is N. in There are N parallel computing units for the output channels. out The number of items is calculated in groups as follows:

[0025] G in =C in ÷N in ;

[0026] G out =C out ÷N out ;

[0027] N dsp ≥N in ×N out ×K×K×η;

[0028] Among them G in and G out C represents the number of groups for the input and output channels, respectively. in and C out These represent the number of input channels and output channels, respectively, N. dsp This indicates the number of DSP48Es in the FPGA, where K×K represents the computing units for the input and output channels, and η represents the utilization efficiency of the DSP48E.

[0029] S23. After the group traversal is completed, the fifth and sixth convolutional layers perform pipelined circulation on the row and column data of the input feature map;

[0030] S24. After the pipelined loop, the convolutional layer performs input channel grouping calculation loop in the seventh layer and output channel traversal loop in the eighth layer.

[0031] Preferably, the process of S3 is as follows:

[0032] S31, Each convolutional processing engine in the multi-level parallel processing engine array receives its required weighted convolutional parameters N from S13. out ×K×K, and cache it in the weighted cache;

[0033] S32, N in K×K feature map pixel data from each input channel are simultaneously input to the corresponding N... in In each convolutional processing engine, the parameters are multiplied and added together with the corresponding weight convolutional parameters;

[0034] S33. Add the results of the multiplication and addition operations of each convolution processing engine and output the sum to obtain N. out The calculation results, i.e., N out Feature map pixel data for each output channel.

[0035] Preferably, the process of further processing in S4 and inputting the result into the output unit is as follows:

[0036] S41, for N out Sum the feature map pixel data of each output channel;

[0037] S42. After completing the summation operation, perform dequantization and normalization operations on the summation result to convert the fixed-point number into a single-precision floating-point number. The process is as follows:

[0038]

[0039] y = x * BN_gamma + BN_beta;

[0040] Where y represents the floating-point number after normalization and dequantization, ε is a fixed value with a default value of 1e-5, γ represents the adjusted variance with an initial value of 1, β represents the adjusted mean with an initial value of 0, x represents the input fixed-point feature value, Var represents the variance, E represents the mean, and BN_gamma and BN_beta are parameters obtained from batch normalization in S14.

[0041] S43. After normalization, use the leaky ReLU activation function to determine the result. Output the data that is greater than 0 directly, and output the data that is less than or equal to 0 after multiplying it by a fixed value. Then, quantize the output of the result.

[0042] S44. Pool the quantized data, input 2×2 data points, and output the largest data point.

[0043] Therefore, this invention adopts a grouped high-parallelism, high-throughput CNN hardware acceleration architecture and acceleration method with the above-mentioned structure, designs a grouped multi-level parallel processing engine array to realize large-size multi-channel data parallel processing, and proposes a high-parallelism, high-throughput CNN hardware acceleration architecture and corresponding method. By processing data in groups, complex CNN object detection algorithms can be implemented with limited edge computing power.

[0044] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0045] Figure 1 This is a flowchart of the hardware acceleration architecture and method for a high-parallelism, high-throughput CNN grouping method according to the present invention.

[0046] Figure 2 This is a structural diagram of the convolutional processing engine of a hardware acceleration architecture and acceleration method for high-parallelism and high-throughput CNN in this invention.

[0047] Figure 3 This is a diagram of a multi-level parallel processing engine array structure for a hardware acceleration architecture and acceleration method for high-parallelism, high-throughput CNN grouping according to the present invention. Detailed Implementation

[0048] Example

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0050] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0051] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0052] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed when in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0053] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," and "connect" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0054] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0055] like Figures 1-3 As shown, the present invention provides a hardware acceleration architecture for grouped high-parallelism and high-throughput CNNs, including an input unit for caching the input data stream, a grouped multi-level convolutional parallel processing engine for implementing parallel processing of large-size multi-channel data streams, a hidden layer high-parallel processing module for further processing of the data stream after parallel processing, an output unit for caching and outputting the data stream from the hidden layer high-parallel processing module, and a configuration unit for adjusting the input unit, the grouped multi-level convolutional parallel processing engine, the hidden layer high-parallel processing module, and the output unit.

[0056] The input unit includes an input weight cache unit that caches the weight parameter data of all neurons required for a single parallel computation and an input feature map cache unit that caches the feature maps. Both the input weight cache unit and the input feature map cache unit include several RAM cache units (FIFO) and a multiplexer.

[0057] The grouped multi-level convolutional parallel processing engine includes a multi-level parallel processing engine array that receives and processes data from input weight cache units and input feature map cache units in parallel, and an adder group that adds the data processed by the multi-level parallel processing engine array. The multi-level parallel processing engine array can support N in ×K×K input feature map pixels and N in ×N out ×K×K weights are processed simultaneously in parallel, and the output is N. out There are N calculation results.in N represents the number of input feature maps. out This indicates the number of output feature maps.

[0058] The multi-level parallel processing engine array includes convolutional layers with the same number of input feature maps. Each convolutional layer consists of a convolutional processing engine equal to half the number of output feature maps and a weight buffer. The addition group includes several adders. Each convolutional processing engine includes 9 multipliers and 4 adders. The multipliers are implemented using a DSP48E, and the adders are implemented using FPGA logic resources. Figure 2 As can be seen, a single convolutional processing engine (CU) can support 9 feature values ​​and 18 weight multiplication-accumulation calculations. Therefore, N features can be processed in parallel within a single CU group. out The calculation of multiplying and adding 3×3 weights and 3×3 eigenvalues ​​only requires N. out / 2 CUs, saving 50% of CUs.

[0059] The hidden layer height parallel processing module includes the same number of hidden processing groups as the output feature maps.

[0060] This invention also provides a hardware acceleration method for grouped high-parallelism, high-throughput CNNs, employing a grouped high-parallelism, high-throughput CNN hardware acceleration architecture, comprising the following steps:

[0061] S1. Use the input unit to cache the weights and feature maps of the input data stream, and send the data stream and cached content to the grouped multi-level convolution parallel processing engine;

[0062] S11. Use the input feature map caching unit to cache N. in Pixel data from ×K×K input feature maps;

[0063] S12, using N in One RAM cache simultaneously feeds the pixel data of the input feature map to the multi-level convolution parallel processing engine;

[0064] S13. Use the input weight cache unit to cache N in the data stream. in ×N out The weight parameter data for ×K×K neurons is obtained through the following process:

[0065] First send N out The weights and convolution parameters corresponding to the first convolutional layer with ×K×K output channels are then sent to N. out The weights and convolution parameters corresponding to the second convolutional layer with ×K×K output channels, up to N. in All the weights and convolution parameters required for each convolutional layer have been sent.

[0066] S14. After obtaining the weight parameter data from the data stream, send it to the multi-level convolution parallel processing engine, and simultaneously send N... out The parameters for batch normalization in each hidden processing group include BN_gamma and BN_beta.

[0067] S2. Optimize the standard convolution in the grouped multi-level convolution parallel processing engine because the limited computing resources of FPGA (such as the Xilinx V7690t's DSP48E, which has only 3600) cannot process large-size, multi-channel data in parallel.

[0068] S21. Perform parallel computation of the K×K convolution kernel in the first and second layers of the convolutional layer.

[0069] S22. After cyclic parallel computation, the input channels are grouped and traversed in the third layer of the convolutional layer, and the output channels are grouped and traversed in the fourth layer. The parallel computation unit for the input channels is N. in There are N parallel computing units for the output channels. out The number of items is calculated in groups as follows:

[0070] G in =C in ÷N in ;

[0071] G out =C out ÷N out ;

[0072] N dsp ≥N in ×N out ×K×K×η;

[0073] Among them G in and G out C represents the number of groups for the input and output channels, respectively. in and C out These represent the number of input channels and output channels, respectively, N. dsp This indicates the number of DSP48Es in the FPGA, where K×K represents the computing units for the input and output channels, and η represents the utilization efficiency of the DSP48E.

[0074] S23. After the group traversal is completed, the fifth and sixth convolutional layers perform pipelined circulation on the row and column data of the input feature map;

[0075] S24. After the pipelined loop, the convolutional layer performs input channel grouping calculation loop in the seventh layer and output channel traversal loop in the eighth layer.

[0076] S3. After optimization, a grouped multi-level convolutional parallel processing engine is used to perform convolutional layer calculations on the data stream to obtain N. out One calculation result;

[0077] S31, Each convolutional processing engine in the multi-level parallel processing engine array receives its required weighted convolutional parameters N from S13. out ×K×K, and cache it in the weighted cache;

[0078] S32, N in K×K feature map pixel data from each input channel are simultaneously input to the corresponding N... in In each convolutional processing engine, the parameters are multiplied and added together with the corresponding weight convolutional parameters;

[0079] S33. Add the results of the multiplication and addition operations of each convolution processing engine and output the sum to obtain N. out The calculation results, i.e., N out Feature map pixel data for each output channel.

[0080] S4, N out Each calculation result is input into the hidden layer height parallel processing module for further processing, and the result is then input into the output unit.

[0081] S41, for N out Sum the feature map pixel data of each output channel;

[0082] S42. After completing the summation operation, perform dequantization and normalization operations on the summation result to convert the fixed-point number into a single-precision floating-point number. The process is as follows:

[0083]

[0084] y = x * BN_gamma + BN_beta;

[0085] Where y represents the floating-point number after normalization and dequantization, ε is a fixed value with a default value of 1e-5, γ represents the adjusted variance with an initial value of 1, β represents the adjusted mean with an initial value of 0, x represents the input fixed-point feature value, Var represents the variance, E represents the mean, and BN_gamma and BN_beta are parameters obtained from batch normalization in S14.

[0086] S43. After normalization, the leaky ReLU activation function is used to determine the result. Data greater than 0 is output directly, and data less than or equal to 0 is multiplied by a fixed value and then output. In this embodiment, the fixed value is set to 0.1. Then, the obtained result is quantized.

[0087] S44. Pool the quantized data, input 2×2 data points, and output the largest data point.

[0088] S5, The output unit simultaneously uses N to process the received result. out The RAM is used for caching, and the final result is output after caching is complete.

[0089] Using the above architecture, a YOLO detection algorithm with 380 GOPs was implemented in a Xilinx XQ7VX690T FPGA. Following the method mentioned above, the final resource utilization is shown in Table 1, and the final throughput reached 770 GOPs.

[0090] Table 1 Resource Utilization Rate

[0091]

[0092] Therefore, this invention adopts a grouped high-parallelism, high-throughput CNN hardware acceleration architecture and acceleration method with the above-mentioned structure, designs a grouped multi-level parallel processing engine array to realize large-size multi-channel data parallel processing, and proposes a high-parallelism, high-throughput CNN hardware acceleration architecture and corresponding method. By processing data in groups, complex CNN object detection algorithms can be implemented with limited edge computing power.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A hardware-accelerated architecture for grouped high-parallelism, high-throughput CNNs, characterized in that: It includes an input unit for buffering the input data stream, a grouped multi-level convolution parallel processing engine for implementing parallel processing of large-size multi-channel data streams, a hidden layer height parallel processing module for further processing of the data stream after parallel processing, an output unit for buffering and outputting the data stream from the hidden layer height parallel processing module, and a configuration unit for controlling the input unit, the grouped multi-level convolution parallel processing engine, the hidden layer height parallel processing module, and the output unit.

2. The hardware acceleration architecture for a high-parallelism, high-throughput grouped CNN according to claim 1, characterized in that: The input unit includes an input weight cache unit that caches the weight parameter data of all neurons required for a single parallel computation and an input feature map cache unit that caches the feature maps. Both the input weight cache unit and the input feature map cache unit include several RAM cache units and a multiplexer.

3. The hardware acceleration architecture for a high-parallelism, high-throughput grouped CNN according to claim 2, characterized in that: The grouped multi-level convolutional parallel processing engine includes a multi-level parallel processing engine array that receives and processes data from input weight cache units and input feature map cache units in parallel, and an adder group that adds the data processed by the multi-level parallel processing engine array.

4. The hardware acceleration architecture for a high-parallelism, high-throughput grouped CNN according to claim 3, characterized in that: The multi-level parallel processing engine array includes the same number of convolutional layers as the input feature maps. Each convolutional layer consists of a convolutional processing engine consisting of half the number of output feature maps and a weight buffer. The addition group includes several adders.

5. The hardware acceleration architecture for a high-parallelism, high-throughput grouped CNN according to claim 4, characterized in that: The hidden layer height parallel processing module includes the same number of hidden processing groups as the output feature maps.

6. A hardware acceleration method for high-parallelism, high-throughput grouped CNNs, characterized in that: The hardware acceleration architecture for a high-parallelism, high-throughput grouped CNN as described in any one of claims 1-5 includes the following steps: S1. Use the input unit to cache the weights and feature maps of the input data stream, and send the data stream and cached content to the grouped multi-level convolution parallel processing engine; S2. Optimize standard convolution in the grouped multi-level convolution parallel processing engine; S3. After optimization, a grouped multi-level convolutional parallel processing engine is used to perform convolutional layer calculations on the data stream to obtain N. out One calculation result; S4, N out Each calculation result is input into the hidden layer height parallel processing module for further processing, and the result is then input into the output unit. S5, The output unit simultaneously uses N to process the received result. out The RAM is used for caching, and the final result is output after caching is complete.

7. A hardware acceleration method for high-parallelism, high-throughput CNNs according to claim 6, characterized in that, The process of S1 is as follows: S11. Use the input feature map caching unit to cache N. in Pixel data from ×K×K input feature maps; S12, using N in One RAM cache simultaneously feeds the pixel data of the input feature map to the multi-level convolution parallel processing engine; S13. Use the input weight cache unit to cache N in the data stream. in ×N out The weight parameter data for ×K×K neurons is obtained through the following process: First send N out The weights and convolution parameters corresponding to the first convolutional layer with ×K×K output channels are then sent to N. out The weights and convolution parameters corresponding to the second convolutional layer with ×K×K output channels, up to N. in All the weights and convolution parameters required for each convolutional layer have been sent. S14. After obtaining the weight parameter data from the data stream, send it to the multi-level convolution parallel processing engine, and simultaneously send N... out The parameters are batch normalized in each hidden processing group.

8. A hardware acceleration method for high-parallelism, high-throughput CNNs according to claim 7, characterized in that, The process of optimizing the standard convolution in S2 is as follows: S21. Perform parallel computation of the K×K convolution kernel in the first and second layers of the convolutional layer. S22. After cyclic parallel computation, the input channels are grouped and traversed in the third layer of the convolutional layer, and the output channels are grouped and traversed in the fourth layer. The parallel computation unit for the input channels is N. in There are N parallel computing units for the output channels. out The number of items is calculated in groups as follows: G in =C in ÷N in ; G out =C out ÷N out ; N dsp ≥N in ×N out ×K×K×η; Among them G in and G out C represents the number of groups for the input and output channels, respectively. in and C out These represent the number of input channels and output channels, respectively, N. dsp This indicates the number of DSP48Es in the FPGA, where K×K represents the computing units for the input and output channels, and η represents the utilization efficiency of the DSP48E. S23. After the group traversal is completed, the fifth and sixth convolutional layers perform pipelined circulation on the row and column data of the input feature map; S24. After the pipelined loop, the convolutional layer performs input channel grouping calculation loop in the seventh layer and output channel traversal loop in the eighth layer.

9. A hardware acceleration method for high-parallelism, high-throughput CNNs according to claim 8, characterized in that, The process of S3 is as follows: S31, Each convolutional processing engine in the multi-level parallel processing engine array receives its required weighted convolutional parameters N from S13. out ×K×K, and cache it in the weighted cache; S32, N in K×K feature map pixel data from each input channel are simultaneously input to the corresponding N... in In each convolutional processing engine, the parameters are multiplied and added together with the corresponding weight convolutional parameters; S33. Add the results of the multiplication and addition operations of each convolution processing engine and output the sum to obtain N. out The calculation results, i.e., N out Feature map pixel data for each output channel.

10. A hardware acceleration method for high-parallelism, high-throughput grouped CNNs according to claim 9, characterized in that, The process of further processing in S4 and inputting the results into the output unit is as follows: S41, for N out Sum the feature map pixel data of each output channel; S42. After completing the summation operation, perform dequantization and normalization operations on the summation result to convert the fixed-point number into a single-precision floating-point number. The process is as follows: y = x * BN_gamma + BN_beta; Where y represents the floating-point number after normalization and dequantization, ε is a fixed value with a default value of 1e-5, γ represents the adjusted variance with an initial value of 1, β represents the adjusted mean with an initial value of 0, x represents the input fixed-point feature value, Var represents the variance, E represents the mean, and BN-gamma and BN_beta are parameters obtained from batch normalization in S14. S43. After normalization, use the leaky ReLU activation function to determine the result. Output the data that is greater than 0 directly, and output the data that is less than or equal to 0 after multiplying it by a fixed value. Then, quantize the output of the result. S44. Pool the quantized data, input 2×2 data points, and output the largest data point.