Method for computing maximum pooling layer hardware of convolutional neural network

By introducing a 32-bit wide and 32-deep FIFO cache structure and alternating odd-even row processing in convolutional neural networks, the problems of computational delay and low resource utilization in the maximum pooling layer are solved, and efficient pooling operations are achieved, which is suitable for high-performance and low-power neural network hardware accelerators.

CN120704638APending Publication Date: 2025-09-26GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510782641.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The hardware implementation of the maximum pooling layer in existing convolutional neural networks has problems such as large computing delay, low resource utilization, and low data processing efficiency. In particular, it is difficult to meet the requirements of efficient parallelization and pipelining in edge computing devices.

Method used

A 32-bit wide and 32-bit deep FIFO cache structure is adopted, combined with alternating processing of odd and even rows and data beating operations to implement time-sharing pipeline processing of the maximum pooling layer. The maximum pooling operation is completed by writing the even-row data into the FIFO through local comparison, and comparing the odd-row data with the FIFO intermediate results.

Benefits of technology

The pooling results are output within two clock cycles, which improves computing speed and efficiency and reduces resource consumption. It is suitable for the design of high-performance and low-power neural network hardware accelerators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704638A_ABST
    Figure CN120704638A_ABST
Patent Text Reader

Abstract

The invention discloses a method for hardware calculation of a maximum pooling layer of a convolutional neural network. The structure of the method comprises a pooling module, an FIFO (First In First Out) cache unit with the width of 32 bits and the depth of 32, a comparator array and a time sequence control circuit. The method is based on a 2 * 2 pooling template, and adopts a row odd-even separation strategy: in an even row, two adjacent activation values are compared and written into an FIFO (First In First Out) after one beat is made; and in odd-numbered lines, reading the maximum value corresponding to the previous line from the FIFO, and comparing the maximum value with the current value to obtain a final pooling result. Through the structural design, the maximum pooling operation is completed in two clock periods, the calculation process is high in speed and continuous, and intermediate storage and data migration processes are remarkably reduced. Compared with the traditional pooling implementation by using a plurality of comparators or a multi-cycle traversal mode, the method has obvious advantages in the aspects of delay control and hardware resource occupation, and is more suitable for the design requirements of low-power consumption and high-performance neural network chips.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence chip design and neural network accelerated computing, and specifically to a hardware calculation method for the pooling layer of a convolutional neural network (CNN). This method is a dedicated hardware optimization technology for the deep learning inference process and is suitable for the design of neural network accelerators in integrated circuits (ASICs), field programmable gate arrays (FPGAs), and high-performance embedded computing platforms. Background Art

[0002] With the rapid development of artificial intelligence (AI), convolutional neural networks (CNNs) have been widely used in a variety of fields, including image recognition, object detection, and natural language processing. Pooling layers, a crucial component of CNN architectures, are used to reduce the spatial dimensionality of feature maps, minimize the number of parameters and computational complexity, and enhance the network's robustness to local transformations such as translation, rotation, and scaling. Max pooling is a commonly used pooling method. By selecting the maximum value within the pooling window, it retains the most significant features in the feature map, thereby enhancing the model's ability to extract local edge and texture information.

[0003] In actual deployments, especially in edge computing devices or embedded AI chips, the efficiency of pooling operations has a significant impact on the reasoning speed of the entire neural network. Compared with convolutional layers and fully connected layers, although pooling layers do not involve a large number of multiplication and addition operations, they involve window sliding, boundary processing, and conditional judgment (such as comparison operations). If not optimized, they will also become a hardware computing bottleneck. In current mainstream hardware implementations, pooling layers often rely on general-purpose processing units (such as CPUs or DSPs) or simple logic modules. They are not deeply customized for the characteristics of maximum pooling operations, and have problems such as low resource utilization, large processing delays, and difficulty in scalability.

[0004] Furthermore, in the design of low-power, high-throughput neural network accelerators, the efficient implementation of the pooling layer is crucial to the overall computing architecture. On the one hand, the number of accesses to external memory should be minimized to reduce data handling overhead; on the other hand, the output data flow characteristics of the convolutional layer should be combined to design adaptive data paths and control mechanisms to achieve parallelization and pipelining of pooling operations, thereby improving overall execution efficiency. Therefore, there is an urgent need for a hardware calculation method specifically for the maximum pooling layer that can balance computational efficiency, resource consumption, and flexibility. This method should maximize the use of the parallel computing capabilities of the hardware while maintaining the correctness of the pooling results. By means of local registers and data flow optimization, the computing performance of the maximum pooling layer can be improved to meet the deployment requirements of high-performance neural networks on edge and embedded platforms.

[0005] To address these issues, we propose an efficient hardware computation method for max pooling operations. This method combines the 2×2 window and stride-size configuration commonly used in pooling layers in convolutional neural networks. By setting up a 32-bit wide and 32-bit deep FIFO buffer structure, and combining it with odd-even row alternation and data paging, we achieve fast comparison and writing of maximum values. Summary of the Invention

[0006] This paper addresses the shortcomings of existing convolutional neural network hardware implementations of maximum pooling operations, such as large computational delays, low resource utilization, and low data processing efficiency. This paper proposes a maximum pooling hardware calculation method based on a FIFO cache structure. This method implements a time-sharing pipeline processing of the maximum pooling operation within a 2×2 pooling window with a step size of 2, by locally comparing even rows of data and writing them into the FIFO, and then comparing and reading odd rows of data with the FIFO intermediate results. This method can output the pooling result in just two clock cycles, effectively improving the speed and efficiency of pooling operations and reducing resource consumption. It is suitable for the design of high-performance, low-power neural network hardware accelerators.

[0007] The technical solution adopted in the present invention is:

[0008] A method for hardware calculation of the maximum pooling layer of a convolutional neural network, a method for hardware calculation of the maximum pooling layer of a convolutional neural network, the structure of which includes: an input data buffer module, a maximum value comparison module, a FIFO cache module for temporarily storing intermediate comparison results, and an output result control module.

[0009] In the above solution, the input data buffer module is used to receive and cache the activation output after the convolution calculation from the previous layer.

[0010] In the above solution, the maximum value comparison module includes a beat register mechanism for performing time sequence alignment on adjacent data and performing a maximum value comparison operation.

[0011] In the above solution, the FIFO buffer module is 32 bits wide and 32 bits deep, and is used to store the middle maximum value after comparing adjacent data in even rows.

[0012] In the above solution, the output result control module is used to read the maximum value of the corresponding position from the FIFO when an odd row arrives and compare it with the current input, and output the final maximum pooling result.

[0013] Compared with the prior art, the present invention has the following advantages:

[0014] (1) By introducing a 32-bit wide and 32-bit deep FIFO cache mechanism, the processing bottleneck caused by insufficient intermediate result cache in traditional pooling operations is improved, and the data processing efficiency is significantly improved.

[0015] (2) By adopting the even / odd row alternating processing strategy, the control complexity caused by simultaneous access to multiple rows of data in the traditional maximum pooling operation is improved, and a simpler control logic is achieved.

[0016] (3) By combining the data beat operation with the adjacent data comparison mechanism, the timing alignment problem in the maximum value calculation is improved, ensuring the synchronization and accuracy of data comparison.

[0017] (4) By decoupling the processing of even and odd rows and combining them with FIFO reading, the reading efficiency of the internal values ​​of the pooling template is improved, avoiding repeated calculations and redundant accesses.

[0018] (5) By constructing a pipeline structure that only requires two clock cycles to obtain the pooling result, the defect of large delay in the traditional method is improved, and a high-speed, low-power pooling computing module design is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a diagram of the overall architecture of the maximum pooling layer hardware calculation in an embodiment of the present invention.

[0020] Figure 2 This is a flowchart of the hardware calculation of the maximum pooling layer according to an embodiment of the present invention.

[0021] Figure 3 This is a timing diagram for writing the maximum value of the even-numbered rows of the convolution result into the FIFO in an embodiment of the present invention.

[0022] Figure 4 This is a timing diagram of reading the maximum value from the FIFO for odd-numbered rows of the convolution result in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific examples.

[0024] like Figure 1 As shown in the figure, the overall architecture of the hardware calculation of the maximum pooling layer requires 3 comparators, 1 FIFO, and 1 input data register.

[0025] like Figure 2As shown in the figure, the hardware calculation process of the maximum pooling layer is as follows: the convolution layer outputs the activation layer (the activation layer is omitted in the figure because it does not affect the amount of data) and sends the calculation result to the input data register. After the data register sends the data to the comparator, the maximum value of the two adjacent data in the even-numbered rows of the convolution layer output after the activation layer is obtained, and the maximum value result is sent to the FIFO. In the odd-numbered rows of the convolution layer output after the activation layer, the maximum values ​​of the two adjacent data are compared through two comparators, and the maximum value of the odd-numbered rows stored in the FIFO is read out at the same time for comparison to obtain the final pooling result.

[0026] like Figure 3 The figure below shows the timing diagram for writing the maximum value of even rows into the FIFO. In the even rows of the convolutional layer's output after the activation layer, two adjacent data points undergo a "beat" operation to ensure data timing alignment. Specifically, the data from the current clock cycle is compared with the data from the previous clock cycle, and the larger value is selected. After one clock cycle, the maximum value is generated and written to a 32-bit deep and 32-bit wide FIFO buffer. This way, the FIFO stores the maximum value of each pair of adjacent data points in the even rows, providing data support for subsequent pooling calculations on the odd rows.

[0027] like Figure 4 The figure below shows the timing diagram for reading the maximum value from the FIFO for odd-numbered rows. When the odd-numbered rows of data from the convolutional layer, after passing through the activation layer, enter the pooling layer, the system simultaneously reads the corresponding maximum value from the previous row from the FIFO. This read FIFO data is compared with the data at the corresponding position in the current odd-numbered row, and the maximum value is selected to obtain the final maximum pooling result for the pooling template area. This FIFO reading ensures the effective coordination of the previous row of data with the current data, enabling fast calculation of the pooling operation.

Claims

1. A method for hardware calculation of the maximum pooling layer of a convolutional neural network, characterized by: This method is applied to a two-dimensional maximum pooling operation with a pooling template of 2×2 and a step size of 2. The activation results output by the previous layer of the convolutional neural network are input row by row and divided into even rows and odd rows according to the parity of the row numbers. In the even rows, two adjacent activation values ​​are compared after a one-beat delay, and the maximum value is taken and written into the FIFO cache module. In the odd rows, the current activation value is compared with the maximum value of the corresponding position in the FIFO to obtain the final maximum pooling result. The maximum value is uniformly clocked by the control logic module, and a pooling operation is completed within two clock cycles.

2. The method according to claim 1, characterized in that The FIFO buffer module is a dual-port FIFO with a width of 32 bits and a depth of 32 bits, and is used to store the maximum value of every two adjacent data in even-numbered rows for use by odd-numbered rows.

3. The method according to claim 1, characterized in that The "beat" operation is implemented by a first-level register, which is used to delay the current cycle data by one clock cycle so as to compare the maximum value with the previous clock cycle data.

4. The method according to claim 1, wherein The maximum value comparison unit includes two input terminals and one output terminal, and is used to compare two activation values ​​or one activation value and a FIFO value, and output a maximum value.

5. The method according to claim 1, wherein The control logic module is responsible for controlling the timing synchronization of the write FIFO, read FIFO and maximum value comparison operations to ensure that each maximum pooling operation is completed within two clock cycles.

6. The method according to claim 1, wherein The pooling result output interface outputs the maximum comparison result of the odd rows as the final maximum pooling value of the 2×2 pooling window.