Max pooling hardware capable of learning and inference acceleration
The maximum value pooling hardware in CNN accelerators addresses inefficiencies in existing systems by enabling parallel operations and supporting multiple CNN models through a matrix array and data comparator design, resulting in improved real-time performance.
Patent Information
- Application Number
- PCT/KR2023/017211
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-01
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-08
AI Technical Summary
Existing maximum pooling hardware in CNN accelerators is inefficient for parallel operations and is fixed with hyperparameters, limiting its ability to support various CNN models without hardware changes, and it struggles with real-time performance improvements.
The proposed maximum value pooling hardware includes a matrix array for storing data in a window size, a purity module with data comparators for identifying maximum values, and line buffers for separating input feature maps, enabling parallel operations and supporting various CNN models without hardware modifications.
This solution improves the maximum value pooling operation rate through parallel processing, supports various CNN models without altering the hardware, and enhances real-time performance by optimizing the learning process in resource-constrained embedded systems.
Smart Images

Figure KR2023017211_08052025_PF_FP_ABST
Abstract
Description
Max pooling hardware capable of accelerating learning and inference
[0001] The present invention relates to a deep learning accelerator, and more particularly, to a maximum pooling hardware structure applicable to a CNN (Convolutional Neural Network) accelerator.
[0002] Conventional CNN hardware accelerators for small embedded systems are designed to be optimized for specific purposes due to limited hardware design conditions such as power and area.
[0003] Optimized hardware designs can be highly efficient when used for their intended purpose, but even a slight change in purpose can result in significant performance degradation or the hardware cannot be reused, requiring a new design.
[0004] Existing max-pooling hardware performs max-pooling serially according to the output of the convolution layer's operation. When the window and stride sizes are set to be the same, it shows high efficiency, but when they are not, there is a problem that the number of times data is read and written to memory increases.
[0005] Additionally, the existing max pooling has a problem in that it cannot run applications using different CNN models because the hyperparameters are fixed.
[0006] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide maximum pooling hardware that improves the maximum pooling operation speed through parallel operation in a CNN accelerator, supports various CNN models without hardware change, and enables performance improvement through real-time learning in the field.
[0007] According to one embodiment of the present invention for achieving the above object, a maximum pooling hardware comprises: a matrix array storing data of a feature map in a window size during an inference process; a forward propagation module including data comparators arranged in a tree structure to sequentially compare two data stored in the matrix array and output a larger value while outputting the maximum value at the final stage;
[0008] The maximum pooling hardware according to the present invention may further include line buffers that store input feature maps by dividing them by row and transfer the stored data to a matrix array in units of window size.
[0009] During the inference process, if the window is smaller than the matrix array, the lowest value can be stored in the index of the matrix array that falls outside the window.
[0010] In the inference process, before the data of the feature map is stored, the lowest value can be stored as the initial value in all indices of the matrix array.
[0011] Comparators can output both the maximum value and the coordinates within the window of the maximum value.
[0012] The matrix array may further include a backpropagation module that stores data of a gradient map and coordinates within a window as a window size during the learning process, and the max pooling hardware may further include a comparator that outputs data whose coordinate within the window is the coordinate of the maximum value, and an adder that adds and outputs data output from the comparators.
[0013] During the learning process, if the window is smaller than the matrix array, the indices of the matrix array that fall outside the window may be stored as 0.
[0014] During the learning process, all indices of the matrix array may be initially stored as 0 before the data of the gradient map and the coordinates within the window are stored.
[0015] By setting the window size to 1×1 and storing 0 as the initial value for all indices of the matrix array, we can also utilize the max pooling hardware as an activation function.
[0016] According to another aspect of the present invention, a maximum pooling method is provided, characterized in that it includes a step of storing data of a feature map in a window size in a matrix array during an inference process; a step of sequentially comparing data stored in the matrix array two by two by data comparators arranged in a tree structure and outputting a larger value while outputting the maximum value at the final stage;
[0017] According to another aspect of the present invention, a deep learning acceleration device is provided, comprising: a convolution operator for generating a feature map by performing a convolution operation on an input image; max pooling hardware for outputting the maximum value among data of the feature map; wherein the max pooling hardware comprises: a matrix array for storing data of the feature map in a window size during an inference process; and a forward propagation module including data comparators arranged in a tree structure and sequentially comparing two data stored in the matrix array to output a larger value while outputting the maximum value at the final stage.
[0018] According to another aspect of the present invention, a deep learning operation method is provided, comprising: a step in which a convolution operator performs a convolution operation on an input image to generate a feature map; a step in which max pooling hardware outputs a maximum value among data of the feature map; and the maximum output step includes a step in which a matrix array of the max pooling hardware stores data of the feature map in a window size during an inference process; and a step in which data comparators arranged in a tree structure in the max pooling hardware sequentially compare data stored in the matrix array two by two and output a larger value, and output the maximum value at a final stage.
[0019] As described above, according to embodiments of the present invention, the speed of max pooling operation can be improved through parallel operation in a CNN accelerator, various CNN models can be supported without changing the structure of max pooling hardware, and performance can be improved through real-time learning in the field.
[0020] Fig. 1. Environmental differences between systems
[0021] Figure 2. Various applications using CNN
[0022] Figure 3. Maximum pooling hardware configuration diagram according to one embodiment of the present invention.
[0023] Figure 4. Max pooling during the inference process
[0024] Figure 5. Max pooling during training
[0025] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0026] CNNs are being utilized in diverse fields, including autonomous driving, facial recognition, and image processing and restoration. To achieve high performance, CNNs are often designed with complex structures featuring deep layers. This increases computational complexity and demands high computing power. Complex and massive CNN models used in server-level systems (left side of Figure 1) that require high-performance application services can utilize multiple high-performance GPUs to perform computations.
[0027] For small embedded systems (right side of Figure 1), which are widely used in real life, there are relatively few cases requiring extremely high precision performance. Therefore, small CNN models optimized for this performance are used. However, the computational load is still significant compared to the performance of existing compute units used in small embedded systems. Because small embedded systems are constrained by power and size, high-performance GPUs are impractical, requiring hardware accelerators to perform deep learning.
[0028] Unlike server-grade systems designed for a variety of purposes, small embedded systems specifically designed to run CNN-based applications (Figure 2) are built for field use, and therefore have the advantage of being able to acquire image data suitable for their intended use in real time on-site. While CNN models typically used in small embedded systems have lower performance than those used in server-grade systems, the ability to train models in real time using the acquired data can further optimize the model for its intended use, thereby improving its performance.
[0029] Furthermore, even within the same system, multiple roles can be performed. Because different CNN models are appropriate for different image processing purposes, performance can be improved by configuring a system that can select and run two or more models appropriately for the specific situation.
[0030] Since the learning process of a CNN model requires more computation than the inference process, rather than performing the learning process internally within a small embedded system, the method of performing the learning process externally using the acquired image data and applying the parameters obtained through learning is mainly used.
[0031] In an embodiment of the present invention, the maximum pooling operation, which is a part of CNN operation, is performed in parallel to shorten the total operation time, thereby improving performance so that inference and learning processes can be performed in a small embedded system.
[0032] Conventional max pooling is performed serially, aligned with the output of the convolutional layer's operations. This serial max pooling method offers high efficiency when the window and stride sizes are set to the same size. However, if the window and stride sizes are not equal, the number of times data is read and written to memory increases, resulting in increased power consumption and longer computation cycles required to output the results.
[0033] Furthermore, embodiments of the present invention support variable hyperparameter application, enabling the operation of applications utilizing different CNN models. Max pooling requires varying computation cycles depending on hyperparameter settings such as window, stride, and padding.
[0034] Since hyperparameters are set differently depending on the CNN model, performing max pooling in parallel as much as the window size can significantly reduce the number of data inputs and outputs from memory, reducing computation cycles and improving performance.
[0035] FIG. 3 is a diagram illustrating a maximum pooling hardware configuration according to an embodiment of the present invention. As illustrated, the maximum pooling hardware according to an embodiment of the present invention is configured to include a DMA module (110), line buffer SRAMs (120), a matrix array (130), a DEMUX (140), a forward propagation module (150), and a backward propagation module (160).
[0036] The DMA module (110) accesses external memory, reads feature map data processed by the convolution operator, stores it in line buffer SRAMs (120), and stores feature map data output from the forward propagation module (150 / 160) in external memory.
[0037] The line buffer SRAMs (120) are buffer memories for storing feature maps by dividing them into rows. Specifically, the line buffer SRAMs (120) are composed of four buffers, and can retrieve and store up to four rows of feature map row data at a time from an external memory.
[0038] Data stored in the line buffer SRAMs (120) are transferred to the matrix array (130) in units of window size. In most small embedded CNN models, a window hyperparameter of 4 or less is used during max pooling. Accordingly, the size of the matrix array (130) is implemented as 4x4, but the size can be changed if necessary.
[0039] DEMUX (140) selectively transmits data stored in the matrix array (130) to the forward propagation module (150) and the reverse propagation module (160).
[0040] The forward propagation module (150) outputs the maximum value and the coordinates within the window of the maximum value among the data stored in the matrix array (130) and performs maximum pooling to reduce the size of the feature map.
[0041] Fig. 4 illustrates in detail the maximum pooling process in the inference process. As illustrated, feature map data is stored in the matrix array (130) in units of window sizes. If the window size (window hyperparameter) is smaller than the size of the matrix array (130), the lowest value (-128) is stored in the index of the matrix array (130) that falls outside the window. This is to prevent it from being selected regardless of which value is compared by the data comparator described later. To this end, the lowest value may be stored as an initial value in all indexes of the matrix array (130) before the feature map data are stored.
[0042] In the forward propagation module (150), data comparators are arranged in a tree structure to sequentially compare data stored in the matrix array (130) two by two, outputting the coordinate values of the larger value and the larger value, and outputting the maximum value and the coordinate values of the maximum value at the final stage.
[0043] A feature map composed of maximum values output by the forward propagation module (150) and a coordinate map thereof are stored in an external memory by the DMA module (110).
[0044] The coordinate map is used by the backpropagation module (160) during the learning process of the CNN model. Fig. 5 illustrates the max pooling process during the learning process in detail. As illustrated, the matrix array (130) stores the data of the gradient map and their coordinates within the window in a window size. If the window size (window hyperparameter) is smaller than the size of the matrix array (130), 0 is stored in the index of the matrix array (130) that falls outside the window. This is to prevent the output value from being affected even when added by the adder described later. To this end, it can be implemented so that 0 is stored as the initial value in all indexes of the matrix array (130) before the data of the gradient map are stored.
[0045] In the backpropagation module (160), the coordinate comparators output data whose coordinates are the maximum coordinates within the window, and the adder adds and outputs data output from the coordinate comparators.
[0046] Max pooling can be expressed as max(x0,x1,x2,x3). Meanwhile, the ReLU activation function can be expressed as max(0,x). If max pooling outputs the maximum value between input values, the ReLU activation function can be said to output the maximum value by comparing the input value with 0. Accordingly, by initializing all indices of the empty matrix array (130) to 0 and setting a 1x1 window hyperparameter to have one input data and performing max pooling, the ReLU activation function can be performed without additional hardware design. The max pooling hardware according to the embodiment of the present invention can also be used as an activation function after the max pooling operation.
[0047] So far, we have described in detail a preferred embodiment of max pooling hardware capable of accelerating learning and inference.
[0048] In an embodiment of the present invention, a maximum pooling hardware is proposed that improves the maximum pooling operation speed through parallel operation in a CNN accelerator, supports various CNN models without hardware change, and enables performance improvement through real-time learning in the field.
[0049] Meanwhile, it goes without saying that the technical idea of the present invention can be applied to a CNN acceleration device to which the proposed maximum pooling hardware is applied, i.e., a CNN acceleration device composed of a convolution operator, maximum pooling hardware, an activation function, etc.
[0050] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. A matrix array that stores data of feature maps in window size during the inference process; A maximum pooling hardware comprising a forward propagation module including data comparators that sequentially compare two data stored in a matrix array in a tree structure and output a larger value while outputting the maximum value at the final stage.
2. In claim 1, A maximum pooling hardware characterized by further including line buffers that store input feature maps by dividing them by row and transfer the stored data to a matrix array in units of window size.
3. In claim 1, During the inference process, if the window is smaller than the matrix array, Max pooling hardware characterized in that the index of the matrix array that leaves the window stores the lowest value.
4. In claim 3, In the inference process, before the data of the feature map is stored, all indices of the matrix array are Max pooling hardware characterized by storing the lowest value as the initial value.
5. In claim 1, The comparators are, Max pooling hardware characterized by outputting the maximum value and the coordinates within the window of the maximum value together.
6. In claim 5, The matrix array is, During the learning process, the data of the gradient map and the coordinates within the window are saved as window size, Max pooling hardware is, A maximum pooling hardware further comprising a backpropagation module including comparators that output data whose coordinates within a window are the coordinates of the maximum value and an adder that adds and outputs data output from the comparators.
7. In claim 6, During the learning process, if the window is smaller than the matrix array, Max pooling hardware characterized in that indices of a matrix array that exit a window are stored as 0.
8. In claim 7, In the learning process, before the data of the gradient map and the coordinates within the window are stored, all indices of the matrix array are Max pooling hardware characterized by having an initial value of 0.
9. In claim 1, Set the window size to 1×1, By storing 0 as the initial value for all indices of the matrix array, Max pooling hardware Max pooling hardware characterized by its use as an activation function.
10. A step in which the matrix array stores data of the feature map in window size during the inference process; A maximum pooling method characterized by including a step in which data comparators arranged in a tree structure sequentially compare two pieces of data stored in a matrix array and output a larger value while outputting the maximum value at the final stage.
11. A convolution operator that generates a feature map by performing a convolution operation on the input image; Includes max pooling hardware that outputs the maximum value among the data of the feature map; Max pooling hardware is, A matrix array that stores data of feature maps in window sizes during the inference process; A deep learning accelerator device characterized by including a forward propagation module including data comparators that sequentially compare two data stored in a matrix array in a tree structure and output a larger value while outputting the maximum value at the final stage.
12. A step in which a convolution operator performs a convolution operation on an input image to generate a feature map; The max pooling hardware includes a step of outputting the maximum value among the data of the feature map; The maximum output step is, A step in which the matrix array of the max pooling hardware stores data of the feature map in window size during the inference process; A deep learning operation method characterized by including a step in which data comparators arranged in a tree structure in maximum pooling hardware sequentially compare two pieces of data stored in a matrix array and output a larger value while outputting the maximum value at the final stage.
Citation Information
Patent Citations
Water leisure boat having means for improving water quality
KR102018062B1
Level measuring method using capacitance
KR1020210116180A
Performing average pooling in hardware
KR102370563B1
Block assembly method providing system
KR102597548B1
KR20200100812A