Neural network processor hardware, with minimized external memory access, for supporting resnet structure of binarized neural network
The NNP hardware addresses the limitations of deep learning hardware by incorporating BNN cores with minimal external memory access, enhancing computational speed and reducing power consumption for ResNet data flow.
Patent Information
- Application Number
- PCT/KR2023/020604
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-19
AI Technical Summary
The processing speed of deep learning hardware is limited by external memory read/write, and external memory access consumes significantly more energy than multiplication operations, leading to high power consumption. Additionally, applying Binarized Neural Networks (BNNs) to Residual Networks (ResNets) generates floating-point data in BatchNorm and Add operations, requiring efficient memory access management.
The proposed Neural Network Processor (NNP) hardware includes a weight memory for binary weights, an input memory for binary inputs, and a plurality of BNN cores that perform binary convolution operations with minimal external memory access. Each BNN core comprises a Processing Element Array (PEA) for binary convolution, a normalization operator for batch-normalization, an ADD operator for accumulation, and an ACT operator for binarization.
This solution minimizes external memory access, reducing power consumption and hardware size while increasing computational speed for ResNet data flow in BNN-based NNP hardware.
Smart Images

Figure KR2023020604_19062025_PF_FP_ABST
Abstract
Description
Neural network processor hardware to support the ResNet structure of binary neural networks with minimal external memory access.
[0001] The present invention relates to a hardware structure of an NNP (Neural Network Processor), and more particularly, to a hardware structure of an NNP that supports a ResNet (Residual Network) of a BNN (Binarized Neural Network).
[0002] The processing speed of deep learning hardware is limited by external memory reads and writes. Furthermore, as shown in Figure 1, external memory access consumes approximately 200 times more energy than multiplication operations, significantly impacting power consumption. Therefore, minimizing memory access is a key factor in increasing the speed and achieving low-power implementations of deep learning hardware.
[0003] Deep learning hardware architectures have an inverse relationship between the number of external memory accesses and the size of the internal memory. However, using a large internal memory size increases hardware cost, resulting in a trade-off relationship. BNNs use convolution operations by converting them into 1-bit operations, reducing memory capacity and access times by 1 / 32. Furthermore, 1-bit operations replace multiplication operations with XOR operations, which has the advantage of reducing the size of the calculation unit.
[0004] However, when applying BNN to ResNet, which shows high accuracy, floating point data is generated in BatchNorm and Add operations as shown in Fig. 2, so floating point data is required except for convolution.
[0005] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide efficient NNP hardware for ResNet data flow that can minimize external memory access.
[0006] According to one embodiment of the present invention for achieving the above object, NNP (Neural Network Processor) hardware includes: a weight memory in which binary weights required for a binary convolution operation are stored; an input memory in which binary inputs required for the binary convolution operation are stored; and a plurality of BNN (Binarized Neural Network) cores for performing a binary convolution operation with binary inputs stored in the input memory and binary weights stored in the weight memory.
[0007] Each BNN core may include a Processing Element Array (PEA) that performs an n×n binary convolution operation; a normalization operator that batch-normalizes the operation results of the PEA; an ADD operator that accumulates the batch-normalized operation results; and an ACT operator that binarizes the accumulated operation results.
[0008] The PEA may include n PERs (PE Rows) that perform n×n binary convolution operations with binary inputs stored in input memory and binary weights stored in weight memory; an ACC operator that adds the outputs of the PERs and the length of pop;
[0009] Each PER may include n PEs (Processing Elements) that perform operations on one row of n×n binary convolution operations with binary inputs stored in input memory and binary weights stored in weight memory.
[0010] Each PE can perform a binary convolution operation on one pixel among the n pixels that make up the row, and pass the operation result to the next PE using pop.
[0011] PE can process k binary inputs for one pixel simultaneously.
[0012] The PE may include a scratch pad storing k binary weights input from a weight memory; k operators performing XOR operations on the k binary inputs input from the input memory and the k binary weights stored in the scratch pad; a popcounter summing and outputting operation results of the operators; an ADD operator accumulating the output of the popcounter with the result of the previous PE; and a delay for delaying the output of the ADD operator and transmitting it to the next PE.
[0013] The PE may further include a delay for delaying k binary inputs and passing them to the next PE.
[0014] BNN can be a BNN that supports the ResNet structure.
[0015] According to another aspect of the present invention, a binary convolution operation method is provided, characterized in that it includes: a step of storing binary weights required for a binary convolution operation in a weight memory; a step of storing binary inputs required for a binary convolution operation in an input memory; and a step of performing a binary convolution operation by a plurality of BNN (Binarized Neural Network) cores using binary inputs stored in the input memory and binary weights stored in the weight memory.
[0016] According to another aspect of the present invention, a BNN hardware is provided, characterized by including a Processing Element Array (PEA) that performs an n×n binary convolution operation; a normalization operator that batch-normalizes the operation results of the PEA; an ADD operator that accumulates the batch-normalized operation results; and an ACT operator that binarizes the accumulated operation results.
[0017] According to another aspect of the present invention, a binary convolution operation method is provided, characterized in that it includes a step of a PEA (Processing Element Array) performing an n×n binary convolution operation; a step of a normalization operator performing a batch normalization operation result; a step of an ADD operator accumulating the batch normalized operation result; and a step of an ACT operator binarizing the accumulated operation result.
[0018] As described above, according to embodiments of the present invention, by minimizing external memory access through efficient NNP hardware for ResNet data flow, power consumption and hardware size can be reduced while increasing computational speed.
[0019] Figure 1. Comparison of power consumption according to operation and memory access.
[0020] Figure 2. ResNet structure using BNN
[0021] Figure 3 is a structure of an NNP according to one embodiment of the present invention.
[0022] Figure 4 shows the detailed structure of PEA.
[0023] Figure 5. 3×3 binary convolution operation
[0024] Figure 6. Detailed structure of PER
[0025] Fig. 7. 1×3 binary convolution operation
[0026] Fig. 8. Detailed structure of PE
[0027] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0028] According to one embodiment of the present invention, a hardware structure and implementation method of a Neural Network Processor (NNP) equipped with a Binarized Neural Network (BNN) that supports a ResNet structure and minimizes external memory access are presented. This technology relates to efficient BNN-based NNP hardware that supports ResNet data flow that can minimize external memory access.
[0029] FIG. 3 is a diagram illustrating the structure of an NNP according to an embodiment of the present invention. As illustrated in FIG. 3, the NNP according to an embodiment of the present invention is configured to include nine BNN cores (100), a binary weight memory (Binaryized Weight Memory, 410), a Con. memory (420), a floating in / out memory (430), and a binary input memory (Binaryized Input Memory, 440).
[0030] The binary weight memory (410) stores binary weights required for binary convolution operations, the binary input memory (440) stores binary inputs required for binary convolution operations, the Con. memory (420) stores parameters required for batch normalization, and the Floating In / Out memory (430) is a memory in which real numbers are stored during the operation of the BNN core (100).
[0031] The BNN core (100) performs a binary convolution operation using binary weights stored in the binary weight memory (410) and binary inputs stored in the binary input memory (440). The BNN cores (100) that perform such a function are configured to include a Processing Element Array (PEA) 110, a batch normalization operator (120), a DD operator (130), and an ACT operator (140), as illustrated.
[0032] PEA (110) performs a 3×3 binary convolution operation, and the batch normalization operator (120) is a floating-point multiplier for batch normalizing the operation result of PEA (110).
[0033] The ADD operator (130) accumulates the operation results that have been batch normalized by the batch normalization operator (120), and the ACT operator (140) binarizes the operation results accumulated by the ADD operator (130).
[0034] Hereinafter, the structure and operation of PEA (110) will be described in detail with reference to FIG. 4. FIG. 4 is a drawing showing the detailed structure of PEA (110). As illustrated, PEA (110) is configured to include three PERs (Processing Element Rows) (210-0, 210-1, 210-2) and an ACC operator (220).
[0035] PERs (210-0, 210-1, 210-2) perform a 3×3 binary convolution operation (Fig. 5) with binary inputs stored in the binary input memory (440) and binary weights stored in the binary weight memory (410).
[0036] The ACC operator (220) adds the outputs of PERs (210-0, 210-1, 210-2) and the length of pop and outputs the result.
[0037] Below, the structure and operation of PERs (210-0, 210-1, 210-2) are described in detail with reference to Fig. 6. Fig. 6 is a drawing illustrating the detailed structure of PERs (210-0, 210-1, 210-2).
[0038] Since PERs (210-0, 210-1, 210-2) have the same structure and operate in the same manner, only one PER (210-0, 210-1, 210-2) is illustrated in FIG. 6, represented by the reference numeral "210".
[0039] PER (210) performs a 1×1 binary convolution operation (Fig. 7), which is a binary convolution operation for one row among 3×3 binary convolution operations, using binary inputs stored in the binary input memory (440) and binary weights stored in the binary weight memory (410).
[0040] PER (210) performing such a function is configured to include three PEs (300-0, 300-1, 300-2) as illustrated in Fig. 6. Each PE performs a binary convolution operation on one pixel among the three pixels constituting one row and passes the operation result to the next PE by popping it.
[0041] Below, the structure and operation of PEs (300-0, 300-1, 300-2) are described in detail with reference to Fig. 8. Fig. 8 is a drawing illustrating the detailed structure of PEs (300-0, 300-1, 300-2).
[0042] Since PEs (300-0, 300-1, 300-2) have the same structure and operate identically, only one PE (300-0, 300-1, 300-2) is represented in FIG. 8 by reference numeral "300".
[0043] PE (300) is a structure capable of simultaneously processing k binary inputs for one pixel. Specifically, PE (300) is configured to include an input delay (310), a scratch pad (320), k XOR operators (330), a popcounter (340), an ADD operator (350), and an output delay (360).
[0044] The input delay (310) delays k binary inputs input from the binary input memory (440) and transmits them to the next PE. The scratch pad (320) stores k binary weights input from the binary weight memory (410).
[0045] The k XOR operators (330) perform an XOR operation on the k binary inputs input from the binary input memory (440) and the k binary weights stored in the scratch pad (320).
[0046] The popcounter (340) adds up the operation results of the XOR operators (330) and outputs them, and the ADD operator (350) accumulates the output of the popcounter (340) with the result of the previous PE and outputs them.
[0047] The output delay (360) delays the output of the ADD operator (350) and transmits it to the next PE.
[0048] So far, we have described in detail a preferred embodiment of NNP hardware for supporting the ResNet structure of BNN with minimal external memory access.
[0049] In the above embodiment, by minimizing external memory access through efficient NNP hardware for ResNet data flow, power consumption and hardware size can be reduced while increasing computational speed.
[0050] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. Weight memory where binary weights required for binary convolution operations are stored; Input memory where binary inputs required for binary convolution operation are stored; and An NNP (Neural Network Processor) hardware comprising a plurality of BNN (Binarized Neural Network) cores that perform binary convolution operations with binary inputs stored in an input memory and binary weights stored in a weight memory.
2. In claim 1, Each BNN core, Processing Element Array (PEA) that performs n×n binary convolution operations; A normalization operator that batch normalizes the results of PEA operations; An ADD operator that accumulates the results of batch normalized operations; A Neural Network Processor (NNP) hardware characterized by including an ACT operator that binarizes accumulated operation results.
3. In claim 2, PEA is, n PERs (PE Rows) that perform n×n binary convolution operations with binary inputs stored in input memory and binary weights stored in weight memory; A Neural Network Processor (NNP) hardware characterized by including an ACC operator that adds the output of PERs and the length of pop.
4. In claim 3, Each PER is, An NNP (Neural Network Processor) hardware comprising: n PEs (Processing Elements) each of which performs an operation on one row of n×n binary convolution operations using binary inputs stored in input memory and binary weights stored in weight memory.
5. In claim 4, Each PE, NNP (Neural Network Processor) hardware characterized by performing a binary convolution operation on one pixel among n pixels forming a row and passing the operation result to the next PE by popping it.
6. In claim 5, PE is, PE is a Neural Network Processor (NNP) hardware characterized by processing k binary inputs for one pixel simultaneously.
7. In claim 6, PE is, A scratch pad storing k binary weights input from weight memory; K operators that perform XOR operations on k binary inputs from input memory and k binary weights stored in the scratch pad; popcounter, which adds up the results of operations of the operators and outputs them; An ADD operator that accumulates the output of the popcounter with the result of the previous PE; and A Neural Network Processor (NNP) hardware characterized by including a delay for delaying the output of an ADD operator and transmitting it to the next PE.
8. In claim 7, PE is, A Neural Network Processor (NNP) hardware characterized by further including a delay for delaying k binary inputs and transmitting them to the next PE.
9. In claim 8, BNN is, NNP (Neural Network Processor) hardware featuring a BNN that supports the ResNet structure.
10. A step in which the weight memory stores binary weights required for binary convolution operations; A step for storing binary inputs required for binary convolution operation in the input memory; and A binary convolution operation method, characterized by including a step in which a plurality of BNN (Binarized Neural Network) cores perform a binary convolution operation using binary inputs stored in an input memory and binary weights stored in a weight memory.
11. Processing Element Array (PEA) that performs n×n binary convolution operations; A normalization operator that batch normalizes the results of PEA operations; An ADD operator that accumulates the results of batch normalized operations; BNN hardware characterized by including an ACT operator for binarizing accumulated operation results.
12. A step in which the PEA (Processing Element Array) performs an n×n binary convolution operation; A step in which the normalization operator batch-normalizes the operation results; A step in which the ADD operator accumulates batch normalized operation results; A binary convolution operation method, characterized in that the ACT operator includes a step of binarizing the accumulated operation result.
Citation Information
Patent Citations
Binary weight convolutional neural network accelerator and RISC-V system on chip
CN115983350A
Mask
KR1020220157822A
Cable tray
KR102303740B1
Processor Accelerating Convolutional Computation in Convolutional Neural Network AND OPERATING METHOD FOR THE SAME
KR102345409B1
one-pot synthesis Method of 1,3-Disubstitued indolizines
KR102765208B1