A method for reconfiguring a systolic array for CNN computing
By reordering and remapping the input feature data and weight data, the compatibility problem of pulsating arrays in different neural networks is solved, and the utilization rate of MAC is improved.
Patent Information
- Application Number
- CN202311758709.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-12-20
AI Technical Summary
In the prior art, since the size of the input feature map is not fixed in different neural networks and different neural network layers, the pulsating array is difficult to compatible with different neural networks, and conventional solutions increase the complexity of back-end layout and routing.
By reordering and remapping the input feature data and weight data, the reconfigurability of the pulsating array is realized, adapting to different operators and feature map input sizes, and improving MAC utilization.
The reconfigurability of the pulsating array is realized, it can be compatible with different operators and different feature map input sizes, and the utilization rate of MAC is improved.
Smart Images

Figure CN117992376B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of systolic arrays, and particularly relates to a method for reconstructing a systolic array for CNN computing. Background Art
[0002] When implementing CNN computing by ASIC, using a systolic array can better solve problems such as overly long broadcast of computing data paths and too high fan-in and fan-out caused by multiple parallel dimensions and high parallel dimensions of data. However, in practical applications, due to the variability of the size of the input feature map in different neural networks and different neural network layers, it is difficult for the systolic array to be compatible with different neural networks in practical applications. The current conventional solution is to use the hardware structure of the systolic array matrix, so that the computing process can be adjusted in real time according to the parameters of the input feature map, but this will greatly increase the complexity of the backend layout and wiring. Summary of the Invention
[0003] In order to solve at least one of the problems mentioned in the above background art, the present invention proposes a method for reconstructing a systolic array for CNN computing.
[0004] A method for reconstructing a systolic array for CNN computing includes the steps of:
[0005] Step S1, after the system controller M0 receives a loading instruction, it starts the DMA module and stores the feature data and weight data in the FSPM and WSPM respectively from the external memory in a certain arrangement order;
[0006] Step S2, after the feature data and weight data are loaded, the system controller M0 sends a start computing instruction to the PEA controller M1, and the systolic array controller M1 transfers the feature data and weight data from the FSPM and WSPM to the BF and BW in the systolic array respectively according to the computing dimensions of the PEA, where
[0007] The PEA is based on the basic processing unit PE and is designed as a 2-row and 8-column PE matrix. Each PE contains 8 independent PEMs, and each PEM contains 16 multipliers and an accumulator.
[0008] The minimum component unit of the BF is 32 feature map input channels, including two groups of BF0 registers and BF1 registers. Each group of registers includes 1 row and ρ columns, where the number of columns ρ satisfies the following formula:
[0009] ρ = w0 + k x -1
[0010] where w0 is the maximum number of output points calculated by the PEA per cycle, and k x is the maximum convolution kernel width supported;
[0011] The calculation process of BF is as follows:
[0012] When calculating the data in the BF0 register, the BF1 register loads the next set of feature data. After the feature data in the BF0 register is calculated, the calculation switches to the feature data in the BF1 register, and at the same time, the feature data in BF0 is updated sequentially;
[0013] The minimum building block of BW is a feature map of 32 input channels, which includes two groups of BW0 registers and BW1 registers. Each group of registers includes 1 row and σ columns, where the number of columns σ satisfies the following formula:
[0014] σ = 16k x
[0015] where k x is the maximum supported convolution kernel width;
[0016] The calculation process of BW is as follows:
[0017] When calculating the weight data in the BW0 register, the BW1 register loads the next set of weight data. After the weight data in the BW0 register is calculated, the calculation switches to the weight data in the BW1 register, and at the same time, the weight data in BW0 is updated sequentially.
[0018] According to the calculation dimensions of PEA, it specifically includes the dimensions according to the conventional CNN convolution and the dimensions according to the Depthwise convolution. Among them, according to the conventional CNN convolution dimensions, the bit width of the input feature data is 8bit, and its convolution dimensions are arranged from low to high as follows:
[0019] C → K x → K y → F → W → H
[0020] According to the dimensions of the Depthwise convolution, the bit width of the input feature data is 8bit and 16bit, and its convolution dimensions are arranged from low to high as follows:
[0021] C → K x → K y → W → H
[0022] In the formula, C is the input channel of the convolution kernel, Kx is the width of the convolution kernel, Ky is the height of the convolution kernel, F is the output channel of the convolution kernel, W is the width of the feature map, and H is the height of the feature map.
[0023] Step S3, perform the remapping and reordering process of the feature data and weight data in BF and BW. Among them, in the reordering process, during the PEA calculation, first, the feature data and weight data are accumulated and multiplied in units of 32 input channels according to the dimension of the input channels. After the polling of the dimension of the input channels is completed, switch to the next Kx dimension. After the polling of the Kx dimension is completed, switch to the next Ky dimension. After the polling of the Ky dimension is completed, switch to the dimension of the next group of 16 output channels;
[0024] The remapping process, according to different calculation modes, the corresponding execution processes are as follows:
[0025] When the PEA calculation mode is conventional CNN convolution and Depthwise + 8bit, the PEA structure is a 2-row and 8-column PE matrix, and the PEA longitudinal input is the feature data of 8 16 input channels;
[0026] When the PEA calculation mode is Depthwise + 16bit, the Depthwise calculation is the multiplication of the input channels of the feature data and the output channels of the weights. Each PE can only process the data of 8 input channels, and it is necessary to remap the output data of 4 16 output channels in BF to the longitudinal feature input registers of 8 PEs, and the longitudinal feature input registers of each PE store the data of 8 input channels.
[0027] Currently, the conventional method is to adopt the hardware structure of the systolic array matrix, so that the calculation process can be adjusted in real time according to the parameters of the input feature map. However, this will greatly increase the complexity of the backend layout and wiring. The present invention proposes to achieve the reconfigurability of the systolic array by reordering and remapping the input feature data and input weight data, which can not only be compatible with different operators and different input sizes of the feature map, but also effectively improve the utilization rate of MAC.
[0028] Step S4, the feature data and weight data are arranged in the timing required for PEA calculation and sent to the PEA for convolution operation. Specifically, for the calculation mode supported by the PEA, when the category of the convolutional network is α, the bit width of the input feature data is β, the number of input channels is C, and the number of output channels is F, the calculation steps are as follows:
[0029] Step A1, M1 loads the data of C input channels of 10 points of the feature data into BF0;
[0030] Step A2, M1 loads the weight data of C input channels and F output channels of K0 - K2 weights into BW0;
[0031] Step A3, BF0 outputs the data of C input channels of β points, and these data are sent to the PEA through the delay register;
[0032] Step A4: Output the K0 data of the number of C input channels and the number of F output channels from BW0. These data are sent to PEA through a delay register.
[0033] Step A5: PEA starts to calculate, and at the same time, M1 starts to load the next set of feature data and weight data into BF1 and BW1.
[0034] Step A6: BF0 shifts the data of the number of C input channels in the Kx direction to Step A3.
[0035] Step A7: BW0 shifts the data of the number of C input channels and the number of F output channels in the Kx direction to Step A4.
[0036] Step A8: Repeat Steps A3 - A7 until the data in BF0 and BW0 are consumed, and switch BF and BW to BF1 and BW1. At the same time, M1 starts to load the data of BF0 and BW0.
[0037] Step A9: Repeat Steps A3 - A8 until all the data is calculated.
[0038] Step S5: After the convolution operation is completed, its calculation result is written back to FSPM by PEA.
[0039] The present invention proposes a method for reconfiguring a systolic array for CNN calculation. Compared with the existing technologies, it has the following beneficial effects:
[0040] The present invention realizes the reconfigurability of the systolic array by reordering and remapping the input feature data and input weight data, which can not only be compatible with different operators and different input sizes of feature maps, but also effectively improve the utilization rate of MAC. Description of the Drawings
[0041] Figure 1 is the flowchart of the present invention;
[0042] Figure 2 is the system architecture diagram of the reconfigurable systolic array of the present invention;
[0043] Figure 3 is the schematic diagram of the data calculation process in BF in the CNN8bitF16C16 calculation mode of the embodiment of the present invention;
[0044] Figure 4 is the schematic diagram of the data calculation process in BW in the CNN8bitF16C16 calculation mode of the embodiment of the present invention. Detailed Embodiments
[0045] In order to make the objectives and features of the present invention more obvious and understandable, the technical solution will be described in detail below through embodiments and in conjunction with the drawings.
[0046] In this embodiment, according to Figure 2 the system architecture shown and Figure 1 the process shown, a systolic array reconstruction method for CNN computing is as follows:
[0047] Step S1, after the system controller M0 receives the loading instruction, it starts the DMA module, and stores the feature data and weight data in the FSPM and WSPM respectively from the external memory according to a certain arrangement order;
[0048] Optionally, the external memory is DDR.
[0049] Specifically, the storage format of the feature data in the FSPM is HWC, where H is the height of the feature map, W is the width of the feature map, and C is the number of input channels of the feature map.
[0050] The storage format of the weight data in the WSPM is KyKxFC, where Ky is the height of the convolution kernel, Kx is the width of the convolution kernel, F is the number of output channels, and C is the number of input channels.
[0051] Step S2, after the feature data and weight data are loaded, the system controller M0 sends a start calculation instruction to the PEA controller M1, and the systolic array controller M1 transfers the feature data and weight data from the FSPM and WSPM to the BF and BW in the systolic array respectively according to the calculation dimension of the PEA, where
[0052] the PEA is based on the basic processing unit PE and is designed as a 2-row and 8-column PE matrix. Each PE contains 8 independent PEMs. In the actual calculation process, each PE performs accumulation calculation internally, and the results between them are not accumulated. Each PEM contains 16 multipliers and an accumulator.
[0053] The minimum component unit of the BF is 32 feature map input channels, including two groups of BF0 registers and BF1 registers. Each group of registers includes 1 row and ρ columns, where the number of columns ρ satisfies the following formula:
[0054] ρ = w0 + k x -1
[0055] where w0 is the maximum number of output points calculated by the PEA per cycle, and k x is the maximum convolution kernel width supported;
[0056] The calculation process of the BF is:
[0057] When the data in the BF0 register is being calculated, the BF1 register loads the next set of feature data. After the feature data in the BF0 register is calculated, the calculation switches to the feature data in the BF1 register, and at the same time, the feature data in BF0 is updated sequentially.
[0058] The minimum building block of BW is a feature map with 32 input channels, which includes two groups of registers, BW0 and BW1. Each group of registers consists of 1 row and σ columns, where the number of columns σ satisfies the following formula:
[0059] σ = 16k x
[0060] where k x is the maximum supported convolution kernel width;
[0061] The calculation process of BW is as follows:
[0062] When the weight data in the BW0 register is being calculated, the BW1 register loads the next set of weight data. After the weight data in the BW0 register is calculated, the calculation switches to the weight data in the BW1 register, and at the same time, the weight data in BW0 is updated sequentially.
[0063] According to the calculation dimensions of PEA, it specifically includes the dimensions according to the conventional CNN convolution and the dimensions according to the Depthwise convolution. Among them, according to the dimensions of the conventional CNN convolution, the bit width of the input feature data is 8bit, and its convolution dimensions are arranged from low to high as follows:
[0064] C → K x → K y → F → W → H
[0065] According to the dimensions of the Depthwise convolution, the bit width of the input feature data is 8bit and 16bit, and its convolution dimensions are arranged from low to high as follows:
[0066] C → K x → K y → W → H
[0067] In the formula, C is the input channel of the convolution kernel, Kx is the width of the convolution kernel, Ky is the height of the convolution kernel, F is the output channel of the convolution kernel, W is the width of the feature map, and H is the height of the feature map.
[0068] Step S3, perform the remapping and reordering processes of the feature data and weight data in BF and BW to meet the supply of the feature data and weight data required for the systolic array calculation under different calculation modes.
[0069] Among them, in the reordering process, during PEA calculation, first, the feature data and weight data are accumulated and multiplied in units of 32 input channels according to the dimension of the input channels. After the dimension of the input channels is polled, it switches to the next Kx dimension. After the Kx dimension is polled, it switches to the next Ky dimension. After the Ky dimension is polled, it switches to the dimension of the next group of 16 output channels;
[0070] In the remapping process, according to different calculation modes, the corresponding execution processes are as follows:
[0071] When the PEA calculation mode is conventional CNN convolution and Depthwise + 8bit, the PEA structure is a 2-row and 8-column PE matrix, and the PEA longitudinal input is the feature data of 8 16-input channels;
[0072] When the PEA calculation mode is Depthwise + 16bit, in the Depthwise calculation, the input channels of the feature data are multiplied by the output channels of the weights. Each PE can only process the data of 8 input channels, and it is necessary to remap the output data of 4 16-output channels in the BF to the longitudinal feature input registers of 8 PEs, and the longitudinal feature input registers of each PE store the data of 8 input channels.
[0073] In step S4, the feature data and weight data are arranged into the timing required for PEA calculation and sent to the PEA for convolution operation. Specifically, for the calculation modes supported by the PEA, when the type of the convolutional network is α, the bit width of the input feature data is β, the number of input channels is C, and the number of output channels is F, the calculation steps are as follows:
[0074] Step A1, M1 loads the data of C input channels of 10 points of feature data into BF0;
[0075] Step A2, M1 loads the C input channels of weight data and the K0 - K2 weight data of F output channels into BW0;
[0076] Step A3, BF0 outputs the data of C input channels of β points, and these data are sent to the PEA through the delay register;
[0077] Step A4, BW0 outputs the K0 data of C input channels and F output channels, and these data are sent to the PEA through the delay register;
[0078] Step A5, the PEA starts to calculate, and at the same time, M1 starts to load the next group of feature data and weight data into BF1 and BW1;
[0079] Step A6, BF0 shifts the data of C input channels in the Kx direction to step A3;
[0080] Step A7, BW0 shifts the data of C input channel numbers and F output channel numbers in the Kx direction to step A4;
[0081] Step A8, repeat steps A3 - A7 until the data in BF0 and BW0 is consumed, and switch BF and BW to BF1 and BW1. Meanwhile, M1 starts to load the data of BF0 and BW0;
[0082] Step A9, repeat steps A3 - A8 until all the data is calculated.
[0083] When the type of the convolutional network is CNN, the bit width of the input feature data is 8 bit, the number of output channels is 16, and the number of input channels is 16. In this embodiment, when the above calculation steps are executed, the internal data calculation processes of BF and BW are respectively as Figure 3 and Figure 4 shown.
[0084] Step S5, after the convolution operation is completed, its calculation result is written back to FSPM by PEA.
[0085] So far, according to the method disclosed in the present invention, one working process of the present invention has been implemented.
[0086] Although the present invention has been described in detail with general descriptions and specific embodiments in this specification, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope claimed by the present invention.
Claims
1. A method for reconfiguring a systolic array for CNN computing, characterized in that, Including the steps: Step S1: After the system controller M0 receives a loading instruction, it starts the DMA module and stores the feature data and weight data in the FSPM and WSPM respectively from the external memory in a certain arrangement order. Among them, the storage format of the feature data in the FSPM is HWC, where H is the height of the feature map, W is the width of the feature map, and C is the number of input channels of the feature map. The storage format of the weight data in the WSPM is KyKxFC, where Ky is the height of the convolution kernel, Kx is the width of the convolution kernel, F is the number of output channels, and C is the number of input channels; Step S2: After the feature data and weight data are loaded, the system controller M0 sends a start calculation instruction to the PEA controller M1. The systolic array controller M1 transfers the feature data and weight data from the FSPM and WSPM to the BF and BW in the systolic array respectively according to the calculation dimension of the PEA. The PEA described in S2 is based on the basic processing unit PE and is designed as a 2-row and 8-column PE matrix. Each PE contains 8 independent PEMs, and each PEM contains 16 multipliers and an accumulator. The minimum component unit of the BF is 32 input channels of the feature map, and the minimum component unit of the BW is the feature map of 32 input channels; Step S3: Execute the remapping and reordering processes of the feature data and weight data in the BF and BW; Among them, for the remapping process described in step S3, the corresponding execution process according to different calculation modes is: When the PEA calculation mode is conventional CNN convolution and Depthwise+8bit, the PEA structure is a 2-row and 8-column PE matrix, and the PEA longitudinal input is 8 pieces of feature data with 16 input channels; When the PEA calculation mode is Depthwise+16bit, the Depthwise calculation is the multiplication of the input channels of the feature data and the output channels of the weight. Each PE can only process data with 8 input channels, and 4 pieces of output data with 16 output channels in the BF need to be remapped to the longitudinal feature input registers of 8 PEs. Each longitudinal feature input register of the PE stores data with 8 input channels; The reordering process described in step S3 is that during PEA calculation, first perform accumulative multiplication of the feature data and weight data in units of 32 input channels according to the dimension of the input channels. After the dimension polling of the input channels is completed, switch to the next Kx dimension. After the Kx dimension polling is completed, switch to the next Ky dimension. After the Ky dimension polling is completed, switch to the next group of 16 output channel dimensions; Step S4: The feature data and weight data are arranged into the timing required for PEA calculation and sent to the PEA for convolution operation; Step S5: After the convolution operation is completed, its calculation result is written back to the FSPM by the PEA.
2. The systolic array reconstruction method for CNN calculation according to claim 1, wherein In step S2, according to the calculation dimension of PEA, it specifically includes the dimension of conventional CNN convolution and the dimension of Depthwise convolution. Among them, according to the dimension of conventional CNN convolution, the bit width of the input feature data is 8bit, and its convolution dimensions are arranged from low to high as follows: C→K x →K y →F→W→H According to the dimension of Depthwise convolution, the bit widths of the input feature data are 8bit and 16bit, and its convolution dimensions are arranged from low to high as follows: C→K x →K y →W→H In the formula, C is the input channel of the convolution kernel, Kx is the width of the convolution kernel, Ky is the height of the convolution kernel, F is the output channel of the convolution kernel, W is the width of the feature map, and H is the height of the feature map.
3. A systolic array reconstruction method for CNN computing according to claim 1, wherein For the BF described in step S2, its minimum constituent unit is a feature map with 32 input channels, including two groups of BF0 registers and BF1 registers. Each group of registers includes 1 row and ρ columns, where the number of columns ρ satisfies the following formula: ρ = w0 + k x -1 where w0 is the maximum number of output points calculated per cycle by PEA, and k x is the maximum convolutional kernel width supported; The calculation process of BF is as follows: When the data in the BF0 register is being calculated, the BF1 register loads the next group of feature data. After the feature data in the BF0 register is calculated, it switches to the feature data in the BF1 register for calculation, and at the same time, the feature data in BF0 is updated sequentially.
4. A systolic array reconstruction method for CNN computing according to claim 1, characterized in that For the BW described in step S2, its minimum constituent unit is a feature map with 32 input channels, including two groups of BW0 registers and BW1 registers. Each group of registers includes 1 row and σ columns, where the number of columns σ satisfies the following formula: σ = 16k x where k x is the maximum convolutional kernel width supported; The calculation process of BW is as follows: When the weight data in the BW0 register is being calculated, the BW1 register loads the next group of weight data. After the weight data in the BW0 register is calculated, it switches to the weight data in the BW1 register for calculation, and at the same time, the weight data in BW0 is updated sequentially.
5. A systolic array reconstruction method for CNN computing according to claim 1, characterized in that In step S4, the transportation to PEA for convolution operation is specifically as follows. When the PEA-supported calculation mode, when the type of the convolution network is α, the bit width of the input feature data is β, the number of input channels is C, and the number of output channels is F, its calculation steps are as follows: Step A1, M1 loads the data of C input channels of 10 points of feature data into BF0; Step A2, M1 loads the C input channels of weight data and the K0-K2 weight data of F output channels into BW0; Step A3, BF0 outputs the data of C input channels of β points, and these data are sent to PEA through a delay register; Step A4, BW0 outputs the data of C input channels and the K0 data of F output channels, and these data are sent to PEA through a delay register; Step A5, PEA starts to calculate, and at the same time M1 starts to load the next group of feature data and weight data into BF1 and BW1; Step A6, BF0 shifts the data of C input channels in the Kx direction to step A3; Step A7, BW0 shifts the data of C input channels and F output channels in the Kx direction to step A4; Step A8, repeat steps A3-A7 until the data in BF0 and BW0 is consumed, and switch BF and BW to BF1 and BW1, and at the same time M1 starts to load the data of BF0 and BW0; Step A9, repeat Steps A3 - A8 until all data is calculated.
Citation Information
Patent Citations
FPGA (Field Programmable Gate Array)-based configurable CNN (Convolutional Neural Network) multiplication accumulator supporting 8-bit and 16-bit data
CN113138748A