A data reorganization method for supporting the training of a convolutional neural network model by a systolic array

By adopting a specific data recombination method in the pulsating array, the problems of large amount of computing and high memory access bandwidth requirements in deep neural network model training are solved, and the computing efficiency is improved, which is suitable for the training of deep neural network models.

CN115375973BActive Publication Date: 2025-07-22JIANGNAN INST OF COMPUTING TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211038910.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-22
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

In the computing process, deep neural network model training has a large amount of computation and high memory access bandwidth requirements. The existing technology requires explicit data layout format conversion, resulting in low computing efficiency.

Method used

The convolutional neural network model is trained using pulsating arrays. Through the data recombination method of forward convolution calculation, reverse calculation of residuals and reverse calculation of weights, the input and output feature maps follow the channel-first format, the convolution kernel follows the convolution kernel number-first format, and is transposed on the height and width dimensions to reduce the data conversion requirements.

Benefits of technology

It improves the spatial locality of the data and improves the computing efficiency. It is suitable for training and calculations of deep neural network models without explicit data conversion during different calculation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375973B_ABST
    Figure CN115375973B_ABST
Patent Text Reader

Abstract

A data reorganization method for supporting the training of a convolutional neural network model by a systolic array, belonging to the technical field of deep neural network model training. The present invention includes the following steps: Step 1, forward convolution calculation: the input and output feature maps follow the channel-first format, and the convolutional kernels follow the convolutional kernel number-first format; Step 2, reverse calculation of residuals: use the residuals of the output feature map in Step 1 as the input feature map, and use the convolutional kernels in Step 1 as the convolutional kernels; the input and output feature maps follow the channel-first format, and the convolutional kernels follow the convolutional kernel number-first format; Step 3, reverse calculation of weights: use the input feature map in Step 1 as the input feature map, and use the residuals of the output feature map in Step 1 as the convolutional kernels; the input and output feature maps follow the channel-first format, and the convolutional kernels follow the channel-first format. The present invention can improve the spatial locality of data, eliminate the need for number arrangement conversion in calculations, and improve the calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep neural network model training, and in particular to a data reorganization method for supporting the training of a convolutional neural network model using a systolic array. Background Art

[0002] During the calculation process of a deep neural network model, the amount of calculation and data access is huge. This not only requires high hardware computing power but also high memory access bandwidth. The characteristic of a systolic array is that data flows between computing units within the array, which can effectively increase the data reuse times, reduce the amount of data access, and thus reduce the bandwidth requirement. Therefore, using a systolic array to accelerate matrix multiplication / convolution operations is one of the common methods in the current academic and industrial circles.

[0003] A two-dimensional systolic array is the most common systolic array structure. During operation, the systolic array uses a data stream-driven method and flows in two directions respectively until the data flows to the position of the other diagonal of the systolic array. Due to the special data stream mechanism of the systolic array, there are also special requirements for the input data format.

[0004] The training of a deep neural network model includes three calculation processes. The three calculation processes are closely related, and the data inputs during calculation are different. It is difficult for the same data storage format to meet the calculation input requirements of the systolic array. Therefore, currently in the academic and industrial circles, explicit data layout format conversions are adopted during different calculation processes of training to meet the calculation input requirements of the systolic array and other components, which greatly reduces the calculation efficiency. Therefore, it is of great significance to find a data storage method with good data spatial locality and a generalizable forward and backward calculation of the network model. Summary of the Invention

[0005] The purpose of the present invention is to solve the problems existing in the above-mentioned prior art, and provide a data reorganization method for supporting the training of a convolutional neural network model using a systolic array, which can improve the spatial locality of data, eliminate the need for data layout conversion in the calculation, and improve the calculation efficiency.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] A data reorganization method for supporting the training of a convolutional neural network model using a systolic array includes the following steps:

[0008] Step 1, forward convolution calculation: the input and output feature maps follow the channel-first format, and the convolutional kernels follow the convolutional kernel number-first format;

[0009] Step 2, reverse calculate the residual: Use the residual of the feature map output in Step 1 as the input feature map, and use the convolutional kernel in Step 1 as the convolutional kernel; the input and output feature maps follow the channel-first format; the convolutional kernel follows the convolutional kernel number-first format and is transposed in the height and width dimensions;

[0010] Step 3, reverse calculate the weight: Use the input feature map in Step 1 as the input feature map, and use the residual of the feature map output in Step 1 as the convolutional kernel; the input and output feature maps follow the channel-first format, and the input feature map is rotated 180 degrees in the height and width dimensions; the convolutional kernel follows the channel-first format.

[0011] Preferably in the present invention, Step 1 specifically includes the following steps:

[0012] Step 1.1, starting from the starting address of the convolutional kernel, read data according to the granularity, continuously read the first points of the granularity number of convolutional kernels, and send them into the systolic array;

[0013] Step 1.2, offset the convolutional kernel reading address by M * R * S points, continue to read the second points of the granularity number of convolutional kernels, and send them into the systolic array; M is the number of convolutional kernels, R is the height of the convolutional kernel, and S is the width of the convolutional kernel;

[0014] Step 1.3, repeat the operation of Step 1.2 until M * C data are read and input into the systolic array; M is the number of convolutional kernels, and C is the channel depth of the convolutional kernel;

[0015] Step 1.4, starting from the starting address of the input feature map, read data according to the granularity, continuously read the data in the channel direction of the granularity number of input feature maps, and send them into the systolic array;

[0016] Step 1.5, offset the input feature map reading address by C points, continue to read the data in the channel direction of the granularity number of input feature maps, and send them into the systolic array; C is the channel depth of the picture;

[0017] Step 1.6, repeat the operation of Step 1.5 until all the input feature map data are read;

[0018] Step 1.7, sequentially replace the convolutional kernels and repeat Steps 1.1 - 1.6; finally, accumulate the intermediate values in the accumulation buffer.

[0019] Preferably in the present invention, Step 2 specifically includes the following steps:

[0020] Step 2.1, starting from the starting address of the convolutional kernel, read data according to the granularity, continuously read the first points of the granularity number of convolutional kernels, and send them into the transpose buffer;

[0021] Step 2.2, the read address of the convolution kernel is offset by M * R * S points, and then the second points of the granularity number of convolution kernels are continuously read and sent into the transpose buffer; M is the number of convolution kernels, R is the height of the convolution kernel, and S is the width of the convolution kernel;

[0022] Step 2.3, repeat the operation in Step 2.2 until M * C data are read and input into the transpose buffer. After realizing the data transpose in the M * C dimension, the data are sent into the systolic array; M is the number of convolution kernels, and C is the channel depth of the convolution kernel;

[0023] Step 2.4, starting from the starting address of the input feature map, read data according to the granularity, continuously read the data in the channel direction of the granularity number of input feature maps, and send them into the systolic array;

[0024] Step 2.5, the read address of the input feature map is offset by M points, and then the data in the channel direction of the granularity number of input feature maps are continuously read and sent into the systolic array; M is the channel depth of the picture;

[0025] Step 2.6, repeat the operation in Step 2.5 until all the input feature map data are read;

[0026] Step 2.7, sequentially replace the convolution kernels and repeat Steps 2.1 - 2.6; finally, accumulate the intermediate values in the accumulation buffer.

[0027] Preferably in the present invention, Step 3 specifically includes the following steps:

[0028] Step 3.1, starting from the starting address of the convolution kernel, read data according to the granularity, continuously read the first points of the granularity number of convolution kernels, and send them into the systolic array;

[0029] Step 3.2, the read address of the convolution kernel is offset by E * F * M points, and then the second points of the granularity number of convolution kernels are continuously read and sent into the systolic array; E is the height of the convolution kernel, F is the width of the convolution kernel, and M is the channel depth of the convolution kernel;

[0030] Step 3.3, repeat the operation in Step 2.2 until M * N data are read and input into the systolic array; M is the channel depth of the convolution kernel, and N is the number of convolution kernels;

[0031] Step 3.4, starting from the starting address of the input feature map, read data according to the granularity, continuously read the data in the N direction of the granularity number of input feature maps, and send them into the systolic array; N is the number of pictures;

[0032] Step 3.5, the read address of the input feature map is offset by N points, and then the data in the N direction of the granularity number of input feature maps are continuously read and sent into the systolic array; N is the number of pictures;

[0033] Step 3.6, repeat the operation in Step 3.5 until all the input feature map data is read;

[0034] Step 3.7, sequentially replace the convolution kernels, and repeat Steps 3.1 - 3.6; finally, accumulate the intermediate values in the accumulation buffer.

[0035] The advantages of the present invention are as follows: it can improve the spatial locality of data, is naturally applicable to the training calculation of deep neural network models, does not require explicit conversion of calculation data in different calculation processes, eliminates the need for number arrangement conversion in calculations, and improves calculation efficiency. Brief Description of the Drawings

[0036] Figure 1 It is a schematic diagram of the calculation process of the forward convolution algorithm in the training of deep learning algorithms;

[0037] Figure 2 It is a schematic diagram of the forward convolution calculation process in a data reorganization method for supporting the training of a convolutional neural network model by a systolic array according to the present invention;

[0038] Figure 3 It is a schematic diagram of the reverse calculation residual process in a data reorganization method for supporting the training of a convolutional neural network model by a systolic array according to the present invention;

[0039] Figure 4 It is a schematic diagram of the reverse calculation weight process in a data reorganization method for supporting the training of a convolutional neural network model by a systolic array according to the present invention;

[0040] Figure 5 It is a summary diagram of the convolution data reorganization method in a data reorganization method for supporting the training of a convolutional neural network model by a systolic array according to the present invention;

[0041] Figure 6 It is a schematic diagram of the forward convolution calculation process of a specific embodiment of the present invention;

[0042] Figure 7 It is a schematic diagram of the reverse calculation residual process of a specific embodiment of the present invention;

[0043] Figure 8 It is a schematic diagram of the reverse calculation weight process of a specific embodiment of the present invention. Specific Embodiment

[0044] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0045] The training of deep learning algorithms is divided into three main processes containing convolution: forward calculation, reverse transmission of residuals, and reverse update of weights. The calculation process of the forward convolution algorithm is as Figure 1 shown. Among them:

[0046] The format of the input feature maps is NHWC, where N is the number of images, H is the height of the image, W is the width of the image, and C is the depth of the image channels;

[0047] The format of the convolution kernels (filters / kernels) is RSCM, where M is the number of convolution kernels, R is the height of the convolution kernel, S is the width of the convolution kernel, and C is the depth of the convolution kernel channels;

[0048] The format of the output feature maps is NEFM, where N is the number of images, E is the height of the image, W is the width of the image, and C is the depth of the image.

[0049] The two reverse calculation processes are similar to the forward calculation process. During the training of the neural network model, the three calculation processes are closely related, and the results of the three calculation processes are mutually input and indispensable. Based on the use of a systolic array in the ShenWei-AI acceleration chip to accelerate the convolution calculation process, the present invention provides a data reorganization method for supporting the training of a convolutional neural network model using a systolic array, including the following steps:

[0050] Step 1, forward convolution calculation: As Figure 2 shown, the input and output feature maps follow the channel-first format, and the convolution kernels follow the convolution-kernel-number-first format;

[0051] Step 2, reverse calculation of the residual: As Figure 3 shown, using the residual of the output feature map in Step 1 as the input feature map and the convolution kernel in Step 1 as the convolution kernel; the input and output feature maps follow the channel-first format; the convolution kernels follow the convolution-kernel-number-first format and are transposed in the height and width dimensions;

[0052] Step 3, reverse calculation of the weight: As Figure 4 shown, using the input feature map in Step 1 as the input feature map and the residual of the output feature map in Step 1 as the convolution kernel; the input and output feature maps follow the channel-first format, and the input feature map is rotated 180 degrees in the height and width dimensions; the convolution kernels follow the channel-first format.

[0053] The convolution data reorganization methods in the above steps are summarized as Figure 5As shown in the figure, combining the systolic array acceleration component and the transpose component of Shenwei-AI, the data storage mode of the convolutional network is channel-first for input and output feature maps, that is, the input fmap dimension is [NHWC], and the output fmap dimension is [NEFM]; the number of convolutional kernels (weights) is direction-first, that is, kernel[CRSM]. When the neural network performs forward calculation, the previous layer network can naturally pass to the subsequent layer network without reshaping the data; when calculating the residual in the reverse direction, the transpose component is used to transpose the convolutional kernel to the C-first dimension for calculation, and the residual of the subsequent layer network can naturally pass to the previous layer network without data reshaping; when calculating the weights in the reverse direction, the transpose component is used to transpose the input feature map, and the convolutional kernel and the output feature map do not require data reshaping.

[0054] Taking a 32x32 two-dimensional systolic array as an example below, the data reorganization methods for the three calculation processes during model training are described in detail.

[0055] As Figure 6 shown, the data reorganization method in the forward convolution calculation process specifically includes:

[0056] The first step: The image is in the NHWC format with channel-first, and the convolutional kernel is in the CRSM format with the number of convolutional kernels first; the storage start address of the input image is ifmap_addr, and the storage address of the convolutional kernel is kernel_addr.

[0057] The second step: Read data from the convolutional kernel start address kernel_addr according to the granularity (32 data), continuously read the first points of 0 to 31 convolutional kernels, and send the data from the north direction of the systolic array; the read address offset of the convolutional kernel is M*R*S points (that is, skip the first channels of the remaining convolutional kernels), and continue to read 32 points (the second points (in the channel direction) of 0 to 31 convolutional kernels); continue the above operation until 32 times of 32 data are read and input to the systolic array.

[0058] The third step: Read data from the image start address ifmap_addr according to the granularity (32 data), continuously read 32 data in the channel direction of the image, and send them to the west side of the systolic array.

[0059] The fourth step: Continue the operation in the third step, continuously send image data, and the address offset of the image read each time is C points. At this time, each column of data in the accumulation buffer is a partial value (not calculated completely) of an image of a channel of the output feature map, and 32 columns of the accumulation buffer are partial values of the images of channels 0 to 31 of the output feature map.

[0060] The fifth step: Continue the operations from the first step to the fourth step, replace the convolutional kernel, and send image data to achieve accumulation in the accumulation buffer.

[0061] As Figure 7 shown, when calculating the residual, the input feature map at this time is the residual of the output feature map in the forward calculation, and the convolution kernel is still the convolution kernel in the forward calculation; the data reorganization method in the reverse calculation of the residual specifically includes:

[0062] The first step: The image is in the NEFM format with channel first, and the convolution kernel is in the CRSM format with the number of convolution kernels first; the storage start address of the input image is ifmap_addr, and the storage address of the convolution kernel is kernel_addr.

[0063] The second step: Read data from the convolution kernel start address kernel_addr in granularity (32 data), continuously read the first points of 0 to 31 convolution kernels, and send them into the transpose buffer; the convolution kernel read address is offset by M * R * S points (that is, cross the first channels of the remaining convolution kernels), and continue to read 32 points (the second points (in the channel direction) of 0 to 31 convolution kernels); continue the above operation until 32 times of 32 data are read and input into the transpose buffer to achieve the transpose of 32x32 dimensional data; after the transpose is completed, load the data from the north direction of the systolic array into the systolic array (the loading method is the same as the forward direction).

[0064] The third step: Read data from the image start address ifmap_addr in granularity (32 data), continuously read 32 data in the channel direction of the image, and send them to the west side of the systolic array.

[0065] The fourth step: Continue the operation in the third step, continuously send image data, and the address of each time reading the image is offset by M data. At this time, each column of data in the accumulation buffer is a partial value (not calculated completely) of an image of a channel of the output feature map, and 32 columns of the accumulation buffer are partial values of the image residuals of channels 0 to 31 of the input feature map A.

[0066] The fifth step: Continue the operations from the first step to the fourth step, realize replacing the convolution kernel, and send image data to perform accumulation in the accumulation buffer.

[0067] As Figure 8 shown, when calculating the weight, the input feature map at this time is the input feature map in the forward calculation, and the "convolution kernel" is the residual of the output feature map in the forward calculation; the data reorganization method in the reverse calculation of the weight specifically includes:

[0068] The first step: The image is in the NHWC format with channel first, and the "convolution kernel" is in the NEFM format with channel first; the storage start address of the input image is ifmap_addr, and the storage address of the "convolution kernel" is kernel_addr.

[0069] Step 2: Read data from the starting address kernel_addr of the "convolution kernel" in granularity (32 data). Continuously read the first points of 0 to 31 "convolution kernels" (i.e., channels 0 to 31 of the first point within the EF dimension of the output feature map residual), and send the data into the systolic array from the north direction; the read address of the "convolution kernel" is offset by EFM points (i.e., skip the first channel of the remaining "convolution kernels"), and continue to read 32 points (the second points (in the channel direction) of 0 to 31 convolution kernels); continue the above operations until 32 times of 32 data are read and input into the systolic array.

[0070] Step 3: Read data from the starting address ifmap_addr of the image in granularity (32 data). Continuously read 32 data in the N direction of the number of images, and send them into the west side of the systolic array.

[0071] Step 4: Continue the operation in Step 3, continuously send image data, and the address of the image read each time is offset by N points. At this time, each column of data in the accumulation buffer is a partial value (not calculated completely) of an image of a channel of the output feature map, and 32 columns of the accumulation buffer are partial values of the weight gradients of channels 0 to 31 of a certain convolution kernel.

[0072] Step 5: Continue the operations from Step 1 to Step 4, implement the replacement of the "convolution kernel", and send image data to perform accumulation in the accumulation buffer.

[0073] As mentioned above, it is only a preferred specific implementation manner of the present invention. This specific implementation manner is an implementation based on the overall concept of the present invention, and the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A data reorganization method for supporting the training of a convolutional neural network model by a systolic array, characterized in that It includes the following steps: Step 1, forward convolution calculation: The input and output feature maps follow the channel-first format, and the convolution kernels follow the convolution kernel number-first format; The specific steps of Step 1 are as follows: Step 1.1, The image is in the NHWC format of channel-first, and the convolution kernels are in the CRSM format of convolution kernel number-first; Starting from the starting address of the convolution kernel, read data in granularity, continuously read the first points of the granularity number of convolution kernels, and send them into the systolic array; Step 1.2, The convolution kernel read address is offset by M*R*S points, continue to read the second points of the granularity number of convolution kernels, and send them into the systolic array; M is the number of convolution kernels, R is the height of the convolution kernel, and S is the width of the convolution kernel; Step 1.3, Repeat the operation of Step 1.2 until M*C data are read and input into the systolic array; M is the number of convolution kernels, and C is the channel depth of the convolution kernel; Step 1.4, Starting from the starting address of the input feature map, read data in granularity, continuously read the data in the channel direction of the granularity number of input feature maps, and send them into the systolic array; Step 1.5, The input feature map read address is offset by C points, continue to read the data in the channel direction of the granularity number of input feature maps, and send them into the systolic array; C is the image channel depth; Step 1.6, Repeat the operation of Step 1.5 until all input feature map data are read; Step 1.7, Replace the convolution kernels in turn, send the image data, and repeat Steps 1.1 - 1.6; Finally, accumulate the intermediate values in the accumulation buffer; Step 2, reverse calculation of the residual: Use the residual of the output feature map in Step 1 as the input feature map, and use the convolution kernels in Step 1 as the convolution kernels; The input and output feature maps follow the channel-first format; The convolution kernels follow the convolution kernel number-first format, and are transposed in the dimensions of height and width; Step 3, reverse calculation of the weight: Use the input feature map in Step 1 as the input feature map, and use the residual of the output feature map in Step 1 as the convolution kernel; The input and output feature maps follow the channel-first format, and the input feature map is rotated 180 degrees in the dimensions of height and width; The convolution kernels follow the channel-first format.

2. The data reorganization method for supporting the training of a convolutional neural network model by a systolic array according to claim 1, wherein The specific steps of Step 2 are as follows: Step 2.1, Starting from the starting address of the convolution kernel, read data in granularity, continuously read the first points of the granularity number of convolution kernels, and send them into the transpose buffer; Step 2.2, The convolution kernel read address is offset by M*R*S points, continue to read the second points of the granularity number of convolution kernels, and send them into the transpose buffer; M is the number of convolution kernels, R is the height of the convolution kernel, and S is the width of the convolution kernel; Step 2.3, Repeat the operation of Step 2.2 until M*C data are read and input into the transpose buffer. After transposing the data in the M*C dimension, send the data into the systolic array; M is the number of convolution kernels, and C is the channel depth of the convolution kernel; Step 2.4, Starting from the starting address of the input feature map, read data in granularity, continuously read the data in the channel direction of the granularity number of input feature maps, and send them into the systolic array; Step 2.5: Shift the read address of the input feature map by M points, continue to read data in the channel direction of the number of granularity input feature maps, and send them into the systolic array; M is the depth of the image channels. Step 2.6: Repeat the operation in Step 2.5 until all the input feature map data is read. Step 2.7: Replace the convolution kernels in sequence and repeat Steps 2.1 - 2.6; finally, accumulate the intermediate values in the accumulation buffer.

3. A data reorganization method for supporting the training of a convolutional neural network model using a systolic array according to claim 1, characterized in that, The specific steps of Step 3 are as follows: Step 3.1: Read data from the starting address of the convolution kernel, read the first points of the number of granularity convolution kernels continuously, and send them into the systolic array. Step 3.2: Shift the read address of the convolution kernel by E * F * M points, continue to read the second points of the number of granularity convolution kernels, and send them into the systolic array; E is the height of the convolution kernel, F is the width of the convolution kernel, and M is the depth of the convolution kernel channels. Step 3.3: Repeat the operation in Step 2.2 until M * N data are input into the systolic array; M is the depth of the convolution kernel channels and N is the number of convolution kernels. Step 3.4: Read data from the starting address of the input feature map, read the data in the N direction of the number of granularity input feature maps continuously, and send them into the systolic array. N is the number of images. Step 3.5: Shift the read address of the input feature map by N points, continue to read the data in the N direction of the number of granularity input feature maps, and send them into the systolic array; N is the number of images. Step 3.6: Repeat the operation in Step 3.5 until all the input feature map data is read. Step 3.7: Replace the convolution kernels in sequence and repeat Steps 3.1 - 3.6; finally, accumulate the intermediate values in the accumulation buffer.

Citation Information

Patent Citations

  • Systolic array calculation structure and method for convolutional neural network

    CN110543934A

  • Data recombination method for systolic array structure

    CN110674927A