An Im2col Acceleration Method for Heterogeneous Many-Core Platforms
By selecting different Im2col acceleration algorithms based on the size of C*Kh on the heterogeneous multi-core platform, and through task division and DMA optimization, the problem of low Im2col computing efficiency in the existing technology is solved, the convolutional computing performance is improved, and the efficient operation of deep neural networks is ensured.
Patent Information
- Application Number
- CN202110349448.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-03-31
AI Technical Summary
The prior art has not yet deeply optimized Im2col computing on the multi-core platform, resulting in the increase in the proportion of Im2col computing when the GEMM performance is high enough, affecting the convolutional computing performance, thereby affecting the overall operation efficiency of deep neural networks.
It provides an Im2col acceleration method for heterogeneous multi-core platforms, selecting different algorithms according to the size of C*Kh, and improving the computing efficiency of Im2col transformation through task division and DMA optimization.
By designing different optimization algorithms for input tensors of different shapes, data utilization is improved, redundant communication is eliminated, the computing efficiency of Im2col transformation is effectively improved, and the efficient operation of convolution operators and convolutional neural networks is ensured.
Smart Images

Figure CN114219065B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an Im2col acceleration method for heterogeneous multi-core platforms, belonging to the technical field of deep learning on heterogeneous multi-core platforms. Background Art
[0002] The Im2col calculation in the convolution operator transforms a three-dimensional matrix input into a two-dimensional matrix to replace the original convolution calculation with an efficient matrix multiplication. Specifically, it transforms a tensor of C*H*W into a matrix of (C*Kh*Kw)*(Ho*Wo), where C is the number of channels, H and W are the height and width of the input respectively, Kh and Kw are the convolution kernel sizes, and Ho and Wo are the height and width of the output tensor. Assuming the convolution kernel size is 3*3, the number of channels is 1, and the input is 4*4, and it slides one grid each time, the shape of the output tensor is 2*2, and the shape of the matrix generated by the Im2col transformation is (1*3*3)*(2*2). The calculation process is as Figure 1 shown.
[0003] In deep learning, as an effective method for feature extraction, convolution calculation occupies a large proportion. Currently, there are various convolution algorithms, including the Im2col algorithm, FFT, Winograd algorithm, etc. Among them, the Im2col algorithm is used the most and has the widest application range. This method transforms the relatively complex and difficult-to-optimize convolution calculation into a matrix calculation, thereby reducing the memory access time and making full use of the already optimized GEMM library to accelerate the convolution calculation. Therefore, this algorithm includes two parts: the tensor expansion of Im2col and the matrix multiplication calculation. When the matrix multiplication performance is high enough, deep multi-core optimization of Im2col can effectively improve the efficiency of convolution calculation, thereby further accelerating the training of deep neural networks.
[0004] Currently, the optimization of convolution calculation on multi-core platforms mainly focuses on the optimization of GEMM, while Im2col has not been deeply optimized. When the GEMM performance is high enough, the proportion of Im2col calculation increases, affecting the convolution calculation performance and further affecting the overall operation efficiency of the deep neural network. Therefore, it is necessary to design a multi-core acceleration algorithm for Im2col that can effectively accelerate different input tensors. Summary of the Invention
[0005] The purpose of the present invention is to provide an Im2col acceleration method for heterogeneous multi-core platforms, which effectively improves the operation efficiency of the Im2col transformation. As a pre-processing process of convolution calculation, it effectively guarantees the efficient operation of convolution operators and convolutional neural networks.
[0006] To achieve the above object, the technical solution adopted by the present invention is: to provide an Im2col acceleration method for heterogeneous many-core platforms. The shape of the matrix after the Im2col transformation of a C*H*W tensor is (C*Kh*Kw)*(Ho*Wo), where C is the number of channels, H and W are the height and width of the input respectively, Kh and Kw are the convolution kernel sizes, and Ho and Wo are the height and width of the output tensor;
[0007] Select different algorithms according to the size of C*Kh: when C*Kh is greater than or equal to 64, starting from the transformed matrix, divide the tasks according to C*Kh; when C*Kh is less than 64, starting from the matrix before the transformation, divide the tasks according to C*H;
[0008] When C*Kh is greater than or equal to 64, select different implementations according to Ho of the output tensor and W of the input tensor:
[0009] When Ho*W is less than the maximum allocable space, the calculation process is as follows:
[0010] S11. Divide the transformed matrix into tasks in units of Kw rows according to C*Kh, and map them to the slave core group;
[0011] S12. For the Kw rows in the transformed matrix, read the corresponding Ho*W data from the input tensor at one time through DMA;
[0012] S13. For the Kw convolutional kernel elements in the same row, Ho*Wo corresponding results can be obtained from the read data respectively;
[0013] S14. Write back the results corresponding to each convolutional kernel to the corresponding position in the main memory in Kw times through DMA.
[0014] When Ho*W is greater than the maximum allocable space, the calculation process is as follows:
[0015] S21. Divide the transformed matrix into tasks in units of Kw rows according to C*Kh, and map them to the slave core group;
[0016] S22. According to the size of the local storage space, calculate the maximum number of rows col_block that can be accommodated when calculating W elements in one row;
[0017] S23. For the Kw rows in the transformed matrix, divide them in the Ho direction, and read them in batches through strided DMA. Each time, read col_block*W data, and the total DMA data volume is Ho*W;
[0018] S24. According to the read col_block*W data, for the Kw convolutional kernel elements in the same row, col_block*Wo results can be obtained;
[0019] S25. Write the results corresponding to each convolution kernel back to the corresponding position in the main memory in Kw times through DMA;
[0020] When C * Kh is less than 64, starting from the input tensor, divide the task according to C * H, and calculate with one row of the input tensor as a unit. The calculation process is as follows:
[0021] S31. Initialize all elements in the transformed matrix to 0;
[0022] S32. Divide the input tensor into tasks in row units according to C * H, and map them to the slave core group;
[0023] S33. Read in one row of input tensor elements through DMA each time;
[0024] S34. For one row in the input matrix, loop the convolution kernel in the column direction to determine the position in the Ho direction of the output matrix;
[0025] S35. Loop the convolution kernel in the row direction to obtain the elements corresponding to each convolution kernel, and write the data of Kw * Wo back to the main memory through strided DMA.
[0026] Due to the application of the above technical solutions, the present invention has the following advantages compared with the prior art:
[0027] The present invention designs different optimization algorithms for input tensors of different shapes, improves data utilization rate, removes redundant communication, effectively improves the operation efficiency of the Im2col transformation, and effectively guarantees the efficient operation of the convolution operator and the convolutional neural network as a preprocessing process for convolution calculation. Brief Description of the Drawings
[0028] Appendix Figure 1 is a schematic diagram of an Im2col acceleration method for a heterogeneous many-core platform according to the present invention Figure 1 ;
[0029] Appendix Figure 2 is a schematic diagram of an Im2col acceleration method for a heterogeneous many-core platform according to the present invention Figure 2 . Detailed Embodiments
[0030] Embodiment: The present invention provides an Im2col acceleration method for a heterogeneous many-core platform. The shape of the matrix after the Im2col transformation of a tensor of C * H * W is (C * Kh * Kw) * (Ho * Wo), where C is the number of channels, H and W are the height and width of the input respectively, Kh and Kw are the convolution kernel sizes, and Ho and Wo are the height and width of the output tensor;
[0031] Select different algorithms according to the value of C*Kh: when C*Kh is greater than or equal to 64, start from the transformed matrix and divide tasks according to C*Kh; when C*Kh is less than 64, start from the matrix before transformation and divide tasks according to C*H;
[0032] When C*Kh is greater than or equal to 64, select different implementations according to Ho of the output tensor and W of the input tensor:
[0033] When Ho*W is less than the maximum allocable space, the calculation process is as follows:
[0034] S11. Divide the transformed matrix into tasks in units of Kw rows according to C*Kh and map them to the slave core group;
[0035] S12. For Kw rows in the transformed matrix, read the corresponding Ho*W data from the input tensor at once through DMA;
[0036] S13. For Kw convolutional kernel elements in the same row, Ho*Wo corresponding results can be obtained from the read data respectively;
[0037] S14. Write the results corresponding to each convolutional kernel back to the corresponding position in the main memory in Kw times through DMA.
[0038] When Ho*W is greater than the maximum allocable space, the calculation process is as follows:
[0039] S21. Divide the transformed matrix into tasks in units of Kw rows according to C*Kh and map them to the slave core group;
[0040] S22. Calculate the maximum number of rows col_block that can be accommodated when calculating W elements in one row according to the size of the local storage space;
[0041] S23. Divide the Kw rows in the transformed matrix in the Ho direction, read them in batches through strided DMA, read col_block*W data each time, and the total DMA data volume is Ho*W;
[0042] S24. According to the read col_block*W data, col_block*Wo results can be obtained for Kw convolutional kernel elements in the same row;
[0043] S25. Write the results corresponding to each convolutional kernel back to the corresponding position in the main memory in Kw times through DMA;
[0044] When C*Kh is less than 64, start from the input tensor, divide tasks according to C*H, and calculate in units of one row of the input tensor. The calculation process is as follows:
[0045] S31. Initialize all elements in the transformed matrix to 0;
[0046] S32. Divide the input tensor by row according to C*H for task partitioning and map it to the slave core group;
[0047] S33. Read in one row of input tensor elements through DMA each time;
[0048] S34. For one row in the input matrix, loop through the convolution kernels in the column direction to determine the position in the Ho direction of the output matrix;
[0049] S35. Loop through the convolution kernels in the row direction to obtain the elements corresponding to each convolution kernel, and write back the Kw*Wo data to the main memory through strided DMA.
[0050] The further explanation of the above embodiments is as follows:
[0051] The present invention realizes the many-core optimization of Im2col calculation and the adaptive selection of algorithms, and can automatically select the optimal strategy according to the input tensor and convolution parameters.
[0052] The shape of the matrix after Im2col transformation is (C*Kh*Kw)*(Ho*Wo), and the input tensors corresponding to the convolution kernels in the same row are the same and can be processed simultaneously. Therefore, C*Kh*Kw is divided into (C*Kh)*Kw. First, different algorithms are selected according to the size of C*Kh: when C*Kh is greater than or equal to 64, starting from the transformed matrix, task partitioning is performed according to C*Kh; when C*Kh is less than 64, starting from the matrix before transformation, task partitioning is performed according to C*H.
[0053] For the case where C*Kh is greater than or equal to 64, different implementations can also be selected according to Ho of the output tensor and W of the input tensor:
[0054] The calculation process when Ho*W is less than the maximum allocable space is as follows:
[0055] 1) Divide the transformed matrix into task partitions with Kw rows as a unit according to C*Kh and map it to the slave core group;
[0056] 2) For the Kw rows in the transformed matrix, read in the corresponding Ho*W data from the input tensor through DMA at one time;
[0057] 3) For the Kw convolution kernel elements in the same row, Ho*Wo corresponding results can be obtained from the read-in data respectively;
[0058] 4) Write back the results corresponding to each convolution kernel to the corresponding positions in the main memory in Kw times through DMA.
[0059] The calculation process when Ho * W is greater than the maximum allocable space is as follows:
[0060] 1) Divide the transformed matrix into tasks with Kw rows as a unit according to C * Kh, and map them to the slave core group;
[0061] 2) Calculate the maximum number of rows col_block that can be accommodated when there are W elements in one row according to the size of the local storage space;
[0062] 3) Divide the Kw rows in the transformed matrix in the Ho direction, and read them in batches through strided DMA. Each time, read data of col_block * W, and the total DMA data volume is Ho * W;
[0063] 4) According to the read data of col_block * W, for the Kw convolutional kernel elements in the same row, col_block * Wo results can be obtained;
[0064] 5) Write back the results corresponding to each convolutional kernel to the corresponding position in the main memory through DMA in Kw times;
[0065] For the case where C * Kh is less than 64, task division according to C * Kh will cause load imbalance. Therefore, starting from the input tensor, task division is performed according to C * H, and calculations are performed with one row of the input tensor as a unit:
[0066] 1) Initialize all elements in the transformed matrix to 0;
[0067] 2) Divide the input tensor into tasks with rows as a unit according to C * H, and map them to the slave core group;
[0068] 3) Read one row of input tensor elements through DMA each time;
[0069] 4) As shown in the following figure, for one row in the input matrix, the output values corresponding to different row convolutional kernels are in a stepped shape. First, loop through the convolutional kernels in the column direction to determine the position in the Ho direction of the output matrix;
[0070] 5) Loop through the convolutional kernels in the row direction to obtain the elements corresponding to each convolutional kernel, and write back the data of Kw * Wo to the main memory through strided DMA.
[0071] When adopting the above-mentioned Im2col acceleration method for heterogeneous many-core platforms, different optimization algorithms are designed for input tensors of different shapes, which improves data utilization rate, eliminates redundant communication, effectively improves the operation efficiency of the Im2col transformation, and effectively guarantees the efficient operation of the convolutional operator and the convolutional neural network as a preprocessing process for convolutional calculation.
[0072] For better understanding of the present invention, the following terms used herein will be briefly explained:
[0073] Im2col: An algorithm for processing tensors, commonly used in convolutional operations of deep learning. It converts convolutional calculations into matrix multiplications, which can significantly accelerate the convolutional speed.
[0074] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. However, the protection scope of the present invention cannot be limited thereby. Any equivalent changes or modifications made according to the spirit of the present invention should be covered within the protection scope of the present invention.
Claims
1. An Im2col acceleration method for heterogeneous many-core platforms. The shape of the matrix after the Im2col transformation of a C*H*W tensor is (C*Kh*Kw)*(Ho*Wo), where C is the number of channels, H and W are the height and width of the input respectively, Kh and Kw are the convolution kernel sizes, and Ho and Wo are the height and width of the output tensor. It is characterized in that: Different algorithms are selected according to the size of C*Kh: when C*Kh is greater than or equal to 64, starting from the transformed matrix, task partitioning is performed according to C*Kh; when C*Kh is less than 64, starting from the matrix before transformation, task partitioning is performed according to C*H. When C*Kh is greater than or equal to 64, different implementations are selected according to Ho of the output tensor and W of the input tensor: When Ho*W is less than the maximum allocable space, the calculation process is as follows: S11. Task partition the transformed matrix by Kw rows as a unit according to C*Kh, and map it to the slave core group. S12. For the Kw rows in the transformed matrix, read Ho*W corresponding data from the input tensor at one time through DMA. S13. For the Kw convolutional kernel elements in the same row, Ho*Wo corresponding results can be obtained from the read data respectively. S14. Write the results corresponding to each convolutional kernel back to the corresponding position in the main memory in Kw times through DMA. When Ho*W is greater than the maximum allocable space, the calculation process is as follows: S21. Task partition the transformed matrix by Kw rows as a unit according to C*Kh, and map it to the slave core group. S22. According to the size of the local storage space, calculate the maximum number of rows col_block that can be accommodated when calculating W elements in one row. S23. For the Kw rows in the transformed matrix, partition in the Ho direction and read in batches through strided DMA, reading col_block*W data each time, and the total DMA data volume is Ho*W. S24. According to the read col_block*W data, for the Kw convolutional kernel elements in the same row, col_block*Wo results can be obtained. S25. Write the results corresponding to each convolutional kernel back to the corresponding position in the main memory in Kw times through DMA. When C*Kh is less than 64, starting from the input tensor, task partitioning is performed according to C*H, and calculations are performed with one row of the input tensor as a unit. The calculation process is as follows: S31. Initialize all elements in the transformed matrix to 0. S32. Task partition the input tensor by rows according to C*H, and map it to the slave core group. S33. Read one row of input tensor elements through DMA each time. S34. For one row in the input matrix, loop through the convolutional kernel in the column direction to determine the position in the Ho direction of the output matrix. S35. Loop through the convolutional kernel in the row direction to obtain the elements corresponding to each convolutional kernel, and write Kw*Wo data back to the main memory through strided DMA.
Citation Information
Patent Citations
Convolutional neural network operation acceleration method and device based on many-core processor
CN111461311A
Convolution acceleration method based on heterogeneous many-core processor
CN112446471A