A Convolution Operation Method Based on Dilated Fetching on a Heterogeneous Many-Core Architecture

By using the method of expanding number fetching and data reusability optimization on heterogeneous multi-core processors, the problems of high memory access pressure caused by convolutional operations and underutilization of computing resources are solved, and significant performance improvement is achieved.

CN114218521BActive Publication Date: 2025-06-17JIANGNAN INST OF COMPUTING TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110452546.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-26
Publication Date
2025-06-17
Estimated Expiration
2041-04-26

AI Technical Summary

Technical Problem

Convolutional operations cause high memory access pressure and underutilization of computing resources on heterogeneous multi-core processors. Traditional optimization methods such as im2col increase system memory pressure.

Method used

The expansion number acquisition method on the heterogeneous multi-core architecture is adopted to read data in one memory access through redundant reading, calculate multiple results, reduce the number of memory accesses, and improve computing efficiency by maximizing data reusability.

Benefits of technology

It significantly reduces the memory memory access requirement of convolutional computing, saves memory bandwidth resources, makes full use of the computing resources of many cores, and improves the performance of convolutional computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218521B_ABST
    Figure CN114218521B_ABST
Patent Text Reader

Abstract

The present invention discloses a convolution operation method based on dilation fetching on a heterogeneous many-core architecture, which includes the following steps: S1. Input input, weight weight, and stride stride, where input is Hi*Wi, weight is K*K, and the shape of the output output is calculated according to the shapes of input and weight to obtain Ho*Wo; S2. According to the shape of output, in the Ho and Wo dimensions, the convolution calculation tasks are evenly distributed to the many cores according to the logical number of each core, and each core processes the calculation task with a size of Ho_BLOCK*Wo_BLOCK; S3. Each core calculates the required input size Hi_BLOCK*Wi_BLOCK according to its own task size, where Hi_BLOCK = Ho_BLOCK*stride + K - 1 and Wi_BLOCK = Wo_BLOCK*stride + K - 1; S4. Each core performs convolution calculation through the obtained input (Hi_BLOCK*Wi_BLOCK) and weight; S5. Repeat S3 and S4 until the calculation is completed, and output output. The present invention saves memory bandwidth resources and can fully utilize the computing resources of the many cores.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a convolution operation method based on dilated fetching on a heterogeneous many-core architecture, belonging to the technical field of deep learning. Background Art

[0002] Convolution is one of the most important concepts in deep learning. During the training and inference processes of the entire convolutional neural network, convolution operations account for the vast majority of the computational workload. High-performance computing platforms usually need to provide specialized solutions for such core operations. For computationally intensive functions, such as convolution in deep learning, how to timely provide sufficient data for powerful computing cores and improve the data reuse rate is a problem that needs to be solved.

[0003] Convolution operation is the core operation of the artificial intelligence CNN network. The data of each K times of convolution operation overlaps but does not repeat, and the rule of fetching data is not strong during the operation (K is the size of the convolution operation kernel). This leads to frequent data fetching in convolution operations. If the characteristics of convolution itself cannot be cleverly utilized and instantiated on a heterogeneous many-core processor according to the traditional algorithm, not only the computing resources of the processor cannot be fully utilized, but also a great pressure on memory access will be caused. Therefore, when the convolution operation fully utilizes the computing resources of the heterogeneous many-core processor, how to reduce the pressure on the system memory access is also a problem that needs to be solved.

[0004] Currently, there are some optimized convolution operation methods. For example, im2col is to convert the convolution operation into a matrix multiplication and optimize the convolution operation using the optimized matrix multiplication. However, this method needs to expand the input to K*K times the original, which will cause additional pressure on the system memory. A heterogeneous many-core processor contains a large number of slave cores with powerful computing capabilities, and the memory access bandwidth is the bottleneck of the system. For such computationally intensive operations, such a method not only cannot utilize the computing resources of the processor, but will instead cause a huge pressure on the system bandwidth, and the optimization effect is not good. Therefore, how to efficiently utilize the memory access bandwidth and reduce the memory access pressure is the key to fully exerting the performance of the processor. Summary of the Invention

[0005] The purpose of the present invention is to provide a convolution operation method based on dilated fetching on a heterogeneous many-core architecture, which...

[0006] To achieve the above purpose, the technical solution adopted by the present invention is: to provide a convolution operation method based on dilated fetching on a heterogeneous many-core architecture, including the following steps:

[0007] S1. Input input, weight weight, and stride stride, where input is Hi*Wi, weight is K*K, calculate the shape of the output output according to the shapes of input and weight, and obtain Ho*Wo;

[0008] S2. According to the shape of the output, in the Ho and Wo dimensions, and based on the logical number of each core, evenly distribute the convolution calculation tasks to multiple cores, with each core processing a calculation task of size Ho_BLOCK * Wo_BLOCK;

[0009] S3. Each core calculates the required input size Hi_BLOCK * Wi_BLOCK according to its own task size, where Hi_BLOCK = Ho_BLOCK * stride + K - 1 and Wi_BLOCK = Wo_BLOCK * stride + K - 1;

[0010] S4. Each core performs convolution calculation through the obtained input (Hi_BLOCK * Wi_BLOCK) and weight;

[0011] S5. Repeat steps S3 and S4 until the tasks assigned to each core are calculated, and then output the output.

[0012] Due to the application of the above technical solutions, the present invention has the following advantages compared with the prior art:

[0013] A convolution operation method based on inflated data fetching on a heterogeneous multi-core architecture of the present invention, without increasing the occupation of additional system memory, through inflated data fetching, not only reduces the memory access pressure of the system, but also improves the data reusability, significantly reduces the memory access requirements of convolution calculation on a heterogeneous multi-core processor, saves memory bandwidth resources, and at the same time can make full use of the computing resources of multiple cores, give full play to the system performance, meet the characteristics of convolution calculation-intensive, and significantly improve the performance of convolution operation. Description of the Drawings

[0014] Att Figure 1 is a schematic diagram of two-dimensional convolution;

[0015] Att Figure 2 is a schematic diagram of the data source required for convolution output;

[0016] Att Figure 3 is a schematic diagram of three-dimensional convolution;

[0017] Att Figure 4 is a schematic diagram of inflated data fetching;

[0018] Att Figure 5 is a block diagram of the operation principle. Detailed Embodiment

[0019] Embodiment: The present invention provides a convolution operation method based on inflated data fetching on a heterogeneous multi-core architecture, which specifically includes the following steps:

[0020] S1. Input, weight, and stride are provided, where the input has the shape Hi*Wi, the weight has the shape K*K, and the shape of the output output is calculated based on the shapes of the input and the weight, resulting in Ho*Wo;

[0021] S2. Based on the shape of the output, in the Ho and Wo dimensions, according to the logical numbering of each kernel, the convolution calculation tasks are evenly distributed among multiple kernels, and each kernel processes a calculation task of size Ho_BLOCK*Wo_BLOCK;

[0022] S3. Each kernel calculates the required input size Hi_BLOCK*Wi_BLOCK according to its own task size, where Hi_BLOCK = Ho_BLOCK*stride + K - 1 and Wi_BLOCK = Wo_BLOCK*stride + K - 1;

[0023] S4. Each kernel performs convolution calculation using the obtained input (Hi_BLOCK*Wi_BLOCK) and weight;

[0024] S5. Repeat steps S3 and S4 until the calculation of the tasks assigned to each kernel is completed, and output the output.

[0025] A further explanation of the above embodiments is as follows:

[0026] The convolution operation is as Figure 1 shown:

[0027] The CPU completes the calculation of the output data block C through the input data block A and the weight data block B, where C = AXB, and X represents the convolution operation.

[0028] The sources of the input data required for each output element are as Figure 2 shown:

[0029] The output of the convolution calculation has four elements, which are represented by different colors. These four elements are calculated from the elements of the input matrix in the corresponding same-color regions and the weights. The output elements are calculated from the elements of the input matrix in the corresponding color regions and the weights, and the same applies to the other three colors.

[0030] As Figure 3 shown:

[0031] The input of the convolution operation input: N*Hi*Wi*Ci, the weight weight: K*K*Ci*Co, the output output: N*Ho*Wo*Co; The convolution operation in NHWC format is based on the two-dimensional convolution of Hi and Wi and is a process of cumulative summation in the Ci dimension.

[0032] Combining the system architecture of heterogeneous many-core processors and the characteristics of convolution operations themselves, without increasing additional memory occupancy, by redundantly reading K data and accessing memory once, K calculations can be performed, achieving a reduction of K memory accesses to 1 memory access. At the same time, the number of calculation results increases from 1 to multiple, improving the data reusability, greatly reducing the memory access pressure of the convolution operation on the system, and fully leveraging the computing resources of the processor, thereby improving the performance of the convolution operation on heterogeneous many-core processors.

[0033] The specific description is as follows:

[0034] 1. Expanded fetching: Usually, to obtain K*K input data (Hi_BLOCK = Wi_BLOCK = K) for one calculation result (Ho_BLOCK = 1 and Wo_BLOCK = 1). To calculate multiple output results simultaneously, let one more number be fetched in the Hi and Wi dimensions respectively. At this time, Hi_BLOCK = Wi_BLOCK = K + 1, so that 2*2 = 4 calculation results can be obtained through one memory access. At this time, Ho_BLOCK = Wo_BLOCK = 2;

[0035] Generally, let K - 1 more numbers be fetched in the Hi and Wi dimensions respectively. At this time, Hi_BLOCK = Wi_BLOCK = 2K - 1, so that K*K calculation results can be obtained through one memory access. At this time, Ho_BLOCK = Wo_BLOCK = K; It can be seen that as long as Hi_BLOCK and Wi_BLOCK are increased, more calculation results can be obtained through one memory access.

[0036] Therefore, on the premise of ensuring sufficient memory space, the present invention obtains larger Hi_BLOCK and Wi_BLOCK by redundantly acquiring data, so as to ensure that more results can be calculated through one memory access.

[0037] From Figure 2 it can be seen that by accessing memory twice to obtain data block 1 and data block 2 and calculating two results respectively, there is an overlap of 6 data between the two data blocks, and 6 identical elements are redundantly read during the two memory accesses, causing waste of limited bandwidth resources. If two data blocks can be read at one time (only 3 more elements than reading one data block), then one memory access is reduced.

[0038] Based on this idea, the present invention proposes the idea of expanded fetching. The schematic diagram of expanded fetching is as follows Figure 4 as shown, and will Figure 2The data blocks 1 and 2 obtained by two memory accesses are combined into one memory access operation. The same applies to data blocks 3 and 4. In this way, two results can be calculated in one memory access. Originally, four memory accesses were required, but now only two are needed. Compared with calculating one result per memory access before, the computing performance has been greatly improved.

[0039] 2. Data reusability: Each input data block Hi_BLOCK * Wi_BLOCK obtained by each CPU has the following characteristics: Hi_BLOCK > K and Wi_BLOCK > K. Since one output result is obtained from K * K input data, (K - 1) * (K - 1) of the data required for this calculation is the same as the data required for the previous calculation. Therefore, theoretically, as long as Hi_BLOCK and Wi_BLOCK each take one more number respectively, multiple computing outputs can be obtained.

[0040] Therefore, when the memory space is sufficient, the present invention tries to make Hi_BLOCK * Wi_BLOCK as large as possible, so that the data reusability will be higher, thus providing sufficient data for convolution calculation and improving the efficiency of convolution operation.

[0041] 3. Convolution operation: Each CPU performs convolution calculation based on the input data block A and the weight data block B it obtains, gets the output data block C, and writes the data block C back to the main memory.

[0042] 4. Repeat the above steps until all data blocks are calculated.

[0043] The above steps 1 and 2 are the core ideas of the present invention. Through the two strategies of inflated data fetching and maximizing data reusability, the actual demand for memory access in convolution operation is reduced, the number of memory accesses is decreased and the memory access pressure is alleviated. At the same time, it is ensured that convolution operation can obtain sufficient data for calculation, giving full play to the computing power of many cores and improving the efficiency of convolution operation.

[0044] When adopting the above convolution operation method based on inflated data fetching on a heterogeneous many-core architecture, without increasing the additional system memory occupation, through inflated data fetching, not only the memory access pressure of the system is reduced, but also the data reusability is improved, significantly reducing the memory access demand of convolution calculation on heterogeneous many-core processors, saving memory bandwidth resources. At the same time, it can make full use of the computing resources of many cores, give full play to the system performance, meet the characteristics of convolution calculation-intensive, and significantly improve the performance of convolution operation.

[0045] To facilitate a better understanding of the present invention, the terms used in this article will be briefly explained below:

[0046] Convolution: Mathematically, it is the result of multiplying two variables within a certain range and then summing them; in deep learning, it refers to the mathematical operation of sliding the input and the convolution kernel to multiply and sum to obtain the output.

[0047] Convolution kernel: The sliding window during the convolution operation, which is essentially a matrix.

[0048] CNN: Convolutional Neural Network, a type of feedforward neural network that contains convolutional calculations and has a deep structure.

[0049] Inflated fetching: Redundant data fetching from memory.

[0050] im2col: An operator for image conversion and storage.

[0051] Logical number: The id of the CPU itself (an integer).

[0052] The above embodiments are only used to illustrate the technical concept and features of the present invention. The purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. A convolution operation method based on dilated fetching on a heterogeneous many-core architecture, characterized in that, It includes the following steps: S1. Input input, weight, and stride, where input is Hi*Wi, weight is K*K. Calculate the shape of the output output based on the shapes of input and weight to obtain Ho*Wo; S2. According to the shape of output, in the Ho and Wo dimensions, evenly distribute the convolution calculation tasks to multiple cores according to the logical number of each core. Each core processes the calculation task with a size of Ho_BLOCK*Wo_BLOCK; S3. Each core calculates the required input size Hi_BLOCK*Wi_BLOCK according to its own task size, where Hi_BLOCK = Ho_BLOCK*stride + K - 1 and Wi_BLOCK = Wo_BLOCK*stride + K - 1; S4. Each core performs convolution calculation through the obtained input (Hi_BLOCK*Wi_BLOCK) and weight; S5. Repeat steps S3 and S4 until the tasks assigned to each core are calculated, and output output.

Citation Information

Patent Citations

  • Convolutional neural network acceleration method and system, terminal and storage medium

    CN111340185A

  • Convolutional neural network operation acceleration method and device based on many-core processor

    CN111461311A