Data mapping method and data mapping device for improving the utilization rate of a memory and computing array

By splitting the convolution kernel into multiple computed arrays and selecting appropriate data mapping methods based on the size relationship between the convolution kernel and the array, the problems of low data multiplexing and low array utilization caused by convolution kernel weight mapping in computed structure chips are solved, and efficient data multiplexing and calculation speed improvements are achieved.

CN115238877BActive Publication Date: 2025-07-01XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210875598.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-25
Publication Date
2025-07-01
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

When computing neural networks using memory structure chips, the weight mapping of the convolution kernel leads to low data multiplexing, increasing the pressure and data transmission volume of the input buffer, thereby increasing the chip delay and power consumption. At the same time, idle rows and columns inside the array cannot be effectively utilized, especially when the output channels of the convolution kernel are small, which is more prominent.

Method used

A data mapping method is proposed, which splits and maps the k×k convolution kernel into k accompaniment arrays, and maps a 1×k×C vector in each cell array. If the number of rows in the cell array is insufficient, it is mapped to the next cell array. According to the relationship between the convolution kernel size and the calculated array size, it is determined that a first data mapping method or a second data mapping method is used. The first data mapping method is suitable for the small convolution kernel, and data multiplexing is achieved by traveling empty in the cell array; the second data mapping method is suitable for the large convolution kernel, and data multiplexing is achieved by mapping the parameters of each input channel into different cell arrays separately.

Benefits of technology

It effectively improves the utilization rate and calculation speed of the computing array, reduces the pressure of the input buffer, reduces the amount of data transmission, and thus reduces chip delay and power consumption. Through this method, the calculation speed can be improved to three times while the number of calculation arrays is unchanged, effectively making use of idle rows and columns inside the array.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115238877B_ABST
    Figure CN115238877B_ABST
Patent Text Reader

Abstract

The present invention provides a data mapping method and a data mapping device for improving the utilization rate of a memory-computation array, relating to the field of chip technology. The method includes: splitting and mapping a k×k convolution kernel into k cell arrays, mapping a 1×k×C vector in each cell array, and if the number of rows of the cell array is not enough, mapping it to the next cell array; determining whether to use a first data mapping method or a second data mapping method for the memory-computation array according to the relationship between the convolution kernel size and the memory-computation array size. The present invention effectively reduces the pressure on the input buffer and the data transmission volume, thereby reducing the chip delay and power consumption. In addition, according to the size of the convolution kernel, the method provided by the present invention can increase the speed / area by up to three times, effectively utilizes the idle rows and columns inside each array, and improves the utilization rate of the memory-computation array and the computing speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chip technology, and in particular, to a data mapping method and a data mapping device for improving the utilization rate of a memory-computation array. Background Art

[0002] In a memory-computation chip or the memory-computation part of a chip, the hierarchical structure of the computing units from top to bottom is as follows: Tile, Processing Element, Array (XBar). There are several rows and columns inside the array, and there is a computing unit (bitcell) at the intersection of each row and column. Each computing unit can perform one calculation per cycle.

[0003] When using a memory-computation structure chip to calculate a neural network, the parameter matrix of the neural network is mapped to the array according to specific rules. One of the main purposes of weight mapping is to achieve the highest possible data reuse rate. Since in a common 3×3 convolution kernel, each number in the input matrix has to perform an average of 9 multiply-accumulate operations, the pressure on the input buffer is increased, and the data transmission volume is large, resulting in increased chip latency and power consumption. In addition, when performing multiple accelerations on a certain layer of the neural network, more processing units and arrays are often used for pipelined acceleration calculations, but the idle rows and columns inside each array cannot be effectively utilized. Especially when the output channels of the convolution kernel are small, this problem is more prominent. How to effectively utilize the idle rows and columns inside each array and improve the utilization rate of the memory-computation array and the computing speed is an urgent problem to be solved. Summary of the Invention

[0004] Based on the above problems, the present invention is proposed to provide a data mapping method and a data mapping device for improving the utilization rate of a memory-computation array that can overcome or at least partially solve the above problems.

[0005] In the first aspect of the embodiments of the present invention, a data mapping method for improving the utilization rate of a memory-computation array is provided. The memory-computation array includes: a plurality of unit arrays. The data mapping method includes:

[0006] Splitting and mapping a k×k convolution kernel into k unit arrays, and mapping a 1×k×C vector into each unit array. If the number of rows of the unit array is not enough, it is mapped to the next unit array, where k×k represents the height×width of the convolution kernel, and C represents the number of input channels of the convolution kernel;

[0007] Determine whether the memory-computation array uses a first data mapping method or a second data mapping method according to the relationship between the convolution kernel size and the memory-computation array size;

[0008] Taking k as 3, the first data mapping method includes:

[0009] Map the first row of the 3×3 convolution kernel to any cell array, where the elements in the first column of the cell array correspond to the arrangement after the first row of the 3×3 convolution kernel is expanded into a column vector;

[0010] In the second and third columns of the cell array, map the same weights as those in the first column of the cell array respectively, and the mapped positions of the elements in the second column of the cell array are vacated by C rows downward, and the mapped positions of the elements in the third column of the cell array are vacated by 2C rows downward;

[0011] Expand the input elements in the input buffer into a 1×5×C column vector;

[0012] In the current cycle, the input buffer outputs 5 elements in the input matrix to the cell array and performs a convolution operation;

[0013] At the end of the current cycle, a complete output element is obtained at the end of each column in the cell array, and then the convolution operation is performed in the next cycle;

[0014] Taking k as 3, the second data mapping method includes:

[0015] Map the parameters on each input channel of the 3×3 convolution kernel to different cell arrays in the memory and computing array respectively, and the three columns in each cell array map the same weights without empty rows, so that each input channel occupies ceil(C / n) cell arrays, a total of ceil(C / n)×3 arrays are occupied in the row dimension, and ceil(M×3 / n) arrays are occupied in the column dimension, where ceil is the ceiling function and n represents the size of the memory and computing array;

[0016] In the current cycle, the input buffer outputs 3 elements in the input matrix to three cell arrays in the memory and computing array and performs a convolution operation to obtain partial results;

[0017] Group and add the partial results obtained in each cycle to obtain the result of the convolution.

[0018] Optionally, the size of the convolution kernel is k×k×C×M, and the size of the memory and computing array is n×n, where M represents the number of output channels of the convolution kernel;

[0019] Determine whether the memory and computing array uses the first data mapping method or the second data mapping method according to the relationship between the size of the convolution kernel and the size of the memory and computing array, including:

[0020] When the convolution kernel is split and mapped to any one of the cell arrays, if the number of rows occupied in any one of the cell arrays is not greater than k / (2k - 1) of the size of the cell array, or the number of columns occupied in any one of the cell arrays is less than 1 / k of the size of the cell array, it is determined that the memory-computation array uses the first data mapping method;

[0021] When the convolution kernel is split and mapped to any one of the cell arrays, if the number of rows occupied in any one of the cell arrays is not greater than k / (2k - 1) of the size of the cell array and the number of columns occupied in any one of the cell arrays is less than 1 / k of the size of the cell array, it is determined that the memory-computation array uses the first data mapping method;

[0022] When the convolution kernel is split and mapped to any one of the cell arrays, if the number of rows occupied in any one of the cell arrays is not less than k times the size of the cell array, then based on the convolution kernel size and the memory-computation array size, the ratio of the first occupied array to the computing speed and the ratio of the second occupied array to the computing speed are calculated. The ratio of the first occupied array to the computing speed is the ratio of the occupied array to the computing speed when the convolution kernel is mapped to the memory-computation array using the traditional method. The ratio of the second occupied array to the computing speed is the ratio of the occupied array to the computing speed obtained by mapping the parameters on each input channel of the convolution kernel to different cell arrays in the memory-computation array and mapping the same weights to k columns in each cell array without empty rows;

[0023] If the value of the ratio of the first occupied array to the computing speed is greater than the value of the ratio of the second occupied array to the computing speed, it is determined that the memory-computation array uses the second data mapping method.

[0024] Optionally, the formula for the ratio of the first occupied array to the computing speed is:

[0025] k × ceil(k × C / n) × ceil(M / n)

[0026] The formula for the ratio of the second occupied array to the computing speed is:

[0027] Ceil(C / n) × k × ceil(M × k / n)

[0028] In the above two formulas, ceil(k×C / n) represents the number of cell arrays occupied by the convolution kernel mapping in the row dimension in the traditional method, ceil(M / n) represents the number of cell arrays occupied by the convolution kernel mapping in the column dimension in the traditional method, ceil(C / n) represents the number of cell arrays occupied by each input channel, ceil(C / n)×k represents the number of cell arrays occupied in the row dimension, and ceil(M×3 / n) represents the number of cell arrays occupied in the column dimension.

[0029] Optionally, in the current cycle, the input buffer outputs 5 elements in the input matrix to the cell array for convolution operation, including:

[0030] Among the 5 elements, the first two elements were partially calculated in the previous cycle and the remaining calculations are completed in the current cycle. The middle element can complete all 3 calculations by being input once in the current cycle. The last two elements complete part of the calculations in the current cycle, and the remaining calculations are completed by placing the last two elements in the positions of the first two elements in the next cycle.

[0031] Optionally, at the end of the current cycle, a complete output element is obtained at the end of each column in the cell array, and then the convolution operation for the next cycle is performed, including:

[0032] At the end of the current cycle, a complete output element is obtained at the end of each column in the cell array. Then, at the beginning of the next cycle, the value-taking area of the input buffer on the input matrix is shifted 3 pixels to the right based on the value-taking area of the input buffer on the input matrix in the current cycle, and then 5 elements are taken and output to the cell array for convolution operation.

[0033] Optionally, the three cell arrays include: the first cell array, the second cell array, and the third cell array;

[0034] Group and add the partial results obtained in each cycle to get the convolution result, including:

[0035] Add the result obtained at the end of the first column in the first cell array, the result obtained at the end of the first column in the second cell array, and the result obtained at the end of the first column in the third cell array in the current cycle to get a complete output element;

[0036] Add the results obtained at the end of the second column, the results obtained at the end of the third column in the first unit array, and the results obtained at the end of the third column in the second unit array in the current cycle to the partial results obtained in the previous cycle respectively to obtain a complete output element;

[0037] Add the results obtained at the end of the second column in the second unit array, the results obtained at the end of the second column in the third unit array, and the results obtained at the end of the third column in the third unit array in the current cycle to the partial results obtained in the next cycle respectively to obtain a complete output element.

[0038] The second aspect of the embodiments of the present invention provides a data mapping device for improving the utilization rate of a memory-computation array. The data mapping device includes:

[0039] A splitting and mapping module, configured to split and map a k×k convolution kernel into k memory-computation arrays, and map a 1×k×C vector into each memory-computation array. If the number of rows of the memory-computation array is not enough, it is mapped to the next memory-computation array, where k×k represents the height×width of the C convolution kernel, and C represents the number of input channels of the convolution kernel;

[0040] A determination method module, configured to determine whether the memory-computation array uses a first data mapping method or a second data mapping method according to the relationship between the convolution kernel size and the memory-computation array size;

[0041] Taking k as 3, the data mapping device further includes: a first method module and a second method module;

[0042] Among them, the first method module includes:

[0043] An unfolding and mapping unit, configured to map the first row of the 3×3 convolution kernel to any unit array in the memory-computation array, and the elements in the first column of this unit array correspond to the arrangement mode after the first row of the 3×3 convolution kernel is unfolded into a column vector;

[0044] An empty row mapping unit, configured to map the same weights as those in the first column in the second column and the third column in the unit array, and the mapping positions corresponding to the elements in the second column of the unit array are emptied by C rows downward, and the mapping positions corresponding to the elements in the third column of the unit array are emptied by 2C rows downward;

[0045] An extended input unit, configured to extend the input data in the input buffer into a 1×5×C column vector;

[0046] A first convolution unit, configured to, in the current cycle, the input buffer outputs 5 elements in the input matrix to the unit array and performs a convolution operation;

[0047] A value output unit, which is used to obtain a complete output element at the end of each column in the cell array at the end of the current cycle, and then perform the next-cycle convolution operation;

[0048] The second method module includes:

[0049] A full-row mapping unit, which is used to map the parameters on each input channel of the 3×3 convolution kernel to different cell arrays in the memory-computation array respectively, and the three columns in each cell array map the same weights without empty rows, so that each input channel occupies ceil(C / n) cell arrays, and a total of ceil(C / n)×3 arrays are occupied in the row dimension, while ceil(M×3 / n) arrays are occupied in the column dimension, where ceil is the ceiling function and n represents the size of the memory-computation array;

[0050] A second convolution unit, which is used to output 3 elements in the input matrix to three cell arrays in the memory-computation array and perform convolution operations to obtain partial results during the current cycle;

[0051] An addition unit, which is used to group and add the partial results obtained in each cycle to obtain the result of convolution.

[0052] Optionally, the size of the convolution kernel is k×k×C×M, and the size of the memory-computation array is n×n, where M represents the number of output channels of the convolution kernel;

[0053] The determination method module is specifically used for:

[0054] When the convolution kernel is split and mapped to any cell array, if the number of rows occupying any cell array is not greater than k / 2k - 1 of the size of the cell array, or the number of columns occupying any cell array is less than 1 / k of the size of the cell array, it is determined that the memory-computation array uses the first data mapping method;

[0055] When the convolution kernel is split and mapped to any cell array, if the number of rows occupying any cell array is not greater than k / 2k - 1 of the size of the cell array and the number of columns occupying any cell array is less than 1 / k of the size of the cell array, it is determined that the memory-computation array uses the first data mapping method;

[0056] When the convolutional kernel is split and mapped to any of the unit arrays, if the number of rows occupied in any of the unit arrays is not less than k times the size of the unit array, then based on the size of the convolutional kernel and the size of the memory-computation array, the ratio of the first occupied array to the computing speed and the ratio of the second occupied array to the computing speed are calculated. The ratio of the first occupied array to the computing speed is the ratio of the occupied array to the computing speed when the convolutional kernel is mapped to the memory-computation array using the traditional method. The ratio of the second occupied array to the computing speed is the ratio of the occupied array to the computing speed obtained by mapping the parameters on each input channel of the convolutional kernel to different unit arrays in the memory-computation array, and mapping the same weights to the three columns in each unit array without empty rows.

[0057] If the value of the ratio of the first occupied array to the computing speed is greater than the value of the ratio of the second occupied array to the computing speed, then it is determined that the second data mapping method is used for the memory-computation array.

[0058] Optionally, the first convolutional unit is specifically configured to:

[0059] For the first two elements among the 5 elements, partial calculations were performed in the previous cycle, and the remaining calculations are completed in the current cycle. The middle element can complete all 3 calculations with a single input in the current cycle. For the last two elements, partial calculations are completed in the current cycle, and the remaining calculations are completed by placing the last two elements in the positions of the first two elements in the next cycle.

[0060] Optionally, the value output unit is specifically configured to:

[0061] At the end of the current cycle, a complete output element is obtained at the end of each column in the unit array. Then, at the beginning of the next cycle, the value-taking area of the input buffer on the input matrix is shifted 3 pixels to the right based on the value-taking area of the input buffer on the input matrix in the current cycle. Then, 5 elements are taken and output to the unit array for convolutional operation.

[0062] Optionally, the three unit arrays include: a first unit array, a second unit array, and a third unit array;

[0063] The addition unit is specifically configured to:

[0064] Add the result obtained at the end of the first column in the first unit array, the result obtained at the end of the first column in the second unit array, and the result obtained at the end of the first column in the third unit array in the current cycle to obtain a complete output element;

[0065] Add the results obtained at the end of the second column in the first unit array, the results obtained at the end of the third column in the first unit array, and the results obtained at the end of the third column in the second unit array in the current cycle to the partial results obtained in the previous cycle respectively to obtain a complete output element.

[0066] Add the results obtained at the end of the second column in the second unit array, the results obtained at the end of the second column in the third unit array, and the results obtained at the end of the third column in the third unit array in the current cycle to the partial results obtained in the next cycle respectively to obtain a complete output element.

[0067] The data mapping method for improving the utilization rate of the memory-computation array provided by the present invention splits a k×k convolution kernel and maps it into k memory-computation arrays. A 1×k×C vector is mapped into each memory-computation array. If the number of rows of the memory-computation array is not enough, it is mapped into the next memory-computation array. According to the relationship between the convolution kernel size and the memory-computation array size, it is determined whether the memory-computation array uses the first data mapping method or the second data mapping method.

[0068] Taking a 3×3 convolution kernel as an example, when the convolution kernel is small, the first data mapping method is used. 80% of the input data needs to be input a total of 2 times, and the remaining 20% only needs to be input 1 time, with an average reuse rate of 1.8. When the convolution kernel is large, the second data mapping method is used, and all input data only needs to be input 1 time, with an average reuse rate of 3. Both of these data mapping methods can effectively reduce the pressure on the input buffer and reduce the data transmission volume, thereby reducing the chip delay and power consumption.

[0069] In addition, according to the size of the convolution kernel, applying the method proposed by the present invention can increase the ratio of the occupied array to the computing speed (abbreviation: speed / area) by up to three times. When using the second mapping method, it will occupy about 3 times the computing array, and at the same time, the computing speed is also increased to 3 times, so the speed / area remains unchanged. When the input or output channels of the convolution kernel are very small, specifically, the number of occupied rows increases to 5 / 3 and the number of occupied columns increases to 3 times, and it still does not exceed the size of a single array. In this case, under the condition that the number of computing arrays remains unchanged, the computing speed is increased to 3 times, so the speed / area is also increased to 3 times, effectively utilizing the idle rows and columns inside each array and improving the utilization rate of the memory-computation array and the computing speed. Description of the Drawings

[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0071] Figure 1 It is a simulation diagram of the mapping method when the convolution kernel is small in the embodiments of the present invention;

[0072] Figure 2 It is a simulation diagram of the mapping method when the convolution kernel is large in the embodiments of the present invention. Specific Embodiments

[0073] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0074] The inventors found that the current method of performing convolution operations using a memory-compute array has problems such as high power consumption and low speed / area ratio. Taking M convolution kernels of size 3×3×C as an example, each number in the input matrix has to perform an average of 9 multiply-accumulate operations. Currently, the common mapping methods for arranging the weight matrix and the flow method of input data are as follows:

[0075] One is to expand each 3×3×C convolution kernel into a one-dimensional column vector and arrange them in the same column, with different columns mapping different convolution kernels; the second is to split the 3×3 convolution kernel and map it into 9 unit arrays, and each column of each unit array maps a 1×C vector; the third is to split the 3×3 convolution kernel and map it into 3 unit arrays, and each unit array maps a 1×3×C column vector. However, these three methods still have the following two obvious disadvantages:

[0076] 1) None of them can achieve the reuse of input data within the array. As mentioned above, in a 3×3 convolution kernel, each input data needs to perform an average of 9 multiply-accumulate operations on average. Therefore, the above methods still need to input each data 9 times into different unit arrays or different rows of the same unit array. The reuse rate of input data is insufficient.

[0077] 2) When performing multiple accelerations on a certain layer of the neural network, more processing units and arrays are often used for accelerated calculations during pipelining. However, the idle rows and columns within each unit array cannot be effectively utilized, especially when the output channels of the convolution kernel are small.

[0078] In view of the above problems, the inventors creatively proposed a data mapping method for improving the utilization rate of the memory-compute array in the present invention. The method of the present invention will be described in detail below.

[0079] Generally, the memory - in - computing array is composed of multiple cell arrays. When applying the data mapping method of the present invention, first, the k×k convolution kernel is split and mapped into k cell arrays, and a 1×k×C column vector is mapped into each cell array. If the number of rows of the cell array is not enough, it continues to be mapped into the next cell array, where k×k represents the height×width of the convolution kernel, and C represents the number of input channels of the convolution kernel.

[0080] After that, according to the relationship between the convolution kernel size and the memory - in - computing array size, it is determined whether the memory - in - computing array uses the first data mapping method or the second data mapping method. The specific way to determine which data mapping method to use can be as follows:

[0081] Suppose the size of the convolution kernel is k (height of the convolution kernel)×k (width of the convolution kernel)×C (number of input channels of the convolution kernel)×M (number of output channels of the convolution kernel), and the size of the memory - in - computing array is n (number of rows of the memory - in - computing array)×n (number of columns of the memory - in - computing array).

[0082] When the convolution kernel is split and mapped to any cell array, if the number of rows occupied by any cell array is no more than k / 2k - 1 of the size of that cell array, or the number of columns occupied by any cell array is less than 1 / k of the size of that cell array, it is determined that the memory - in - computing array uses the first data mapping method.

[0083] If the number of rows occupied by any cell array is no more than k / 2k - 1 of the size of that cell array, but the number of columns occupied by any cell array is not less than 1 / k of the size of that cell array, then the speed / area corresponding to the data mapping method of the present invention can achieve a reduction of 1 to k / 2 times compared to the speed / area corresponding to the existing mapping method described above; if the number of columns occupied by any cell array is less than 1 / k of the size of that cell array, but the number of rows occupied by any cell array is greater than k / 2k - 1 of the size of that cell array, then the speed / area corresponding to the data mapping method of the present invention can achieve a reduction of k / 2 to k times compared to the speed / area corresponding to the existing mapping method described above.

[0084] When the convolution kernel is split and mapped to any cell array, if the number of rows occupied by any cell array is no more than k / 2k - 1 of the size of that cell array, and the number of columns occupied by any cell array is less than 1 / k of the size of that cell array, it is also determined that the memory - in - computing array uses the first data mapping method. In this case, a reduction of k times in speed / area can be fully achieved.

[0085] The foregoing method essentially targets the case when the convolution kernel is small. When the convolution kernel needs to occupy multiple cell arrays to completely map all data, it is the case of a large convolution kernel. In this case, when the convolution kernel is split and mapped to any cell array, the number of rows occupied by any cell array is not less than k times the size of the cell array. Then, based on the convolution kernel size and the memory-computation array size, the ratio of the first occupied array to the computing speed and the ratio of the second occupied array to the computing speed are calculated, that is, the first speed / area and the second speed / area are calculated.

[0086] The so-called ratio of the first occupied array to the computing speed refers to the ratio of the occupied array to the computing speed when mapping the convolution kernel to the memory-computation array using the traditional method; the so-called ratio of the second occupied array to the computing speed refers to the ratio of the occupied array to the computing speed obtained by mapping the parameters on each input channel of the convolution kernel to different cell arrays in the memory-computation array respectively, and mapping the same weights to k columns in each cell array without empty rows. The meaning of the empty row will be explained below and will not be elaborated here.

[0087] After obtaining the ratio of the first occupied array to the computing speed and the ratio of the second occupied array to the computing speed, if the value of the ratio of the first occupied array to the computing speed is greater than the value of the ratio of the second occupied array to the computing speed, it is determined that the second data mapping method is used for the memory-computation array.

[0088] Among them, the formula for the ratio of the first occupied array to the computing speed can be:

[0089] k×ceil(k×C / n)×ceil(M / n)

[0090] The formula for the ratio of the second occupied array to the computing speed can be:

[0091] Ceil(C / n)×k×ceil(M×k / n)

[0092] In the above two formulas, ceil(k×C / n) represents that in the traditional method, the convolution kernel needs to occupy ceil(k×C / n) cell arrays in the row dimension, ceil(M / n) represents that in the traditional method, the convolution kernel needs to occupy ceil(M / n) cell arrays in the column dimension, ceil(C / n) represents that in the method of the present invention, each input channel occupies ceil(C / n) cell arrays, ceil(C / n)×k represents that in the method of the present invention, it occupies ceil(C / n)×k cell arrays in the row dimension, and ceil(M×3 / n) represents that in the method of the present invention, it occupies ceil(M×k / n) cell arrays in the column dimension.

[0093] If the value calculated by k × ceil(k × C / n) × ceil(M / n) is greater than the value calculated by Ceil(C / n) × k × ceil(M × k / n), the data mapping method of the present invention is used.

[0094] Taking a 3×3 convolution kernel as an example, the content of the first data mapping method and the second data mapping method in the embodiments of the present invention will be described below.

[0095] Combined with Figure 1 The simulation diagram of the mapping method when the convolution kernel is small as shown, the first data mapping method includes: mapping the first row of the 3×3 convolution kernel to any unit array, and the elements in the first column of the unit array correspond to the arrangement mode after the first row of the 3×3 convolution kernel is expanded into a column vector. For example Figure 1 As shown, the 3×3 convolution kernel exemplarily shows X1~Xc, Y1~Yc, Z1~Zc; the matrix data INPUT in the input matrix 01 、INPUT 02 ,INPUT 01 includes: IN 01 ~IN 0c 、IN 11 ~IN 1c 、IN 21 ~IN 2c 、IN 31 ~IN 3c 。

[0096] In the first cycle, the input buffer will input the corresponding matrix data INPUT 01 into the memory-computation array and calculate the first element OUT 01 of the output matrix. In the current traditional mapping method, the input buffer will input INPUT 02 in the next cycle and obtain the second element OUT 02 of the output matrix at the end of the first column of the memory-computation array. However, as shown in Figure 1 the upper left corner, two-thirds of the input data in adjacent two cycles is repeated. In order to effectively utilize these repeatedly input data in the first cycle, the same weights as the first column are respectively mapped in the second column and the third column of the memory-computation array, but the mapped positions are successively downward, leaving C rows empty each time. That is, the position of X1 in the second column is aligned with the position of Y1 in the first column, and the position of X1 in the third column is aligned with the position of Y1 in the second column and aligned with the position of Z1 in the first column.

[0097] Expand the input elements of the input buffer into a column vector of 1×5×C, so that three elements OUT 01 、OUT 02 、OUT03 , the computing speed is increased to three times. It can be seen that among the input data, one-fifth of the data is calculated three times, two-fifths of the data is calculated twice, and the remaining two-fifths of the data is calculated once. The average reuse rate of the data is 1.8. At the same time, the price paid is that the number of rows occupied inside the array is increased to 5 / 3 of the original, the number of columns occupied is expanded to 3 times, and the amount of data input by the input buffer per cycle is increased to 5 / 3.

[0098] That is, for the 5 input elements, the first two elements have been partially calculated in the previous cycle, and the remaining calculations are completed in this cycle; the middle element only needs to be input once in this cycle to complete all 3 calculations; the last two elements complete part of the calculations in this cycle, and the remaining calculations are completed by placing these two elements in the positions of the first two elements in the next cycle. Under this data mapping method, after each cycle, a complete output element can be obtained at the end of each column. Then, the value-taking area of the input buffer on the input matrix is shifted 3 pixels to the right. Therefore, under this data mapping method, the amount of data input by the input buffer per cycle is increased to 5 / 3 times, but the number of calculation cycles becomes one-third of the original.

[0099] Combined with Figure 2 the simulation diagram of the mapping method when the convolution kernel is large as shown, the second data mapping method includes: mapping the parameters on each input channel of the 3×3 convolution kernel to different cell arrays in the memory-computation array, and the three columns in each cell array map the same weights without empty rows. For example Figure 2 as shown, the 3×3 convolution kernel exemplarily shows X1~Xc, Y1~Yc, Z1~Zc; the matrix data INPUT in the input matrix m1 , INPUT m1 includes: IN m1 ~IN mc , IN (m+1)1 ~IN (m+1)c , IN (m+2)1 ~IN (m+2)c .

[0100] Map the parameters on each input channel to three cell arrays respectively. The three columns in each cell array map the same weights without empty lines. The parameters mapped by each column in the three cell arrays are different from those in other cell arrays. For example, in the first cell array at the top of the storage array, the parameters mapped by the first column are X1 to Xc. Then, in the second cell array in the middle, the parameters mapped by the first column are Y1 to Yc. In the third cell array at the bottom, the parameters mapped by the first column are Z1 to Zc. And the output part of the result is at the end of each column of each cell array. Group and add the partial results obtained in each period to get the result of convolution. In this case, the average reuse rate of the data reaches 3, and compared with the first data mapping, it does not occupy extra rows, and the input buffer does not need to input extra data per cycle.

[0101] That is, add the result OUT01 obtained at the end of the first column of the first cell array in the current period, the result OUT04 obtained at the end of the first column of the second cell array, and the result OUT07 obtained at the end of the first column of the third cell array to get a complete output element.

[0102] Add the result OUT02 obtained at the end of the second column of the first cell array in the current period, the result OUT03 obtained at the end of the third column of the first cell array, and the result OUT06 obtained at the end of the third column of the second cell array to the partial results obtained in the previous period respectively to get a complete output element.

[0103] Add the result OUT05 obtained at the end of the second column of the second cell array in the current period, the result OUT08 obtained at the end of the second column of the third cell array, and the result OUT09 obtained at the end of the third column of the third cell array to the partial results obtained in the next period respectively to get a complete output element. It can be seen that each element in the input matrix only needs to be input once to complete all calculations. The input data of the input buffer remains unchanged per cycle, and the number of calculation cycles becomes one-third of the original.

[0104] Taking a 3×3 convolution kernel as an example, when the convolution kernel is small, using the first data mapping method, 80% of the input data needs to be input a total of 2 times, and the remaining 20% only needs to be input 1 time, and the average reuse rate is 1.8. When the convolution kernel is large, using the second data mapping method, all input data only needs to be input 1 time, and the average reuse rate is 3. Both of these data mapping methods can effectively reduce the pressure on the input buffer, reduce the data transmission volume, and thus reduce the chip delay and power consumption.

[0105] In addition, according to the size of the convolution kernel, the method proposed in the present invention can increase the speed / area by up to three times. When using the second mapping method, it will occupy about three times the computing array, and at the same time, the computing speed is also increased to three times, so the speed / area remains unchanged. When the input or output channels of the convolution kernel are very small, specifically, the number of rows it occupies increases to 5 / 3, and the number of columns it occupies increases to three times, and it still does not exceed the size of a single array. In this case, with the number of computing arrays unchanged, the computing speed is increased to three times, so the speed / area is also increased to three times, effectively utilizing the idle rows and columns inside each array, and improving the utilization rate of the memory-computation array and the computing speed.

[0106] Based on the above data mapping method for improving the utilization rate of the memory-computation array, an embodiment of the present invention further provides a data mapping device for improving the utilization rate of the memory-computation array. The data mapping device includes:

[0107] A split mapping module, configured to split and map a k×k convolution kernel into k memory-computation arrays, and map a 1×k×C vector into each memory-computation array. If the number of rows of the memory-computation array is not enough, it is mapped to the next memory-computation array, where k×k represents the height×width of the C convolution kernel, and C represents the number of input channels of the convolution kernel;

[0108] A determination method module, configured to determine whether the memory-computation array uses the first data mapping method or the second data mapping method according to the relationship between the convolution kernel size and the memory-computation array size;

[0109] Taking k as 3 as an example, the data mapping device further includes: a first method module and a second method module;

[0110] Among them, the first method module includes:

[0111] An expansion mapping unit, configured to map the first row of the 3×3 convolution kernel to any unit array in the memory-computation array, and the elements in the first column of this unit array correspond to the arrangement after the first row of the 3×3 convolution kernel is expanded into a column vector;

[0112] An empty row mapping unit, configured to map the same weights as those in the first column in the second column and the third column of the unit array, and the mapping position corresponding to the elements in the second column of the unit array is vacated by C rows downward, and the mapping position corresponding to the elements in the third column of the unit array is vacated by 2C rows downward;

[0113] An extended input unit, configured to extend the input data in the input buffer into a 1×5×C column vector;

[0114] A first convolution unit, configured to, in the current cycle, the input buffer outputs 5 elements in the input matrix to the unit array for convolution operation;

[0115] An output unit for obtaining a complete output element at the end of each column in the cell array at the end of the current cycle, and then performing the next cycle of convolution operation;

[0116] The second method module includes:

[0117] A full-row mapping unit for mapping the parameters on each input channel of the 3×3 convolution kernel to different cell arrays in the memory-computation array respectively, and mapping the same weights to three columns in each cell array without empty rows, so that each input channel occupies ceil(C / n) cell arrays, a total of ceil(C / n)×3 arrays in the row dimension, and ceil(M×3 / n) arrays in the column dimension, where ceil is the ceiling function and n represents the size of the memory-computation array;

[0118] A second convolution unit for, in the current cycle, the input buffer outputting 3 elements in the input matrix to three cell arrays in the memory-computation array and performing convolution operation to obtain partial results;

[0119] An addition unit for grouping and adding the partial results obtained in each cycle to obtain the convolution result.

[0120] Optionally, the size of the convolution kernel is k×k×C×M, and the size of the memory-computation array is n×n, where M represents the number of output channels of the convolution kernel;

[0121] The determination method module is specifically used for:

[0122] When the convolution kernel is split and mapped to any one of the cell arrays, if the number of rows occupied by any one of the cell arrays is not greater than k / 2k - 1 of the size of the cell array, or the number of columns occupied by any one of the cell arrays is less than 1 / k of the size of the cell array, it is determined that the memory-computation array uses the first data mapping method;

[0123] When the convolution kernel is split and mapped to any one of the cell arrays, if the number of rows occupied by any one of the cell arrays is not greater than k / 2k - 1 of the size of the cell array and the number of columns occupied by any one of the cell arrays is less than 1 / k of the size of the cell array, it is determined that the memory-computation array uses the first data mapping method;

[0124] When the convolutional kernel is split and mapped to any of the unit arrays, if the number of rows occupied in any of the unit arrays is not less than k times the size of the unit array, then based on the size of the convolutional kernel and the size of the memory-computation array, the ratio of the first occupied array to the computing speed and the ratio of the second occupied array to the computing speed are calculated. The ratio of the first occupied array to the computing speed is the ratio of the occupied array to the computing speed when the convolutional kernel is mapped to the memory-computation array using the traditional method. The ratio of the second occupied array to the computing speed is the ratio of the occupied array to the computing speed obtained by mapping the parameters on each input channel of the convolutional kernel to different unit arrays in the memory-computation array respectively, and mapping the same weights to the three columns in each unit array without empty rows.

[0125] If the value of the ratio of the first occupied array to the computing speed is greater than the value of the ratio of the second occupied array to the computing speed, then it is determined that the second data mapping method is used for the memory-computation array.

[0126] Optionally, the first convolutional unit is specifically configured to:

[0127] For the first two elements among the 5 elements, partial calculations were performed in the previous cycle, and the remaining calculations are completed in the current cycle. The middle element can complete all 3 calculations with one input in the current cycle. For the last two elements, partial calculations are completed in the current cycle, and the remaining calculations are completed by placing the last two elements in the positions of the first two elements in the next cycle.

[0128] Optionally, the value output unit is specifically configured to:

[0129] At the end of the current cycle, a complete output element is obtained at the end of each column in the unit array. Then, at the beginning of the next cycle, the value-taking area of the input buffer on the input matrix is shifted 3 pixels to the right based on the value-taking area of the input buffer on the input matrix in the current cycle. Then, 5 elements are taken and output to the unit array for convolutional operation.

[0130] Optionally, the three unit arrays include: a first unit array, a second unit array, and a third unit array;

[0131] The addition unit is specifically configured to:

[0132] Add the result obtained at the end of the first column in the first unit array, the result obtained at the end of the first column in the second unit array, and the result obtained at the end of the first column in the third unit array in the current cycle to obtain a complete output element;

[0133] Add the results obtained at the end of the second column in the first unit array, the results obtained at the end of the third column in the first unit array, and the results obtained at the end of the third column in the second unit array in the current cycle to the partial results obtained in the previous cycle respectively to obtain a complete output element;

[0134] Add the results obtained at the end of the second column in the second unit array, the results obtained at the end of the second column in the third unit array, and the results obtained at the end of the third column in the third unit array in the current cycle to the partial results obtained in the next cycle respectively to obtain a complete output element.

[0135] Through the above example, in the data mapping method for improving the utilization rate of the memory - computing array of the present invention, a k×k convolution kernel is split and mapped into k memory - computing arrays, and a 1×k×C vector is mapped into each memory - computing array. If the number of rows of the memory - computing array is not enough, it is mapped into the next memory - computing array. According to the relationship between the convolution kernel size and the memory - computing array size, it is determined whether the memory - computing array uses the first data mapping method or the second data mapping method.

[0136] Taking a 3×3 convolution kernel as an example, when the convolution kernel is small, the first data mapping method is used. 80% of the input data needs to be input a total of 2 times, and the remaining 20% only needs to be input 1 time, with an average reuse rate of 1.8. When the convolution kernel is large, the second data mapping method is used, and all input data only needs to be input 1 time, with an average reuse rate of 3. Both of these data mapping methods can effectively relieve the pressure on the input buffer, reduce the data transmission volume, thereby reducing the chip delay and power consumption.

[0137] In addition, according to the size of the convolution kernel, applying the method proposed in the present invention can increase the ratio of the occupied array to the computing speed (abbreviation: speed / area) by up to three times. When using the second mapping method, it will occupy about 3 times the computing array, and at the same time, the computing speed is also increased to 3 times, so the speed / area remains unchanged. When the input or output channels of the convolution kernel are very small, specifically, the number of occupied rows increases to 5 / 3 and the number of occupied columns increases to 3 times, still not exceeding the size of a single array. In this case, with the number of computing arrays unchanged, the computing speed is increased to 3 times, so the speed / area is also increased to 3 times, effectively utilizing the idle rows and columns inside each array, and improving the utilization rate of the memory - computing array and the computing speed.

[0138] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.

[0139] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.

Claims

1. A data mapping method for improving the utilization rate of a memory-computation array, characterized in that, The memory - in - computing array includes: a plurality of cell arrays, and the data mapping method includes: Split - map a k×k convolution kernel into k cell arrays, and map a 1×k×C vector into each cell array. If the number of rows of a cell array is not enough, map it to the next cell array, where k×k represents the height×width of the convolution kernel, and C represents the number of input channels of the convolution kernel; Determine whether the memory - in - computing array uses the first data mapping method or the second data mapping method according to the relationship between the convolution kernel size and the memory - in - computing array size; The first data mapping method includes: When k = 3, map the first row of the 3×3 convolution kernel to any cell array, and the elements in the first column of this cell array correspond to the arrangement of the first row of the 3×3 convolution kernel expanded into a column vector; In the second and third columns of the cell array, map the same weights as those in the first column of the cell array respectively. And the mapping positions corresponding to the elements in the second column of the cell array are vacated by C rows downward, and the mapping positions corresponding to the elements in the third column of the cell array are vacated by 2C rows downward; Expand the input elements in the input buffer into a 1×5×C column vector; In the current cycle, the input buffer outputs 5 elements in the input matrix to the cell array for convolution operation; At the end of the current cycle, a complete output element is obtained at the end of each column in the cell array, and then the convolution operation of the next cycle is performed; The second data mapping method includes: When k = 3, map the parameters on each input channel of the 3×3 convolution kernel to different cell arrays in the memory - in - computing array respectively, and the three columns in each cell array map the same weights without empty rows, so that each input channel occupies ceil(C / n) cell arrays, occupies ceil(C / n)×3 arrays in the row dimension in total, and occupies ceil(M×3 / n) arrays in the column dimension, where ceil is the ceiling function and n represents the size of the memory - in - computing array; In the current cycle, the input buffer outputs 3 elements in the input matrix to three cell arrays in the memory - in - computing array for convolution operation to obtain partial results; Group and add the partial results obtained in each cycle to get the result of convolution.

2. The data mapping method according to claim 1, wherein The size of the convolution kernel is k×k×C×M, and the size of the memory - in - computing array is n×n, where M represents the number of output channels of the convolution kernel; Determine whether the memory - in - computing array uses the first data mapping method or the second data mapping method according to the relationship between the convolution kernel size and the memory - in - computing array size, including: When the convolution kernel is split - mapped to any cell array, if the number of rows occupied by it is not greater than k / (2k - 1) of the size of this cell array, or the number of columns occupied by it is less than 1 / k of the size of this cell array, determine that the memory - in - computing array uses the first data mapping method; When the convolution kernel is split and mapped to any one of the cell arrays, if the number of rows occupied in any one of the cell arrays is not greater than k / (2k - 1) of the size of the cell array, and the number of columns occupied in any one of the cell arrays is less than 1 / k of the size of the cell array, it is determined that the memory - computing array uses the first data mapping method; When the convolution kernel is split and mapped to any one of the cell arrays, if the number of rows occupied in any one of the cell arrays is not less than k times the size of the cell array, then based on the convolution kernel size and the memory - computing array size, the ratio of the first occupied array to the computing speed and the ratio of the second occupied array to the computing speed are calculated. The ratio of the first occupied array to the computing speed is the ratio of the occupied array to the computing speed when the convolution kernel is mapped to the memory - computing array using the traditional method. The ratio of the second occupied array to the computing speed is the ratio of the occupied array to the computing speed obtained by mapping the parameters on each input channel of the convolution kernel to different cell arrays in the memory - computing array, and mapping the k columns in each cell array to the same weights without empty rows; If the value of the ratio of the first occupied array to the computing speed is greater than the value of the ratio of the second occupied array to the computing speed, it is determined that the memory - computing array uses the second data mapping method.

3. The data mapping method according to claim 2, wherein The formula for the ratio of the first occupied array to the computing speed is: k×ceil(k×C / n)×ceil(M / n) The formula for the ratio of the second occupied array to the computing speed is: Ceil(C / n)×k×ceil(M×k / n) In the above two formulas, ceil(k×C / n) represents that in the traditional method, the convolution kernel needs to occupy ceil(k×C / n) cell arrays in the row dimension, ceil(M / n) represents that in the traditional method, the convolution kernel needs to occupy ceil(M / n) cell arrays in the column dimension, ceil(C / n) represents that each input channel occupies ceil(C / n) cell arrays, ceil(C / n)×k represents that in the row dimension, it occupies ceil(C / n)×k cell arrays, and ceil(M×3 / n) represents that in the column dimension, it occupies ceil(M×k / n) cell arrays.

4. The data mapping method according to claim 1, wherein In the current cycle, the input buffer outputs 5 elements in the input matrix to the cell array for convolution operation, including: Among the 5 elements, the first two elements were partially calculated in the previous cycle and the remaining calculations are completed in the current cycle. The middle element can complete all 3 calculations by being input once in the current cycle. The last two elements complete part of the calculations in the current cycle, and the remaining calculations are completed by placing the last two elements in the positions of the first two elements in the next cycle.

5. The data mapping method according to claim 4, wherein At the end of the current cycle, a complete output element is obtained at the end of each column in the cell array, and then the next - cycle convolution operation is performed, including: At the end of the current cycle, a complete output element is obtained at the end of each column in the cell array. Then, at the beginning of the next cycle, the value-taking area of the input buffer on the input matrix is shifted 3 pixels to the right based on the value-taking area of the input buffer on the input matrix in the current cycle. Then, 5 elements are taken and output to the cell array for convolution operation.

6. The data mapping method according to claim 2, wherein The three cell arrays include: a first cell array, a second cell array, and a third cell array; Group and add the partial results obtained in each cycle to obtain the convolution result, including: Add the result obtained at the end of the first column in the first cell array, the result obtained at the end of the first column in the second cell array, and the result obtained at the end of the first column in the third cell array in the current cycle to obtain a complete output element; Add the result obtained at the end of the second column in the first cell array, the result obtained at the end of the third column in the first cell array, and the result obtained at the end of the third column in the second cell array in the current cycle to the partial results obtained in the previous cycle respectively to obtain a complete output element; Add the result obtained at the end of the second column in the second cell array, the result obtained at the end of the second column in the third cell array, and the result obtained at the end of the third column in the third cell array in the current cycle to the partial results obtained in the next cycle respectively to obtain a complete output element.

7. A data mapping device for improving the utilization rate of a memory-computation array, characterized in that, The data mapping device includes: A split mapping module for splitting and mapping a k×k convolution kernel into k memory and computing arrays, mapping a 1×k×C vector in each memory and computing array. If the number of rows of the memory and computing array is not enough, it is mapped to the next memory and computing array, where k×k represents the height×width of the C convolution kernel, and C represents the number of input channels of the convolution kernel; A determination method module for determining whether the memory and computing array uses the first data mapping method or the second data mapping method according to the relationship between the convolution kernel size and the memory and computing array size; The data mapping device further includes: a first method module and a second method module; Among them, the first method module includes: An expansion mapping unit for mapping the first row of a 3×3 convolution kernel to any cell array in the memory and computing array when k is 3, and the elements in the first column of this cell array correspond to the arrangement after the first row of the 3×3 convolution kernel is expanded into a column vector; A blank row mapping unit for mapping the same weights as those in the first column in the second column and the third column in the cell array respectively, and leaving C rows empty downward for the mapping positions corresponding to the elements in the second column of the cell array, and leaving 2C rows empty downward for the mapping positions corresponding to the elements in the third column of the cell array; An extended input unit for extending the input data in the input buffer into a 1×5×C column vector; A first convolution unit for, in the current cycle, the input buffer outputting 5 elements in the input matrix to the cell array for convolution operation; A value output unit, which is used to obtain a complete output element at the end of each column in the cell array at the end of the current cycle, and then perform the next-cycle convolution operation; The second method module includes: A full-row mapping unit, which is used when k is 3 to map the parameters on each input channel of the 3×3 convolution kernel to different cell arrays in the memory and computing array respectively, and the three columns in each cell array map the same weights without empty rows, so that each input channel occupies ceil(C / n) cell arrays, a total of ceil(C / n)×3 arrays are occupied in the row dimension, and ceil(M×3 / n) arrays are occupied in the column dimension, where ceil is the ceiling function and n represents the size of the memory and computing array; A second convolution unit, which is used in the current cycle to output 3 elements in the input matrix to three cell arrays in the memory and computing array and perform convolution operations to obtain partial results; An addition unit, which is used to group and add the partial results obtained in each cycle to obtain the result of the convolution.

8. The data mapping device according to claim 7, wherein The size of the convolution kernel is k×k×C×M, and the size of the memory and computing array is n×n, where M represents the number of output channels of the convolution kernel; The determination method module is specifically used for: When the convolution kernel is split and mapped to any cell array, if the number of rows occupied by any cell array is not greater than k / 2k - 1 of the size of the cell array, or the number of columns occupied by any cell array is less than 1 / k of the size of the cell array, it is determined that the memory and computing array uses the first data mapping method; When the convolution kernel is split and mapped to any cell array, if the number of rows occupied by any cell array is not greater than k / 2k - 1 of the size of the cell array and the number of columns occupied by any cell array is less than 1 / k of the size of the cell array, it is determined that the memory and computing array uses the first data mapping method; When the convolution kernel is split and mapped to any cell array, if the number of rows occupied by any cell array is not less than k times the size of the cell array, the ratio of the first occupied array to the computing speed and the ratio of the second occupied array to the computing speed are calculated based on the size of the convolution kernel and the size of the memory and computing array. The ratio of the first occupied array to the computing speed is the ratio of the occupied array to the computing speed when the convolution kernel is mapped to the memory and computing array by the traditional method, and the ratio of the second occupied array to the computing speed is the ratio of the occupied array to the computing speed obtained by the method of mapping the parameters on each input channel of the convolution kernel to different cell arrays in the memory and computing array respectively and mapping the same weights to three columns in each cell array without empty rows; If the value of the ratio of the first occupied array to the computing speed is greater than the value of the ratio of the second occupied array to the computing speed, it is determined that the memory and computing array uses the second data mapping method.

9. The data mapping device according to claim 7, characterized in that The first convolution unit is specifically used for: Among the five elements, the first two elements have been partially calculated in the previous cycle and the remaining calculations are completed in the current cycle. The middle element can complete all three calculations with one input in the current cycle. The last two elements complete partial calculations in the current cycle, and the remaining calculations are completed by placing the last two elements in the positions of the first two elements in the next cycle.

10. The data mapping device according to claim 9, wherein The value output unit is specifically configured to: At the end of the current cycle, a complete output element is obtained at the end of each column in the cell array. Then, at the beginning of the next cycle, the value-taking area of the input buffer on the input matrix is shifted 3 pixels to the right based on the value-taking area of the input buffer on the input matrix in the current cycle. Then, five elements are taken and output to the cell array for convolution operation.