Scaling operator calculation method based on deep learning acceleration kernel
By expanding the input tensor into one-dimensional eigenvectors in parallel to calculate the weight of the output tensor, the problem of low computing efficiency of the scaling operator during hardware deployment is solved, efficient scaling operator inference is realized, and hardware performance is significantly improved.
Patent Information
- Application Number
- CN202510448639.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the scaling operator has low computing efficiency when deploying hardware, especially in neural network models, which requires a large number of multiplication and addition operations, resulting in low computing efficiency.
Expand the two-dimensional eigenvectors of each channel of the input tensor into one-dimensional eigenvectors as rows of matrix A, use the number of channels C of the input tensor as rows of matrix A, and use the number of features H*W of each channel of the input tensor as columns of matrix A; determine the size of matrix B based on the size of the input and output tensors, and calculate the weight of each pixel point of the output tensor, load it into the input_buffer and weight_buffer of the deep learning acceleration kernel for matrix operations, and use the MAC array for parallel calculation.
This greatly reduces the time-consuming process of the model inference stage and improves the computing efficiency of the scaling operator. It is especially obvious for feature maps with a large number of channels. The inference time is about 0.025ms on the deep learning acceleration core, which significantly improves hardware performance.
Smart Images

Figure CN120297345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a neural network processor NPU, and particularly to a method for calculating a scaling operator based on a deep learning acceleration kernel. Background Art
[0002] The scaling operator, also known as the resize operator, is a deep learning algorithm layer whose main function is to adjust the size of the input tensor to the target size. The specific implementation method is to use an interpolation algorithm to calculate the values of the output tensor. There are many interpolation methods for the scaling operator, among which bilinear interpolation is the most commonly used interpolation method, which can effectively adjust the size of the output tensor while ensuring the output quality.
[0003] However, the above method needs to traverse the positions of each pixel point during the inference process and then calculate the interpolated values. In a neural network model, the two-dimensional feature map of the input tensor often has tens of thousands of pixel points. Therefore, a large number of multiplication and addition operations are required on the hardware, and each interpolation calculation needs to repeatedly fetch values from the input tensor, which leads to the problem of low operation efficiency when this method is deployed on the hardware. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] In view of the above-mentioned drawbacks of the prior art, the present invention provides a method for calculating a scaling operator based on a deep learning acceleration kernel, which can effectively overcome the defect of low calculation efficiency of the scaling operator existing in the prior art.
[0006] (2) Technical Solutions
[0007] To achieve the above object, the present invention is realized through the following technical solutions:
[0008] A method for calculating a scaling operator based on a deep learning acceleration kernel, which expands the two-dimensional feature vectors of each channel of the input tensor into one-dimensional feature vectors as the rows of matrix A, takes the number of channels C of the input tensor as the number of rows of matrix A, and takes the number of features H*W of each channel of the input tensor as the number of columns of matrix A;
[0009] Determine the size of matrix B according to the size of the input tensor and the target size of the output tensor. Take the number of columns H*W of matrix A as the number of rows of matrix B, take the number of features DST_H*DST_W of each channel of the output tensor as the number of columns of matrix B. The columns of matrix B correspond to each pixel point of the output tensor, calculate the weights required for calculating each pixel point of the output tensor, and place the calculated weights at the corresponding positions of each column of matrix B;
[0010] Load matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration core respectively for matrix operations, and the calculation result of the scaling operator can be obtained;
[0011] Among them, the size of the input tensor is [C, H, W], where C is the number of channels of the input tensor, H is the feature height of the input tensor, and W is the feature width of the input tensor; the target size of the output tensor is [C, DST_H, DST_W], where DST_H is the feature height of the output tensor and DST_W is the feature width of the output tensor.
[0012] Preferably, calculating the weights required for each pixel point of the output tensor and placing the calculated weights at the corresponding positions in each column of matrix B includes:
[0013] S1. Set matrix B as a matrix of all 0s, and the size of matrix B is [H * W, DST_H * DST_W];
[0014] S2. Traverse each pixel point of the output tensor and calculate the weights of the four nearest pixel points around the corresponding position of each pixel point of the output tensor in the input tensor;
[0015] S3. Place the calculated weights at the corresponding positions in each column of matrix B according to each pixel point of the output tensor.
[0016] Preferably, in S2, traversing each pixel point of the output tensor and calculating the weights of the four nearest pixel points around the corresponding position of each pixel point of the output tensor in the input tensor includes:
[0017] S21. Calculate the height scaling ratio scale H and the width scaling ratio scale W :
[0018]
[0019] S22. Calculate the corresponding position of the pixel point of the output tensor in the input tensor according to the height scaling ratio scale H and the width scaling ratio scale W :
[0020] src x = x * scale W ;
[0021] src y = y * scale H ;
[0022] Among them, (x, y) is the pixel point position of the output tensor, (src x , srcy ) is the corresponding position in the input tensor of the pixel of the output tensor;
[0023] S23. Calculate the positions of the four nearest pixels around the corresponding position in the input tensor of the pixel of the output tensor:
[0024]
[0025] x2 = x1 + 1;
[0026]
[0027] y2 = y1 + 1;
[0028] Among them, (x1, y1), (x1, y2), (x2, y1), and (x2, y2) are the positions of the four nearest pixels around the corresponding position in the input tensor of the pixel of the output tensor, is floor function;
[0029] S24. Calculate the weights of the four nearest pixels around the corresponding position in the input tensor of the pixel of the output tensor:
[0030] ω1 = (x2 - src x ) * (y2 - src y );
[0031] ω2 = (src x - y1) * (y2 - src y );
[0032] ω3 = (x2 - src x ) * (src y - y1);
[0033] ω4 = (src x - x1) * (scr y - y1);
[0034] Among them, ω1, ω2, ω3, and ω4 are the weights of the four nearest pixels at the positions (x1, y1), (x1, y2), (x2, y1), and (x2, y2) in the input tensor of the pixel of the output tensor respectively.
[0035] Preferably, in S3, according to each pixel of the output tensor, the calculated weights are placed at the corresponding positions in each column of matrix B, including:
[0036] For the nearest pixel at the position (x1, y1) in the input tensor of the pixel of the output tensor, place its weight ω1 at the position y1 * W + x1 in the first column of matrix B;
[0037] For the nearest pixel of the pixel point of the output tensor at the position (x1, y2) in the input tensor, place its weight ω2 at the position y2*W + x1 in the first column of matrix B.
[0038] Preferably, by loading matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration core respectively for matrix operations, the calculation result of the scaling operator can be obtained, including:
[0039] Set the starting position and step size parameters of the uop, load matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration core respectively, execute the matrix instruction to multiply matrix A and matrix B, and export the multiplication result from the output_buffer to the DDR to obtain the calculation result of the scaling operator.
[0040] Preferably, based on the characteristic that the two-dimensional convolution results of elements in different regions in the two-dimensional feature map plane can be calculated in parallel, the deep learning acceleration core realizes that all processing elements PE in the MAC array calculate the two-dimensional convolution results of each region in the two-dimensional feature map in parallel when processing a single batch of two-dimensional convolution calculations;
[0041] Among them, the MAC array is the execution unit of the matrix instruction.
[0042] Preferably, the MAC array is provided with data by the on-chip memory. The on-chip memory is divided into three major blocks, which are called blocks. One block of memory provides data to the MAC array in the row direction, denoted as input_buffer; one block of memory provides data to the MAC array in the column direction, denoted as weight_buffer; one block of memory receives the calculation results of the MAC array, denoted as output_buffer;
[0043] Each block of memory is further divided into multiple small blocks. Each small block of memory transmits data to a row of processing elements PE or a column of processing elements PE. These small blocks of memory can transmit data to the MAC array simultaneously, denoted as bank;
[0044] The data of each bank in the input_buffer is broadcast to each row of the MAC array, the data of each bank in the weight_buffer is broadcast to each column of the MAC array, and each processing element PE performs the multiply-accumulate calculation of the input data in the row direction and the input data in the column direction. The calculation results of a row of processing elements PE are concatenated and output to the corresponding bank of the output_buffer;
[0045] Among them, the number of banks of the input_buffer and the output_buffer is the same as the number of rows R of the processing elements PE in the MAC array, and the number of banks of the weight_buffer is the same as the number of columns C of the processing elements PE in the MAC array.
[0046] Preferably, when processing a single batch of two-dimensional convolution calculations, the data of each region in the two-dimensional feature map is respectively stored in each bank of the input_buffer, the convolution kernel weights for each output channel are stored in each bank of the weight_buffer, and the two-dimensional convolution results of each region are stored in each bank of the output_buffer;
[0047] The two-dimensional feature map is divided into multiple regions, the data of each region is respectively stored in each bank of the input_buffer, and is respectively broadcast to all the processing elements PE in each row of the MAC array. Each row of processing elements PE calculates the two-dimensional convolution result of one region respectively, so as to realize the parallel calculation of the two-dimensional convolution results of each region in the two-dimensional feature map by all the processing elements PE in the MAC array when processing a single batch of two-dimensional convolution calculations;
[0048] When the calculation results are output, the calculation results of each row of processing elements PE are concatenated and output to the corresponding bank of the output_buffer, so as to achieve the effect that the calculation results of the output feature Figure 1 in each region are still in one bank in the output_buffer. When performing two-dimensional convolution calculation on the calculation results again, data rearrangement is not required;
[0049] Among them, the matrix A, the matrix B, and the output tensor are arranged in the corresponding buffer in the mode of bank-channel first and column first;
[0050] Denote the width of the two-dimensional feature map, that is, the number of columns of the two-dimensional feature map, as FC, and the number of banks of the input_buffer as R. Then each bank of the input_buffer can store at most columns of data, and only the first banks store columns of data, and the next bank stores the remaining columns of data, and the subsequent banks are not filled with data;
[0051] Among them, is the ceiling function, is the floor function.
[0052] Preferably, when processing a single batch of two-dimensional convolution calculations, uops are used to specify the addresses for fetching data from the input_buffer and weight_buffer, and the address for storing data to the output_buffer.
[0053] Among them, one uop is used to specify the addresses for fetching data from the input_buffer and weight_buffer during one calculation, and the address for storing data to the output_buffer. Multiple uops are combined to complete the calculation of a two-dimensional convolution sliding window. The set of uops is stored in the uop_buffer. The addresses specified by multiple uops required in one operation can be continuous or discontinuous.
[0054] Preferably, based on the characteristic that the moving step size of the sliding window remains fixed in the row direction or column direction during the two-dimensional convolution process, a quadruple loop is used to control the movement of the two-dimensional convolution sliding window on the two-dimensional feature map.
[0055] The use of a quadruple loop to control the movement of the two-dimensional convolution sliding window on the two-dimensional feature map includes:
[0056] The innermost loop traverses the data of multiple input channels corresponding to the same position on the two-dimensional feature map. The second innermost loop flexibly specifies the addresses of each element in the sliding window participating in the convolution calculation through uops. The number of uops in the second innermost loop is the same as the number of elements in the two-dimensional convolution sliding window. The innermost loop and the second innermost loop complete the calculation of the two-dimensional convolution result at one output feature map position.
[0057] The outermost two loops respectively control the traversal of the two-dimensional feature map in the row direction or column direction.
[0058] The innermost loop traverses the data of multiple input channels corresponding to the same position on the two-dimensional feature map. The second innermost loop flexibly specifies the addresses of each element in the sliding window participating in the convolution calculation through uops. The outermost two loops respectively control the traversal of the two-dimensional feature map in the row direction or column direction, including:
[0059] Use the number of times of the innermost loop inner_loop_cnt and the stride of the innermost loop inner_input_stride to specify the addresses of the data of multiple input channels corresponding to the same position on the two-dimensional feature map.
[0060] Use multiple uops to flexibly specify the addresses of each element in the sliding window participating in the convolution calculation.
[0061] Use the number of times of the second outermost loop outer_loop_cnt, the stride of the second outermost loop Outer_input_stride, the number of times of the outermost loop top_loop_cnt, and the stride of the outermost loop Top_Input_stride to control the traversal of the two-dimensional feature map in the row direction or column direction.
[0062] (III) Beneficial Effects
[0063] Compared with the prior art, a method for calculating a scaling operator based on a deep learning acceleration core provided by the present invention has the following beneficial effects:
[0064] 1) By using the method proposed by the present invention, the scaling operator can be deployed on the deep learning acceleration core, and the data required for matrix calculation can be prepared in the model initialization stage. In the inference stage, only one matrix instruction is needed to complete the calculation of the scaling operator, which greatly reduces the time consumption in the model inference stage;
[0065] 2) By using the method proposed by the present invention, parallel calculation of the scaling operator can be realized, accelerating the inference of the scaling operator, especially for feature maps with a large number of channels. Since matrix multiplications of multiple channels can be executed in parallel, and regardless of the number of channels of the feature map, matrix B only needs to be calculated once and there is only one matrix, so the calculation efficiency is very high, and the more the number of channels, the higher the calculation efficiency. Through experimental verification, the inference time of the scaling operator on the deep learning acceleration core is about 0.025 ms, and the inference time using the CPU of the hxai100 chip is about 40.12 ms, which shows that the inference performance of the deployed scaling operator has been improved well. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0067] Figure 1 It is a schematic diagram of the construction of matrix A in the present invention;
[0068] Figure 2 It is a schematic diagram of the multiplication of matrix A and matrix B in the present invention;
[0069] Figure 3 It is a schematic diagram of the data transmission relationship between three types of on-chip memories and the MAC array of the deep learning acceleration core in the present invention;
[0070] Figure 4Schematic diagram of arranging matrix A in the bank of the input_buffer in the pattern of bank channel priority and column priority according to the present invention;
[0071] Figure 5 Schematic diagram of the execution time sequence of the matrix instruction in the double - layer loop according to the present invention;
[0072] Figure 6 Schematic diagram of the execution unit of the matrix instruction according to the present invention;
[0073] Figure 7 Schematic diagram of the calculation of a MAC array in the execution unit of the matrix instruction according to the present invention;
[0074] Figure 8 Schematic diagram of the calculation of all MAC arrays in the execution unit of the matrix instruction according to the present invention. Detailed implementation manners
[0075] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0076] In the technical solution of this application, based on the characteristic that the two - dimensional convolution results of elements in different regions within the two - dimensional feature map plane can be calculated in parallel, the deep - learning acceleration core realizes parallel calculation of the two - dimensional convolution results of each region in the two - dimensional feature map by all processing elements PE in the MAC array (the execution unit of the matrix instruction) when processing a single - batch two - dimensional convolution calculation.
[0077] As Figure 3 shown, the MAC array is provided with data by the on - chip memory. The on - chip memory is divided into three major blocks, which are called blocks. One block of memory provides data to the MAC array in the row direction, denoted as input_buffer (for storing matrix A); one block of memory provides data to the MAC array in the column direction, denoted as weight_buffer (for storing matrix B); one block of memory receives the calculation results of the MAC array, denoted as output_buffer (for storing the output tensor);
[0078] Each block of memory is further divided into multiple small blocks. Each small - block memory transmits data to a row of processing elements PE or a column of processing elements PE. These small - block memories can transmit data to the MAC array simultaneously, denoted as bank.
[0079] The data of each bank in the input_buffer is broadcast to each row of the MAC array, and the data of each bank in the weight_buffer is broadcast to each column of the MAC array. Each processing element PE performs the multiply-accumulate calculation of the input data in the row direction and the input data in the column direction. The calculation results of one row of processing elements PE are concatenated and output to the corresponding bank of the output_buffer;
[0080] Among them, the number of banks in the input_buffer and output_buffer is the same as the number of rows R of the processing elements PE in the MAC array, and the number of banks in the weight_buffer is the same as the number of columns C of the processing elements PE in the MAC array.
[0081] When processing a single batch of two-dimensional convolution calculations, the data in each region of the two-dimensional feature map is stored in each bank of the input_buffer, the convolution kernel weights for each output channel are stored in each bank of the weight_buffer, and the two-dimensional convolution results of each region are stored in each bank of the output_buffer;
[0082] The two-dimensional feature map is divided into multiple regions, and the data in each region is stored in each bank of the input_buffer and broadcast to all processing elements PE in each row of the MAC array respectively. Each row of processing elements PE calculates the two-dimensional convolution result of one region respectively, realizing the parallel calculation of the two-dimensional convolution results of each region in the two-dimensional feature map by all processing elements PE in the MAC array when processing a single batch of two-dimensional convolution calculations;
[0083] When the calculation results are output, the calculation results of each row of processing elements PE are concatenated and output to the corresponding bank of the output_buffer, realizing the effect that the calculation results of each region of the output feature Figure 1 are still in one bank in the output_buffer. When performing two-dimensional convolution calculation on the calculation results again, data rearrangement is not required;
[0084] Among them, as Figure 4 shown, the arrangement modes of matrix A, matrix B and the output tensor in the corresponding buffer are in the mode of bank-channel first and column first.
[0085] As Figure 4 shown, denote the width of the two-dimensional feature map, that is, the number of columns of the two-dimensional feature map as FC, and the number of banks in the input_buffer as R. Then each bank of the input_buffer stores at most columns of data, and only the first banks store columns of data, and the next bank stores the remaining For column data, the subsequent bank does not fill in data;
[0086] Among them, is for ceiling, is for floor.
[0087] When processing single-batch two-dimensional convolution calculation, use uop to specify the addresses for fetching data from input_buffer and weight_buffer, and the address for storing to output_buffer;
[0088] Among them, one uop is used to specify the addresses for fetching data from input_buffer and weight_buffer during one calculation, and the address for storing to output_buffer. Multiple uops are combined to complete the calculation of a two-dimensional convolution sliding window. The uop set is stored in uop_buffer, and the addresses specified by multiple uops required in one operation are either continuous or discontinuous.
[0089] Such as Figure 5 shown, based on the characteristic that during the two-dimensional convolution process, the moving step in the row direction or column direction of the sliding window remains fixed, use a quadruple loop (double loop) to control the movement of the two-dimensional convolution sliding window on the two-dimensional feature map, and accordingly design to use one instruction to control the hardware state machine to execute the following pseudocode to drive the MAC array to complete the entire two-dimensional convolution calculation:
[0090]
[0091]
[0092] Use a quadruple loop (double loop) to control the movement of the two-dimensional convolution sliding window on the two-dimensional feature map, including:
[0093] The innermost loop traverses the data of multiple input channels corresponding to the same two-dimensional feature map position. The second-innermost loop flexibly specifies the addresses of each element in the sliding window participating in the convolution calculation through uop. The number of uops in the second-innermost loop is the same as the number of elements in the two-dimensional convolution sliding window. The innermost loop and the second-innermost loop complete the calculation of the two-dimensional convolution result at one output feature map position;
[0094] The outermost two loops respectively control traversing the two-dimensional feature map in the row direction or column direction.
[0095] The innermost loop traverses the data of multiple input channels corresponding to the same two-dimensional feature map position. The second-innermost loop flexibly specifies the addresses of each element in the sliding window participating in the convolution calculation through uop. The outermost two loops respectively control traversing the two-dimensional feature map in the row direction or column direction, including:
[0096] Use the number of times of the innermost loop inner_loop_cnt and the innermost loop stride inner_input_stride to specify the data addresses of multiple input channels corresponding to the same position in the two-dimensional feature map;
[0097] Use multiple uops to flexibly specify the addresses of each element in the sliding window participating in the convolution calculation;
[0098] Use the number of times of the second outermost loop outer_loop_cnt, the second outermost loop stride Outer_input_stride, the number of times of the outermost loop top_loop_cnt, and the outermost loop stride Top_Input_stride to control the traversal of the two-dimensional feature map in the row direction or column direction.
[0099] The matrix instruction is used to control the MAC array to perform matrix multiplication operations and two-dimensional convolution calculations. The matrix instruction and the set_matrix instruction together complete the matrix multiplication and convolution calculation processes. The execution process of the matrix instruction includes four layers of loops and can complete the matrix multiplication calculation of two matrix blocks; by reasonably setting the uop, the positions and loop increments of each buffer in the four layers of loops, this instruction can also perform two-dimensional convolution calculations with multiple input channels and multiple output channels.
[0100] The execution unit of the matrix instruction is a two-dimensional MAC array. The number of rows of this array is denoted as R, and the number of columns is denoted as C. The input in the row direction of each processing element PE is P*8-bit data, and the input in the column direction is also P*8-bit data. Each processing element PE can perform P 8-bit multiplications internally, or perform P / 2 16-bit multiplications. R should be the same as the number of banks of the input_buffer, and C should be the same as the number of banks of the weight_buffer.
[0101] Generally, when the matrix instruction is executed, the data fetched from one bank of the input_buffer is broadcast to a fixed row of the MAC array, and the data fetched from one bank of the weight_buffer is broadcast to a fixed column of the MAC array. The output bit width of each bank of the Input_buffer and each bank of the weight_buffer should be P*8. The calculation results of all processing elements PE in a row are concatenated together and sent to one bank of the output_buffer. Each processing element PE outputs 8-bit calculation results per cycle, and all processing elements PE in a row output C*8-bit calculation results per cycle to one bank of the output_buffer. If the calculation results to be output are greater than 8 bits, they are output in multiple cycles.
[0102] Taking 8-bit calculation as an example, at the intersection of each row and column, that is, inside each processing element PE, the 8-bit data transmitted by P input_buffers is multiplied by the 8-bit data transmitted by P weight_buffers respectively, and then accumulated to complete the calculation of an 8-bit vector dot product. The calculation result of each processing element PE is accumulated with the value of ACC and temporarily stored back in ACC. The bit width of ACC is 32 bits. When Output_mode is 2'b00 and Output_bitwidth is 2'b00, whenever the uop loop is completed, the value in ACC is quantized and then output. Each ACC is quantized into 8-bit data for output. The output bit width of a row of processing elements PE is C * 8 bits and is sent to one bank of output_buffer. When Output_mode is 2'b01 and Output_bitwidth is 2'b11, the output bit width of ACC is 32 bits. Whenever the uop loop is completed, the total value of a row of ACC, which is 32 * C bits, is sent to one bank of output_buffer in multiple cycles.
[0103] Generally, the input_buffer has a total of R banks, the bit width of each bank is P * 8, and the total bit width is R * P * 8; the weight_buffer has a total of C banks, the bit width of each bank is P * 8, and the total bit width is C * P * 8; the output_buffer has a total of R banks, the bit width of each bank is C * 8, and the total bit width is R * C * 8.
[0104] The matrix multiplication operation is performed by multiple MAC arrays. As Figures 6 - 8 shown, after the input from the input_buffer and the weight_buffer is multiplied and added, it is accumulated with the ACC stored inside the processing element PE. After the uop loop ends, it is then exported to the output_buffer.
[0105] A method for calculating a scaling operator based on a deep learning acceleration core. As Figure 1 shown, the two-dimensional feature vector of each channel of the input tensor is expanded into a one-dimensional feature vector as the rows of matrix A. The number of channels C of the input tensor is used as the number of rows of matrix A, and the number of features H * W of each channel of the input tensor is used as the number of columns of matrix A.
[0106] Determine the size of matrix B according to the size of the input tensor and the target size of the output tensor. Take the number of columns H*W of matrix A as the number of rows of matrix B, and take the number of features DST_H*DST_W of each channel of the output tensor as the number of columns of matrix B. The columns of matrix B correspond to each pixel of the output tensor. Calculate the weights required for each pixel of the output tensor, and place the calculated weights at the corresponding positions of each column of matrix B;
[0107] Load matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration core respectively for matrix operations. As Figure 2 shown, the calculation result of the scaling operator can be obtained;
[0108] Among them, the size of the input tensor is [C, H, W], where C is the number of channels of the input tensor, H is the feature height of the input tensor, and W is the feature width of the input tensor; the target size of the output tensor is [C, DST_H, DST_W], where DST_H is the feature height of the output tensor and DST_W is the feature width of the output tensor.
[0109] ① Calculate the weights required for each pixel of the output tensor, and place the calculated weights at the corresponding positions of each column of matrix B, including:
[0110] S1. Set matrix B as a matrix of all 0s, and the size of matrix B is [H*W, DST_H*DST_W];
[0111] S2. Traverse each pixel of the output tensor, and calculate the weights of the four nearest pixels around the corresponding position of each pixel of the output tensor in the input tensor;
[0112] S3. Place the calculated weights at the corresponding positions of each column of matrix B according to each pixel of the output tensor.
[0113] 1) In S2, traverse each pixel of the output tensor, and calculate the weights of the four nearest pixels around the corresponding position of each pixel of the output tensor in the input tensor, including:
[0114] S21. Calculate the height scaling ratio scale H and the width scaling ratio scale W :
[0115]
[0116] S22. Calculate the corresponding position of the pixel of the output tensor in the input tensor according to the height scaling ratio scale H and the width scaling ratio scale W :
[0117] src x = x * scale W ;
[0118] src y = y * scale H ;
[0119] Where (x, y) is the pixel position of the output tensor, and (src x , src y ) is the corresponding position of the pixel of the output tensor in the input tensor;
[0120] S23. Calculate the positions of the four nearest pixels around the corresponding position of the pixel of the output tensor in the input tensor:
[0121]
[0122] x2 = x1 + 1;
[0123]
[0124] y2 = y1 + 1;
[0125] Where (x1, y1), (x1, y2), (x2, y1), and (x2, y2) are the positions of the four nearest pixels around the corresponding position of the pixel of the output tensor in the input tensor, is rounding down;
[0126] S24. Calculate the weights of the four nearest pixels around the corresponding position of the pixel of the output tensor in the input tensor:
[0127] ω1 = (x2 - src x ) * (y2 - src y );
[0128] ω2 = (src x - y1) * (y2 - src y );
[0129] ω3 = (x2 - src x ) * (src y - y1);
[0130] ω4 = (src x - x1) * (scr y - y1);
[0131] Where ω1, ω2, ω3, and ω4 are the weights of the four nearest pixels at the positions (x1, y1), (x1, y2), (x2, y1), and (x2, y2) of the pixel of the output tensor in the input tensor, respectively.
[0132] 2) In S3, place the calculated weights at the corresponding positions in each column of matrix B according to each pixel point of the output tensor, including:
[0133] For the nearest pixel point of the pixel point of the output tensor at the position (x1, y1) in the input tensor, place its weight ω1 at the position y1*W + x1 in the first column of matrix B;
[0134] For the nearest pixel point of the pixel point of the output tensor at the position (x1, y2) in the input tensor, place its weight ω2 at the position y2*W + x1 in the first column of matrix B.
[0135] ② Load matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration kernel respectively for matrix operations, and then the calculation result of the scaling operator can be obtained, including:
[0136] Set the starting position and step parameter of the uop, load matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration kernel respectively, execute the matrix instruction to multiply matrix A and matrix B. As Figure 2 shown, export the multiplication result from the output_buffer to the DDR, and then the calculation result of the scaling operator can be obtained.
[0137] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A calculation method for a scaling operator based on a deep learning acceleration kernel, characterized in that: Unfold the two-dimensional feature vectors of each channel of the input tensor into one-dimensional feature vectors as the rows of matrix A. Take the number of channels C of the input tensor as the number of rows of matrix A, and take the number of features H*W of each channel of the input tensor as the number of columns of matrix A; Determine the size of matrix B according to the size of the input tensor and the target size of the output tensor. Take the number of columns H*W of matrix A as the number of rows of matrix B, and take the number of features DST_H*DST_W of each channel of the output tensor as the number of columns of matrix B. The columns of matrix B correspond to each pixel point of the output tensor. Calculate the weights required for each pixel point of the output tensor and place the calculated weights at the corresponding positions of each column of matrix B; Load matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration kernel respectively for matrix operations, and then the calculation result of the scaling operator can be obtained; Among them, the size of the input tensor is [C, H, W], where C is the number of channels of the input tensor, H is the feature height of the input tensor, and W is the feature width of the input tensor; the target size of the output tensor is [C, DST_H, DST_W], where DST_H is the feature height of the output tensor and DST_W is the feature width of the output tensor.
2. The scaling operator calculation method based on a deep learning acceleration core according to claim 1, wherein: The step of calculating the weights required for each pixel point of the output tensor and placing the calculated weights at the corresponding positions of each column of matrix B includes: S1. Set matrix B as a matrix of all 0s, and the size of matrix B is [H*W, DST_H*DST_W]; S2. Traverse each pixel point of the output tensor and calculate the weights of the four nearest pixel points around the corresponding position of each pixel point of the output tensor in the input tensor; S3. Place the calculated weights at the corresponding positions of each column of matrix B according to each pixel point of the output tensor.
3. The scaling operator calculation method based on the deep learning acceleration core according to claim 2, characterized in that: In S2, when traversing each pixel point of the output tensor and calculating the weights of the four nearest pixel points around the corresponding position of each pixel point of the output tensor in the input tensor, it includes: S21. Calculate the height scaling ratio scale H , the width scaling ratio scale W : S22. Scale according to the height scale factor scale H and the width scale factor scale W to calculate the corresponding position of the pixel points of the output tensor in the input tensor: src x = x * scale W ; src y = y * scale H ; where (x, y) is the pixel position of the output tensor, and (src x , src y ) is the corresponding position of the pixel of the output tensor in the input tensor; S23. Calculate the positions of the four nearest pixel points around the corresponding position of the pixel point of the output tensor in the input tensor: x2 = x1 + 1; y2 = y1 + 1; Among them, (x1, y1), (x1, y2), (x2, y1), and (x2, y2) are the positions of the four nearest pixels around the corresponding positions of the pixels of the output tensor in the input tensor. is rounding down; S24. Calculate the weights of the four nearest pixel points around the corresponding position of the pixel point of the output tensor in the input tensor: ω1 = (x2 - src x ) * (y2 - src y ); ω2 = (src x - y1) * (y2 - src y ); ω3 = (x2 - src x ) * (src y - y1); ω4 = (src x - x1)*(scr y - y1); Among them, ω1, ω2, ω3, and ω4 are the weights of the four nearest pixel points at the positions (x1, y1), (x1, y2), (x2, y1), and (x2, y2) of the pixel point of the output tensor in the input tensor respectively.
4. The method for calculating a scaling operator based on a deep learning acceleration core according to claim 2, wherein: In S3, placing the calculated weights at the corresponding positions of each column of matrix B according to each pixel point of the output tensor includes: For the nearest pixel point at the position (x1, y1) of the pixel point of the output tensor in the input tensor, place its weight ω1 at the position y1*W + x1 in the first column of matrix B; For the nearest pixel point at the position (x1, y2) of the pixel point of the output tensor in the input tensor, place its weight ω2 at the position y2*W + x1 in the first column of matrix B.
5. The method for calculating a scaling operator based on a deep learning acceleration core according to claim 1, wherein: Loading matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration core respectively for matrix operations can obtain the calculation result of the scaling operator, including: Setting the starting position and step parameters of the uop, loading matrix A and matrix B into the input_buffer and weight_buffer of the deep learning acceleration core respectively, executing the matrix instruction to multiply matrix A and matrix B, and exporting the multiplication result from the output_buffer to the DDR, then the calculation result of the scaling operator can be obtained.
6. The method for calculating a scaling operator based on a deep learning acceleration kernel according to any one of claims 1 to 5, characterized in that: Based on the feature that the two-dimensional convolution results of elements in different regions within the two-dimensional feature map plane can be calculated in parallel, the deep learning acceleration core realizes parallel calculation of the two-dimensional convolution results of each region in the two-dimensional feature map by all processing elements PE in the MAC array when processing a single batch of two-dimensional convolution calculations; Among them, the MAC array is the execution unit of the matrix instruction.
7. The scaling operator calculation method based on the deep learning acceleration core according to claim 6, characterized in that: The MAC array is provided with data by the on-chip memory. The on-chip memory is divided into three major blocks, which are called blocks. One block of memory provides data to the MAC array in the row direction, denoted as input_buffer; one block of memory provides data to the MAC array in the column direction, denoted as weight_buffer; one block of memory receives the calculation results of the MAC array, denoted as output_buffer; Each block of memory is further divided into multiple small blocks. Each small block of memory transmits data to a row of processing elements PE or a column of processing elements PE. These small blocks of memory can transmit data to the MAC array simultaneously, denoted as bank; The data of each bank in the input_buffer is broadcast to each row of the MAC array, the data of each bank in the weight_buffer is broadcast to each column of the MAC array, and each processing element PE performs the multiply-accumulate calculation of the input data in the row direction and the input data in the column direction. The calculation results of a row of processing elements PE are concatenated and output to the corresponding bank of the output_buffer; Among them, the number of banks in the input_buffer and output_buffer is the same as the number of rows R of the processing elements PE in the MAC array, and the number of banks in the weight_buffer is the same as the number of columns C of the processing elements PE in the MAC array.
8. The method for calculating a scaling operator based on a deep learning acceleration core according to claim 7, wherein: When processing a single batch of two-dimensional convolution calculations, the data of each region in the two-dimensional feature map is stored in each bank of the input_buffer, the convolution kernel weights for each output channel are stored in each bank of the weight_buffer, and the two-dimensional convolution results of each region are stored in each bank of the output_buffer; The two-dimensional feature map is divided into multiple regions, and the data of each region is stored in each bank of the input_buffer respectively, and is broadcast to all the processing elements PE in each row of the MAC array. Each row of processing elements PE calculates the two-dimensional convolution result of one region respectively, so as to realize the parallel calculation of the two-dimensional convolution results of each region in the two-dimensional feature map by all the processing elements PE in the MAC array when processing a single batch of two-dimensional convolution calculations; When the calculation results are output, the calculation results of each row of processing elements PE are concatenated and output to the corresponding bank of the output_buffer, so as to realize the effect that the calculation results of one region of the output feature map are still in one bank in the output_buffer. When performing the two-dimensional convolution calculation on the calculation results again, there is no need to perform data rearrangement; Among them, matrix A, matrix B and the output tensor are arranged in the data in the corresponding buffer in the mode of bank channel priority and column priority; Let the width of the two-dimensional feature map, i.e., the number of columns of the two-dimensional feature map, be FC, and the number of banks of the input_buffer be R. Then, each bank of the input_buffer can store at most columns of data, and only the first banks store columns of data. The next bank stores the remaining columns of data, and the subsequent banks are not filled with data; wherein, is the ceiling function, is the floor function.
9. The method for calculating a scaling operator based on a deep learning acceleration core according to claim 8, wherein: When processing a single batch of two-dimensional convolution calculations, uop is used to specify the addresses for fetching data from the input_buffer and the weight_buffer, and the address for storing to the output_buffer; Among them, one uop is used to specify the addresses for fetching data from the input_buffer and the weight_buffer during one calculation, and the address for storing to the output_buffer. Multiple uops are combined to complete the calculation of a two-dimensional convolution sliding window. The uop set is stored in the uop_buffer, and the addresses specified by multiple uops required in one operation are continuous or discontinuous.
10. The method for calculating a scaling operator based on a deep learning acceleration core according to claim 9, wherein: Based on the characteristic that the moving step length of the sliding window is fixed in the row direction or the column direction during the two-dimensional convolution process, a four-layer loop is adopted to control the movement of the two-dimensional convolution sliding window on the two-dimensional feature map; The adoption of a four-layer loop to control the movement of the two-dimensional convolution sliding window on the two-dimensional feature map includes: The innermost loop traverses the data of multiple input channels corresponding to the same position of the two-dimensional feature map. The second innermost loop flexibly specifies the addresses of each element in the sliding window participating in the convolution calculation through uop. The number of uops in the second innermost loop is the same as the number of elements in the two-dimensional convolution sliding window. The innermost loop and the second innermost loop complete the calculation of the two-dimensional convolution result of one position of the output feature map; The outermost two loops respectively control the traversal of the two-dimensional feature map in the row direction or the column direction; The innermost loop traverses the data of multiple input channels corresponding to the same position of the two-dimensional feature map. The second innermost loop flexibly specifies the addresses of each element in the sliding window participating in the convolution calculation through uop. The outermost two loops respectively control the traversal of the two-dimensional feature map in the row direction or the column direction, including: Use the innermost loop count inner_loop_cnt and the innermost loop stride inner_input_stride to specify the addresses of the data of multiple input channels corresponding to the same position of the two-dimensional feature map; Flexibly specify the addresses of each element in the sliding window participating in the convolution calculation by using multiple uops; Use the number of times of the second outermost loop outer_loop_cnt, the stride of the second outermost loop Outer_input_stride, the number of times of the outermost loop top_loop_cnt, and the stride of the outermost loop Top_Input_stride to control traversing the two-dimensional feature map in the row direction or the column direction.