Data processing method and device

By using intermediate-level memory units to save the output gradient sub-tenster and perform basic convolution operations in deconvolution processing, the problems of high storage space requirements and large data read and write overhead are solved, and the execution efficiency of deconvolution processing and the performance of neural network models are improved.

CN120234152APending Publication Date: 2025-07-01广州壁仞智能科技有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510420950.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art has problems such as high storage space requirements and large data read and write overhead in deconvolution processing, resulting in low execution efficiency.

Method used

The intermediate level memory unit is used to save the output gradient tensor. By dividing the output gradient sub-tensor and performing basic convolution operations, the storage space is recycled to reduce the read and write data in external memory and improve execution efficiency.

Benefits of technology

By recycling the storage space of intermediate-level memory units, the data access time is shortened, the execution efficiency of deconvolution processing is improved, and the performance of neural network models is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234152A_ABST
    Figure CN120234152A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, and the method comprises the steps: a, determining the size of an output gradient sub-tensor stored at a time based on an available storage space of a first middle-level memory unit used for storing an output gradient tensor and a parameter of PADDING data of the output gradient tensor; receiving and inputting an output gradient tensor of a non-1-step-length deconvolution module; b, selecting a current output gradient sub-tensor from the current unprocessed output gradient tensors, and storing to-be-processed data composed of the current output gradient sub-tensor and sub PADDING data of the current output gradient sub-tensor into a first middle-level memory unit; c, performing basic convolution processing on the to-be-processed data, outputting a result of the basic convolution processing to a position, corresponding to the current output gradient sub-tensor, in the external memory, and releasing the to-be-processed data from the first middle-level memory unit; and returning to the step b until all the output gradient tensors are processed. By applying the method and the device, the execution efficiency of deconvolution processing can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence technology, and particularly to data processing methods and devices. Background Art

[0002] With the continuous development of artificial intelligence technology, neural network models are increasingly widely used.

[0003] Deconvolution is an operator often used in the training and inference of neural network models, usually implemented in a deconvolution module. The implementation method and execution efficiency of the deconvolution module become the key to improving the performance of the entire neural network model.

[0004] Deconvolution operations involve many optional configuration parameters: kernel size, stride, dilation coefficient, padding, etc. Currently, in the processing of non-1 stride convolution modules, the requirements for the size of the storage space are relatively high, requiring additional memory overhead and additional data read / write overhead. Summary of the Invention

[0005] This application provides data processing methods and devices that can effectively improve the execution efficiency of deconvolution processing.

[0006] To achieve the above objective, this application adopts the following technical solutions: A data processing method, including: a. Determine the size of the output gradient sub-tensor stored once based on the available storage space of the first intermediate-level memory unit for storing the output gradient tensor and the parameters of the PADDING data of the output gradient tensor; receive the output gradient tensor input to the non-1 stride deconvolution module; b. Select data equal in amount to the size of the output gradient sub-tensor from the currently unprocessed output gradient tensor as the current output gradient sub-tensor, determine the sub-PADDING data of the current output gradient sub-tensor, and save the data to be processed composed of the current output gradient sub-tensor and its sub-PADDING data to the first intermediate-level memory unit; c. Perform basic convolution processing on the data to be processed, output the result of the basic convolution processing to the corresponding position in the external memory for the current output gradient sub-tensor, and then release the data to be processed from the first intermediate-level memory unit; return to step b until all of the output gradient tensor is processed.

[0007] Preferably, the performing basic convolution processing on the data to be processed includes: Using each weight subtensor obtained by splitting the weight tensor, successively perform basic convolution operations on the data to be processed, and merge the results of the basic convolution operations to obtain the result of the basic convolution processing.

[0008] Preferably, the merging of the results of the basic convolution operations includes: Save the results of each basic convolution operation in the second intermediate-level memory unit, and perform merging processing in the second intermediate-level memory unit in a set format, and save the merged result at the position corresponding to the current output gradient subtensor.

[0009] Preferably, after outputting the result of the basic convolution processing to an external memory, the method further includes: releasing the results of each basic convolution operation from the second intermediate-level memory unit.

[0010] Preferably, determining the sub-PADDING data of the current output gradient subtensor includes: After processing of selecting the current output gradient subtensor, supplement sub-PADDING data around the current output gradient subtensor, and determine the supplemented sub-PADDING data as the sub-PADDING data of the current output gradient subtensor; Wherein, the method of supplementing the sub-PADDING data includes: For the current output gradient subtensor, for any PADDING direction of any edge data, if there is data in the output gradient tensor in the any PADDING direction of the any edge data, use the M data closest to the any edge data in the any PADDING direction as the sub-PADDING data of the any edge data in the any PADDING direction; if there is no data in the output gradient tensor in the any PADDING direction of the any edge data, set the sub-PADDING data of the any edge data in the any PADDING direction to a preset value; Wherein, the PADDING direction is the direction of the additional sub-PADDING data relative to its adjacent edge data, and the M is the number of sub-PADDING data in the any PADDING direction.

[0011] Preferably, before the first selection of the current output gradient subtensor, the method further includes: supplementing PADDING data to the received output gradient tensor to generate a corrected output gradient tensor; Determining the sub-PADDING data of the current output gradient subtensor includes: In the corrected output gradient tensor, determine a data block centered on the current output gradient sub-tensor with dimensions (H1 + H2×2, W1 + W2×2), and determine the data other than the current output gradient sub-tensor in the determined data block as the sub-PADDING data; Among them, (H1, W1) is the data dimension of the current output gradient sub-tensor, and H2×2 and W2×2 are the numbers of sub-PADDING data of the current output gradient sub-tensor in the vertical and horizontal directions respectively.

[0012] A data processing device includes: a receiving unit, a sub-tensor selection unit, a basic convolution unit, an output unit, and a first intermediate-level memory unit; The receiving unit is configured to determine the size of the output gradient sub-tensor stored once based on the available storage space of the first intermediate-level memory unit for storing the output gradient tensor and the parameters of the PADDING data of the output gradient tensor; receive the output gradient tensor of the input deconvolution module with a non-1 stride; The sub-tensor selection unit is configured to select the current output gradient sub-tensor from the currently unprocessed output gradient tensor based on the size of the output gradient sub-tensor, determine the sub-PADDING data of the current output gradient sub-tensor, and save the data to be processed composed of the current output gradient sub-tensor and its sub-PADDING data to the first intermediate-level memory unit; The basic convolution unit is configured to perform basic convolution processing on the data to be processed to obtain the result of the basic convolution processing; The output unit is configured to output the result of the basic convolution processing to the position corresponding to the current output gradient sub-tensor in the external memory, and then notify the first intermediate-level memory unit to release the data to be processed; notify the sub-tensor selection unit to perform the next selection of the current sub-tensor until all the input tensors are processed; The first intermediate-level memory unit is configured to save the data to be processed and release the data to be processed after receiving the notification from the basic convolution unit.

[0013] Preferably, the device includes a second intermediate-level memory unit; In the basic convolution unit, the performing basic convolution processing on the data to be processed includes: Using each weight subtensor obtained by splitting the weight tensor, perform basic convolution operations with the data to be processed in sequence, save the results of each of the basic convolution operations in the second intermediate-level memory unit, and perform merging processing in a set format in the second intermediate-level memory unit, and save the merged result as the result of the basic convolution processing at the position corresponding to the current output gradient subtensor.

[0014] Preferably, the output unit is further configured to, after outputting the result of the basic convolution processing to an external memory, notify the second intermediate-level memory unit to release the results of each of the basic convolution operations.

[0015] A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the data processing method described in any one of the above can be implemented.

[0016] A computer program product, including computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data processing method described in any one of the above is implemented.

[0017] As can be seen from the above technical solution, the present application performs deconvolution processing for non-1 stride. First, based on the available storage space of the first intermediate-level memory unit for storing the output gradient tensor and the parameters of the PADDING data, the size of the output gradient sub-tensor stored once is determined, and the input tensor of the convolution module with non-1 stride is received; in the output gradient tensor, data equal in amount to the size of the output gradient sub-tensor is selected from the unprocessed data as the current output gradient sub-tensor for performing a sub-convolution (sub conv) operation. The sub conv operation is looped multiple times until all output gradient tensors are processed; in each sub conv operation, sub-PADDING data is determined for the selected current output gradient sub-tensor, and the data to be processed composed of the current output gradient sub-tensor and its sub-PADDING data is stored in the first intermediate-level memory unit; the stored data to be processed is subjected to a basic convolution operation, and the operation result is output to the position in the global memory corresponding to the split data of the current output gradient sub-tensor to be output, and then the storage space in the first intermediate-level memory unit for storing the data to be processed is released, thus completing one sub conv operation. After looping the sub conv operation multiple times, the complete deconvolution operation of the output gradient tensor and the output of the deconvolution result are completed. In the above processing, by storing the output gradient sub-tensor data in the first intermediate-level memory unit in each subconv operation and releasing the corresponding storage space after the sub conv operation, on the one hand, the deconvolution operation with non-1 stride is completed through looping the sub conv operation multiple times, and on the other hand, although the storage space of the intermediate-level memory unit is limited, the function of storing the output gradient sub-tensor is realized through the cyclic use of the storage space of the corresponding memory unit, so that the intermediate-level memory unit closer to the core processing part of the convolution module is used for data storage during the deconvolution operation, greatly shortening the time required for data access and effectively improving the execution efficiency of the deconvolution processing, and further effectively improving the performance of the neural network model where the deconvolution processing is located. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic flowchart of the basic process of the data processing method in the present application; Figure 2 is a block diagram of the overall architecture of the data processing method in the present application; Figure 3 is a schematic diagram of dividing the output gradient tensor input to the deconvolution module into four output gradient sub-tensors; Figure 4 is a schematic diagram of the output gradient tensor of the input deconvolution module and its PADDING data; Figure 5a and Figure 5b are respectivelyFigure 4 Schematic diagrams of two output gradient sub-tensors numbered 1 and 2 and their sub-PADDING data; Figure 6 Schematic diagram for supplementing sub-PADDING data to the selected output gradient sub-tensors in the output gradient tensor; Figure 7 Schematic diagram for splitting the weight tensor to obtain weight sub-tensors; Figure 8 Schematic diagram for data merging and storage of the basic conv operation results; Figure 9 Schematic diagram for performing the basic conv operation on the output gradient sub-tensor and the weight sub-tensor; Figure 10 Basic structural schematic diagram of the data processing device in this application. Detailed implementation manner

[0019] In order to make the purpose, technical means and advantages of this application clearer and more understandable, the following further elaborates on this application in conjunction with the accompanying drawings.

[0020] For the deconvolution operation with a non-1 stride, it can currently be implemented through a combination of basic convolution (Conv) operators (hereinafter referred to as Method A). Specifically, the deconvolution operation is used in the reverse propagation process of the model and is carried out on the basis of the completion of the convolution operation. Therefore, the input of the deconvolution operation is the gradient result of the output of the convolution operation, which is usually called the output gradient tensor (gradoutput). That is to say, the input tensor of the deconvolution operation is the output gradient tensor of the convolution operation. Specifically, in Method A, the weight tensor (weight) of the convolution is decomposed into multiple weight sub-tensors (sub weight), and then the output gradient tensor (grad output) of the convolution operation is respectively subjected to a basic convolution operation with the weight sub-tensors (sub weight), and the operation results are combined together in a certain format to obtain the deconvolution result. Considering the limited size of the intermediate-level memory and the inability to completely store the output gradient tensor, when performing the basic Conv operation with the weight sub-tensor, it is necessary to completely store the output gradient tensor of the convolution operation in an external memory (such as global memory), and then the complete non-1 stride deconv operation is achieved by repeatedly executing the basic Conv operator. Since the output gradient tensor of the convolution operation used in the processing process is stored in the external memory, this will result in additional memory overhead and additional data read / write overhead.

[0021] Considering the above-mentioned problem of memory overhead, in the present application, the deconvolution process is improved on the basis of the above-mentioned processing method, and the intermediate-level memory unit is used for data transfer, so as to effectively reduce the time for data reading and writing during the deconvolution process, and improve the execution efficiency of the deconvolution process and the performance of the system where the deconvolution process is located.

[0022] Figure 1 It is a schematic diagram of the basic process of the data processing method in the present application. As Figure 1 shown, the method includes: Step 101, determine the size of the output gradient sub-tensor stored once based on the available storage space of the first intermediate-level memory unit and the parameters of the PADDING data of the output gradient tensor; receive the output gradient tensor input to the deconvolution module with a non-1 stride.

[0023] The processing of the present application is carried out for the deconvolution process with a non-1 stride, that is, the processing of the deconvolution module with a non-1 stride is optimized. This deconvolution process can be carried out in the deconvolution processing module of various related processing systems (such as neural network models). Therefore, this step receives the gradient tensor input to such a deconvolution module with a non-1 stride. Among them, the input gradient tensor is the input data of the corresponding deconvolution module, which can specifically be various physical data in the neural network model. For example, it can be the image data processed in the image neural network model, or it can also be the text data processed in the text neural network model, etc. In fact, the input data of the deconvolution module is the gradient of the output tensor of the convolution process. Hereinafter, to avoid ambiguity, the input data of the deconvolution module is uniformly referred to as the output gradient tensor (grad output). Usually, the output gradient tensor can be received from the external memory into the register for transfer.

[0024] In the implementation process of the above-mentioned method A mentioned above, since the data volume of the output gradient tensor is relatively large, it is necessary to use the external memory to store it, which leads to frequent data reading and writing in the external memory during the basic convolution operation processing, resulting in additional memory overhead and data reading and writing overhead, and then causing the execution efficiency of the deconvolution operation to be relatively low. Based on this, in the present application, the intermediate-level memory unit is used to store the output gradient tensor. Since the intermediate-level memory unit is closer to the core processing of the convolution operation than the external memory, it can effectively save the data reading and writing time and overhead and improve the execution efficiency of the convolution operation.

[0025] However, the storage space of the intermediate-level memory cells is limited and cannot completely save the output gradient tensor. Based on this, in this application, the output gradient tensor is divided into multiple output gradient sub-tensors (sub grad output), and only part of the data in the output gradient tensor is subjected to the basic convolution operation each time to reduce the amount of data to be saved and ensure that the intermediate-level memory cells can provide enough storage space to save the output gradient sub-tensors. This step is used to determine the size of the output gradient sub-tensor saved in the intermediate-level memory cells each time, that is, to determine how much data in the input tensor is saved each time for the basic Conv operation.

[0026] Among them, the size of the output gradient sub-tensor can be determined based on the available storage space size of the intermediate-level memory cells. Specifically, the size of the output gradient sub-tensor needs to ensure that: a single output gradient sub-tensor and its PADDING data (to distinguish it from the PADDING data of the output gradient tensor, the PADDING data of a single output gradient sub-tensor is hereinafter referred to as sub-PADDING data) can be completely saved in the available storage space of the intermediate-level memory cells, and the remaining space is as small as possible. More specifically, on the one hand, in order to complete the convolution operation as soon as possible, the available storage space of the intermediate-level memory cells should be used as much as possible to store the output gradient sub-tensors, and the output gradient sub-tensors should be as large as possible, so that the number of loops is the least and the time used is the shortest; on the other hand, the output gradient sub-tensors need to be saved after adding sub-PADDING data, and the amount of sub-PADDING data added is determined by the dimension of the PADDING data of the output gradient tensor, and the dimension of the PADDING data of the output gradient tensor is determined by the PADDING parameter of the output gradient tensor. Therefore, the parameter of the PADDING data of the output gradient tensor is related to the amount of data to be saved; thus, considering the above two factors comprehensively, the size of the output gradient sub-tensor is determined. Among them, the parameters of the known PADDING data usually give the direct parameters in the convolution process, that is, the dimension of the additional data in the convolution process. For the transposed convolution process, it is necessary to first convert based on the given parameters of the PADDING data to obtain the dimension of the PADDING data added to the output gradient tensor in the transposed convolution operation. The specific conversion method is the same as the existing one.

[0027] The processing of specifically determining the size of the output gradient sub-tensor can be completed in advance before the entire transposed convolution process, or even before all the processes of the system (such as a neural network model) where the transposed convolution process is located. Of course, it can also be completed after receiving the output gradient tensor of the input transposed convolution module. This application does not limit the specific position of this operation.

[0028] In addition, the intermediate-level memory units can be various storage units that are closer to the convolution processing distance relative to the external memory, such as shared buffer or shared memory. For the convenience of description, the intermediate-level memory unit used to save the output gradient sub-tensor is hereinafter referred to as the first intermediate-level memory unit, or tensor buffer, to distinguish it from other intermediate-level memory units.

[0029] Step 102: Select data equal in size to the output gradient sub-tensor from the currently unprocessed output gradient tensor as the current output gradient sub-tensor.

[0030] One execution process from this step to step 105 completes one sub conv operation. As mentioned above, to use the intermediate-level memory unit with limited storage space to save the output gradient tensor, for each sub conv operation, part of the data in the output gradient tensor is read, the basic convolution operation is performed, and the operation result is stored. As Figure 2 shown, where the sub conv for each page is performed on one output gradient sub-tensor, so as to realize the close-range storage of the output gradient tensor through the storage space reuse (buffer reuse) of the intermediate-level memory unit.

[0031] In the first sub conv operation (i.e., the first loop processing of steps 102 - 104), part of the data is selected from the input output gradient tensor based on the size of the output gradient sub-tensor as the current output gradient sub-tensor; while for subsequent sub conv operations, data is selected from the remaining unprocessed output gradient tensor as the current output gradient sub-tensor. For example, if the size of the output gradient sub-tensor is 1 / 4 of the data size of the input output gradient tensor, then in the first subconv operation, the first 1 / 4 of the data is selected, as Figure 3 shown by the data block numbered 1 in; in the second sub conv operation, another 1 / 4 of the data is selected from the remaining unprocessed data, and so on, until the last 1 / 4 of the data is selected in the fourth sub conv operation. Among them, the selection order of the output gradient sub-tensors is not restricted and can be set according to actual needs. For example, for Figure 3 the 4 output gradient sub-tensors shown, they can be selected in the order of 1, 2, 3, 4, or in the order of 4, 3, 2, 1, as long as all the output gradient sub-tensors that make up the output gradient tensor are processed.

[0032] For each sub-conv operation, the current output gradient sub-tensor is selected from the register used to transfer the output gradient tensor, and is ready to be saved into the first intermediate-level storage unit. Compared with the above-mentioned output gradient tensor saving method of the present application, in the implementation process of the foregoing method A, the saving of the output gradient tensor is performed on the output gradient tensor of the input deconvolution module. Therefore, it cannot be completely saved in the intermediate-level memory unit and can only be saved in the external memory. It can be seen that the method of implementing close-range storage by dividing the output gradient tensor into multiple sub-tensors in the present application can effectively improve the execution efficiency of the deconvolution process.

[0033] Step 103: Determine the sub-PADDING data of the current output gradient sub-tensor, and save the data to be processed composed of the current output gradient sub-tensor and its sub-PADDING data into the first intermediate-level memory unit.

[0034] The selected current output gradient sub-tensor needs to perform a basic Conv operation subsequently. Therefore, it is necessary to determine the PADDING data of the current output gradient sub-tensor. To distinguish it from the PADDING data of the output gradient tensor, the PADDING data of the output gradient sub-tensor is hereinafter referred to as sub-PADDING data. After determining the sub-PADDING data of the current output gradient sub-tensor, a set of data to be processed composed of the current output gradient sub-tensor and its sub-PADDING data is saved in the first intermediate-level memory unit.

[0035] There are two specific ways to determine the sub-PADDING data: Method 1: Supplement PADDING data to the output gradient tensor, and determine the sub-PADDING data of the current output gradient sub-tensor based on the result after supplementing PADDING data to the output gradient tensor; Specifically, after receiving the output gradient tensor and before the first selection of the output gradient sub-tensor, PADDING data can be supplemented to the output gradient tensor in the existing manner. As Figure 4 shown, the data block with the digital part represents the output gradient tensor, and the dotted-line box data block around the output gradient tensor represents the PADDING data of the output gradient tensor. For the convenience of subsequent description, the data block formed after supplementing PADDING data to the output gradient tensor is called the corrected output gradient tensor.

[0036] After selecting the current output gradient sub-tensor, a data block centered on the current output gradient sub-tensor with dimensions (H1 + H2×2, W1 + W2×2) is determined in the corrected output gradient tensor, and the data in the determined data block excluding the current output gradient sub-tensor is determined as the sub-PADDING data of the current output gradient sub-tensor. Among them, (H1, W1) is the data dimension of the current output gradient sub-tensor, and H2×2 and W2×2 are the numbers of sub-PADDING data in the vertical and horizontal directions of the current output gradient sub-tensor respectively. For example, Figure 4 in the corrected output gradient tensor shown, the output gradient sub-tensor numbered 1 and its sub-PADDING data are as Figure 5a shown, and the output gradient sub-tensor numbered 2 and its sub-PADDING data are as Figure 5b shown, where the solid data block represents the output gradient sub-tensor, the dashed data block represents the sub-PADDING data, and the complete data block represents the data to be processed.

[0037] Method 2: Directly supplement sub-PADDING data for the output gradient sub-tensor; Specifically, after receiving the output gradient tensor, instead of supplementing PADDING data for the output gradient tensor, the current output gradient sub-tensor can be directly selected from the output gradient tensor, and then sub-PADDING data is supplemented around the output gradient sub-tensor. For example Figure 6 shown, the left output gradient tensor is divided into four output gradient sub-tensors. For the output gradient sub-tensor numbered 1, the sub-PADDING data supplemented for it is as Figure 6 shown in the right data block, where the solid data block represents the output gradient sub-tensor and the dashed data block represents the supplemented sub-PADDING data.

[0038] The specific way to supplement sub-PADDING data can be carried out according to the existing method. In addition, to ensure the accuracy of the convolution operation, the present application provides a preferred way to supplement sub-PADDING data: 1) Determine the PADDING direction for any edge data Y of the current output gradient sub-tensor; The PADDING direction here refers to the direction of supplementing sub-PADDING data, that is, the direction of the additional sub-PADDING data relative to its adjacent edge data. For example Figure 7 for the upper left edge data of the orange data block in, there are three PADDING directions, left, upper left, and upper. For the edge data in the first row and second column of the orange data block, there is only one PADDING direction, upper; 2) For any PADDING direction X, determine whether there is data in the sub-PADDING direction X of the edge data Y in the output gradient tensor; if there is, perform step 3), otherwise perform step 4); In the implementation method of the foregoing method A, the output gradient tensor of the input deconvolution module is used as a data block for the basic Conv operation. Therefore, the PADDING data is supplemented as a whole for this data block. In this processing method, the PADDING data are all new data and can be set to a preset value, such as 0; while in the processing of this application, each basic Conv operation is performed on part of the data of the output gradient tensor (that is, the output gradient sub-tensor), and some edge data of each output gradient sub-tensor is actually not real edge data from the perspective of the output gradient tensor, but intermediate data of the output gradient tensor. The following will refer to this type of edge data as the first type of edge data. For example, Figure 6 the edge data marked with the number 4 in the output gradient sub-tensor numbered 1 in Figure 6 is the first type of edge data, and some edge data of the output gradient sub-tensor is also edge data from the perspective of the output gradient tensor. The following will refer to this type of edge data as the second type of edge data. For example, the edge data marked with the number 1 in the output gradient sub-tensor numbered 1 in is the second type of edge data; In the current output gradient sub-tensor, for the above two different types of edge data, different strategies are adopted when supplementing the sub-PADDING data; the first type of edge data is supplemented with sub-PADDING data by the processing method of step 3), and the second type of edge data is supplemented with sub-PADDING data by the processing method of step 4); 3) Use the M data closest to the edge data Y in the PADDING direction X as the sub-PADDING data of the edge data Y in the PADDING direction X; That is, for the first type of edge data, adjacent data is used as the sub-PADDING data, and this processing can ensure the accuracy of the convolution operation; where M is the preset number of sub-PADDING data supplemented in the PADDING direction X; 4) Set the sub-PADDING data of the edge data Y in the PADDING direction X to a preset value; That is, for the second type of edge data, the existing method of supplementing PADDING data is adopted, that is, the sub-PADDING data is set to a preset value, such as 0.

[0039] Through the method of supplementing sub-PADDING data provided by this application above, it can be ensured that the basic convolution processing performed on the current output gradient sub-tensor can accurately simulate the movement process of the convolution kernel inside the output gradient tensor and ensure the accuracy of the convolution operation.

[0040] Step 104, perform basic convolution processing on the data to be processed.

[0041] For the data to be processed, perform basic convolution processing. Specifically, each weight subtensor obtained by splitting the weight tensor is successively subjected to basic convolution operations with the data to be processed, and the results of each basic convolution operation are merged to obtain the result of basic convolution processing.

[0042] More specifically, splitting the weight tensor to obtain weight subtensors can be as Figure 7 shown. The specific process includes: 1) Read the weight tensor (weight) from external storage (such as global memory); 2) Split the weight tensor using a specific data splitting method (knit split). Here, the knit split varies according to the stride parameter passed to the conv operator interface used by the user. Taking the case of stride = 2 as an example, the weight tensor will be split into 4 weight subtensors according to the corresponding splitting rules of the coordinates. The splitting process is completed through the register; 3) Write the split data saved in the register to an external storage unit (such as an extra global memory applied). Taking the case of stride = 2 as an example, finally, the original weight tensor will be split into 4 weight subtensors, and the data sizes of each weight subtensor are different. Figure 7 Taking a weight tensor of [1, 3, 3, 3] (kernel size is 3x3) as an example, the splitting method of the weight subtensors is shown. After splitting, four weight subtensors of [1, 3, 2, 2], [1, 3, 2, 1], [1, 3, 1, 2], and [1, 3, 1, 1] are obtained. Among them, in the size [K, C, R, S] of the weight tensor, K represents the number of output channels of the convolution operation, C represents the number of input channels of the convolution operation, and R and S represent the height and width of the convolution kernel respectively. The specific operation of determining the number of weight subtensors and performing the specific splitting is the same as the existing method, and will not be elaborated here.

[0043] For each weight subtensor obtained after splitting, perform basic conv operations with the data to be processed, as Figure 8As shown, multiple basic conv operation results (sub grad input1, sub grad input2, sub grad input3, and sub grad input4) can be obtained. All the basic conv operation results are combined together according to the set format to obtain the basic convolution processing result (sub grad input) of the data to be processed. Specifically, each weight sub-tensor performs a basic conv operation with the data to be processed to obtain a basic conv operation result. These operation results need to be combined according to the combination rule (knit merge) under different convolution strides, such as Figure 8 the combination rule shown

[0044] In fact, the process of performing basic convolution operations on the data to be processed and then combining them is the same as the process of performing basic convolution operations on the existing modified output gradient tensor and then combining them, except for the data dimensions

[0045] As described above, compared with the data volume of the existing modified output gradient tensor, the data volume of the data to be processed in this application becomes smaller. Therefore, the number of basic conv operation results obtained from the basic convolution processing of the data to be processed also becomes smaller. To further improve the execution efficiency of the convolution processing, each basic conv operation result can be stored in an intermediate-level memory unit, which is referred to as the second intermediate-level memory unit here. Additionally, the first intermediate-level memory unit and the second intermediate-level memory unit can be the same storage unit or different storage units; if the two are the same storage unit, different storage spaces can be pre-allocated to separately store the data to be processed and the basic conv operation results

[0046] Step 105: Output the result of the basic convolution processing to the corresponding position in the external memory for the current output gradient sub-tensor, and then release the data to be processed from the first intermediate-level memory unit

[0047] In this step, the basic convolution processing result obtained in step 104 is output to the external memory. This external memory is used to store the deconvolution processing result of the output gradient tensor. Among them, if the basic convolution processing result obtained in step 104 is the deconvolution processing result of the current output gradient sub-tensor targeted by the current loop of the sub conv operation, then when storing this deconvolution processing result in the external memory, it is stored at the storage position corresponding to the current output gradient sub-tensor, such as Figure 8Here, the storage location corresponding to the data to be split of the current output gradient sub-tensor is a storage space pre-set for each output gradient sub-tensor for storing the corresponding deconvolution result. In this way, after multiple cycles, the deconvolution processing results of each cycle are stored in the corresponding storage location to form the deconvolution processing result of the complete output gradient tensor.

[0048] After outputting the deconvolution processing result of the current sub conv operation, it means that the sub conv operation of the current output gradient sub-tensor is completed, and the storage space in the first intermediate level memory unit for storing the data to be processed and the storage space in the second intermediate level memory unit for storing the basic conv operation results can be released, so that in the next cycle, the corresponding storage space can be reused to save the data to be processed and the corresponding basic conv operation results in the next cycle. Therefore, by recycling the first intermediate level memory unit and the second intermediate level memory unit, the output gradient tensor and the results of the basic conv operation in the entire deconvolution processing operation of the present application can be saved, the execution efficiency of the deconvolution operation can be improved, and then the performance of the system (such as a neural network model) where the entire deconvolution operation is located can be improved.

[0049] At this point, a cycle of steps 102 to 105 is completed, that is, the current output gradient sub-tensor is processed. Figure 2 A sub conv operation is shown.

[0050] Step 106, return to step 102, and perform the next sub conv operation until all output gradient tensors are processed.

[0051] Through this step, we enter the sub conv operation of the next loop. The loop ends when all the output gradient tensors are processed, that is, all the data input to the deconvolution module has been processed, all the deconvolution operations have been completed, and the process ends.

[0052] The above is a specific implementation of the data processing method of the present application. In the above processing of the present application, the output gradient tensor is saved multiple times in the intermediate level memory unit loop, thereby shortening the storage location of the output gradient tensor, effectively improving the execution efficiency of the deconvolution operation, and thus improving the processing performance of the entire deconvolution processing system (such as a neural network model).

[0053] To more clearly illustrate the data processing method in this application, a specific example is given below. Figure 9As shown, taking two-dimensional convolution as an example, assuming that the shape of the output gradient tensor grad_output is [1, 1, 8, 8], and the shape of the output gradient sub-tensor sub_grad_output is [1, 1, 4, 4], then 4 sub-convolutions are required, that is, it is divided into four output gradient sub-tensors. Among the dimensions [N, K, P, Q] of the output gradient tensor, N represents the number of batches, K represents the number of convolution output channels, and P and Q represent the height and width of the output gradient tensor respectively; the weight tensor of [1, 3, 3, 3] is split to obtain four weight sub-tensors subweight1~subweight4 of [1, 3, 2, 2], [1, 3, 2, 1], [1, 3, 1, 2], [1, 3, 1, 1]; for any output gradient sub-tensor, it is necessary to perform a basic conv operation with each weight sub-tensor; then the results of the basic conv operations are merged and stored at the corresponding positions. The basic conv operation here refers to the convolution operation with a stride of 1 between the output gradient sub-tensor (sub_grad_output) and one of the weight sub-tensors (sub_weight).

[0054] It should be noted that due to the characteristics of transposed convolution, the sub-PADDING data used in each basic conv operation is different, and the specific sub-PADDING data used can be determined according to the dimensions of the weight sub-tensor after splitting. For example, Figure 9In the example shown, it is assumed that the convolution operation parameters are stride = 2, padding = (0, 0), kernel size = (3, 3), dilation = 1, and group = 1. After calculation, the dimensions of the deconvolution PADDING data of the output gradient tensor are BpaPadH = BpaPadW = 2. However, since the weight tensor is split into the four weight sub-tensors subweight1 to subweight4 mentioned above, when the output gradient sub-tensor performs the basic conv operation with subweight1, the sub-PADDING data used corresponds to (BpaPadH, BpaPadW) = (1, 1), that is, the lengths of the sub-PADDING data used in the vertical and horizontal directions are both 1; when the output gradient sub-tensor performs the basic conv operation with subweight2, the sub-PADDING data used corresponds to (BpaPadH, BpaPadW) = (1, 0), that is, the lengths of the sub-PADDING data used in the vertical and horizontal directions are 1 and 0 respectively; when the output gradient sub-tensor performs the basic conv operation with subweight3, the sub-PADDING data used corresponds to (BpaPadH, BpaPadW) = (0, 1), that is, the lengths of the sub-PADDING data used in the vertical and horizontal directions are 0 and 1 respectively; when the output gradient sub-tensor performs the basic conv operation with subweight4, the sub-PADDING data used corresponds to (BpaPadH, BpaPadW) = (0, 0), that is, the lengths of the sub-PADDING data used in the vertical and horizontal directions are both 0, as detailed in Figure 9 shown.

[0055] The above is the specific implementation of the data processing method of the present application. The present application also provides a data processing device that can be used to implement the above processing method. Figure 10 It is a schematic diagram of the basic structure of the data processing device in the present application. As Figure 10 shown, the device includes: a receiving unit, a sub-tensor selection unit, a basic convolution unit, an output unit, and a first intermediate-level memory unit.

[0056] Among them, the receiving unit is used to determine the size of the output gradient sub-tensor stored once based on the available storage space of the first intermediate-level memory unit for storing the output gradient tensor and the parameters of the PADDING data of the output gradient tensor; receive the output gradient tensor of the input deconvolution module with a non-1 stride; receive the output gradient tensor of the input deconvolution module with a non-1 stride; A sub-tensor selection unit, configured to select a current output gradient sub-tensor from the currently unprocessed output gradient tensor based on the size of the output gradient sub-tensor, determine the sub-PADDING data of the current output gradient sub-tensor, and save the data to be processed composed of the current output gradient sub-tensor and its sub-PADDING data to the first intermediate-level memory unit; A basic convolution unit, configured to perform basic convolution processing on the data to be processed to obtain the result of the basic convolution processing; An output unit, configured to output the result of the basic convolution processing to the corresponding position in the external memory for the current output gradient sub-tensor, and then notify the first intermediate-level memory unit to release the data to be processed; notify the sub-tensor selection unit to perform the next selection of the current sub-tensor until all input tensors are processed; The first intermediate-level memory unit is configured to save the data to be processed and release the data to be processed after receiving the notification from the basic convolution unit.

[0057] Optionally, in the basic convolution unit, the operation of performing basic convolution processing on the data to be processed may specifically include: Using each weight sub-tensor obtained by splitting the weight tensor, sequentially performing basic convolution operations with the data to be processed, and merging the results of the basic convolution operations, and taking the merged result as the result of the basic convolution processing.

[0058] Optionally, the device may include a second intermediate-level memory unit; In the basic convolution unit, the process of merging the results of the basic convolution operations may specifically include: Saving the results of each basic convolution operation in the second intermediate-level memory unit, and performing a merging process in the second intermediate-level memory unit in a set format, and saving the merged result at the position corresponding to the current output gradient sub-tensor.

[0059] Optionally, the output unit is further configured to notify the second intermediate-level memory unit to release the results of each basic convolution operation after outputting the result of the basic convolution processing to the external memory.

[0060] Optionally, in the sub-tensor selection unit, the process of determining the sub-PADDING data of the current output gradient sub-tensor specifically includes: After the process of selecting the current output gradient sub-tensor, supplement sub-PADDING data around the current output gradient sub-tensor, and determine the supplemented sub-PADDING data as the sub-PADDING data of the current output gradient sub-tensor; Among them, the method of supplementing sub-PADDING data may include: For the current output gradient sub-tensor, for any PADDING direction of any edge data, if there is data in any PADDING direction of any edge data in the output gradient tensor, then the M data closest to any edge data in any PADDING direction are used as the sub-PADDING data of any edge data in any PADDING direction; if there is no data in any PADDING direction of any edge data in the output gradient tensor, then the sub-PADDING data of any edge data in any PADDING direction is set to a preset value. Wherein, the PADDING direction is the direction of the additional sub-PADDING data relative to its adjacent edge data, and M is the number of sub-PADDING data in any PADDING direction.

[0061] Optionally, the receiving unit is further configured to supplement PADDING data for the received output gradient tensor to generate a corrected output gradient tensor. In the sub-tensor selection unit, the method for determining the sub-PADDING data of the current output gradient sub-tensor may specifically include: In the corrected output gradient tensor, determine a data block centered on the current output gradient sub-tensor with a dimension of (H1 + H2×2, W1 + W2×2), and determine the data other than the current output gradient sub-tensor in the determined data block as the sub-PADDING data. Wherein, (H1, W1) is the data dimension of the current output gradient sub-tensor, and H2×2 and W2×2 are the numbers of sub-PADDING data of the current output gradient sub-tensor in the vertical and horizontal directions, respectively.

[0062] The data processing method or data processing device provided by at least one embodiment of the present application can be applied to different systems or devices, such as being applied to an electronic device. The electronic device may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an AR device, a VR device, a vehicle-mounted terminal, etc., or may also be a server, etc. The data processing method provided by at least one embodiment of the present application can be applied to scenarios related to data processing such as CPU, high performance computing (HPC), and artificial intelligence (AI) in an electronic device, such as a deconvolution processing unit. Of course, the present application is not limited thereto, and any scenario, device, device, etc. involving non-1-step deconvolution processing can adopt the data processing method or data processing device provided by at least one embodiment of the present application.

[0063] In some embodiments, the data processing device provided by at least one embodiment of the present application may be a chip. For example, the chip is a System-on-a-Chip (SoC). The system-on-a-chip includes a processor, which may be a single-core processor or a multi-core processor, a memory, and an I / O interface, etc. After loading the data and application programs in the memory, the processor can process the data, for example, perform deconvolution processing involving non-unit stride.

[0064] For example, when an electronic device implements the data processing method provided by at least one embodiment of the present disclosure, the processing device performs various appropriate actions and processes according to the non-transitory computer-readable instructions stored in the memory to implement the data processing method. For example, input parameters (such as an output gradient tensor and a weight tensor) are stored in the memory, such as in a register, a cache, or a memory. When non-unit stride deconvolution processing is required, the input parameters are read from the memory, and the processing device processes the input output gradient tensor according to steps 101 to 106 in the data processing method described in at least one embodiment of the present application disclosure to obtain a deconvolution result. The deconvolution result can be transmitted again to a corresponding operator for use, or transmitted to a memory (such as a high-bandwidth memory) for storage.

[0065] In addition, it should be noted that the data type of the data processed by the data processing method or the data processing device provided by at least one embodiment of the present application may have different specific physical meanings according to different application scenarios. For example, the data processing method provided by at least one embodiment of the present application can be applied in fields such as speech processing, image processing, text processing, and video processing.

[0066] For example, in the field of speech processing, the data processed can be any parameter used, input, or generated in tasks such as feature extraction, speech enhancement, and speech recognition, such as a speech feature vector, a filtering parameter, etc.

[0067] For example, in the field of image processing, the data processed can be data for image preprocessing, feature extraction, image segmentation, object detection, etc.

[0068] The present application also provides a computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed by a processor, can execute the steps in the above-mentioned data processing method. In practical applications, the computer-readable medium can be included in each device / device / system of the above embodiments, or can exist separately without being assembled into the device / device / system. Among them, instructions are stored in the computer-readable storage medium, and the stored instructions can execute the steps in the above-mentioned data processing method when executed by a processor.

[0069] According to the embodiments disclosed in the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above, but is not used to limit the scope of protection of the present application. In the embodiments disclosed in the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device.

[0070] The present application also provides a computer program product, including computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the convolutional module processing method in the neural network model as described above.

[0071] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A data processing method, characterized in that: include: a. Determine the size of the output gradient sub-tensor stored at a single time based on the available storage space of the first intermediate level memory unit for storing the output gradient tensor and the parameters of the PADDING data of the output gradient tensor; receive the output gradient tensor of the deconvolution module with non-1 step size; b. Selecting data of the same size as the output gradient sub-tensor from the current unprocessed output gradient tensor as the current output gradient sub-tensor, determining sub-PADDING data of the current output gradient sub-tensor, and saving the data to be processed consisting of the current output gradient sub-tensor and its sub-PADDING data to the first intermediate-level memory unit; c. performing basic convolution processing on the data to be processed, and outputting the result of the basic convolution processing to a position corresponding to the current output gradient sub-tensor in the external memory, and then releasing the data to be processed from the first intermediate-level memory unit; Return to step b until all the output gradient tensors are processed.

2. The method according to claim 1, characterized in that The performing basic convolution processing on the data to be processed includes: The weight sub-tensors obtained by splitting the weight tensor are used to perform basic convolution operations with the data to be processed in turn, and the results of the basic convolution operations are merged to obtain the results of the basic convolution processing.

3. The method according to claim 2, characterized in that The step of merging the results of the basic convolution operation comprises: The results of each of the basic convolution operations are stored in the second intermediate-level memory unit, and are merged in the second intermediate-level memory unit according to a set format, and the merged result is stored in the position corresponding to the current output gradient sub-tensor.

4. The method according to claim 3, characterized in that: After outputting the result of the basic convolution processing to the external memory, the method further includes: releasing the result of each basic convolution operation from the second intermediate level memory unit.

5. The method according to claim 1, characterized in that The determining of the sub-PADDING data of the current output gradient sub-tensor comprises: After the process of selecting the current output gradient sub-tensor, supplementing sub-PADDING data around the current output gradient sub-tensor, and determining the supplemented sub-PADDING data as the sub-PADDING data of the current output gradient sub-tensor; The method of supplementing the sub-PADDING data includes: For the current output gradient sub-tensor, for any PADDING direction of any edge data, if there is data in the output gradient tensor in the any PADDING direction of the any edge data, then the M data closest to the any edge data in the any PADDING direction are used as the sub-PADDING data of the any edge data in the any PADDING direction; if there is no data in the output gradient tensor in the any PADDING direction of the any edge data, then the sub-PADDING data of the any edge data in the any PADDING direction is set to a preset value; The PADDING direction is the direction of the additional sub-PADDING data relative to its adjacent edge data, and M is the number of sub-PADDING data in any PADDING direction.

6. The method according to claim 1, characterized in that Before selecting the current output gradient sub-tensor for the first time, the method further comprises: supplementing the received output gradient tensor with PADDING data to generate a modified output gradient tensor; The determining of the sub-PADDING data of the current output gradient sub-tensor comprises: In the modified output gradient tensor, a data block with the current output gradient sub-tensor as the center and a dimension of (H1+H2×2, W1+W2×2) is determined, and data other than the current output gradient sub-tensor in the determined data block is determined as the sub-PADDING data; Wherein, (H1, W1) is the data dimension of the current output gradient sub-tensor, and H2×2 and W2×2 are the numbers of sub-PADDING data of the current output gradient sub-tensor in the vertical and horizontal directions, respectively.

7. A data processing device, characterized in that: include: Receiving unit, subtensor selection unit, basic convolution unit, output unit and first intermediate level memory unit; The receiving unit is used to determine the size of the output gradient sub-tensor stored at a single time based on the available storage space of the first intermediate level memory unit for storing the output gradient tensor and the parameters of the PADDING data of the output gradient tensor; receive the output gradient tensor of the deconvolution module with a non-1 step size; The sub-tensor selection unit is used to select a current output gradient sub-tensor from a currently unprocessed output gradient tensor based on the size of the output gradient sub-tensor, determine the sub-PADDING data of the current output gradient sub-tensor, and save the to-be-processed data consisting of the current output gradient sub-tensor and its sub-PADDING data to the first intermediate-level memory unit; The basic convolution unit is used to perform basic convolution processing on the data to be processed to obtain a result of the basic convolution processing; The output unit is used to output the result of the basic convolution processing to a position corresponding to the current output gradient sub-tensor in the external memory, and then notify the first intermediate level memory unit to release the data to be processed; Notifying the sub-tensor selection unit to select the current sub-tensor next time until all the input tensors are processed; The first intermediate-level memory unit is used to store the data to be processed and release the data to be processed after receiving a notification from the basic convolution unit.

8. The device according to claim 7, characterized in that The device includes a second intermediate level memory unit; In the basic convolution unit, performing basic convolution processing on the data to be processed includes: Using each weight sub-tensor obtained by splitting the weight tensor, basic convolution operations are performed with the data to be processed in sequence, and the results of each basic convolution operation are saved in the second intermediate-level memory unit. Merging processing is performed in the second intermediate-level memory unit according to a set format, and the merged result is saved as the result of the basic convolution processing at the position corresponding to the current output gradient sub-tensor.

9. The device according to claim 8, characterized in that The output unit is further used to notify the second intermediate level memory unit to release the results of each basic convolution operation after outputting the results of the basic convolution processing to the external memory.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed by a processor, the data processing method described in any one of claims 1 to 6 can be implemented.

11. A computer program product, characterized in that The method comprises computer executable instructions, which, when executed by a processor, implement the data processing method according to any one of claims 1 to 6.