Data processing method and device
By splitting data blocks in the intermediate level memory unit and performing basic convolution operations, the storage space requirements of non-1-step convolution modules are solved, and the execution efficiency and system performance of convolution processing are improved.
Patent Information
- Application Number
- CN202510420947.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, non-1-step convolution module processing requires a large amount of storage space, resulting in additional memory overhead and data read and write overhead, reducing the execution efficiency of convolution operations.
By splitting data blocks in the intermediate level memory unit and performing basic convolution operations, combining data merging and freeing of storage space, reducing the read and write data in external memory, and using the limited storage space of the intermediate level memory unit for data transit, improving execution efficiency.
It effectively improves the execution efficiency of convolution processing, reduces data reading and writing time and overhead, and improves the performance of convolution processing system.
Smart Images

Figure CN120336012A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and particularly to data processing methods and devices. Background Art
[0002] With the continuous development of artificial intelligence technology, neural network models are increasingly widely used.
[0003] Convolution is an operator often used in the training and inference of neural network models, usually implemented in a convolution module. The implementation method and execution efficiency of the convolution module become the key to improving the performance of the entire neural network model.
[0004] Convolution operations involve many optional configuration parameters: convolution kernel size, stride, dilation coefficient, padding, etc. Currently, in the processing of non-1 stride convolution modules, the requirements for the size of the storage space are relatively high, requiring additional memory overhead and additional data read / write overhead. Summary of the Invention
[0005] This application provides data processing methods and devices, which can effectively improve the execution efficiency of convolution processing.
[0006] To achieve the above object, this application adopts the following technical solutions:
[0007] A data processing method, comprising:
[0008] a. Determine the size of the data block to be split based on the stride of the current convolution operation and the available storage space of the first intermediate-level memory unit for storing the split data; receive the input tensor of the non-1 stride convolution module;
[0009] b. Based on the size of the data block to be split, select the current data to be split from the current unsplit input tensor in a set order for data splitting to obtain N groups of current sub-tensors; generate N groups of data to be processed by supplementing padding data to the N groups of split current sub-tensors in groups, and save them to the first intermediate-level memory unit; where N is a positive integer determined based on the stride and the data splitting method;
[0010] c. For the N groups of data to be processed, perform basic convolution processing on each group separately, merge the results of the basic convolution processing, and output them to the corresponding positions in the external memory for the current data to be split, and then release the N groups of data to be processed from the first intermediate-level memory unit; return to step b until all the input tensors are split.
[0011] Preferably, the merging of the results of the basic convolution operation includes:
[0012] Directly write the result of the basic convolution process of the first group of data to be processed after splitting to the specified storage location in the second intermediate-level memory unit, and accumulate and sum the result of the basic convolution process of other groups of data to be processed except the first group of data to be processed with the content currently stored at the specified storage location, and then write it to the specified storage location.
[0013] Preferably, after outputting to the external memory, the method further includes: releasing the data at the specified storage location from the second intermediate-level memory unit.
[0014] Preferably, the step of supplementing PADDING data for each of the N groups of current subtensors by group includes:
[0015] For each of the N groups of current subtensors, for any PADDING direction of any edge data of the current subtensor of the group, if there is data of the same type in the input tensor in the any PADDING direction of the any edge data, use the M data of the same type closest to the any edge data in the any PADDING direction as the PADDING data of the any edge data in the any PADDING direction; if there is no data of the same type in the input tensor in the any PADDING direction of the any edge data, set the PADDING data of the any edge data in the any PADDING direction to a preset value.
[0016] Wherein, the PADDING direction is the direction of the additional PADDING data relative to its adjacent edge data, the data of the same type is the data located at the same position after data splitting, and M is the number of PADDING data in the any PADDING direction.
[0017] Preferably, the step of performing basic convolution process for each group separately includes:
[0018] For each group of the data to be processed, use each weight subtensor obtained by splitting the weight tensor to perform basic convolution operations with the data to be processed in the group in sequence, and splice the operation results together according to a set format to obtain the result of the basic convolution process of the data to be processed in the group.
[0019] A data processing device includes: a receiving unit, a data splitting unit, a basic convolution unit, a data merging unit, and a first intermediate-level memory unit;
[0020] The receiving unit is configured to determine the size of the data block to be split based on the stride of the current convolution operation and the available storage space of the first intermediate-level memory unit for storing the split data, and is further configured to receive the input tensor of the convolution module with a stride other than 1.
[0021] The data splitting unit is configured to, based on the size of the data block to be split, select the current data to be split from the currently unsplit input tensor in a set order for data splitting; generate N sets of data to be processed by supplementing PADDING data for each of the N sets of current sub-tensors split out, and store them in the first intermediate-level memory unit; where N is a positive integer determined based on the stride and the data splitting method.
[0022] The basic convolution unit is configured to perform basic convolution processing on each of the N sets of data to be processed in units of groups.
[0023] The data merging unit is configured to merge the results of the basic convolution processing and output them to the corresponding positions in the external memory for the currently split data, and then notify the first intermediate-level memory unit to release the N sets of data to be processed; notify the data splitting unit to perform the next data splitting until all the input tensors are split.
[0024] The first intermediate-level memory unit is configured to store the N sets of current sub-tensors and release the N sets of current sub-tensors after receiving the notification from the data merging unit.
[0025] Preferably, the device further includes a second intermediate-level memory unit for storing the processing results of the basic convolution operation.
[0026] In the data merging unit, the merging of the results of the basic convolution operation includes:
[0027] Directly write the result of the basic convolution processing of the first set of data to be processed after splitting to the specified storage location in the second intermediate-level memory unit, and accumulate and sum the results of the basic convolution processing of the other sets of data to be processed except the first set of data to be processed with the currently stored content at the specified storage location and then write to the specified storage location.
[0028] Preferably, the data merging unit is further configured to notify the second intermediate-level memory unit to release the data at the specified storage location after merging the results of the basic convolution processing and outputting them to the corresponding positions in the external memory for the currently split data.
[0029] A computer-readable storage medium stores computer instructions thereon, and when the instructions are executed by a processor, the data processing method described in any one of the above can be implemented.
[0030] A computer program product includes computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data processing method described in any one of the above is implemented.
[0031] As can be seen from the above technical solutions, the present application performs convolution processing for a non-1 stride. First, based on the stride of the current convolution operation and the available storage space of the first intermediate-level memory unit for storing the split data, the size of the data block to be split is determined, and the input tensor of the non-1 stride convolution module is received; in the input tensor, based on the size of the data block to be split, the current data to be split is sequentially selected from the unsplit data in a set order for performing a sub-convolution (sub conv) operation, and the sub conv operation is looped multiple times until all the input tensors are split and processed; in each sub conv operation, the selected current data block to be split is split and PADDING data is supplemented and then stored in the first intermediate-level memory unit; the N groups of current sub-tensors stored are respectively subjected to a basic convolution operation in units of groups, and the results of the N groups of convolution operations are merged and output to the corresponding position in the external memory for the current data to be split, and then the storage space in the first intermediate-level memory unit for storing the N groups of current sub-tensors is released, and thus a sub conv operation is completed. After looping the sub conv operation multiple times, the complete convolution operation of the input tensor and the output of the convolution result are completed. In the above processing, by storing the split data in the first intermediate-level memory unit in each sub conv operation and releasing the corresponding storage space after the sub conv operation, on the one hand, the convolution operation with a non-1 stride is completed through looping the sub conv operation multiple times, and on the other hand, although the storage space of the intermediate-level memory unit is limited, the function of storing the split data is realized through the cyclic use of the storage space of the corresponding memory unit, so that the intermediate-level memory unit closer to the core processing part of the convolution module is used for data storage during the convolution operation, greatly shortening the time required for data access and effectively improving the execution efficiency of the convolution processing, and further effectively improving the performance of the processing system where the convolution processing is located. Description of the Drawings
[0032] Figure 1 It is a schematic diagram of the basic process of the data processing method in the present application;
[0033] Figure 2 It is a schematic diagram of each sub conv operation in the present application;
[0034] Figure 3Schematic diagram for data block division;
[0035] Figure 4 Schematic diagram for splitting the current data to be split using a specific data splitting method (knit split) in this application;
[0036] Figure 5 Schematic diagram for splitting an input tensor using a specific data splitting method (knit split) in the related art;
[0037] Figure 6 Schematic diagram for data merging and data output;
[0038] Figure 7a Schematic diagram for the arrangement of the original input tensor data;
[0039] Figure 7b Schematic diagram for the data arrangement of 4 different data groups after knit split according to the method of Method A;
[0040] Figure 7c Schematic diagram for the knit split data splitting method of the convolutional kernel;
[0041] Figure 7d Schematic diagram for the process of performing the corresponding subconv on a group of current subtensors obtained after knit split in this application;
[0042] Figure 8 Schematic diagram for the basic structure of the convolutional module processing device in this application. Detailed implementation manners
[0043] In order to make the purpose, technical means and advantages of this application clearer, the following further elaborates on this application with reference to the accompanying drawings.
[0044] For non-unit stride convolutional operations, currently it can be achieved through a combination of basic Conv operators (hereinafter referred to as Method A). Specifically, through data preprocessing, the data to be processed is pre-divided into sub-data suitable for the basic Conv operator. Considering that the memory size at the intermediate level may be limited and unable to completely store the divided sub-data, the divided sub-data is completely stored in an external memory (such as global memory), and then the complete non-unit stride Conv operation is achieved by executing the basic Conv operator multiple times. However, since the divided sub-data used in the processing is stored in the external memory, this will result in additional memory overhead and additional data read / write overhead.
[0045] Considering the above memory overhead problem, in the present application, convolution processing is improved on the basis of the above processing method, and an intermediate-level memory unit is used for data transfer, so as to effectively reduce the time for data reading and writing during convolution processing, thereby improving the execution efficiency of convolution processing and the performance of the system where the convolution processing is located.
[0046] Figure 1 This is a schematic diagram of the basic process of the data processing method in the present application. As Figure 1 shown, the method includes:
[0047] Step 101, determine the size of the data block to be split based on the stride of the current convolution operation and the available storage space of the first intermediate-level memory unit; receive the input tensor of the convolution module with a non-1 stride.
[0048] The processing of the present application is carried out for the convolution processing with a non-1 stride, that is, the processing of the convolution module with a non-1 stride is optimized. This convolution processing can be carried out in the convolution processing modules of various related processing systems (such as neural network models). Therefore, this step receives the input tensor of such a convolution module with a non-1 stride. Among them, the input tensor is the input data of the corresponding convolution module, and specifically can be various physical data in the neural network model. For example, it can be the image data processed in the image neural network model, or it can also be the text data processed in the text neural network model, etc.
[0049] In the implementation manner of the non-1 stride convolution operation mentioned above, since the amount of data after splitting is relatively large, an external memory needs to be used to store the data after splitting. As a result, during the convolution operation processing, data reading and writing need to be frequently performed in the external memory, resulting in additional memory overhead and data reading and writing overhead, and then causing the execution efficiency of the convolution operation to be relatively low. Based on this, in the present application, an intermediate-level memory unit is used to store the data after splitting. Since the intermediate-level memory unit is closer to the core processing of the convolution operation than the external memory, it can effectively save the data reading and writing time and overhead, and improve the execution efficiency of the convolution operation.
[0050] However, the storage space of the intermediate-level memory unit is limited and cannot completely store the data after splitting the input tensor. Based on this, in the present application, the input tensor is split into multiple times for data splitting, and only part of the data in the input tensor is split each time to reduce the amount of data after splitting, so as to ensure that the intermediate-level memory unit can provide sufficient storage space to store the data after splitting. This step is used to determine the size of the data block for each data splitting, that is, to determine how much data in the input tensor is split each time, that is, the size of the data block to be split.
[0051] Among them, the size of the data block to be split can be determined based on the stride of the current convolution operation and the available storage space of the intermediate-level memory unit, ensuring that the split data of a single data block to be split can be stored in the available storage space of the intermediate-level memory unit. In practical applications, the size of the data block to be split can be designed according to actual requirements. For example, preferably, the determined size of the data block to be split can ensure that: the split data obtained by splitting a single data block to be split according to the stride of the current convolution operation can be completely stored in the available storage space of the intermediate-level memory unit, and the remaining space is as small as possible. More specifically, on the one hand, in order to complete the convolution operation as soon as possible, the available storage space of the intermediate-level memory unit is used as much as possible to store the split data, so that the number of loops is the least and the time used is the shortest; on the other hand, the stride of the current convolution operation is related to the data splitting method and determines how many groups of split data there are; thus, considering the above two factors comprehensively, determining the size of the data block to be split can ensure that the split data of a single data block to be split can be completely stored in the intermediate-level memory unit and the remaining space is the smallest.
[0052] The process of specifically determining the size of the data block to be split can be completed in advance before the entire convolution process, and even can be completed before all processes of the system where the convolution process is located (such as a neural network model). Of course, it can also be completed after receiving the input tensor of the convolution module. This application does not limit the specific position of this operation.
[0053] In addition, the intermediate-level memory unit can be various storage units that are closer to the convolution process than the external memory, such as shared buffer or shared memory. For the convenience of description, the intermediate-level memory unit used to store the split data is hereinafter referred to as the first intermediate-level memory unit, or tensor buffer, to distinguish it from other intermediate-level memory units.
[0054] Step 102: Based on the size of the data block to be split, select the current data to be split from the current unsplit input tensor in a set order, and perform data splitting to obtain N groups of current sub-tensors.
[0055] From this step to one execution of step 105, a sub conv operation is completed. As mentioned above, in order to use the intermediate-level memory unit with limited storage space to store the split data, each sub conv operation reads part of the data in the input tensor, performs data splitting, basic convolution operation, data merging and data storage, as Figure 2 shown.
[0056] In the first sub-conv operation (i.e., the first loop processing in steps 102 to 105), part of the data is selected from the complete input tensor as the current data to be split based on the size of the data block to be split; while in subsequent sub-conv operations, data is selected from the remaining unsplit input tensor as the current data to be split. For example, if the size of the data block to be split is 1 / 4 of the size of the input tensor data, then in the first sub-conv operation, the first 1 / 4 of the data is selected, as shown by the data block with serial number 1 in Figure 3 ; in the second sub-conv operation, another 1 / 4 of the data is selected from the remaining unsplit data. The specific selection method can be carried out in a pre-set order. For example, the data blocks with serial numbers 2 or 3 as shown in Figure 3 can be selected, and so on, until in the fourth sub-conv operation, the last 1 / 4 of the data is selected, as shown by the data block with serial number 4 in Figure 3 .
[0057] For each sub-conv operation, the current data to be split is read into a register and split according to the set splitting rules to obtain N groups of current sub-tensors, where N is a positive integer determined according to the convolution stride and the data splitting method. The specific data splitting method can be carried out in various feasible ways. For example, a specific data splitting method (knit split) is used to split the current data to be split. Here, the knit split varies according to the stride parameter of the convolution operation. Taking the scenario with stride stride = 2 as an example, for the current data to be split, it can be split into 4 groups of current sub-tensors according to the parity characteristics of the data coordinates, that is, N = 4, as shown in Figure 4 . That is to say, after the convolution stride and the data splitting method are determined, N is a determined value. The specific corresponding relationship between the convolution stride, the data splitting method, and N is well-known in the art.
[0058] Compared with the data splitting in the present application above, in the non-1-step convolution operation carried out by the aforementioned method A, data splitting is carried out on the complete input tensor. Still taking the scenario with stride stride = 2 as an example, 4 groups of sub-tensors are also split, as shown in Figure 5 . Among them, since data splitting is carried out on the complete input tensor, the amount of data split is much larger than the 4 groups of current sub-tensors split in the present application and cannot be completely stored in the intermediate-level memory unit and can only be stored in the external memory, which requires consuming more resources to obtain data and the execution efficiency of the convolution process is not high. While in the present application, part of the input tensor is selected as the data to be split and saved in the nearby intermediate-level memory unit, effectively improving the execution efficiency of the convolution process.
[0059] Step 103: Generate N groups of data to be processed by supplementing PADDING data to the N groups of current subtensors in groups, and save them in the first intermediate-level memory unit.
[0060] The N groups of current subtensors split out need to perform basic conv operations in groups subsequently. Therefore, it is necessary to supplement PADDING data to each group of current subtensors to obtain a group of data to be processed, and save the N groups of data to be processed in the first intermediate-level memory unit in groups.
[0061] The specific way to supplement PADDING data can be carried out according to the existing method. In addition, to ensure the accuracy of the convolution operation, the present application provides a preferred way to supplement PADDING data. Among them, for each group of current subtensors, the way to supplement PADDING data is the same. Taking the processing of a group of current subtensors as an example, the way to supplement PADDING data provided by the present application is described as follows:
[0062] 1) Determine the PADDING direction for any edge data A of the current subtensor;
[0063] The PADDING direction here refers to the direction of supplementing PADDING data, that is, the direction of the additional PADDING data relative to its adjacent edge data. For example, Figure 4 for the upper left edge data of the orange data block in, there are three PADDING directions: left, upper left, and above. For the edge data in the first row and second column of the orange data block, there is only one PADDING direction: above;
[0064] 2) For any PADDING direction B, judge whether there is the same type of data in each PADDING direction of the edge data A in the input tensor; if so, execute step 3), otherwise execute step 4);
[0065] The same type of data refers to the data located in the same position after data splitting. For example, Figure 4 and Figure 5 in, the data on all the same-color data blocks are the same type of data;
[0066] In the non-one-step convolution operation performed by the aforementioned method A, the data of the complete input tensor is split at one time. Therefore, data of the same type is split into the same data block. When padding data is added to this data block, the padding data are all newly added data and can be set to a preset value, such as 0. In the processing of the present application, however, each data split is performed on a part of the input tensor. Therefore, data of the same type is split into different data blocks. Then, for some data blocks, the edge data is not actually the real edge data from the perspective of the complete input tensor, and there are adjacent data of the same type. The following will refer to this type of edge data as the first type of edge data. For some data blocks, the edge data is also the edge data from the perspective of the complete input tensor. The following will refer to this type of edge data as the second type of edge data. For example, if Figure 4 are the 4 groups of current sub-tensors obtained after the first subconv operation, then the edge data in the lower right corner of the orange data block is not the real edge data from the perspective of the complete input tensor, while the edge data in the upper left corner of the orange data block is also the edge data from the perspective of the complete input tensor;
[0067] In the current sub-tensor, for the above two different types of edge data, different strategies are adopted when adding padding data; the first type of edge data is padded with padding data using the processing method of step 3), and the second type of edge data is padded with padding data using the processing method of step 4);
[0068] 3) Use the M data of the same type that are closest to the edge data A in the padding direction B as the padding data of the edge data A in the padding direction;
[0069] That is, for the first type of edge data, use the adjacent data of the same type as the padding data. Such processing can ensure the accuracy of the convolution operation; where M is the preset number of padding data added in the padding direction B;
[0070] 4) Set the padding data of the edge data A in the padding direction B to a preset value;
[0071] That is, for the second type of edge data, use the existing method of adding padding data, that is, set the padding data to a preset value, such as 0.
[0072] By the method of adding padding data provided by the present application above, it can be ensured that the basic convolution processing performed on the current sub-tensor can accurately simulate the movement process of the convolution kernel inside the input tensor and ensure the accuracy of the convolution operation.
[0073] Step 104: For N groups of data to be processed, perform basic convolution processing on each group separately, and merge the results of the basic convolution processing.
[0074] As Figure 2 shown, for each group of data to be processed, perform basic convolution processing in sequence, and merge the results of the basic convolution processing of N groups.
[0075] The specific basic convolution processing can be carried out in the existing manner. Specifically, for each group of data to be processed, each weight subtensor obtained by splitting the weight tensor can be used to perform basic convolution operations with the data of this group in sequence, and the operation results are spliced together in a set format to obtain the result of the basic convolution processing of this group of data to be processed.
[0076] When merging the results of the basic convolution processing of multiple groups, the result merging method after data splitting in the aforementioned method A can be followed. Specifically, for the first group of data to be processed, the result after its basic convolution processing is directly written to the specified storage location in the storage unit; for the second group of data to be processed, the result after its basic convolution processing is added to and summed with the data saved at the specified storage location in the storage unit and then saved at this specified storage location; for the subsequent groups of data to be processed, the basic convolution processing results are processed in the same way as the second group of data to be processed until the basic convolution processing results of the last group of data to be processed are completed. Then, the data saved at the specified storage location in the storage unit is the convolution processing result corresponding to the current data to be split. Among them, to further improve the execution efficiency of the convolution processing, the storage unit can be an intermediate-level memory unit, which is called the second intermediate-level memory unit here. In addition, the first intermediate-level memory unit and the second intermediate-level memory unit can be the same storage unit or different storage units; if the two are the same storage unit, different storage spaces can be pre-allocated to save the split data and the results of the basic convolution processing respectively.
[0077] Step 105: Output the data merging result to the corresponding position in the external memory for the current data to be split, and then release the N groups of data to be processed from the first intermediate-level memory unit.
[0078] In this step, the data merging result is output to an external memory. The external memory is used to store the convolution processing result of the input tensor. Among them, the data merging result obtained in step 104 is the convolution processing result of the current data to be split for the current sub-conv operation. When storing this convolution processing result in the external memory, it is stored at the storage location corresponding to the current data to be split. Here, the storage location corresponding to the current data to be split is a storage space set in advance for each data to be split to store the corresponding convolution result. Specifically, the relationship between the storage locations corresponding to different data to be split is consistent with the selection order of the current data to be split from the input tensor in the aforementioned step 102. In this way, after multiple loops, the convolution processing results of each time can be stored at the corresponding storage locations to form the complete convolution processing result of the input tensor.
[0079] As Figure 6 shown in the schematic diagram of data merging and data output, where the thick solid lines represent data directly written into storage units, the dashed lines represent data accumulated and summed and written into storage units, and the small squares in the external memory represent the storage locations corresponding to the current data to be split. In this way, after multiple loops, the convolution processing results obtained in each loop are sequentially output to the corresponding storage locations, and finally the complete convolution processing result of the input tensor can be spliced.
[0080] After outputting the convolution processing result of the current sub-conv operation, it means that the sub-conv operation on the current data to be split is completed, and the storage space in the first intermediate-level memory unit for storing N groups of data to be processed and the storage space in the second intermediate-level memory unit for storing the merged result of the basic convolution operation can be released. Thus, in the next loop, the corresponding storage space can be reused to store the split data and the merged result of the basic convolution operation in the next loop. In this way, by recycling the first intermediate-level memory unit and the second intermediate-level memory unit, the split data and the merged result of the basic convolution operation in the entire convolution processing operation of this application can be stored, improving the execution efficiency of the convolution operation, and further improving the performance of the entire system where the convolution processing is located (such as a neural network model).
[0081] So far, one loop processing of steps 102 to 105 ends, that is, a sub-conv operation as Figure 2 shown is performed.
[0082] Step 106, return to step 102 to perform the next sub-conv operation until the entire input tensor is split.
[0083] Enter the sub-conv operation of the next cycle through this step. The loop end condition is that all input tensors are split, that is, all data input to the convolutional module this time has been processed, and all convolutional operations have been completed, and the process ends.
[0084] The above is the specific implementation of the data processing method in this application. This method can be applied, for example, in the convolutional processing module of a neural network model. In the processing of this application above, by storing the split data of multiple data splits in the middle-level memory unit loop, the storage locations of the split data are brought closer, effectively improving the execution efficiency of the convolutional operation, and thus improving the processing performance of the entire system where the convolutional processing is located (such as a neural network model).
[0085] To more clearly illustrate the method of storing the split data in this application, the following shows the specific methods of storing the split data in this application and in the previous method A processing in a comparative manner. Among them, the case of conv2d, kernel size = 3, padding = 1, stride = 2 is used as an example for illustration. Figure 7a Shows the arrangement of the original input tensor data. After being split into different data groups through knit split, the data points in different data groups are represented by different colors, and the dotted line on the outer circle represents the padding data; Figure 7b Shows for Figure 7a the input tensor shown, the data arrangements of the 4 different data groups after knit split according to the method of the previous method A. Similarly, the outer circle of each data group is the padding data; Figure 7c Shows the corresponding knit split data splitting method of the convolutional kernel. Taking the convolutional kernel (kernel size) of (3, 3) as an example, it will be split into four sub-kernels of (2, 2), (2, 1), (1, 2), and (1, 1) after knit split; Figure 7d Shows the process of performing the corresponding sub-conv on a group of current sub-tensors obtained after being split by knit split in this application.
[0086] Due to the characteristics of the Conv operation itself, in addition to reading the data block of the convolution itself, it is also necessary to read the padding data additionally. As Figure 7b shown, in the method of the previous method A, the complete split data will be stored in the external memory, and for the outer circle padding (shown as the dotted line in the figure), additional processing is required. And for the padding involved in the movement of the internal convolutional kernel, the hardware can automatically complete the reading of the relevant data.
[0087] In the method proposed in this application, since the intermediate-level memory cells cannot store the data after splitting the entire input tensor completely, additional processing is also required when dealing with the outer padding; when dealing with the original internal padding, for example Figure 7d when padding data points 4, 12, 20, 28, etc. shown, this part of the internal padding may not be in the corresponding data saved in the current intermediate-level memory cell. Therefore, it is necessary to read an additional part of the internal padding data into the intermediate-level memory cell as padding. Take Figure 7d as an example. The actual data size to be processed is 4x4, but the intermediate-level memory cell needs to store data of (4 + padding * 2) * (4 + padding * 2).
[0088] Correspondingly, each time a sub conv operation performs a partial convolution operation, it will use (padding, padding) as the starting coordinate of the data, (4 + padding, 4 + padding) as the ending coordinate, slide the convolution window and perform the corresponding convolution operation.
[0089] The above is the specific implementation of the data processing method in this application. This application also provides a data processing device that can be used to implement the above processing method. Figure 8 This is the basic structure schematic diagram of the data processing device in this application. As Figure 8 shown, the device includes: a receiving unit, a data splitting unit, a basic convolution unit, a data merging unit, and a first intermediate-level memory cell.
[0090] Among them, the receiving unit is used to determine the size of the data block to be split based on the stride of the current convolution operation and the available storage space of the first intermediate-level memory cell for storing the split data; it is also used to receive the input tensor of the convolution module with a non-1 stride;
[0091] The data splitting unit is used to select the current data to be split from the current unsplit input tensor in a set order based on the size of the data block to be split; after supplementing PADDING data to the N groups of current sub-tensors split out in groups, generate N groups of data to be processed and save them to the first intermediate-level memory cell; where N is a positive integer determined based on the stride and the data splitting method;
[0092] The basic convolution unit is used to perform basic convolution operations on the N groups of data to be processed in groups;
[0093] A data merging unit, configured to merge the results of the convolution operations and output them to the corresponding positions in the external memory for the currently to-be-split data, and then notify the first intermediate-level memory unit to release N groups of to-be-processed data; notify the data splitting unit to perform the next data splitting until all the input tensors are split.
[0094] The first intermediate-level memory unit is configured to store N groups of current sub-tensors and release the N groups of current sub-tensors after receiving the notification from the data merging unit.
[0095] Optionally, the device may further include a second intermediate-level memory unit, configured to store the processing results of the basic convolution operations.
[0096] In the data merging unit, the processing of merging the results of the basic convolution operations may specifically include:
[0097] Directly write the results of the basic convolution operations of the first group of to-be-processed data after splitting into the specified storage location in the second intermediate-level memory unit, and accumulate and sum the results of the basic convolution operations of the other groups of to-be-processed data except the first group of to-be-processed data with the currently stored content at the specified storage location and then write it into the specified storage location.
[0098] Optionally, after outputting to the external memory, the method further includes: releasing the data at the specified storage location from the second intermediate-level memory unit.
[0099] Optionally, in the data splitting unit, the processing of supplementing PADDING data for the N groups of current sub-tensors in groups may specifically include:
[0100] For each of the N groups of current sub-tensors, for any PADDING direction of any edge data of the group of current sub-tensors, if there is the same type of data in any PADDING direction of any edge data in the input tensor, then use the same type of data closest to any edge data in any PADDING direction as the PADDING data of any edge data in any PADDING direction; if there is no same type of data in any PADDING direction of any edge data in the input tensor, then set the PADDING data of any edge data in any PADDING direction to a preset value.
[0101] Wherein, the PADDING direction is the direction of the additional PADDING data relative to its adjacent edge data, and the same type of data is the data located at the same type of position after data splitting.
[0102] Optionally, the processing of performing the basic convolution operations in groups may specifically include:
[0103] For each group of data to be processed, each weight subtensor obtained by splitting the weight tensor is used to perform a basic convolution operation with the group of data to be processed in sequence, and the operation results are concatenated together in a set format to obtain the result of the basic convolution operation for the group of data to be processed.
[0104] The data processing method or data processing device provided by at least one embodiment of the present application can be applied to different systems or devices, such as being applied to an electronic device. The electronic device can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an AR device, a VR device, a vehicle-mounted terminal, etc., or can also be a server, etc. The data processing method provided by at least one embodiment of the present application can be applied to scenarios related to data processing such as CPU, high-performance computing (HPC), and artificial intelligence (AI) in an electronic device, such as a convolution processing unit. Of course, the present application is not limited thereto, and any scenario, device, device, etc. involving non-1-step convolution processing can adopt the data processing method or data processing device provided by at least one embodiment of the present application.
[0105] In some embodiments, the data processing device provided by at least one embodiment of the present application can be a chip, for example, the chip is a system-on-a-chip (SoC). The system-on-a-chip includes a processor, and the processor can be a single-core processor or a multi-core processor, a memory, and an I / O interface, etc. After the processor loads the data and application programs in the memory, it processes the data, such as performing non-1-step convolution processing.
[0106] For example, when an electronic device implements the data processing method provided by at least one embodiment of the present application, the processing device included therein performs various appropriate actions and processes according to the non-temporary computer-readable instructions stored in the memory to implement the data processing method. For example, the input tensor is stored in the memory, such as a register, a cache, or a memory. When non-1-step convolution processing is required, the input tensor is read from the memory, and the processing device processes the input tensor according to steps 101 to 106 in the data processing method disclosed by at least one embodiment of the present application to obtain a convolution result. The convolution result can be transmitted to the corresponding operator for use again, or transmitted to the memory (such as a high-bandwidth memory) for storage.
[0107] In addition, it should be noted that the data type of the data processed by the data processing method or data processing device provided by at least one embodiment of the present application can have different specific physical meanings according to different application scenarios. For example, the data processing method provided by at least one embodiment of the present application can be applied in fields such as speech processing, image processing, text processing, and video processing.
[0108] For example, in the field of speech processing, the data to be processed can be any parameter used, input, or generated in tasks such as feature extraction, speech enhancement, and speech recognition, such as speech feature vectors, filtering parameters, etc.
[0109] For example, in the field of image processing, the data to be processed can be image preprocessing, feature extraction, image segmentation, and object detection.
[0110] The present application also provides a computer-readable storage medium that stores instructions, which can execute the steps in the data processing method as described above when executed by a processor. In practical applications, the computer-readable medium can be included in each device / device / system of the above embodiments, or can exist separately without being assembled into the device / device / system. Among them, instructions are stored in the computer-readable storage medium, and the stored instructions can execute the steps in the data processing method as described above when executed by a processor.
[0111] According to the embodiments disclosed in the present application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, for example, it can include but is not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above, but is not used to limit the scope of protection of the present application. In the embodiments disclosed in the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used or combined with an instruction execution system, device, or device.
[0112] The present application also provides a computer program product, including computer-executable instructions, which implement the data processing method as described above when executed by a processor.
[0113] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A data processing method, characterized in that including: a. determining the size of the data block to be split based on the stride of the current convolution operation and the available storage space of the first intermediate-level memory unit for storing the split data; receiving the input tensor of the convolution module with a non-1 stride; b. based on the size of the data block to be split, selecting the current data to be split from the currently unsplit input tensor in a set order for data splitting to obtain N groups of current sub-tensors; generating N groups of data to be processed by supplementing PADDING data to the N groups of split current sub-tensors in units of groups and storing them in the first intermediate-level memory unit; where N is a positive integer determined based on the stride and the data splitting method; c. for the N groups of data to be processed, performing basic convolution processing on each group in units of groups, merging the results of the basic convolution processing, and outputting them to the corresponding positions in the external memory for the current data to be split, and then releasing the N groups of data to be processed from the first intermediate-level memory unit; returning to step b until all the input tensors are split.
2. The method according to claim 1, wherein The merging of the results of the basic convolution operation includes: directly writing the result of the basic convolution processing of the first group of data to be processed after splitting to the specified storage location in the second intermediate-level memory unit, and adding and summing the results of the basic convolution processing of the other groups of data to be processed except the first group of data to be processed with the content currently saved at the specified storage location and then writing it to the specified storage location.
3. The method according to claim 2, wherein After the output to the external memory, the method further includes: releasing the data at the specified storage location from the second intermediate-level memory unit.
4. The method according to claim 1, characterized in that The supplementing of PADDING data to the N groups of split current sub-tensors in units of groups includes: for each of the N groups of current sub-tensors, for any PADDING direction of any edge data of the group of current sub-tensors, if there is the same type of data in the input tensor in the any PADDING direction of the any edge data, taking the M nearest same type of data in the any PADDING direction to the any edge data as the PADDING data of the any edge data in the any PADDING direction; if there is no same type of data in the input tensor in the any PADDING direction of the any edge data, setting the PADDING data of the any edge data in the any PADDING direction to a preset value; where the PADDING direction is the direction of the additional PADDING data relative to its adjacent edge data, the same type of data is the data located at the same type of positions after data splitting, and M is the number of PADDING data in the any PADDING direction.
5. The method according to claim 1, characterized in that, The performing of basic convolution processing on each group in units of groups includes: For each group of the data to be processed, use each weight sub-tensor obtained by splitting the weight tensor to perform a basic convolution operation with the data of this group in sequence, and splice the operation results together in a set format to obtain the result of the basic convolution processing of the data of this group.
6. A data processing device, characterized in that, It includes: a receiving unit, a data splitting unit, a basic convolution unit, a data merging unit, and a first intermediate-level memory unit; The receiving unit is used to determine the size of the data block to be split based on the stride of the current convolution operation and the available storage space of the first intermediate-level memory unit for storing the split data; and is further used to receive the input tensor of the convolution module with a non-1 stride. The data splitting unit is used to select the current data to be split from the currently unsplit input tensor in a set order based on the size of the data block to be split; generate N groups of data to be processed by supplementing PADDING data to the N groups of current sub-tensors split out in groups, and save them to the first intermediate-level memory unit; where N is a positive integer determined based on the stride and the data splitting method. The basic convolution unit is used to perform basic convolution processing on the N groups of data to be processed in groups. The data merging unit is used to merge the results of the basic convolution processing and output them to the corresponding position in the external memory for the currently split data, and then notify the first intermediate-level memory unit to release the N groups of data to be processed; notify the data splitting unit to perform the next data splitting until all the input tensors are split. The first intermediate-level memory unit is used to store the N groups of current sub-tensors and release the N groups of current sub-tensors after receiving the notification from the data merging unit.
7. The device according to claim 6, wherein The device further includes a second intermediate-level memory unit for storing the processing results of the basic convolution operation. In the data merging unit, the merging of the results of the basic convolution operation includes: Directly write the result of the basic convolution processing of the first group of data to be processed after splitting to the specified storage location in the second intermediate-level memory unit, and accumulate and sum the results of the basic convolution processing of the other groups of data to be processed except the first group of data to be processed with the content currently stored at the specified storage location and then write it to the specified storage location.
8. The device according to claim 7, characterized in that, The data merging unit is further used to notify the second intermediate-level memory unit to release the data at the specified storage location after outputting the merged results of the basic convolution processing to the corresponding position in the external memory for the currently split data.
9. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the instruction is executed by a processor, it can implement the data processing method described in any one of claims 1 to 5.
10. A computer program product, characterized in that, It includes computer-executable instructions, and when the computer-executable instructions are executed by a processor, they implement the data processing method described in any one of claims 1 to 5.