Data processing method, electronic device, storage medium and program product
By converting the deconvolution operation into a convolution operation, the existing convolution calculation unit is used to perform deconvolution operations, which solves the problem of complex design of the deconvolution calculation unit in the prior art, reduces development costs and improves data processing efficiency.
Patent Information
- Application Number
- CN202510328128.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, it is difficult to adapt to the structure and parameters of different deconvolution neural networks, resulting in increased implementation complexity and development costs.
By converting the deconvolution operation into a convolution operation, the deconvolution operation is performed using the existing convolution calculation unit, and the target convolution kernel is determined based on the parameter information of the first convolution kernel and the deconvolution operation.
It reduces the complexity and development cost of deconvolution operations, provides feasibility under limited memory and computing resources, and improves data processing efficiency.
Smart Images

Figure CN120179971A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and particularly to a data processing method, an electronic device, a storage medium, and a program product. Background Art
[0002] In a deep neural network, deconvolution is a common neural network structure for extracting spatial information, which is used to restore a low-dimensional feature map to a high-dimensional space, and is usually applied to tasks such as image generation, super-resolution reconstruction, and semantic segmentation.
[0003] In some existing technologies, a deconvolution calculation unit needs to be designed separately to implement deconvolution operations. However, different deconvolution neural networks may have different structures and parameters, and the separately designed deconvolution calculation unit may be difficult to meet the requirements of different tasks, thus increasing the implementation complexity and development cost. Summary of the Invention
[0004] To solve the above technical problems, embodiments of this application provide a data processing method, an electronic device, a storage medium, and a program product. The following introduces this application from multiple aspects, and the implementation manners and beneficial effects of the following multiple aspects can be referred to each other.
[0005] In a first aspect, an embodiment of this application provides a data processing method, including: obtaining data to be processed, where the data to be processed is data to be subjected to a deconvolution operation; performing a convolution operation on the data to be processed based on a target convolution kernel to obtain first output data; performing a memory rearrangement operation on the first output data to obtain target output data; where the target convolution kernel is determined based on a first convolution kernel and parameter information of the deconvolution operation, and the target output data is the same as the output data obtained by performing a deconvolution operation on the data to be processed based on the first convolution kernel.
[0006] The data processing method provided by the embodiments of this application can convert a deconvolution operation into a convolution operation, providing feasibility for performing a deconvolution operation under the condition of limited memory and computing resources of an electronic device. The electronic device can use an existing convolution calculation unit to implement a deconvolution operation without separately designing a deconvolution operation unit, reducing the implementation complexity and development cost of the deconvolution operation.
[0007] In a possible implementation of the first aspect, the data of the first convolution kernel includes a width dimension, a height dimension, an input channel dimension, and an output channel dimension. The parameter information of the deconvolution operation includes the width of the first convolution kernel, the height of the first convolution kernel, the stride of the width dimension, the stride of the height dimension, and the number of input channels. Moreover, the first convolution kernel includes at least one output channel, and each output channel of the first convolution kernel includes at least one second convolution kernel, and the number of the at least one second convolution kernel is equal to the number of input channels. When the stride of the width dimension is greater than 1 or the stride of the height dimension is greater than 1, the target convolution kernel is determined in the following manner: For each output channel of the first convolution kernel, based on the width of the first convolution kernel, the height of the first convolution kernel, the stride of the width dimension, and the stride of the height dimension, the second convolution kernel is divided into M first sub-blocks. Wherein, the width of the first sub-block is equal to the stride of the width dimension, the height of the first sub-block is equal to the stride of the height dimension, the positions of the data elements of each first sub-block among the M first sub-blocks correspond one by one, and M is an integer greater than or equal to 1. Based on the data elements at the corresponding positions in the M first sub-blocks, N second sub-blocks are determined. The N second sub-blocks are respectively horizontally flipped and vertically flipped along the first plane to determine N sub-convolution kernels. Wherein, the first plane is the plane formed by the width dimension and the height dimension. The N sub-convolution kernels corresponding to each second convolution kernel in each output channel of the first convolution kernel are concatenated along the output channel dimension to obtain the target convolution kernel.
[0008] It can be understood that if k h represents the height of the first convolution kernel, k w represents the width of the first convolution kernel, s h represents the stride of the height dimension, s w represents the stride of the width dimension, c in represents the number of input channels of the first convolution kernel, then M=(k h / s h )*(k w / s w ), and each output channel of the first convolution kernel includes c in second convolution kernels. For each second convolution kernel in each output channel of the first convolution kernel, it is necessary to first divide it into M first sub-blocks, then determine N second sub-blocks based on the data elements at the corresponding positions in each of the M first sub-blocks, and then horizontally flip and vertically flip the N second sub-blocks along the first plane to determine N sub-convolution kernels, where N is equal to the product of the stride of the width dimension and the stride of the height dimension (s w *s h ).
[0009] In a possible implementation of the first aspect, the first sub-block includes N data elements, and the N data elements are respectively located at different positions in the first sub-block, where N is an integer greater than or equal to 1; determining N second sub-blocks based on the data elements at the corresponding positions in the M first sub-blocks includes: placing the data elements at the corresponding positions among the N data elements of each first sub-block in the M first sub-blocks in the same second sub-block to determine N second sub-blocks; wherein, the positions of the N data elements in the second sub-block correspond to the positions of the first sub-blocks where the N data elements are located in the second convolution kernel.
[0010] It can be understood that after dividing the second convolution kernel, each second convolution kernel (i.e., at least one second convolution kernel in each output channel of the first convolution kernel) is divided into M first sub-blocks, and each output channel of the first convolution kernel includes c in * M first sub-blocks. If c out represents the number of output channels of the first convolution kernel, then for all output channels of the first convolution kernel, there are c out * c in * M first sub-blocks.
[0011] In a possible implementation of the first aspect, dividing the second convolution kernel into M first sub-blocks based on the width of the first convolution kernel, the height of the first convolution kernel, the width dimension stride, and the height dimension stride includes: when the width dimension stride cannot be divided evenly by the width of the first convolution kernel, performing data padding on the width dimension of the second convolution kernel to divide the second convolution kernel into M first sub-blocks; when the height dimension stride cannot be divided evenly by the height of the first convolution kernel, performing data padding on the height dimension of the second convolution kernel to divide the second convolution kernel into M first sub-blocks.
[0012] It can be understood that to ensure the accuracy of the calculation result, when performing data padding (i.e., filling) on the width dimension or the height dimension of the second convolution kernel, the filled data is 0. Moreover, performing data padding on the width dimension or the height dimension of the second convolution kernel can ensure that after dividing the second convolution kernel into M first sub-blocks, the width of each first sub-block is equal to the width dimension stride, and the height of each first sub-block is equal to the height dimension stride. If the padded second convolution kernel is denoted as kernel_pad, the size of kernel_pad is (d h * s h , d w * s w ), where d h = ceil(k h / s h ), d w = ceil(k w / s w), where ceil is the ceiling function. Assume that k w = k h = 8, s w = s h = 3, then d h = ceil(k h / s h ) = 3, d w = ceil(k w / s w ) = 3, the size of kernel_pad is (9, 9), M = d h * d w = 9.
[0013] In a possible implementation of the first aspect, when the stride in the width dimension and the stride in the height dimension are both equal to 1, the target convolution kernel is determined as follows:
[0014] For each output channel of the first convolution kernel, the second convolution kernel is horizontally and vertically flipped along the first plane to obtain the target convolution kernel.
[0015] It can be understood that when the stride in the width dimension and the stride in the height dimension are both equal to 1, there is no need to divide the second convolution kernel.
[0016] In a possible implementation of the first aspect, the target convolution kernel includes Q sub-convolution kernels, where Q is equal to the product of the number of output channels, the number of input channels of the first convolution kernel, and N; the above-mentioned convolution operation on the data to be processed based on the target convolution kernel to obtain the first output data includes: performing a convolution operation on the data to be processed in parallel based on the target parameter information and Q sub-convolution kernels to obtain the first output data; where the target parameter information is determined based on the parameter information of the deconvolution operation.
[0017] It can be understood that in the embodiments of the present application, since the convolution operations of the Q sub-convolution kernels and the data to be processed are run in parallel, that is, the time-consuming of this convolution operation is equal to the time-consuming of one convolution operation, this convolution operation can be regarded as one convolution operation. When the electronic device performs the convolution operation, it only needs to start the hardware resources once and load the weights (target parameter information) once, which has a relatively small load on the hardware and a relatively short time-consuming.
[0018] In a possible implementation of the first aspect, performing a memory rearrangement operation on the first output data to obtain the target output data includes: performing a memory rearrangement operation on the first output data to obtain the second output data; performing an interception operation on the second output data to obtain the target output data.
[0019] It can be understood that the parameter information of the deconvolution operation further includes the number of padding (pad) layers. Among them, the number of pad layers represents the number of zero-padding layers added to the edge of the data to be processed, which can include the number of left pad layers, the number of right pad layers, the number of upper pad layers, and the number of lower pad layers. If at least one of the numbers of pad layers is not 0, then during the deconvolution operation, after obtaining the deconvolution result, a cropping operation is performed again. Therefore, after converting the deconvolution operation into a convolution operation, in order to obtain the target output data that is the same as the output data obtained by the deconvolution operation, it is necessary to perform a cropping operation after completing the memory rearrangement operation.
[0020] It should be noted that if the numbers of pad layers are all 0, then the target output data obtained after performing the cropping operation is the same as the second output data.
[0021] In some embodiments, the above-mentioned memory rearrangement operation and cropping operation can be combined into one memory rearrangement operation, that is, after obtaining the first output data, the memory rearrangement operation is performed to obtain the target output data.
[0022] In a second aspect, an embodiment of the present application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the data processing method mentioned in the first aspect and any possible implementation of the first aspect.
[0023] In a third aspect, an embodiment of the present application provides a readable storage medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device executes the data processing method mentioned in the first aspect and any possible implementation of the first aspect.
[0024] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by an electronic device, the electronic device executes the computer program code of the data processing method mentioned in the first aspect and any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 According to some embodiments of the present application, a schematic diagram of an application scenario is shown;
[0026] Figure 2A According to some embodiments of the present application, an example diagram of a convolution operation process is shown;
[0027] Figure 2B According to some embodiments of the present application, a schematic diagram of a deconvolution operation process is shown;
[0028] Figure 3According to some embodiments of the present application, a schematic diagram of D2S memory rearrangement operation is shown;
[0029] Figure 4A According to some embodiments of the present application, a schematic diagram of another deconvolution operation process is shown;
[0030] Figure 4B According to some embodiments of the present application, a Figure 4A schematic diagram of converting the shown convolution process into a deconvolution operation process is shown;
[0031] Figure 5 According to some embodiments of the present application, a flowchart of a data processing method is shown;
[0032] Figure 6 According to some embodiments of the present application, a schematic diagram of a deconvolution operation process is shown when both the width dimension stride and the height dimension stride are 1 and both the input channels and the output channels of the first convolution kernel are greater than 1;
[0033] Figure 7 According to some embodiments of the present application, a flowchart of a method for determining a target convolution kernel is shown;
[0034] Figure 8A According to some embodiments of the present application, a schematic diagram of data elements corresponding in position in a second convolution kernel is shown;
[0035] Figure 8B According to some embodiments of the present application, a schematic diagram of data padding in the width dimension and the height dimension of a second convolution kernel is shown;
[0036] Figure 9A According to some embodiments of the present application, a schematic diagram of M first sub - blocks is shown;
[0037] Figure 9B According to some embodiments of the present application, a schematic diagram of the process of determining N second sub - blocks is shown;
[0038] Figure 10 According to some embodiments of the present application, a schematic diagram of the process of determining N sub - convolution kernels is shown;
[0039] Figure 11 According to some embodiments of the present application, taking the input channels number of the first convolution kernel equal to 1, the output channels number equal to 1, the width dimension stride equal to 3, and the height dimension stride equal to 3 as an example, a schematic diagram of the process of the data processing method is shown;
[0040] Figure 12 According to some embodiments of the present application, a schematic diagram of the principle of dividing a convolution kernel is shown;
[0041] Figure 13 According to some embodiments of the present application, a schematic diagram of D2S memory rearrangement operation is shown when the number of output channels of a first convolution kernel is not equal to 1;
[0042] Figure 14 According to some embodiments of the present application, a schematic structural diagram of an electronic device is shown. Detailed implementation manners
[0043] Illustrative embodiments of the present application include, but are not limited to, a data processing method, an electronic device, a storage medium, and a program product. The data processing method of the present application will be introduced below in conjunction with specific embodiments.
[0044] The application scenarios of the data processing method provided in the embodiments of the present application can refer to Figure 1 .
[0045] In some embodiments, as shown in (1) of Figure 1 , in response to a user's click operation on the high-definition restoration control 11, an application program (such as an image processing application) in the electronic device processes the low-resolution image P1 based on the method provided in the present application to obtain a high-resolution image P2 as shown in (2) of Figure 1 .
[0046] In some embodiments, the application scenarios further include scenarios such as image segmentation and image generation. For example, in the image segmentation scenario, the electronic device processes an image through the data processing method provided in the present application to obtain a segmented image. The present application does not limit the application scenarios of the data processing method.
[0047] To more clearly illustrate the solutions of the embodiments of the present application, the relevant domain terms involved in the present application will be explained below. For the convenience of description, the position of the i-th row and j-th column in the data block mentioned in the present application is denoted as position (i - 1, j - 1).
[0048] Convolution: In a deep neural network, convolution is a common neural network structure for extracting spatial information. The convolution operation is a downsampling operation mainly used to extract local spatial features from the input data, reduce the spatial size, and can be applied to tasks such as image classification, object detection, and semantic segmentation. During the convolution operation, the convolution kernel slides over the input data, and the convolution kernel is multiplied by the part of the input data covered by the convolution kernel, and then the products corresponding to the data elements at each position of the convolution kernel are accumulated to obtain the data elements at multiple positions in the output data. The input data of the convolution operation is usually a tensor, which can include three dimensions, namely the height dimension, the width dimension, and the channel dimension, and the convolution kernel includes four dimensions, namely the height dimension, the width dimension, the input channel dimension, and the output channel dimension. Among them, the number of input channels of the convolution kernel is the same as the number of channels of the input data, and the number of output channels is the number of feature maps generated after performing the convolution operation.
[0049] Exemplarily, Figure 2A shows an example diagram of a convolution operation process. As Figure 2A shown, the number of channels of the input data of the convolution operation is 1, the width (i.e., the size of the input data in the width dimension) is 4, the height (i.e., the size of the input data in the height dimension) is 4, and the input data can be represented as a 4×4 tensor Conv_in([[0, 1, 4, 5], [2, 3, 6, 7], [8, 9, 12, 13], [10, 11, 14, 15]]). And, the number of output channels and the number of input channels of the convolution kernel of the convolution operation are both 1, the size of the convolution kernel is 3×3, and this convolution kernel can be represented as a tensor K_c([[0, 1, 4], [2, 3, 6], [8, 9, 12]]).
[0050] During the convolution operation, when the convolution kernel K_c is at the position Conv_in1 on the input data, the convolution kernel K_c can be multiplied by the part of the input data covered by the convolution kernel K_c ([[0, 1, 4], [2, 3, 6], [8, 9, 12]]), and the products are accumulated to obtain the data element (i.e., 355) at the position (0, 0) in the output data Conv_out. Similarly, when the convolution kernel K_c is at the position Conv_in2 on the input data, the data element (i.e., 426) at the position (0, 1) in the output data Conv_out can be obtained; when the convolution kernel K_c is at the position Conv_in3 on the input data, the data element (i.e., 489) at the position (1, 0) in the output data Conv_out can be obtained; when the convolution kernel K_c is at the position Conv_in4 on the input data, the data element (i.e., 560) at the position (1, 1) in the output data Conv_out can be obtained.
[0051] Deconvolution: The difference between deconvolution and convolution is that deconvolution operation is an upsampling operation used to restore spatial information and increase the spatial dimension. Deconvolution can be considered as the inverse operation of convolution, but through the deconvolution operation, the data can only be restored to the size before the convolution operation is performed, and the values of the data cannot be fully restored. During the deconvolution operation, by multiplying the convolutional kernel with each data element in the input data, multiple data blocks equal in number to the data elements in the input data are obtained. Then, these multiple data blocks are overlapped according to the positions of the corresponding data elements in the input data, and the elements at the overlapping positions are accumulated to obtain the output data. The input data of the deconvolution operation is usually a tensor, which can include three dimensions, namely the height dimension, the width dimension, and the channel dimension, and the convolutional kernel includes four dimensions, namely the height dimension, the width dimension, the input channel dimension, and the output channel dimension. Among them, the number of input channels of the convolutional kernel is the same as the number of channels of the input data, and the number of output channels is the number of feature maps generated after performing the deconvolution operation. The parameters used in the deconvolution operation also include the stride, which includes the stride in the width dimension and the stride in the height dimension. Among them, the stride in the width dimension is the degree of expansion of the input data in the width dimension, and the stride in the height dimension represents the degree of expansion of the input data in the height dimension. When the stride in the width dimension and the stride in the height dimension of the deconvolution operation are both 2, it means that one 0 element is added between adjacent elements in the input data before performing the deconvolution operation. It should be noted that during the deconvolution operation, the stride of the convolutional kernel movement must be 1.
[0052] Exemplarily, Figure 2B Fig. shows an example diagram of a deconvolution operation process. As Figure 2B shown, the number of channels of the input data of the deconvolution operation is 1, the width is 2, and the height is 2. The stride of the deconvolution operation is 1, and this input data can be represented as a 2×2 tensor deconv_in([[0, 1], [2, 3]]). And, Figure 2B the number of output channels and the number of input channels of the convolutional kernel shown in are both 1, and the size of the convolutional kernel is 2×2, and this convolutional kernel can be represented as the tensor k_d([[0, 1], [2, 3]]).
[0053] During the deconvolution process, the convolution kernel k_d is multiplied by the data element at position (0, 0) in the input data to obtain the data block deconv_in1([[0, 0], [0, 0]]); the convolution kernel k_d is multiplied by the data element at position (0, 1) in the input data to obtain the data block deconv_in2([[0, 1], [2, 3]]); the convolution kernel k_d is multiplied by the data element at position (1, 0) in the input data to obtain the data block deconv_in3([[0, 2], [4, 6]]); the convolution kernel k_d is multiplied by the data element at position (1, 1) in the input data to obtain the data block deconv_in4([[0, 3], [6, 9]]). Then, since the data element corresponding to the data block deconv_in1 is at position (0, 0) in the input data, the data element at position (0, 0) in the data block deconv_in1 is aligned with the position (0, 0) of the output data, that is, the data block deconv_in1 covers the positions (0, 0), (0, 1), (1, 0), and (1, 1) of the output data. Similarly, the data block deconv_in2 covers the positions (0, 1), (0, 2), (1, 1), and (1, 2) of the output data; deconv_in3 covers the positions (1, 0), (1, 1), (2, 0), and (2, 1) of the output data; deconv_in4 covers the positions (1, 1), (1, 2), (2, 1), and (2, 2) of the output data. After that, the elements at the overlapping positions are accumulated to obtain the output data, and the output data can be represented as the tensor deconv_out([[0, 0, 1], [1, 4, 6], [4, 12, 9]]).
[0054] Depth to Space (D2S) memory rearrangement operation: An operation that rearranges the input data from the depth dimension to the space dimension, where the input data is usually a tensor and can include three dimensions, namely the height dimension, the width dimension, and the channel dimension.
[0055] Exemplarily, such as Figure 3As shown, the input data has a height of 2, a width of 2, and 4 channels. The data in each channel can be a data block, and each data block corresponding to a channel includes four positions. Taking channel C1 as an example, the data block corresponding to channel C1 can include positions (0, 0), (0, 1), (1, 0), and (1, 1). The data element at position (0, 0) is 0, the data element at position (0, 1) is 4, the data element at position (1, 0) is 8, and the data element at position (1, 1) is 12. When performing the D2S memory rearrangement operation on this input data, the data elements at the corresponding positions of each channel can be placed into an output data block.
[0056] Specifically, the data element at position (0, 0) in the data block corresponding to channel C1 (i.e., 0), the data element at position (0, 0) in the data block corresponding to channel C2 (i.e., 1), the data element at position (0, 0) in the data block corresponding to channel C3 (i.e., 2), and the data element at position (0, 0) in the data block corresponding to channel C4 (i.e., 3) can be placed in an output data block in sequence. And the width of this output data block (i.e., block_size_x) is equal to the size of the input data in the W dimension, and the height of this output data block (i.e., block_size_y) is equal to the size of the input data in the H dimension.
[0057] It can be understood that since the width and height of the output data block are determined, when placing the data elements in the output data block, they can be placed along the width dimension first and then along the height dimension. For example, the data elements corresponding to channel C1 can be placed at position (0, 0) of the output data block, the data elements corresponding to channel C2 can be placed at position (0, 1) of the output data block, the data elements corresponding to channel C3 can be placed at position (1, 0) of the output data block, and the data elements corresponding to channel C4 can be placed at position (1, 1) of the output data block.
[0058] Then, this output data block can be placed at position (0, 0) of the output data. Similarly, the output data block obtained based on the data elements at position (0, 1) in the data blocks corresponding to each channel is placed at position (0, 1) of the output data, the output data block obtained based on the data elements at position (1, 0) in the data blocks corresponding to each channel is placed at position (1, 0) of the output data, and the output data block obtained based on the data elements at position (1, 1) in the data blocks corresponding to each channel is placed at position (1, 1) of the output data. In this way, the output data as shown in Figure 3 can be obtained, that is, the data after the memory rearrangement operation. The height of this output data is 4, the width is 4, and the number of channels is 1.
[0059] As described above, since the structures and parameters of different deconvolution neural networks may vary, it may be difficult to design a separate deconvolution calculation unit to meet the requirements of different tasks, which will increase the implementation complexity and development cost.
[0060] Therefore, the embodiments of the present application provide a data processing method. In this method, an electronic device can obtain data to be processed, where the data to be processed is data to be deconvolved, and then perform a convolution operation on the data to be processed based on a target convolution kernel to obtain target output data. Among them, the target convolution kernel is determined based on the first convolution kernel and the parameter information of the deconvolution operation, and the target output data is the same as the output data obtained by performing a deconvolution operation on the data to be processed based on the first convolution kernel. Moreover, the first convolution kernel is the convolution kernel used in the deconvolution operation.
[0061] In this way, the data processing method provided by the present application can convert the deconvolution operation into a convolution operation, providing feasibility for performing the deconvolution operation under the condition of limited memory and computing resources of the electronic device. The electronic device can use the existing convolution calculation unit to implement the deconvolution operation without designing a separate deconvolution calculation unit, reducing the implementation complexity and development cost of the deconvolution operation.
[0062] In some embodiments, in the data processing method mentioned in the embodiments of the present application, the target convolution kernel can be determined by: performing one horizontal flip and one vertical flip on the first convolution kernel along the first plane formed by the width dimension and the height dimension (that is, performing one horizontal flip based on the height dimension axis and then one vertical flip based on the width dimension axis) to obtain the target convolution kernel. Among them, the target convolution kernel can be obtained before data processing or when receiving the data to be processed.
[0063] In some embodiments, during the convolution process of the data to be processed and the target convolution kernel, zero padding can be added to the edge of the data to be processed first, and then the convolution is performed based on the padded data to be processed and the target convolution kernel. Among them, the number of zero-padding layers in the width dimension is 1 less than the width of the target convolution kernel, and the number of zero-padding layers in the height dimension is 1 less than the height of the target convolution kernel. For example, if the width of the target convolution kernel is 2, then the number of zero-padding layers in the width dimension is 1 layer. Another example, if the height of the target convolution kernel is 2, then the number of zero-padding layers in the height dimension is 1 layer.
[0064] For ease of understanding, first, the following will be combined with Figure 4A and Figure 4B to describe the data processing method mentioned in the embodiments of the present application.
[0065] Figure 4ATaking the case where the stride of the width dimension of the deconvolution operation is 1 and the stride of the height dimension is 1 as an example, a schematic diagram of a deconvolution operation process is shown, where the positions of the respective data elements represent the data elements.
[0066] As Figure 4A shown, the data to be processed can be represented as the tensor deconv_in([[x(0, 0), x(0, 1)], [x(1, 0), x(1, 1)]]), and the first convolution kernel can be represented as the tensor k_d([[k(0, 0), k(0, 1)], [k(1, 0), k(1, 1)]]).
[0067] In this deconvolution operation process, x(0, 0), x(0, 1), x(1, 0), and x(1, 1) in the data to be processed deconv_in are respectively multiplied by the convolution kernel k_d to obtain the data blocks x(0, 0)*k_d, x(0, 1)*k_d, x(1, 0)*k_d, and x(1, 1)*k_d. It can be understood that the sizes of the data blocks x(0, 0)*k_d, x(0, 1)*k_d, x(1, 0)*k_d, and x(1, 1)*k_d are all 2×2.
[0068] Then, the data element at position (0, 0) in the data block x(0, 0)*k_d is aligned with the position (0, 0) of the output data, the data element at position (0, 0) in the data block x(0, 1)*k_d is aligned with the position (0, 1) of the output data, the data element at position (0, 0) in the data block x(1, 0)*k_d is aligned with the position (1, 0) of the output data, and the data element at position (0, 0) in the data block x(1, 1)*k_d is aligned with the position (1, 1) of the output data. Then, the elements at the overlapping positions are added.
[0069] In this way, the data element at position (0, 0) in the output data can be represented as: x(0, 0)*k(0, 0), the data element at position (0, 1) in the output data can be represented as: x(0, 1)*k(0, 0) + x(0, 0)*k(0, 1), the data element at position (0, 2) in the output data can be represented as: x(0, 1)*k(0, 1), and so on.
[0070] According to the data processing method provided in the embodiments of the present application, the process of Figure 4A converting the deconvolution operation shown into a convolution operation is as shown in 4B.
[0071] Referring to Figure 4B, the above convolution kernel \(k_d\) is horizontally flipped once along the height dimension axis to obtain the convolution kernel \(k_d'\), and then vertically flipped once along the width dimension axis (i.e., horizontally flipped once and vertically flipped once based on the first plane composed of the height dimension and the width dimension), to obtain the convolution kernel \(k_d''([[k(1, 1), k(1, 0)], [k(0, 1), k(0, 0)]])\). Then, add 1 layer of zero padding to the edge of the above data to be processed \(deconv\_in\) to obtain the padded data to be processed, which can be represented as the tensor \(deconv\_in'([[0, 0, 0, 0], [0, x(0, 0), x(0, 1), 0], [0, x(1, 0), x(1, 1), 0], [0, 0, 0, 0]])\). Based on the convolution kernel \(k_d''\) and the padded data to be processed \(deconv\_in'\) for convolution, in the obtained convolution result (i.e., the output data), the data element at the position \((0, 0)\) can be represented as: \(x(0, 0) * k(0, 0)\), the data element at the position \((0, 1)\) can be represented as: \(x(0, 1) * k(0, 0)+x(0, 0) * k(0, 1)\), the data element at the position \((0, 2)\) can be represented as: \(x(0, 1) * k(0, 1)\), and so on.
[0072] It can be seen from the represented output data that the output data obtained based on the data processing method provided in the embodiments of the present application is the same as the output data obtained based on the deconvolution operation. The output data can be represented as the tensor \(deconv\_out([y(0, 0), y(0, 1), y(0, 2)], [y(1, 0), y(1, 1), y(1, 2)], [y(2, 0), y(2, 1), y(2, 2)])\). Thus, through the method provided in the embodiments of the present application, the Figure 4A deconvolution operation shown in can be converted into a convolution operation.
[0073] Next, in combination with Figure 5 the flowchart corresponding to the data processing method shown, the data processing method provided in the embodiments of the present application will be described.
[0074] It should be noted that the data to be processed and the target output data both include a width dimension, a height dimension, and a channel dimension. The data of the first convolution kernel includes a width dimension, a height dimension, an input channel dimension, and an output channel dimension. The parameter information of the deconvolution operation includes the width of the data to be processed, the height of the data to be processed, the number of channels of the data to be processed, the width of the first convolution kernel, the height of the first convolution kernel, the step size of the width dimension, the step size of the height dimension, and the number of layers of padding (pad), etc. Among them, the number of layers of pad represents the number of layers of zero padding added to the edge of the data to be processed, and can include the number of left pad layers, the number of right pad layers, the number of upper pad layers, and the number of lower pad layers.
[0075] For ease of description, the size of the data to be processed is represented as (h in , w in , c in ), where h in represents the height of the data to be processed, w in represents the width of the data to be processed, and c in represents the number of channels of the data to be processed and the number of input channels of the first convolution kernel. The size of the target output data is represented as (h out , w out , c out ), where h out represents the height of the target output data, w out represents the width of the target output data, and c out represents the number of channels of the target output data and the number of output channels of the first convolution kernel. The size of the first convolution kernel is represented as (c out , k h , k w , c in ), where k h represents the height of the first convolution kernel, and k w represents the width of the first convolution kernel. The size of the stride is represented as (s h , s w ), where s h represents the stride in the height dimension, and s w represents the stride in the width dimension. The number of padding layers can be represented in the order of left, right, up, and down as (pad left , pad right , pad top , pad bottom ), where pad left represents the number of padding layers on the left side, pad right represents the number of padding layers on the right side, pad top represents the number of padding layers on the upper side, and pad bottom represents the number of padding layers on the lower side.
[0076] As Figure 5 shown, the data processing method includes:
[0077] S501, obtain the data to be processed, which is the data to be deconvolved.
[0078] In some embodiments, when an electronic device executes the data processing method, it first obtains the data to be processed, which is the data to be deconvolved.
[0079] It can be understood that in some embodiments, the data to be processed may be a low-resolution image (or the feature map corresponding to the image). This data processing method can process the low-resolution image to obtain a high-resolution image. The data corresponding to each channel of the data to be processed usually represents the abstract representation of the feature map in different feature dimensions. For example, the data corresponding to different channels of the data to be processed can represent different features of the image (such as edges, textures, colors, etc.); in the semantic segmentation task, the data corresponding to different channels of the data to be processed can represent image features of different semantic granularities, etc. In the embodiments of the present application, the data to be processed may be image data, audio data, etc., and the present application does not make any limitations.
[0080] S502, perform a convolution operation on the data to be processed based on the target convolution kernel to obtain the first output data.
[0081] In some embodiments, after the electronic device obtains the data to be processed, it performs a convolution operation on the data to be processed based on the target convolution kernel to obtain the first output data, where the target convolution kernel is determined based on the parameter information of the first convolution kernel and the transposed convolution operation. The method for determining the target convolution kernel based on the parameter information of the first convolution kernel and the transposed convolution operation will be described later. Among them, the width of the target convolution kernel is d w , and the height is d h , the target convolution kernel may include multiple sub-convolution kernels and the widths of the multiple sub-convolution kernels are equal to the width of the target convolution kernel, the heights of the multiple sub-convolution kernels are equal to the height of the target convolution kernel, and any one channel of the data to be processed corresponds to at least one of the multiple sub-convolution kernels.
[0082] In some embodiments, performing a convolution operation on the data to be processed based on the target convolution kernel to obtain the first output data may include: performing a convolution operation on the data to be processed in parallel based on the target parameter information and the multiple sub-convolution kernels in the target convolution kernel to obtain the first output data. Among them, the target parameter information is determined based on the parameter information of the transposed convolution operation.
[0083] Specifically, the target parameter information includes the number of layers of the convolution pad. The number of layers of the convolution pad represents the number of layers of zero padding added to the edge of the data to be processed during the convolution operation, and may include the number of left pad layers, the number of right pad layers, the number of upper pad layers, and the number of lower pad layers. During the convolution operation, the number of layers of the convolution pad is determined based on the width dimension stride and the height dimension stride of the transposed convolution operation, and the size of the number of layers of the convolution pad is (d w - 1, d w - 1, d h - 1, d h - 1).
[0084] During the convolution operation, it is first necessary to pad the data to be processed based on the number of layers of the convolution pad to obtain the padded data to be processed.
[0085] The above-mentioned parallel convolution operation on the data to be processed based on the target parameter information and multiple sub-convolution kernels in the target convolution kernel to obtain the first output data may include padding the data to be processed based on the number of layers of the convolution pad to obtain the padded data to be processed, and parallelly performing a convolution operation on the padded data to be processed based on multiple sub-convolution kernels in the target convolution kernel to obtain the first output data.
[0086] Denote the first output data as ACT out ’, the padded data to be processed as ACT in , and the target convolution kernel as K new . Then the convolution operation can be expressed as the following formula (1):
[0087]
[0088] Where, represents the convolution operation. The stride in the width dimension of the convolution operation is 1, and the stride in the height dimension is 1. The convolution operation of the data in any one channel of the padded data to be processed with the corresponding at least one sub-convolution kernel can be performed in parallel, and moreover, the convolution operations based on the data in each channel of the padded data to be processed can also be performed in parallel.
[0089] It can be understood that the size of the first output data ACT out ’ is ((pad top + h out + pad bottom ) / s h , (pad left + w out + pad right ) / s w , s h * s w * c out ).
[0090] S503. Perform a memory rearrangement operation based on the first output data to obtain the target output data.
[0091] In some embodiments, after the electronic device performs a convolution operation to obtain the first output data, it may perform a memory rearrangement operation based on the first output data to obtain the target output data, where the target output data is the same as the output data obtained by performing a deconvolution operation on the data to be processed based on the first convolution kernel.
[0092] It can be understood that when the data to be processed is image data, the target output data mentioned in this application can be the feature map corresponding to a high-resolution picture. When the data to be processed is audio data, the target output data mentioned in this application can be the processed audio data.
[0093] In some embodiments, the first output data can be converted into the target output data through D2S memory rearrangement operation. Based on Figure 3 the example shown, each output channel in the first output data can correspond to Figure 3 each output channel in the input data in []. Through the D2S memory rearrangement operation, the data elements in the position (0, 0) of each output channel in the first output data can be placed in the output data block at the position (0, 0) of the target output data, so as to obtain the target output data.
[0094] Specifically, denoting the output result of the D2S memory rearrangement operation as ACT out ”, then this D2S memory rearrangement operation can be expressed as the following formula (2):
[0095] ACT out ” = D2S(ACT out ’) (2)
[0096] where block_size_x = s w and block_size_y = s h .
[0097] In other embodiments, the first output data can also be converted into the target output data through other memory rearrangement operation methods, and this application does not make any limitations in this regard.
[0098] In some embodiments, after obtaining the first output data based on the convolution operation, a memory rearrangement operation is performed on the first output data to obtain the second output data, and then a crop operation is performed based on the second output data to obtain the target output data.
[0099] Specifically, denoting the target output data as ACT out , this crop operation can be expressed as the following formula (3):
[0100] ACT out = crop(ACT out ”)
[0101] = ACT out ”[pad top :(pad top + h out ), pad left :(pad left + wout ), 0: c out (3)
[0102] It can be understood that, as can be seen from the above formula (3), performing the crop operation means intercepting the padded layers corresponding in the second output data. During the deconvolution operation, if any of the padded layers (including the left pad layer, right pad layer, upper pad layer, and lower pad layer) is greater than 0, a crop operation needs to be performed after obtaining the deconvolution result. Therefore, in order to obtain the target output data identical to the output data obtained by the deconvolution operation, a crop operation needs to be performed again after completing the D2S memory rearrangement operation.
[0103] It should be noted that if all the padded layers are 0, the target output data obtained after performing the crop is the same as the second output data.
[0104] In some embodiments, the above memory rearrangement operation and crop operation can be combined into one memory rearrangement operation, that is, perform the memory rearrangement operation after obtaining the first output data to obtain the target output data.
[0105] In this way, the data processing method provided by the embodiments of the present application can convert the deconvolution operation into a convolution operation, without the need to separately design a deconvolution calculation unit. During the data processing of the electronic device, the hardware resources only need to be started once, the load on the hardware is relatively small, and the time consumption is relatively short, providing feasibility for performing the deconvolution operation under the condition of limited memory and computing resources of the electronic device, and reducing the complexity and development cost of implementing the deconvolution operation.
[0106] It can be understood that the number of output channels of the first convolution kernel is greater than or equal to 1. The data processing method described in the embodiments of the present application in combination with Figure 4A and Figure 4B is described by taking the number of channels of the data to be processed in the deconvolution operation as 1, the number of input channels of the first convolution kernel as 1, and the number of output channels of the first convolution kernel as 1 as an example. For the case where the number of channels of the data to be processed is greater than 1, the number of input channels of the first convolution kernel is greater than 1, and the number of output channels of the first convolution kernel is greater than 1, it is similar to the case where the number of channels of the data to be processed is 1, the number of input channels of the first convolution kernel is 1, and the number of output channels of the first convolution kernel is 1.
[0107] Next, in combination with Figure 6 the data processing method described in the embodiments of the present application will be introduced for the case where the number of channels of the data to be processed is greater than 1, the number of input channels of the first convolution kernel is greater than 1, and the number of output channels of the first convolution kernel is greater than 1.
[0108] For the sake of convenient description, it is described by taking the stride in the width dimension and the stride in the height dimension as 1 as an example.
[0109] Specifically, as Figure 6 shown, each output channel of the first convolution kernel corresponds to a group of second convolution kernels, and the number of second convolution kernels corresponding to each output channel is equal to the number of input channels (Cin) of the first convolution kernel. For example, if the number of input channels of the first convolution kernel is 4, then each output channel of the first convolution kernel corresponds to 4 second convolution kernels.
[0110] During the transposed convolution operation, the data in each channel of the data to be processed is subjected to a transposed convolution operation with the corresponding second convolution kernels in each output channel of the first convolution kernel, and moreover, the output data corresponding to each output channel is obtained by performing weighted summation on the results obtained by performing transposed convolution operations on the data in the corresponding channels of the data to be processed based on the corresponding second convolution kernels in that output channel.
[0111] For example, referring to Figure 6 , assume that each output channel of the first convolution kernel corresponds to Cin second convolution kernels, the number of channels of the data to be processed is Cin, and the Cin second convolution kernels correspond one-to-one to the Cin channels of the data to be processed. Then, in the transposed convolution operation, the Cin second convolution kernels in output channel C out _1 perform transposed convolution operations with the data in the corresponding channels of the data to be processed respectively, obtaining Cin transposed convolution results. The output data deconv_out1 is obtained by performing weighted summation on the Cin transposed convolution results corresponding to output channel C out _1.
[0112] It can be understood that, in the dimension of each output channel of the first convolution kernel, the transposed convolution operations performed based on each output channel of the first convolution kernel are independent, and for an output channel of the first convolution kernel, the transposed convolution operations performed based on the second convolution kernels in that output channel are also independent.
[0113] According to the data processing method provided by the embodiments of the present application, when converting the transposed convolution operation Figure 6 shown into a convolution operation, the process shown in Figure 4B can be referred to. And, Figure 4B the first convolution kernel k_d in Figure 6 can be any one of the second convolution kernels in any one of the output channels of the first convolution kernel in Figure 6 . That is, for any second convolution kernel in the first convolution kernel in Figure 6 , a horizontal flip and a vertical flip are performed once based on the first plane to obtain a target convolution kernel. Then, after adding zero padding to the edge of the data to be processed, convolution is performed based on the target convolution kernel and the padded data to be processed. The target convolution kernel includes sub-convolution kernels with the same number as the second convolution kernels, and the process of each sub-convolution kernel performing convolution with the data to be processed is the same as that in Figure 4BThe process shown is the same. In this way, the data processing method provided by the embodiments of this application can Figure 6 convert the deconvolution operation shown into a convolution operation.
[0114] Next, the method for determining the target convolution kernel based on the parameter information of the first convolution kernel and the deconvolution operation mentioned in S502 above will be described.
[0115] In some embodiments, the first convolution kernel includes at least one output channel, and each output channel of the first convolution kernel includes at least one second convolution kernel, and the number of the at least one second convolution kernel is equal to the number of input channels of the first convolution kernel. That is, c out is an integer greater than or equal to 1, c in is an integer greater than or equal to 1, and each output channel of the first convolution kernel includes c in second convolution kernels.
[0116] It can be understood that in the case where the first convolution kernel includes at least one output channel, it can be considered that the part corresponding to each output channel of the first convolution kernel is at least one second convolution kernel.
[0117] In the case where the stride in the width dimension is greater than 1 or the stride in the height dimension is greater than 1, the method for determining the target convolution kernel can refer to Figure 7 the flowchart shown. As Figure 7 shown, this method includes:
[0118] S701, for each output channel of the first convolution kernel, based on the width of the first convolution kernel, the height of the first convolution kernel, the stride in the width dimension, and the stride in the height dimension, divide each second convolution kernel into M first sub-blocks.
[0119] Among them, the width of the first sub-block is equal to s w , the height of the first sub-block is equal to s h , the positions of the data elements of each first sub-block in the M first sub-blocks correspond one by one, and M is an integer greater than or equal to 1, M = (k h / s h ) * (k w / s w ).
[0120] Exemplarily, if k w / s w = 3, k h / s h = 3, as Figure 8A shown, for example, k w = k h = 9, s w = s h = 3, then for each output channel of the first convolution kernel, the corresponding c inAny one of the second convolutional kernels in the second convolutional kernels is kernel, and the size of kernel is (9, 9). For each output channel of the first convolutional kernel, kernel can be divided into 9 first sub-blocks. Moreover, each of the first sub-blocks includes 9 data elements, and the positions of the data elements in each first sub-block correspond to each other one by one. Figure 8A The data elements at the same-color positions in each of the first sub-blocks are the data elements with corresponding positions. For example, the data elements at the upper-left corner positions in each of the first sub-blocks are the data elements corresponding to the upper-left corner positions, etc.
[0121] In some embodiments, in the case where the width dimension stride cannot be divided evenly by the width of the first convolutional kernel, data padding (i.e., filling) is performed in the width dimension of the second convolutional kernel so that the width dimension stride can be divided evenly by the width of the first convolutional kernel, and then the second convolutional kernel is divided into M first sub-blocks; in the case where the height dimension stride cannot be divided evenly by the height of the first convolutional kernel, filling is performed in the height dimension of the second convolutional kernel so that the height dimension stride can be divided evenly by the height of the first convolutional kernel, and then the second convolutional kernel is divided into M first sub-blocks. Among them, the data for data padding is 0.
[0122] As Figure 8B shown, the padded second convolutional kernel is denoted as kernel_pad, and the size of kernel_pad is (d h *s h , d w *s w ), where d h = ceil(k h / s h ), d w = ceil(k w / s w ), and ceil is the ceiling function. Assume that k w = k h = 8, s w = s h = 3, then d h = ceil(k h / s h ) = 3, d w = ceil(k w / s w ) = 3, and the size of kernel_pad is (9, 9).
[0123] It can be understood that since a group of second convolutional kernels corresponding to each input channel of the first convolutional kernel is independent, and the second convolutional kernels corresponding to each input channel are also independent, the size of the padded first convolutional kernel is (c out , d h * sh , d w *s w , c in )。
[0124] S702. Determine N second sub - blocks based on the data elements at corresponding positions in the M first sub - blocks.
[0125] Among them, the first sub - block includes N data elements, and the N data elements are respectively located at different positions in the first sub - block, and N is an integer greater than or equal to 1. Taking the Figure 8A shown example as an example, each first sub - block includes 9 data elements, that is, N is equal to 9.
[0126] It can be understood that N is equal to the product of the width - dimension step size and the height - dimension step size (s w *s h ).
[0127] In some embodiments, among the N data elements of each first sub - block in the M first sub - blocks, the data elements with corresponding positions are placed in the same second sub - block to determine N second sub - blocks, where the positions of the N data elements in the second sub - block correspond to the positions of the first sub - blocks where the N data elements are located in the second convolution kernel.
[0128] Based on Figure 8B the shown example, as Figure 9A shown, the M first sub - blocks can be expressed as k(i, j), and k(i, j)=kernel_pad[i*s h :(i + 1)*s h , j*s w :(j + 1)*s w , where i = 0, 1, …, d h - 1; j = 0, 1, …, d w - 1. Based on Figure 8A the shown example, after dividing the second convolution kernel kernel into 9 first sub - blocks, the 9 first sub - blocks can be expressed as k(i, j), and k(i, j)=kernel[i*3:(i + 1)*3, j*3:(j + 1)*3], where i = 0, 1, …, 2; j = 0, 1, …, 2. Taking k(0, 0) as an example, k(0, 0) represents the first sub - block located at position (0, 0) in kernel, and can be expressed as kernel[0:3, 0:3], including the data elements from the 1st row to the 3rd row and from the 1st column to the 3rd column in kernel.
[0129] It can be understood that after dividing each second convolution kernel, each second convolution kernel is divided into M first sub - blocks, and each output channel of the first convolution kernel includes c in*M first sub - blocks, for all output channels of the first convolutional kernel, including c out *c in *M first sub - blocks, which can be expressed as k(i, j) = kernel_pad[c out , i*s h :(i + 1)*s h , j*s w :(j + 1)*s w , ci n .
[0130] Continuing based on Figure 8B the example shown, referring to Figure 9B , for the M first sub - blocks k(i, j) corresponding to a second convolutional kernel, the positions of the data elements in each first sub - block can be expressed as (m, n), and m = 0, 1, …, s h - 1; n = 0, 1, …, s w - 1. The data element at the position (m, n) in each of the M first sub - blocks is the data element corresponding to the position. Placing the data elements corresponding to the positions in the same second sub - block, N second sub - blocks can be determined. These N second sub - blocks can be expressed as k_scatter [m,n] , and the size of k_scatter [m,n] is (d h , d w ), and the relationship between k_scatter [m,n] and k(i, j) can be expressed as: k_scatter [m,n] [i, j] = k(i, j)[m, n]. For example, when m = 0 and n = 0, k_scatter [0,0] includes the N data elements at the position (0, 0) in the M first sub - blocks, and the positions of the N data elements in k_scatter [0,0] correspond to the positions of the first sub - blocks where these N data elements are located in the second convolutional kernel. That is, the data element at the position (0, 0) in k_scatter [0,0] is at the position (0, 0) of the first sub - block k(0, 0), and the data element at the position (0, 1) in k_scatter [0,0] is at the position (0, 0) of the first sub - block k(0, 1), and so on. It can be understood that the number N of the second sub - blocks is equal to the number of data elements in the first sub - block, that is, equal to s h *s w .
[0131] Exemplarily, based on Figure 8AIn the example shown, the nine data elements at position (0, 0) in the nine first sub-blocks k(i, j) can be placed in a second sub-block. That is, the data element at position (0, 0) in k(0, 0) is placed at position (0, 0) of the second sub-block, the data element at position (0, 0) in k(0, 1) is placed at position (0, 1) of the second sub-block, the data element at position (0, 0) in k(0, 2) is placed at position (0, 2) of the second sub-block, the data element at position (0, 0) in k(1, 0) is placed at position (1, 0) of the second sub-block, the data element at position (0, 0) in k(1, 1) is placed at position (1, 1) of the second sub-block, the data element at position (0, 0) in k(1, 2) is placed at position (1, 2) of the second sub-block, the data element at position (0, 0) in k(2, 0) is placed at position (2, 0) of the second sub-block, the data element at position (0, 0) in k(2, 1) is placed at position (2, 1) of the second sub-block, and the data element at position (0, 0) in k(2, 2) is placed at position (2, 2) of the second sub-block.
[0132] It can be understood that for all output channels and input channels of the first convolutional kernel, N second sub-blocks k_scatter [m,n] have a size of (c out , d h , d w , c in ).
[0133] It should be noted that in the case where the width of the first convolutional kernel is equal to the width dimension stride and the height of the first convolutional kernel is equal to the height dimension stride, through the above S701, the second convolutional kernel can be divided into 1 first sub-block. Through the above S702, the data elements at each position in this first sub-block can be placed in a second sub-block, and N second sub-blocks are determined. Among them, each second sub-block only includes 1 data element. That is to say, in this case, the data elements at the corresponding positions in the above M first sub-blocks only include 1 data element.
[0134] S703, horizontally and vertically flip the above N second sub-blocks along the first plane respectively to determine N sub-convolutional kernels.
[0135] Specifically, the N sub-convolutional kernels can be represented as k_reverse [m,n] , k_reverse [m,n] have a size of (d h , d w ). Based on Figure 9B the example shown, as Figure 10 shown, for k_scatter[m,n] Performing a horizontal flip along the first plane, k_reverse1 can be obtained. [m,n] Then, [m,n] performing a vertical flip of k_reverse1 along the first plane, k_reverse can be obtained. [m,n] .
[0136] It can be understood that when m and n take different values, k_reverse [m,n] can represent different sub-convolution kernels. Moreover, k_reverse [m,n] represents N sub-convolution kernels corresponding to any one of the input channels in any one of the output channels of the first convolution kernel.
[0137] It should be noted that in the case where the width of the first convolution kernel is equal to the width dimension stride and the height of the first convolution kernel is equal to the height dimension stride, since each of the second sub-blocks only contains 1 data element, the N sub-convolution kernels obtained by horizontally and vertically flipping the N second sub-blocks along the first plane are the same as the second sub-blocks before flipping.
[0138] S704, concatenating the N sub-convolution kernels corresponding to each second convolution kernel in each output channel of the first convolution kernel along the output channel dimension to obtain the target convolution kernel.
[0139] Specifically, for one input channel in one output channel of the first convolution kernel, the N sub-convolution kernels k_reverse corresponding to the second convolution kernel corresponding to this input channel [m,n] are concatenated along the output channel dimension (i.e., the c out axis) in the order of first n increasing from zero and then m increasing from 0 to obtain a new vector k new . The size of k new is (s h * s w , d h , d w ). k new can be represented by the following formula (4), where || represents concatenating the sub-convolution kernels along the output channel dimension. For example, if s w = d h = d w = s h = 3, the concatenated sub-convolution kernel can be k Figure 11 as shown. new .
[0140]
[0141] Perform the above operations on the N sub-convolution kernels corresponding to each input channel in each output channel of the first convolution kernel to obtain the target convolution kernel, denoted as K new , the size of the target convolution kernel is (s h *s w *c out , d h , d w , c in ). It can be understood that the number of sub-convolution kernels included in the target convolution kernel is equal to the number of channels of the target convolution kernel, that is, equal to s h *s w *c out *c in , and any one channel in the data to be processed corresponds to s h *s w *c out *c in of the s h *s w *c out sub-convolutions.
[0142] It can be understood that during the convolution operation, the convolution operation between the convolution kernel corresponding to each output channel of the convolution kernel and the data to be processed can be run in parallel. After splicing the N sub-convolution kernels corresponding to each second convolution kernel in each output channel of the above first convolution kernel along the output channel dimension, during the subsequent convolution operation, the convolution operations of each sub-convolution kernel and the data to be processed can all be run in parallel. In this way, the time consumption of the convolution operation process can be reduced and the data processing efficiency can be improved.
[0143] Taking the number of input channels c in of the first convolution kernel equal to 1, the number of output channels c out equal to 1, the stride s w in the width dimension equal to 3, and the stride s h in the height dimension equal to 3 as an example, as Figure 11 shown, the size of the data to be processed act in is 3×3, the size of the target convolution kernel k new is s h ×s w (i.e., 3×3), the number of output channels of the target convolution kernel is s w *s h *c out = 9 (i.e., Figure 11 the channels out_c[0] to out_c[8] in), after performing the convolution operation between the data to be processed and the target convolution kernel, in the first output data act out ', the size of the output data in each channel is 5×5. When processing the first output data actout After performing the D2S memory rearrangement operation, the second output data act can be obtained. out ”, the size of the second output data act out ” is 5*s h ×5*s w , and the second output data act out ” is divided into 25 data blocks, and the size of each data block is 3×3 (equal to the size of the target convolution kernel). Then, each data block includes the data at the corresponding positions in the output data of each channel in the first output data.
[0144] For example, the data block at the position (0, 0) in the second output data act out ” is the data block act out ”(1). The data block act out ”(1) includes the data element at the position (0, 0) in the channel out_c[0] and this data element is at the position (0, 0) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[1] and this data element is at the position (0, 1) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[2] and this data element is at the position (0, 2) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[3] and this data element is at the position (1, 0) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[4] and this data element is at the position (1, 1) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[5] and this data element is at the position (1, 2) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[6] and this data element is at the position (2, 0) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[7] and this data element is at the position (2, 1) of the data block act out ”(1), the data element at the position (0, 0) in the channel out_c[8] and this data element is at the position (2, 2) of the data block act out ”(1).
[0145] Then, perform the crop operation on the second output data to obtain the target output data ACT out .
[0146] It can be understood that, since the deconvolution operations performed based on the output channels of the first convolution kernel are independent in terms of the dimensions of each output channel of the first convolution kernel, and for an output channel of the first convolution kernel, the deconvolution operations performed based on the second convolution kernels in that output channel are also independent. When both the number of input channels and the number of output channels are greater than 1, in Figure 11 the example shown, the data to be processed act in can be regarded as the data to be processed corresponding to any one channel of the data to be processed ACT in , and the target convolution kernel k new can be regarded as the target convolution kernel obtained based on all the second convolution kernels in any one output channel of the first convolution kernel. Then, when the data to be processed act in is convolved with the target convolution kernel k new , the first output data act out ' is the output data corresponding to any one output channel of the first convolution kernel in the first output data.
[0147] The following is combined with Figure 12 to explain the reason why when determining the target convolution kernel as shown above, each second convolution kernel in the first convolution kernel needs to be divided into M second sub - blocks. Figure 7 For the case where the stride in the width dimension is greater than 1 or the stride in the height dimension is greater than 1, during the deconvolution operation, at least one 0 element needs to be added between adjacent elements in the data to be processed.
[0148] Specifically, for the case where the stride in the width dimension is greater than 1, at least one 0 element needs to be added between adjacent elements in the width dimension of the data to be processed, and the number of these 0 elements is 1 less than the stride in the width dimension. For the case where the stride in the height dimension is greater than 1, at least one 0 element needs to be added between adjacent elements in the height dimension of the data to be processed, and the number of these 0 elements is 1 less than the stride in the height dimension.
[0149] For the sake of easy description, take the data to be processed with a width of 2, a height of 2, and a channel number of 1, a convolution kernel with a width of 4, a height of 4, an output channel number of 1, an input channel number of 1, a stride in the width dimension (denoted as s
[0150] ) of 2, and a stride in the height dimension (denoted as s w ) of 2 as an example for explanation. h ) as 2 for example to illustrate.
[0151] Based on the above analysis, since the stride in the width dimension and the stride in the height dimension in this example are both greater than 1, during the deconvolution operation, first, 1 zero element needs to be added between adjacent elements in the width dimension of the data to be processed, and 1 zero element needs to be added between adjacent elements in the height dimension, resulting in the data to be processed "deconv_in" as shown in Figure 12 shown.
[0152] In this way, the data element at position (0, 0) in the output data of the deconvolution operation can be expressed as: x(0, 0)*k(0, 0), the data element at position (0, 1) in the output data can be expressed as: x(0, 0)*k(0, 1), the data element at position (0, 2) in the output data can be expressed as: x(0, 1)*k(0, 0)+x(0, 0)*k(0, 2), the data element at position (0, 3) in the output data can be expressed as: x(0, 1)*k(0, 1)+x(0, 0)*k(0, 3), the data element at position (0, 4) in the output data can be expressed as: x(0, 1)*k(0, 2), the data element at position (0, 5) in the output data can be expressed as: x(0, 1)*k(0, 3), and so on.
[0153] It can be seen that the output data can be expressed as: output(0, 0)[m, n] = x(0, 0)*k(0, 0)[m, n], output(0, 1)[m, n] = x(0, 1)*k(0, 0)[m, n]+x(0, 0)*k(0, 1)[m, n], output(0, 2) = x(0, 1)*k(0, 1)[m, n], and so on. Among them, output(0, 0)[m, n], output(0, 1)[m, n], output(0, 2)[m, n], k(0, 0)[m, n], k(0, 1)[m, n] all represent vectors with width n and height m, and m = 0 or 1, n = 0 or 1.
[0154] Based on the above expression of the output data, it can be seen that this expression is the same as the expression of the result of the convolution operation. When m and n take different values, different convolution kernels can be represented. The above output data can be regarded as different convolution results obtained from the data to be processed based on these different convolution kernels.
[0155] In order to obtain the different convolution kernels represented when m and n take different values, as shown in Figure 12 shown, when performing the data processing method mentioned in the embodiments of the present application, the process of obtaining the target convolution kernel can be as follows: The convolution kernel k_d can be divided into four sub-blocks, each sub-block has a width of 2 (equal to s w ), and a height of 2 (equal to s h) When m = 0 and n = 0, it means to take the element at position (0, 0) in the sub-block corresponding to position (0, 0) in the convolution kernel k_d, that is, take k(0, 0), k(0, 2), k(2, 0), k(2, 2) to obtain the convolution kernel k_d1'; when m = 0 and n = 1, it means to take the element at position (0, 1) in the sub-block corresponding to position (0, 1) in the convolution kernel k_d, that is, take k(0, 1), k(0, 3), k(2, 1), k(2, 3) to obtain the convolution kernel k_d2'; when m = 1 and n = 0, it means to take the element at position (1, 0) in the sub-block corresponding to position (1, 0) in the convolution kernel k_d, that is, take k(1, 0), k(1, 2), k(3, 0), k(3, 2) to obtain the convolution kernel k_d3'; when m = 1 and n = 1, it means to take the element at position (1, 1) in the sub-block corresponding to position (1, 1) in the convolution kernel k_d, that is, take k(1, 1), k(1, 3), k(3, 1), k(3, 3) to obtain the convolution kernel k_d4'.
[0156] After that, perform horizontal flipping and vertical flipping (along the first plane composed of the width dimension and the height dimension) on the convolution kernels k_d1', k_d2', k_d3' and k_d4' respectively, and the convolution kernels k_d1, k_d2, k_d3 and k_d4 can be obtained. From the above expressions of the output data, it can be seen that after convolving the convolution kernel k_d1 with the data to be processed, partial output data can be obtained. That is to say, after convolving the convolution kernels k_d1, k_d2, k_d3 and k_d4 with the data to be processed, the complete output data can be obtained, and the output data obtained by convolving each convolution kernel with the data to be processed is discrete.
[0157] The following combines Figure 13 to explain the reason for performing the memory rearrangement operation mentioned in S503 above.
[0158] Since the first output data obtained by directly performing convolution operation on the data to be processed based on the target convolution kernel is discrete compared with the output data obtained by deconvolution operation, therefore, the memory rearrangement operation can be performed based on the first output data to convert the first output data into the target output data.
[0159] When the number of output channels of the first convolution kernel is equal to 1, the process of performing the D2S memory rearrangement operation on the first output data can refer to the relevant description of Figure 3 above. When the number of output channels of the first convolution kernel is not equal to 1, in the output data, the convolution results of the sub-convolution kernels corresponding to each output channel of the first convolution kernel are arranged together. As Figure 13 shown, the first output data ACT out’ Among them, the parts with the same color are the convolution results of the sub-convolution kernels corresponding to the same output channel based on the first convolution kernel. During the execution of the D2S memory rearrangement operation, first, the channels of the first output data are evenly divided into N (equal to s w *s h ) parts, and the length of each part of the data is equal to the number of output channels (c out ) of the first convolution kernel. Then, in each part of the data, the data at the corresponding positions are placed in the same data block, where the size of each data block is s w ×s h , to obtain the second output data. For example, as Figure 13 shown, the stride s w in the width dimension is 2, and the stride s h in the height dimension is 2. Then, during the execution of the D2S memory rearrangement operation, first, the first output data ACT out ’ is divided into 4 parts, and the length of each part of the data is c out . Then, the data at the corresponding positions in each part of the data are placed in the same data block, where the size of each data block is 2×2, to obtain the second output data ACT out ”. Taking the data at the position (0, 0) as an example, the data at the position (0, 0) in the first output data ACT out ’ can be placed in the data block at the position (0, 0).
[0160] It can be understood that compared with the first output data, the number of channels of the second output data is 1 / N of the first output data.
[0161] In some embodiments, when the stride in the width dimension and the stride in the height dimension are both equal to 1, the target convolution kernel is determined in the following manner:
[0162] Horizontally and vertically flip at least one second convolution kernel corresponding to each output channel of the first convolution kernel along the first plane to obtain the target convolution kernel.
[0163] It can be understood that based on the foregoing analysis of the deconvolution operation process, when the stride in the width dimension and the stride in the height dimension are both equal to 1, when determining the target convolution kernel, there is no need to divide the target convolution kernel into sub-blocks. It only needs to flip each second convolution kernel corresponding to each input channel in each input channel corresponding to each output channel of the first convolution kernel.
[0164] In some embodiments, the process of determining the target convolution kernel described above can be processed offline.
[0165] It can be understood that determining the target convolution kernel offline does not need to occupy the running inference time of the data processing chip in the electronic device, will not excessively increase the software complexity, and can improve the data processing efficiency of the electronic device.
[0166] In some embodiments, the target convolution kernel includes Q sub-convolution kernels (that is, the N sub-convolution kernels corresponding to at least one second convolution kernel included in each output channel of the first convolution kernel in the target convolution kernel), where Q is equal to the product of the number of output channels (c out ) of the first convolution kernel, the number of input channels (c in ) and N. The convolution operations of the Q sub-convolution kernels and the data to be processed can be run in parallel.
[0167] It can be understood that since the convolution operations of each sub-convolution kernel and the data to be processed are run in parallel, that is, the time consumed by this convolution operation is equal to the time consumed by one convolution operation, this convolution operation can be regarded as one convolution operation.
[0168] The data processing method provided by the embodiments of the present application can convert the deconvolution operation into one convolution operation and a memory rearrangement operation, and can directly implement the deconvolution operation using the hardware unit of the convolution in the electronic device without separately designing a deconvolution calculation unit. Moreover, in the data processing method provided by the present application, the deconvolution operation can be converted into one convolution operation, so that only the hardware resources need to be started once and the weights (target parameter information) need to be loaded once, and the load on the hardware is relatively small and the time consumption is relatively short.
[0169] It can be understood that for the above S704, if the sub-convolution kernels are not concatenated along the output channel dimension, it is equivalent to converting the transposed convolution operation into N convolution operations. In the process of convolution operation, multiple hardware resources need to be started and the weights need to be loaded multiple times, which takes a relatively long time (for the convenience of description, this solution is abbreviated as other solutions). In actual engineering implementation, when implementing a single convolution operation in hardware, it is usually necessary to perform alignment operations on the input and output channel dimensions. For example, the number of input channels and the number of output channels are aligned upward to be divisible by 32. Taking 32 alignment as an example, if the stride of the transposed convolution operation is (3, 3) and the number of output channels is 1, the data processing method provided by the present application can convert the transposed convolution operation into a convolution operation with 9 output channels, while other solutions will convert this transposed convolution operation into 9 convolution operations with 1 output channel. Due to the hardware alignment requirements, the data processing method of the present application needs to align the convolution with 9 output channels into a convolution with 32 output channels, where the first 9 output channels are valid data and the other output channels are invalid data. While other solutions need to align 9 convolution operations with 1 output channel into 9 convolution operations with 32 output channels, where the first 1 output channel is valid data and the other output channels are invalid data. It can be seen from this that the data processing method of the present application only needs to start a convolution with 32 output channels once in hardware, while other solutions need to start 9 convolutions with 32 output channels. The hardware computing resources and running time required by the data processing method of the present application are only 11.11% of other solutions. Moreover, if the data processing method of the present application and other solutions are both applied to the fast super-resolution convolutional neural network (FSRCNN), the frame rate of other solutions is 70.67fps and the memory read / write (footprint) is 107.92MB, while the method of the present application has a frame rate of 266.72fps and a memory read / write of 107.92MB in the same environment. It can be seen that the method of the present application has obvious effects in improving the running frame rate of the network in hardware and optimizing the memory read / write.
[0170] The embodiment of the present application also provides a readable storage medium, on which instructions are stored. When the instructions are executed on an electronic device, the electronic device executes the data processing method mentioned in the present application.
[0171] The embodiment of the present application also provides a computer program product, which includes computer instructions. When executed by an electronic device, the electronic device executes the computer program code of the data processing method mentioned in the present application.
[0172] An embodiment of the present application further provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the data processing method mentioned in the present application.
[0173] The method provided by the embodiments of the present application can be applied to any electronic device with a convolution operation unit, including but not limited to a mobile station (MS), a mobile terminal (MT), etc. For example, the electronic device can be a mobile phone, a smart TV, a wearable device, a tablet computer (Pad), a desktop computer, a laptop computer, a virtual reality (VR) device, an augmented reality (AR) device, a terminal in industrial control, a terminal in self-driving, a terminal in remote medical surgery, a terminal in a smart grid, a terminal in transportation safety, a terminal in a smart city, a terminal in a smart home, and so on. The embodiments of the present application do not limit the specific form of the electronic device.
[0174] Figure 14 According to some embodiments of the present application, a schematic structural diagram of an electronic device 100 is shown. As Figure 14 shown, the electronic device 100 includes one or more processors 1001, a system memory 1002, a non-volatile memory (NVM) 1003, a communication interface 1004, an input / output (I / O) device 1005, and a system control logic 1006 for coupling the processor 1001, the system memory 1002, the non-volatile memory 1003, the communication interface 1004, and the input / output device 1005. Among them:
[0175] The processor 1001 may include one or more processing units. For example, it may include a data processing unit or processing circuit such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, a field programmable gate array (FPGA), a neural-network processing unit (NPU), etc., which may include one or more single-core or multi-core processors. In some embodiments, the processor 1001 may include a convolution operation unit.
[0176] The system memory 1002 is volatile memory, such as random-access memory (RAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions.
[0177] The non-volatile memory 1003 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 1003 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 1003 may also be a removable storage medium, such as a secure digital (SD) memory card, etc.
[0178] In particular, the system memory 1002 and the non-volatile memory 1003 may respectively include: a temporary copy and a permanent copy of the instruction 1007. The instruction 1007 may include: when executed by at least one of the processors 1001, enabling the electronic device 100 to implement the data processing methods provided in the embodiments of the present application.
[0179] The communication interface 1004 may include a transceiver for providing a wired or wireless communication interface for the electronic device 100, and thus communicating with any other suitable device via one or more networks. In some embodiments, the communication interface 1004 may be integrated with other components of the electronic device 100. For example, the communication interface 1004 may be integrated in the processor 1001.
[0180] The input / output device 1005 may include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. The user may interact with the electronic device 100 through the input / output device 1005.
[0181] The system control logic 1006 may include any suitable interface controller to provide any suitable interface to other modules of the electronic device 100. For example, in some embodiments, the system control logic 1006 may include one or more memory controllers to provide an interface to the system memory 1002 and the non-volatile memory 1003.
[0182] In some embodiments, at least one of the processors 1001 may be packaged with the logic of one or more controllers for the system control logic 1006 to form a system in package (SiP). In other embodiments, at least one of the processors 1001 may also be integrated with the logic of one or more controllers for the system control logic 1006 on the same chip to form a system-on-chip (SoC).
[0183] It can be understood that Figure 14 The structure of the illustrated electronic device 100 is only an example. In other embodiments, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0184] It can be understood that, as used herein, the term "module" may refer to or include an application-specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or grouped) that executes one or more software or firmware programs, and / or a memory, combinational logic circuits, and / or other suitable hardware components that provide the described functionality, or may be a part of these hardware components.
[0185] It can be understood that in various embodiments of the present application, the processor may be a microprocessor, a digital signal processor, a microcontroller, etc., and / or any combination thereof. According to another aspect, the processor may be a single-core processor, a multi-core processor, etc., and / or any combination thereof.
[0186] The various embodiments disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0187] The program code can be applied to the input instructions to perform the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0188] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When necessary, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in the present application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0189] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions may be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, magneto-optical discs, read only memory (ROM), random access memory (RAM), electrically programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in the form of electrical, optical, acoustic, or other propagated signals using the Internet. Thus, machine-readable media includes any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0190] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0191] It should be noted that the various units / modules mentioned in the device embodiments of the present application are all logical units / modules. Physically, a logical unit / module may be a physical unit / module, a part of a physical unit / module, or may be implemented as a combination of multiple physical units / modules. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above device embodiments.
[0192] It should be noted that in the examples and description of this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0193] Although this application has been illustrated and described by reference to certain embodiments thereof, those of ordinary skill in the art should understand that various changes may be made therein in form and detail without departing from the scope of this application.
Claims
1. A data processing method, characterized in that: include: Acquire data to be processed, where the data to be processed is data to be subjected to a deconvolution operation; Performing a convolution operation on the data to be processed based on a target convolution kernel to obtain first output data; Performing a memory rearrangement operation based on the first output data to obtain target output data; Among them, the target convolution kernel is determined based on the parameter information of the first convolution kernel and the deconvolution operation, and the target output data is the same as the output data obtained by performing a deconvolution operation on the data to be processed based on the first convolution kernel.
2. The method according to claim 1, characterized in that The data of the first convolution kernel includes a width dimension, a height dimension, an input channel dimension, and an output channel dimension, and the parameter information of the deconvolution operation includes a width of the first convolution kernel, a height of the first convolution kernel, a width dimension step, a height dimension step, and the number of input channels; and the first convolution kernel includes at least one output channel, and each output channel of the first convolution kernel includes at least one second convolution kernel, and the number of the at least one second convolution kernel is equal to the number of input channels; When the width dimension step size is greater than 1 or the height dimension step size is greater than 1, the target convolution kernel is determined by: For each output channel of the first convolution kernel, based on the width of the first convolution kernel, the height of the first convolution kernel, the width dimension step, and the height dimension step, the second convolution kernel is divided into M first sub-blocks; wherein the width of the first sub-block is equal to the width dimension step, the height of the first sub-block is equal to the height dimension step, the position of the data element of each first sub-block in the M first sub-blocks corresponds one to one, and M is an integer greater than or equal to 1; Determine N second sub-blocks based on data elements at corresponding positions in the M first sub-blocks; Flipping the N second sub-blocks horizontally and vertically along a first plane respectively to determine N sub-convolution kernels; wherein the first plane is a plane formed by the width dimension and the height dimension; The N sub-convolution kernels corresponding to each of the second convolution kernels in each output channel of the first convolution kernel are spliced along the output channel dimension to obtain the target convolution kernel.
3. The method according to claim 2, characterized in that The first sub-block includes N data elements, the N data elements are respectively located at different positions in the first sub-block, and N is an integer greater than or equal to 1; The determining N second sub-blocks based on the data elements at corresponding positions in the M first sub-blocks includes: Placing data elements with corresponding positions among the N data elements of each of the M first sub-blocks in the same second sub-block to determine the N second sub-blocks; The positions of the N data elements in the second sub-block correspond to the positions of the first sub-block where the N data elements are located in the second convolution kernel.
4. The method according to claim 2, characterized in that: The dividing the second convolution kernel into M first sub-blocks based on the width of the first convolution kernel, the height of the first convolution kernel, the width dimension step size, and the height dimension step size includes: When the width dimension step size cannot be divided by the width of the first convolution kernel, data padding is performed on the width dimension of the second convolution kernel to divide the second convolution kernel into the M first sub-blocks; When the height dimension step cannot be divided by the height of the first convolution kernel, data padding is performed on the height dimension of the second convolution kernel to divide the second convolution kernel into the M first sub-blocks.
5. The method according to claim 2, characterized in that: When the width dimension step size and the height dimension step size are both equal to 1, the target convolution kernel is determined by: The at least one second convolution kernel corresponding to each output channel of the first convolution kernel is horizontally flipped and vertically flipped along the first plane to obtain the target convolution kernel.
6. The method according to any one of claims 2 to 5, characterized in that The target convolution kernel includes Q sub-convolution kernels, wherein Q is equal to the product of the number of output channels of the first convolution kernel, the number of input channels, and N; The step of performing a convolution operation on the data to be processed based on the target convolution kernel to obtain first output data includes: Based on the target parameter information and the Q sub-convolution kernels, performing convolution operations on the data to be processed in parallel to obtain the first output data; Wherein, the target parameter information is determined based on parameter information of the deconvolution operation.
7. The method according to claim 6, characterized in that The performing a memory rearrangement operation based on the first output data to obtain target output data includes: Performing a memory rearrangement operation based on the first output data to obtain second output data; An interception operation is performed based on the second output data to obtain the target output data.
8. An electronic device, characterized in that: include: A memory, used to store instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, used to execute the data processing method described in any one of claims 1 to 7.
9. A readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the data processing method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The computer program product comprises computer instructions, and when executed by an electronic device, the electronic device executes the computer program code of the data processing method according to any one of claims 1 to 7.