A Depthwise Fast Convolution Method and Related Devices
By dividing the image data into a subconvolution matrix in depthwise convolution and using parallel single branch computing subunits to perform parallel convolution operations, the problems of waste of computing resources and large IO operation overhead in the prior art are solved, and efficient computing resource utilization and the effect of reducing IO operation is achieved.
Patent Information
- Application Number
- CN202211710182.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-12-29
AI Technical Summary
The prior art cannot maximize the utilization of matrix computing units in depthwise convolution, resulting in wasted computing resources. When the DepthwiseConv calculation is combined with PointwiseConv calculation, the overhead brought by IO operations seriously affects efficiency.
By obtaining the image data to be convolutionized, it is divided into several image blocks, and converting the image blocks into a matrix to be convolutionized, and dividing them into a preset number of sub-convolution matrices according to the row direction. Then, using the parallel single branch computing subunit in the DepthwiseConv matrix calculation unit, parallel convolution operations are performed on the subconvolution matrix to improve the calculation efficiency, and store the results in the cache to reduce IO operations.
The utilization rate of DepthwiseConv matrix calculation unit is improved, the waste of computing power is avoided, and the impact of IO operations is effectively reduced, and the overall computing efficiency is improved.
Smart Images

Figure CN115952387B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field, and in particular, to a Depthwise fast convolution method and related device. Background Art
[0002] The AI compiler provides computing power through the computing units in the AI Core. Among them, the computing units in the AI Core include matrix computing units, vector computing units, scalar computing units, accumulators, etc. Among them, the matrix computing unit generally adopts a 16*16*16 structure, and one instruction can complete the matrix multiplication calculation of two 16*16 matrices. However, in depthwise convolution, each convolution kernel only processes one channel, resulting in the inability to maximize the utilization of the matrix computing unit and wasting computing resources.
[0003] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The technical problem to be solved by the present application is to provide a Depthwise fast convolution method and related device in view of the deficiencies of the prior art.
[0005] To solve the above technical problem, in the first aspect of the embodiments of the present application, a Depthwise fast convolution method is provided. The method includes:
[0006] Obtain the image data to be convolved, and divide the image data into a plurality of image blocks. Among them, the number of channels of each image block in the plurality of image blocks is equal to a preset number, and the image size of the image block is equal to the size of the convolution kernel;
[0007] Convert each image block into a matrix to be convolved, and divide each matrix to be convolved obtained by conversion into a preset number of sub-convolution matrices in the row direction. Among them, the number of rows of the matrix to be convolved is a preset number, and the number of rows of each sub-convolution matrix is 1;
[0008] Perform parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved in sequence through the DepthwiseConv matrix computing unit to obtain the target image data corresponding to the image data. Among them, the DepthwiseConv matrix computing unit includes a preset number of parallel single-branch computing sub-units.
[0009] In the Depthwise fast convolution method, the preset number is 16.
[0010] In the Depthwise fast convolution method, the step of obtaining the image data to be convolved and dividing the image data into a plurality of image blocks specifically includes:
[0011] Obtain the image data to be convolved and read the number of channels of the image data;
[0012] When the number of channels is equal to the preset number, divide the image data into several image blocks;
[0013] When the number of channels is not equal to the preset number, adjust the number of channels of the image data to the preset number and divide the adjusted image data into several image blocks.
[0014] The Depthwise fast convolution method, wherein the parallel convolution operation is sequentially performed on the preset number of sub-convolution matrices included in each convolution matrix to be convolved through the DepthwiseConv matrix calculation unit to obtain the target image data corresponding to the image data, specifically including:
[0015] Obtain the calculation order of each convolution matrix to be convolved;
[0016] Input the preset number of sub-convolution matrices included in each convolution matrix to be convolved into the DepthwiseConv matrix calculation unit in sequence according to the calculation order, wherein each single-branch calculation subunit in the DepthwiseConv matrix calculation unit corresponds to a sub-convolution matrix;
[0017] Perform parallel convolution calculations on the respective corresponding sub-convolution matrices through each single-branch calculation subunit to obtain the target image data corresponding to the image data.
[0018] The Depthwise fast convolution method, wherein the obtaining of the calculation order of each convolution matrix to be convolved specifically includes:
[0019] Obtain the position information of the image blocks corresponding to each convolution matrix to be convolved in the image data;
[0020] Determine the calculation order of each convolution matrix to be convolved according to the position information.
[0021] The Depthwise fast convolution method, wherein the parallel convolution operation is sequentially performed on the preset number of sub-convolution matrices included in each convolution matrix to be convolved through the DepthwiseConv matrix calculation unit to obtain the target image data corresponding to the image data, specifically including:
[0022] Perform parallel convolution operations on the preset number of sub-convolution matrices included in each convolution matrix to be convolved through the DepthwiseConv matrix calculation unit in sequence, and store the convolution results of each single-branch calculation subunit in the cache;
[0023] When the convolution result in the cache meets the calculation requirements of PointwiseConv, input the convolution result in the cache into the PointwiseConv matrix calculation unit;
[0024] Write the calculation result of the PointwiseConv matrix calculation unit into the memory to obtain the target image data corresponding to the image data.
[0025] The Depthwise fast convolution method, wherein the PointwiseConv matrix calculation unit is parallel to the DepthwiseConv matrix calculation unit.
[0026] The second aspect of the embodiments of the present application provides a Depthwise fast convolution device, and the device includes:
[0027] An acquisition unit, configured to acquire image data to be convolved, and divide the image data into a plurality of image blocks, wherein the number of channels of each image block in the plurality of image blocks is equal to a preset number, and the image size of the image block is equal to the size of the convolution kernel;
[0028] A division unit, configured to convert each image block into a matrix to be convolved, and divide each matrix to be convolved obtained by conversion into a preset number of sub-convolution matrices in the row direction, wherein the number of rows of the matrix to be convolved is a preset number, and the number of rows of each sub-convolution matrix is 1;
[0029] A DepthwiseConv matrix calculation unit, configured to perform parallel convolution operations on a preset number of sub-convolution matrices included in each matrix to be convolved to obtain the target image data corresponding to the image data, wherein the DepthwiseConv matrix calculation unit includes a preset number of parallel single-branch calculation sub-units.
[0030] The third aspect of the embodiments of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any one of the above-mentioned Depthwise fast convolution methods.
[0031] The fourth aspect of the embodiments of the present application provides a terminal device, which includes: a processor, a memory, and a communication bus; a computer-readable program executable by the processor is stored on the memory;
[0032] The communication bus realizes the connection and communication between the processor and the memory;
[0033] When the processor executes the computer-readable program, the steps in any one of the above-mentioned Depthwise fast convolution methods are implemented.
[0034] Beneficial effects: Compared with the prior art, the present application provides a Depthwise fast convolution method and related device. The method includes obtaining image data to be convolved, and dividing the image data into a plurality of image blocks; converting each image block into a to-be-convolved matrix, and dividing each obtained to-be-convolved matrix into a preset number of sub-convolution matrices in the row direction; sequentially performing parallel convolution operations on the preset number of sub-convolution matrices included in each to-be-convolved matrix through a DepthwiseConv matrix calculation unit to obtain target image data corresponding to the image data. In the present application, by setting a preset number of single-branch sub-calculation units in the DepthwiseConv matrix calculation unit, and then performing parallel calculations on the preset number of channels of the DepthwiseConv convolution through the preset number of single-branch sub-calculation units, the utilization rate of the DepthwiseConv matrix calculation unit can be improved, and the waste of computing power can be avoided. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative labor, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is a flowchart of the Depthwise fast convolution method provided by the present application.
[0037] Figure 2 It is a schematic flow chart of the combination of the Depthwise fast convolution method provided by the present application and the PointwiseConv calculation.
[0038] Figure 3 It is a schematic structural diagram of the Depthwise fast convolution device provided by the present application.
[0039] Figure 4 It is a schematic structural diagram of the terminal device provided by the present application. Detailed Embodiments
[0040] The present application provides a Depthwise fast convolution method and related device. To make the purpose, technical solutions and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0041] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0042] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0043] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0044] Through research, it is found that the AI compiler provides computing power through the computing units in the AI Core. Among them, the computing units in the AI Core include matrix computing units, vector computing units, scalar computing units, accumulators, etc. Among them, the matrix computing unit generally adopts a 16*16*16 structure, and one instruction can complete the matrix multiplication calculation of two 16*16. However, in depthwise convolution, each convolution kernel only processes one channel, resulting in the inability to maximize the utilization of the matrix computing unit and wasting computing resources. In addition, DepthwiseConv calculations are generally used in combination with PointwiseConv calculations. The calculation results of DepthwiseConv will be stored in the IO, and then the data will be read from the IO as the input of PointwiseConv. The processes of storing in the IO and reading data from the IO will bring a large amount of overhead, especially in the case of a large amount of data, which will seriously affect the efficiency of DepthwiseConv calculations.
[0045] In order to solve the above problems, in an embodiment of the present application, the image data to be convolved is obtained, and the image data is divided into a number of image blocks; each image block is converted into a matrix to be convolved, and each converted matrix to be convolved is divided into a preset number of sub-convolution matrices in the row direction; the DepthwiseConv matrix calculation unit sequentially performs parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved, so as to obtain the target image data corresponding to the image data. The present application sets a preset number of single-branch sub-computing units in the DepthwiseConv matrix calculation unit, and then performs parallel calculations on a preset number of channels of the DepthwiseConv convolution through the preset number of single-branch sub-computing units, which can improve the utilization rate of the DepthwiseConv matrix calculation unit and avoid waste of computing power. In addition, when the DepthwiseConv matrix calculation unit performs parallel convolution operations on a preset number of sub-convolution matrices included in the convolution matrix, the calculated results are stored in the cache. When the time in the cache is sufficient for PointwiseConv calculation, the calculation results in the cache are input into the PointwiseConv matrix calculation unit, and the PointwiseConv calculation is performed synchronously. This avoids the process of storing IO and reading data from IO, and effectively solves the impact of IO operations on DepthwiseConv.
[0046] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.
[0047] This embodiment provides a Depthwise fast convolution method, such as Figure 1 As shown, the method comprises:
[0048] S10, obtaining image data to be convolved, and dividing the image data into a plurality of image blocks.
[0049] Specifically, the image data is the image to be calculated by DepthwiseConv, which can be the output item of other convolution layers, or the original image after convolution calculation.
[0050] The image data is a 64*64*16 feature map. The image blocks are obtained by dividing the image data, wherein the number of channels of each image block in a plurality of 5 image blocks is equal to the preset number, and the image size of the image block is equal to the size of the convolution kernel. That is to say, when dividing the image data, the image data is divided according to the size of the convolution kernel so that the image size of each image block is the same as the size of the convolution kernel, and at the same time, it is divided according to the preset number in the channel direction so that the number of channels of each image block obtained by division is equal to the preset number.
[0051] The preset number is determined based on the number of parallel sub-computation units in the DepthwiseConv matrix calculation unit for performing DepthwiseConv calculation, wherein the preset number is equal to the number of parallel sub-computation units in the DepthwiseConv matrix calculation unit. In other words, the number of channels of the image block is equal to the number of parallel sub-computation units in the DepthwiseConv matrix calculation unit.
[0052] The number of units is such that the image block can be calculated along the channel and 5 rows by the DepthwiseConv matrix calculation unit. In one implementation, the preset number is 16. Of course, in practical applications, the preset number can also be other numbers, such as 8, 32, etc.
[0053] Furthermore, since the number of channels of the image data may not be equal to the preset number, for example, the number of channels may be greater than the preset number, or the number of channels may be less than the preset number. Therefore, when dividing the image data into a plurality of image blocks, it is possible to first detect whether the number of channels of the image data is equal to the preset number, and adjust the image data according to the detection result before dividing. Based on this, in one implementation, the obtaining of the image data to be convolved and the dividing of the image data into a plurality of image blocks specifically include:
[0054] S11, obtaining image data to be convolved, and reading the number of channels of the image data;
[0055] S12, when the number of channels is equal to a preset number, dividing the image data into a plurality of image blocks;
[0056] S13. When the number of channels is not equal to a preset number, the number of channels of the image data is adjusted to a preset number, and the adjusted image data is divided into a plurality of image blocks.
[0057] Specifically, when the number of channels is not equal to the preset number, the number of channels is greater than the preset number, or the number of channels is less than the preset number. Among them, when the number of channels is greater than the preset number, the image data is divided along the channel direction based on the preset number, so that the number of channels of each divided image data block is the preset number, and then each divided image data block is further divided to obtain a number of image blocks; conversely, when the number of channels is less than the preset number, the width or height of the image data can be folded onto the channels so that the number of channels of the folded image data is equal to the preset number. Of course, it should be noted that when dividing the image data blocks according to the preset number when the number of channels is greater than the preset number, if there are image data blocks that do not meet the preset number, it can also be adjusted by folding the width or height onto the channels. In addition, after adjusting the number of channels of the image data to the preset number, the image data is divided based on the size of the convolutional kernel, so that the image size of each divided image block is equal to the size of the convolutional kernel and the number of channels is equal to the preset number.
[0058] S20. Convert each image block into a matrix to be convolved, and divide the obtained matrices to be convolved into a preset number of sub-convolution matrices in the row direction.
[0059] Specifically, the matrix to be convolved can be obtained by performing an img2col operation on the image block. Among them, the number of columns of the matrix to be convolved is equal to the product of the sizes of the convolutional kernel, and the number of rows is equal to the preset number, that is, the number of rows is equal to the number of channels of the image block. For example, if the size of the convolutional kernel is 3*3 and the preset number is 16, then the number of rows of the matrix to be convolved is 16 and the number of columns is 9. In addition, after obtaining the matrix to be convolved, the matrix to be convolved is divided into a preset number of sub-convolution matrices, where the number of rows of the sub-convolution matrix is equal to the number of rows of the matrix to be convolved and the number of columns is 1. For example, if the matrix to be convolved is a 16*9 matrix, then the matrix to be convolved will be divided into 16 sub-convolution matrices of 1*9.
[0060] S30. Sequentially perform parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved through the DepthwiseConv matrix calculation unit to obtain the target image data corresponding to the image data.
[0061] Specifically, the DepthwiseConv matrix calculation unit includes a preset number of parallel single-branch calculation sub-units. Each single-branch calculation sub-unit supports one-way calculation, that is, each single-branch calculation sub-unit is used to calculate one channel, and the preset number of single-branch calculation sub-units can perform parallel calculations on the preset number of channels. That is to say, for the input data input into the DepthwiseConv matrix calculation unit, the preset number of single-branch calculation sub-units can be used to perform parallel convolution calculations on the preset number of channels of the input data, improving the calculation efficiency of the DepthwiseConv matrix calculation unit and avoiding waste of the computing power of the DepthwiseConv matrix calculation unit.
[0062] In one implementation, since DepthwiseConv calculation and PointwiseConv calculation are generally performed jointly, and PointwiseConv calculation is used to perform calculations along the channel direction. Therefore, when performing DepthwiseConv calculation, the image blocks corresponding to the same image position are calculated to obtain the calculation results corresponding to each channel at this image position, so that the DepthwiseConv calculation and the PointwiseConv calculation can be synchronized.
[0063] Based on this, when calculating each convolution matrix to be calculated in sequence through the DepthwiseConv matrix calculation unit, the calculation order of each convolution matrix to be calculated can be obtained first, and then the calculation is performed according to the calculation order. Therefore, the process of sequentially performing parallel convolution operations on the preset number of sub-convolution matrices included in each convolution matrix to be calculated through the DepthwiseConv matrix calculation unit to obtain the target image data corresponding to the image data specifically includes:
[0064] Obtain the calculation order of each convolution matrix to be calculated;
[0065] Sequentially input the preset number of sub-convolution matrices included in each convolution matrix to be calculated into the DepthwiseConv matrix calculation unit according to the calculation order, where each single-branch calculation sub-unit in the DepthwiseConv matrix calculation unit corresponds to a sub-convolution matrix;
[0066] Perform parallel convolution calculations on the sub-convolution matrices corresponding to each single-branch calculation sub-unit to obtain the target image data corresponding to the image data.
[0067] Specifically, the calculation order is determined according to the sequence of calculation results required for PointwiseConv calculation, so as to parallelize the DepthwiseConv calculation and the PointwiseConv calculation. In this embodiment, the calculation order can be determined according to the position information of the image patches in the image data. Specifically, the specific process of the calculation order can be as follows:
[0068] Obtain the position information of the image patches corresponding to each convolution matrix to be calculated in the image data;
[0069] Determine the calculation order corresponding to each convolution matrix to be calculated according to the position information.
[0070] Specifically, the position information can be the position information of a pixel point in the image patch. The position information includes the horizontal coordinate, vertical coordinate of the pixel, and the channel position of the pixel point. The pixel point can be the upper left pixel point, the upper right pixel point, etc. After obtaining the position information, first sort by the horizontal coordinate as the priority, secondly, for the image patches in the same order, sort by the vertical coordinate of the pixel as the priority, and finally, for the image patches in the same order, sort by the channel position as the priority to obtain the calculation order. For example, if the position information of image patch A is (1, 1, 1) and the position information of image patch B is (2, 1, 1), then the calculation order of image patch A takes precedence over that of image patch B. Of course, in practical applications, other sorting methods can also be used, for example, sorting in turn with the vertical coordinate of the pixel - the horizontal coordinate of the pixel - the pixel position as the priority, etc.
[0071] In one implementation, as Figure 2 shown, when the DepthwiseConv calculation and the PointwiseConv calculation are used in combination, the process of the DepthwiseConv matrix calculation unit sequentially performing parallel convolution operations on a preset number of sub-convolution matrices included in each convolution matrix to be calculated to obtain the target image data corresponding to the image data specifically includes:
[0072] The DepthwiseConv matrix calculation unit sequentially performs parallel convolution operations on a preset number of sub-convolution matrices included in each convolution matrix to be calculated, and stores the convolution results of each single-branch calculation sub-unit in the cache;
[0073] When the convolution results in the cache meet the PointwiseConv calculation requirements, input the convolution results in the cache into the PointwiseConv matrix calculation unit;
[0074] Write the calculation results of the PointwiseConv matrix calculation unit into the memory to obtain the target image data corresponding to the image data.
[0075] Specifically, the PointwiseConv calculation requirement is that all convolution results in the channel direction of a pixel position are stored in the cache. That is to say, after storing the convolution results of each single-branch calculation subunit in the cache, it will be detected in real time whether all convolution results in the channel direction of a pixel position are stored in the cache. When all convolution results in the channel direction of a pixel position are stored in the cache, it means that the convolution results in the cache meet the PointwiseConv calculation requirement. At this time, the convolution results that meet the PointwiseConv calculation requirement are read, and then the read convolution results are calculated by the PointwiseConv matrix calculation unit, and the calculation results are written into the content. In addition, when the convolution results are input into the PointwiseConv matrix calculation unit and the calculation structure is obtained through the PointwiseConv matrix calculation unit, the convolution results input into the PointwiseConv matrix calculation unit are deleted from the cache. In this way, on the one hand, the occupation of the cache by the convolution results can be reduced, and on the other hand, the problem of repeated PointwiseConv matrix calculations of the convolution results can be avoided.
[0076] In this embodiment, by storing the convolution results in the cache and then detecting in real time whether there are convolution results that meet the PointwiseConv calculation in the cache, when there are convolution results that meet the PointwiseConv calculation, the PointwiseConv calculation is performed in parallel. In this way, on the one hand, the calculation efficiency of the combined calculation of DepthwiseConv calculation and PointwiseConv calculation can be improved. On the other hand, the convolution structure obtained by performing DepthwiseConv calculation on each channel of the image data is not directly stored in the DDR, but cached in the UB. When the output of the calculation is sufficient for the PointwiseConv calculation of one point, it is input into the corresponding calculation unit of the PointwiseConv for calculation, which can avoid reducing the IO operation and effectively solve the influence of the IO operation on the DepthwiseConv, especially the influence brought by the IO bottleneck under a large amount of data.
[0077] In summary, this embodiment provides a Depthwise fast convolution method. The method includes obtaining image data to be convolved and dividing the image data into several image blocks; converting each image block into a matrix to be convolved, and dividing each obtained matrix to be convolved into a preset number of sub-convolution matrices in the row direction; performing parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved in sequence through a DepthwiseConv matrix calculation unit to obtain target image data corresponding to the image data. In this application, a preset number of single-branch sub-calculation units are set in the DepthwiseConv matrix calculation unit, and then the preset number of channels of the DepthwiseConv convolution are calculated in parallel through the preset number of single-branch sub-calculation units, which can improve the utilization rate of the DepthwiseConv matrix calculation unit and avoid waste of computing power. In addition, when performing parallel convolution operations on the preset number of sub-convolution matrices included in the matrix to be convolved through the DepthwiseConv matrix calculation unit, the calculated results will be stored in the cache. When the time in the cache allows for PointwiseConv calculation, the calculated results in the cache are input into the PointwiseConv matrix calculation unit to perform PointwiseConv calculation synchronously, which can avoid the process of storage IO and reading data from IO, and effectively solve the influence of IO operations on DepthwiseConv.
[0078] Based on the above Depthwise fast convolution method, this embodiment provides a Depthwise fast convolution device, as Figure 3 shown, the device includes:
[0079] An acquisition unit 100, configured to obtain image data to be convolved and divide the image data into several image blocks, where the number of channels of each image block in the several image blocks is equal to a preset number, and the image size of the image block is equal to the size of the convolution kernel;
[0080] A division unit 200, configured to convert each image block into a matrix to be convolved, and divide each obtained matrix to be convolved into a preset number of sub-convolution matrices in the row direction, where the number of rows of the matrix to be convolved is a preset number, and the number of rows of each sub-convolution matrix is 1;
[0081] A DepthwiseConv matrix calculation unit 300, configured to perform parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved to obtain target image data corresponding to the image data, where the DepthwiseConv matrix calculation unit includes a preset number of parallel single-branch calculation sub-units.
[0082] Based on the above Depthwise fast convolution method, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the Depthwise fast convolution method as described in the above embodiment.
[0083] Based on the above Depthwise fast convolution method, this application also provides a terminal device, as Figure 4 shown, which includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiment.
[0084] In addition, when the logical instructions in the above-mentioned memory 22 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0085] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the method in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, to implement the method in the above embodiment.
[0086] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes can also be a transient storage medium.
[0087] In addition, the specific processes of loading and executing multiple instructions by the above storage medium and the instruction processor in the terminal device have been described in detail in the above method, and will not be repeated here one by one.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A Depthwise fast convolution method, characterized in that The method includes: Obtain image data to be convolved, and divide the image data into a plurality of image blocks, where the number of channels of each image block among the plurality of image blocks is equal to a preset number, and the image size of the image block is equal to the size of the convolution kernel; Convert each image block into a matrix to be convolved, and divide each matrix to be convolved obtained by conversion into a preset number of sub-convolution matrices in the row direction, where the number of rows of the matrix to be convolved is a preset number, and the number of rows of each sub-convolution matrix is 1; Perform parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved in sequence through a DepthwiseConv matrix calculation unit, so as to obtain target image data corresponding to the image data, where the DepthwiseConv matrix calculation unit includes a preset number of parallel single-branch calculation sub-units; Among them, the step of performing parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved in sequence through a DepthwiseConv matrix calculation unit to obtain target image data corresponding to the image data specifically includes: Perform parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved in sequence through a DepthwiseConv matrix calculation unit, and store the convolution results of each single-branch calculation sub-unit in a cache; When the convolution results in the cache meet the PointwiseConv calculation requirements, input the convolution results in the cache into the PointwiseConv matrix calculation unit; Write the calculation results of the PointwiseConv matrix calculation unit into memory to obtain target image data corresponding to the image data; The step of obtaining image data to be convolved and dividing the image data into a plurality of image blocks specifically includes: Obtain image data to be convolved, and read the number of channels of the image data; When the number of channels is equal to the preset number, divide the image data into a plurality of image blocks; When the number of channels is not equal to the preset number, adjust the number of channels of the image data to the preset number, and divide the adjusted image data into a plurality of image blocks. Among them, when the number of channels is greater than the preset number, divide the image data along the channel direction based on the preset number, so that the number of channels of each divided image data block is the preset number, and then divide each divided image data block to obtain a plurality of image blocks; conversely, when the number of channels is less than the preset number, fold the width or height of the image data onto the channels so that the number of channels of the folded image data is equal to the preset number.
2. The Depthwise fast convolution method according to claim 1, wherein The preset number is 16.
3. The Depthwise fast convolution method according to claim 1, wherein The step of performing parallel convolution operations on the preset number of sub-convolution matrices included in each matrix to be convolved in sequence through a DepthwiseConv matrix calculation unit to obtain target image data corresponding to the image data specifically includes: Obtain the calculation order of each matrix to be convolved; Sequentially input the preset number of sub-convolution matrices included in each convolution matrix to be processed into the DepthwiseConv matrix calculation unit, where each single-branch calculation subunit in the DepthwiseConv matrix calculation unit corresponds to a sub-convolution matrix; Perform parallel convolution calculations on the sub-convolution matrices corresponding to each single-branch calculation subunit respectively to obtain the target image data corresponding to the image data.
4. The Depthwise fast convolution method according to claim 3, wherein, The specific method for obtaining the calculation order of each convolution matrix to be processed includes: Obtain the position information of the image blocks corresponding to each convolution matrix to be processed in the image data; Determine the calculation order of each convolution matrix to be processed according to the position information.
5. The Depthwise fast convolution method according to claim 1, wherein , The PointwiseConv matrix calculation unit is parallel to the DepthwiseConv matrix calculation unit.
6. A Depthwise fast convolution device, characterized in that The device includes: An acquisition unit, configured to acquire the image data to be convolved, and divide the image data into several image blocks, where the number of channels of each image block in the several image blocks is equal to the preset number, and the image size of the image block is equal to the size of the convolution kernel; A division unit, configured to convert each image block into a convolution matrix to be processed, and divide each convolution matrix obtained by conversion into a preset number of sub-convolution matrices in the row direction, where the number of rows of the convolution matrix to be processed is the preset number, and the number of rows of each sub-convolution matrix is 1; A DepthwiseConv matrix calculation unit, configured to perform parallel convolution operations on the preset number of sub-convolution matrices included in each convolution matrix to be processed to obtain the target image data corresponding to the image data, where the DepthwiseConv matrix calculation unit includes a preset number of parallel single-branch calculation subunits; Among them, performing parallel convolution operations on the preset number of sub-convolution matrices included in each convolution matrix to be processed to obtain the target image data corresponding to the image data specifically includes: Perform parallel convolution operations on the preset number of sub-convolution matrices included in each convolution matrix to be processed in sequence, and store the convolution results of each single-branch calculation subunit in a cache; When the convolution results in the cache meet the PointwiseConv calculation requirements, input the convolution results in the cache into the PointwiseConv matrix calculation unit; Write the calculation results of the PointwiseConv matrix calculation unit into the memory to obtain the target image data corresponding to the image data; The specific method for acquiring the image data to be convolved and dividing the image data into several image blocks includes: Acquire the image data to be convolved, and read the number of channels of the image data; When the number of channels is equal to the preset number, divide the image data into several image blocks; When the number of channels is not equal to the preset number, adjust the number of channels of the image data to the preset number, and divide the adjusted image data into a plurality of image blocks. Specifically, when the number of channels is greater than the preset number, divide the image data along the channel direction based on the preset number, so that the number of channels of each obtained image data block is the preset number, and then divide each obtained image data block to obtain a plurality of image blocks; conversely, when the number of channels is less than the preset number, fold the width or height of the image data onto the channels so that the number of channels of the folded image data is equal to the preset number.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the Depthwise fast convolution method according to any one of claims 1-6.
8. A terminal device, characterized in that, Including: A processor, a memory, and a communication bus; a computer-readable program executable by the processor is stored on the memory; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, it implements the steps in the Depthwise fast convolution method according to any one of claims 1-6.
Citation Information
Patent Citations
Convolution calculation method, system and equipment and storage medium
CN113870091A
Dethwise fast convolution operation method and system
CN114329323A