Depthwise Fast Convolution Operation Method and System
By dividing the data of the matrix to be convolution into several target matrix data groups and performing convolution operations in parallel, the problem of wasted parallelism in depthwise convolution is solved, and the multiplier resource utilization and computing efficiency are improved.
Patent Information
- Application Number
- CN202111481789.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-06
AI Technical Summary
In depthwise convolution operation, the parallelism of the channel direction is wasted, resulting in the waste of multiplier resources and affecting the operation efficiency.
By dividing the data of the matrix to be convolution into several target matrix data groups, performing convolution operations in parallel, and using multiple multipliers and accumulators to perform parallel convolution, the parallelism in the channel direction and the resource utilization rate of the multiplier are improved.
Improve the efficiency of depthwise convolution operation, reduce the waste of multiplier resources, and improve the computing performance.
Smart Images

Figure CN114329323B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision, and particularly to a depthwise fast convolution operation method, system, and computer-readable storage medium. Background Art
[0002] In traditional data stream convolution hardware acceleration designs, in order to consider generality (the size of the spatial dimension of the convolution kernel may vary, and there may even be a spatial dimension of 1*1), parallel operations are not often performed on the spatial dimension within a convolution kernel, but rather parallel operations are performed on the channel dimension / between multiple convolution kernels (calculate the first point in multiple channels simultaneously, then calculate the second point in multiple channels simultaneously...). This is an efficient approach in ordinary convolutions (the number of channels is very large in the middle layer). However, in depthwise convolutions, each convolution kernel only processes one channel, resulting in waste of parallelism in the channel direction, and since among the hardware resources used in convolution operations, the multiplier is the operation resource that occupies the most resources (area, power consumption). Therefore, waste of multipliers often means a large amount of resource waste.
[0003] The above content is only used to assist in understanding the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of the present invention is to provide a depthwise fast convolution operation method, system, and computer-readable storage medium, aiming to improve the utilization rate of multiplier resources in depthwise convolution operations.
[0005] To achieve the above objective, the present invention provides a depthwise fast convolution operation method, and the steps of the depthwise fast convolution operation method include:
[0006] Obtain the matrix data to be convolved;
[0007] Determine a plurality of target matrix data groups corresponding to the data parallelism according to the data parallelism;
[0008] Successively perform parallel convolution operations on each of the target matrix data groups and the convolution kernel to obtain target convolution results corresponding to each of the target matrix data groups;
[0009] Determine the target convolution matrix data corresponding to the matrix to be convolved according to the target convolution results.
[0010] Optionally, the step of determining a plurality of target matrix groups corresponding to the data parallelism according to the data parallelism includes:
[0011] Determine a plurality of sub - matrices corresponding to the size in the matrix data to be convolved according to the size of the convolution kernel;
[0012] Divide the sub - matrices into several groups of target matrix data according to the data parallelism, wherein the number of sub - matrices in each group of target matrix data is equal to the data parallelism.
[0013] Optionally, the group of target matrix data includes the first group of target matrix data, the second group of target matrix data to the Nth group of target matrix data. The step of sequentially performing parallel convolution operations on each group of target matrix data and the convolution kernel to obtain the target convolution results corresponding to each group of target matrix data includes:
[0014] Perform parallel convolution operations on each first sub - matrix of the first group of target matrix data and the convolution kernel to obtain the first group of target convolution results corresponding to the first group of target matrix data;
[0015] After obtaining the first group of target convolution results, perform parallel convolution operations on each second sub - matrix of the second group of target matrix data and the convolution kernel to obtain the second group of target convolution results corresponding to the second group of target matrix data;
[0016] Sequentially perform parallel convolution operations to obtain the Nth group of target convolution results corresponding to the Nth group of target matrix data;
[0017] Determine the target convolution results according to the first group of target convolution results, the second group of target convolution results to the Nth group of target convolution results.
[0018] Optionally, the first group of target convolution results includes the convolution results corresponding to each first sub - matrix in the first group of target matrix data, the second group of target convolution results includes the convolution results corresponding to each second sub - matrix in the second group of target matrix data, and the Nth group of target convolution results includes the convolution results corresponding to each Nth sub - matrix in the Nth group of target matrix data.
[0019] Optionally, the step of performing parallel convolution operations on each first sub - matrix of the first group of target matrix data and the convolution kernel to obtain the first group of target convolution results corresponding to the first group of target matrix data includes:
[0020] In one time period, determine the first group of data corresponding to the first group of target matrix data and the first weight coefficients corresponding to the convolution kernel, and simultaneously perform multiplication operations on each first data in the first group of data and the first weight coefficients to obtain the first processing results corresponding to the first group of data. The first group of data includes the first data of each first sub - matrix in the first group of target matrix data;
[0021] After obtaining the first processing result, extract the second set of data and the second weight coefficient corresponding to the target matrix data group, and perform a multiplication operation on each second data in the second set of data with the second weight coefficient simultaneously to obtain the second processing result corresponding to the second set of data, where the second set of data includes the second data of each first sub-matrix in the first target matrix data group;
[0022] Extract the data of the target matrix group and obtain the processing result in sequence to obtain the Nth processing result;
[0023] Determine the first target convolution result group according to the first processing result, the second processing result to the Nth processing result.
[0024] Optionally, the step of determining the first target convolution result according to the first processing result, the second processing result to the Nth processing result includes:
[0025] After obtaining the first processing result and the second processing result, superimpose the second processing result onto the first processing result respectively to update the first processing result;
[0026] After obtaining the third processing result, superimpose the third processing result onto the updated first processing result respectively to update the first processing result again;
[0027] Update the first processing result in sequence;
[0028] Determine the first processing result after the last update as the first target convolution result group.
[0029] Optionally, the step of determining the target convolution matrix data corresponding to the matrix to be convolved according to the target convolution result includes:
[0030] Determine the convolution result corresponding to each sub-matrix according to the target convolution result and the data of the matrix to be convolved;
[0031] Determine the target convolution matrix data according to the convolution result corresponding to each sub-matrix.
[0032] In addition, to achieve the above object, the present invention also provides a depthwise fast convolution operation system, where the depthwise fast convolution operation system includes: a memory, a processor, and a depthwise fast convolution operation program stored on the memory and executable on the processor, and when the depthwise fast convolution operation program is executed by the processor, the steps of the depthwise fast convolution operation method as described above are implemented.
[0033] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, on which a depthwise fast convolution operation program is stored. When the depthwise fast convolution operation program is executed by a processor, the steps of the depthwise fast convolution operation method as described above are implemented.
[0034] A depthwise fast convolution operation method, system and computer-readable storage medium proposed by an embodiment of the present invention, after obtaining data of a matrix to be convolved, determine each target matrix data group in the data of the matrix to be convolved according to the parallelism. The number of matrices in each target matrix data group is equal to the data parallelism. After determining each target matrix data group, sequentially perform parallel convolution operations on each target matrix data group and a convolution kernel, that is, simultaneously perform parallel convolution operations on each matrix in a target matrix data group and the convolution kernel. After calculating a target matrix data group, simultaneously perform parallel convolution operations on each matrix in the next target matrix data group and the convolution kernel until all parallel convolution operations on all target matrix data groups and the convolution kernel are completed, thereby obtaining target convolution results corresponding to each target matrix data group, and then determining target convolution data corresponding to the data of the matrix to be convolved according to the target convolution results. By performing parallel convolution operations on each target matrix data group and the convolution kernel, the utilization rate of the parallelism in the channel direction is improved, thereby improving the resource utilization rate of the multiplier and improving the convolution operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic diagram of the terminal structure of the hardware operating environment related to the solution of the embodiment of the present invention;
[0036] Figure 2 is a schematic flowchart of the first embodiment of the depthwise fast convolution operation method of the present invention;
[0037] Figure 3 is an architecture diagram of the convolution operation device of the present invention;
[0038] Figure 4 is a schematic detailed flowchart of step S20 of the first embodiment of the depthwise fast convolution operation method of the present invention;
[0039] Figure 5 is an example diagram of obtaining a sub-matrix in the first embodiment of the depthwise fast convolution operation method of the present invention;
[0040] Figure 6 is an example diagram of obtaining a target matrix data group in the first embodiment of the depthwise fast convolution operation method of the present invention;
[0041] Figure 7 It is a schematic diagram of the refined process of step S30 in the first embodiment of the depthwise fast convolution operation method of the present invention;
[0042] Figure 8 It is a schematic diagram of the refined process of step S31 in the second embodiment of the depthwise fast convolution operation method of the present invention.
[0043] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0044] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0045] The main solution of the embodiment of the present invention is: obtaining the data of the matrix to be convolved; determining a plurality of target matrix data groups corresponding to the data parallelism according to the data parallelism; sequentially performing parallel convolution operations on each of the target matrix data groups and the convolution kernel to obtain the target convolution results corresponding to each of the target matrix data groups; determining the target convolution matrix data corresponding to the data of the matrix to be convolved according to the target convolution results.
[0046] As Figure 1 shown, Figure 1 It is a schematic diagram of the terminal structure of the hardware operating environment involved in the embodiment solution of the present invention.
[0047] The terminal in the embodiment of the present invention can be a PC, or a terminal device with processing functions such as a smart phone, a tablet computer, and a portable computer.
[0048] As Figure 1 shown, the terminal may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the foregoing processor 1001.
[0049] Those skilled in the art can understand, Figure 1The terminal structure shown does not limit the terminal, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0050] As Figure 1 shown, in the memory 1005 as a computer storage medium, an operating system, a network communication module, a user interface module, and a depthwise fast convolution operation program may be included.
[0051] In Figure 1 the terminal shown, the network interface 1004 is mainly used to connect to the background server and communicate with the background server for data; the user interface 1003 is mainly used to connect to the client (user side) and communicate with the client for data; and the processor 1001 may be used to call the depthwise fast convolution operation program stored in the memory 1005 and perform the following operations:
[0052] Obtain the data of the matrix to be convolved;
[0053] Determine a plurality of target matrix data groups corresponding to the data parallelism according to the data parallelism;
[0054] Successively perform parallel convolution operations on each of the target matrix data groups and the convolution kernel to obtain the target convolution results corresponding to each of the target matrix data groups;
[0055] Determine the target convolution matrix data corresponding to the matrix data to be convolved according to the target convolution results.
[0056] Further, the processor 1001 may call the depthwise fast convolution operation program stored in the memory 1005 and also perform the following operations:
[0057] Determine a plurality of sub-matrices corresponding to the size in the matrix data to be convolved according to the size of the convolution kernel;
[0058] Divide the sub-matrices into several target matrix data groups according to the data parallelism, where the number of matrices in the sub-matrices in the target matrix data group is equal to the data parallelism.
[0059] Further, the processor 1001 may call the depthwise fast convolution operation program stored in the memory 1005 and also perform the following operations:
[0060] Perform parallel convolution operations on each of the first sub-matrices of the first target matrix data group and the convolution kernel to obtain the first target convolution result group corresponding to the first target matrix data group;
[0061] After obtaining the first target convolution result group, perform parallel convolution operations on each second sub-matrix of the second target matrix data group and the convolution kernel to obtain a second target convolution result group corresponding to the second target matrix data group;
[0062] Perform parallel convolution operations in sequence to obtain an Nth target convolution result group of the Nth target matrix data group;
[0063] Determine the target convolution result based on the first target convolution result group, the second target convolution result group to the Nth target convolution result group.
[0064] Further, the processor 1001 can call the depthwise fast convolution operation program stored in the memory 1005 and also perform the following operations:
[0065] Within one time period, determine the first group of data corresponding to the first target matrix data group and the first weight coefficient corresponding to the convolution kernel, and perform multiplication operations on each first data in the first group of data and the first weight coefficient simultaneously to obtain a first processing result corresponding to the first group of data. The first group of data includes the first data of each first sub-matrix in the first target matrix data group;
[0066] After obtaining the first processing result, extract the second group of data corresponding to the target matrix data group and the second weight coefficient, and perform multiplication operations on each second data in the second group of data and the second weight coefficient simultaneously to obtain a second processing result corresponding to the second group of data. The second group of data includes the second data of each second sub-matrix in the first target matrix data group;
[0067] Extract the data of the target matrix group and obtain processing results in sequence to obtain the Nth processing result;
[0068] Determine the first target convolution result group based on the first processing result, the second processing result to the Nth processing result.
[0069] Further, the processor 1001 can call the depthwise fast convolution operation program stored in the memory 1005 and also perform the following operations:
[0070] After obtaining the first processing result and the second processing result, superimpose the second processing result onto the first processing result respectively to update the first processing result;
[0071] After obtaining the third processing result, superimpose the third processing result onto the updated first processing result respectively to update the first processing result again;
[0072] Update the first processing result sequentially;
[0073] Determine the first processing result after the last update as the first target convolution result group.
[0074] Further, the processor 1001 can call the depthwise fast convolution operation program stored in the memory 1005 and further perform the following operations:
[0075] Determine the convolution results corresponding to each sub-matrix according to the target convolution result to determine the data of the matrix to be convolved;
[0076] Determine the target convolution matrix data according to the convolution results corresponding to each sub-matrix.
[0077] Refer to Figure 2 , the first embodiment of the depthwise fast convolution operation method of the present invention provides a depthwise fast convolution operation method, and the steps of the depthwise fast convolution operation method include:
[0078] Step S10, obtain the data of the matrix to be convolved;
[0079] Step S20, determine a plurality of target matrix data groups corresponding to the data parallelism according to the data parallelism;
[0080] Step S30, sequentially perform parallel convolution operations on each of the target matrix data groups and the convolution kernel to obtain the target convolution results corresponding to each of the target matrix data groups;
[0081] Step S40, determine the target convolution matrix data corresponding to the data of the matrix to be convolved according to the target convolution result.
[0082] In the embodiment of the present application, the depthwise fast convolution operation method is applied to a depthwise fast convolution operation system. The depthwise fast convolution operation system includes an input device for inputting the data of the matrix to be convolved. The depthwise fast convolution operation system further includes a data preprocessing device for determining a plurality of target matrix data groups corresponding to the data parallelism according to the data parallelism. The depthwise fast convolution operation system further includes a convolution operation device for sequentially performing parallel convolution operations on each of the target matrix data groups and the convolution kernel to obtain the target convolution results corresponding to each of the target matrix data groups. The convolution operation device is composed of a plurality of multipliers and accumulators. Refer to Figure 3 , Figure 3 shows the architecture diagram of the convolution operation device, as Figure 3As shown, the convolution operation device includes 8 multipliers and 8 accumulators. The 8 multipliers are used to calculate 8 data simultaneously. It can be understood that the number of multipliers and accumulators can be user-defined settings, and the number of multipliers is equal to the number of accumulators, and the accumulators correspond to the multipliers one by one.
[0083] Optionally, the matrix data to be convolved includes several matrix elements. The matrix data to be convolved can be voice matrix data, text matrix data, image matrix data, etc. The voice matrix data can be obtained by encoding voice information into the matrix space. The above-mentioned text matrix data can be obtained by encoding text information into the matrix space. The above-mentioned image matrix data can be the pixel matrix of the image itself, or can be obtained by encoding the pixel matrix of the image itself into the matrix space.
[0084] Optionally, the data parallelism is determined according to the hardware resources. Optionally, the data parallelism is obtained according to the number of multipliers of the convolution operation device. When the number of multipliers is k, the data parallelism is k. Optionally, the data parallelism can be less than k.
[0085] Optionally, in the embodiments of the present application, taking the data parallelism of 8 as an example for analysis.
[0086] Optionally, when determining a plurality of target matrix data groups corresponding to the data parallelism according to the data parallelism, each target matrix data group includes 8 sub-matrices. The sub-matrices are determined according to the matrix data to be convolved and the size of the convolution kernel. The number of target matrix data groups is determined by the data volume of the matrix data to be convolved.
[0087] Optionally, referring to Figure 4 , the step S20 includes:
[0088] Step S21, determining a plurality of sub-matrices corresponding to the size in the matrix data to be convolved according to the size of the convolution kernel;
[0089] Step S22, dividing the sub-matrices into several target matrix data groups according to the data parallelism, wherein the number of matrices of the sub-matrices in the target matrix data group is equal to the data parallelism.
[0090] Optionally, the size of the convolution kernel can be 3*3, or can be 5*5. The size of the convolution kernel can be user-defined settings. In the embodiments of the present application, taking the convolution kernel size of 3*3 as an example for analysis.
[0091] Optionally, the size of the sub-matrix is equal to the size of the convolution kernel. When the size of the convolution kernel is 3*3, the size of the sub-matrix is 3*3. The specific implementation of determining multiple sub-matrices corresponding to the matrix data to be convolved according to the size of the convolution kernel is as follows: Starting from the leftmost side of the first row data of the matrix data to be convolved, extract the first 3*3 sub-matrix, move it to the right according to a preset step length, and extract the second 3*3 sub-matrix. After the first row data is completely extracted, start from the leftmost side of the second row data of the data to be convolved and extract sub-matrices in sequence until all matrix elements in the matrix data to be convolved are extracted. Refer to Figure 5 , Figure 5 shows an example diagram of determining multiple sub-matrices corresponding to the matrix data to be convolved.
[0092] It can be understood that the number of sub-matrices is related to the size of the matrix data to be convolved and the convolution kernel. For example, when the size of the matrix data to be convolved is 6*6 and the size of the convolution kernel is 3*3, the number of sub-matrices is 4*4 = 16.
[0093] Optionally, after obtaining multiple sub-matrices, divide the sub-matrices into several target matrix data groups according to the data parallelism. Each target matrix data group includes multiple sub-matrices, and the number of sub-matrices in each target matrix data group is equal to the data parallelism. For example, when the data parallelism is 8 and the sub-matrices are respectively "DW01, DW02, DW03...DW16", that is, when the number of sub-matrices is 16, 16 / 8 = 2 target matrix data groups can be obtained. The first target matrix data group includes "QW01, DW02, DW03...DW08", and the second target matrix data group includes "QW09, DW10, DW11...DW16". Refer to Figure 6 , Figure 6 shows an example diagram of obtaining the target matrix data group.
[0094] Optionally, after obtaining each target matrix data group, perform parallel convolution operations on each target matrix data group and the convolution kernel respectively to obtain the target convolution results corresponding to each target matrix data group.
[0095] It can be understood that when the matrix data to be convolved includes at least one target matrix data group, for example, when the target matrix data group includes the first target matrix data group, the second target matrix data group to the Nth target matrix data group, the target convolution results include the target convolution results corresponding to each target matrix data group respectively.
[0096] Optionally, refer to Figure 7 , the S30 includes:
[0097] Step S31: Perform parallel convolution operations on each first sub - matrix of the first target matrix data group with the convolution kernel to obtain a first target convolution result group corresponding to the first target matrix data group;
[0098] Step S32: After obtaining the first target convolution result group, perform parallel convolution operations on each sub - matrix of the second target matrix data group with the convolution kernel to obtain a second target convolution result group corresponding to the second target matrix data group;
[0099] Step S33: Perform parallel convolution operations in sequence to obtain an N - th target convolution result group corresponding to the N - th target matrix data group;
[0100] Step S34: Determine the target convolution result based on the first target convolution result group, the second target convolution result group, up to the N - th target convolution result group.
[0101] Optionally, when the target matrix data group includes the first target matrix data group, the second target matrix data group, up to the N - th target matrix data group, N is determined according to the size of the matrix data to be convolved, the size of the convolution kernel, and the data parallelism. For example, when the size of the matrix data to be convolved is 6*6, the size of the convolution kernel is 3*3, and the data parallelism is 8, the number of sub - matrices is 4*4 = 16, and at this time N is 16 / 8 = 2.
[0102] Optionally, when the data parallelism is k, each target matrix data group includes k sub - matrices. After obtaining the first target matrix data group, the first target matrix data group includes k first sub - matrices. Perform convolution operations on the k first sub - matrices in the first target matrix data group with the convolution kernel simultaneously to obtain a first target convolution result group corresponding to the first target matrix data group.
[0103] Optionally, after obtaining the first target convolution result group, extract the second target matrix data group. The second target matrix data group includes k second sub - matrices. Perform convolution operations on the k second sub - matrices with the convolution kernel simultaneously to obtain a second target convolution result group corresponding to the second target matrix data group.
[0104] Optionally, and so on, sequentially retrieve the target matrix data group and perform parallel convolution operations on the target matrix data group to obtain an N - th target convolution result group corresponding to the N - th target matrix data group.
[0105] It can be understood that the first target convolution result group includes the convolution results respectively corresponding to each sub-matrix in the first target matrix data group, the second target convolution result group includes the convolution results respectively corresponding to each sub-matrix in the second target matrix data group, and the Nth target convolution result group includes the convolution results respectively corresponding to each sub-matrix in the Nth target matrix data group. That is, the first target convolution result group includes the convolution results respectively corresponding to k first sub-matrices, the second target convolution result group includes the convolution results respectively corresponding to k second sub-matrices, and the Nth target convolution result group includes the convolution results respectively corresponding to k Nth sub-matrices.
[0106] Optionally, after obtaining the first target convolution result group, the second target convolution result group to the Nth target convolution result group, the first target convolution group, the second target convolution result group to the Nth target convolution result group are determined as the target convolution results.
[0107] Optionally, after obtaining the target convolution results, the target convolution matrix data corresponding to the matrix data to be convolved is determined according to the target convolution results. Specifically: the convolution results corresponding to each sub-matrix in the matrix data to be convolved are determined according to the target convolution results, and the target convolution matrix data is determined according to the convolution results corresponding to each sub-matrix.
[0108] Optionally, the target convolution results include the convolution results corresponding to each sub-matrix. After obtaining the target convolution results, the convolution results corresponding to each sub-matrix are directly determined according to the target convolution results, and then the target convolution matrix data is determined according to the convolution results corresponding to each sub-matrix.
[0109] In an embodiment of the present application, sub-matrix data corresponding to the matrix data to be convolved is determined according to the size of the convolution kernel, and then the sub-matrix data is divided into several target matrix data groups according to the data parallelism. The target matrix data groups include a first target matrix data group, a second target matrix data group to an Nth target matrix data group. The number of matrices of the sub-matrices in each target matrix data group is equal to the data parallelism. Then, convolution operations are first simultaneously performed on each first sub-matrix of the first target matrix data group and the convolution kernel, and then convolution operations are simultaneously performed on each second sub-matrix of the second target matrix data group and the convolution kernel, and then convolution operations are sequentially performed on each sub-matrix in each target matrix data group and the convolution kernel to obtain target convolution results corresponding to each target matrix data group. Then, convolution results corresponding to each sub-matrix in the matrix data to be convolved are determined according to the target convolution results, and then target convolution matrix data corresponding to the matrix data to be convolved is determined according to the convolution results corresponding to each sub-matrix. In the embodiment of the present application, by performing parallel convolution operations on each target matrix data group and the convolution kernel, the utilization rate of the parallelism in the channel direction is improved, and then the resource utilization rate of the multiplier is improved, and the convolution operation efficiency is also improved.
[0110] Optionally, referring to Figure 8 , based on the first embodiment, step S31 includes:
[0111] Step S311, within one time period, determine the first group of data corresponding to the first target matrix data group and the first weight coefficient corresponding to the convolution kernel, and simultaneously perform multiplication operations on each first data in the first group of data and the first weight coefficient to obtain a first processing result corresponding to the first group of data. The first group of data includes the first data of each first sub-matrix in the first target matrix data group;
[0112] Step S312, after obtaining the first processing result, extract the second group of data corresponding to the target matrix data group and the second weight coefficient, and simultaneously perform multiplication operations on each second data in the second group of data and the second weight coefficient to obtain a second processing result corresponding to the second group of data. The second group of data includes the second data of each first sub-matrix in the first target matrix data group;
[0113] Step S313, sequentially extract the data of the first target matrix data group and obtain processing results to obtain an Nth processing result;
[0114] Step S314, determine the first target convolution result group according to the first processing result, the second processing result to the Nth processing result.
[0115] In this embodiment, when the size of the convolution kernel is 3*3, the first target matrix data group includes k first sub-matrices, and the size of each first sub-matrix is 3*3. The first group of data includes the first data corresponding to each of the k first sub-matrices, the second group of data includes the second data corresponding to each of the k first sub-matrices, and so on. The Nth group of data includes the Nth data corresponding to the k first sub-matrices. When the size of the first sub-matrix is 3*3, N = 9.
[0116] Optionally, the first weight coefficient is the first weight corresponding to the convolution kernel, the second weight coefficient is the second weight corresponding to the convolution kernel, and so on. The Nth weight coefficient is the Nth weight corresponding to the convolution kernel. When the size of the convolution kernel is 3*3, N = 9.
[0117] Optionally, within one time period, multiply each of the first data in the first group of data by the first weight coefficient simultaneously to obtain the product corresponding to each of the first data, and then determine the first processing result based on the products corresponding to each of the first data.
[0118] Optionally, during actual operation, within one time period, the specific implementation of multiplying each of the first data in the first group of data by the first weight coefficient simultaneously is as follows: input each of the first data into the corresponding multiplier, and input the first weight coefficient into each of the multipliers simultaneously, so that each multiplier can perform a product operation on the input first data and the first weight coefficient respectively to obtain the product corresponding to each of the first data. Refer to Figure 3 , the depthwise fast convolution operation system includes multipliers. When the depthwise fast convolution operation system includes 8 multipliers, the first group of data includes 8 first data, and the first processing result includes 8 products.
[0119] Optionally, after obtaining the first processing result, in the next time period, extract the second group of data and the second weight coefficient of the convolution kernel, input each of the second data corresponding to the second group of data into the corresponding multiplier, and input the second weight coefficient into each of the multipliers simultaneously, so that each multiplier can perform a product operation on the input second data and the second weight coefficient respectively to obtain the product corresponding to each of the second data.
[0120] By analogy, the third set of data, the fourth set of data, and the ninth set of data are extracted in sequence, and then the third set of data and the third weight coefficient are input into each multiplier in sequence, the fourth set of data and the fourth weight coefficient are input into each multiplier, until the ninth set of data and the ninth weight coefficient are input into each multiplier to obtain the third processing result, the fourth processing result to the ninth processing result.
[0121] Optionally, as Figure 3 shown, the depthwise fast convolution operation system further includes an accumulator, and the step of determining the first target convolution result group according to the first processing result, the second processing result to the Nth processing result includes:
[0122] After obtaining the first processing result and the second processing result, the second processing result is respectively superimposed into the first processing result to update the first processing result;
[0123] After obtaining the third processing result, the third processing result is respectively superimposed into the updated first processing result to update the first processing result again;
[0124] Update the first processing result in sequence;
[0125] The first processing result after the last update is determined as the first target convolution result group.
[0126] Optionally, after obtaining the first processing result and the second processing result, the first processing result includes the product of the first data corresponding to each of the first sub-matrices, and the second processing result includes the product of the second data corresponding to each of the first sub-matrices. The second processing result is respectively superimposed into the first processing result. For example, when there are 8 multipliers, the first processing result includes "Q011, Q021, Q031...Q081", and the second processing result includes "Q012, Q022, Q032...Q082". The second processing result is respectively superimposed into the first processing result as "Q012+Q011, Q022+Q021, Q032+Q031...Q08 2+ Q081".
[0127] Optionally, within a time period, after obtaining the first processing result through the multiplier, the first processing result is respectively input into corresponding accumulators for the accumulators to store the corresponding first processing results; after the next time period, after obtaining the second processing result through the multiplier, the second processing result is respectively input into the corresponding accumulators for the accumulators to accumulate the second processing result to the first processing result according to the second processing result to update the first processing result, and store the updated first processing result in the accumulators.
[0128] Optionally, in the next time period, after obtaining the third processing result through the multiplier, the third processing result is respectively superimposed on the updated first processing result to update the first processing result again. For example, the updated first processing result is "Q012 + Q011, Q022 + Q021, Q032 + Q031... Q08 2+ Q081", the third processing result is "Q013, Q023, Q033... Q083", and superimposing the third processing result on the updated first processing result respectively gives "Q013 + Q012 + Q011, Q023 + Q022 + Q021, Q033 + Q032 + Q031... Q083 + Q08 2+ Q081".
[0129] Optionally, within each time period, the corresponding processing results are sequentially obtained, and then the processing results are sequentially superimposed to sequentially update the first processing result.
[0130] Optionally, after superimposing all the processing results, the last updated first processing result is obtained, and then the last updated first processing result is determined as the first target convolution result group, and the first target convolution result group includes the convolution results corresponding to each of the first matrices.
[0131] Optionally, after obtaining the first target convolution result group, the second target convolution result group to the Nth target convolution result group are sequentially obtained in the same manner as obtaining the first target convolution result group.
[0132] Optionally, in each time period, 8 multipliers are simultaneously called to perform a multiplication operation on the currently extracted data and the current weight coefficients of the convolution kernel to obtain the current processing result, and then the current processing result is sent to the corresponding accumulator for the accumulator to superimpose the current processing result on the previous processing result.
[0133] In an embodiment of the present application, within one time period, the first data corresponding to each of the first matrices and the first weight coefficients of the convolution kernels are extracted. The multiplier performs parallel multiplication operations on the first data and the first weight coefficients to obtain a first processing result, and the first processing result is sent to the corresponding accumulator. Then, in the next time period, the second data and the second weight coefficients are obtained, and the multiplier performs parallel multiplication operations on the second data and the second weight coefficients to obtain a second processing result. Then, the second processing result is sent to the accumulator for the accumulator to superimpose the second processing result on the first processing result. And so on, in each subsequent time period, the processing results are sequentially obtained and sequentially superimposed to obtain the first processing result updated for the last time. Then, the first processing result is determined as the first target convolution result group of the first target matrix data group. By using multiple multipliers and multiple accumulators, the present application embodiment can perform parallel convolution operations on the target matrix data group within the same time period, improving the resource utilization rate of the multiplier and further improving the operation efficiency of the depth convolution operation.
[0134] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a depthwise fast convolution operation program is stored. When the depthwise fast convolution operation program is executed by a processor, the steps of the above-described various embodiments are implemented.
[0135] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0136] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0138] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A depthwise fast convolution operation method, characterized in that, The steps of the depthwise fast convolution operation method include: Obtain the data of the matrix to be convolved; Determine multiple sub-matrices corresponding to the size in the data of the matrix to be convolved according to the size of the convolution kernel. Wherein, the size of the convolution kernel is set by the user, and the size of the sub-matrix is equal to the size of the convolution kernel, and the number of sub-matrices is related to the data of the matrix to be convolved and the size of the convolution kernel; Divide the sub-matrices into several groups of target matrix data according to the data parallelism. Wherein, the number of sub-matrices in each group of target matrix data is equal to the data parallelism; Successively perform parallel convolution operations on each group of target matrix data and the convolution kernel to obtain the target convolution results corresponding to each group of target matrix data; Determine the target convolution matrix data corresponding to the data of the matrix to be convolved according to the target convolution results; Wherein, the groups of target matrix data include the first group of target matrix data, the second group of target matrix data to the Nth group of target matrix data. The step of successively performing parallel convolution operations on each group of target matrix data and the convolution kernel to obtain the target convolution results corresponding to each group of target matrix data includes: Perform parallel convolution operations on each first sub-matrix of the first group of target matrix data and the convolution kernel to obtain the first group of target convolution results corresponding to the first group of target matrix data; After obtaining the first group of target convolution results, perform parallel convolution operations on each second sub-matrix of the second group of target matrix data and the convolution kernel to obtain the second group of target convolution results corresponding to the second group of target matrix data; Successively perform parallel convolution operations to obtain the Nth group of target convolution results corresponding to the Nth group of target matrix data; Determine the target convolution results according to the first group of target convolution results, the second group of target convolution results to the Nth group of target convolution results.
2. The depthwise fast convolution operation method according to claim 1, wherein The first group of target convolution results includes the convolution results corresponding to each first sub-matrix in the first group of target matrix data, the second group of target convolution results includes the convolution results corresponding to each second sub-matrix in the second group of target matrix data, and the Nth group of target convolution results includes the convolution results corresponding to each Nth sub-matrix in the Nth group of target matrix data.
3. The depthwise fast convolution operation method according to claim 1, wherein The step of performing parallel convolution operations on each first sub-matrix of the first group of target matrix data and the convolution kernel to obtain the first group of target convolution results corresponding to the first group of target matrix data includes: Within one time period, determine the first group of data corresponding to the first group of target matrix data and the first weight coefficients corresponding to the convolution kernel, and simultaneously perform multiplication operations on each first data in the first group of data and the first weight coefficients to obtain the first processing results corresponding to the first group of data. The first group of data includes the first data of each first sub-matrix in the first group of target matrix data; After obtaining the first processing result, extract the second set of data corresponding to the target matrix data group and the second weight coefficient, and perform a multiplication operation on each second data in the second set of data with the second weight coefficient simultaneously to obtain the second processing result corresponding to the second set of data, where the second set of data includes the second data of each first sub-matrix in the first target matrix data group; Extract the data of the target matrix data group and obtain the processing results in sequence to obtain the Nth processing result; Determine the first target convolution result group according to the first processing result, the second processing result to the Nth processing result.
4. The depthwise fast convolution operation method according to claim 3, wherein The step of determining the first target convolution result according to the first processing result, the second processing result to the Nth processing result includes: After obtaining the first processing result and the second processing result, superimpose the second processing result into the first processing result respectively to update the first processing result; After obtaining the third processing result, superimpose the third processing result into the updated first processing result respectively to update the first processing result again; Update the first processing result in sequence; Determine the first processing result after the last update as the first target convolution result group.
5. The depthwise fast convolution operation method according to any one of claims 1-4, characterized in that The step of determining the target convolution matrix data corresponding to the matrix to be convolved according to the target convolution result includes: Determine the convolution result corresponding to each sub-matrix in the matrix data to be convolved according to the target convolution result; Determine the target convolution matrix data according to the convolution result corresponding to each sub-matrix.
6. A depthwise fast convolution operation system, characterized in that The depthwise fast convolution operation system includes: a memory, a processor, and a depthwise fast convolution operation program stored on the memory and executable on the processor. When the depthwise fast convolution operation program is executed by the processor, it implements the steps of the depthwise fast convolution operation method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, A depthwise fast convolution operation program is stored on the computer-readable storage medium. When the depthwise fast convolution operation program is executed by the processor, it implements the steps of the depthwise fast convolution operation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Lightweight neural network hardware accelerator based on depth separable convolution
CN113033794A
Mapping convolution to a channel convolution engine
CN113344172A