Method, apparatus, electronic device, and storage medium for processing feature maps
By obtaining the computing power information of the feature map and the weight matrix and reconstructing the multiplication mode between them, the repeated reading and repeated calculation problems in convolutional calculations in deep learning are solved, and more efficient calculations are achieved.
Patent Information
- Application Number
- CN202111661159.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-30
AI Technical Summary
Convolutional calculations have problems with repeated reading and repeated calculations in deep learning, resulting in waste of computing power and low computing efficiency.
By obtaining the computing power information of the feature map to be processed and the weight matrix, the multiplication mode between them is reconstructed from the computing power level, and the computing mode is dynamically adjusted to make full use of the hardware computing power to avoid repeated operations.
Given the computing power, the effective computing power of the hardware is fully released, the computing time is reduced, and the computing efficiency is improved.
Smart Images

Figure CN114444656B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a method, apparatus, electronic device, and storage medium for processing feature maps. Background Art
[0002] Deep learning is a learning structure of a multi-layer perceptron with multiple hidden layers. Deep learning can form more abstract high-level representations of attribute categories or features by combining low-level features to discover the distributed feature representations of data. Convolution and matrix calculations are the most important operations in current deep learning. These calculations have enabled neural network models based on deep learning to achieve rapid development in the fields of image recognition, speech recognition, natural language processing, intelligent game challenges, etc. The principle of convolution calculation is mainly to obtain the convolution result by sliding a convolution weight (in matrix form) over a feature map. Specifically, the implementation of convolution calculation is generally to expand the weight matrix and the input feature matrix into two two-dimensional matrices, and then perform matrix multiplication on the two-dimensional matrices to obtain the final convolution result. Such a calculation method will have problems of repeated reading and repeated calculation, resulting in waste of computing power and low computing efficiency. Summary of the Invention
[0003] An embodiment of the present invention provides a method for processing a feature map. By obtaining the computing power information of the calculation area in the to-be-processed feature map and the computing power information of the weight matrix, the multiplication mode between the to-be-processed feature map and the weight matrix is reconstructed from the computing power level. Since this multiplication mode is reconstructed according to the computing power of the calculation area and the weight matrix, the effective computing power of the hardware can be fully released under a given computing power, avoiding repeated reading and repeated calculation, thereby reducing the calculation time and improving the calculation efficiency.
[0004] In a first aspect, an embodiment of the present invention provides a method for processing a feature map, the method including:
[0005] Obtain the computing power information of the weight matrix, where the computing power information of the weight matrix includes the first-row computing power information and the first-column computing power information;
[0006] Determine the calculation area corresponding to the weight matrix in the to-be-processed feature map, and obtain the computing power information of the calculation area, where the computing power information of the calculation area includes the second-row computing power information and the second-column computing power information;
[0007] Dynamically reconstruct the multiplication mode of the calculation area and the weight matrix according to the first-row computing power information, the first-column computing power information, the second-row computing power information, the second-column computing power information, and a preset multiplication calculation unit;
[0008] Perform multiplication operation processing on the to-be-processed feature map and the weight matrix through the reconstructed multiplication mode.
[0009] Optionally, before dynamically reconstructing the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit, the method further includes:
[0010] Obtaining basic calculation unit information under the current hardware conditions;
[0011] Determining a first multiplication matrix and a second multiplication matrix according to the basic calculation unit information, where the first multiplication matrix includes third row computing power information and third column computing power information, the second multiplication matrix includes fourth row computing power information and fourth column computing power information, and the third column computing power information is equal to the fourth row computing power information;
[0012] Constructing the multiplication calculation unit for performing the matrix multiplication of the first multiplication matrix and the second multiplication matrix.
[0013] Optionally, the dynamically reconstructing the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit includes:
[0014] Calculating a first number of loops according to the first row computing power information and the third row computing power information;
[0015] Calculating a second number of loops according to the first column computing power information and the third column computing power information, or according to the second row computing power information and the fourth row computing power information;
[0016] Calculating a third number of loops according to the second column computing power information, the fourth column computing power information, and the column tiling number;
[0017] Determining a first multiplication pattern of the calculation region and the weight matrix according to the first number of loops, the second number of loops, and the third number of loops;
[0018] Reconstructing the first multiplication pattern through the multiplication calculation unit.
[0019] Optionally, the dynamically reconstructing the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit includes:
[0020] Calculating a fourth number of loops according to the first row computing power information, the third row computing power information, and the row tiling number;
[0021] Calculate the fifth number of cycles based on the computing power information in the second column and the computing power information in the fourth column;
[0022] Determine the second multiplication pattern of the calculation area and the weight matrix according to the fourth number of cycles, the second number of cycles, and the fifth number of cycles;
[0023] Reconstruct the second multiplication pattern through the multiplication calculation unit.
[0024] Optionally, the first multiplication pattern includes the first total number of cycles, the second multiplication pattern includes the second total number of cycles, and the dynamic reconstruction of the multiplication pattern of the calculation area and the weight matrix according to the computing power information in the first row, the computing power information in the first column, the computing power information in the second row, the computing power information in the second column, and a preset multiplication calculation unit further includes:
[0025] If the first total number of cycles is less than the second total number of cycles, determine that the multiplication pattern of the calculation area and the weight matrix is the first multiplication pattern;
[0026] If the first total number of cycles is greater than the second total number of cycles, determine that the multiplication pattern of the calculation area and the weight matrix is the second multiplication pattern;
[0027] If the first total number of cycles is equal to the second total number of cycles, determine that the multiplication pattern of the calculation area and the weight matrix is randomly one of the first multiplication pattern and the second multiplication pattern.
[0028] Optionally, the dynamic reconstruction of the multiplication pattern of the calculation area and the weight matrix according to the computing power information in the first row, the computing power information in the first column, the computing power information in the second row, the computing power information in the second column, and a preset multiplication calculation unit further includes:
[0029] When batch processing information is detected, determine the third multiplication pattern of the calculation area and the weight matrix according to the first number of cycles, the second number of cycles, and the fifth number of cycles;
[0030] Reconstruct the third multiplication pattern through the multiplication calculation unit.
[0031] Optionally, the batch processing information includes the first batch processing information of the weight matrix and / or the second batch processing information of the feature map to be processed. Determining the third multiplication pattern of the calculation area and the weight matrix according to the first number of cycles, the second number of cycles, and the fifth number of cycles includes:
[0032] Determine the third multiplication pattern between the calculation region and the weight matrix according to the first batch of processing information and / or the second batch of processing information, in combination with the first number of cycles, the second number of cycles, and the fifth number of cycles.
[0033] In a second aspect, an embodiment of the present invention provides a processing device for a feature map. The device includes:
[0034] A first acquisition module, configured to acquire the computing power information of a weight matrix, where the computing power information of the weight matrix includes first row computing power information and first column computing power information;
[0035] A second acquisition module, configured to determine a calculation region corresponding to the weight matrix in the to-be-processed feature map, and acquire the computing power information of the calculation region, where the computing power information of the calculation region includes second row computing power information and second column computing power information;
[0036] A reconstruction module, configured to reconstruct the multiplication pattern between the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, and the second column computing power information;
[0037] A processing module, configured to perform a multiplication operation on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps in the method for processing a feature map provided by the embodiment of the present invention are implemented.
[0039] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for processing a feature map provided by the embodiment of the invention are implemented.
[0040] In an embodiment of the present invention, computing power information of a weight matrix is obtained. The computing power information of the weight matrix includes first-row computing power information and first-column computing power information. A calculation area corresponding to the weight matrix in a to-be-processed feature map is determined, and computing power information of the calculation area is obtained. The computing power information of the calculation area includes second-row computing power information and second-column computing power information. According to the first-row computing power information, the first-column computing power information, the second-row computing power information, the second-column computing power information, and a preset multiplication calculation unit, a multiplication pattern between the calculation area and the weight matrix is dynamically reconstructed. A multiplication operation is performed on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern. By obtaining the computing power information of the to-be-processed feature map and the computing power information of the weight matrix, and reconstructing the multiplication pattern between the to-be-processed feature map and the weight matrix from the computing power level, since the multiplication pattern is reconstructed according to the computing power conditions of the to-be-processed feature map and the weight matrix, the effective computing power of the hardware can be fully released under a given computing power, avoiding the situations of repeated reading and repeated calculation, thereby reducing the calculation time and improving the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 is a flowchart of a method for processing a feature map provided by an embodiment of the present invention;
[0043] Figure 1a is a schematic diagram of a multiplication calculation unit provided by an embodiment of the present invention;
[0044] Figure 2 is Figure 1 an implementation flowchart of step 103 in the embodiment;
[0045] Figure 2a is a single-loop schematic diagram of a first multiplication pattern provided by an embodiment of the present invention;
[0046] Figure 2b is a multi-loop schematic diagram of a first multiplication pattern provided by an embodiment of the present invention;
[0047] Figure 3 is Figure 1 another implementation flowchart of step 103 in the embodiment;
[0048] Figure 3aIt is a single - loop schematic diagram of a second multiplication mode provided by an embodiment of the present invention;
[0049] Figure 3b It is a multi - loop schematic diagram of a second multiplication mode provided by an embodiment of the present invention;
[0050] Figure 3c It is a single - loop schematic diagram of a third multiplication mode provided by an embodiment of the present invention;
[0051] Figure 3d It is a multi - loop schematic diagram of a second multiplication mode provided by an embodiment of the present invention;
[0052] Figure 4 It is a structural schematic diagram of a processing device for a feature map provided by an embodiment of the present invention;
[0053] Figure 5 It is a structural schematic diagram of another processing device for a feature map provided by an embodiment of the present invention;
[0054] Figure 6 It is a structural schematic diagram of a reconstruction module provided by an embodiment of the present invention;
[0055] Figure 7 It is a structural schematic diagram of another reconstruction module provided by an embodiment of the present invention;
[0056] Figure 8 It is a structural schematic diagram of another reconstruction module provided by an embodiment of the present invention;
[0057] Figure 9 It is a structural schematic diagram of another reconstruction module provided by an embodiment of the present invention;
[0058] Figure 10 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.
[0060] Please refer to Figure 1 , Figure 1 It is a flowchart of a method for processing a feature map provided by an embodiment of the present invention. As Figure 1 shown, the method for processing the feature map includes the following steps:
[0061] 101. Obtain the computing power information of the weight matrix.
[0062] In an embodiment of the present invention, the computing power information of the weight matrix includes the computing power information of the first row and the computing power information of the first column. The above weight matrix may be the weight matrix of each computing layer in a deep neural network, and the above weight matrix may also be referred to as a weight parameter. The above weight matrix may be a trained weight matrix. The weight matrix includes n1 rows and m1 columns, and the weight matrix may generally include n1*m1 matrix units. In an embodiment of the present invention, the computing power information of the first row is denoted as a_h, and a_h represents the tiled computing power of the n1 rows in the weight matrix. The computing power information of the first column is denoted as a_w, and a_w represents the tiled computing power of the m1 columns in the weight matrix.
[0063] 102. Determine the computing area corresponding to the weight matrix in the to-be-processed feature map, and obtain the computing power information of the computing area.
[0064] In an embodiment of the present invention, the to-be-processed feature map may be the feature map of the input of each computing layer in a deep neural network, and the to-be-processed feature map may also be referred to as a to-be-processed feature matrix. The to-be-processed feature map includes n2 rows and m2 columns, and the to-be-processed feature map generally includes n2*m2 matrix units.
[0065] It should be noted that the weight matrix slides in the to-be-processed feature map for convolution calculation. The convolution calculation mainly includes the matrix multiplication operation between the weight matrix and the computing area. Therefore, the above computing area is also a matrix. Further, the above computing area is also a matrix including n3 rows and m3 columns, and the to-be-processed feature map generally includes n3*m3 matrix units. Among them, n3 is less than or equal to n2, and m3 is less than or equal to m2. When the weight matrix has the same number of rows and columns as the to-be-processed feature map, the computing area is the entire to-be-processed feature map. At this time, n1 = n2 = n3, and m1 = m2 = m3.
[0066] It should also be noted that in the matrix multiplication operation between the weight matrix n1*m1 and the computing area n3*m3, it is necessary to satisfy m1 = n3. When m1 = n3, the weight matrix (n1*m1) is left-multiplied by the computing area (n3*m3) to obtain the result matrix (n1*m3).
[0067] The above calculation region can be determined according to the above weight matrix, so that the above calculation region and the above weight matrix can perform matrix multiplication. When the weight matrix is multiplied by the calculation region on the left, the number of rows of the calculation region is equal to the number of columns of the weight matrix. When the calculation region is multiplied by the weight matrix on the left, the number of rows of the weight matrix is equal to the number of columns of the weight matrix. The computing power information of the above calculation region includes the second row computing power information and the second column computing power information. In the embodiment of the present invention, the above second row computing power information is denoted as b_h, and b_h represents the tiled computing power of the n3 rows in the calculation region. The above second column computing power information is denoted as b_w, and b_w represents the tiled computing power of the m3 columns in the calculation region.
[0068] 103. Dynamically reconstruct the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit.
[0069] In the embodiment of the present invention, taking the weight matrix multiplied by the calculation region on the left as an example, the first row computing power information is denoted as a_h, the first column computing power information is denoted as a_w, the second row computing power information is denoted as b_h, and the second column computing power information is denoted as b_w. Among them, the first column computing power information a_w is equal to the second row computing power information b_h.
[0070] The multiplication pattern of the above to-be-processed feature image and the weight matrix can be determined by the matrix multiplication pattern of the calculation region and the weight matrix, and the above to-be-processed feature image and the above weight matrix are multiplied through the multiplication pattern. Further, the multiplication pattern of the above to-be-processed feature image and the weight matrix can be understood as the matrix multiplication pattern of the above calculation region and the weight matrix.
[0071] In the embodiment of the present invention, the above matrix multiplication is performed based on a multiplication calculation unit in a hardware environment. The multiplication calculation unit can implement the matrix multiplication of two basic multiplication matrices. By tiling the multiplication calculation unit, the matrix multiplication pattern of the above calculation region and the weight matrix is dynamically constructed, and then the multiplication pattern of the above to-be-processed feature image and the weight matrix is determined according to the matrix multiplication pattern of the above calculation region and the weight matrix.
[0072] 104. Perform multiplication operation processing on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern.
[0073] In the embodiment of the present invention, the above to-be-processed feature map can be split into at least one calculation region according to the weight matrix. Through the computing power information of the calculation region and the computing power information of the weight matrix, the to-be-processed feature map and the weight matrix are reconstructed, and the multiplication calculation unit of the hardware executes the corresponding multiplication pattern, thereby completing the multiplication operation processing on the to-be-processed feature map and the weight matrix.
[0074] In an embodiment of the present invention, computing power information of a weight matrix is obtained. The computing power information of the weight matrix includes first row computing power information and first column computing power information. A calculation area corresponding to the weight matrix in a to-be-processed feature map is determined, and computing power information of the calculation area is obtained. The computing power information of the calculation area includes second row computing power information and second column computing power information. According to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit, dynamic reconstruction is performed on a multiplication pattern of the calculation area and the weight matrix. A multiplication operation is performed on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern. By obtaining the computing power information of the to-be-processed feature map and the computing power information of the weight matrix, the multiplication pattern between the to-be-processed feature map and the weight matrix is reconstructed at the computing power level. Since the multiplication pattern is reconstructed according to the computing power conditions of the to-be-processed feature map and the weight matrix, the effective computing power of the hardware can be fully released under a given computing power, avoiding situations of repeated reading and repeated calculation, thereby reducing the calculation time and improving the calculation efficiency.
[0075] Optionally, before step 103, the method for processing a feature map according to an embodiment of the present invention may further include: obtaining basic computing unit information under current hardware conditions; determining a first multiplication matrix and a second multiplication matrix according to the basic computing unit information, where the first multiplication matrix includes third row computing power information and third column computing power information, the second multiplication matrix includes fourth row computing power information and fourth column computing power information, and the third column computing power information is equal to the fourth row computing power information; constructing a multiplication calculation unit for performing matrix multiplication of the first multiplication matrix and the second multiplication matrix.
[0076] Further, the above-mentioned hardware can be any processing chip, such as a CPU (Central Processing Unit / Processor), a GPU (Graphics Processing Unit), an FPGA (Field-Programmable Gate Array), and other self-developed chips for computing and processing. When different processing chips are designed and manufactured, there are corresponding basic computing units, and one basic computing unit is used to process a basic calculation. When performing matrix multiplication, one multiplication computing unit is used to process a basic matrix multiplication calculation. The basic computing unit information includes the parameters of the matrix multiplication that the multiplication computing unit can handle. Further, it includes the first multiplication matrix parameters and the second multiplication matrix parameters that a multiplication computing unit can handle. Through the above first multiplication matrix parameters and second multiplication matrix parameters, the first multiplication matrix and the second multiplication matrix can be determined. The above first multiplication matrix includes the third row computing power information and the third column computing power information, the above second multiplication matrix includes the fourth row computing power information and the fourth column computing power information, and the above third column computing power information is equal to the above fourth row computing power information so that the above first multiplication matrix is multiplied by the second multiplication matrix on the left. Of course, in a possible embodiment, the above fourth column computing power information can also be equal to the above third row computing power information. At this time, the above second multiplication matrix can be multiplied by the first multiplication matrix on the left.
[0077] Denote the above third row computing power information as alpha, the above third column computing power information as beta, the above fourth row computing power information as beta, and the above fourth column computing power information as gamma. Then, a corresponding multiplication computing unit can be constructed on the current hardware according to the above third row computing power information alpha, the above third column computing power information beta, the above fourth row computing power information beta, and the above fourth column computing power information gamma. The constructed multiplication computing unit can perform the multiplication calculation of [alpha, beta]×[beta, gamma]. It can be understood that the computing power of the above multiplication computing unit is alpha*beta*gamma multiply-accumulate units. Specifically, as shown in Figure 1a shown. Figure 1a is a schematic diagram of a multiplication computing unit provided by an embodiment of the present invention. In Figure 1a shown, the above multiplication computing unit performs matrix multiplication of the first multiplication matrix A0 and the second multiplication matrix B0. The matrix multiplication of the weight matrix and the matrix in the calculation area can be split into multiple multiplication computing units to perform multiplication calculations, so as to complete the matrix multiplication of the weight matrix and the matrix in the calculation area, and further complete the multiplication operation processing of the to-be-processed feature map and the weight matrix.
[0078] In the embodiments of the present invention, by using the basic computing unit information under the current hardware conditions, a corresponding multiplication computing unit is constructed, and the multiplication mode of the computing area and the weight matrix can be dynamically reconstructed flexibly through the multiplication computing unit.
[0079] Optionally, in the embodiments of the present invention, the above multiplication mode may include at least one of three multiplication modes, and the above three multiplication modes can be reconstructed through the above multiplication computing unit. The above three multiplication modes specifically include a first multiplication mode, a second multiplication mode, and a third multiplication mode. Specifically, the required multiplication mode can be determined by using the above first row computing power information, first column computing power information, second row computing power information, second column computing power information, and a preset multiplication computing unit.
[0080] Further, please refer to Figure 2 , Figure 2 which Figure 1 is Figure 2 a flowchart of an implementation of step 103 in an embodiment. As Figure 1 shown, on the basis of the embodiments, the step of dynamically reconstructing the multiplication mode of the computing area and the weight matrix according to the first row computing power information, first column computing power information, second row computing power information, second column computing power information, and a preset multiplication computing unit specifically includes:
[0081] 201. Calculate the first number of cycles according to the first row computing power information and the third row computing power information.
[0082] In the embodiments of the present invention, after obtaining the first row computing power information a_h and the third row computing power information alpha, the first row computing power information a_h can be divided by the third row computing power information alpha, and the integer greater than or equal to the result value is used as the first number of cycles. Specifically, the first number of cycles cyc_alpha can be calculated by the following formula:
[0083] cyc_alpha = ceil(a_h / alpha)
[0084] In the above formula, ceil(a_h / alpha) represents taking the smallest integer greater than or equal to a_h / alpha. For example, if a_h is 5 and alpha is 4, then cyc_alpha = ceil(5 / 4) = 2.
[0085] 202. Calculate the second number of cycles according to the first column computing power information and the third column computing power information, or according to the second row computing power information and the fourth row computing power information.
[0086] In an embodiment of the present invention, after obtaining the first column computing power information a_w and the third column computing power information beta, the first column computing power information a_w can be divided by the third column computing power information beta, and the integer greater than or equal to the result value is used as the second loop count. Specifically, the second loop count cyc_beta can be calculated by the following formula:
[0087] cyc_beta = ceil(a_w / beta)
[0088] In the above formula, ceil(a_w / beta) represents taking the smallest integer greater than or equal to a_w / beta. For example, if a_w is 5 and beta is 4, then cyc_beta = ceil(5 / 4) = 2.
[0089] It should be noted that since the computing power information of the second row is the same as that of the first column, and the computing power information of the fourth row is the same as that of the third column, the second loop count can also be calculated by the above calculation method using the computing power information of the second row and the fourth row.
[0090] 203. Calculate the third loop count according to the second column computing power information, the fourth column computing power information, and the column tiling quantity.
[0091] In an embodiment of the present invention, the above column tiling quantity can be preset. The above tiling can also be referred to as parallel setting. Denote the above column tiling quantity as num_pe. The above column tiling quantity means that num_pe multiplication calculation units are tiled in one row. After obtaining the second column computing power information b_w and the fourth column computing power information gamma, the second column computing power information b_w can be divided by the total computing power information of the fourth column num_pe * gamma, and the integer greater than or equal to the result value is used as the third loop count. Specifically, the third loop count cyc_num_gamma can be calculated by the following formula:
[0092] cyc_num_gamma = ceil(b_w / (num_pe * gamma))
[0093] In the above formula, ceil(b_w / (num_pe * gamma)) represents taking the smallest integer greater than or equal to b_w / (num_pe * gamma). For example, if a_w is 5, num_pe is 4, and gamma is 4, then cyc_beta = ceil(5 / 16) = 1.
[0094] 204. Determine the first multiplication mode of the calculation area and the weight matrix according to the first loop count, the second loop count, and the third loop count.
[0095] In an embodiment of the present invention, when the first loop count, the second loop count, and the third loop count meet a preset condition, it is determined that the first multiplication mode of the calculation region and the weight matrix is matrix multiplication of n first multiplication matrices and a row of second multiplication matrices, and a row of second multiplication matrices includes num_pe second multiplication matrices.
[0096] Furthermore, the product of the first loop count, the second loop count, and the third loop count can be calculated as the first loop total amount cyc_ti. When the first loop total amount cyc_ti is less than a preset value, it is determined that the first multiplication mode of the calculation region and the weight matrix is matrix multiplication of n first multiplication matrices and a row of second multiplication matrices, and a row of second multiplication matrices includes num_pe second multiplication matrices. Among them, the combination of n first multiplication matrices should meet the rule of matrix multiplication with the combination of num_pe second multiplication matrices, that is, the computing power information of the third column of n is equal to the computing power information of the fourth row of num_pe. Specifically, the above first loop total amount cyc_ti can be calculated by the following formula:
[0097] cyc_ti = cyc_alpha * cyc_beta * cyc_num_gamma
[0098] 205. Reconstruct the first multiplication mode through the multiplication calculation unit.
[0099] In an embodiment of the present invention, after determining that the multiplication mode of the weight matrix and the calculation region is the first multiplication mode, the first multiplication mode is reconstructed under the current hardware conditions through the multiplication calculation unit.
[0100] Furthermore, taking the left multiplication of the weight matrix by the calculation region as an example, one first multiplication matrix is tiled in the weight matrix, and num_pe second multiplication matrices are tiled in the corresponding row of the calculation region. Specifically, reference can be made to Figure 2a , Figure 2a is a single-loop schematic diagram of a first multiplication mode provided by an embodiment of the present invention. In Figure 2a , taking num_pe = 4 as an example, one first multiplication matrix A0 is tiled in the weight matrix A, and 4 second multiplication matrices [B0, B1, B2, B3] are tiled in the corresponding row of the calculation region B. It can be seen that each time the first multiplication mode is executed, its corresponding computing power is alpha * beta * gamma * num_pe multiply-accumulate units. Combining with reference Figure 2b , Figure 2b is a multi-loop schematic diagram of a first multiplication mode provided by an embodiment of the present invention. In Figure 2bWhen cyc is 0, 1, 2, or 3, it represents the first, second, third, and fourth executions of the first multiplication mode respectively. In multiple loops, when cyc is 0, the first multiplication mode is executed for the first time, such that the first multiplication matrix A0 corresponding to the first row and first column of the weight matrix A performs matrix multiplication with the second multiplication matrices [B0, B1, B2, B3] corresponding to the first column to the fourth column of the first row of the calculation region B; when cyc is 1, the second multiplication mode is executed, such that the first multiplication matrix A0 corresponding to the first row and second column of the weight matrix A performs matrix multiplication with the second multiplication matrices [B0, B1, B2, B3] corresponding to the first column to the fourth column of the second row of the calculation region B; and so on. By executing the above first multiplication mode in multiple loops, the matrix multiplication of the weight matrix A and the calculation region B is completed. By completing the matrix multiplication of the weight matrix A and the calculation region B multiple times, the multiplication operation processing of the weight matrix and the feature map to be processed is completed.
[0101] In an embodiment of the present invention, based on the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit, the first multiplication mode of the weight matrix and the calculation region is determined, and the multiplication calculation unit is used to reconstruct the first multiplication mode, such that, given a certain computing power, the effective computing power of the hardware is fully released for the weight matrix and the calculation region, avoiding the situations of repeated reading and repeated calculation, thereby reducing the calculation time and improving the calculation efficiency.
[0102] Optionally, please refer to Figure 3 , Figure 3 is Figure 1 Another implementation flowchart of step 103 in the embodiment, as shown in Figure 3 shown. Based on the embodiments of Figure 1 and Figure 2 the step of dynamically reconstructing the multiplication mode of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit specifically includes:
[0103] 301. Calculate the fourth loop count according to the first row computing power information, the third row computing power information, and the row tiling number.
[0104] In an embodiment of the present invention, the above column tiling quantity can be preset. Denote the above row tiling quantity as num_pe. The above column tiling quantity indicates that num_pe multiplication calculation units are tiled in one column. After obtaining the first row computing power information a_h and the third row computing power information alpha, the first row computing power information a_h can be divided by the total computing power information alpha * num_pe of the third row to obtain an integer greater than or equal to the result value as the first loop count. Specifically, the fourth loop count cyc_alpha can be calculated through the following formula:
[0105] cyc_num_alpha = ceil(a_h / (alpha * num_pe))
[0106] In the above formula, ceil(a_h / (alpha * num_pe)) represents taking the smallest integer greater than or equal to a_h / (alpha * num_pe). For example, if a_h is 5, alpha is 4, and num_pe is 4, then cyc_num_alpha = ceil(5 / 16) = 1.
[0107] 302. Calculate the fifth loop count according to the second column computing power information and the fourth column computing power information.
[0108] In an embodiment of the present invention, after obtaining the second column computing power information b_w and the fourth column computing power information gamma, the second column computing power information a_w can be divided by the fourth column computing power information gamma to obtain an integer greater than or equal to the result value as the fifth loop count. Specifically, the second loop count cyc_gamma can be calculated through the following formula:
[0109] cyc_gamma = ceil(b_w / gamma)
[0110] In the above formula, ceil(b_w / gamma) represents taking the smallest integer greater than or equal to b_w / gamma. For example, if b_w is 5, gamma is 4, then cyc_gamma = ceil(5 / 4) = 2.
[0111] 303. Determine the second multiplication mode of the calculation area and the weight matrix according to the fourth loop count, the second loop count, and the fifth loop count.
[0112] In an embodiment of the present invention, when the fourth loop count, the second loop count, and the fifth loop count meet the preset conditions, it can be determined that the second multiplication mode of the calculation area and the weight matrix is a matrix multiplication of one column of first multiplication matrices and m second multiplication matrices. One column of first multiplication matrices includes num_pe first multiplication matrices.
[0113] Further, the product of the fourth loop count, the second loop count, and the fifth loop count can be calculated as the second loop total amount cyc_co. When the second loop total amount cyc_co is less than a preset value, it is determined that the second multiplication mode of the calculation area and the weight matrix is a matrix multiplication of one column of the first multiplication matrix and m second multiplication matrices. One column of the first multiplication matrix includes num_pe first multiplication matrices. Among them, the combination of the m second multiplication matrices should satisfy the rule of matrix multiplication with the combination of the num_pe first multiplication matrices, that is, the computing power information of the fourth row of the m second multiplication matrices is equal to the computing power information of the third column of the num_pe first multiplication matrices. Specifically, the above-mentioned second loop total amount cyc_co can be calculated by the following formula:
[0114] cyc_co = cyc_num_alpha * cyc_beta * cyc_gamma
[0115] 304. Reconstruct the second multiplication mode through the multiplication calculation unit.
[0116] In the embodiment of the present invention, after determining that the multiplication mode of the weight matrix and the calculation area is the second multiplication mode, the multiplication calculation unit is used to reconstruct the second multiplication mode under the current hardware conditions.
[0117] Further, taking the left multiplication of the weight matrix by the calculation area as an example, one second multiplication matrix is tiled in the calculation area, and num_pe first multiplication matrices are tiled in the corresponding column of the weight matrix. Specifically, reference can be made to Figure 3a , Figure 3a is a single-loop schematic diagram of a second multiplication mode provided by an embodiment of the present invention. In Figure 3a , taking num_pe = 4 as an example, the weight matrix A is tiled with a column of first multiplication matrices [A0, A1, A2, A3], and the corresponding column of the calculation area B is tiled with one second multiplication matrix B0. It can be seen that each time the second multiplication mode is executed, its corresponding computing power is num_pe * alpha * beta * gamma multiply-accumulate units. Combining with reference Figure 3b , Figure 3b is a multi-loop schematic diagram of a second multiplication mode provided by an embodiment of the present invention. In Figure 3bWhen cyc is 0, 1, 2, or 3, it represents the first, second, third, and fourth executions of the second multiplication mode respectively. In multiple loops, when cyc is 0, the first multiplication matrix [A0, A1, A2, A3] corresponding to the first column of the weight matrix A performs matrix multiplication with the second multiplication matrix B0 corresponding to the first element in the first row of the calculation area B; when cyc is 1, the second multiplication mode is executed, so that the first multiplication matrix [A0, A1, A2, A3] corresponding to the second column of the weight matrix A performs matrix multiplication with the second multiplication matrix B0 corresponding to the first element in the second row of the calculation area B; and so on. By executing the above second multiplication mode in multiple loops, the matrix multiplication of the weight matrix A and the calculation area B is completed. By completing the matrix multiplication of the weight matrix A and the calculation area B multiple times, the multiplication operation between the weight matrix and the feature map to be processed is completed.
[0118] In the embodiment of the present invention, through the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit, the second multiplication mode between the weight matrix and the calculation area is determined, and the multiplication calculation unit reconstructs the second multiplication mode, so that the weight matrix and the calculation area can fully release the effective computing power of the hardware under the given computing power, avoiding the situation of repeated reading and repeated calculation, thereby reducing the calculation time and improving the calculation efficiency.
[0119] It should be noted that in a possible embodiment, the above first multiplication mode and the above second multiplication mode can be set by choosing one of the two.
[0120] Specifically, if the above first loop total is less than the above second loop total, the multiplication mode between the above calculation area and the above weight matrix is determined to be the above first multiplication mode; if the above first loop total is greater than the above second loop total, the multiplication mode between the above calculation area and the above weight matrix is determined to be the above second multiplication mode; if the above first loop total is equal to the above second loop total, the multiplication mode between the above calculation area and the above weight matrix is determined to be a random one of the above first multiplication mode and the above second multiplication mode.
[0121] More specifically, the selection of the above first multiplication mode and the above second multiplication mode can be to compare the above second loop total cyc_ti with the second loop total cyc_co. When cyc_ti < cyc_co, the multiplication mode between the calculation area and the weight matrix is determined to be the first multiplication mode. When cyc_ti > cyc_co, the multiplication mode between the calculation area and the weight matrix is determined to be the second multiplication mode. When cyc_ti = cyc_co, a random selection can be made between the first multiplication mode and the second multiplication mode.
[0122] In an embodiment of the present invention, based on the magnitude relationship between the first total number of cycles and the second total number of cycles, the multiplication mode of the calculation region and the weight matrix is determined to be the first multiplication mode or the second multiplication mode. One of the multiplication modes with a smaller total number of cycles can be selected for reconstruction, so that when the weight matrix and the calculation region are under a given computing power, the calculation time can be further reduced and the calculation efficiency can be improved.
[0123] Optionally, in an embodiment of the present invention, when batch processing information is detected, based on the first number of cycles, the second number of cycles, and the fifth number of cycles, the third multiplication mode of the calculation region and the weight matrix is determined; and the third multiplication mode is reconstructed by the multiplication calculation unit.
[0124] It should be noted that the above third multiplication mode is used for batch processing of matrix multiplication. For example, when there is a weight matrix A performing matrix multiplication with multiple calculation regions (such as four calculation regions B_0, B_1, B_2, and B_3 at the same time), the third multiplication mode can be used to calculate the matrix multiplication. Of course, it can also be that the calculation region B performs matrix multiplication with multiple weight matrices at the same time.
[0125] Optionally, the above batch processing information includes the first batch processing information of the weight matrix and / or the second batch processing information of the to-be-processed feature map. Based on the first batch processing information and / or the second batch processing information, combined with the first number of cycles, the second number of cycles, and the fifth number of cycles, the third multiplication mode of the calculation region and the weight matrix is determined.
[0126] Further, the product of the first number of cycles, the second number of cycles, and the fifth number of cycles can be calculated as the third total number of cycles cyc_ba. When the third total number of cycles cyc_ba is less than a preset value, the third multiplication mode of the calculation region and the weight matrix is determined to be that j first multiplication matrices among multiple matrices perform matrix multiplication with k corresponding second multiplication matrices. The combination of j first multiplication matrices should satisfy the rule of performing matrix multiplication with the combination of k first multiplication matrices, that is, the computing power information of the third column of j is equal to the computing power information of the fourth row of k. Specifically, the above third total number of cycles cyc_ba can be calculated by the following formula:
[0127] cyc_ba = cyc_alpha * cyc_beta * cyc_gamma
[0128] Further, taking the left multiplication of the weight matrix by multiple calculation regions as an example, one second multiplication matrix is tiled in the calculation region, and one first multiplication matrix A0 corresponding to the column of the weight matrix A is tiled. Specifically, reference can be made to Figure 3c , Figure 3cThis is a single-cycle schematic diagram of a third multiplication mode provided by an embodiment of the present invention. In Figure 3c , taking 4 calculation regions as an example, the calculation regions are B_0, B_1, B_2, and B_3 respectively. The weight matrix A tiles a first multiplication matrix A0, and the calculation regions B_0, B_1, B_2, and B_3 each tile a second multiplication matrix B0 at the corresponding positions. It can be seen that each time the third multiplication mode is executed, its corresponding computing power is B_x * alpha * beta * gamma multiply-accumulate units. Combining with reference to Figure 3d , Figure 3d This is a multi-cycle schematic diagram of a second multiplication mode provided by an embodiment of the present invention. In Figure 3d , when cyc is 0, 1, 2, or 3, it represents the 1st, 2nd, 3rd, and 4th executions of the third multiplication mode respectively. In the multi-cycle, when cyc is 0, the 1st second multiplication mode is executed, so that the first multiplication matrix A0 in the first row and first column of the weight matrix A performs matrix multiplication with the second multiplication matrix B0 in the first row and first column of the calculation regions B_0, B_1, B_2, and B_3 at the same time; when cyc is 1, the 2nd third multiplication mode is executed, so that the first multiplication matrix corresponding to the first row and second column of the weight matrix A performs matrix multiplication with the second multiplication matrix B0 in the second row and first column of the calculation regions B_0, B_1, B_2, and B_3; and so on. By executing the above third multiplication mode multiple times in a loop, the batch matrix multiplication of the weight matrix A and multiple calculation regions B_0, B_1, B_2, and B_3 is completed. By completing the matrix multiplication of the weight matrix A and the calculation regions B_0, B_1, B_2, and B_3 multiple times, the multiplication operation processing of the weight matrix and the feature map to be processed is completed.
[0129] In the embodiment of the present invention, through the first batch of processing information and / or the second batch of processing information, combined with the first loop count, the second loop count, and the fifth loop count, the third multiplication mode of the weight matrix and the calculation region is determined, and the third multiplication mode is reconstructed by the multiplication calculation unit, so that the weight matrix and the calculation region can fully release the effective computing power of the hardware during the batch processing under the given computing power, avoiding the situation of repeated reading and repeated calculation during the batch processing, thereby reducing the time of batch processing calculation and improving the batch processing calculation efficiency.
[0130] It should be noted that the method for processing the feature map provided by the embodiment of the present invention can be applied to devices such as smart phones, computers, and servers that can process feature maps.
[0131] Optionally, please refer to Figure 4 , Figure 4 This is a structural schematic diagram of a device for processing a feature map provided by an embodiment of the present invention. As shown in Figure 4 , the device includes:
[0132] The first acquisition module 401 is configured to acquire the computing power information of the weight matrix, where the computing power information of the weight matrix includes the first-row computing power information and the first-column computing power information;
[0133] The second acquisition module 402 is configured to determine the calculation area corresponding to the weight matrix in the to-be-processed feature map, and acquire the computing power information of the calculation area, where the computing power information of the calculation area includes the second-row computing power information and the second-column computing power information;
[0134] The reconstruction module 403 is configured to dynamically reconstruct the multiplication pattern of the calculation area and the weight matrix according to the first-row computing power information, the first-column computing power information, the second-row computing power information, the second-column computing power information, and a preset multiplication calculation unit;
[0135] The processing module 404 is configured to perform a multiplication operation on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern.
[0136] Optionally, as Figure 5 shown, the apparatus further includes:
[0137] The third acquisition module 405 is configured to acquire the basic computing unit information under the current hardware conditions;
[0138] The determination module 406 is configured to determine a first multiplication matrix and a second multiplication matrix according to the basic computing unit information, where the first multiplication matrix includes the third-row computing power information and the third-column computing power information, the second multiplication matrix includes the fourth-row computing power information and the fourth-column computing power information, and the third-column computing power information is equal to the fourth-row computing power information;
[0139] The construction module 407 is configured to construct the multiplication calculation unit that executes the matrix multiplication of the first multiplication matrix and the second multiplication matrix.
[0140] Optionally, as Figure 6 shown, the reconstruction module 403 includes:
[0141] The first calculation unit 4031 is configured to calculate a first number of loops according to the first-row computing power information and the third-row computing power information;
[0142] The second calculation unit 4032 is configured to calculate a second number of loops according to the first-column computing power information and the third-column computing power information, or according to the second-row computing power information and the fourth-row computing power information;
[0143] The third calculation unit 4033 is configured to calculate a third number of loops according to the second-column computing power information, the fourth-column computing power information, and the column tiling quantity;
[0144] The first determination unit 4034 is configured to determine a first multiplication mode of the calculation region and the weight matrix according to the first number of cycles, the second number of cycles, and the third number of cycles;
[0145] The first reconstruction unit 4035 is configured to reconstruct the first multiplication mode through the multiplication calculation unit.
[0146] Optionally, as Figure 7 shown, the reconstruction module 403 further includes:
[0147] The fourth calculation unit 4036 is configured to calculate a fourth number of cycles according to the first row computing power information, the third row computing power information, and the row tiling quantity;
[0148] The fifth calculation unit 4037 is configured to calculate a fifth number of cycles according to the second column computing power information and the fourth column computing power information;
[0149] The second determination unit 4038 is configured to determine a second multiplication mode of the calculation region and the weight matrix according to the fourth number of cycles, the second number of cycles, and the fifth number of cycles;
[0150] The second reconstruction unit 4039 is configured to reconstruct the second multiplication mode through the multiplication calculation unit.
[0151] Optionally, as Figure 8 shown, the first multiplication mode includes a first total number of cycles, the second multiplication mode includes a second total number of cycles, and the reconstruction module 403 further includes:
[0152] The third determination unit 40310 is configured to determine that the multiplication mode of the calculation region and the weight matrix is the first multiplication mode if the first total number of cycles is less than the second total number of cycles;
[0153] The fourth determination unit 40311 is configured to determine that the multiplication mode of the calculation region and the weight matrix is the second multiplication mode if the first total number of cycles is greater than the second total number of cycles;
[0154] The fifth determination unit 40312 is configured to determine that the multiplication mode of the calculation region and the weight matrix is randomly one of the first multiplication mode and the second multiplication mode if the first total number of cycles is equal to the second total number of cycles.
[0155] Optionally, as Figure 9 shown, the reconstruction module 403 further includes:
[0156] The sixth determination unit 40313 is configured to determine a third multiplication pattern of the calculation region and the weight matrix according to the first number of cycles, the second number of cycles, and the fifth number of cycles when batch processing information is detected.
[0157] The third reconstruction unit 40314 is configured to reconstruct the third multiplication pattern through the multiplication calculation unit.
[0158] Optionally, the batch processing information includes first batch processing information of the weight matrix and / or second batch processing information of the feature map to be processed. The sixth determination unit 40313 is further configured to determine a third multiplication pattern of the calculation region and the weight matrix according to the first batch processing information and / or the second batch processing information, in combination with the first number of cycles, the second number of cycles, and the fifth number of cycles.
[0159] It should be noted that the feature map processing device provided in the embodiments of the present invention can be applied to devices such as smart phones, computers, and servers that can process feature maps.
[0160] The feature map processing device provided in the embodiments of the present invention can implement each process implemented by the feature map processing method in the above method embodiments, and can achieve the same beneficial effects. To avoid repetition, it will not be elaborated here.
[0161] See Figure 10 , Figure 10 is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. As Figure 10 shown, it includes: a memory 1002, a processor 1001, and a computer program for the feature map processing method stored on the memory 1002 and executable on the processor 1001, where:
[0162] The processor 1001 is configured to call the computer program stored in the memory 1002 and execute the following steps:
[0163] Obtain computing power information of the weight matrix, where the computing power information of the weight matrix includes first row computing power information and first column computing power information;
[0164] Determine a calculation region corresponding to the weight matrix in the feature map to be processed, and obtain computing power information of the calculation region, where the computing power information of the calculation region includes second row computing power information and second column computing power information;
[0165] Dynamically reconstruct a multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit;
[0166] Perform a multiplication operation on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern.
[0167] Optionally, before the processor 1001 performs dynamic reconstruction of the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit, the method further includes:
[0168] Obtain the basic calculation unit information under the current hardware conditions;
[0169] According to the basic calculation unit information, determine a first multiplication matrix and a second multiplication matrix. The first multiplication matrix includes third row computing power information and third column computing power information, and the second multiplication matrix includes fourth row computing power information and fourth column computing power information. The third column computing power information is equal to the fourth row computing power information;
[0170] Construct the multiplication calculation unit that executes the matrix multiplication of the first multiplication matrix and the second multiplication matrix.
[0171] Optionally, the dynamic reconstruction of the multiplication pattern of the calculation region and the weight matrix by the processor 1001 according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit includes:
[0172] Calculate a first number of loops according to the first row computing power information and the third row computing power information;
[0173] Calculate a second number of loops according to the first column computing power information and the third column computing power information, or according to the second row computing power information and the fourth row computing power information;
[0174] Calculate a third number of loops according to the second column computing power information, the fourth column computing power information, and the column tiling quantity;
[0175] Determine a first multiplication pattern of the calculation region and the weight matrix according to the first number of loops, the second number of loops, and the third number of loops;
[0176] Reconstruct the first multiplication pattern through the multiplication calculation unit.
[0177] Optionally, the dynamic reconstruction of the multiplication pattern of the calculation region and the weight matrix by the processor 1001 according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit includes:
[0178] Calculate a fourth loop count based on the first row computing power information, the third row computing power information, and the row tiling count;
[0179] Calculate a fifth loop count based on the second column computing power information and the fourth column computing power information;
[0180] Determine a second multiplication pattern of the calculation region and the weight matrix based on the fourth loop count, the second loop count, and the fifth loop count;
[0181] Reconstruct the second multiplication pattern through the multiplication calculation unit.
[0182] Optionally, the first multiplication pattern includes a first loop total, the second multiplication pattern includes a second loop total, and the dynamic reconstruction of the multiplication pattern of the calculation region and the weight matrix by the processor 1001 according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit further includes:
[0183] If the first loop total is less than the second loop total, determine that the multiplication pattern of the calculation region and the weight matrix is the first multiplication pattern;
[0184] If the first loop total is greater than the second loop total, determine that the multiplication pattern of the calculation region and the weight matrix is the second multiplication pattern;
[0185] If the first loop total is equal to the second loop total, determine that the multiplication pattern of the calculation region and the weight matrix is either the first multiplication pattern or the second multiplication pattern randomly.
[0186] Optionally, the dynamic reconstruction of the multiplication pattern of the calculation region and the weight matrix by the processor 1001 according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit further includes:
[0187] When batch processing information is detected, determine a third multiplication pattern of the calculation region and the weight matrix according to the first loop count, the second loop count, and the fifth loop count;
[0188] Reconstruct the third multiplication pattern through the multiplication calculation unit.
[0189] Optionally, the batch processing information executed by the processor 1001 includes the first batch processing information of the weight matrix and / or the second batch processing information of the feature map to be processed. Determining the third multiplication mode of the calculation area and the weight matrix according to the first number of cycles, the second number of cycles, and the fifth number of cycles includes:
[0190] Determine the third multiplication mode of the calculation area and the weight matrix according to the first batch processing information and / or the second batch processing information, in combination with the first number of cycles, the second number of cycles, and the fifth number of cycles.
[0191] It should be noted that the electronic device provided in the embodiments of the present invention can be applied to devices such as smartphones, computers, and servers that can perform processing of feature maps.
[0192] The electronic device provided in the embodiments of the present invention can implement each process implemented by the feature map processing method in the above method embodiments, and can achieve the same beneficial effects. To avoid repetition, it will not be elaborated here.
[0193] The embodiments of the present invention also provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements each process of the feature map processing method or the application-side feature map processing method provided in the embodiments of the present invention, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0194] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0195] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A method for processing a feature map, characterized in that, it includes the following steps: Obtain the computing power information of the weight matrix, where the computing power information of the weight matrix includes the first row computing power information and the first column computing power information; Determine the calculation area corresponding to the weight matrix in the to-be-processed feature map, and obtain the computing power information of the calculation area, where the computing power information of the calculation area includes the second row computing power information and the second column computing power information; Dynamically reconstruct the multiplication pattern of the calculation area and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit; Perform a multiplication operation on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern; Before dynamically reconstructing the multiplication pattern of the calculation area and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit, the method further includes: Obtain the basic calculation unit information under the current hardware conditions; According to the basic calculation unit information, determine a first multiplication matrix and a second multiplication matrix, where the first multiplication matrix includes the third row computing power information and the third column computing power information, the second multiplication matrix includes the fourth row computing power information and the fourth column computing power information, and the third column computing power information is equal to the fourth row computing power information; Construct the multiplication calculation unit for performing the matrix multiplication of the first multiplication matrix and the second multiplication matrix.
2. The method according to claim 1, characterized in that, dynamically reconstructing the multiplication pattern of the calculation area and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit includes: Calculate a first loop count according to the first row computing power information and the third row computing power information; Calculate a second loop count according to the first column computing power information and the third column computing power information, or according to the second row computing power information and the fourth row computing power information; Calculate a third loop count according to the second column computing power information, the fourth column computing power information, and the column tiling number; Determine the first multiplication pattern of the calculation area and the weight matrix according to the first loop count, the second loop count, and the third loop count; Reconstruct the first multiplication pattern through the multiplication calculation unit.
3. The method according to claim 2, characterized in that, dynamically reconstructing the multiplication pattern of the calculation area and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit includes: Calculate a fourth loop count according to the first row computing power information, the third row computing power information, and the row tiling number; Calculate a fifth loop count according to the second column computing power information and the fourth column computing power information; Determine a second multiplication pattern of the calculation region and the weight matrix according to the fourth number of cycles, the second number of cycles, and the fifth number of cycles; Reconstruct the second multiplication pattern through the multiplication calculation unit.
4. The method according to claim 3, wherein, the first multiplication pattern includes a first total number of cycles, the second multiplication pattern includes a second total number of cycles, and the dynamic reconstruction of the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit further includes: If the first total number of cycles is less than the second total number of cycles, determine the multiplication pattern of the calculation region and the weight matrix as the first multiplication pattern; If the first total number of cycles is greater than the second total number of cycles, determine the multiplication pattern of the calculation region and the weight matrix as the second multiplication pattern; If the first total number of cycles is equal to the second total number of cycles, determine the multiplication pattern of the calculation region and the weight matrix as a random one of the first multiplication pattern and the second multiplication pattern.
5. The method according to claim 3, wherein, the dynamic reconstruction of the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit further includes: When batch processing information is detected, determine a third multiplication pattern of the calculation region and the weight matrix according to the first number of cycles, the second number of cycles, and the fifth number of cycles; Reconstruct the third multiplication pattern through the multiplication calculation unit.
6. The method according to claim 5, wherein, the batch processing information includes first batch processing information of the weight matrix and / or second batch processing information of the to-be-processed feature map, and determining the third multiplication pattern of the calculation region and the weight matrix according to the first number of cycles, the second number of cycles, and the fifth number of cycles includes: Determine the third multiplication pattern of the calculation region and the weight matrix according to the first batch processing information and / or the second batch processing information, in combination with the first number of cycles, the second number of cycles, and the fifth number of cycles.
7. A processing device for a feature map, wherein, the device includes: A first acquisition module, configured to acquire computing power information of a weight matrix, where the computing power information of the weight matrix includes first row computing power information and first column computing power information; A second acquisition module, configured to determine a calculation region corresponding to the weight matrix in the to-be-processed feature map and acquire computing power information of the calculation region, where the computing power information of the calculation region includes second row computing power information and second column computing power information; A reconstruction module, configured to dynamically reconstruct the multiplication pattern of the calculation region and the weight matrix according to the first row computing power information, the first column computing power information, the second row computing power information, the second column computing power information, and a preset multiplication calculation unit; A processing module, configured to perform a multiplication operation on the to-be-processed feature map and the weight matrix through the reconstructed multiplication pattern; The apparatus further includes: A third acquisition module, configured to acquire basic computing unit information under current hardware conditions; A determination module, configured to determine a first multiplication matrix and a second multiplication matrix according to the basic computing unit information, where the first multiplication matrix includes third row computing power information and third column computing power information, the second multiplication matrix includes fourth row computing power information and fourth column computing power information, and the third column computing power information is equal to the fourth row computing power information; A construction module, configured to construct the multiplication calculation unit that executes the matrix multiplication of the first multiplication matrix and the second multiplication matrix.
8. An electronic device, characterized in that, it includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps in the method for processing a feature map according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the method for processing a feature map according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Accelerating 2d convolutional layer mapping on a dot product architecture
US20210182025A1