Grouping convolution processing optimization device and grouping convolution processing optimization method
By adjusting weight matrices to align with memory addressing constraints and inserting zero components, the grouped convolution processing optimization method enhances the performance of AI chips with restricted memory addressing, allowing for efficient execution of grouped convolutions.
Patent Information
- Application Number
- JP2023206216
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-18
AI Technical Summary
AI chips with restricted memory addressing constraints face performance degradation and slow execution speeds when performing grouped convolutions, as they struggle to satisfy memory addressing constraints during the grouped convolution process.
The proposed solution involves a grouped convolution processing optimization method that adjusts the weight matrices used in convolution layers to align with the AI chip's memory addressing constraints, achieving this without explicit memory copying. This is done by inserting zero components into the weight matrices, allowing the AI chip to execute grouped convolutions efficiently.
The optimization method enables AI chips with memory addressing constraints to perform grouped convolutions at high speeds, reducing the overhead associated with memory adjustments and maintaining the accuracy of the convolution results.
Smart Images

Figure 2025091144000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a grouped convolution processing optimization device, a grouped convolution processing optimization method, and a grouped convolution processing optimization program.
Background Art
[0002] A convolutional neural network (CNN) is a feed-forward neural network having a structure in which two types of layers, a convolutional layer and a pooling layer, are alternately stacked. Hereinafter, the convolutional neural network is also simply referred to as CNN.
[0003] FIG. 12 is an explanatory diagram showing an example of a convolutional neural network. In the CNN shown in FIG. 12, a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer are alternately stacked.
[0004] Also, C1 and C2 shown in FIG. 12 each represent a convolutional calculation. For example, a convolutional calculation C1 is executed on the input image input to the first convolutional layer.
[0005] Note that an image is an example of the input data. The data input to the CNN may be data other than an image.
[0006] Also, P1 and P2 shown in FIG. 12 each represent a pooling calculation. For example, a pooling calculation P1 is executed on the convolutional calculation result input to the first pooling layer.
[0007] Also, F shown in FIG. 12 represents a fully connected network. The fully connected network F has the function of a fully connected layer that connects all the nodes of the second pooling layer and the nodes of the output layer. Finally, the output of the CNN is obtained from the output layer.
[0008] The following specifically explains the calculation of convolution in a CNN. FIG. 13 is an explanatory diagram showing an example of the calculation of convolution in a CNN. Note that the example of the calculation of convolution shown in FIG. 13 corresponds to the convolution calculation C1 shown in FIG. 12.
[0009] The input image shown in FIG. 13 is the image input to the CNN. The input image shown in FIG. 13 is composed of the first channel to the C in channel (C in is an integer of 2 or more) arranged in order. That is, C in means the number of input channels. Also, as shown in FIG. 13, the vertical size of the image constituting the input image is H, and the horizontal size is W.
[0010] For simplicity of explanation, as the input X for the calculation of convolution, consider an image with a vertical size of 1, a horizontal size of 1, and a channel number of C in marked with a grid pattern as shown in FIG. 13. At the bottom of FIG. 13, the input X when viewed from the height direction is shown. Also, the symbols below the input X shown in FIG. 13 are the channel identification numbers (the same applies to other figures).
[0011] That is, in the example of the calculation of convolution shown in FIG. 13, the kernel size is "1×1". However, the content of the following explanation is the same even if the kernel size is a size other than "1×1" (for example, "3×3" or "5×5").
[0012] In the convolution calculation shown in FIG. 13, the input X is multiplied by the weight W represented by a hatched pattern with a horizontal size of C out and a vertical size of C in . As a result of the multiplication, an output Y0 which is an image with a channel number of C out represented by the black rectangle shown in FIG. 13 is obtained. That is, C out means the number of output channels.
[0013] Note that the convolution calculation shown in Fig. 13 corresponds to the multiplication of matrices. That is, in the convolution calculation shown in Fig. 13, the weight W is treated as a matrix. Also, the input X and the output Y0 are each treated as a vector, which is a type of matrix. The "weight" in this specification is exactly a "weight matrix", but for simplicity, it is also simply referred to as "weight".
[0014] Note that in this specification, consistently, the input X and the output Y0 are expressed as row vectors (vectors with components arranged horizontally), and the weight is expressed as a matrix. Also, the convolution operation is expressed as the product of multiplying the input row vector from the left by the weight matrix.
[0015] In this specification, the expressions "row" and "column" are used under the above premise. That is, when read as another mathematically equivalent expression, the above "row" and "column" are also read correspondingly.
[0016] Also, the CNNs shown in Figs. 12 to 13 are pre-trained models. That is, the weight W shown in Fig. 13 is also a weight obtained by performing pre-training.
[0017] As the above method of convolution calculation, the number of CNNs using grouped convolution is increasing. In grouped convolution, C in and C out are divided into G (G is an integer greater than or equal to 2) groups for calculation processing. For example, Non-Patent Document 1 describes grouped convolution.
[0018] The amount of computation in grouped convolution is less than that in ordinary convolution. Also, it has been experimentally shown that the recognition accuracy of ResNeXt using grouped convolution is higher than that of ResNet (Residual Neural Networks) using ordinary convolution, which is a type of CNN.
[0019] The following specifically describes the calculation of grouped convolution in a CNN. FIG. 14 is an explanatory diagram showing an example of the calculation of grouped convolution in one grouped convolution layer.
[0020] Note that the example of the calculation of grouped convolution shown in FIG. 14 is an example of the case where grouped convolution is applied instead of the calculation of convolution shown in FIG. 13. Also, in this example, consider the case where an AI (Artificial Intelligence) chip executes the calculation of grouped convolution.
[0021] In the example of the calculation of grouped convolution shown in FIG. 14, the number of groups is defined as "8". That is, the grouped convolution layer used in the example shown in FIG. 14 is a layer that is pre - defined to perform convolution calculation by dividing the input X into 8 groups in the channel direction and is pre - learned.
[0022] Note that the input X, which is an image, is an example of the input data. The data input to the grouped convolution layer may be data other than images.
[0023] Therefore, as shown in FIG. 14, in the calculation of grouped convolution, weights W with a vertical size of (C in / 8) and a horizontal size of (C out / 8) are prepared. There can be various implementation methods for the AI chip to execute the calculation of grouped convolution. In this example, assume that the AI chip operates as follows. a ~weights W h There can be various methods for implementing the AI chip for performing the calculation of grouped convolution. In this example, assume that the AI chip operates as follows, for example.
[0024] The AI chip multiplies the weight W a with the image (with the number of channels (C in / 8)) composed of the divided first channel to the C in / 8 channels. As a result of the multiplication, the AI chip obtains an image with the number of channels (C out / 8) as the output.
[0025] As shown in FIG. 14, the AI chip performs the above calculations for each of the weights W b ~weight W h as well. That is, the AI chip divides the input X into 8 parts in the channel direction, and for the image composed of the ((i - 1)×C in / 8 + 1)-th channel to the (i×C in / 8)-th channel (i = 1 to 8), it uses the i-th weight matrix to perform the convolution calculation for each i from 1 to 8. Note that the first weight matrix to the eighth weight matrix respectively correspond to the weights W a ~weight W h . As a result of each calculation, the AI chip obtains 8 images with the number of channels being (C out / 8).
[0026] Finally, the AI chip places the obtained images at the same positions as the positions of the 8-divided input X used in the calculation. After placing each of the obtained 8 images, the AI chip combines (the "concat" shown in FIG. 14) each image.
[0027] By combining, the AI chip obtains the output Y, which is an image with the number of channels being C out , corresponding to the calculation result of the normal convolution. Note that the output Y obtained by the grouped convolution calculation shown in FIG. 14 is not equivalent to the output Y0 obtained by the convolution calculation shown in FIG. 13.
[0028] The amount of calculation for convolution is proportional to the size of the weights. For example, the amount of calculation for the convolution shown in FIG. 13 is proportional to the size of the weight W, which is (C in ×C out ). Also, the amount of calculation for the grouped convolution shown in FIG. 14 is proportional to the sum of the sizes of each of the weights W a ~weight W h , which is {((C in / 8)×(C out / 8)×8)}. That is, the amount of calculation for the grouped convolution shown in FIG. 14 is 1 / 8 of the amount of calculation for the convolution shown in FIG. 13.
[0029] Generally, when the input X is divided into G groups and the grouped convolution calculation is performed, the computational complexity of the grouped convolution shown in FIG. 14 is { (C in / G) × (C out / G) × G} = (C in × C out ) / G, which is proportional to. That is, the computational complexity of the grouped convolution shown in FIG. 14 is 1 / G of the computational complexity of the convolution shown in FIG. 13.
[0030] FIG. 15 is an explanatory diagram showing an example of other calculations of grouped convolution in one grouped convolution layer. From the upper left to the lower right of the weight W shown in FIG. 15, the weights W a to W h shown in FIG. 14 are arranged in the order of the weights W a to W h on the diagonal. Note that the values of the components of the weight W other than the components where the weights W a to W h are arranged (the components at the dotted pattern locations shown in FIG. 15) are arbitrary values.
[0031] As shown in FIG. 15, the weight W has a vertical size of C in and a horizontal size of C out . In the calculation of the grouped convolution shown in FIG. 15, the AI chip performs the calculation of multiplying the weight W and the input X only once. By performing the calculation only once, the AI chip obtains the output Y.
[0032] For an AI chip that is not suitable for grouped convolution, for example, the calculation process of grouped convolution may be implemented as a calculation process of multiple convolutions. Therefore, when an AI chip in which the calculation process of multiple convolutions is implemented performs the grouped convolution calculation, it is affected by the overhead associated with the call of the convolution calculation G times. When the AI chip is affected by the overhead G times, the calculation speed of the grouped convolution decreases.
[0033] In the calculation of grouped convolution shown in FIG. 15, the number of times the AI chip is affected by overhead is minimized (once). Therefore, the calculation speed of grouped convolution may be improved.
Prior Art Documents
Non-Patent Documents
[0034]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0035] As other possibilities for the AI chip to be unsuitable for grouped convolution, there may be the following possibilities. Some AI chips for CNNs have strong constraints on the number of rows and columns of weights corresponding to the number of channels and their arrangement in memory, that is, memory addressing.
[0036] Note that "constraints are imposed" means that as long as the constraints are adhered to, the performance of the calculation speed can be fully exerted, but if the constraints are violated, the performance of the calculation speed will be significantly reduced. The reason for imposing constraints on memory addressing is to reduce the weight of the memory access circuit of the AI chip.
[0037] For example, addressing by equal multiples of a specific constant is a constraint imposed on memory addressing. When the memory address is represented by 32 bits, for example, the leading address of the array must be arranged at a multiple of 16, ignoring the lower 4 bits, that is, access to the array is possible without wiring the circuit for the lower 4 bits.
[0038] When constraints are imposed on the starting address of the weights used in grouped convolution, the convolution operation is executed for each group. Therefore, the constraints are substantially C for each group in *C out or C in *C out *(kernel size). Fig. 16 is an explanatory diagram showing an example in which the weights used in grouped convolution are arranged in memory.
[0039] As shown in Fig. 16, the weights used in grouped convolution are generally held as a continuous block of data in memory. The size per weight shown in Fig. 16 is ((C in / 8)*(C out / 8)). Since the number of groups is 8, the weight size is ((C in / 8)*(C out / 8))*8. Software such as a compiler that determines the memory layout of the weights arranges the weights in memory so that the starting address of the weights with a size of ((C in / 8)*(C out / 8))*8 satisfies the constraints imposed on the AI chip.
[0040] However, in grouped convolution, as shown in Fig. 14, the input and weights can be divided and executed. At this time, it is required to satisfy the constraints for each division unit of the input and weights. For example, the size per weight shown in Fig. 16 is ((C in / 8)*(C out / 8)). Therefore, when obtaining each weight one by one, the starting addresses of the weights are 0*((C in / 8)*(C out / 8)), 1*((C in / 8)*(C out / 8)), ···, so it is necessary for the starting addresses of each of the 8 groups to satisfy the constraints.
[0041] However, the number of elements in each group ((C in / 8)*(C outDepending on the value of ( / 8)), the starting addresses of all groups may not satisfy the constraints.
[0042] Depending on the design of the AI chip, if the constraints are not satisfied, not only will the performance degrade, but it may not even operate properly. In such a case, it may be necessary to add a process to adjust the arrangement on the memory so that the constraints are satisfied. As a result of adding the process to adjust the arrangement on the memory, the overall processing by the AI chip may become slower.
[0043] For example, in the grouped convolution shown in FIG. 14, when C in = 160 and G = 16, the number of channels of the divided input X is C in / G = 10. That is, for an AI chip composed of hardware with the constraint that C in should be a multiple of 16, the condition that C in is a multiple of 16 is not satisfied. Also, the starting addresses of the weights on the memory other than the weight W a are not multiples of 16.
[0044] Therefore, in order to satisfy the constraints, for example, it is required to perform adjustment of the starting memory addresses by memory copy or the like for the input X and the weight W a before each convolution is executed. The overhead associated with the memory copy is likely to cause a decrease in the speed of the grouped convolution.
[0045] Also, an AI chip specialized for convolution may not have a circuit for memory copy. In this case, the memory copy may be realized by a slow circuit such as a host CPU (Central Processing Unit) or an auxiliary CPU that can execute processing flexibly. When the memory copy is realized by the host CPU or the auxiliary CPU, the speed of the grouped convolution further decreases.
[0046] As a method to satisfy the constraints, zero-padding can be considered. However, it is required to perform zero-padding intermittently between the divided units of the input X instead of at the end of the input X. Therefore, in grouped convolution, even when zero-padding is performed, the overhead associated with memory copy is large.
[0047] Figure 17 is an explanatory diagram showing an example of performing zero-padding on the input X for grouped convolution. Figure 17 shows an example where, in grouped convolution with C in = 160 and G = 8, the start address of each divided unit of the input X is required to be aligned to a multiple of 16.
[0048] As shown in the upper part of Figure 17, C in / G = 20. Therefore, the start address of each divided unit of the input X is a multiple of 20 and not a multiple of 16.
[0049] Among the multiples of 16, the smallest multiple greater than or equal to 20 is 32. Therefore, consider performing zero-padding so that the start address of each divided unit of the input X becomes a multiple of 32 (hereinafter referred to as the "32-aligned state").
[0050] As shown in the lower part of Figure 17, the 32-aligned state is realized by performing zero-padding between the divided units of the input X. The white rectangles shown in Figure 17 represent regions where the components are zero. Also, the number of input channels is changed to C in ’ = 256.
[0051] However, if zero-padding is designed to be executed before each grouped convolution, multiple memory copies are also added in the design. The additional memory copy is, for example, a process of newly securing an area of C in ’ = 256 in the memory and copying each divided unit of the input X to a location where the address is a multiple of 32. That is, since it is a process with a relatively high load, the memory copy may cause a decrease in the speed of grouped convolution.
[0052] A technique that can solve the problem of a decrease in the speed at which a hardware circuit with restricted memory addressing executes grouped convolution is not described in Non-Patent Document 1 either.
[0053] Therefore, an object of the present disclosure is to provide a grouped convolution process optimization device, a grouped convolution process optimization method, and a grouped convolution process optimization program that enable a hardware circuit with restricted memory addressing to execute grouped convolution at high speed.
Means for Solving the Problems
[0054] For a learned convolutional neural network in which an input convolution is defined in which an N×N input weight matrix is used to perform convolution calculation on data composed of the first channel to the Nth channel (N is an integer of 2 or more) arranged in order, a grouped convolution in which the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and an N / G×N / G weight matrix of the i-th is used for data composed of the ((i - 1)×N / G + 1)-th channel to the (i×N / G)-th channel (i = 1 to G) to perform convolution calculation for each i = 1 to G, and an output convolution in which an N×N output weight matrix is used to perform convolution calculation on the result of the grouped convolution are defined, when a constraint is imposed such that convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each i = 1 to G, and a first insertion unit that executes at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the lower end of the i-th weight matrix for each i = 1 to G, an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / G-th column of the input weight matrix for each i = 1 to G, and a second insertion unit that executes at least one of an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G, and the second insertion unit executes the input weight matrix insertion process when the row insertion process is executed, and executes the output weight matrix insertion process when the column insertion process is executed.
[0055] The grouped convolution processing optimization method according to the present disclosure is for input convolution in which an input weight matrix of N rows and N columns is used to perform convolution calculation on data composed of the first channel to the Nth channel (N is an integer of 2 or more) arranged in order, and the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and for data composed of the {(i - 1)×N / G + 1}th channel to the (i×N / G)th channel (i = 1 to G) after division, the ith weight matrix of N / G rows and N / G columns is used to perform convolution calculation respectively for i = 1 to G. For a learned convolutional neural network in which output convolution in which an output weight matrix of N rows and N columns is used to perform convolution calculation on the result of the grouped convolution is defined respectively, when a constraint is imposed that convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, at least one of a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the ith weight matrix respectively for i = 1 to G, and a row insertion process of inserting (M - N / G) rows of rows with zero components at the bottom end of the ith weight matrix respectively for i = 1 to G is executed, an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components on the right side of the (i×N / G)th column of the input weight matrix respectively for i = 1 to G, and an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the (i×N / G)th row of the output weight matrix respectively for i = 1 to G is executed. When the row insertion process is executed, the input weight matrix insertion process is executed, and when the column insertion process is executed, the output weight matrix insertion process is executed.
[0056] The grouped convolution processing optimization program according to the present disclosure causes a computer to perform an input convolution in which an input weight matrix of N rows and N columns is used to perform a convolution calculation on data configured by arranging the first channel to the Nth channel (N is an integer of 2 or more) in order, and the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and the i-th weight matrix of N / G rows and N / G columns is used for the data composed of the {(i - 1)×N / G + 1}-th channel to the (i×N / G)-th channel (i = 1 to G) to perform the convolution calculation for each i = 1 to G, and an output convolution in which an output weight matrix of N rows and N columns is used to perform a convolution calculation on the result of the grouped convolution. Regarding the learned convolutional neural network in which each is defined, when a constraint is imposed that the convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each i = 1 to G, and a first insertion process of performing at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the bottom end of the i-th weight matrix for each i = 1 to G, and an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components on the right side of the i×N / G-th column of the input weight matrix for each i = 1 to G, and an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G. A grouped convolution processing optimization program for causing the execution, in the second insertion process, to execute the input weight matrix insertion process when the row insertion process is executed, and to execute the output weight matrix insertion process when the column insertion process is executed.
Effect of the Invention
[0057] According to the present disclosure, a hardware circuit with constraints on memory addressing can execute grouped convolution at high speed.
Brief Description of the Drawings
[0058]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Mode for Carrying Out the Invention
[0059] [Explanation of Configuration] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the present disclosure, the drawings are associated with one or more embodiments.
[0060] FIG. 1 is a block diagram showing a configuration example of a grouped convolution processing optimization device according to the present disclosure. The grouped convolution processing optimization device 100 shown in FIG. 1 is communicably connected to a pre-change CNN model storage unit 200 and a post-change CNN model storage unit 300.
[0061] In the pre-change CNN model storage unit 200, the learned CNN model described above is stored, including the weights W a ~weights W h shown in FIGS. 14 to 16. The learned CNN model stored in the pre-change CNN model storage unit 200 is a model learned after grouped convolution is defined.
[0062] Further, in the post-change CNN model storage unit 300, the learned CNN model stored in the pre-change CNN model storage unit 200, which has been optimized by the grouped convolution processing optimization device 100, is stored.
[0063] Also, the AI chip 400 is communicably connected to the modified CNN model storage unit 300. The AI chip 400 is a chip that performs convolution calculations using the learned CNN model stored in the modified CNN model storage unit 300.
[0064] Also, as shown in FIG. 1, the grouped convolution process optimization device 100 includes an optimizable structure detection unit 110, an optimal channel number determination unit 120, a forward convolution weight correction unit 130, a grouped convolution weight correction unit 140, and a backward convolution weight correction unit 150.
[0065] The grouped convolution process optimization device 100 of the present embodiment assumes that grouped convolution is not used alone and that there are other convolutions before and after the grouped convolution. FIG. 2 is an explanatory diagram showing an example of the convolution process in ResNeXt described in Non-Patent Document 1.
[0066] As shown in FIG. 2, in ResNeXt, a representative learning model that uses grouped convolution, a convolution using a 1×1 kernel, a grouped convolution with G = 32 using a 3×3 kernel, and a convolution using a 1×1 kernel are used as a set. That is, there is always a convolution using a 1×1 kernel before and after the grouped convolution.
[0067] The grouped convolution process optimization device 100 of the present embodiment is characterized by realizing the 32-line state shown at the bottom of FIG. 17 without explicitly performing a memory copy. Specifically, the grouped convolution process optimization device 100 adjusts the calculation content of the convolutions before and after the grouped convolution so that the input to the grouped convolution becomes the 32-line state shown at the bottom of FIG. 17.
[0068] Also, the grouped convolution process optimization device 100 converts the weights used in the grouped convolution into weights on the premise that the input to the grouped convolution is in the 32-line state shown at the bottom of FIG. 17.
[0069] Even if the grouped convolution processing optimization device 100 of this embodiment adjusts the weights, the calculation result of the grouped convolution does not change. That is, the grouped convolution processing optimization device 100 realizes the acceleration of the grouped convolution by adjusting the weights without adding new processing.
[0070] The optimizable structure detection unit 110 has a function of detecting an optimizable grouped convolution among the grouped convolutions existing in the structure of the learned CNN model stored in the pre-change CNN model storage unit 200.
[0071] The optimizable structure detection unit 110 of this embodiment is characterized by considering the optimization of the grouped convolution in a set with the convolutions existing before and after the grouped convolution.
[0072] The optimal channel number determination unit 120 has a function of determining the new input channel number and the new output channel number of the detected grouped convolution of the CNN model after optimization.
[0073] The forward convolution weight correction unit 130 has a function of correcting the weights used in the convolution existing before the detected grouped convolution.
[0074] The grouped convolution weight correction unit 140 has a function of correcting each weight used in the detected grouped convolution.
[0075] The backward convolution weight correction unit 150 has a function of correcting the weights used in the convolution existing after the detected grouped convolution.
[0076] First, an example in which a grouped convolution and the convolutions existing before and after the grouped convolution are executed as a set will be described. FIG. 3 is an explanatory diagram showing an example in which a grouped convolution and the convolutions existing before and after the grouped convolution are executed as a set.
[0077] As shown in FIG. 3, first, for the data of C = 160 with horizontal stripes, convolution is performed using weights of 160 rows and 160 columns. Here, C represents the number of channels (the same applies in other figures).
[0078] Next, as shown in FIG. 3, grouped convolution is performed on the data of C = 160 with a lattice pattern. In the example shown in FIG. 3, G = 8. Therefore, for each data with a channel number of 160 / 8 = 20 that constitutes the data of C = 160 with a lattice pattern, convolutions using weights of 20 rows and 20 columns are each performed.
[0079] Note that in FIG. 3, for the sake of convenience, each weight used in the grouped convolution is arranged as in the example shown in FIG. 15. However, the arrangement method of each weight is not limited to the example shown in FIG. 15.
[0080] Next, as shown in FIG. 3, for the data of C = 160 in black, convolution is performed using weights of 160 rows and 160 columns. As a result of the convolution, data of C = 160 with vertical stripes is output.
[0081] The memory address indicated by the straight line in the data of C = 160 with a lattice pattern shown in FIG. 3 is 3 * 20 = 60. That is, the starting address of each data constituting the lattice pattern data is not a multiple of 16. An example of converting the lattice pattern data so that the starting address becomes a multiple of 16 is shown in FIG. 4.
[0082] FIG. 4 is an explanatory diagram showing another example in which grouped convolution and the convolutions existing before and after the grouped convolution are executed as a set. As shown in FIG. 4, by performing zero-padding shown at the bottom of FIG. 17 on the data of C = 160 with a lattice pattern, a state of 32 alignment is realized.
[0083] Specifically, the memory address indicated by the straight line in the C = 256 data of the lattice pattern shown in FIG. 4 is 3 * 32 = 96. That is, the starting address of each data constituting the lattice pattern data is a multiple of 16. However, as described above, the load imposed on the conversion process of the lattice pattern data shown in FIG. 4 is high.
[0084] Note that the grouped convolution shown in FIG. 4 and the convolution existing after the grouped convolution do not reflect the changes associated with the conversion of the lattice pattern data.
[0085] FIG. 5 is an explanatory diagram showing another example in which the grouped convolution and the convolutions existing before and after the grouped convolution are executed as a set. Each weight shown in FIG. 5 is a weight corrected by the forward convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the backward convolution weight correction unit 150 of the present embodiment, respectively.
[0086] For example, as shown in FIG. 5, first, for the data of C = 160 with a horizontal stripe pattern, a convolution using a weight of 160 rows and 256 columns is executed. The forward convolution weight correction unit 130 generates the weight of 160 rows and 256 columns by inserting columns with zero components 12 (= 32 - 20) columns to the right of the i × 20 (= 160 / 8) -th column of the weight of 160 rows and 160 columns shown in FIG. 3 for each i = 1 to 8.
[0087] By executing the convolution using the weight of 160 rows and 256 columns, as shown in FIG. 5, data of the lattice pattern of C = 256 in a 32 - aligned state is output. The memory address indicated by the straight line in the C = 256 data of the lattice pattern shown in FIG. 5 is 3 * 32 = 96. That is, the starting address of each data constituting the lattice pattern data is a multiple of 16.
[0088] Next, as shown in FIG. 5, grouped convolution is performed on the data of C = 256 with a lattice pattern. In the example shown in FIG. 5, for each data with 32 channels that constitutes the data of C = 256 with a lattice pattern, each convolution using weights of 32 rows and 32 columns is respectively performed.
[0089] The grouped convolution weight modification unit 140 generates the weights of 32 rows and 32 columns by inserting 12 (= 32 - 20) columns with zero components at the right end of the weights of 20 (= 160 / 8) rows and 20 (= 160 / 8) columns shown in FIG. 3 and inserting 12 (= 32 - 20) rows with zero components at the lower end of the weights of 20 rows and 20 columns. The grouped convolution weight modification unit 140 performs the above-described insertion processing of zero columns and zero rows for each of the eight 20-row and 20-column weights shown in FIG. 3.
[0090] By performing each convolution using weights of 32 rows and 32 columns, as shown in FIG. 5, black data of C = 256 in a state of 32 alignments is output.
[0091] Next, as shown in FIG. 5, convolution using weights of 256 rows and 160 columns is performed on the black data of C = 256. The backward convolution weight modification unit 150 generates the weights of 256 rows and 160 columns by inserting 12 (= 32 - 20) rows with zero components below the i×20 (= 160 / 8) -th row of the weights of 160 rows and 160 columns shown in FIG. 3 for each i = 1 to 8.
[0092] By performing convolution using weights of 256 rows and 160 columns, as shown in FIG. 5, data of C = 160 with a vertical stripe pattern is output. The data of C = 160 with a vertical stripe pattern shown in FIG. 5 is equal to the data of C = 160 with a vertical stripe pattern shown in FIG. 3.
[0093] The forward convolution weight modification unit 130, the grouped convolution weight modification unit 140, and the backward convolution weight modification unit 150 store the optimized learned CNN model, including the modified weights, in the modified CNN model storage unit 300.
[0094] In addition, in the present embodiment, when the grouped convolution weight modification unit 140 inserts a zero row into each weight, the forward convolution weight modification unit 130 inserts a zero column into the weight. Further, when the grouped convolution weight modification unit 140 inserts a zero column into each weight, the backward convolution weight modification unit 150 inserts a zero row into the weight.
[0095] [Description of Operations] Hereinafter, the operations of the grouped convolution processing optimization device 100 of the present embodiment will be described with reference to FIGS. 6 to 7. FIG. 6 is a flowchart showing the operations of the optimization processing by the grouped convolution processing optimization device 100 according to the present disclosure.
[0096] First, the optimizable structure detection unit 110 of the grouped convolution processing optimization device 100 acquires a learned CNN model from the pre-change CNN model storage unit 200 (step S110).
[0097] Next, the optimizable structure detection unit 110 determines whether grouped convolution exists in the structure of the acquired CNN model (step S120). If grouped convolution does not exist (No in step S120), the grouped convolution processing optimization device 100 ends the optimization processing.
[0098] If grouped convolution exists (Yes in step S120), the optimizable structure detection unit 110 determines whether C in / G or C out / G is a value suitable for the device (step S130).
[0099] Note that the device in this example is the AI chip 400. Also, a value suitable for the device is, for example, "32" in the state of 32 lines shown in the example of FIG. 5. C in / G and C out / G is a value suitable for the device (No in step S130), the optimizable structure detection unit 110 proceeds to the process of step S170.
[0100] Grouped convolution C in / G or C out / G is not a value suitable for the device (Yes in step S130), the optimizable structure detection unit 110 determines whether there is a convolution with G = 1 before and after the grouped convolution (step S140). If there is no convolution with G = 1 before and after the grouped convolution (No in step S140), the optimizable structure detection unit 110 proceeds to the process of step S170.
[0101] If there is a convolution with G = 1 before and after the grouped convolution (Yes in step S140), the optimizable structure detection unit 110 determines whether there is a processing layer other than the processing layer where element-wise operations are performed between the two convolutions and the grouped convolution (step S150). If there is a processing layer other than the processing layer where element-wise operations are performed between the two convolutions and the grouped convolution (No in step S150), the optimizable structure detection unit 110 proceeds to the process of step S170.
[0102] If there is no processing layer other than the processing layer where element-wise operations are performed between the two convolutions and the grouped convolution (Yes in step S150), the grouped convolution processing optimization device 100 executes the grouped convolution optimization process (step S160).
[0103] Next, the optimizable structure detection unit 110 determines whether all the grouped convolutions in the structure of the obtained CNN model have been confirmed (step S170). If there is an unconfirmed grouped convolution among the grouped convolutions in the structure of the obtained CNN model (No in step S170), the optimizable structure detection unit 110 executes the process of step S130 again.
[0104] When all the grouped convolutions in the structure of the obtained CNN model have been confirmed (Yes in step S170), the forward convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the backward convolution weight correction unit 150 store the optimized learned CNN model in the modified CNN model storage unit 300, including each corrected weight (step S180). After storing, the grouped convolution process optimization device 100 ends the optimization process.
[0105] Next, the grouped convolution optimization process in step S160, which is a sub-process constituting the optimization process shown in FIG. 6, will be described with reference to FIG. 7. FIG. 7 is a flowchart showing the operation of the grouped convolution optimization process by the grouped convolution process optimization device 100 according to the present disclosure.
[0106] First, the optimal channel number determination unit 120 determines M, which is the minimum value among the values greater than C / G or C / G of the original grouped convolution and suitable for the device (step S161). in / G or C out / G and is suitable for the device (step S161).
[0107] Next, the optimal channel number determination unit 120 sets the new input channel number C' of the grouped convolution of the optimized CNN model to C' = M * G (step S162). in ’ to C in ’ = M * G (step S162).
[0108] Next, the optimal channel number determination unit 120 sets the new output channel number C' of the grouped convolution of the optimized CNN model to C' = M * G (step S163). out ’ to C out ’ = M * G (step S163).
[0109] Note that, as will be described later, the optimal channel number determination unit 120 executes at least one of the processes in step S162 and step S163.
[0110] Next, the front convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the back convolution weight correction unit 150 set the number of input channels of the grouped convolution to C in ’, and set the number of output channels to C out ’ and respectively generate a new model structure (step S164).
[0111] Next, the front convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the back convolution weight correction unit 150 respectively initialize the weight values of the generated model structure to zero (step S165).
[0112] Next, the front convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the back convolution weight correction unit 150 copy parameters such as weight values from the original CNN model as they are for the layers unrelated to the optimization process of the generated model structure (step S166).
[0113] Next, the front convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the back convolution weight correction unit 150 copy the original weight values to the weights of the new model structure for the layers to be optimized in the generated model structure (step S167).
[0114] In particular, the grouped convolution weight correction unit 140 copies the original weight values for each weight matrix. As described above, the front convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the back convolution weight correction unit 150 realize the insertion process shown in FIG. 5 by copying the original weight values to the weights of the new model structure.
[0115] After copying the weight values, the grouped convolution processing optimization device 100 returns to the optimization process shown in FIG. 6.
[0116] Note that the grouped convolution process optimization device 100 according to this embodiment may modify only any one of the weights used in the convolutions before and after the grouped convolution. FIG. 8 is an explanatory diagram showing an example in which grouped convolution and the convolution existing before the grouped convolution are executed as a set.
[0117] In the example shown in FIG. 8, the convolution existing after the grouped convolution is omitted. The example shown in FIG. 8 is an example in which the optimal channel number determination unit 120 determines that since only the input channel number C in affects the performance and thus needs to satisfy the constraint, and since the output channel number C out does not affect the performance and thus does not need to satisfy the constraint.
[0118] In the above case, as shown in FIG. 8, the forward convolution weight modification unit 130 modifies the weights in the same manner as the example shown in FIG. 5. Also, the backward convolution weight modification unit 150 does not modify the weights.
[0119] Also, the grouped convolution weight modification unit 140 generates the 32-row 20-column weights by inserting 12 (= 32 - 20) rows of rows with zero components at the lower end of the 20-row 20-column weights shown in FIG. 3. The grouped convolution weight modification unit 140 executes the above zero-row insertion process for each of the eight 20-row 20-column weights shown in FIG. 3.
[0120] FIG. 9 is an explanatory diagram showing an example in which grouped convolution and the convolution existing after the grouped convolution are executed as a set. In the example shown in FIG. 9, the convolution existing before the grouped convolution is omitted. The example shown in FIG. 9 is an example in which the optimal channel number determination unit 120 determines that since only the output channel number C out affects the performance and thus needs to satisfy the constraint, and since the input channel number C in does not affect the performance and thus does not need to satisfy the constraint.
[0121] In the above case, as shown in FIG. 9, the backward convolution weight correction unit 150 corrects the weights in the same manner as the example shown in FIG. 5. Also, the forward convolution weight correction unit 130 does not correct the weights.
[0122] Also, the grouped convolution weight correction unit 140 generates the weights of 20 rows and 32 columns by inserting 12 (= 32 - 20) columns with zero components at the right end of the weights of 20 rows and 20 columns shown in FIG. 3. The grouped convolution weight correction unit 140 executes the above-described zero-column insertion process for each of the eight 20-row and 20-column weights shown in FIG. 3.
[0123] As described above, the grouped convolution weight correction unit 140 of the present embodiment is for an input convolution in which an N×N input weight matrix is used to perform a convolution calculation on data in which the first channel to the Nth channel (N is an integer of 2 or more) are arranged in order, a grouped convolution in which the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and a convolution calculation is performed for each of the data composed of the {(i - 1)×N / G + 1}th channel to the (i×N / G)th channel (i = 1 to G) using the i-th weight matrix of N / G rows and N / G columns for i = 1 to G, and an output convolution in which a convolution calculation is performed using an N×N output weight matrix for the result of the grouped convolution. Regarding a learned convolutional neural network in which each is defined, when a constraint is imposed such that a convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, at least one of a column insertion process of inserting (M - N / G) columns with zero components at the right end of the i-th weight matrix for each i = 1 to G and a row insertion process of inserting (M - N / G) rows with zero components at the bottom end of the i-th weight matrix for each i = 1 to G is executed.
[0124] In addition, the forward convolution weight modification unit 130 of the present embodiment executes an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / G-th column of the input weight matrix for each i = 1 to G. Further, the backward convolution weight modification unit 150 executes an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G.
[0125] In the grouped convolution process optimization device 100 of the present embodiment, at least one of the input weight matrix insertion process and the output weight matrix insertion process is executed. Further, the forward convolution weight modification unit 130 executes the input weight matrix insertion process when the row insertion process is executed. Further, the backward convolution weight modification unit 150 executes the output weight matrix insertion process when the column insertion process is executed.
[0126] In addition, M is the smallest multiple (for example, 32) among multiples larger than a predetermined number (for example, 16) of N / G. Further, in grouped convolution, a matrix in which the i-th weight matrix is arranged diagonally in the order of i = 1 to G from the upper left to the lower right may be used as the weight matrix.
[0127] In addition, the optimal channel number determination unit 120 of the present embodiment determines at least one of the input channel number, which is the channel number of the data input as the target of the constrained grouped convolution, and the output channel number, which is the channel number of the data output as the result of the grouped convolution, to be M*G.
[0128] The grouped convolution weight modification unit 140 executes the row insertion process when the input channel number is determined to be M*G, and executes the column insertion process when the output channel number is determined to be M*G.
[0129] In addition, the optimizable structure detection unit 110 of the present embodiment determines whether the input convolution, the grouped convolution, and the output convolution are respectively defined in the learned convolutional neural network.
[0130] [Description of Effects] Some AI chips have constraints on memory addressing, resulting in slow execution speeds for grouped convolutions. The forward convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the backward convolution weight correction unit 150 of the grouped convolution processing optimization device 100 according to the present embodiment adjust the weight matrix used in the convolution layer that constitutes the CNN model.
[0131] By adjusting the weight matrix, the grouped convolution processing optimization device 100 converts the processing executed by the CNN model into processing for an AI chip with constraints. When converted to processing for an AI chip, although the amount of calculation increases, the processing that the AI chip is good at is replaced, so the AI chip can execute grouped convolutions at high speed. That is, the grouped convolution processing optimization device 100 enables the acceleration of grouped convolution processing by an AI chip with constraints without changing the processing result.
[0132] Hereinafter, a specific example of the hardware configuration of the grouped convolution processing optimization device 100 according to the present embodiment will be described. FIG. 10 is an explanatory diagram showing an example of the hardware configuration of the grouped convolution processing optimization device 100 according to the present disclosure.
[0133] The grouped convolution processing optimization device 100 shown in FIG. 10 includes a CPU 11, a main memory unit 12, a communication unit 13, and an auxiliary storage unit 14. It also includes an input unit 15 for the user to operate and an output unit 16 for presenting the processing result or the progress of the processing content to the user.
[0134] The grouped convolution processing optimization device 100 is realized by software by the CPU 11 shown in FIG. 10 executing a program that provides the functions of each component.
[0135] That is, the CPU 11 loads the program stored in the auxiliary storage unit 14 into the main storage unit 12 and executes it, and controls the operation of the grouped convolution process optimization device 100, whereby each function is realized by software.
[0136] Note that the grouped convolution process optimization device 100 shown in FIG. 10 may include a DSP (Digital Signal Processor) instead of the CPU 11. Alternatively, the grouped convolution process optimization device 100 shown in FIG. 10 may include both the CPU 11 and the DSP.
[0137] The main storage unit 12 is used as a work area for data and a temporary storage area for data. The main storage unit 12 is, for example, a RAM (Random Access Memory).
[0138] The communication unit 13 has a function of inputting and outputting data to and from peripheral devices via a wireless network (information communication network).
[0139] The auxiliary storage unit 14 is a non-temporary tangible storage medium. Examples of non-temporary tangible storage media include magnetic disks, magneto-optical disks, CD-ROMs (Compact Disk Read Only Memories), DVD-ROMs (Digital Versatile Disk Read Only Memories), and semiconductor memories.
[0140] The input unit 15 has a function of inputting data and processing instructions. The input unit 15 is, for example, an input device such as a keyboard, a mouse, or a touch panel.
[0141] The output unit 16 has a function of outputting data. The output unit 16 is, for example, a display device such as a liquid crystal display device, a touch panel, or a printing device such as a printer.
[0142] Also, as shown in FIG. 10, in the grouped convolution process optimization device 100, each component is connected to the system bus 17.
[0143] In the grouped convolution process optimization device 100, the auxiliary storage unit 14 stores a program for realizing the optimizable structure detection unit 110, the optimal channel number determination unit 120, the forward convolution weight correction unit 130, the grouped convolution weight correction unit 140, and the backward convolution weight correction unit 150.
[0144] Note that the grouped convolution process optimization device 100 may be implemented with a circuit including hardware components such as an LSI (Large Scale Integration) that realizes functions as shown in FIG. 1 inside, for example.
[0145] Also, the grouped convolution process optimization device 100 may be realized by hardware that does not include a computer function using elements such as a CPU. For example, some or all of each component may be realized by general-purpose circuitry or dedicated circuitry, a processor, etc., or a combination thereof. These may be configured by a single chip (for example, the above LSI), or may be configured by a plurality of chips connected via a bus. Some or all of each component may be realized by a combination of the above-described circuitry, etc. and a program.
[0146] Also, some or all of each component of the grouped convolution process optimization device 100 may be configured by one or a plurality of information processing devices having an arithmetic unit and a storage unit.
[0147] When some or all of each component is realized by a plurality of information processing devices, circuitry, etc., the plurality of information processing devices, circuitry, etc. may be centrally arranged or may be distributed. For example, the information processing devices, circuitry, etc. may be realized in a form in which each is connected via a communication network, such as a client and server system, a cloud computing system, etc.
[0148] Next, an overview of the present disclosure will be described. FIG. 11 is a block diagram showing an overview of a grouped convolution processing optimization apparatus according to the present disclosure. The grouped convolution processing optimization apparatus 20 according to the present disclosure includes an input convolution in which an input weight matrix of N rows and N columns is used to perform convolution calculation on data configured by arranging the first channel to the Nth channel (N is an integer of 2 or more) in order, and the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and for data configured by the {(i - 1)×N / G + 1}th channel to the (i×N / G)th channel (i = 1 to G) after division, a grouped convolution in which a weight matrix of the i-th of N / G rows and N / G columns is used to perform convolution calculation for each i = 1 to G, and an output convolution in which an output weight matrix of N rows and N columns is used to perform convolution calculation on the result of the grouped convolution are respectively defined. Regarding the learned convolutional neural network, when a constraint is imposed such that convolution calculation is performed for each data configured by M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each i = 1 to G, and at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the bottom end of the i-th weight matrix for each i = 1 to G are executed by a first insertion unit 21 (for example, a grouped convolution weight correction unit 140), an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components on the right side of the i×N / G-th column of the input weight matrix for each i = 1 to G, and an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G are executed by a second insertion unit 22 (for example, a forward convolution weight correction unit 130 and a backward convolution weight correction unit 150). The second insertion unit 22 executes the input weight matrix insertion process when the row insertion process is executed, and executes the output weight matrix insertion process when the column insertion process is executed.
[0149] When a grouped convolution process optimization device with such a configuration is used, a hardware circuit with restrictions imposed on memory addressing can execute grouped convolution at high speed.
[0150] Also, M may be the smallest multiple among multiples greater than a predetermined number of N / G.
[0151] Also, in grouped convolution, a matrix in which the i-th weight matrix is arranged diagonally in the order of i = 1 to G from the upper left to the lower right may be used as the weight matrix.
[0152] With such a configuration, the grouped convolution process optimization device can reduce the influence of overhead associated with the execution of grouped convolution.
[0153] Also, the grouped convolution process optimization device 20 includes a determination unit (for example, the optimal channel number determination unit 120) that determines at least one of the input channel number, which is the number of channels of the data input as the target of the constrained grouped convolution, and the output channel number, which is the number of channels of the data output as the result of the grouped convolution, to be M*G. The first insertion unit 21 may execute row insertion processing when the input channel number is determined to be M*G, and execute column insertion processing when the output channel number is determined to be M*G.
[0154] With such a configuration, the grouped convolution process optimization device can adjust the weights to be modified.
[0155] Also, the grouped convolution process optimization device 20 may include a determination unit (for example, the optimizable structure detection unit 110) that determines whether input convolution, grouped convolution, and output convolution are respectively defined in the learned convolutional neural network.
[0156] With such a configuration, the grouped convolution process optimization device can determine whether to modify the weights used in each convolution.
[0157] Also, some or all of the above embodiments can be described as follows, but are not limited thereto.
[0158] (Appendix 1) Regarding a trained convolutional neural network in which an input convolution is performed using an N×N input weight matrix for data configured by arranging the first channel to the Nth channel (N is an integer of 2 or more) in order, a grouped convolution in which the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and a convolution calculation is performed for each of the data configured by the {(i - 1)×N / G + 1}th channel to the (i×N / G)th channel (i = 1 to G) using an N / G×N / G weight matrix for the ith weight matrix for i = 1 to G, and an output convolution in which a convolution calculation is performed using an N×N output weight matrix for the result of the grouped convolution, when a constraint is imposed such that a convolution calculation is performed for each data configured by M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the ith weight matrix for each i = 1 to G, and a first insertion unit that executes at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the lower end of the ith weight matrix for each i = 1 to G, an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / Gth column of the input weight matrix for each i = 1 to G, and a second insertion unit that executes at least one of an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / Gth row of the output weight matrix for each i = 1 to G, the second insertion unit executes the input weight matrix insertion process when the row insertion process is executed, and executes the output weight matrix insertion process when the column insertion process is executed A grouped convolution processing optimization device characterized by the above.
[0159] (Appendix 2) M is the smallest multiple that is a multiple greater than a predetermined number of N / G The grouped convolution process optimization device described in Appendix 1
[0160] (Appendix 3) In grouped convolution, a matrix in which the i-th weight matrix is arranged diagonally in the order of i = 1 to G from the upper left to the lower right is used as the weight matrix The grouped convolution process optimization device described in Appendix 1 or Appendix 2
[0161] (Appendix 4) A determination unit that determines at least one of the number of input channels, which is the number of channels of the data input as the target of the constrained grouped convolution, and the number of output channels, which is the number of channels of the data output as the result of the grouped convolution, to be M*G The first insertion unit When the number of input channels is determined to be M*G, perform row insertion processing When the number of output channels is determined to be M*G, perform column insertion processing The grouped convolution process optimization device described in any one of Appendix 1 to Appendix 3
[0162] (Appendix 5) Comprising a determination unit that determines whether input convolution, grouped convolution, and output convolution are respectively defined in a learned convolutional neural network The grouped convolution process optimization device described in any one of Appendix 1 to Appendix 4
[0163] (Appendix 6) For an input convolution in which a convolution calculation is performed using an N×N input weight matrix on data composed of the first channel to the Nth channel (N is an integer of 2 or more) arranged in order, and for a grouped convolution in which the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and a convolution calculation is performed for each of the data composed of the {(i - 1)×N / G + 1}th channel to the (i×N / G)th channel (i = 1 to G) using an N / G×N / G weight matrix for the i-th time for i = 1 to G, and for an output convolution in which a convolution calculation is performed using an N×N output weight matrix on the result of the grouped convolution, in a learned convolutional neural network in which each is defined, when a constraint is imposed that a convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each i = 1 to G, and at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the bottom end of the i-th weight matrix for each i = 1 to G is executed, an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / G-th column of the input weight matrix for each i = 1 to G, and at least one of an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G is executed, when the row insertion process is executed, the input weight matrix insertion process is executed, when the column insertion process is executed, the output weight matrix insertion process is executed A method for optimizing grouped convolution processing, characterized by the above.
[0164] (Appendix 7) M is the smallest multiple among multiples larger than a predetermined number of N / G The method for optimizing grouped convolution processing according to Appendix 6.
[0165] (Appendix 8) In grouped convolution, a matrix in which the i-th weight matrix is arranged diagonally from the upper left to the lower right in the order of i = 1 to G is used as the weight matrix. The grouped convolution process optimization method described in Appendix 6 or Appendix 7.
[0166] (Appendix 9) Determine at least one of the input channel number, which is the number of channels of the data input as the target of the constrained grouped convolution, and the output channel number, which is the number of channels of the data output as the result of the grouped convolution, to be M * G. Execute row insertion processing when the input channel number is determined to be M * G. Execute column insertion processing when the output channel number is determined to be M * G. The grouped convolution process optimization method described in any one of Appendices 6 to 8.
[0167] (Appendix 10) Determine whether input convolution, grouped convolution, and output convolution are respectively defined in the pre-trained convolutional neural network. The grouped convolution process optimization method described in any one of Appendices 6 to 9.
[0168] (Appendix 11) To the computer For an input convolutional operation in which an N×N input weight matrix is used to perform a convolution calculation on data composed of the first channel to the Nth channel (N is an integer of 2 or more) arranged in order, and for a grouped convolution operation in which the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and an N / G×N / G i-th weight matrix is used to perform a convolution calculation for each of the data composed of the {(i - 1)×N / G + 1}-th channel to the (i×N / G)-th channel (i = 1 to G) for i = 1 to G, and an output convolutional operation in which an N×N output weight matrix is used to perform a convolution calculation on the result of the grouped convolution, for a learned convolutional neural network in which each is defined, when a constraint is imposed such that a convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each of i = 1 to G, and a first insertion process of performing at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the bottom end of the i-th weight matrix for each of i = 1 to G, and An input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / G-th column of the input weight matrix for each of i = 1 to G, and a grouped convolution process optimization program for executing at least one of an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each of i = 1 to G, In the second insertion process, when the row insertion process is executed, the input weight matrix insertion process is executed, when the column insertion process is executed, the output weight matrix insertion process is executed Grouped convolution process optimization program.
[0169] (Appendix 12) M is the smallest multiple among multiples larger than a predetermined number of N / G The grouped convolution processing optimization program described in Supplementary Note 11.
[0170] (Supplementary Note 13) In grouped convolution, a matrix in which the i-th weight matrix is arranged diagonally in the order of i = 1 to G from the upper left to the lower right is used as the weight matrix. The grouped convolution processing optimization program described in Supplementary Note 11 or Supplementary Note 12.
[0171] (Supplementary Note 14) Cause the computer to Execute a determination process to determine at least one of the number of input channels, which is the number of channels of the data input as the target of the grouped convolution with constraints, and the number of output channels, which is the number of channels of the data output as the result of the grouped convolution, to be M * G. In the first insertion process, Execute row insertion processing when the number of input channels is determined to be M * G. Execute column insertion processing when the number of output channels is determined to be M * G. The grouped convolution processing optimization program described in any one of Supplementary Notes 11 to 13.
[0172] (Supplementary Note 15) Cause the computer to Execute a determination process to determine whether input convolution, grouped convolution, and output convolution are respectively defined in the learned convolutional neural network. The grouped convolution processing optimization program described in any one of Supplementary Notes 11 to 14.
[0173] As described above, the present disclosure has been described with reference to the embodiments, but the present disclosure is not limited to the above embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure. And each embodiment can be combined with other embodiments as appropriate.
Explanation of Reference Numerals
[0174] 11 CPU 12 Main memory unit 13 Communication unit 14 Auxiliary memory unit 15 Input unit 16 Output unit 17 System bus 20, 100 Grouped convolution processing optimization device 21 First insertion unit 22 Second insertion unit 110 Optimizable structure detection unit 120 Optimal channel number determination unit 130 Forward convolution weight correction unit 140 Grouped convolution weight correction unit 150 Backward convolution weight correction unit 200 Pre-change CNN model memory unit 300 Post-change CNN model memory unit 400 AI chip
Claims
1. For input convolution in which a convolution calculation is performed using an input weight matrix of N rows and N columns on data configured by arranging the first channel to the Nth channel (N is an integer of 2 or more) in order, and the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and for data configured by the {(i - 1)×N / G + 1}th channel to the (i×N / G)th channel (i = 1 to G) thus divided, a convolution calculation is performed using the i-th weight matrix of N / G rows and N / G columns for each i = 1 to G in grouped convolution, and for output convolution in which a convolution calculation is performed using an output weight matrix of N rows and N columns on the result of the grouped convolution, in a learned convolutional neural network in which each is defined, when a constraint is imposed such that a convolution calculation is performed for each data composed of M channels (M is an integer greater than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each i = 1 to G, and a first insertion unit that executes at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the lower end of the i-th weight matrix for each i = 1 to G, An input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / G-th column of the input weight matrix for each i = 1 to G, and a second insertion unit that executes at least one of an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G, The second insertion unit is, When the row insertion process is executed, the input weight matrix insertion process is executed, When the column insertion process is executed, the output weight matrix insertion process is executed A grouped convolution process optimization device characterized by this.
2. M is the smallest multiple among multiples greater than a predetermined number of N / G The grouped convolution process optimization device according to Claim 1.
3. In grouped convolution, a matrix in which the i-th weight matrix is arranged diagonally from the upper left to the lower right in the order of i = 1 to G is used as the weight matrix. The grouped convolution process optimization device according to claim 1.
4. A determination unit that determines at least one of the number of input channels, which is the number of channels of the data input as the target of the constrained grouped convolution, and the number of output channels, which is the number of channels of the data output as a result of the grouped convolution, to be M*G. The first insertion unit executes row insertion processing when the number of input channels is determined to be M*G, and executes column insertion processing when the number of output channels is determined to be M*G. The grouped convolution process optimization device according to claim 1.
5. It includes a determination unit that determines whether input convolution, grouped convolution, and output convolution are respectively defined in a pre-trained convolutional neural network. The grouped convolution process optimization device according to any one of claims 1 to 4.
6. For an input convolutional operation in which an N×N input weight matrix is used to perform a convolution calculation on data composed of the first channel to the Nth channel (N is an integer of 2 or more) arranged in order, and for a grouped convolution in which the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction and an N / G×N / G weight matrix for the i-th is used to perform a convolution calculation for each of the data composed of the {(i - 1)×N / G + 1}-th channel to the (i×N / G)-th channel (i = 1 to G) for i = 1 to G, and for an output convolution in which an N×N output weight matrix is used to perform a convolution calculation on the result of the grouped convolution, respectively defined, regarding a learned convolutional neural network, when a constraint is imposed such that a convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each i = 1 to G, and at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the bottom end of the i-th weight matrix for each i = 1 to G is executed, an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / G-th column of the input weight matrix for each i = 1 to G, and at least one of an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G is executed, when the row insertion process is executed, the input weight matrix insertion process is executed, when the column insertion process is executed, the output weight matrix insertion process is executed A grouped convolution process optimization method characterized by this.
7. On a computer, For an input convolution in which a convolution calculation is performed using an input weight matrix of N rows and N columns for data configured by arranging the first channel to the Nth channel (N is an integer of 2 or more) in order, and for a grouped convolution in which the result of the input convolution is divided into G (G is an integer of 2 or more) in the channel direction, and a convolution calculation is performed for each of the data configured by the {(i - 1)×N / G + 1}th channel to the (i×N / G)th channel (i = 1 to G) using the i-th weight matrix of N / G rows and N / G columns, and for an output convolution in which a convolution calculation is performed using an output weight matrix of N rows and N columns for the result of the grouped convolution, for a learned convolutional neural network in which each is defined, when a constraint is imposed that a convolution calculation is performed for each data composed of M channels (M is an integer larger than N / G) for the grouped convolution, a column insertion process of inserting (M - N / G) columns of columns with zero components at the right end of the i-th weight matrix for each i = 1 to G, and a first insertion process of performing at least one of a row insertion process of inserting (M - N / G) rows of rows with zero components at the bottom end of the i-th weight matrix for each i = 1 to G, and A grouped convolution process optimization program for causing execution of a second insertion process of performing at least one of an input weight matrix insertion process of inserting (M - N / G) columns of columns with zero components to the right of the i×N / G-th column of the input weight matrix for each i = 1 to G, and an output weight matrix insertion process of inserting (M - N / G) rows of rows with zero components below the i×N / G-th row of the output weight matrix for each i = 1 to G, In the second insertion process, When the row insertion process is executed, the input weight matrix insertion process is caused to be executed, When the column insertion process is executed, the output weight matrix insertion process is caused to be executed Grouped convolution process optimization program.