Convolution layer conversion device, convolution layer conversion method, and program
By decomposing large convolutional kernels into smaller ones and aggregating their results, the method enhances the execution speed of neural network convolutional layers, addressing the speed limitations of large kernels and optimizing performance.
Patent Information
- Application Number
- JP2023574933
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-01-19
AI Technical Summary
The execution speed of convolutional layers in neural networks using large kernels (e.g., 7×7 or 5×5) is slower due to optimization issues and design limitations, leading to reduced performance compared to smaller kernels like 3×3.
A convolutional layer conversion method that decomposes large kernels into smaller kernels and aggregates their results, utilizing optimization methods like Winograd optimization and sparsity-utilizing acceleration mechanisms to enhance execution speed.
This approach significantly improves the execution speed of convolutional layers by leveraging optimized circuits and algorithms for smaller kernels, thereby increasing computational efficiency.
Smart Images

Figure 0007700880000001 
Figure 0007700880000002 
Figure 0007700880000003
Abstract
Description
Technical Field
[0001] The present invention relates to a convolutional layer conversion device, a convolutional layer conversion method, and a program.
Background Art
[0002] In the convolutional layer used in a neural network (NN) model, kernels of various sizes are used. Recently, the use of kernels of sizes 1×1 and 3×3 has been mainstream, and the use of kernels of sizes 7×7 and 5×5 has tended to be rare. Also, a 7×7 size kernel or a 5×5 size kernel in one convolutional layer tends to be used instead by a plurality of consecutive 3×3 size kernels. However, for example, when changing the structure to two 3×3 size kernels instead of one 7×7 size kernel, although they may be considered semantically close in some cases, the calculation content and results are often not equivalent.
[0003] On the other hand, there may be cases where a 7×7 size kernel is used in the convolutional layer. This is because a larger size kernel may be easier to learn, and in the case of using an old neural network with well-known achievement accuracy, kernels of sizes 7×7 and 5×5 are used.
[0004] Patent Document 1 relates to an information processing device that efficiently performs generation processing of neighboring matrix image data for convolutional operations.
[0005] Patent Document 2 relates to an apparatus for detecting mutant malicious codes based on neural network learning.
[0006] Patent Document 3 relates to a DNN quantization device that enables efficient quantization of convolutional layers included in a CNN.
[0007] Patent Document 4 relates to a device for generating a learning model of a neural network.
[0008] Patent Document 5 relates to a neural network device.
Prior Art Documents
Patent Documents
[0009]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0010] The following analysis is provided by the present invention.
[0011] However, when using a kernel with a size of 7×7 or a kernel with a size of 5×5 in the convolutional layer, the execution speed when implementing the convolutional layer may be slow. That is, for a kernel with a small size such as 3×3, a speed reduction in implementation may occur more than the increase in the amount of calculation due to using a kernel with a size of 7×7 or a kernel with a size of 5×5. This is due to, for example, the degree of optimization of the kernel, such as that a simple and well-known 3×3 kernel has a higher degree of optimization, or problems in the design of the device to be implemented or the software library.
[0012] An object of the present invention is to provide a convolutional layer conversion device, a convolutional layer conversion method, and a program that contribute to improving the execution speed when implementing a convolutional layer of a neural network model.
Means for Solving the Problem
[0013] According to a first aspect of the present invention, a convolutional layer detection unit that detects a convolutional layer including a large kernel having a kernel size equal to or larger than a predetermined size from an input neural network model structure, and the convolutional layer including the large kernel is decomposed into a combination of a plurality of small kernels having a kernel size smaller than the predetermined size, and a convolutional layer including the combination of the plurality of small kernels, and an aggregated convolutional layer that aggregates the convolution results of the convolutional layer including the combination of the plurality of small kernels, and outputs a neural network model structure in which the convolutional layer including the large kernel is converted. A convolutional layer decomposition unit can be provided.
[0014] According to a second aspect of the present invention, a step of detecting a convolutional layer including a large kernel having a kernel size equal to or larger than a predetermined size from an input neural network model structure, which is executed by a computer including a processor and a storage device, and the convolutional layer including the large kernel is decomposed into a combination of a plurality of small kernels having a kernel size smaller than the predetermined size, and a convolutional layer including the combination of the plurality of small kernels, and a step of outputting a neural network model structure in which the convolutional layer including the large kernel is converted by aggregating the convolution results of the convolutional layer including the combination of the plurality of small kernels. A convolutional layer conversion method can be provided.
[0015] According to a third aspect of the present invention, a computer is caused to execute a process of detecting a convolutional layer including a large kernel having a kernel size equal to or greater than a predetermined size from an input neural network model structure, converting the convolutional layer including the large kernel into a convolutional layer including a combination of a plurality of small kernels having a kernel size smaller than the predetermined size obtained by decomposing the large kernel, and an aggregated convolutional layer that aggregates the convolution results of the convolutional layer including the combination of the plurality of small kernels, and outputting a neural network model structure in which the convolutional layer including the large kernel is converted. Note that this program can be recorded on a computer-readable storage medium. The storage medium can be a non-transient one such as a semiconductor memory, a hard disk, a magnetic recording medium, or an optical recording medium. The present invention can also be embodied as a computer program product.
Advantages of the Invention
[0016] According to the present invention, it is possible to provide a convolutional layer conversion device, a convolutional layer conversion method, and a program that contribute to improving the execution speed at the time of implementing the convolutional layer of a neural network model.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Embodiments for Carrying Out the Invention
[0018] First, an overview of an embodiment of the present invention will be described with reference to the drawings. Note that the reference numerals in the drawings appended to this overview are for convenience of each element as an example to assist understanding, and are not intended to limit the present invention to the illustrated embodiments. Also, the connection lines between blocks in the drawings and the like referred to in the following description include both bidirectional and unidirectional ones. For the one-way arrow, it schematically shows the flow of the main signal (data) and does not exclude bidirectionality.
[0019] FIG. 1 is a diagram showing an example of the configuration of a convolutional layer conversion device according to an embodiment of the present invention. Referring to FIG. 1, a convolutional layer conversion device 100 according to an embodiment of the present invention includes a large kernel convolution (Conv) layer detection unit 110 (convolutional layer detection unit) and a convolutional layer decomposition unit 120. The large kernel convolution (Conv) layer detection unit 110 detects a convolutional layer including a large kernel having a kernel size of a predetermined size or more from the input neural network (NN) model structure 10. The convolutional layer decomposition unit 120 converts a convolutional layer including a large kernel into a convolutional layer including a combination of a plurality of small kernels having a kernel size smaller than a predetermined size obtained by decomposing the large kernel and an aggregation convolutional layer that aggregates the convolutional results of the convolutional layer including the combination of the plurality of small kernels, and outputs a neural network model structure 20 in which the convolutional layer including the large kernel is converted.
[0020] In this way, since the convolutional layer including the large kernel is converted into a convolutional layer and an aggregated convolutional layer including a combination of a plurality of small kernels obtained by decomposing the large kernel, for the convolution of the input data and each small kernel, for example, optimization methods for convolution with a kernel of size 3×3, such as winograd optimization that can be accelerated at double speed, can be utilized. Further, circuits and implementations for high-speed execution of the convolutional layer, such as a hardware circuit and a software library optimized for convolution with a kernel of size 3×3, can be maximally utilized. Furthermore, if a sparsity-utilizing acceleration mechanism that skips multiplication by zero values can be used, it can also be utilized. Thereby, the execution speed of convolution can be increased.
[0021] As described above, according to an embodiment of the present invention, a convolutional layer conversion device that contributes to improving the execution speed during the implementation of the convolutional layer of a neural network model can be provided.
[0022] [First Embodiment] Next, the convolutional layer conversion device according to the first embodiment of the present invention will be described with reference to the drawings. FIG. 2 is a diagram showing an example of the configuration of the convolutional layer conversion device according to the first embodiment of the present invention. In FIG. 2, components denoted by the same reference numerals as in FIG. 1 are assumed to be the same components, and the description thereof will be omitted.
[0023] Referring to FIG. 2, the convolutional layer conversion device 100 according to the first embodiment of the present invention includes a large kernel convolution (Conv) layer detection unit 110 (convolutional layer detection unit) and a convolutional layer decomposition unit 120. The convolutional layer decomposition unit 120 includes a decomposition method selection unit 121 and a layer decomposition application unit 122. The convolutional layer conversion device 100 takes a neural network (NN) model structure 10 as an input and outputs a converted neural network model structure 20. Further, a target device information storage unit 30 is connected to the decomposition method selection unit 121 of the convolutional layer decomposition unit 120, and target device information is input thereto.
[0024] The large kernel convolutional layer detector 110 detects a convolutional layer including a large kernel having a kernel size equal to or larger than a predetermined size from the input neural network (NN) model structure 10. FIG. 3 is a diagram showing an example of a large kernel according to the first embodiment of the present invention and small kernels having a plurality of small kernel sizes obtained by decomposing the large kernel.
[0025] In FIG. 3, it is assumed that the input neural network (NN) model structure 10 includes a convolutional layer including a kernel having a kernel size of 7×7 (hereinafter referred to as a large kernel 300). Here, when the predetermined size is the 7×7 kernel size, the large kernel convolutional layer detector 110 detects a convolutional layer including the large kernel 300 having a kernel size of 7×7 from the input neural network (NN) model structure 10.
[0026] First, the convolution of the convolutional layer including the large kernel 300 of the neural network model structure 10 without performing the decomposition of the convolutional layer according to the present invention will be described.
[0027] FIG. 4 is a diagram showing an example of convolution between input data 400 and a convolution layer of a large kernel 300 without performing decomposition of the convolution layer according to the present invention. Assume that the input data 400 is two-dimensional data including elements from 1 to 100. Note that the numbers from 1 to 100 on the input data 400 indicate positions on the input data 400, not the values of each element. As an example, the upper left corner is 1, increasing in the rightward and downward directions, and becoming 100 at the lower right corner. Convolution is performed between the data located in the range 405 (positions 12 to 18, 22 to 28, 32 to 38, 42 to 48, 52 to 58, 62 to 68, 72 to 78) on the input data 400 and the large kernel 300. In FIG. 4, a symbol having an "x" inside a "〇" represents convolution. The same shall apply to other figures. The result of the convolution shown in FIG. 4 is output to position 45 on the output data 450 corresponding to position 45 on the input data 400 that is multiplied by the center 301 of the large kernel 300. That is, as an example, the convolution between the input data 400 and the large kernel 300 is calculated as an output corresponding to the center 301 of the large kernel 300 while moving the center 301 of the large kernel 300 from position 1 to position 100 on the input data 400.
[0028] Next, the operation of the decomposition method selection unit 121 of the convolutional layer decomposition unit 120 of the convolutional layer conversion device 100 according to the first embodiment of the present invention will be described. The decomposition method selection unit 121 selects a decomposition method for the large kernel 300 of the convolutional layer including the detected large kernel 300 (kernel with a kernel size of 7×7). The decomposition method is selected based on the target device information in the target device information storage unit 30. The target device information in the target device information storage unit 30 includes execution speed information indicating the execution speed when a convolutional layer (decomposition candidate decomposed by the decomposition method) including a combination of a plurality of small kernels is executed on the target device, or memory usage information indicating the memory usage amount when a convolutional layer (decomposition candidate decomposed by the decomposition method) including a combination of a plurality of small kernels is executed on the target device. Since the memory during execution may increase when the large kernel is decomposed into a plurality of small kernels, and it may be necessary to keep the memory usage amount constant, in addition to the execution speed information, considering the memory usage information and selecting the decomposition method can keep the memory usage amount constant and increase the execution speed.
[0029] In the first embodiment of the present invention, as an example, according to the target device information, a decomposition method is selected in which the large kernel 300 is decomposed into a plurality of small kernels with a kernel size smaller than 7×7, which is a predetermined size, as shown in FIG. 3.
[0030] Next, an example of the operation of the layer decomposition application unit 122 of the convolutional layer decomposition unit 120 will be described. The layer decomposition application unit 122 decomposes the convolutional layer including the large kernel 300 into a plurality of small kernels 310, 320, 330, 340 as shown in FIG. 3 according to the decomposition method selected by the decomposition method selection unit 121 for the detected large kernel 300.
[0031] Referring to FIG. 3, the large kernel 300 is a kernel with a kernel size of 7×7 having a center 301. The layer decomposition application unit 122 decomposes the large kernel 300 into a small kernel 310 with a kernel size of 4×4 having a center 311, a small kernel 320 with a kernel size of 3×3 having a center 321, a small kernel 330 with a kernel size of 3×3 having a center 331, and a small kernel 340 with a kernel size of 4×4 having a center 341. Note that for the small kernel 340, so that its elements do not overlap with those of the small kernel 310, either the upper left corner part does not exist or 0 is assigned. Note that the decomposition method for decomposing the large kernel of the present invention into a plurality of small kernels is not limited to the above decomposition method.
[0032] Next, the layer decomposition application unit 122 outputs a neural network model structure 20 in which a convolutional layer including the large kernel 300 is converted into a convolutional layer including a convolutional layer including a combination of the decomposed plurality of small kernels 310, 320, 330, 340 and an aggregation convolutional layer for aggregating the results of the convolutional layer including the combination of the plurality of small kernels.
[0033] FIG. 5 is a diagram showing an example of the configuration of a converted convolutional layer in which a convolutional layer including the large kernel 300 of the first embodiment of the present invention includes a convolutional layer 500 including a combination of a plurality of small kernels 310, 320, 330, 340 and an aggregation convolutional layer 550 for aggregating the results of the convolutional layer including the combination of the plurality of small kernels 310, 320, 330, 340.
[0034] Next, the configuration and operation of the convolutional layer 500 including a combination of the plurality of small kernels 310, 320, 330, 340 and the aggregation convolutional layer 550 will be described with reference to the drawings.
[0035] Note that the convolution result by the input data 400, the convolutional layer 500 including a combination of the plurality of small kernels, and the aggregation convolutional layer 550 coincides with the convolution result of the input data 400 and the large kernel 300 except for a part of the peripheral portion of the output data.
[0036] Referring to FIG. 5, the convolutional layer 500 including a combination of a plurality of small kernels includes convolutional parts 510, 520, 530, 540 of the small kernels, and the aggregated convolutional layer 550 includes aggregated convolutional parts 511, 521, 531, 541 and addition parts 560, 570, 580. Here, for ease of explanation, block 501 includes the convolutional part 510 of the small kernel and the aggregated convolutional part 511, block 502 includes the convolutional part 520 of the small kernel and the aggregated convolutional part 521, block 503 includes the convolutional part 530 of the small kernel and the aggregated convolutional part 531, and block 504 includes the convolutional part 540 of the small kernel and the aggregated convolutional part 541.
[0037] Next, the configuration and operation of each block 501, 502, 503, 504 in FIG. 5 will be described with reference to FIGS. 6 to 9.
[0038] FIG. 6 is a diagram showing the configuration and operation of block 501. Block 501 includes a convolutional part 510 of a small kernel including a small kernel 310 having a center 311, and an aggregated convolutional part 511 including an aggregated kernel 350. The circled center 301 described for the small kernel 310 in FIG. 6 indicates the relative position of the center 301 of the large kernel 300 before decomposition with respect to the center 311 of the small kernel 310 after decomposition.
[0039] The convolution of the input data 400 and the small kernel 310 is performed, for example, in the same manner as the convolution of the input data 400 and the large kernel 300. On the input data 400, while moving the center 311 of the small kernel 310 from position 1 to position 100 on the input data 400, it is calculated as the output corresponding to the center 311 of the small kernel 310.
[0040] For example, when the convolution is performed on the data located in the range 401 (positions 12 to 15, 22 to 25, 32 to 35, 42 to 45) on the input data 400 and the small kernel 310 having the center 311, the result of the convolution shown in FIG. 6 is output to the position 23 on the output data 410 of the convolution by the small kernel 310 corresponding to the position 23 of the center 311 of the small kernel 310 on the input data 400.
[0041] As described above, the convolution output data 410 by the small kernel 310 shown in FIG. 6 shows all the results when the input data 400 is convolved with the small kernel 310.
[0042] Here, when calculating the convolution when the center 301 of the large kernel 300 shown in FIG. 3 is at the position 45 of the input data 400, using the convolution results of the decomposed small kernels 310, 320, 330, and 340, the result of the convolution by the small kernel 310 requires the data at the position 23 on the convolution output data 410 by the small kernel 310. The data at the position 23 on this output data 410 is selected and output by the convolution by the aggregated convolution unit 511.
[0043] Next, the configuration and operation of the aggregated convolution unit 511 will be described with reference to FIG. 6. The aggregated convolution unit 511 has an aggregated kernel 350 with a kernel size of 5×5 having a center 351, a value 1 is arranged at the position 352 of the aggregated kernel 350, and values 0 (zero) are arranged at other positions.
[0044] The convolution between the aggregated kernel 350 and the convolution output data 410 by the small kernel 310 is calculated as the output corresponding to the center 351 of the aggregated kernel 350 while moving the center of the aggregated kernel 350 from position 1 to position 100, for example, on the convolution output data 410 by the small kernel 310.
[0045] As shown in FIG. 6, when the center 351 of the aggregated kernel 350 becomes the position 45 on the convolution output data 410 by the small kernel 310, the convolution between the data in the range 411 (positions 23 to 27, 33 to 37, 43 to 47, 53 to 57, 63 to 67) on the convolution output data 410 by the small kernel 310 and the aggregated kernel 350 of the aggregated convolution unit 511 is performed.
[0046] As a result, the data at the position 23 on the convolution output data 410 by the small kernel 310 is output as the result of the block 501.
[0047] That is, the input data 400 is convolved with the small kernel 310, and the aggregation kernel 350 is convolved with the output data 410 of the convolution by the small kernel 310 of the result, so that the necessary convolution result of the input data 400 and the small kernel 310 is output as the result of block 501.
[0048] FIG. 7 is a diagram showing the configuration and operation of block 502. Block 502 includes a convolution section 520 of a small kernel including a small kernel 320 having a center 321, and an aggregation convolution section 521 including an aggregation kernel 360. The circled center 301 described for the small kernel 320 in FIG. 7 indicates the relative position of the center 301 of the large kernel 300 before decomposition with respect to the center 321 of the small kernel 320 after decomposition.
[0049] The convolution of the input data 400 and the small kernel 320 is performed, for example, in the same manner as in the case of the convolution of the input data 400 and the large kernel 300. On the input data 400, while moving the center 321 of the small kernel 320 from position 1 to position 100 on the input data 400, it is calculated as the output corresponding to the center 321 of the small kernel 320.
[0050] For example, when the convolution of the data located in the range 402 (positions 16 to 18, 26 to 28, 36 to 38) on the input data 400 and the small kernel 320 having the center 321 is performed, the result of the convolution shown in FIG. 7 corresponds to the position 27 of the center 321 of the small kernel 320 on the input data 400, and is output to the position 27 on the output data 420 of the convolution by the small kernel 320.
[0051] The output data 420 of the convolution by the small kernel 320 shown in FIG. 7 shows all the results when the convolution of the input data 400 and the small kernel 320 is performed as described above.
[0052] Here, when calculating the convolution in the case where the center 301 of the large kernel 300 shown in FIG. 3 is at the position 45 of the input data 400, to calculate it using the convolution results of the decomposed small kernels 310, 320, 330, 340, the result of the convolution by the small kernel 320 requires the data at the position 27 on the output data 420 of the convolution by the small kernel 320. The data at the position 27 on this output data 420 is selected and output by the convolution performed by the aggregated convolution unit 521.
[0053] Next, the configuration and operation of the aggregated convolution unit 521 will be described with reference to FIG. 7. The aggregated convolution unit 521 has an aggregated kernel 360 with a kernel size of 5×5 having a center 361. A value 1 is arranged at the position 362 of the aggregated kernel 360, and a value 0 (zero) is arranged at other positions.
[0054] The convolution of the output data 420 of the convolution by the aggregated kernel 360 and the small kernel 320 is calculated as the output corresponding to the center 361 of the aggregated kernel 360 while moving the center of the aggregated kernel 360 from position 1 to position 100 on the output data 420 of the convolution by the small kernel 320, for example.
[0055] As shown in FIG. 7, when the center 361 of the aggregated kernel 360 becomes the position 45 on the output data 420 of the convolution by the small kernel 320, the convolution between the data in the range 421 (positions 23 to 27, 33 to 37, 43 to 47, 53 to 57, 63 to 67) on the output data 420 of the convolution by the small kernel 320 and the aggregated kernel 360 of the aggregated convolution unit 521 is performed.
[0056] As a result, the data at the position 27 on the output data 420 of the convolution by the small kernel 320 is output as the result of the block 502.
[0057] That is, the input data 400 is convolved with the small kernel 320, and the aggregation kernel 360 is convolved with the output data 420 of the convolution by the small kernel 320 of the result, so that the necessary convolution result by the input data 400 and the small kernel 320 is output as the result of block 502.
[0058] FIG. 8 is a diagram showing the configuration and operation of block 503. Block 503 includes a convolution section 530 of a small kernel including a small kernel 330 having a center 331, and an aggregation convolution section 531 including an aggregation kernel 370. The circled center 301 described for the small kernel 330 in FIG. 8 indicates the relative position of the center 301 of the large kernel 300 before decomposition with respect to the center 331 of the small kernel 330 after decomposition.
[0059] The convolution of the input data 400 and the small kernel 330 is performed, for example, in the same manner as in the case of the convolution of the input data 400 and the large kernel 300. On the input data 400, while moving the center 331 of the small kernel 330 from position 1 to position 100 on the input data 400, it is calculated as the output corresponding to the center 331 of the small kernel 330.
[0060] For example, when the convolution is performed between the data located in the range 403 (positions 52 to 54, 62 to 64, 72 to 74) on the input data 400 and the small kernel 330 having the center 331, the result of the convolution shown in FIG. 8 corresponds to the position 63 of the center 331 of the small kernel 330 on the input data 400 and is output to the position 63 on the output data 430 of the convolution by the small kernel 330.
[0061] The output data 430 of the convolution by the small kernel 330 shown in FIG. 8 shows all the results when the convolution is performed between the input data 400 and the small kernel 330 as described above.
[0062] Here, when the center 301 of the large kernel 300 shown in FIG. 3 is at the position 45 of the input data 400, to calculate the convolution using the convolution results of the decomposed small kernels 310, 320, 330, 340, the result of the convolution by the small kernel 330 requires the data at the position 63 on the output data 430 of the convolution by the small kernel 330. The data at the position 63 on this output data 430 is selected and output by the convolution by the aggregation convolution unit 531.
[0063] Next, the configuration and operation of the aggregation convolution unit 531 will be described with reference to FIG. 8. The aggregation convolution unit 531 has an aggregation kernel 370 with a kernel size of 5×5 having a center 371, a value 1 is arranged at the position 372 of the aggregation kernel 370, and values 0 (zeros) are arranged at other positions.
[0064] The convolution of the aggregation kernel 370 and the output data 430 of the convolution by the small kernel 330 is calculated as the output corresponding to the center 371 of the aggregation kernel 370 while moving the center of the aggregation kernel 370 from position 1 to position 100 on the output data 430 of the convolution by the small kernel 330, for example.
[0065] As shown in FIG. 8, when the center 371 of the aggregation kernel 370 becomes the position 45 on the output data 430 of the convolution by the small kernel 330, the convolution of the data in the range 431 (positions 23 to 27, 33 to 37, 43 to 47, 53 to 57, 63 to 67) on the output data 430 of the convolution by the small kernel 330 and the aggregation kernel 370 of the aggregation convolution unit 531 is performed.
[0066] As a result, the data at the position 63 on the output data 430 of the convolution by the small kernel 330 is output as the result of the block 503.
[0067] That is, the input data 400 is convolved with the small kernel 330, and the aggregation kernel 370 is convolved with the output data 430 of the convolution by the small kernel 330 as a result, so that the necessary convolution result of the input data 400 and the small kernel 330 is output as the result of block 503.
[0068] FIG. 9 is a diagram showing the configuration and operation of block 504. Block 504 includes a convolution section 540 of a small kernel including a small kernel 340 having a center 341, and an aggregation convolution section 541 including an aggregation kernel 380. The circled center 301 described for the small kernel 340 in FIG. 9 indicates the relative position of the center 301 of the large kernel 300 before decomposition with respect to the center 341 of the small kernel 340 after decomposition.
[0069] The convolution of the input data 400 and the small kernel 340 is performed in the same way as in the case of the convolution of the input data 400 and the large kernel 300 as an example. On the input data 400, while moving the center 341 of the small kernel 340 from position 1 to position 100 on the input data 400, it is calculated as the output corresponding to the center 341 of the small kernel 340.
[0070] For example, when the convolution of the data located in the range 404 (positions 45 to 48, 55 to 58, 65 to 68, 75 to 78) on the input data 400 and the small kernel 340 having the center 341 is performed, the result of the convolution shown in FIG. 9 corresponds to the position 56 of the center 341 of the small kernel 340 on the input data 400, and is output to the position 56 on the output data 440 of the convolution by the small kernel 340.
[0071] The output data 440 of the convolution by the small kernel 340 shown in FIG. 9 shows all the results when the convolution of the input data 400 and the small kernel 340 is performed as described above.
[0072] Here, when calculating the convolution in the case where the center 301 of the large kernel 300 shown in FIG. 3 is at the position 45 of the input data 400, to calculate it using the convolution results of the decomposed small kernels 310, 320, 330, 340, the result of the convolution by the small kernel 340 requires the data at the position 56 on the output data 440 of the convolution by the small kernel 340. The data at the position 56 on this output data 440 is selected and output by the convolution by the aggregation convolution unit 541.
[0073] Next, the configuration and operation of the aggregation convolution unit 541 will be described with reference to FIG. 9. The aggregation convolution unit 541 has an aggregation kernel 380 with a kernel size of 5×5 having a center 381, and a value 1 is arranged at the position 382 of the aggregation kernel 380, and a value 0 (zero) is arranged at other positions.
[0074] The convolution of the aggregation kernel 380 and the output data 440 of the convolution by the small kernel 340 is calculated as the output corresponding to the center 381 of the aggregation kernel 380 while moving the center of the aggregation kernel 380 from position 1 to position 100, for example, on the output data 440 of the convolution by the small kernel 340.
[0075] As shown in FIG. 9, when the center 381 of the aggregation kernel 380 becomes the position 45 on the output data 440 of the convolution by the small kernel 340, the convolution of the data in the range 441 (positions 23 to 27, 33 to 37, 43 to 47, 53 to 57, 63 to 67) on the output data 440 of the convolution by the small kernel 340 and the aggregation kernel 380 of the aggregation convolution unit 541 is performed.
[0076] As a result, the data at the position 56 on the output data 440 of the convolution by the small kernel 340 is output as the result of the block 504.
[0077] That is, the input data 400 is convolved with the small kernel 340, and the aggregation kernel 380 is convolved with the output data 440 of the convolution by the small kernel 340 as a result, so that the necessary convolution result of the input data 400 and the small kernel 340 is output as the result of the block 504.
[0078] By adding the results of each block 501, 502, 503, and 504 output as described above by the adders 560, 570, and 580 in the aggregation convolution layer 550 in FIG. 5, when the center 301 of the large kernel 300 shown in FIG. 3 is at the position 45 of the input data 400, the calculation of the convolution can be calculated using the convolution results of the decomposed small kernels 310, 320, 330, and 340. Note that the output data 455 in FIG. 5 coincides with the output data 450 in FIG. 4 except for a part of the peripheral part.
[0079] In this way, since the convolution layer including the large kernel is converted into a convolution layer and an aggregation convolution layer including a combination of a plurality of small kernels obtained by decomposing the large kernel, for the convolution of the input data and each small kernel, for example, optimization methods related to the convolution of a kernel of size 3×3 such as winograd optimization that can be accelerated at double speed can be utilized. In addition, circuits and implementations that can execute the convolution layer of the 3×3 size kernel at high speed can be utilized to the maximum extent. Furthermore, if a sparsity-utilizing acceleration mechanism that skips multiplication by zero values can be used, it can also be utilized. Thereby, the execution speed of the convolution can be increased.
[0080] As described above, according to the first embodiment of the present invention, a convolution layer conversion device can be provided that contributes to improving the execution speed when implementing the convolution layer of the neural network model.
[0081] [Second Embodiment] Next, the convolution layer conversion device according to the second embodiment of the present invention will be described with reference to the drawings. FIG. 10 is a diagram showing an example of the configuration of the convolution layer conversion device according to the second embodiment of the present invention. In FIG. 10, components denoted by the same reference numerals as those in FIG. 2 are assumed to be the same components, and the description thereof will be omitted.
[0082] Referring to FIG. 10, the convolutional layer conversion device 100 of the second embodiment of the present invention includes a large kernel convolution (Conv) layer detection unit 110 (convolutional layer detection unit) and a convolutional layer decomposition unit 120. The convolutional layer decomposition unit 120 includes a decomposition method selection unit 121, a layer decomposition application unit 122, and an adjustment unit 125. The convolutional layer conversion device 100 takes the neural network (NN) model structure 10 as an input and outputs the converted neural network model structure 20. Further, a target device information storage unit 30 is connected to the decomposition method selection unit 121 of the convolutional layer decomposition unit 120, and target device information is inputted.
[0083] FIG. 11 is a diagram showing an example of the configuration of the converted convolutional layer of the second embodiment of the present invention in which padding processing units 590 and 595 are respectively provided in the convolutional layer 500 including a combination of a plurality of small kernels and the aggregated convolutional layer 550 by the adjustment unit 125.
[0084] Next, the operation of the adjustment unit 125 of the convolutional layer decomposition unit 120 will be described with reference to the drawings. The adjustment unit 125 has a function of providing padding processing units 590 and 595 for adjusting the degree of mismatch between the output data 455 output from the aggregated convolutional layer 550 and the output data 450 of the convolution result by the convolutional layer including the large kernel 300 shown in FIG. 4, respectively, in the convolutional layer 500 including a combination of a plurality of small kernels and the aggregated convolutional layer 550 shown in FIG. 11. The padding processing units 590 and 595 provided by the adjustment unit 125 in the convolutional layer 500 and the aggregated convolutional layer 550 including a combination of a plurality of small kernels respectively execute the following operations in order to adjust the degree of mismatch with the output data 450.
[0085] Next, the operations of the padding processing unit 590 and the padding processing unit 595 provided in the convolutional layer 500 and the aggregated convolutional layer 550 including a combination of a plurality of small kernels respectively by the adjustment unit 125 will be described.
[0086] FIG. 12 is a diagram showing an example of operations of a padding processing unit 590 provided in a convolutional layer 500 in which an adjustment unit according to a second embodiment of the present invention includes a combination of small kernels, and a padding processing unit 595 provided in an aggregated convolutional layer 550.
[0087] FIG. 12 shows an example of operations of the padding processing unit 590 and the padding processing unit 595 with respect to block 502 of FIG. 11. Similar operations may be performed on other blocks 501, 503, and 504.
[0088] As an example, the padding processing unit 590 adds padding data with a value of 0 (zero) to each position from p1 to p21 with respect to positions 1 to 100 of the input data 400 shown in FIG. 12.
[0089] Next, when the padding processing unit 590 calculates the convolution of the small kernel 320 when the center 301 of the large kernel 300 corresponds to position 12 on the input data 400, the convolution unit 520 of the small kernel controls the convolution unit 520 of the small kernel so as to calculate the convolution in which the center 321 of the small kernel 320 is at position p4. Specifically, the padding processing unit 590 controls the convolution unit 520 of the small kernel in block 502 so that the convolution of the data at positions 3, 4, and 5 on the input data 400 and the elements of the three kernels in the bottom row of the small kernel 320 is performed and output to position p4 on the output data 420 of the convolution by the small kernel 320.
[0090] When the padding processing unit 595 calculates the convolution of the small kernel 320 when the center 301 of the large kernel 300 corresponds to position 12 on the input data 400, the padding processing unit 595 controls the aggregated convolution unit 521 so that the center 361 of the aggregated kernel 360 in the aggregated convolution unit 521 performs convolution with the data range 421 that becomes position 12 on the output data 420 of the convolution by the small kernel 320. Specifically, the data at position p4 on the output data 420 of the convolution by the small kernel 320 is convolved with the value 1 at position 362 of the aggregated kernel 360 in the aggregated convolution unit 521 and output as a result of block 502.
[0091] Similarly to the above, for the other blocks 501, 503, and 504, depending on the positional relationship between the decomposed small kernels and the large kernel, the padding processing units 590 and 595 perform padding processing on the upper, lower, left, and right peripheral portions of the input data 400. By doing so, also at the peripheral portions of the input data 400, the output data 455 in FIG. 11 can be made to match the output data 450 that is the result of convolution by the convolutional layer including the large kernel 300 shown in FIG. 4. Note that the number of padding data to be added changes depending on the size of the decomposed small kernel and the positional relationship between the large kernel before decomposition and the decomposed small kernel.
[0092] By adjusting, by the adjustment unit 125 of the convolutional layer conversion device 100 according to the second embodiment of the present invention, the sizes of the padding data added and processed by the padding processing unit 590 and the padding processing unit 595, it is also possible to adjust the degree of mismatch with the convolution result by the large kernel in the peripheral portion of the image. Further, when it is possible to allow a mismatch in the peripheral portion of the image, the adjustment unit 125 may not provide the padding processing units 590 and 595.
[0093] [Third Embodiment] Next, another decomposition method for decomposing the large kernel according to the third embodiment of the present invention into a plurality of small kernels will be described with reference to the drawings. FIG. 13 is a diagram showing another example of a decomposition method for decomposing the large kernel according to the third embodiment of the present invention into a plurality of small kernels. In FIG. 13, components denoted by the same reference numerals as in FIG. 3 are assumed to be the same components, and the description thereof will be omitted.
[0094] FIG. 13 is a diagram showing an example of a decomposition method for decomposing a large kernel 300 with a 7×7 kernel size according to the third embodiment of the present invention into seven small kernels 710, 720, 730, 740, 750, 760, and 770. The small kernel 710 having a center 711, the small kernel 720 having a center 721, the small kernel 730 having a center 731, and the small kernel 740 having a center 741 are small kernels with a 3×3 kernel size corresponding to the upper left corner, upper right corner, lower left corner, and lower right corner portions of the large kernel 300.
[0095] The small kernel 750 having a center 751 has the values at positions a, b, c, d, and e on the large kernel 300 at the corresponding positions indicated by positions a, b, c, d, and e on the small kernel 750, respectively, and the positions where 0 (zero) is described have zero values.
[0096] The small kernel 760 having a center 761 has the values at positions p, q, r, and s on the large kernel 300 at the corresponding positions indicated by positions p, q, r, and s of the small kernel 760, respectively, the positions where 0 (zero) is described have zero values, and the hatched portions have no values.
[0097] The small kernel 770 having a center 771 has the values at positions w, x, y, and z on the large kernel 300 at the corresponding positions indicated by positions w, x, y, and z of the small kernel 770, respectively, the positions where 0 (zero) is described have zero values, and the hatched portions have no values.
[0098] The small kernel 750 is a small kernel with a 3×3 kernel size.
[0099] The small kernel 760 is a small kernel with a 5×5 kernel size, but by setting the dilation during convolution to 2, it can be expressed as a small kernel 760A having a center 761 with a 3×3 kernel size with the hatched portions removed.
[0100] The small kernel 770 is a small kernel with a kernel size of 7×7. However, by setting the dilation during convolution to 3, it can be expressed as a small kernel 770A having a central 771 with a kernel size of 3×3, where the hatched part is removed.
[0101] Thus, the small kernels 710, 720, 730, 740, 750, 760, 770 of the third embodiment can all be expressed as small kernels with a kernel size of 3×3 by appropriately setting the dilation during convolution.
[0102] FIG. 14 is a diagram showing an example of each small kernel 710, 720, 730, 740 of a convolutional layer including a combination of disassembled small kernels of the third embodiment of the present invention, and each aggregated kernel 1410, 1420, 1430, 1440 of the aggregated convolutional layer corresponding to each small kernel 710, 720, 730, 740.
[0103] The sizes of the aggregated kernels 1410, 1420, 1430, 1440 have a size of 5×5 as an example, and the centers are 1411, 1421, 1431, 1441 respectively.
[0104] Also, FIG. 15 is a diagram showing an example of each small kernel 750, 760 (760A, dilation = 2), 770 (770A, dilation = 3) of a convolutional layer including a combination of disassembled small kernels of the third embodiment of the present invention, and each aggregated kernel 1450, 1460, 1470 of the aggregated convolutional layer corresponding to each small kernel 750, 760, 770. The sizes of the aggregated kernels 1450, 1460, 1470 have a size of 1×1 as an example.
[0105] When calculating the convolution of the input data 400 and the large kernel 300 using the convolutions of the input data 400 and each of the small kernels 710, 720, 730, 740, 750, 760 (760A, dilation = 2), 770 (770A, dilation = 3), the convolution results of each of the small kernels 710, 720, 730, 740 need to be obtained and added respectively from positions corresponding to the centers 711, 721, 731, 741 of each of the small kernels 710, 720, 730, 740 with respect to the center 301 of the large kernel 300.
[0106] The convolution results of each of the necessary small kernels 710, 720, 730, 740 can be obtained from the convolution results of the input data and each of the small kernels 710, 720, 730, 740 respectively, and the convolution results with each of the aggregation kernels 1410, 1420, 1430, 1440 described in FIG. 14.
[0107] The convolution results of each of the necessary small kernels 750, 760, 770 can be obtained from the convolution results of the input data and each of the small kernels 750, 760, 770 respectively, and the convolution results with each of the aggregation kernels 1450, 1460, 1470 described in FIG. 15. Note that each of the aggregation kernels 1450, 1460, 1470 is, as an example, of size 1×1. That is, since the center 301 of the large kernel 300 and the centers 751, 761, 771 of each of the small kernels 750, 760, 770 coincide, the convolution by the aggregation kernels 1450, 1460, 1470 is unnecessary.
[0108] FIG. 16 is a diagram showing an example of a computational graph before and after decomposition of the large kernel 300 according to the third embodiment of the present invention. The computational graph 1600 by the large kernel 300 before decomposition on the left side of FIG. 16 includes a convolution 1602 of the large kernel 300 having a size of 7×7. On the other hand, the computational graph 1605 after decomposition on the right side includes convolutions 1610, 1620, 1630, 1640 that execute convolutions of the small kernels 710, 720, 730, 740, and aggregation convolutions 1611, 1621, 1631, 1641 that execute convolutions by the corresponding aggregation kernels 1410, 1420, 1430, 1440 for the respective convolution results, and additions 1681 to 1686 that add the computational results of the convolutions 1650, 1660, 1670 that execute convolutions of the small kernels 750, 760, 770. With the configuration of FIG. 16, by sequentially adding each computational result, it is possible to execute the convolution after decomposition into small kernels.
[0109] FIG. 17 is a diagram showing another example of a computational graph before and after decomposition of the large kernel 300 according to the third embodiment of the present invention. The computational graph 1600 by the large kernel 300 before decomposition on the left side of FIG. 17 is the same as the computational graph 1600 by the large kernel 300 before decomposition on the left side of FIG. 16. On the other hand, in the computational graph 1606 after decomposition on the right side, the computational results of the convolutions 1610, 1620, 1630, 1640 that execute convolutions of the small kernels 710, 720, 730, 740 and the aggregation convolutions 1611, 1621, 1631, 1641 that execute convolutions by the corresponding aggregation kernels 1410, 1420, 1430, 1440 for the respective convolution results, and the computational results of the convolutions 1650, 1660, 1670 that execute convolutions of the small kernels 750, 760, 770 are calculated in parallel, and by performing parallel addition of each result by the additions 1701 to 1706, it is possible to execute the convolution after decomposition into small kernels.
[0110] In this way, since the convolutional layer of the large kernel 300 is converted into a convolutional layer and an aggregated convolutional layer including a combination of a plurality of small kernels 710, 720, 730, 740, 750, 760, 770, for the convolution of the input data and each small kernel, for example, optimization methods related to the convolution of a kernel of size 3×3, such as winograd optimization that can be accelerated at double speed, can be utilized. Also, circuits and implementations capable of executing the convolutional layer of a 3×3-sized kernel can be maximally utilized. Furthermore, if a sparsity-utilizing acceleration mechanism that skips multiplication by zero values can be used, it can also be utilized. As a result, the execution speed of the convolution can be increased.
[0111] [Fourth Embodiment] Next, a fourth embodiment of the present invention will be described with reference to the drawings. FIG. 18 is a diagram showing an example of the configuration of a decomposition method selection unit of a convolutional layer decomposition unit of a convolutional layer conversion device according to the fourth embodiment of the present invention. The decomposition method selection unit 121 in FIG. 18 shows an example of the configuration of the decomposition method selection unit 121 of the convolutional layer decomposition unit 120 of the convolutional layer conversion device 100 according to the first embodiment of the present invention shown in FIG. 2. In FIG. 18, components denoted by the same reference numerals as those in FIG. 2 are assumed to be the same components, and the description thereof will be omitted.
[0112] The decomposition method selection unit 121 according to the fourth embodiment of the present invention includes a decomposition candidate enumeration unit 1801, an execution parameter investigation unit 1802 for each candidate, and a candidate selection unit 1803. The execution parameter investigation unit 1802 for each candidate is connected to a target device information storage unit 30. The target device information storage unit 30 includes an on-device execution speed database (DB) 31 and a target device specifying unit 32.
[0113] The execution speed of convolution varies depending on the method of accelerating the calculations of the device that executes the convolution. For example, in the case of convolution with a 3×3 kernel, as an example, optimization methods for convolution with a 3×3 kernel, such as winograd optimization that can be accelerated at double speed, can be utilized. Further, circuits and implementations capable of executing the convolution layer with a 3×3 kernel at high speed can be maximally utilized. Moreover, if a sparsity-utilizing acceleration mechanism that skips multiplication with zero values can be used, that can also be utilized. Thereby, the execution speed of convolution can be increased.
[0114] However, since the method of accelerating the calculations of the device that executes the convolution may vary for each device, for example, as a decomposition candidate, when measuring the execution speed when executing the convolution of a convolution layer including a combination of a plurality of small kernels obtained by decomposing a 7×7 large kernel into a 4×4 kernel and a 3×3 kernel on each device, or as another decomposition candidate, when measuring the execution speed when executing the convolution by a convolution layer including a combination of a plurality of small kernels obtained by decomposing a 7×7 large kernel into a 3×3 kernel, a 3×3 (dilation 2) kernel, and a 3×3 (dilation 3) kernel on each device, etc., execution speed information corresponding to the device is stored in advance in the on-device execution speed database 31 as target device information. Note that the decomposition candidates for measuring the execution speed are not limited to the above examples, and combinations of kernels of other sizes or formats may also be used. Further, the on-device execution speed database 31 may store memory usage information indicating the memory usage when executing a convolution layer including a combination of a plurality of small kernels on the target device.
[0115] FIG. 19 shows an example of target device information stored in the on-device execution speed database 31. Referring to FIG. 19, column 1901 indicates the target device, column 1902 indicates the layer type, column 1903 indicates the layer parameters, column 1904 indicates the execution time, and column 1905 indicates the memory usage information. The execution time can be obtained by previously executing the process indicated by each layer parameter on the target device and measuring the time required for the process. The execution speed information is indicated by the execution time in column 1904, and a smaller execution time corresponds to the execution speed information indicating a higher execution speed.
[0116] Next, the operation of the decomposition method selection unit 121 of the convolution layer decomposition unit 120 of the convolution layer conversion device 100 according to the fourth embodiment of the present invention will be described with reference to FIG. 18. Referring to FIG. 18, the convolution layer including a large kernel having a kernel size equal to or larger than a predetermined size detected by the large kernel convolution layer detection unit 110 is input to the decomposition candidate enumeration unit 1801 of the decomposition method selection unit 121 from the input neural network (NN) model structure 10.
[0117] The decomposition candidate enumeration unit 1801 enumerates decomposition candidates for decomposing the input convolution layer including a large kernel having a size of, for example, 7×7 into a convolution layer including a combination of a plurality of small kernels. For example, an example of the decomposition candidate is described in the above embodiment, but the decomposition candidate is not limited to the above example, and a combination of kernels of other sizes or formats may be used. The enumerated decomposition candidates are sent to the execution parameter investigation unit 1802 for each candidate.
[0118] On the one hand, the target device designator 32 of the target device information storage unit 30 designates a target device on which convolution is to be executed. The device execution speed database (DB) 31 sends, to the execution parameter investigation unit 1802 for each candidate, execution speed information indicating the execution speed when a convolution layer (decomposition candidate decomposed by a decomposition method) including a combination of a plurality of small kernels stored for the target device designated by the target device designator 32 is executed on the target device. Also, when memory usage information is stored, the device execution speed database (DB) 31 sends, to the execution parameter investigation unit 1802 for each candidate, memory usage information indicating the memory usage when a convolution layer (decomposition candidate decomposed by a decomposition method) including a combination of a plurality of small kernels is executed on the target device.
[0119] The execution parameter investigation unit 1802 for each candidate investigates the execution speed of the enumerated decomposition candidates using the execution speed information for each decomposition candidate, and designates the fastest decomposition candidate.
[0120] The candidate selection unit 1803 instructs the designated fastest decomposition candidate to the layer decomposition application unit 122.
[0121] As a result, the layer decomposition application unit 122 can decompose the convolution layer using the decomposition candidate that can execute convolution fastest on the designated device.
[0122] Also, when the memory usage information is also sent to the execution parameter investigation unit 1802 for each candidate, in addition to the execution speed information, the memory usage information may also be referred to, and a decomposition candidate that meets a predetermined selection criterion for both the execution speed information and the memory usage information may be selected. For example, a decomposition candidate whose execution speed is equal to or higher than a predetermined value and whose memory usage is equal to or lower than a predetermined value may be selected. Alternatively, a decomposition candidate with the least memory usage may be selected based on the memory usage information.
[0123] Since the decomposition method selection unit 121 of the fourth embodiment of the present invention can select a decomposition candidate that can execute convolution at the fastest speed on a specified device, the convolution layer conversion device 100 shown in FIG. 2 can convert a convolution layer including a large kernel into a convolution layer including a combination of a plurality of small kernels with the fastest execution speed on the specified device and an aggregation convolution layer that aggregates the results of the convolution layer including the combination of the plurality of small kernels, and can output a neural network model structure 20 in which the convolution layer including the large kernel is converted.
[0124] Also, both the execution speed information and the memory usage information are used to select a decomposition candidate that meets a predetermined selection criterion, and a convolution layer including a large kernel is converted into a convolution layer including a combination of a plurality of small kernels and an aggregation convolution layer that aggregates the results of the convolution layer including the combination of the plurality of small kernels, and a neural network model structure 20 in which the convolution layer including the large kernel is converted can be output.
[0125] Furthermore, a decomposition candidate with the least memory usage is selected based on the memory usage information, and a convolution layer including a large kernel is converted into a convolution layer including a combination of a plurality of small kernels and an aggregation convolution layer that aggregates the results of the convolution layer including the combination of the plurality of small kernels, and a neural network model structure 20 in which the convolution layer including the large kernel is converted can be output.
[0126] As described above, each embodiment of the present invention has been described. However, the present invention is not limited to the above-described embodiments, and further modifications, substitutions, and adjustments can be made without departing from the basic technical idea of the present invention. For example, the system configuration shown in each drawing, the configuration of each element, and the expression form of the message are examples for helping the understanding of the present invention, and are not limited to the configurations shown in these drawings. Also, in the following description, "A and / or B" is used to mean at least one of A or B.
[0127] Also, the procedures shown in the above-described first to fourth embodiments can be realized by a program that causes a computer (9000 in FIG. 20) functioning as a convolutional layer conversion device to realize the functions of the convolutional layer conversion device. Such a computer is exemplified by a configuration including a CPU (Central Processing Unit) 9010, a communication interface 9020, a memory 9030, and an auxiliary storage device 9040 in FIG. 20. That is, the CPU 9010 in FIG. 20 may execute a convolutional layer conversion program and perform an update process on each calculation parameter held in the auxiliary storage device 9040 or the like.
[0128] The memory 9030 is a RAM (Random Access Memory), a ROM (Read Only Memory), or the like.
[0129] That is, each part (processing means, function) of the convolutional layer conversion device shown in the above-described first to fourth embodiments can be realized by a computer program that causes a processor of the computer to execute each of the above-described processes using the hardware thereof.
[0130] Finally, the preferred forms of the present invention are summarized. [First Form] (Refer to the convolutional layer conversion device from the above first perspective) [Second Form] In the convolutional layer conversion device of the first form, the convolutional layer decomposition unit preferably further includes an adjustment unit that respectively provides a padding processing unit for adjusting the degree of mismatch between the aggregation result of the aggregation convolutional layer and the result of convolution by the convolutional layer including the large kernel to the convolutional layer including the combination of the plurality of small kernels and the aggregation convolutional layer. [Third Form] The convolutional layer conversion device of the first or second form, the convolutional layer decomposition unit refers to target device information and selects a decomposition method for decomposing the large kernel into the plurality of small kernels, a decomposition method selection unit, A layer decomposition application unit that generates a convolutional layer including a combination of the plurality of small kernels and the aggregated convolutional layer according to the selected decomposition method; It is preferable to include. [Fourth form] In the convolutional layer conversion device of the third form, the decomposition method selection unit A decomposition candidate enumeration unit that enumerates decomposition candidates of the decomposition method; For each of the enumerated decomposition candidates, referring to the target device information, an execution parameter investigation unit that investigates execution parameters on the target device; It is preferable to include a decomposition candidate selection unit that selects a decomposition candidate with optimal execution parameters. [Fifth form] In the convolutional layer conversion device of the fourth form, the target device information includes execution speed information indicating the execution speed when a convolutional layer including a combination of the plurality of small kernels is executed on the target device, or memory usage information indicating the memory usage amount when a convolutional layer including a combination of the plurality of small kernels is executed on the target device. The execution parameter investigation unit investigates the execution speed on the target device with reference to the execution speed information for each of the enumerated decomposition candidates, or investigates the memory usage amount on the target device with reference to the memory usage information for each of the enumerated decomposition candidates. It is preferable that the decomposition candidate selection unit selects a decomposition candidate in which at least one of the execution speed and the memory usage amount meets a predetermined selection criterion. [Sixth form] In the convolutional layer conversion device of the fifth form, the predetermined selection criterion for the execution speed is the fastest execution speed, and the predetermined selection criterion for the memory usage amount is the minimum memory usage amount, which is preferable. [Seventh form] In the convolutional layer conversion device of the fifth form, it is preferable that the decomposition candidate selection unit selects a decomposition candidate in which both the execution speed and the memory usage amount meet a predetermined selection criterion. [Eighth embodiment] In the convolutional layer conversion device according to the seventh embodiment, it is preferable that the predetermined selection criterion is that the execution speed is a speed equal to or higher than a predetermined value and the memory usage amount is an amount equal to or less than a predetermined value. [Ninth embodiment] (Refer to the convolutional layer conversion method from the second perspective above) [Tenth embodiment] (Refer to the program from the third perspective above) Note that the ninth and tenth embodiments can be expanded to the second to eighth embodiments in the same manner as the first embodiment.
[0131] Note that each disclosure of the above patent documents is incorporated herein by reference. Within the scope of the entire disclosure of the present invention (including the claims), further modifications and adjustments of the embodiments or examples can be made based on the basic technical idea. Also, within the scope of the disclosure of the present invention, various combinations or selections of various disclosure elements (including each element of each claim, each element of each embodiment or example, each element of each drawing, etc.) are possible. That is, the present invention naturally includes all disclosures including the claims and various modifications and corrections that could be made by those skilled in the art according to the technical idea. In particular, for the numerical ranges described in this document, any numerical value or small range included within the range should be construed as specifically described even without separate description.
Explanation of reference numerals
[0132] 10 Neural network (NN) model structure 20 Converted neural network (NN) model structure 30 Target device information storage unit 31 Device execution speed database (DB) 32 Target device designating unit 100 Convolutional layer conversion device 110 Large kernel convolution (Conv) layer detection unit 120 Convolutional layer decomposition unit 121 Decomposition method selection unit 122-layer decomposition application part 125 adjustment part 300 large kernel 301 center 310, 320, 330, 340 small kernels 311, 321, 331, 341 centers 350, 360, 370, 380 aggregation kernels 351, 361, 371, 381 centers 400 input data 450, 455 output data 500 convolutional layer including combinations of multiple small kernels 501, 502, 503, 504 blocks 510, 520, 530, 540 convolutional parts of small kernels 511, 521, 531, 541 aggregation convolutional parts 550 aggregation convolutional layer 560, 570, 580 addition parts 590, 595 padding processing parts 710, 720, 730, 740, 750 small kernels 760, 760A, 770, 770A small kernels 711, 721, 731, 741 centers 751, 761, 771 centers 1410, 1420, 1430, 1440, 1450, 1460, 1470 aggregation kernels 1411, 1421, 1431, 1441 centers 1600 calculation graph before decomposition 1605 calculation graph after decomposition 1606 calculation graph after decomposition 1801 decomposition candidate enumeration part 1802 execution parameter investigation part 1803 candidate selection part 9000 computer 9010 CPU 9020 communication interface 9030 memory 9040 auxiliary storage device
Claims
1. A convolutional layer detection unit that detects a convolutional layer including a large kernel having a kernel size equal to or greater than a predetermined size from an input neural network model structure; A convolutional layer decomposition unit that converts the convolutional layer including the large kernel into a convolutional layer including a combination of a plurality of small kernels having a kernel size smaller than the predetermined size obtained by decomposing the large kernel, and an aggregation convolutional layer that aggregates the convolutional results of the convolutional layer including the combination of the plurality of small kernels, and outputs a neural network model structure in which the convolutional layer including the large kernel is converted; comprising The aggregation convolutional layer includes an aggregation convolutional unit that performs convolution on the convolutional result of the small kernel, and an addition unit that simultaneously adds the outputs of all the aggregation convolutional units. The aggregation convolutional unit is a convolutional layer conversion device in which a value 0 is arranged at positions other than a predetermined one position based on a relative position between the corresponding small kernel and the large kernel.
2. The convolutional layer decomposition unit further includes an adjustment unit that respectively provides a padding processing unit for adjusting the degree of mismatch between the aggregation result of the aggregation convolutional layer and the convolution result by the convolutional layer including the large kernel to the convolutional layer including the combination of the plurality of small kernels and the aggregation convolutional layer. The convolutional layer conversion device according to claim 1.
3. The convolutional layer decomposition unit A decomposition method selection unit that selects a decomposition method for decomposing the large kernel into the plurality of small kernels with reference to target device information; A layer decomposition application unit that generates a convolutional layer including the combination of the plurality of small kernels and the aggregation convolutional layer according to the selected decomposition method. The convolutional layer conversion device according to claim 1 or 2. comprising
4. The decomposition method selection unit A decomposition candidate enumeration unit that enumerates decomposition candidates of the decomposition method; An execution parameter investigation unit that investigates execution parameters on a target device with reference to the target device information for each of the enumerated decomposition candidates; A decomposition candidate selection unit that selects a decomposition candidate with the optimal execution parameter. The convolutional layer conversion device according to claim 3.
5. The target device information includes execution speed information indicating the execution speed when a convolutional layer including a combination of the plurality of small kernels is executed on the target device, or memory usage information indicating the memory usage amount when a convolutional layer including a combination of the plurality of small kernels is executed on the target device. The execution parameter investigation unit investigates the execution speed on the target device with reference to the execution speed information for each of the enumerated decomposition candidates, or investigates the memory usage amount on the target device with reference to the memory usage information for each of the enumerated decomposition candidates. The decomposition candidate selection unit selects a decomposition candidate in which at least one of the execution speed and the memory usage amount meets a predetermined selection criterion. The convolutional layer conversion device according to claim 4.
6. The predetermined selection criterion for the execution speed is the fastest execution speed, and the predetermined selection criterion for the memory usage amount is the minimum memory usage amount. The convolutional layer conversion device according to claim 5.
7. The decomposition candidate selection unit selects a decomposition candidate in which both the execution speed and the memory usage amount meet a predetermined selection criterion. The convolutional layer conversion device according to claim 5.
8. The predetermined selection criterion is that the execution speed is a speed equal to or higher than a predetermined value, and the memory usage amount is an amount equal to or less than a predetermined value. The convolutional layer conversion device according to claim 7.
9. Executed by a computer including a processor and a storage device, Detecting a convolutional layer including a large kernel with a kernel size equal to or larger than a predetermined size from the input neural network model structure; Converting the convolutional layer including the large kernel into a convolutional layer including a combination of a plurality of small kernels with a kernel size smaller than the predetermined size obtained by decomposing the large kernel and an aggregation convolutional layer that aggregates the convolutional results of the convolutional layer including the combination of the plurality of small kernels, and outputting a neural network model structure in which the convolutional layer including the large kernel is converted; including The aggregation convolutional layer includes an aggregation convolutional unit that performs convolution on the convolutional results of the small kernels, and an addition unit that simultaneously adds the outputs of all the aggregation convolutional units. The convolutional layer conversion method in which a value of 0 is arranged at positions other than a predetermined one position based on the relative position between the corresponding small kernel and the large kernel in the aggregation convolutional part.
10. A computer, a process of detecting a convolutional layer including a large kernel having a kernel size equal to or larger than a predetermined size from an input neural network model structure, a process of converting the convolutional layer including the large kernel into a convolutional layer including a combination of a plurality of small kernels having a kernel size smaller than the predetermined size obtained by decomposing the large kernel, and an aggregation convolutional layer that aggregates the convolution results of the convolutional layer including the combination of the plurality of small kernels, and outputting a neural network model structure in which the convolutional layer including the large kernel is converted, A program for causing the computer to execute the above processes, wherein the aggregation convolutional layer includes an aggregation convolutional part that performs convolution on the convolution result of the small kernel, corresponding to each of the small kernels, and an addition part that simultaneously adds the outputs of all the aggregation convolutional parts, The program, wherein the aggregation convolutional part has a value of 0 arranged at positions other than a predetermined one position based on the relative position between the corresponding small kernel and the large kernel.
Citation Information
Patent Citations
Remote sensing image target extraction method based on deep neural network
CN112712500A
Processor, information processing device, and operation method for processor
JP2018120549A
Apparatus for detecting variant malicious code based on neural network learning, method therefor, and computer-readable recording medium having recorded thereon a program for executing the method
JP2019527447A
DNN weight saving device
JP2020087288A
Learning model generation device, learning model generation method, and program
JP2020107042A