Neural network optimization method, program, and machine learning device
The neural network optimization method addresses accuracy issues in pruning by evaluating and pruning layers with dependency relationships using statistical scores, ensuring effective model size reduction across diverse network models.
Patent Information
- Application Number
- JP2021209309
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2041-12-23
AI Technical Summary
Existing neural network pruning methods fail to consider dependency relationships between layers, leading to significant accuracy degradation when applied to neural networks with such dependencies, and are not generalizable to network models other than ResNet.
A method for neural network optimization that evaluates and prunes layers with dependency relationships by calculating scores, treating them as evaluation units, and using statistical methods like average, maximum, or median scores to determine pruning targets, while normalizing layers collectively.
This approach maintains accuracy by appropriately pruning layers with dependency relationships, applicable to various network models, reducing model size without significant accuracy loss.
Smart Images

Figure 0007715036000001 
Figure 0007715036000002 
Figure 0007715036000003
Abstract
Description
Technical Field
[0001] The present invention relates to a neural network optimization method, a program, and a machine learning device.
Background Art
[0002] Deep learning has come to be widely used in various fields. In the analysis process using deep learning, a method has been used in which sensing data from edge devices arranged on a production line or a vehicle is analyzed by a machine learning model on a cloud server, and the analysis result is sent to the edge device.
[0003] Under such circumstances, from the viewpoint of communication delay between the edge device and the cloud server, there is a desire to perform analysis on the edge device side. In order to perform analysis with limited resources on the edge device side, it is necessary to compress and reduce the size of the machine learning model. As a weight reduction algorithm for that purpose, there is pruning.
[0004] In the technique disclosed in Non-Patent Document 1, pruning of a neural network is performed by a three-step method of first learning the importance of connections, second deleting unimportant connections, and finally re-training the remaining connections. And in this second step, connections having weights below a threshold value are removed from the neural network.
[0005] In the technique disclosed in Non-Patent Document 2, in a convolutional neural network, importance is determined by the total value of the absolute values of the weights of filters, and filters having a low total value and not important and the feature maps related thereto are pruned.
[0006] In the technology disclosed in Non-Patent Document 3, in ResNet disclosed in Non-Patent Document 4, the importance is calculated for each residual unit, and the residual units determined to be unimportant according to the importance are deleted.
Prior Art Documents
Non-Patent Documents
[0007]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0008] The structure of a neural network is composed of multiple layers, and there may be layers that have a dependency relationship with each other among the multiple layers. Layers with a dependency relationship need to match the number of channels even after pruning. In Non-Patent Documents 1 and 2, pruning regarding such layers with a dependency relationship is not considered. For example, if the rules of Non-Patent Documents 1 and 2 are mechanically applied to the pruning of a neural network including such layers with a dependency relationship, there may be a significant decrease in accuracy. For example, when pruning the filters or weights of a certain layer (hereinafter referred to as the first layer), the corresponding filters or weights (connections) of this layer and the layer with a dependency relationship (hereinafter referred to as the second layer) will be mechanically pruned. In such a case, if the importance of the filters or weights of the mechanically pruned second layer is high, a decrease in accuracy will occur.
[0009] Also, the technology of Non-Patent Document 3 performs pruning in units of residual units and comprehensively evaluates including the dependency relationship associated with the add process (skip connection), but it cannot be applied to network models other than ResNet and is not general-purpose.
[0010] The present invention has been made to solve such problems. That is, it is an object of the present invention to provide a neural network optimization method and a machine learning device that can be generally applied to a neural network including layers having a dependency relationship without being limited to a specific network model and can suppress deterioration in accuracy.
Means for Solving the Problems
[0011] The above problems of the present invention are solved by the following means.
[0012] (1) A method for performing pruning on a neural network, comprising: step (a) of obtaining a neural network; step (b) of calculating a score for each layer included in the neural network by scoring; when the layer has a dependency relationship, step (c) of evaluating with a plurality of layers having a dependency relationship based on the scoring result of step (b); step (d) of determining a pruning target included in the layer from the evaluation result of step (c); step (e) of pruning the target determined in step (d), and executing a process including the above steps, a neural network optimization method.
[0013] (2) The plurality of layers having the dependency relationship are: layers that have a relationship that requires processing with the same number of channels in the processing between the plurality of layers or between subsequent layers of each of the plurality of layers. The neural network optimization method according to (1) above.
[0014] (3) The target is a filter included in each of a plurality of layers having a dependency relationship. The neural network optimization method according to (1) or (2) above.
[0015] (4) A neural network optimization method according to any one of (1) to (3), wherein in step (c), a set of dependent layers is treated as a set of evaluation units, and the evaluation is performed between the evaluation units.
[0016] (5) The neural network optimization method according to (4) above, wherein in step (c), evaluation is performed using at least one of the average, maximum, minimum, total, and median of the scores calculated for each evaluation unit.
[0017] (6) further comprising a step (f) of normalizing the layer; A neural network optimization method according to any one of (1) to (5) above, wherein in step (f), normalization of the plurality of dependent layers is performed collectively for the plurality of dependent layers or for the components of the plurality of layers.
[0018] (7) A program for causing a computer to execute any one of the neural network optimization methods (1) to (6) above.
[0019] (8) an acquisition unit for acquiring a neural network; a scoring unit that calculates a score for each layer included in the neural network by scoring; an evaluation unit that evaluates a plurality of layers that are in a dependent relationship based on a scoring result of the scoring unit when the layers have a dependent relationship; a determination unit that determines a pruning target included in the layer based on the evaluation result by the evaluation unit; a pruning unit that prunes the target determined by the determination unit.
[0020] (9) The machine learning device described in (8) above, wherein the scoring unit treats multiple layers that are dependent on each other as a set of evaluation units and performs the evaluation between the evaluation units.
[0021] (10) The scoring unit is the machine learning device according to (9) above, which evaluates using at least any one of the average value, maximum value, minimum value, total value, and median value of the scores calculated for each evaluation unit.
[0022] (11) The scoring unit is the machine learning device according to any one of (8) to (10) above, which performs normalization of a plurality of layers in the dependency relationship collectively for the plurality of layers in the dependency relationship or for the components of the plurality of layers.
[0023] According to the neural network optimization method and the machine learning device of the present invention, for each layer included in the acquired neural network, a score is calculated by scoring. When the layer has a dependency relationship, based on the calculated scoring result, it is evaluated with a plurality of layers in the dependency relationship, and from the evaluation result, a pruning target included in the layer is determined, and a process of pruning the determined target is executed. By doing so, it is possible to provide a neural network optimization method and a machine learning device that can be generally applied to a neural network including layers with a dependency relationship in the structure without being limited to a specific network model and can suppress deterioration in accuracy.
Brief Description of the Drawings
[0024] [Figure 1] It is a block diagram showing a schematic configuration of an information processing apparatus according to an embodiment of the present invention. [Figure 2] It is a functional block diagram showing the configuration of a control unit. [Diagram 3] It is a schematic diagram for explaining pruning for a filter. [Figure 4] It is a schematic diagram for explaining layers in a dependency relationship. [Diagram 5] It is an example of scores of a plurality of filters in a plurality of layers that are mutually in a dependency relationship. [Figure 6]It is a schematic diagram for explaining filters to be pruned in a comparative example and an example in the case of the score of FIG. 5. [Figure 7] It is a flowchart showing a neural network optimization method according to the first embodiment. [Figure 8] It is a subroutine flowchart showing the evaluation process of step S23 in FIG. 7. [Figure 9] It is a subroutine flowchart showing the evaluation process of step S23 in FIG. 7 according to a modification. [Figure 10] It is a flowchart showing a neural network optimization method according to the second embodiment. [Figure 11] It is a graph showing the evaluation results of the accuracy in the models of each embodiment and the comparative example. [Figure 12] It is a graph showing the same evaluation results as those in FIG. 11 in another expression. [Figure 13] It is a graph showing the evaluation results of the accuracy of each embodiment.
Mode for Carrying Out the Invention
[0025] Hereinafter, embodiments of the present invention will be described with reference to the attached drawings. However, the scope of the present invention is not limited to the disclosed embodiments. In the description of the drawings, the same reference numerals are given to the same elements, and duplicate descriptions are omitted. Also, the dimensional ratios in the drawings are exaggerated for convenience of explanation and may be different from the actual ratios.
[0026] FIG. 1 is a block diagram showing a schematic configuration of an information processing apparatus 10 according to the present embodiment. The information processing apparatus 10 functions as a machine learning apparatus. The information processing apparatus 10 includes a control unit 11, a storage unit 12, an input / output unit 13, and a communication unit 14. These are mutually connected via signal lines such as a bus for signal exchange.
[0027] The control unit 11 also functions as a machine learning unit, includes multiple CPUs, multiple GPUs (Graphics Processing Units), RAM, ROM, etc., and controls each device and performs machine learning according to a program. The information processing device 10 may be an on-premise server or a cloud server using a commercial cloud service. Furthermore, some of the functions of the information processing device 10 (for example, only the functions of the machine learning unit) may be implemented by the cloud server.
[0028] The storage unit 12 is composed of a semiconductor memory or a magnetic memory such as a hard disk that stores various programs and data in advance. A learning model 200 (also referred to as a trained model) that is trained, generated, and updated by machine learning on a cloud server or on the information processing device 10 itself is stored in the storage unit 12. The learning model 200 has a neural network structure such as CNN. When performing pruning processing or a neural network optimization method, the control unit 11 or the communication unit 14 acquires the structure of the learning model stored in the storage unit 12 or acquired from the cloud server. The control unit 11, or the control unit 11 working in cooperation with the communication unit 14, functions as an acquisition unit.
[0029] The input / output unit 13 is, for example, a touch panel display that displays various information and accepts various inputs from the user. For example, the user can set a compression ratio via the input / output unit 13. Here, the compression ratio is the number of elements, such as filters, after pruning relative to the number before pruning.
[0030] The communication unit 14 is an interface for transmitting and receiving data via a network, and performs communication according to standards such as Ethernet, Bluetooth (registered trademark), and IEEE802.11 (Wi-Fi).
[0031] FIG. 2 is a functional block diagram showing the functions of the control unit 11. The control unit 11 that functions as a machine learning unit further functions as a scoring unit 111, an evaluation unit 112, a determination unit 113, and a pruning unit 114 as sub-functions.
[0032] The scoring unit 111 calculates a score by scoring for each layer included in the learning model 200 which is a neural network. The scoring unit 111 calculates a score by scoring for the entire layer, or for a weight (connection) which is an element of the layer, or for a filter which is a set of weights having spatial information. For example, in the case of a convolutional neural network (CNN), if the element of the layer is a filter which is a set of weights having spatial information, the total value of the weights of the filter is calculated as the score.
[0033] Also, in order to mitigate the influence by the size of the filter (the number of weights constituting it), well-known normalization may be performed during this score calculation. For example, the L2 norm is applied and each weight value of the filter is divided by the square root of the sum of the squares of each weight value (Euclidean distance). In the following, such normalization by a well-known method is also simply referred to as "normalization" or "general normalization" in comparison with the normalization of the present invention in the second embodiment described later.
[0034] The evaluation unit 112 analyzes the structure of the target neural network, extracts a plurality of layers in a dependency relationship, and evaluates them based on the scores calculated by the scoring unit 111. Here, the plurality of layers in a dependency relationship refers to the relationship between layers that require the same number of channels. More specifically, it is a relationship between layers that require processing with the same number of channels in the processing between multiple layers or between subsequent layers of each of these layers. For example, when performing arithmetic operations such as addition on the lower layer side, it is necessary to match the number of channels in the upper layer. Specific examples of the plurality of layers in a dependency relationship will be described later (Figure 4 described later). The evaluation unit 112 evaluates the entire layers or elements such as filters of a plurality of these layers as a set of evaluation units in the plurality of layers in a dependency relationship, and performs the evaluation between the evaluation units using an evaluation index. For example, if a set of evaluation units is a filter, at least one of the average value, maximum value, minimum value, total value, and median value of the scores of each filter is used as the evaluation index for a set of evaluation units. As the evaluation index, any one can be used alone or in combination. For example, the average value is combined with the maximum value or the minimum value. In this case, after excluding the evaluation units with a maximum value greater than or equal to the threshold value, sorting is performed in ascending order using the average value as the evaluation index. Alternatively, after excluding the evaluation units with a minimum value greater than or equal to the threshold value, sorting is performed in ascending order using the average value as the evaluation index.
[0035] The determination unit 113 determines the pruning target included in the layer based on the evaluation result of the evaluation unit 112. For example, if the evaluation unit 112 evaluates using the average value for each evaluation unit as the evaluation index, it sorts each evaluation unit in ascending order by that average value, and determines the evaluation units up to the rank corresponding to the compression rate in ascending order of the lowest average value as the pruning target. For example, when the compression rate is 90% (10% pruning), the corresponding filters of the plurality of layers in a dependency relationship are used as the evaluation units, and the filters up to the 10% rank of the lowest importance in ascending order from the lower average value are determined as the pruning target.
[0036] The pruning unit 114 executes a pruning process on the target determined by the determination unit 113 to reconstruct the neural network. After the reconstructed neural network (machine learning model) is stored in the storage unit 12, it is re-learned on the information processing device 10 itself or on a cloud server. When re-learning, it may be learned with a smaller number of epochs than the number of epochs during normal learning, that is, when learning the machine learning model before pruning.
[0037] (A plurality of layers in a dependency relationship) Next, with reference to FIGS. 3 and 4, a specific example of a plurality of layers in a dependency relationship will be described. In the following, as an example for pruning a filter and an addition (add) process as a process that requires the same number of channels, the description will be given, but it is not limited to this. As the object of pruning, the entire layer or weights may be used. Also, as a process that requires the same number of channels, for example, arithmetic operations other than addition may be used.
[0038] FIG. 3 is a schematic diagram for explaining pruning for a filter. In FIG. 3 (and subsequent figures as well), a part of the layers (n - 1, n, n + 1) of a CNN composed of a plurality of layers from the first to the (n + x)th layer is shown. Also, in FIG. 3 (and subsequent figures as well), a cross (×) mark represents a convolution operation, and a dashed filter or a dashed channel represents an element to be deleted that is the object of pruning, or a channel or filter deleted corresponding to this deleted element.
[0039] FIGS. 3(a) and (b) show the configurations of the CNN before and after pruning. As shown in FIG. 3(b), when a filter of the (n - 1)th layer, which is one layer above the nth layer, is deleted, the same number of channels of the nth layer are also deleted, and the corresponding filter of the nth layer is also deleted.
[0040] FIG. 4 is a schematic diagram for explaining layers in a dependency relationship. The upper network 210 and the lower network 220 in FIG. 4 are networks having the same structure, and addition processing is performed on the lower layer side (the subsequent stage side). The network in FIG. 3 described above corresponds to the upper and lower networks in FIG. 4. When there is addition processing on such a lower layer side, it is necessary to match the number of channels in each of the upper and lower networks 210 and 220. Therefore, when deleting the filters of one network 210 (or 220), the corresponding filters of the other network 220 (or 210) are also deleted together.
[0041] In FIGS. 3 and 4, a plurality of layers in a parallel relationship in which arithmetic operations are performed in the lower layer are shown as representative examples of a plurality of layers in a dependency relationship, but it is not limited to such a structure. There are also a plurality of layers in a dependency relationship in ResNet. In this ResNet, the last convolutional layer of each residual unit (corresponding to the building block of ResNet-18, 34 or the bottleneck of ResNet-50, 101, 152) corresponds to a plurality of layers in a dependency relationship. For example, in the example of ResNet-34 in FIG. 3 of Non-Patent Document 4 "1512.03385.pdf", among the 6 convolutional layers (3×3 conv64) in the second stage (conv2_x), the 2nd, 4th, and 6th convolutional layers immediately before the addition processing with the skip connection correspond to a plurality of layers in a dependency relationship with each other. In this case, each of the 64 filters in the 2nd, 4th, and 6th convolutional layers becomes a set of evaluation units. That is, there are 64 sets of evaluation units each composed of 3 filters. Note that when deleting the filters of a plurality of layers in a dependency relationship (the convolutional layer immediately before the addition processing), finally, a process is performed to align the number of channels on the input side of the stage (conv2_x) including these convolutional layers so that the number of channels becomes the same.
[0042] (Evaluation between Evaluation Units in a Plurality of Layers in a Dependency Relationship) Next, referring to FIGS. 5 and 6, the evaluation units in a plurality of layers in a dependency relationship and the evaluation between these evaluation units will be described. FIG. 5 is an example of scores of a plurality of filters in a plurality of layers in a dependency relationship with each other. FIG. 6 is a schematic diagram for explaining filters to be pruned in a comparative example and an embodiment in the case of the scores in FIG. 5. Layers A and B in FIGS. 5 and 6 respectively correspond to corresponding filters of a part of the networks 210 and 220 shown in FIG. 3 and the like. The filters to be applied are N filters numbered from 1 to N at the time before pruning, and N channels are output after the convolution process in this case. In the example shown in FIG. 5, scores calculated by adding the normalized weights for each of the 1st to Nth are shown. The score can take a value in the range of 0 to 1.
[0043] (Comparative Example) Referring to FIG. 5, the filter with the lowest score alone is the No. 1 filter of layer B, and the score is 0.1 (the value surrounded by the broken line). In the known pruning process disclosed by Hao L et al. in Non-Patent Document 2 described above, as shown in the comparative examples of FIGS. 5 and 6, the No. 1 filter of layer B with the lowest importance is the pruning target, and accordingly, the No. 1 filter of the corresponding layer A is also a deletion target (note that as shown in FIG. 4, when there is a processing layer of a convolution operation on the lower layer side, that filter is also deleted in the same manner). However, the No. 1 filter of layer A is not deleted after considering the importance, but is mechanically deleted based on the rule. Such a deleted filter may have a high importance. In the example shown in FIG. 5, the No. 1 filter of layer A has a score of 0.9 and a relatively high importance. If such a filter is deleted, there is a risk that the accuracy will be significantly reduced.
[0044] (Embodiment) On the one hand, in the first embodiment, as shown in the examples of FIGS. 5 and 6, corresponding filters of layers in a dependency relationship are used as a set of evaluation units, and evaluation is performed between the evaluation units. In the first embodiment, the average value of the scores is used as the evaluation index. The filter with the lowest importance among the evaluation indexes is the set of No. 2 filters, and the average value of the scores is 0.35 (the value surrounded by the dashed line). As shown in the comparative examples of FIGS. 5 and 6, the set of No. 2 filters of both layers A and B with the lowest importance as a result of comparing the evaluation indexes becomes the pruning target. Hereinafter, the specific content of such pruning processing will be described.
[0045] (Neural Network Optimization Method) FIG. 7 is a flowchart showing a neural network optimization method according to the first embodiment.
[0046] (Step S21) First, the control unit 11 functioning as an acquisition unit acquires the structure information of the neural network. The structure information of the learning model stored in the storage unit 12 or the like is acquired.
[0047] (Step S22) Next, the scoring unit 111 performs scoring on each layer or each element of each layer. For example, for a filter, normalization is performed on the filter weights and then the total value is calculated.
[0048] (Step S23) The evaluation unit 112 executes evaluation on layers in a dependency relationship. FIG. 8 is a subroutine flowchart showing the evaluation process of step S23 in FIG. 7.
[0049] (Step S301) The evaluation unit 112 sets the corresponding filters of a plurality of layers in a dependency relationship as a set of evaluation units. A set of evaluation units includes the corresponding filters of two or more layers. For example, in the example of FIG. 5, a pair of No. 1 filters for each of layers A and B becomes a set of evaluation units. In the example of FIG. 5, N sets of evaluation units from No. 1 to No. N are set.
[0050] (Step S302) The evaluation unit 112 statistically processes the scores for each set of evaluation units. In the first embodiment, as shown in FIG. 5, an average value is calculated as the statistical process. Thus, the process of FIG. 8 ends, and the process returns to the process of FIG. 7.
[0051] (Step S24) The determination unit 113 determines the pruning target included in the layer based on the evaluation result of the evaluation unit 112. Based on the evaluation index calculated by the evaluation unit 112, a plurality of sets of evaluation units are sorted in ascending order, and the evaluation units with low importance are extracted. Then, the evaluation units up to the rank corresponding to the compression rate are determined as the pruning targets.
[0052] (Step S25) The pruning unit 114 deletes the pruning targets determined in Step S24 and related lower-level elements. For example, as shown in FIG. 3(b), when it is determined to delete some of the filters in the (n - 1)th layer, the channels in the lower nth layer are deleted accordingly. Also, the filters in the nth layer are deleted along with the deletion of these channels. Thus, the pruning process ends, and then the neural network is reconstructed. After the reconstruction, the information processing device 10 retrains the neural network (learning model) with the reconstructed configuration. This retraining may be performed not by the information processing device 10 but by sending the neural network to a cloud server and having the cloud server perform it.
[0053] (Variant example) In the first embodiment shown in FIGS. 5 to 8, the average value was used as the evaluation index of the evaluation unit. However, as shown below, the maximum value may be used as the statistical process. FIG. 9 is a subroutine flowchart showing the evaluation process of step S23 in FIG. 7 according to the modification.
[0054] (Step S311) The evaluation unit 112 sets, as a set of evaluation units, the corresponding filters of a plurality of layers in a dependency relationship, in the same manner as in step S301 described above.
[0055] (Step S312) The evaluation unit 112 statistically processes the scores for each set of evaluation units and calculates the evaluation index. In this modification, the maximum value is calculated as the statistical process. For example, in the example of FIG. 5, the maximum value of the "No. 1 filter" is 0.9, the maximum value of the "No. 2 filter" is 0.4, and the maximum value of the "No. 3 filter" is 0.5. With the above, the process of FIG. 9 ends, returns to the process of FIG. 7 (return), and performs the same process as in the first embodiment. For example, the determination unit 113 sorts a plurality of sets of evaluation units in ascending order based on the evaluation index (maximum value) calculated by the evaluation unit 112, extracts the evaluation units with low importance, and determines the pruning target.
[0056] Thus, in the neural network optimization method according to this embodiment, for each layer included in the neural network, there are steps of: (a) obtaining the neural network; (b) calculating a score by scoring; (c) when the layer has a dependency relationship, evaluating among a plurality of layers in the dependency relationship based on the scoring result of step (b); (d) determining a pruning target included in the layer from the evaluation result of step (c); and (e) pruning the target determined in step (d). Thereby, a neural network optimization method can be provided that can be generally applied to a neural network including layers with a dependency relationship in the structure without being limited to a specific network model, and can suppress accuracy degradation. In particular, in the first embodiment, corresponding filters of a plurality of layers in the dependency relationship are used as a set of evaluation units, and evaluation is performed using the average value of the scores as evaluation indicators calculated for each evaluation unit. Thereby, appropriate pruning can be performed while suppressing a decrease in accuracy, and the neural network can be optimized.
[0057] (Second Embodiment) FIG. 10 is a flowchart showing a neural network optimization method according to the second embodiment. In the second embodiment described below, normalization of a plurality of layers in the dependency relationship is performed by grouping together a plurality of layers in the dependency relationship or components of the plurality of layers and performing normalization. In the flowchart of FIG. 10, for the same processing as in the flowchart of FIG. 7 in the first embodiment, the description is omitted by assigning the same step numbers.
[0058] (Step S21) This is the same processing as step S21 in FIG. 7. Analyze the structure of the neural network.
[0059] (Step S221) Here, the evaluation unit 112 sets the corresponding filters of the layers in the dependency relationship as a set of evaluation units. This process is the same as the process in step S301 in FIG. 8 and the like.
[0060] (Step S222) The scoring unit 111 summarizes, normalizes, and then scores the weights of the filters in a set of evaluation units. For example, in the example of FIG. 5, the pair of No. 1 filters of layers A and B are summarized, normalized, and then scored (hereinafter, such normalization is also referred to as "summation normalization"). For example, if one filter is 5×5 and there are a total of 50 weights for two filters, the square root of the sum of the squares of the 50 weight values of the filters is used to divide each weight value of each filter. Then, the scores are calculated by adding the weights after "summation normalization".
[0061] (Steps S23 to S25) The process here is the same as the process in FIGS. 7 and 8 (or FIG. 9) except that the normalization method is different, determining the pruning target and deleting the determined target and related elements.
[0062] In this way, the same effects as those of the first embodiment can also be obtained in the second embodiment. Also, as shown in the examples described later, the second embodiment can suppress the decrease in accuracy due to pruning more than the first embodiment.
[0063] (Examples) Hereinafter, the evaluation results of the accuracy when applying each neural network optimization method in the above-described first embodiment (FIGS. 7 and 8), modification example (FIG. 9), and second embodiment (FIG. 10) will be described. In the following examples, the evaluation index in the second embodiment uses the average value in FIG. 8. For evaluation, the above three types of pruning were performed on five types of structures of ResNet-18, 34, 50, 101, and 152, and the accuracy was evaluated with the reconfigured structure. After reconfiguration, a model retrained with a smaller number of epochs than the number of epochs when the machine learning model before pruning was trained was used.
[0064] (Test conditions) Evaluation values to compare: Top1 and Top5 output accuracy (correctness rate), Datasets used: CIFAR-100, Number of tests: 3 times Pruning compression rate: Average of three types: 90%, 75%, and 50% Pruning techniques: (1) No pruning (before pruning) (2) Example 1: The first embodiment described below, (3) Example 2: Modification (4) Example 3: Second embodiment (5) Comparative Example: The method of Non-Patent Document 2 (Hao Li et al.) shown in FIG. Evaluation method: A learning model was used that had been pruned and retrained using the three compression rates of methods (2) to (4) above. This learning model was tested three times using each dataset, and the accuracy was evaluated based on the match rate with the correct label assigned to the dataset. Accuracy was evaluated using TOP1, which evaluates the match rate between the candidate with the highest likelihood and the correct label, and TOP5, which evaluates the match rate between the candidates with the highest likelihood from the 1st to 5th highest likelihood and the correct label.
[0065] (Test results) Fig. 11 is a graph showing the evaluation results of the accuracy of the models of each embodiment and the comparative example. Fig. 12 is a graph showing the same evaluation results as Fig. 11 in a different manner. Fig. 12 shows the degree of accuracy improvement (difference in accuracy) compared to the comparative example.
[0066] The correspondence between the legends in the graphs of the evaluation results in Figures 11 and 12 and the pruning methods is as follows: Note that @1 at the end of the legend indicates the accuracy (correctness rate) of the output TOP1, and @5 indicates the accuracy of the TOP5. Before: The structure before pruning (non-pruning), Normal: (Comparative example) The comparative example of FIG. 6 corresponding to the above-mentioned Non-Patent Document 2 (Hao Li et al.), Average: First Embodiment (corresponding to FIGS. 7 and 8), Maximum: Modified Example (corresponding to FIGS. 7 and 9).
[0067] As shown in FIGS. 11 and 12, the structure with pruning according to the first embodiment (average value) and the modified example (maximum value) shows an improvement in accuracy compared to the comparative example of the conventional method. As shown in FIG. 11, in the system with fewer layers of ResnNet-18 and 34, the improvement effect is significant compared to the comparative example, but in the system with more layers such as ResNet-50, 101, and 152, the improvement effect on accuracy tends to be small.
[0068] FIG. 13 is a graph showing the accuracy evaluation results of each embodiment. The differences due to the normalization methods are compared. The accuracy degradation (difference) in the structure of each embodiment with respect to the accuracy before pruning is plotted. The correspondence between the legends in the graphs of each evaluation result in FIG. 13 and the pruning methods is as follows. Note that in the legends, @1 and @5 indicate the accuracies of output TOP1 and 5, respectively, similar to FIGS. 11 and the like. Without Norm: First Embodiment, applying "general normalization", With Norm: Second Embodiment (FIG. 10), applying "sum normalization".
[0069] As shown in FIG. 13, it can be seen that the structure of the second embodiment applying sum normalization has a smaller degree of accuracy degradation before and after pruning and better performance compared to the structure applying general normalization.
[0070] The configurations of the machine learning device and the information processing device described above explain the main configurations in describing the features of the above embodiments, and are not limited to the above configurations, and various modifications can be made within the scope of the claims. Also, it does not exclude the configurations provided in general machine learning devices or information processing devices.
[0071] In addition, in the above flowchart, some steps may be omitted, and other steps may be added. Also, some parts of each step may be reordered, executed simultaneously, or one step may be split into multiple steps for execution.
[0072] Moreover, the means and methods for performing various processes in the information processing apparatus 10 described above can be realized by either a dedicated hardware circuit or a programmed computer. The above program may be provided by a computer-readable recording medium such as a USB memory or a DVD (Digital Versatile Disc)-ROM, or may be provided online via a network such as the Internet. In this case, the program recorded on the computer-readable recording medium is usually transferred and stored in a storage unit such as a hard disk. Also, the above program may be provided as a single application software or incorporated into the software of the apparatus as one of its functions.
Explanation of Reference Numerals
[0073] 10 Information processing apparatus (machine learning apparatus) 11 Control unit 111 Scoring unit 112 Evaluation unit 113 Decision unit 114 Pruning unit 12 Storage unit 13 Input / output unit 14 Communication unit 200 Learning model
Claims
1. A method for pruning a neural network by a computer, comprising: Step (a) of obtaining a neural network; Step (b) of calculating a score for each layer included in the neural network by scoring; When the layers have a dependency relationship, step (c) of evaluating among a plurality of layers in the dependency relationship based on the scoring result of step (b); Step (d) of determining a pruning target included in the layer from the evaluation result of step (c); Step (e) of pruning the target determined in step (d), and executing a process including the steps, a neural network optimization method.
2. The plurality of layers in the dependency relationship are: The neural network optimization method according to claim 1, wherein the layers are in a relationship that requires processing with the same number of channels in the processing between the plurality of layers or between subsequent layers of each of the plurality of layers.
3. The neural network optimization method according to claim 1 or 2, wherein the target is a filter included in each of a plurality of layers in a dependency relationship.
4. In step (c), the neural network optimization method according to any one of claims 1 to 3, wherein a plurality of layers in a dependency relationship are used as one set of evaluation units, and the evaluation is performed between the evaluation units.
5. In step (c), the neural network optimization method according to claim 4, wherein the evaluation is performed using at least any one of an average value, a maximum value, a minimum value, a total value, and a median value of the scores calculated for each evaluation unit.
6. Further comprising step (f) of normalizing the layer, In step (f), the neural network optimization method according to any one of claims 1 to 5, wherein the normalization of the plurality of layers in the dependency relationship is performed collectively between the plurality of layers in the dependency relationship or between components of the plurality of layers.
7. A program for causing a computer to execute the neural network optimization method according to any one of claims 1 to 6.
8. An acquisition unit for acquiring a neural network; For each layer included in the neural network, a scoring unit that calculates a score by scoring, when the layer has a dependency relationship, an evaluation unit that evaluates among a plurality of layers in the dependency relationship based on the scoring result of the scoring unit, a determination unit that determines a pruning target included in the layer from the evaluation result by the evaluation unit, a pruning unit that prunes the target determined by the determination unit, a machine learning device.
9. The scoring unit uses a plurality of layers in a dependency relationship as a set of evaluation units and executes the evaluation between the evaluation units. The machine learning device according to claim 8.
10. The scoring unit according to claim 9, wherein the scoring unit evaluates using at least any one of an average value, a maximum value, a minimum value, a total value, and a median value of the scores calculated for each evaluation unit.
11. The scoring unit according to any one of claims 8 to 10, wherein the normalization of the plurality of layers in the dependency relationship is collectively executed between the plurality of layers in the dependency relationship or between the components of the plurality of layers.
Citation Information
Patent Citations
Learning device, learning method, and learning program
JP2019185275A
Method and system for performing machine learning
JP2021522574A
Data processing system comprising neural network
WO2019135274A1