Pruning method, pruning device, machine learning method, and trained model
The proposed pruning method for neural networks addresses the issue of performance degradation by selectively retaining high-importance kernels during pruning, ensuring the model's accuracy and efficiency are maintained.
Patent Information
- Application Number
- JP2023194179
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-27
AI Technical Summary
When pruning a neural network on a per-channel basis, there is a risk of removing kernels with high importance, leading to performance degradation of the learned model.
A pruning method that determines the removal target for each filter of the convolutional layer on a per-channel basis, extracts kernels with high importance from the targeted channels for removal, and reconstructs the filter using these extracted kernels and those from non-targeted channels.
This approach effectively suppresses performance degradation of the model associated with pruning by ensuring that important kernels are retained, thereby maintaining the model's accuracy and efficiency.
Smart Images

Figure 2025080845000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for pruning a neural network.
Background Art
[0002] Conventionally, a neural network (trained model) trained by deep learning has been used for image recognition such as image classification and object detection. The neural network tends to obtain high performance such as high inference accuracy by making its configuration complex. However, when the configuration of the neural network becomes complex, an increase in the number of operations and the memory size becomes a problem when the neural network is executed by a computer.
[0003] Pruning (branch cutting) is known as a technique for reducing the number of operations (i.e., speeding up) and reducing the memory size (i.e., reducing the model size). Various pruning techniques have been proposed conventionally. For example, Patent Document 1 discloses performing channel removal as a pruning technique.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] When pruning is performed to remove weight data on a per-channel basis, kernels included in multiple channels are grouped and removed. In channels that are targeted for pruning because they have low importance when viewed as a whole, there may be kernels that have high importance when viewed individually. That is, when pruning is performed on a per-channel basis, there is a possibility of removing kernels with high importance. As a result, the performance of the learned model obtained after pruning may be degraded.
[0006] In view of the above points, an object of the present invention is to provide a technique capable of suppressing performance degradation of a model associated with pruning.
Means for Solving the Problems
[0007] An exemplary pruning method of the present invention is a pruning method for a neural network including a convolutional layer, which determines a removal target for each filter of the convolutional layer on a per-channel basis, extracts kernels with high importance from among a plurality of kernels included in the channels targeted for removal, and reconstructs the filter using the extracted kernels and the kernels included in the channels that are left without being targeted for removal in the filter.
Effects of the Invention
[0008] According to an exemplary aspect of the present invention, performance degradation of a model associated with pruning can be suppressed.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0010] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings.
[0011] <1. Machine Learning System> FIG. 1 is a block diagram showing a schematic configuration of a machine learning system 100 according to an embodiment of the present invention. As shown in FIG. 1, the machine learning system 100 includes a machine learning device 1 and an edge device 2.
[0012] The machine learning device 1 is a computer device, which is a so-called server device. Note that the server device may be a physical server or a cloud server. The machine learning device 1 generates a learned model 3. The machine learning device 1 provides the generated learned model 3 to the edge device 2 connected via a communication network such as the Internet. The edge device 2 acquires the learned model 3 from the machine learning device 1 by so-called downloading.
[0013] In addition, the learned model 3 generated by the machine learning device 1 may be recorded on a recording medium such as an optical recording medium or a magnetic recording medium, and provided to the edge device 2 via the recording medium.
[0014] The edge device 2 is, for example, an in-vehicle device, a smartphone, a personal computer, an IoT (Internet of Things) home appliance, or the like. The edge device 2 stores the learned model 3 obtained from the machine learning device 1 in a memory (not shown) and uses it as appropriate. The memory stores the structure and parameters of the learned model, as well as the code instructions for executing the model. The learned model 3 is, for example, an AI (Artificial Intelligence) model for performing image recognition such as image classification or object detection.
[0015] The above-mentioned AI model for image recognition is provided in, for example, an in-vehicle device. The in-vehicle device inputs, for example, a captured image of the surroundings of the vehicle into the AI model and attempts to detect a preset detection target. The detection target is, for example, a vehicle, a person, or the like. The in-vehicle device performs driving control of the vehicle, notification control for notifying of danger, etc. according to the detection result of the detection target by the AI model.
[0016] <2. Machine Learning Device> As shown in FIG. 1, the machine learning device 1 includes a controller 11 and a memory 12.
[0017] The controller 11 includes an arithmetic circuit that performs arithmetic processing. The arithmetic circuit is specifically composed of a processor. The processor includes, for example, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The controller 11 may be composed of one processor or a plurality of processors. When composed of a plurality of processors, those processors are provided so as to be communicable with each other.
[0018] The memory 12 is composed of a volatile memory and a non-volatile memory. The volatile memory is specifically a RAM (Random Access Memory). The non-volatile memory is specifically a ROM (Read Only Memory). The non-volatile memory may also be, for example, a flash memory or a hard disk drive. Programs and data that can be read by a computer are stored in the non-volatile memory.
[0019] As shown in FIG. 1, as its functions, the controller 11 includes an acquisition unit 111, a learning unit 112, a pruning unit 113, and an output unit 114. The functions of the controller 11 are realized by the processor executing arithmetic processing according to a program stored in the memory 12. The number of programs for realizing the functions of the controller 11 may be singular or plural.
[0020] Note that the program stored in the memory 12 is a computer program for causing a computer to realize the functions of the controller 11. Such a computer program may be provided, for example, by a computer-readable non-volatile recording medium. The non-volatile recording medium may be, for example, in addition to the above-mentioned non-volatile memory, an optical recording medium (such as an optical disk), a magneto-optical recording medium (such as a magneto-optical disk), a USB memory, or an SD card. As another example, the computer program may be provided from a program-providing server via a communication line such as the Internet, that is, provided by so-called downloading.
[0021] Further, each of the functional units 111 to 114 may be implemented by one program, or may be implemented by separate programs for each functional unit, for example. Also, as described above, each of the functional units 111 to 114 may be implemented by causing a processor to execute a program, that is, by software, but may also be implemented by other methods. Each of the functional units 111 to 114 may be implemented using, for example, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or the like. That is, each of the functional units 111 to 114 may be implemented by hardware using a dedicated IC or the like. Also, each of the functional units 111 to 114 may be implemented by using a combination of software and hardware. Also, each of the functional units 111 to 114 is a conceptual component. The functions executed by one component may be distributed among a plurality of components. Also, the functions of a plurality of components may be integrated into one component.
[0022] The acquisition unit 111 acquires an unlearned model and machine learning data, and stores them in the memory 12 as appropriate. The acquisition unit 111 receives, for example, the unlearned model and the machine learning data from a device provided outside the machine learning device 1 via a communication network (not shown). Note that at least one of the unlearned model and the machine learning data may be acquired not from outside the machine learning device 1 but by being generated by the machine learning device 1 itself.
[0023] The unlearned model is a neural network including unlearned parameters such as weight data. Specifically, the neural network is a convolutional neural network including a plurality of convolutional layers. The structure of the convolutional neural network is appropriately determined by a designer, a developer, or the like. The machine learning data is a learning dataset used for machine learning (training) of the unlearned model. As an example, assuming machine learning of a neural network for performing image recognition, the machine learning data includes a plurality of pairs of learning image data and teacher data including correct labels for the learning image data.
[0024] The learning unit 112 performs learning of an unlearned model (a learning model before learning is completed) using machine learning data. Further, the learning unit 112 performs re-learning on the model (weight-reduced model) obtained after the pruning process described later using machine learning data. The machine learning data used for the first learning and the machine learning data used for the re-learning may be the same or different.
[0025] In this specification, the learned model obtained after the above-described re-learning is simply referred to as the "learned model". In order to avoid confusion with the "learned model", the learned model obtained by the machine learning device 1 before pruning is referred to as the "pre-learned model". The learned model obtained by learning the unlearned model in this embodiment is a pre-learned model.
[0026] Specifically, the pre-learned model is obtained by updating parameters such as weight data included in the unlearned model. The pre-learned model may be regarded as a model that is the result of parameter changes through machine learning processing from the unlearned model. Further, specifically, the learned model is obtained by updating parameters such as weight data included in the weight-reduced model obtained after the pruning process (details will be described later). The learned model may be regarded as a model that is the result of parameter changes through machine learning processing from the weight-reduced model. The machine learning processing may be realized by various known methods such as the error backpropagation method.
[0027] In this embodiment, the pre-learned model is configured to be obtained by the learning unit 112 performing learning on the unlearned model, but this is an example. The learning unit 112 may be configured to perform only re-learning and not perform learning of the unlearned model. When such a configuration is adopted, the acquisition unit 111 may obtain the pre-learned model from an external device.
[0028] The pruning unit 113 performs pruning processing on a neural network including a convolutional layer. That is, in this embodiment, the machine learning device 1 also serves as a pruning device. However, the pruning processing executed by the pruning unit 113 may be executed by a device provided separately from the machine learning device including the learning unit. In the case of such a configuration, the machine learning device and the pruning device that executes the pruning processing may be provided so as to be communicable with each other and configured to exchange information with each other.
[0029] Specifically, the pruning unit 113 performs pruning processing on a pre-trained model. More specifically, the pruning unit 113 determines a target for removal in units of channels with respect to the weight data of the convolutional layer included in the pre-trained model. In response to the determination of the target for removal, for example, a part of the weight data of the pre-trained model stored in the memory 12 is deleted. Hereinafter, a method for determining a target for removal in units of channels with respect to the weight data will be described. Before that, a schematic configuration of the convolutional neural network constituting the pre-trained model will be described with reference to FIG. 2. FIG. 2 is a diagram for explaining a schematic configuration of the convolutional neural network.
[0030] FIG. 2 shows a partial configuration of the convolutional neural network. As shown in FIG. 2, the convolutional neural network includes a plurality of convolutional layers. In the example shown in FIG. 2, the convolutional neural network includes three convolutional layers of the (n−1)th layer, the nth layer, and the (n + 1)th layer. In the example shown in FIG. 2, n is an integer greater than 2.
[0031] For example, the (n-1)th layer performs a convolution process using the feature map 10 obtained from the immediately preceding convolution layer (not shown) as input, and obtains a new feature map 10. Note that the feature map 10 is tensor data output by a convolution operation and is data including feature amounts extracted by the convolution operation. The nth layer performs a convolution process using the feature map 10 obtained from the (n-1)th layer as input, and obtains a new feature map. The (n+1)th layer performs a convolution process using the feature map 10 obtained from the nth layer as input, and obtains a new feature map. Thereafter, the same process is repeated according to the number of convolution layers. Note that when the (n-1)th layer is the first layer (the first layer), in image recognition, the input is not a feature map but an image (image data).
[0032] As shown in FIG. 2, each convolution layer includes a filter 4 having weight data in which Cout×Cin kernels 20 of size K×K are arranged. The kernel 20 of size K×K has K×K weights (parameters). Cout indicates the number of output channels, which is the number of feature maps 10 output by performing a convolution process on the input feature map 10, and matches the number of feature maps 10 input in the next convolution layer. Cin indicates the number of input channels, which matches the number of feature maps 10 input to each convolution layer.
[0033] In each convolution layer, feature maps 10 of the number of output channels are generated by a convolution process using the filter 4. Specifically, for each output channel, each convolution operation using each feature map 10 and the kernel 20 prepared for each feature map 10 (for each input channel), and the addition process of each operation result are performed, and one feature map is generated.
[0034] In the example shown in FIG. 2, the filter 4 in the (n - 1)th layer has (Cout, Cin, K, K) = (5, 3, 3, 3). The kernel 20 has a size of 3×3 and has nine weights. For each of the five output channels, a convolution operation is performed using each of the three feature maps 10 and the three kernels 20 prepared for each feature map 10 (input channel), and the results of each operation are added together, generating one feature map 10. Since the number of channels of the output channels is five, five feature maps 10 are output from the (n - 1)th layer. In other words, five feature maps 10 are input to the nth layer.
[0035] The filter 4 in the nth layer has (Cout, Cin, K, K) = (3, 5, 3, 3). The kernel 20 has a size of 3×3. For each of the three output channels, a convolution operation is performed using each of the five feature maps 10 and the five kernels 20 prepared for each feature map 10 (input channel), and the results of each operation are added together, generating one feature map 10. Since the number of channels of the output channels is three, three feature maps 10 are output from the nth layer. In other words, three feature maps 10 are input to the (n + 1)th layer.
[0036] The filter 4 in the (n + 1)th layer has (Cout, Cin, K, K) = (4, 3, 3, 3). The kernel 20 has a size of 3×3. For each of the four output channels, a convolution operation is performed using each of the three feature maps 10 and the three kernels 20 prepared for each feature map 10 (input channel), and the results of each operation are added together, generating one feature map 10. Since the number of channels of the output channels is four, four feature maps 10 are output from the (n + 1)th layer.
[0037] FIG. 3 is a diagram for explaining an example of a method for removing the filter 4 in units of channels. FIG. 3 assumes removing the filter 4 in units of channels for the convolutional neural network shown in FIG. 2. The example shown in FIG. 3 is a method (hereinafter referred to as the first pruning method) of removing the filter 4 of the convolutional layer in units of input channels (the smallest unit). That is, in FIG. 3, the determination of the removal target in units of channels for the filter 4 is performed using the first pruning method of removing in units of input channels. Note that the first pruning method is a so-called channel pruning method.
[0038] In the first pruning, attention is paid to the input channels of the filter 4 included in each convolutional layer of the neural network, and an index value indicating the importance is obtained for each input channel, which is a group of kernels that are convolutionally operated on each feature map. Then, for all the input channels for which the index value of the importance has been obtained, ranking is performed such that the higher the index value, the more important the channel. The index value indicating the importance is, for example, the L1 norm, the L2 norm, or the maximum value of the absolute value, etc. The input channels with low importance are targeted for removal as being unnecessary. The input channels targeted for removal are, for example, a certain ratio starting from the lower ranks of the importance. As another example, the input channels targeted for removal may be the input channels with an index value indicating importance smaller than a preset threshold.
[0039] Note that the index value of the importance of each input channel is specifically obtained, for example, as follows. For each kernel 20, the sum of the absolute values of the weights is calculated. For example, when the size of the kernel 20 is 3×3, the sum of the absolute values of 9 weights is calculated. For each input channel of each convolutional layer, the total sum of the absolute value sums of the obtained kernels (the total sum of Cout absolute value sums) is calculated, and the calculated value is used as the index value of the importance of the input channel.
[0040] In the example shown in FIG. 3, in the n-th layer (convolutional layer), the first input channel 4-1 (corresponding to the feature map and input channel indicated by the dashed line) to which the convolutional operation is performed on the first feature map 10-1 has a low importance and is a target for removal. That is, in the n-th layer, the number of input channels is reduced by one by the first pruning, and the size of the filter 4 (Cout, Cin, K, K) changes from (3, 5, 3, 3) to (3, 4, 3, 3).
[0041] Since one input channel of the filter 4 is removed in the n-th layer, the feature map 10-1 input to the n-th layer becomes unnecessary data. Accordingly, one output channel 4-2 for outputting the feature map 10-1 in the (n-1)-th layer is also removed. That is, in the (n-1)-th layer, the number of output channels is reduced by one due to the reduction of one input channel in the n-th layer, and the size of the filter 4 (Cout, Cin, K, K) changes from (5, 3, 3, 3) to (4, 3, 3, 3). In FIG. 3, the data of the kernel 20 group indicated by the dashed line in the (n-1)-th layer corresponds to the data of the output channel 4-2 to be removed.
[0042] FIG. 4 is a diagram for explaining another example of a method for removing the filter 4 in units of channels. FIG. 4 assumes removing the filter 4 in units of channels for the convolutional neural network shown in FIG. 2. The example shown in FIG. 4 is a method (hereinafter referred to as the second pruning method) of removing the filter 4 of the convolutional layer in units of output channels (the smallest unit). That is, in FIG. 4, the determination of the removal target in units of channels for the filter 4 is performed using the second pruning method of removing in units of output channels. The second pruning method is a so-called filter pruning method. The dashed line 4-3 in FIG. 4 indicates the output channel to be removed, and the output channel is composed of a group of kernels in which the kernels 20 are arranged in the direction of the input channels.
[0043] In the second pruning, attention is paid to the output channels of the filters 4 of each convolutional layer included in the neural network, and an index value indicating the importance is obtained for each output channel, which is a group of kernels that are convolutionally operated to calculate each output feature map. Then, for all the output channels for which the index value of the importance has been obtained, ranking is performed such that the channels with larger index values are more important. The index value indicating the importance is, for example, the L1 norm, the L2 norm, or the maximum value of the absolute values. Output channels with low importance are targeted for removal as being unnecessary. The output channels targeted for removal are, for example, a certain ratio starting from the ones with lower importance ranks. As another example, the output channels targeted for removal may be those with an index value indicating importance smaller than a preset threshold value.
[0044] Note that the index value of the importance of each output channel is specifically obtained, for example, as follows. For each kernel 20, the sum of the absolute values of the weights is calculated. For example, when the size of the kernel 20 is 3×3, the sum of the absolute values of 9 weights is calculated. For each output channel of each convolutional layer, the total sum of the absolute value sums of the obtained kernels (the total sum of Cin absolute value sums) is calculated, and this calculated value is used as the index value of the importance of the output channel.
[0045] In the example shown in FIG. 4, in the n-th layer (convolutional layer), the first output channel (corresponding to the output channel indicated by the broken line 4-3), which is a group of kernels that are convolutionally operated to calculate the first output feature map 10-2 (corresponding to the feature map indicated by the broken line input to the (n + 1)-th layer), has low importance and is targeted for removal. That is, in the n-th layer, the number of output channels is reduced by one by the second pruning, and the size of the filter 4 (Cout, Cin, K, K) changes from (3, 5, 3, 3) to (2, 5, 3, 3).
[0046] In the n-th layer, since one output channel of the filter 4 is removed, the number of feature maps 10 input to the (n + 1)-th layer is reduced by one. Accordingly, the number of kernel groups (input channels) convolved with the feature map 10-2 in the (n + 1)-th layer is also reduced by one. That is, in the (n + 1)-th layer, the number of input channels is reduced by one due to the reduction of one output channel in the n-th layer, and the size (Cout, Cin, K, K) of the filter 4 changes from (4, 3, 3, 3) to (4, 2, 3, 3). In FIG. 4, the data of the kernel 20 indicated by the dashed line in the (n + 1)-th layer corresponds to the data of the input channel 4-4 to be removed.
[0047] In the method of determining the pruning target in units of channels (the first pruning method and the second pruning method) described above, a plurality of kernels 20 included in each channel are removed together. In a channel that is determined to be a pruning target because its overall importance is low, there may be a kernel 20 with high importance when viewed individually. That is, if the determination of the pruning target is performed in units of channels, it may cause a performance degradation of the model as a result of removing a kernel 20 with high importance.
[0048] In the present embodiment, in order to suppress the performance degradation of the model as described above, the pruning unit 113 reconstructs the filter 4 for each convolutional layer using the kernels 20 with high importance included in the channels to be removed. As a result, the important kernels 20 included in the channels to be removed can be finally left, and the performance degradation of the model after the pruning process can be suppressed. Details of the reconstruction of the filter 4 will be described later.
[0049] Returning to FIG. 1, the output unit 114 transmits the output data to the edge device 2. The output unit 114 may store the output data in the memory 12 and manage it so that the edge device 2 can acquire it. The output data is data of a learned model obtained by re-training a lightweight model obtained by performing pruning processing on a pre-trained model. The data of the learned model includes, for example, the structure and parameters of the learned model, as well as code instructions for executing the model.
[0050] <3. Machine Learning Method> Next, the machine learning method executed by the machine learning device 1 will be described. Specifically, the machine learning method is a machine learning method for a neural network including a plurality of convolutional layers.
[0051] [3-1. Overview] FIG. 5 is a flowchart showing an overview of the machine learning method according to an embodiment of the present invention. The flow shown in FIG. 5 starts at an appropriate timing after the acquisition unit 111 acquires an unlearned model and machine learning data.
[0052] In step S1, the learning unit 112 learns the unlearned model. That is, the machine learning method of the present embodiment includes learning an unlearned neural network. The learning is performed using a known method such as the error backpropagation method using machine learning data including teacher data as described above. When the learning is completed and a pre-trained model is obtained, the process proceeds to the next step S2.
[0053] In step S2, the pruning unit 113 executes a pruning process. The pruning process includes a process of determining a removal target for each channel with respect to the filter 4 in the pre-trained model. That is, the machine learning method of the present embodiment includes determining a removal target for each channel with respect to the filter 4 in the neural network after learning. Specifically, the first pruning or the second pruning described above is used as a method for determining the removal target for each channel with respect to the filter 4. Further, the pruning process includes the reconstruction process of the filter 4 that is performed after the determination of the removal target in the filter 4 described above. Details of the reconstruction process of the filter 4 will be described later. When the pruning process is completed, the process proceeds to the next step S3.
[0054] In step S3, the learning unit 112 re-learns the lightweight model obtained by pruning the pre-trained model. The re-learning is performed using a known method such as the error backpropagation method using machine learning data including teacher data, similar to the learning in step S1. When the re-learning is completed, the process proceeds to the next step S4.
[0055] In step S4, the pruning unit 113 determines whether to end the pruning process. The pruning unit 113 determines to end the pruning process when the processing accuracy of the learned model obtained after re-learning is equal to or lower than a preset threshold. Note that the processing accuracy is, for example, the classification accuracy of image classification or the detection accuracy of object detection. The threshold is obtained by experiments or the like and is set to a value slightly higher than the minimum required classification accuracy or detection accuracy determined according to the usage purpose of the model, for example. When it is determined to end the pruning process (Yes in step S4), the flow shown in FIG. 5 ends. Thereby, the learned model, which is the target product in the machine learning method shown in FIG. 5, is obtained. The generated learned model is appropriately distributed to the edge device 2 by the output unit 114. When it is determined not to end the pruning process (No in step S4), the process returns to step S2 and the processes after step S2 are performed.
[0056] Note that the method for determining the end of the pruning process may be other than the method shown above. For example, the pruning unit 113 may determine to end the pruning process when the execution time of the task (such as image classification) of the learned model obtained after re-learning is equal to or less than a preset target value.
[0057] [3-2. Pruning Method] Next, the pruning method of the present embodiment will be described in detail. The pruning method of the present embodiment is a pruning method for a neural network including a convolutional layer, and is a method for realizing the above-described pruning process. Hereinafter, the pruning method will be described by giving two examples, a first example and a second example.
[0058] (3-2-1. First Example) The first example assumes a case where the first pruning method is used as the pruning method. FIG. 6 is a flowchart showing the flow of the pruning method according to the first example. Note that the process shown in FIG. 6 is a detailed example of the process of step S2 shown in FIG. 5. Among the processes shown in FIG. 6, the processes of step S21 and step S22 are the same as the content of the first pruning described above with reference to FIG. 3, and thus the description thereof will be simplified.
[0059] In step S21, the pruning unit 113 determines the importance for each input channel of the filter 4 included in each convolutional layer included in the neural network (pre-trained model) and ranks them. Specifically, the higher the importance, the higher the rank. When the ranking of the importance of all input channels in the neural network is completed, the process proceeds to the next step S22.
[0060] In step S22, the pruning unit 113 determines that the input channels with the lowest N% importance are to be removed. Note that N% is appropriately determined through experiments or simulations. The data of the input channels determined to be removed is deleted from the weight data of the pre-trained model stored in the memory 12. However, the data of the channels determined to be removed is separately stored in the memory 12 in consideration of subsequent processing.
[0061] Note that the data of the input channels determined to be removed may simply be stored in the memory 12 without being deleted from the weight data of the pre-trained model stored in the memory 12. In this embodiment, for the sake of easy understanding of the processing flow, it is assumed that the data of the input channels determined to be removed is deleted from the weight data of the pre-trained model stored in the memory 12. Also, as described above, with the removal of the input channels, a removal target occurs in the output channels in the filter 4 of the immediately preceding convolutional layer. The data of the output channels determined to be removed is also stored in the memory 12 in consideration of subsequent processing. When the channels to be removed are determined, the processing proceeds to the next step S23.
[0062] In step S23, the pruning unit 113 sets the variable α to X. X is the number of convolutional layers included in the pre-trained model. The variable α corresponds to the number of the convolutional layer that reconstructs the filter 4. In this example, the numbers of the convolutional layers are sequentially numbered from the input side to the output side of the neural network. The number of the convolutional layer located closest to the input side is "1". The number of the convolutional layer located closest to the output side is "X". Through the processing in step S23, since α = X, the convolutional layer existing closest to the output side becomes the convolutional layer that first reconstructs the filter 4. When the process of setting the variable α to X is performed, the processing proceeds to the next step S24.
[0063] In step S24, the pruning unit 113 determines the importance of each kernel 20 included in the input channels that were targeted for removal in the previous first pruning in the α-th layer (convolutional layer), and ranks them. Specifically, the importance of the kernel 20 is obtained by determining an index value of importance. The index value of importance is, for example, the L1 norm, the L2 norm, or the maximum value of the absolute values. For example, when the size of the kernel 20 is 3×3, the index value of importance of the kernel 20 may be the sum of the absolute values of nine weights. The higher the index value of importance, the higher the ranking of importance. When the ranking of the importance of the kernel 20 is completed, the process proceeds to the next step S25.
[0064] In step S25, the pruning unit 113 extracts the kernels 20 with the top M% importance among the plurality of kernels 20 whose importance was ranked in step S24. That is, in the pruning method of this example, it includes extracting the kernels 20 with high importance from among the plurality of kernels 20 included in the channels targeted for removal. More specifically, the pruning method of this example includes extracting, as the kernels 20 with high importance, the kernels 20 with a certain top percentage of importance from among the plurality of kernels 20 included in the channels targeted for removal. Note that the above-mentioned M% is appropriately determined by experiments or simulations. When the extraction of the kernels 20 with high importance is completed, the process proceeds to the next step S26.
[0065] In step S26, the pruning unit 113 reconstructs the filter 4 (the filter 4 of the α-th layer) by putting the extracted kernel 20 out of the plurality of kernels 20 included in the channel to be removed into the filter 4. More specifically, in the pruning method of this example, at least a part of the extracted kernels 20 is used to form the filter 4 after reconstruction. The reconstruction of the filter 4 is performed using the input channels that are left without being targeted for removal in the filter 4 and the kernels with high importance included in the input channels targeted for removal. With this configuration, according to the conventionally known first pruning method, the important kernels 20 included in the input channels targeted for removal can be left in the filter 4, and the performance degradation of the model after pruning can be suppressed.
[0066] Note that there may be a convolutional layer in which no input channel to be removed is generated during the first pruning executed by steps S21 and S22. For such a convolutional layer, since there is no input channel to be removed, the processes of steps S24, S25, and S26 are substantially skipped.
[0067] Here, a specific example will be given to explain the reconstruction of the filter 4 in step S26.
[0068] FIG. 7 is a schematic diagram showing the state of the filter 4 of the X-th layer (the convolutional layer closest to the output side) at each stage. In FIG. 7, (a) shows the state of the filter 4 before the first pruning in the X-th layer. (b) shows the state of the filter 4 after the first pruning in the X-th layer. (b) shows the state immediately after the first pruning, before the kernels 20 with high importance are returned to the filter 4. (c) shows the state of the filter 4 after the reconstruction in the X-th layer. The state of the filter 4 changes in the order of (a), (b), and (c).
[0069] In the example shown in FIG. 7, the input channels are reduced by the first pruning, and the size (Cout, Cin, K, K) of the filter 4 in the X-th layer changes from (64, 256, 3, 3) to (64, 16, 3, 3). That is, 240 (= 256 - 16) input channels are targeted for removal by the first pruning. In such a case, an explanation will be given as to how the above-mentioned filter 4 is reconstructed.
[0070] When 240 input channels are targeted for removal, the size (Cout, Cin, K, K) of the data to be removed is (64, 240, 3, 3). That is, the number of kernels 20 to be removed is 15360 (= 240 × 64). Among the kernels 20 to be removed, the top M% of the kernels 20 with high importance extracted in the process of step S25 are candidates to be returned to the filter 4 including the remaining channels without being removed.
[0071] Here, assuming that M% is 10%, 1536 (= 15630 × 0.1) kernels 20 are candidates to be returned to the filter 4 including the remaining channels without being removed. When returning the kernels 20 to the filter 4, it is necessary to make the number of kernels 20 returned the same for each output channel. For this purpose, 24 (= 1536 ÷ 64) kernels are returned for each output channel of the remaining filter 4. That is, the reconstruction of the filter 4 includes a process of distributing the extracted kernels 20 so that the number is the same for each output channel. As a result, the size (Cout, Cin, K, K) of the reconstructed filter 4 is (64, 40, 3, 3). Here, the number of input channels "40" is the sum of the number of input channels "16" after the first pruning is performed and the number of kernels returned to the filter 4 "24".
[0072] Note that, in the above example, all of the kernels 20 (candidate kernels) that are extracted in the process of step S25 and are candidates to be returned to filter 4 were returned to filter 4, but this is merely an illustration. If the number of candidate kernels is not divisible by the number of output channels, then some of the candidate kernels will be removed without being returned to filter 4. As another example, when extracting kernel 20 in step S25, the kernel 20 may be extracted so that no remainder occurs.
[0073] Also, when returning kernel 20 to filter 4 for the reconstruction of filter 4, for example, the kernel 20 may be returned to filter 4 so that the kernels 20 with high importance do not bias towards a specific output channel. For example, the kernel 20 may be returned to filter 4 while distributing them to each output channel in order from the kernel 20 with high importance. However, it is not limited to this, and the kernel 20 may be returned so that the kernels 20 with high importance bias towards a specific output channel.
[0074] Also, when returning kernel 20 to each output channel, for example, without changing the order of the kernels 20 that were left without being removed in the first pruning, the kernels 20 to be returned may be arranged in order at the head side or the tail side of the order. However, a configuration may also be adopted in which the order of the kernels 20 that were left without being removed in the first pruning is changed and the kernel 20 is returned.
[0075] Also, in a convolutional neural network, the number of input channels of filter 4 in a certain convolutional layer must be the same as the number of output channels of filter 4 in the convolutional layer immediately before the said convolutional layer. For this reason, if the number of input channels is increased by reconstructing filter 4, then the reconstruction of filter 4 in the immediately preceding convolutional layer is also necessary. A specific example will be given and explained below.
[0076] FIG. 8 is a diagram showing the relationship of filter 4 between the X-th layer and the (X-1)-th layer. Note that the numbers shown in parentheses in FIG. 8 indicate the size (Cout, Cin, K, K) of filter 4. The X-th layer is the same as the X-th layer in FIG. 7 and is the convolutional layer existing on the most output side of the convolutional neural network. The (X-1)-th layer is the convolutional layer immediately before the X-th layer. Also, in FIG. 8, (a) before the first pruning is performed, (b) after the first pruning is performed, and (c) after the reconstruction is performed mean the same state as in the case of FIG. 7 described above.
[0077] As shown in FIG. 8, in any of the cases before the first pruning is performed, after the first pruning is performed, and after the reconstruction is performed, the number of input channels of filter 4 in the X-th layer is the same as the number of output channels of filter 4 in the immediately preceding (X-1)-th layer. In the (X-1)-th layer, the size (Cout, Cin, K, K) of filter 4 after the first pruning is (16, 32, 3, 3). That is, the number of output channels of filter 4 in the (X-1)-th layer after the first pruning is "16", which does not match the number of input channels "40" of filter 4 in the X-th layer after the reconstruction. For this reason, a reconstruction of filter 4 is performed to change the number of output channels in the (X-1)-th layer from "16" to "40".
[0078] In the (X-1)-th layer, in order to increase the number of output channels from "16" to "40", since the number of input channels is "32", 768 (= (40 - 16) × 32) kernels are required. These kernels 20 are replenished using the kernels 20 removed during the first pruning. In the (X-1)-th layer, 7680 (= (256 - 16) × 32) kernels 20 are targeted for removal by the first pruning. From among these kernels 20 targeted for removal, 768 kernels 20 are secured in order from the kernels 20 with high importance and reconstructed into a filter 4 with a size (Cout, Cin, K, K) of (40, 32, 3, 3).
[0079] When reconstructing the filter 4, it is necessary to create 16 groups of 32 kernels arranged in the input channel direction for the kernel 20. At this time, for example, the kernels 20 with high importance may be dispersed so as not to be biased towards a specific group of kernels, that is, 16 groups of kernels may be created while distributing them to each group of kernels in order from the kernels with high importance. As another example, the kernels 20 with high importance may be extracted in order to create one group of kernels in order, and 16 groups of kernels may be obtained.
[0080] Returning to FIG. 6, when the reconstruction of the filter 4 of the α-th layer in step S26 is completed, the process proceeds to step S27.
[0081] In step S27, the pruning unit 113 performs a process of subtracting "1" from the variable α. When the subtraction process is performed, the process proceeds to step S28.
[0082] In step S28, the pruning unit 113 determines whether the variable α is zero. If the variable α is zero, it can be determined that the processing related to the reconstruction of the filter 4 has been completed for all convolutional layers. For this reason, when the variable α is 0 (Yes in step S28), the pruning unit 113 ends the pruning process shown in FIG. 6. On the other hand, when the variable α is not 0 (No in step S28), the pruning unit 113 returns the process to step S24. As a result, the processing after step S24 is repeated.
[0083] Note that, by the end of the pruning process (the end of the process shown in FIG. 6), a lightweight model with a reduced model scale compared to the pre-trained model is obtained. After that, re-training of the lightweight model (the process of step S3 in FIG. 5) is performed. When the trained model obtained after re-training has the desired performance, machine learning is completed, and the trained model, which is the target product, is obtained. Since a part of the weight data is deleted in the trained model compared to the pre-trained model, it is a lightweight and high-speed model. And among the kernels 20 that were to be removed by the implementation of the first pruning, the important kernels 20 are left in the filter 4 without being removed, so that the performance degradation of the model can be suppressed.
[0084] Also, as can be seen from the above description, in this example, for each convolutional layer, the extraction of the kernels 20 and the reconstruction of the filter 4 are performed in order from the layer on the output side to the layer on the input side of the neural network. For this reason, the filter 4 can be efficiently reconstructed in an appropriate procedure according to the use of the first pruning method.
[0085] (3-2-2. Second Embodiment) The second embodiment assumes the case of using the second pruning method as the pruning method. FIG. 9 is a flowchart showing the flow of the pruning method according to the second embodiment. Note that the process shown in FIG. 9 is a detailed example of the process of step S2 shown in FIG. 5. Among the processes shown in FIG. 9, the processes of step S21A and step S22A are the same as the content of the second pruning described above with reference to FIG. 4, so the description thereof will be simplified.
[0086] In step S21A, the pruning unit 113 determines the importance for each output channel of the filter 4 included in each convolutional layer (pre-trained model) of the neural network and ranks them. Specifically, the higher the importance of the output channel, the higher the rank. When the ranking of the importance of all output channels in the neural network is completed, the process proceeds to the next step S22A.
[0087] In step S22A, the pruning unit 113 determines that the output channels with the lowest N% importance are to be removed. Note that N% is appropriately determined by experiments or simulations. The data of the output channels determined to be removed is deleted from the weight data of the pre-trained model stored in the memory 12. However, the data of the channels determined to be removed is separately stored in the memory 12 in consideration of later processing.
[0088] Note that the data of the output channels determined to be removed may simply be stored in the memory 12 without being deleted from the weight data of the pre-trained model stored in the memory 12. In this embodiment, for the sake of easy understanding of the processing flow, the data of the output channels determined to be removed is set to be deleted from the weight data of the pre-trained model stored in the memory 12. Also, as described above, with the removal of the output channels, the input channels in the filter 4 of the immediately subsequent convolutional layer have removal targets. The data of the input channels determined to be removed is also stored in the memory 12 in consideration of later processing. When the channels to be removed are determined, the processing proceeds to the next step S23A.
[0089] In step S23A, the pruning unit 113 sets the variable α to 1. The variable α corresponds to the number of the convolutional layer that reconstructs the filter 4. In this example, the numbers of the convolutional layers are sequentially assigned in order from the input side to the output side of the neural network. Also in this example, similar to the first embodiment, the number of convolutional layers included in the pre-trained model is X. The number of the convolutional layer located closest to the input side is "1". The number of the convolutional layer located closest to the output side is "X". By the processing of step S23A, since α = 1, the convolutional layer existing closest to the input side becomes the convolutional layer that first reconstructs the filter 4. When the process of setting the variable α to 1 is performed, the processing proceeds to the next step S24A.
[0090] In step S24A, the pruning unit 113 determines the importance of each kernel 20 included in the output channels that were targeted for removal in the previous second pruning in the α-th layer (convolutional layer), and ranks them. Specifically, the importance of the kernel 20 is obtained by determining an importance index value. The importance index value is, for example, the L1 norm, the L2 norm, or the maximum value of the absolute values. For example, when the size of the kernel 20 is 3×3, the importance index value of the kernel 20 may be the sum of the absolute values of the nine weights. The higher the importance index value, the higher the rank of the importance. When the ranking of the importance of the kernel 20 is completed, the process proceeds to the next step S25A.
[0091] In step S25A, the pruning unit 113 extracts the kernels 20 with the top M% importance among the plurality of kernels 20 whose importance was ranked in step S24. That is, in the pruning method of this example, it includes extracting the kernels 20 with high importance from among the plurality of kernels 20 included in the channels targeted for removal. More specifically, the pruning method of this example includes extracting the kernels 20 with a certain top percentage of importance from among the plurality of kernels 20 included in the channels targeted for removal as the kernels 20 with high importance. Note that the above-mentioned M% is appropriately determined by experiments or simulations. When the extraction of the kernels 20 with high importance is completed, the process proceeds to the next step S26A.
[0092] In step S26A, the pruning unit 113 reconstructs the filter 4 (the filter 4 of the α-th layer) by putting the extracted kernel 20 out of the plurality of kernels 20 included in the channel to be removed into the filter 4. More specifically, in the pruning method of this example, at least a part of the extracted kernels 20 is used to form the filter 4 after reconstruction. The reconstruction of the filter 4 is performed using the output channels that are left without being targeted for removal in the filter 4 and the kernels with high importance included in the output channels targeted for removal. With such a configuration, according to the conventionally known second pruning method, important kernels 20 included in the output channels targeted for removal can be left in the filter 4, and deterioration of the performance of the model after pruning can be suppressed.
[0093] Note that when performing the second pruning executed by steps S21A and S22A, there may be a convolutional layer in which no output channel to be removed is generated. For such a convolutional layer, since there is no output channel to be removed, the processes of steps S24A, S25A, and S26A are substantially skipped.
[0094] Here, a specific example will be given to explain the reconstruction of the filter 4 in step S26A.
[0095] FIG. 10 is a schematic diagram showing the state of the filter 4 of the first layer (the convolutional layer existing on the most input side) at each stage. In FIG. 10, (a) shows the state of the filter 4 before the second pruning is performed in the first layer. (b) shows the state of the filter 4 after the second pruning is performed in the first layer. (b) shows the state immediately after the second pruning is performed, before the kernel 20 with high importance is returned to the filter 4. (c) shows the state of the filter 4 after the reconstruction is performed in the first layer. The state of the filter 4 changes in the order of (a), (b), and (c).
[0096] In the example shown in FIG. 10, the output channels are reduced by the second pruning, and the size (Cout, Cin, K, K) of the filter 4 in the first layer changes from (64, 32, 3, 3) to (20, 32, 3, 3). That is, 44 (= 64 - 20) output channels are targeted for removal by the second pruning. In such a case, an explanation will be given as to how the reconstruction of the filter 4 described above is performed.
[0097] When 44 output channels are targeted for removal, the size (Cout, Cin, K, K) of the data to be removed is (44, 32, 3, 3). That is, the number of kernels 20 to be removed is 1408 (= 44 × 32). Among the kernels 20 to be removed, the top M% of the kernels 20 with high importance extracted in the process of step S25A are candidates to be returned to the filter 4 including the channels that remain without being removed.
[0098] Here, assuming that M% is 10%, 141 (≒ 1408 × 0.1) kernels 20 are candidates to be returned to the filter 4 including the channels that remain without being removed. When returning the kernels 20 to the filter 4, it is necessary to return a group of kernels composed of the same number as the number of kernels (the same as the number of input channels) of each output channel that remains without being removed. In the example shown in FIG. 10, the number of kernels (number of input channels) of each output channel of the filter 4 including the channels that remain without being removed is 32. For this reason, 4 (≒ 141 ÷ 32) groups of kernels are returned to the filter 4. As can be seen from the above, the reconstruction of the filter 4 includes a process of arranging the extracted kernels 20 in the direction of the input channels to construct the output channels. In this example, the size (Cout, Cin, K, K) of the filter 4 after reconstruction is (24, 32, 3, 3). Here, the number of output channels "24" is the sum of the number of output channels "20" after the second pruning is performed and the number of filters "4" to be returned to the filter 4. Note that in this example, some of the kernels 20 extracted as candidates to be returned to the filter 4 will be removed without being returned to the filter 4.
[0099] Also, when constructing a kernel group for reconstructing the filter 4, for example, the kernel group may be formed so that the highly important kernel 20 does not bias towards a specific kernel group. For example, a predetermined number of kernel groups may be formed while distributing the kernel groups to be formed in order from the highly important kernels. However, it is not limited to this, and the kernel group may be formed so that the highly important kernel 20 biases towards a specific kernel group.
[0100] Also, when returning the kernel group to the filter 4, for example, without changing the order of the kernel groups left without being removed by the second pruning, the kernel groups to be returned may be arranged in order on the head side or the tail side of the order. However, it may also be configured to change the order of the kernel groups left without being removed by the second pruning and return the newly formed kernel group to the filter 4.
[0101] Also, in a convolutional neural network, the number of output channels of the filter 4 in a certain convolutional layer must be the same as the number of input channels of the filter 4 in the convolutional layer immediately after that convolutional layer. For this reason, when the number of output channels is increased by reconstructing the filter 4, reconstruction of the filter 4 in the immediately following convolutional layer is also required. A specific example will be given to explain this.
[0102] FIG. 11 is a diagram showing the relationship between the filters 4 of the first layer and the second layer. The numbers shown in parentheses in FIG. 11 indicate the size (Cout, Cin, K, K) of the filter 4. The first layer is the same as the first layer in FIG. 10 and is the convolutional layer existing on the most input side of the convolutional neural network. The second layer is the convolutional layer immediately after the first layer. Also, in FIG. 10, (a) before the second pruning is performed, (b) after the second pruning is performed, and (c) after the reconstruction is performed mean the same state as in the case of FIG. 10 described above.
[0103] As shown in FIG. 11, before the second pruning, after the second pruning, and after the reconstruction, the number of output channels of the filter 4 in the first layer is the same as the number of input channels of the filter 4 in the immediately subsequent second layer. In the second layer, the size (Cout, Cin, K, K) of the filter 4 after the second pruning is (192, 20, 3, 3). That is, the number of input channels of the filter 4 in the second layer after the second pruning is "20", which does not match the number of output channels "24" of the filter 4 in the first layer after the reconstruction. For this reason, the filter 4 is reconstructed to change the number of input channels in the second layer from "20" to "24".
[0104] In the second layer, in order to increase the number of input channels from "20" to "24", since the number of output channels is "192", 768 (= 192×(24 - 20)) kernels are required. These 20 kernels are replenished using the 20 kernels removed during the second pruning. In the second layer, 8536 (= 192×(64 - 20)) kernels 20 are targeted for removal by the second pruning. From these 20 kernels targeted for removal, 768 kernels 20 are secured in order from the kernels 20 with high importance and reconstructed into a filter 4 with a size (Cout, Cin, K, K) of (192, 24, 3, 3).
[0105] Note that for the reconstruction of the filter 4 that increases the number of input channels, a method similar to the method for reconstructing the filter 4 that increases the number of input channels in the first embodiment described above may be used.
[0106] Returning to FIG. 9, when the reconstruction of the filter 4 in the α-th layer in step S26A is completed, the process proceeds to step S27A.
[0107] In step S27A, the pruning unit 113 performs a process of adding "1" to the variable α. When this addition process is performed, the process proceeds to step S28A.
[0108] In step S28A, the pruning unit 113 determines whether the variable α is X. If the variable α is X, it can be determined that the processing related to the reconstruction of filter 4 for all convolutional layers has been completed. For this reason, if the variable α is X (Yes in step S28A), the pruning unit 113 ends the pruning process shown in FIG. 9. On the other hand, if the variable α is not X (No in step S28A), the pruning unit 113 returns the process to step S24A. As a result, the processes after step S24A are repeated.
[0109] Note that by ending the pruning process (ending the process shown in FIG. 10), a lightweight model with a reduced model scale compared to the pre-trained model is obtained, and thereafter, re-learning of the lightweight model (the process of step S3 in FIG. 5) is performed. When the learned model obtained after re-learning has the desired performance, the machine learning is completed, and the learned model, which is the target product, is obtained. Since a part of the weight data of the learned model is deleted compared to the pre-trained model, it is a lightweight and high-speed model. And among the kernels 20 that were to be removed by the implementation of the second pruning, the important kernels 20 are returned to filter 4 without being removed, so that deterioration of the model performance can be suppressed.
[0110] Also, as can be seen from the above description, in this example, the extraction of kernel 20 and the reconstruction of filter 4 for each convolutional layer are performed in order from the layer on the input side to the layer on the output side of the neural network. For this reason, the reconstruction of filter 4 can be efficiently performed in an appropriate procedure according to the use of the second pruning method.
[0111] <4. Precautions, etc.> Various technical features disclosed in the embodiments for carrying out the invention in this specification can be variously modified without departing from the gist of the technical creation. Also, a plurality of embodiments and modifications disclosed in the embodiments for carrying out the invention in this specification may be implemented in combination within the possible range.
Explanation of Reference Numerals
[0112] 1 ··· Machine learning device (pruning device) 3 ··· Trained model 4 ··· Filter 20 ··· Kernel
Claims
1. A method for pruning a neural network including a convolutional layer, comprising: determining a removal target for each channel of the filters of the convolutional layer; extracting kernels with high importance from among a plurality of kernels included in the channels determined to be the removal targets; reconstructing the filters using the extracted kernels and the kernels included in the channels of the filters that are left without being the removal targets; a pruning method.
2. The convolutional layer is included in the neural network in plurality, For each convolutional layer, the extraction of the kernels and the reconstruction of the filters are performed, the pruning method according to claim 1.
3. The determination of the removal target is performed using a first pruning method that removes in units of input channels, For each convolutional layer, the extraction of the kernels and the reconstruction of the filters are performed in order from the layer on the output side to the layer on the input side of the neural network, the pruning method according to claim 2.
4. The reconstruction of the filters includes a process of distributing the extracted kernels so that the number is the same for each output channel, the pruning method according to claim 3.
5. The determination of the removal target is performed using a second pruning method that removes in units of output channels, For each convolutional layer, the extraction of the kernels and the reconstruction of the filters are performed in order from the layer on the input side to the layer on the output side of the neural network, the pruning method according to claim 2.
6. The reconstruction of the filters includes a process of constructing the output channels by arranging the extracted kernels in the direction of the input channels, the pruning method according to claim 5.
7. Kernels with a certain upper ratio of importance among the plurality of kernels included in the channels determined to be the removal targets are extracted as the kernels with high importance, The pruned filter after reconstruction is formed using at least a part of the extracted kernels, the pruning method according to any one of claims 1 to 6.
8. An apparatus for pruning a neural network including a convolutional layer, comprising: determining a removal target for each channel of the filters of the convolutional layer; Extract a kernel with a high degree of importance from among a plurality of kernels included in the channel to be removed, Reconstruct the filter using the extracted kernel and the kernels included in the channels that are left without being targeted for removal among the filters, Pruning device.
9. A machine learning method for a neural network including a plurality of convolutional layers, Perform learning of the neural network that has not been learned, Determine a removal target for each channel with respect to the filters in the neural network after the learning, For each convolutional layer, reconstruct the filter using the kernels included in the channels that are left without being targeted for removal among the filters and the kernels with a high degree of importance included in the channels targeted for removal, Perform re-learning of the neural network obtained after the reconstruction of the filter to generate a learned model, Machine learning method.
10. Determine a removal target for each channel with respect to the filters in the convolutional neural network after learning, Reconstruct the filter using the kernels included in the channels that are left without being targeted for removal among the filters and the kernels with a high degree of importance included in the channels targeted for removal, A learned model configured by re-learning the neural network obtained after the reconstruction of the filter. Learned model.
Citation Information
Patent Citations
Method and apparatus for pruning neural network
JP2021047854A