Convolution kernel integer pruning method and system of convolutional neural network structure
Through adaptive selection mask strategy and sparse training method, convolution kernel integer pruning of convolutional neural network structure is realized, solving the problem of difficult to take into account both hardware friendliness and model accuracy in the existing technology, and significantly improving the computing efficiency of the model.
Patent Information
- Application Number
- CN202510036233.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing convolutional network pruning methods are difficult to achieve hardware-friendly and maintain model accuracy at the same time, especially in structured pruning, where the accuracy loss is high.
A convolution kernel integer pruning method for convolutional neural network structure is proposed. Through adaptive selection of mask strategies, a candidate mask set is defined, and a weighted sum calculation is performed using the product representation of normalized learnable parameters and temperature. Finally, the final mask of each convolution kernel is obtained through sparse training, and the pruning is completed.
It realizes pruning on the basis of maintaining the model accuracy, significantly reducing the computational amount of convolutional neural network, improving the computing efficiency of the model, and supporting the convolutional network structure of any tasks such as image classification, detection, and segmentation.
Smart Images

Figure CN119940420A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a convolution kernel integer pruning method and system for a convolutional neural network structure. Background Art
[0002] With the development of deep learning technology, CNN (Convolutional Neural Network) has been widely used in image recognition, detection and other fields. However, in order to improve the effect of the model, the number of CNN model parameters has become huge, which leads to high consumption of computing resources.
[0003] In order to improve the reasoning efficiency of the model and reduce the consumption of computing resources, researchers have proposed a variety of model compression and acceleration methods. Among them, pruning is an effective means, which can prune redundant parameters in the network through a certain measurement method, so that the network can maintain a high relative accuracy while reducing the number of network parameters, and ultimately achieve model compression and computational acceleration.
[0004] Existing convolutional network pruning methods can be divided into two categories: unstructured pruning and structured pruning. Unstructured pruning refers to suppressing redundant parameters in the weight tensor to 0. These redundant parameters do not participate in the calculation during inference and are stored in a special format during deployment, which can achieve inference acceleration and volume compression.
[0005] The most advanced unstructured pruning of convolutional networks can achieve 90% pruning with almost no loss of accuracy. However, this method has an obvious flaw: it requires specific hardware design and storage design of 0 values, otherwise it cannot achieve inference acceleration and volume compression. These two designs are difficult to implement in practice, so many scholars turn to structured pruning methods.
[0006] Convolutional network structured pruning refers to removing redundant structures in the network structure, such as removing certain channels of the convolution kernel. This method can directly compress the model size and improve the inference speed without the need for special hardware design and parameter storage design. Therefore, it is a hardware-friendly pruning method. However, since its operation is coarse-grained, the accuracy loss of the model will be relatively high. At present, the most advanced convolutional network structured pruning has an accuracy loss of more than 1% when pruning 33%. Summary of the invention
[0007] The purpose of the present invention is to overcome the problem that it is difficult to simultaneously achieve hardware friendliness and maintain model accuracy in convolutional network pruning, and propose a convolution kernel integer pruning method and system for a convolutional neural network structure.
[0008] In order to achieve the above object, the present invention adopts the following technical scheme: A convolution kernel integer pruning method for a convolutional neural network structure, comprising the following steps: For each convolution kernel to be shaped and pruned in the convolutional network, a candidate mask set is defined according to the pruning rate by an adaptive mask selection strategy to obtain a selected candidate mask set; Assigning a normalized weight to each mask in the selected candidate mask set to obtain a normalized mask weight, where the normalized mask weight is represented by the product of the normalized learnable parameter and the temperature; Use the normalized mask weights to weight the candidate mask set and calculate the mask of the convolution kernel to be integer pruned, and obtain the weight of the convolution kernel after masking; The convolution network is sparsely trained according to the masked convolution kernel weights until convergence to obtain the final mask of each convolution kernel. The final mask of each convolution kernel is the mask with the largest weight of each convolution kernel. The final mask of each convolution kernel is used to replace the weight of the convolution kernel to be integer pruned to complete the pruning.
[0009] Further, the candidate mask set includes a plurality of masks with regular structures or masks with semi-regular structures; A well-structured mask is one in which each element in the mask is surrounded by at least two elements with the same value, and the same values are connected horizontally and vertically; A semi-regular mask is one in which only one element value around each element in the mask is discontinuous, and the rest of the elements are continuous. When the width of the convolution kernel is 3 and the height is 3, each mask in the candidate mask set is a matrix of shape 3×3, and the matrix elements are 0 or 1.
[0010] Furthermore, the normalized mask weight is:
[0011]
[0012] in, is the normalized weight, softmax() is the normalization function, x is the current mask, e is the base of the natural logarithm, is the exponential function value of x, is a learnable parameter, i is an intermediate variable, For temperature.
[0013] Furthermore, the masked convolution kernel weight is:
[0014]
[0015] in, is the convolution kernel weight after masking, weight is the original weight, For the mask, is the normalized weight, softmax() is the normalization function, is a learnable parameter, i is an intermediate variable, is the temperature, is the mask matrix.
[0016] Furthermore, during the training process of the sparse training, the learnable parameters of different masks are learned separately, and the temperature gradually increases. The gap between the learnable parameters is expanded by the temperature value after multiplying with the temperature value. After passing the softmax function, a mask close to 1 other than 0 is selected as the final mask of the convolution kernel.
[0017] Furthermore, the adaptive mask selection strategy is: when defining the integer pruning convolution kernel, after analyzing the sensitivity of each convolution layer to the final accuracy, a different mask type is configured for the convolution kernel of each convolution layer; The goal of the adaptive mask selection strategy is to assign masks with low pruning rates to convolutional layers with high sensitivity, and assign masks with high pruning rates to those with low sensitivity; The sensitivity evaluation index is the accuracy loss of the model after the convolution kernel of a single convolutional layer is pruned.
[0018] Furthermore, the construction of the adaptive selection mask strategy utilizes a dynamic programming algorithm, specifically: Set the target pruning rate and convert the target pruning rate into the FLOPs upper limit of the pruning model; Calculate the model accuracy of each convolutional layer at different pruning rates, output the model accuracy as a sensitivity table, and construct the FLOPs of each convolutional layer at different pruning rates and output the table; The knapsack problem is constructed using the FLOPs of convolutional layers with different pruning rates. The optimal mask combination is obtained by solving the knapsack problem. The closer the sum of the FLOPs of convolutional layers with different pruning rates is to the FLOPs upper limit of the pruning model, the smaller the overall sensitivity. The optimal mask combination is used to configure the corresponding mask for the convolution kernel of each convolution layer.
[0019] In a second aspect, the present invention provides a convolution kernel shaping pruning system for a convolutional neural network structure, comprising: A candidate mask set module is defined, which is used to define a candidate mask set according to the pruning rate for each convolution kernel to be pruned in the convolutional network through an adaptive mask selection strategy to obtain a selected candidate mask set; A normalized mask weight module is used to assign a normalized weight to each mask in the selected candidate mask set to obtain a normalized mask weight, where the normalized mask weight is represented by the product of the normalized learnable parameter and the temperature; The mask calculation convolution kernel weight module is used to use the normalized mask weight to weight the candidate mask set to calculate the mask of the convolution kernel to be integer pruned, and obtain the convolution kernel weight after masking; The training converges to obtain the final mask module, which is used to perform sparse training on the convolution network according to the masked convolution kernel weights until convergence, and obtain the final mask of each convolution kernel. The final mask of each convolution kernel is the mask with the largest weight of each convolution kernel. The pruning module is completed, which is used to replace the weight of the convolution kernel to be integer pruned with the final mask of each convolution kernel to complete the pruning.
[0020] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the convolution kernel integer pruning method for a convolutional neural network structure when executing the computer program.
[0021] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, provides a convolution kernel integer pruning method for a convolutional neural network structure.
[0022] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention proposes a convolution kernel integer pruning method for a convolutional neural network structure, which is a fine-grained semi-structured pruning method between structured pruning and unstructured pruning. According to the characteristics of the convolution kernel of the convolution network, a mask construction method mainly for 3×3 convolution kernels is designed. Through training, each layer of convolution kernels autonomously learns a mask, and the convolution kernels are pruned according to the learned mask. Compared with the existing unstructured pruning method, the present invention is hardware compatible, and only specific convolution calculations need to be configured in compilation. The present invention is a more fine-grained pruning method, with less precision loss under the same sparsity rate. The present invention supports convolutional network structures for any tasks such as image classification, detection, and segmentation. The present invention significantly reduces the computational complexity of the convolutional neural network and improves the computational efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are for explanation purposes only and are not intended to limit the scope of the present invention in any way. In addition, the shapes and proportional dimensions of the components in the drawings are only for illustration purposes to help understand the present invention, and are not intended to specifically limit the shapes and proportional dimensions of the components of the present invention. In the drawings: Figure 1 The present invention provides a flow chart of a convolution kernel integer pruning method for a convolutional neural network structure.
[0024] Figure 2 This is a structural diagram of a convolution kernel integer pruning system for a convolutional neural network structure of the present invention.
[0025] Figure 3 This is a diagram of an electronic device for a convolution kernel integer pruning method for a convolutional neural network structure of the present invention.
[0026] Figure 4 This is a flow chart of the use of a convolution kernel integer pruning method for a convolutional neural network structure according to an embodiment of the present invention.
[0027] Figure 5 A mask style diagram with a regular structure defined in a convolution kernel integer pruning method for a convolutional neural network structure of the present invention.
[0028] Figure 6 This is a mask pattern with a shape of 2×3 in a convolution kernel integer pruning method of a convolutional neural network structure of the present invention.
[0029] Figure 7 A convolution kernel calculation optimization diagram of a convolution kernel integer pruning method for a convolutional neural network structure of the present invention.
[0030] Figure 8 This is a comparison chart of the inference efficiency of a convolution kernel integer pruning method for a convolutional neural network structure of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0032] Embodiment 1 See also Figure 1 , a convolution kernel integer pruning method for a convolutional neural network structure, comprising the following steps: For each convolution kernel to be shaped and pruned in the convolutional network, a candidate mask set is defined according to the pruning rate by an adaptive mask selection strategy to obtain a selected candidate mask set; Assigning a normalized weight to each mask in the selected candidate mask set to obtain a normalized mask weight, where the normalized mask weight is represented by the product of the normalized learnable parameter and the temperature; Use the normalized mask weights to weight the candidate mask set and calculate the mask of the convolution kernel to be integer pruned, and obtain the weight of the convolution kernel after masking; The convolution network is sparsely trained according to the masked convolution kernel weights until convergence to obtain the final mask of each convolution kernel. The final mask of each convolution kernel is the mask with the largest weight of each convolution kernel. The final mask of each convolution kernel is used to replace the weight of the convolution kernel to be integer pruned to complete the pruning.
[0033] This embodiment can significantly reduce the number of non-zero elements in the convolution kernel through integer pruning, thereby reducing the storage requirements of the model. This is particularly important for deploying deep learning models on resource-constrained devices, such as mobile devices, embedded systems, etc. When the pruned model performs reasoning, the calculation speed can be accelerated and the delay can be reduced due to the reduction of the calculation of non-zero elements. This has significant advantages for real-time applications such as autonomous driving and video surveillance. By adaptively selecting a mask strategy and representing learnable parameters with temperature, pruning can be performed on the basis of maintaining model accuracy. This means that while reducing the model size and the amount of calculation, too much prediction performance will not be sacrificed. During the pruning process, the weights of the model can be further optimized by sparsely training the convolutional network, making it more compact and efficient. This helps to improve the generalization ability and robustness of the model. The adaptive selection of the mask strategy makes the pruning process flexible and can be adjusted according to different pruning rates. The integer pruning method is relatively simple and easy to implement and deploy in the existing deep learning framework. This makes the method highly feasible and practical in practical applications.
[0034] The candidate mask set includes a number of masks with regular structures or masks with semi-regular structures; A well-structured mask is one in which each element in the mask is surrounded by at least two elements with the same value, and the same values are connected horizontally and vertically; A semi-regular mask is one in which only one element value around each element in the mask is discontinuous, and the rest of the elements are continuous. When the width of the convolution kernel is 3 and the height is 3, each mask in the candidate mask set is a 3×3 matrix with matrix elements of 0 or 1. This method is not only applicable to convolution kernels with a width of 3 and a height of 3, but can also be extended to convolution kernels of other sizes according to actual needs.
[0035] Using a candidate mask set containing well-structured masks and semi-regular masks for integer pruning can significantly improve the pruning efficiency of convolutional neural networks, reduce computational complexity, reduce memory usage, and maintain the performance and accuracy of the model to a certain extent. Well-structured masks ensure that each non-zero element is connected to at least two identical elements around it, which helps to maintain the internal structural characteristics of the convolution kernel. This feature preservation is crucial to the performance of the model because the structure of the convolution kernel is closely related to the features it extracts. Semi-regular masks allow only one non-contiguous element around each element. This slight "irregularity" can reduce the damage of pruning to the model structure to a certain extent. Compared with completely random pruning, this method is more likely to retain weights that are important to model performance. Since the masks in the candidate mask set have specific structural characteristics, this makes the pruning process more efficient. The algorithm can find masks that meet the pruning requirements faster and reduce unnecessary searches and calculations. Well-structured masks and semi-regular masks usually have fewer non-zero elements, which reduces the computational complexity of convolution operations. During the inference phase, this helps reduce the amount of computation and increase the running speed of the model. The masks in the candidate mask set have clear structural features, which makes them easier to implement and optimize in the algorithm. For example, these features can be used to speed up the search and selection process during pruning. By selecting an appropriate set of candidate masks, efficient pruning can be achieved while maintaining model accuracy. This is because masks with regular structures and semi-regular masks are more likely to retain weights that are important to model performance, thereby reducing the negative impact on model accuracy. Since the pruned convolution kernels have fewer non-zero elements, they also take up less space in memory. This is especially important for deploying deep learning models on resource-constrained devices.
[0036] The normalized mask weight is:
[0037]
[0038] in, is the normalized weight, softmax() is the normalization function, x is the current mask, e is the base of the natural logarithm, is the exponential function value of x, is a learnable parameter, i is an intermediate variable, For temperature.
[0039] The weight of the convolution kernel after masking is:
[0040]
[0041] in, is the convolution kernel weight after masking, weight is the original weight, For the mask, is the normalized weight, softmax() is the normalization function, is a learnable parameter, i is an intermediate variable, is the temperature, is the mask matrix.
[0042] During the sparse training process, the learnable parameters of different masks are learned separately, and the temperature gradually increases. The gap between the learnable parameters is multiplied by the temperature value and expanded by the temperature value. After the softmax function, the mask close to 1 other than 0 is selected as the final mask of the convolution kernel.
[0043] The adaptive mask selection strategy is: when defining the integer pruning convolution kernel, after analyzing the sensitivity of each convolution layer to the final accuracy, different mask types are configured for the convolution kernel of each convolution layer; The goal of the adaptive mask selection strategy is to assign masks with low pruning rates to convolutional layers with high sensitivity, and assign masks with high pruning rates to those with low sensitivity; The sensitivity evaluation indicator is the accuracy loss of the model after pruning the convolution kernel of a single convolutional layer.
[0044] The construction of the adaptive selection mask strategy uses a dynamic programming algorithm, specifically: Set the target pruning rate and convert the target pruning rate into the FLOPs upper limit of the pruning model; The target pruning rate in this step refers to the proportion of model parameters or computational complexity that you want to reduce. Based on the FLOPs of the original model and the target pruning rate, the upper limit of the FLOPs after pruning is calculated.
[0045] Calculate the model accuracy of each convolutional layer at different pruning rates, output the model accuracy as a sensitivity table, and construct the FLOPs of each convolutional layer at different pruning rates and output the table; This step calculates the model accuracy and FLOPs of each convolutional layer at different pruning rates, applies different pruning rates to each convolutional layer, and evaluates the model accuracy at each pruning rate. Calculate the FLOPs at each pruning rate.
[0046] The knapsack problem is constructed using the FLOPs of convolutional layers with different pruning rates. The optimal mask combination is obtained by solving the knapsack problem. The closer the sum of the FLOPs of convolutional layers with different pruning rates is to the FLOPs upper limit of the pruning model, the smaller the overall sensitivity. In this step, the FLOPs of each convolutional layer at different pruning rates are regarded as items in the knapsack problem, with a weight of FLOPs and a value of the inverse of the sensitivity (because the goal is to minimize the sensitivity). Solve the knapsack problem and find a combination that makes the total FLOPs closest to the FLOPs upper limit while minimizing the overall sensitivity.
[0047] The optimal mask combination is used to configure the corresponding mask for the convolution kernel of each convolution layer.
[0048] This step applies the corresponding pruning rate to each convolutional layer according to the optimal mask combination.
[0049] Through the above steps, CNN pruning based on the target pruning rate can be achieved.
[0050] Embodiment 2 See also Figure 2 , a convolution kernel integer pruning system for a convolutional neural network structure, comprising: A candidate mask set module is defined, which is used to define a candidate mask set according to the pruning rate for each convolution kernel to be pruned in the convolutional network through an adaptive mask selection strategy to obtain a selected candidate mask set; A normalized mask weight module is used to assign a normalized weight to each mask in the selected candidate mask set to obtain a normalized mask weight, where the normalized mask weight is represented by the product of the normalized learnable parameter and the temperature; The mask calculation convolution kernel weight module is used to use the normalized mask weight to weight the candidate mask set to calculate the mask of the convolution kernel to be integer pruned, and obtain the convolution kernel weight after masking; The training converges to obtain the final mask module, which is used to perform sparse training on the convolution network according to the masked convolution kernel weights until convergence, and obtain the final mask of each convolution kernel. The final mask of each convolution kernel is the mask with the largest weight of each convolution kernel. The pruning module is completed, which is used to replace the weight of the convolution kernel to be integer pruned with the final mask of each convolution kernel to complete the pruning.
[0051] Embodiment 3 See also Figure 3 , an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the convolution kernel integer pruning method of a convolutional neural network structure is implemented.
[0052] Embodiment 4 A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a convolution kernel integer pruning method for a convolutional neural network structure is described.
[0053] Embodiment 5 See also Figure 4 , using a convolution kernel integer pruning method with a convolutional neural network structure, first the user selects a convolutional neural network based on a specific application scenario, and then deploys reasoning. If the reasoning speed does not meet the requirements of the actual scenario, select the method of this embodiment for pruning. The user configures the pruning strategy and pruning rate according to actual needs, then defines the integer pruned convolution kernel and replaces the original convolution kernel. Next, the redefined convolution network is sparsely trained, and finally the training converged model is pruned. Because the pruning method is structured, the pruned model can be directly deployed in the original environment and make reasoning decisions.
[0054] The detailed steps of the method are as follows: S1. Define integer pruned convolution kernel S101. Define a candidate mask set for each convolution kernel of size 3×3 (height×width represents the shape) in the convolutional network. Each mask in the set is a matrix of shape 3×3, and the matrix elements are 0 or 1. This mask set contains several masks with regular or semi-regular structures. Regular structure means that each element in the mask is surrounded by at least two elements with the same value as itself, that is, the same values are connected horizontally and vertically. Semi-regular means that only one element is discontinuous, and the rest are continuous. For example Figure 5 The difference between the three masks is shown. The reason for constructing a mask with a regular structure is that it can improve the efficiency of the convolution calculation of the convolution kernel.
[0055] Specifically, the masks with regular structures are as follows: 2 2×3, 2 3×2, 4 2×2, 3 1×3, 3 3×1, 6 1×2, 6 2×1, and 9 1×1 masks.
[0056] For example, a 2×3 mask such as Figure 6 , where 2×3 refers to the distribution of 1 in the mask, Figure 6 The distribution of 1 in the mask is 2 rows and 3 columns, so it is abbreviated as 2×3. In order to increase the diversity of the mask, the 2×2 mask and the 1×1 mask are combined to obtain 20 semi-regular 2×2+1×1 masks, such as Figure 6 Middle image; By combining the 1×3 and 1×1 masks, we can get 36 1×3+1×1 masks.
[0057] When defining a candidate mask set, it is not necessary to select all of the above masks as candidates. Instead, the user can customize a certain type of mask as a candidate, such as selecting 14 masks with the shape of {2×3, 2×2, 1×3}.
[0058] S102: assign a learnable parameter to each mask of the selected candidate mask set . In addition, define a temperature , which is multiplied by the parameter p to represent the weight of the mask. At this time, the weight is not normalized, so the softmax function is used for normalization to obtain the normalized weight ,
[0059] in:
[0060] If there is a mask Then the corresponding parameters Weight .
[0061] S103, define a new convolution kernel weight. Let the original weight be weight, then the new weight after masking is:
[0062]
[0063] S2. Train the newly defined convolutional network.
[0064] The model training makes each convolution kernel learn a mask, which is also called sparse training. The training configuration of the network remains the same as the original training method. During the training process, the parameters p of different masks are learned separately, and the temperature Gradually increases, resulting in learnable parameters The difference between them is multiplied by T and expanded by T times. Then after passing through the softmax function, a will approach 1, and the rest will approach 0. A mask is selected as the final mask of the convolution kernel, that is:
[0065] S3. Use mask to prune the convolution kernel.
[0066] Train until the model converges and each convolution kernel learns a mask , the weights of the convolution kernel are replaced:
[0067] The final convolution kernel weight is as follows: Figure 7 shown.
[0068] The meanings of symbols are shown in Table 1: Table 1 Symbols
[0069] The implementation of inference acceleration in this embodiment is mainly to reduce the number of floating-point operations (FLOPs) of convolution calculations in each convolution layer, and ultimately reduce the inference time of the entire convolution network. The following takes the convolution kernel with a 2×2 mask as an example to illustrate the principle of reducing convolution calculation FLOPs: like Figure 7 In the first row, the left side is the convolution kernel with a 2×2 mask, and the right side is a feature map of size 4×4. When performing convolution normally, assuming that the movement stride is 1, the window is slid from left to right and from top to bottom.
[0070] It is found that since the last column and the last row of the convolution kernel are 0, the values at these positions multiplied by the feature map are all 0, which has no effect on the final result. Therefore, the calculated result can be equivalent to Figure 7 The second row shows the result of sliding the convolution kernel of size 2×2 on the feature map of size 3×3.
[0071] This reduces the FLOPs (≈0.56).
[0072] It is exactly the ratio of 0 elements on the mask, that is, the FLOPs pruning rate (pr for short) of the convolution kernel is the ratio of 0 elements on the mask. The pruning rates corresponding to different masks are shown in Table 2: Table 2 Pruning rates corresponding to different masks
[0073] The FLOPs of the convolutional network are mainly concentrated in the convolution calculation, so the pruning rate of the entire convolutional network FLOPs is approximately equal to the average pruning rate of each convolutional layer:
[0074] Figure 8 The following table 3 is the comparison result of the inference speed after the mask of a single convolution kernel measured on the zebu simulation platform; Table 3 Comparison of the inference speed of a single convolution kernel after masking measured on the zebu simulation platform
[0075] 3×3 represents the original convolution kernel. It is found that the convolution kernels with different shape masks have a significant speed improvement compared with the original convolution kernel, which shows that the method of this embodiment has a significant performance improvement.
[0076] There are two defects in the method of this embodiment when implementing step S101 to define candidate masks: (1) The set of candidate masks defined by the convolution kernels of each convolution layer is the same, so the pruning rate corresponding to the learned mask will be relatively fixed. For example, if the candidate mask set is {2×3}, then the final model pruning rate is 33.33%. Or if the candidate mask set is {2×3, 2×2, 1×3}, it is unpredictable which mask each convolution kernel will learn in the end. Therefore, it is difficult to specify pruning with an arbitrary target pruning rate.
[0077] (2) It does not take into account that the impact of pruning each convolutional layer on model accuracy is different. That is, the sensitivity of each convolutional layer is different. If some layers are pruned too much, it will cause a large loss in model accuracy.
[0078] Therefore, the method of this embodiment proposes a scheme of adaptively selecting masks. Adaptive selection of masks means that when defining integer pruned convolution kernels, different mask types are configured for the convolution kernels of each convolution layer after analyzing the sensitivity of each convolution layer to the final accuracy. The evaluation index of sensitivity is the accuracy loss of the model after the convolution kernel of a single convolution layer is pruned. The greater the loss, the higher the sensitivity of the convolution layer. The goal of adaptive selection of masks is to assign masks with low pruning rates to convolution layers with high sensitivity, or even not prune them, and to assign masks with higher pruning rates to those with low sensitivity, so as to minimize the accuracy loss of the model while achieving the target pruning rate.
[0079] The strategy for adaptively selecting masks is to construct a knapsack problem and use a dynamic programming algorithm to solve the mask combination assigned to each convolutional layer of the model. The specific construction method is as follows: (1) Set the target pruning rate (e.g., 0.28) and convert it into the FLOPs upper limit of the pruning model; (2) Calculate the model accuracy at different pruning rates for each convolutional layer as sensitivity and output the table. At the same time, construct the FLOPs at different pruning rates for each convolutional layer and output the table.
[0080] (3) Construct the optimal mask combination for solving the knapsack problem: when the sum of the FLOPs of the convolutional layers with different pruning rates is as close as possible to the FLOPs upper limit of the pruning model, the overall sensitivity is minimized.
[0081] (4) Obtain the optimal combination and configure the corresponding mask for the convolution kernel of each convolutional layer.
[0082] The strategy for providing candidate masks based on the pruning rate is shown in Table 4: Table 4 Comparison of pruning rate and candidate mask strategies
[0083] Advantages of adaptive selection mask: (1) A mask combination that meets any target pruning rate can be selected; (2) At the same pruning rate, the adaptive mask selection pruning scheme is better than the global mask selection scheme. Table 5 below shows some comparison results. Table 5 Comparison results of the pruning scheme of adaptive selection mask and the scheme of global selection mask under the same pruning rate
[0084] The following is a further description of the method of this embodiment in combination with experiments and experimental results: During the experiment, the model hyperparameter temperature T is initialized to 1, and the value range is [1, 256]. The adjustment of temperature T is linear and obeys the function:
[0085] The learnable parameters The initial value follows a uniform distribution: , Indicates the number of masks, which are updated together with the rest of the model parameters during sparse training.
[0086] The test results of the method in this embodiment on different models with different pruning rates are shown in Table 6; Table 6 Test results of the method in this embodiment on different models with different pruning rates
[0087] It can be seen that the method of this embodiment has almost no accuracy loss when pruning 33% on most models of image classification and segmentation, and only about 1% accuracy loss when pruning 50%.
[0088] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, read-only optical disk, optical storage, etc.) containing computer-usable program code.
[0089] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0090] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention is described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation of the present invention can still be modified or replaced by equivalents, and any modification or equivalent replacement that does not deviate from the spirit and scope of the present invention should be included in the protection scope of the present invention.
Claims
1. A convolution kernel integer pruning method for a convolutional neural network structure, characterized in that: The following steps are involved: For each convolution kernel to be shaped and pruned in the convolutional network, a candidate mask set is defined according to the pruning rate by an adaptive mask selection strategy to obtain a selected candidate mask set; Assigning a normalized weight to each mask in the selected candidate mask set to obtain a normalized mask weight, where the normalized mask weight is represented by the product of the normalized learnable parameter and the temperature; Use the normalized mask weights to weight the candidate mask set and calculate the mask of the convolution kernel to be integer pruned, and obtain the weight of the convolution kernel after masking; The convolution network is sparsely trained according to the masked convolution kernel weights until convergence to obtain the final mask of each convolution kernel. The final mask of each convolution kernel is the mask with the largest weight of each convolution kernel. The final mask of each convolution kernel is used to replace the weight of the convolution kernel to be integer pruned to complete the pruning.
2. The convolution kernel shaping and pruning method of a convolutional neural network structure according to claim 1, characterized in that: The candidate mask set includes a plurality of masks with regular structures or masks with semi-regular structures; A well-structured mask is one in which each element in the mask is surrounded by at least two elements with the same value, and the same values are connected horizontally and vertically; A semi-regular mask is one in which only one element value around each element in the mask is discontinuous, and the rest of the elements are continuous. When the width of the convolution kernel is 3 and the height is 3, each mask in the candidate mask set is a matrix of shape 3×3, and the matrix elements are 0 or 1.
3. The convolution kernel shaping and pruning method of a convolutional neural network structure according to claim 1, characterized in that: The normalized mask weight is: in, is the normalized weight, softmax() is the normalization function, x is the current mask, e is the base of the natural logarithm, is the exponential function value of x, is a learnable parameter, i is an intermediate variable, For temperature.
4. The convolution kernel shaping and pruning method of a convolutional neural network structure according to claim 3, characterized in that: The masked convolution kernel weight is: in, is the convolution kernel weight after masking, weight is the original weight, For the mask, is the normalized weight, softmax() is the normalization function, is a learnable parameter, i is an intermediate variable, is the temperature, is the mask matrix.
5. The convolution kernel shaping and pruning method of a convolutional neural network structure according to claim 1, characterized in that: During the training process of the sparse training, the learnable parameters of different masks are learned separately, and the temperature gradually increases. The gap between the learnable parameters is multiplied by the temperature value and expanded by the temperature value. After the softmax function, the mask close to 1 other than 0 is selected as the final mask of the convolution kernel.
6. The convolution kernel shaping and pruning method of a convolutional neural network structure according to claim 1, characterized in that: The adaptive mask selection strategy is: when defining the integer pruning convolution kernel, after analyzing the sensitivity of each convolution layer to the final accuracy, a different mask type is configured for the convolution kernel of each convolution layer; The goal of the adaptive mask selection strategy is to assign masks with low pruning rates to convolutional layers with high sensitivity, and assign masks with high pruning rates to those with low sensitivity; The sensitivity evaluation index is the accuracy loss of the model after the convolution kernel of a single convolutional layer is pruned.
7. The convolution kernel shaping and pruning method of a convolutional neural network structure according to claim 1, characterized in that: The construction of the adaptive selection mask strategy utilizes a dynamic programming algorithm, specifically: Set the target pruning rate and convert the target pruning rate into the FLOPs upper limit of the pruning model; Calculate the model accuracy of each convolutional layer at different pruning rates, output the model accuracy as a sensitivity table, and construct the FLOPs of each convolutional layer at different pruning rates and output the table; The knapsack problem is constructed using the FLOPs of convolutional layers with different pruning rates. The optimal mask combination is obtained by solving the knapsack problem. The closer the sum of the FLOPs of convolutional layers with different pruning rates is to the FLOPs upper limit of the pruning model, the smaller the overall sensitivity. The optimal mask combination is used to configure the corresponding mask for the convolution kernel of each convolution layer.
8. A convolution kernel integer pruning system for a convolutional neural network structure, characterized in that: include: A candidate mask set module is defined, which is used to define a candidate mask set according to the pruning rate for each convolution kernel to be pruned in the convolutional network through an adaptive mask selection strategy to obtain a selected candidate mask set; A normalized mask weight module is used to assign a normalized weight to each mask in the selected candidate mask set to obtain a normalized mask weight, where the normalized mask weight is represented by the product of the normalized learnable parameter and the temperature; The mask calculation convolution kernel weight module is used to use the normalized mask weight to weight the candidate mask set to calculate the mask of the convolution kernel to be integer pruned, and obtain the convolution kernel weight after masking; The training converges to obtain the final mask module, which is used to perform sparse training on the convolution network according to the masked convolution kernel weights until convergence, and obtain the final mask of each convolution kernel. The final mask of each convolution kernel is the mask with the largest weight of each convolution kernel. The pruning module is completed, which is used to replace the weight of the convolution kernel to be integer pruned with the final mask of each convolution kernel to complete the pruning.
9. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a convolution kernel integer pruning method for a convolutional neural network structure as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the convolution kernel integer pruning method of a convolutional neural network structure described in any one of claims 1 to 7.
Citation Information
Cited By
Image data processing method and system based on lightweight sparse neural network
CN121190950A