Multi-hardware energy-consumption-oriented channel pruning method and related product
By combining feature distribution and multi-objective evolutionary solution model, pruning solutions are provided for multiple hardware devices, which solves the problem of low pruning efficiency in the existing technology and realizes more efficient CNN model deployment.
Patent Information
- Application Number
- PCT/CN2023/134605
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-05
AI Technical Summary
The existing hardware-oriented pruning method can only generate an energy efficient CNN model for a specific energy budget of a hardware device during a pruning process, resulting in low pruning efficiency.
By combining the feature distribution of the original network model and the multi-objective evolutionary solution model, pruning solutions are provided for multiple hardware devices in a single pruning process to optimize the importance of filters and energy consumption.
The pruning efficiency of the convolutional neural network model is improved and the CNN model can be deployed more efficiently in cross-platform dynamic deployment scenarios.
Smart Images

Figure CN2023134605_05062025_PF_FP_ABST
Abstract
Description
A multi-hardware energy consumption-oriented channel pruning method and related products Technical Field
[0001] The present application relates to the field of model compression technology, and in particular to a multi-hardware energy consumption-oriented channel pruning method and related products. Background Art
[0002] Convolutional Neural Networks (CNNs) are a class of feedforward neural networks with deep structures that incorporate convolutional computations. They are one of the representative algorithms for deep learning. CNNs have demonstrated outstanding performance in various computer vision tasks, but the extensive computation and data movement of CNN models results in high energy consumption, which poses a challenge to their deployment in energy-constrained environments. Therefore, reducing the energy consumption of CNN models is crucial for their deployment on battery-powered edge devices.
[0003] Among existing model compression techniques, structured pruning has gained increasing attention due to its ease of deployment on general-purpose hardware. However, current hardware-specific pruning methods only generate an energy-efficient CNN model for a specific energy budget on a hardware device during a single pruning process, resulting in low pruning efficiency.
[0004] Therefore, how to improve pruning efficiency is an urgent problem that those skilled in the art need to solve.
[0005] Summary of the Invention
[0006] Based on the above problems, the present application provides a multi-hardware energy consumption-oriented channel pruning method and related products. Through a multi-objective evolutionary solution model, a pruning solution is provided for multiple hardware devices in one pruning process, solving the problem of low pruning efficiency in the existing technology.
[0007] In a first aspect, the present application provides a multi-hardware energy consumption-oriented channel pruning method, comprising:
[0008] Combined with the feature distribution of the original network model, the feature distribution difference evaluation model is used to rank the importance of the filters in the pruned convolutional neural network model, and the filters with the lowest importance ranking are deleted to obtain the candidate first pruned model;
[0009] Determining the energy consumption of the candidate first pruning model using an energy consumption estimation model based on actual measurement data;
[0010] Using a multi-objective evolutionary solution model to weigh the importance of the filter in the candidate first pruning model and the energy consumption, and obtain a low-energy pruning solution corresponding to each hardware device;
[0011] The low-energy pruning scheme is used to prune the convolutional neural network model to be pruned to obtain a second pruned model corresponding to each hardware device.
[0012] Optionally, the step of combining the feature distribution of the original network model and ranking the importance of filters in the to-be-pruned convolutional neural network model using a feature distribution difference evaluation model includes:
[0013] Combined with the evaluation image data, the feature distribution of the feature map of each layer in the convolutional neural network model to be pruned is determined;
[0014] In combination with the feature distribution, based on the feature distribution difference evaluation model, a maximum mean difference function is used to perform a difference evaluation on the feature distribution to obtain a difference evaluation result value;
[0015] The importance of filters corresponding to each feature graph of the convolutional neural network model to be pruned is sorted based on the evaluation result value.
[0016] Optionally, determining the energy consumption of the candidate first pruning model by using an energy consumption estimation model based on actual measurement data includes:
[0017] Build a lookup table for each hardware device based on actual measurement data;
[0018] Based on the lookup table, an energy consumption estimation model is used to determine the energy consumption of the candidate first pruning model.
[0019] Optionally, determining the energy consumption of the candidate first pruning model by using an energy consumption estimation model includes:
[0020] Determining the energy consumption value of each layer of the candidate first pruning model using an energy consumption estimation model;
[0021] The energy consumption values are summed to determine the energy consumption of the candidate first pruning model.
[0022] Optionally, the multi-objective evolutionary solution model is used to weigh the importance of the filter in the candidate first pruning model and the energy consumption to obtain a low-energy pruning solution corresponding to each hardware device, including:
[0023] Constructing a multi-objective evolutionary solution model based on the importance of the filters in the candidate first pruning model and the energy consumption;
[0024] Solving the multi-objective evolutionary solution model using a layer-by-layer pruning strategy to explore energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned;
[0025] Based on the energy-efficient pruning solution, a low-energy pruning solution corresponding to each hardware device is determined.
[0026] Optionally, the multi-objective evolutionary solution model is solved by adopting a layer-by-layer pruning strategy to explore an energy-efficient pruning solution for each layer of the convolutional neural network model to be pruned, including:
[0027] Based on the multi-objective evolutionary solution model, construct a single-layer objective evolutionary solution model for each layer of the convolutional neural network model to be pruned;
[0028] The single-layer objective evolution solution model is used to explore energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned.
[0029] Optionally, the method further includes:
[0030] When new hardware devices are introduced, obtain the hardware characteristics of the new hardware devices;
[0031] Identifying a hardware device in an existing hardware device repository that has characteristics most similar to the hardware;
[0032] A hardware device with characteristics most similar to the hardware is used as a proxy to determine a pruning solution for the new hardware device.
[0033] In a second aspect, the present application provides a multi-hardware energy consumption-oriented channel pruning device, comprising:
[0034] A deletion module is used to combine the feature distribution of the original network model and use the feature distribution difference evaluation model to rank the importance of the filters in the pruned convolutional neural network model, and delete the filter with the lowest importance ranking to obtain the candidate first pruned model;
[0035] a determination module, configured to determine the energy consumption of the candidate first pruning model using an energy consumption estimation model based on actual measurement data;
[0036] a processing module, configured to use a multi-objective evolutionary solution model to weigh the importance of filters in the candidate first pruning model and the energy consumption, and obtain a low-energy pruning solution corresponding to each hardware device;
[0037] A pruning module is used to prune the convolutional neural network model to be pruned using the low-energy pruning solution to obtain a second pruned model corresponding to each hardware device.
[0038] In a third aspect, the present application provides a multi-hardware energy consumption-oriented channel pruning device, including:
[0039] Memory for storing computer programs;
[0040] A processor is configured to implement the steps of any of the above-mentioned multi-hardware energy consumption-oriented channel pruning methods when executing the computer program.
[0041] In a fourth aspect, the present application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-hardware energy consumption-oriented channel pruning method as described in any one of the above items.
[0042] It can be seen from the above technical solutions that compared with the existing technology, this application has the following advantages:
[0043] This application first combines the feature distribution of the original network model, and uses the feature distribution difference evaluation model to rank the importance of the filters in the pruned convolutional neural network model, and deletes the filters with the lowest importance ranking to obtain the candidate first pruned model. Then, based on the actual measurement data, the energy consumption estimation model is used to determine the energy consumption of the candidate first pruned model, and the importance and energy consumption of the filters in the candidate first pruned model are weighed using the multi-objective evolutionary solution model to obtain the low-energy pruning scheme corresponding to each hardware device. Finally, the low-energy pruning scheme is used to prune the pruned convolutional neural network model to obtain the second pruning model corresponding to each hardware device. In this way, through the multi-objective evolutionary solution model, a pruning scheme is provided for multiple hardware devices in one pruning process, thereby improving the pruning efficiency of the convolutional neural network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] FIG1 is a flow chart of a multi-hardware energy consumption-oriented channel pruning method provided by the present application;
[0045] FIG2 is a schematic structural diagram of an all-in-one energy consumption-oriented channel pruning framework provided by the present application;
[0046] FIG3 is a schematic diagram of the pruning results of VGG-16 on the CIFAR-10 dataset provided by this application;
[0047] FIG4 is a schematic diagram of the pruning results of ResNet-56 on the CIFAR-10 dataset provided by this application;
[0048] FIG5 is a schematic diagram of the pruning results of ResNet-50 on the ImageNet dataset provided by this application;
[0049] FIG6 is a schematic diagram of the pruning results of MobileNet-V2 on the ImageNet dataset provided by this application;
[0050] FIG7 is a schematic structural diagram of a multi-hardware energy consumption-oriented channel pruning device provided in this application. DETAILED DESCRIPTION
[0051] As mentioned above, existing hardware-oriented pruning methods suffer from low pruning efficiency. Specifically, hardware-oriented pruning methods are better at reducing energy consumption than model-oriented pruning methods. However, the increasing number of hardware devices and the significantly different energy budgets between different devices pose new challenges to existing hardware-oriented pruning methods. In cross-platform dynamic deployment scenarios, existing hardware-oriented pruning methods can only generate an energy-efficient CNN model for a specific energy budget of a hardware device in one pruning process. When dealing with the various requirements of cross-platform dynamic deployment scenarios involving numerous energy budgets and hundreds of different device types, the pruning cost of existing hardware-oriented pruning methods increases linearly with the energy budget and the number of hardware devices, and the pruning efficiency also decreases accordingly.
[0052] To solve the above problems, the present application provides a multi-hardware energy consumption-oriented channel pruning method, including: first, combining the feature distribution of the original network model, using the feature distribution difference evaluation model to rank the importance of the filters in the pruned convolutional neural network model, and deleting the filters with the lowest importance ranking to obtain a candidate first pruned model. Then, based on the actual measurement data, the energy consumption of the candidate first pruned model is determined using the energy consumption estimation model, and the importance and energy consumption of the filters in the candidate first pruned model are weighed using the multi-objective evolutionary solution model to obtain a low-energy pruning solution corresponding to each hardware device. Finally, the low-energy pruning solution is used to prune the pruned convolutional neural network model to obtain a second pruned model corresponding to each hardware device.
[0053] In this way, through the multi-objective evolutionary solution model, pruning solutions are provided for multiple hardware devices in one pruning process, thereby improving the pruning efficiency of the convolutional neural network model.
[0054] It should be noted that the multi-hardware energy consumption-oriented channel pruning method and related products provided in this application can be applied to the field of model compression technology. The above is only an example and does not limit the application field of the multi-hardware energy consumption-oriented channel pruning method and related products provided in this application.
[0055] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0056] FIG1 is a flow chart of a multi-hardware energy consumption-oriented channel pruning method provided by the present application. In conjunction with FIG1 , the multi-hardware energy consumption-oriented channel pruning method provided by the present application may include:
[0057] S101: Combined with the feature distribution of the original network model, the feature distribution difference evaluation model is used to rank the importance of the filters in the pruned convolutional neural network model, and the filter with the lowest importance ranking is deleted to obtain the candidate first pruned model.
[0058] In practical applications, existing hardware-oriented pruning methods can only generate an energy-efficient CNN model for a specific energy budget of a hardware device in one pruning process. When dealing with various requirements of cross-platform dynamic deployment scenarios involving numerous energy budgets and hundreds of different device types, the pruning cost of existing hardware-oriented pruning methods will increase linearly with the energy budget and the number of hardware devices, and the pruning efficiency will also decrease accordingly. Therefore, the present application provides a multi-hardware energy consumption-oriented channel pruning method, which provides pruning solutions for multiple hardware devices in one pruning process through a multi-objective evolutionary solution model to improve pruning efficiency. Figure 2 is a structural schematic diagram of an all-in-one energy consumption-oriented channel pruning framework provided by the present application. In conjunction with Figure 2, the present application proposes an all-in-one energy consumption-oriented channel pruning framework (AEP) as a multi-hardware energy consumption-oriented channel pruning device to implement the multi-hardware energy consumption-oriented channel pruning method proposed in the present application, and the AEP framework consists of a feature distribution difference evaluator (FDD), an energy consumption estimator (ECE) and a multi-objective evolutionary solver (MOES). FDD evaluates filter importance from the perspective of feature distribution and then removes filters with minimal impact on feature distribution. For dynamic, cross-platform deployments across multiple devices, FDD uses a feature distribution difference assessment model to assess filter importance on a small amount of evaluation data. This determines the importance of each filter in the CNN to the feature distribution of the current layer. After determining the results, all filters are ranked in descending order of importance. The lowest-ranked filter, representing the least important filter, is then removed to obtain the candidate first pruning model.
[0059] In addition, since the ranking methods for the importance of filters in the pruned convolutional neural network model are different, this application can illustrate a possible ranking method.
[0060] In one embodiment, the importance of filters in a pruned convolutional neural network model is ranked. Accordingly, the importance of filters in the pruned convolutional neural network model is ranked using a feature distribution difference evaluation model based on the feature distribution of the original network model, including:
[0061] Combined with the evaluation image data, the feature distribution of the feature map of each layer in the convolutional neural network model to be pruned is determined;
[0062] In combination with the feature distribution, based on the feature distribution difference evaluation model, a maximum mean difference function is used to perform a difference evaluation on the feature distribution to obtain a difference evaluation result value;
[0063] The importance of filters corresponding to each feature graph of the convolutional neural network model to be pruned is sorted based on the evaluation result value.
[0064] For convolutional neural networks, if a feature map has almost no effect on the feature distribution of the layer, then the importance of the filter corresponding to this feature map to the current hardware device can be considered unimportant. Therefore, deleting these feature maps rarely affects the network capacity. Based on this idea, first, combined with a small amount of evaluation image data, the feature distribution of each feature map of the convolutional neural network model to be pruned is determined by the FDD in the AEP, and then combined with the feature distribution, the maximum mean difference function is used to measure the difference between the original feature distribution and the pruned feature distribution to obtain the difference evaluation result. Specifically, let m and n represent the original feature distribution and the pruned feature distribution, respectively. The maximum mean difference function based on the feature distribution difference evaluation model is defined in the data space Z as follows:
[0065] This function can be used as a feature distribution difference evaluation model, where F is a function class: f: Z→R. Let F represent the unit ball in the universal reproducing kernel Hilbert space (RKHS), denoted by H. In RKHS, f(a) can be expressed as: f(a) =<f,θ(a)> H , where θ:Z→H represents the feature space mapping from Z to H. In addition, FDD can be rewritten as:
[0066] Let O={o 1 ,…o b} and P = {p 1 ,…p b} denote independent and identically distributed samples drawn from feature distributions m and n, respectively. Here b denotes the number of images used to evaluate the importance of the filter. i ∈R C×R and p i ∈R C×R , where C and R represent the number of output channels and resolution of the feature map, respectively. The empirical estimate of FDD can be expressed as:
[0067] Then the kernel technique is introduced, and the above formula can be expressed as:
[0068] in and k(.,.) is a kernel function used to map sample vectors to high-dimensional feature space. and Respectively represent o i and p i Then the polynomial kernel function is used to project the sample vector into the high-dimensional feature space, which can be defined as k(x, y) = (x T y+c) d In this way, if we empirically set c = 0 and d = 2, we can get the FDD value, which is the difference evaluation result value. It should be noted that AEP ranks the importance of filters according to the FDD value, and smaller FDD values correspond to higher filter importance.
[0069] S102: Based on actual measurement data, determine the energy consumption of the candidate first pruning model using an energy consumption estimation model.
[0070] In practice, the ECE in the AEP evaluates the energy consumption of each candidate first pruned model on each hardware device to be deployed. Specifically, the ECE obtains actual measurement data for each first pruned model and then determines the energy consumption of the current candidate first pruned model based on the energy consumption estimation model.
[0071] In addition, since there are different ways to determine the energy consumption of the candidate first pruning model using the energy consumption estimation model based on actual measurement data, this application can illustrate one possible determination method.
[0072] In one case, regarding how to determine the energy consumption of the candidate first pruning model based on actual measurement data, S102: based on the actual measurement data, using the energy consumption estimation model to determine the energy consumption of the candidate first pruning model, may specifically include:
[0073] Build a lookup table for each hardware device based on actual measurement data;
[0074] Based on the lookup table, an energy consumption estimation model is used to determine the energy consumption of the candidate first pruning model.
[0075] In practical applications, the ECE developed within AEP can efficiently estimate the energy consumption of CNN models using a small amount of real-world measurement data. Unlike previous approaches that designed hardware-specific energy consumption models, ECE eliminates the need for specialized hardware-related knowledge, thereby enhancing the AEP framework's adaptability to a wide range of hardware devices. Specifically, ECE first constructs a corresponding lookup table for each hardware device using real-world measurement data. Based on the energy consumption estimation model, these constructed lookup tables are then used to estimate the energy consumption of candidate first-pruning models.
[0076] In addition, since there are different ways to determine the energy consumption of the candidate first pruning model, this application may describe a possible determination method.
[0077] In one embodiment, the energy consumption of the candidate first pruning model is determined by using an energy consumption estimation model, including:
[0078] Determining the energy consumption value of each layer of the candidate first pruning model using an energy consumption estimation model;
[0079] The energy consumption values are summed to determine the energy consumption of the candidate first pruning model.
[0080] In practical applications, the energy consumption of the candidate first pruning model can be regarded as the sum of the energy consumption of each layer, as an energy consumption estimation model, as follows:
[0081] By using the accumulation algorithm, the energy consumption value of each layer of the candidate first pruning model is determined and summed up to determine the total energy consumption of the candidate first pruning model. represents the pruning weight, represents the energy consumption of pruned weights. The energy consumption of pruned weights can be evaluated using existing monitoring tools.
[0082] S103: Using a multi-objective evolutionary solution model, the importance of the filter in the candidate first pruning model and the energy consumption are weighed to obtain a low-energy pruning solution corresponding to each hardware device.
[0083] In practical applications, to simultaneously optimize energy consumption on multiple hardware devices, the multi-hardware energy-oriented channel pruning problem is modeled as a multi-objective optimization problem, aiming to find a series of Pareto-optimal solutions that achieve an excellent trade-off between filter importance and energy consumption across multiple hardware devices. Specifically, the MOES in AEP is used to quickly obtain a series of energy-efficient solutions for several energy budgets on multiple hardware devices, and a multi-objective evolutionary solution model is constructed, as shown below:
[0084] The importance of the filters in the candidate first pruning model and the energy consumption are weighed by the multi-objective evolutionary solution model to obtain the pruning scheme corresponding to each hardware device. In the above formula, L(·,·) represents the filter importance evaluation function, Represents the training images. N represents the number of training images, and m represents the number of hardware devices. and W ={W 1 , W 2 , W 3 ,…,W L} represent the pruned weight and original weight respectively. E i (W) and Denote the energy consumption of the original model and the pruned model on the i-th hardware device, respectively. In this way, based on the multi-objective evolutionary solution model, the filter importance of the pruned model is maximized on all hardware devices, while the energy consumption of the pruned model is minimized, resulting in the optimal pruning solution for each hardware device.
[0085] In addition, since the methods for obtaining the pruning solutions corresponding to various hardware devices are different, this application may describe one possible method for obtaining the solutions.
[0086] In one case, regarding how to obtain a low-energy pruning solution corresponding to each hardware device, S103: using a multi-objective evolutionary solution model to weigh the importance of the filter in the candidate first pruning model and the energy consumption, to obtain a low-energy pruning solution corresponding to each hardware device, specifically including:
[0087] Constructing a multi-objective evolutionary solution model based on the importance of the filters in the candidate first pruning model and the energy consumption;
[0088] Solving the multi-objective evolutionary solution model using a layer-by-layer pruning strategy to explore energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned;
[0089] Based on the energy-efficient pruning solution, a low-energy pruning solution corresponding to each hardware device is determined.
[0090] In practical applications, MOES in AEP uses a layer-by-layer pruning strategy to iteratively explore energy-efficient pruning solutions for each layer of the CNN model, significantly improving the efficiency of multi-objective optimization problems. Specifically, for each hardware device, MOES first extracts the importance and energy consumption of the filters in the candidate first pruning model, and then constructs a multi-objective evolutionary solution model based on this. Through a layer-by-layer pruning strategy, the multi-objective evolutionary solution model is converted into an optimization model for each CNN layer and solved, thereby achieving the goal of exploring energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned. Finally, based on the energy-efficient pruning solution, the corresponding low-energy pruning solution for each hardware device is determined.
[0091] In addition, since the methods for exploring energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned are different, this application can illustrate a possible exploration method.
[0092] In one embodiment, an energy-efficient pruning solution for each layer of a convolutional neural network model to be pruned is explored. Accordingly, the multi-objective evolutionary solution model is solved using a layer-by-layer pruning strategy to explore an energy-efficient pruning solution for each layer of the convolutional neural network model to be pruned, including:
[0093] Based on the multi-objective evolutionary solution model, construct a single-layer objective evolutionary solution model for each layer of the convolutional neural network model to be pruned;
[0094] The single-layer objective evolution solution model is used to explore energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned.
[0095] In practical applications, energy consumption across different hardware devices does not introduce conflicts. When the energy consumption of a CNN model on one hardware device is reduced, it may also lead to a reduction in energy consumption of the same model on other hardware devices. In addition, because the impact of deleting a filter mainly depends on its independence from other filters in the same layer, the layer-by-layer pruning strategy has almost no effect on the performance of the pruned model and can greatly improve the solution efficiency of the multi-objective evolutionary solution model. Based on the multi-objective evolutionary solution model, this application constructs a single-layer target evolutionary solution model for each layer of the candidate first pruning model. For the i-th layer, there is a model as shown below:
[0096] where e i (W i )and where represents the energy consumption of the original and pruned weights of the i-th layer on the j-th target hardware device, respectively. The single-layer objective evolutionary solution model is based on an improved multi-objective evolutionary algorithm (NSGA-III). It uses a segment-by-segment budget selection strategy to efficiently search for energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned. Specifically, for a new individual Q and a parent individual P, MOES evenly divides the combined individual set P∪Q into T segments based on their energy consumption. For each segment, MOES uses an elitist selection strategy to select the P / T best individuals from all individuals in that segment to form the next generation. Finally, MOES obtains a range of energy-efficient pruning solutions across a wide range of energy consumption levels.
[0097] S104: Pruning the convolutional neural network model to be pruned using the low-energy pruning solution to obtain a second pruned model corresponding to each hardware device.
[0098] In actual applications, the above trade-offs are used to determine the pruning solution for each hardware device. Each model is then pruned using the pruning solution corresponding to each hardware device, resulting in a second pruned model for each hardware device, which is then deployed.
[0099] In addition, since the methods for obtaining the model pruning solution for the newly introduced hardware device are different, this application can describe a possible acquisition method.
[0100] In one embodiment, the method further includes:
[0101] When new hardware devices are introduced, obtain the hardware characteristics of the new hardware devices;
[0102] Identifying a hardware device in an existing hardware device repository that has characteristics most similar to the hardware;
[0103] A hardware device with characteristics most similar to the hardware is used as a proxy to determine a pruning solution for the new hardware device.
[0104] In practical applications, the AEP framework can simultaneously provide energy-efficient pruning solutions for multiple different hardware devices, as well as multiple energy budgets for a single hardware device, thereby meeting the diverse needs of dynamic, cross-platform deployment scenarios. If a new hardware device is introduced after all hardware devices have been deployed, the AEP framework analyzes the hardware characteristics of the newly introduced device and identifies a similar device from a repository of existing hardware devices. This similar device is then used as a proxy to determine the appropriate pruning solution for the newly introduced device, enabling rapid, energy-efficient deployment of a wide range of hardware devices.
[0105] In addition, the AEP framework is compared with various SOTA pruning methods, including model-oriented pruning methods and hardware-oriented pruning methods, through comparative experiments. Specifically, the classification performance is evaluated on the CIFAR-10 and ImageNet datasets. Since previous methods mainly focus on reducing FLOPs and parameters, we also use the number of FLOPs and the number of parameters to evaluate the performance of pruning methods. In addition, we evaluate energy consumption on six hardware devices. For a fair comparison, for all other methods, we recreate the pruned models according to their official implementations and then measure their energy consumption. Figure 3 is a schematic diagram of the pruning results of VGG-16 on the CIFAR-10 dataset provided in this application. Figure 4 is a schematic diagram of the pruning results of ResNet-56 on the CIFAR-10 dataset provided in this application. Combined with Figures 3 and 4, the final pruning results on the CIFAR-10 dataset are as follows: For VGG-16, AEP consistently outperforms state-of-the-art pruning methods at three different pruning ratios. For ResNet-56, AEP reduces 55.61% FLOPs and 51.76% of parameters while improving the accuracy of the original ResNet-56 model by 0.44%. Furthermore, AEP leads in top-1 accuracy and FLOPs at both low and high pruning ratios. Figure 5 is a schematic diagram of the pruning results of ResNet-50 on the ImageNet dataset provided by this application. Figure 6 is a schematic diagram of the pruning results of MobileNet-V2 on the ImageNet dataset provided by this application. Combined with Figures 5 and 6, the pruning results on the ImageNet dataset are as follows: Compared to the original ResNet-50 model, AEP reduces 58.83% FLOPs and 48.89% of parameters, while improving top-1 accuracy by 0.06%. Furthermore, AEP outperforms SOTA pruning methods in terms of both FLOPs count and classification performance at all three pruning ratios. AEP improves the top-1 accuracy of the original MobileNet-V2 model by 0.11% while reducing FLOPs by 26.60%. Furthermore, AEP achieves a better balance between classification performance and FLOPs count at both low and high pruning ratios, outperforming SOTA pruning methods. In particular, at low pruning ratios, AEP achieves higher accuracy compared to the recent SOTA pruning method, HALP. Therefore, AEP can efficiently compress network architectures with depthwise separable convolutions. Furthermore, for several energy budgets on a single hardware device, AEP also achieves a better trade-off between accuracy and energy consumption, outperforming SOTA pruning methods and significantly reducing pruning costs, thereby facilitating the efficient deployment of CNN models in cross-platform dynamic deployment scenarios.
[0106] In summary, the present application first combines a small amount of evaluation image data, uses a feature distribution difference evaluation model to rank the importance of the filters in the pruned convolutional neural network model, and deletes the filters with the lowest importance ranking to obtain a candidate first pruned model. Then, based on the actual measurement data, the energy consumption estimation model is used to determine the energy consumption of the candidate first pruned model, and the importance and energy consumption of the filters in the candidate first pruned model are weighed using a multi-objective evolutionary solution model to obtain a low-energy pruning solution corresponding to each hardware device. Finally, the low-energy pruning solution is used to prune the pruned convolutional neural network model to obtain a second pruning model corresponding to each hardware device. In this way, through the multi-objective evolutionary solution model, a pruning solution is provided for multiple hardware devices in one pruning process, thereby improving the pruning efficiency of the convolutional neural network model.
[0107] Based on the multi-hardware energy consumption-oriented channel pruning method provided in the above embodiment, the present application also provides a multi-hardware energy consumption-oriented channel pruning device. The multi-hardware energy consumption-oriented channel pruning device is described below with reference to the embodiments and drawings.
[0108] FIG7 is a schematic diagram of the structure of a multi-hardware energy consumption-oriented channel pruning device provided in an embodiment of the present application. Referring to FIG7 , the multi-hardware energy consumption-oriented channel pruning device 200 provided in an embodiment of the present application includes:
[0109] A deletion module 201 is configured to rank the importance of filters in the to-be-pruned convolutional neural network model using a feature distribution difference evaluation model in combination with the feature distribution of the original network model, and delete the filter with the lowest importance ranking to obtain a candidate first pruned model;
[0110] A determination module 202 is configured to determine the energy consumption of the candidate first pruning model using an energy consumption estimation model based on actual measurement data;
[0111] A processing module 203 is configured to use a multi-objective evolutionary solution model to weigh the importance of the filters in the candidate first pruning model and the energy consumption to obtain a low-energy pruning solution corresponding to each hardware device;
[0112] The pruning module 204 is used to prune the convolutional neural network model to be pruned using the low-energy pruning solution to obtain a second pruned model corresponding to each hardware device.
[0113] As an implementation method, regarding how to use the feature distribution difference evaluation model to rank the importance of filters in the convolutional neural network model, the deletion module 201 can be specifically used to:
[0114] Combined with the evaluation image data, the feature distribution of the feature map of each layer in the convolutional neural network model to be pruned is determined;
[0115] In combination with the feature distribution, based on the feature distribution difference evaluation model, a maximum mean difference function is used to perform a difference evaluation on the feature distribution to obtain a difference evaluation result value;
[0116] The importance of filters corresponding to each feature graph of the convolutional neural network model to be pruned is sorted based on the evaluation result value.
[0117] As an embodiment, regarding how to determine the energy consumption of the candidate first pruning model using the energy consumption estimation model based on actual measurement data, the determination module 202 is specifically configured to:
[0118] Build a lookup table for each hardware device based on actual measurement data;
[0119] Based on the lookup table, an energy consumption estimation model is used to determine the energy consumption of the candidate first pruning model.
[0120] As an embodiment, regarding how to use the energy consumption estimation model to determine the energy consumption of the candidate first pruning model, the determination module 202 may be specifically configured to:
[0121] Determining the energy consumption value of each layer of the candidate first pruning model using an energy consumption estimation model;
[0122] The energy consumption values are summed to determine the energy consumption of the candidate first pruning model.
[0123] As an embodiment, regarding how to use a multi-objective evolutionary solution model to weigh the importance and energy consumption of filters in a candidate first pruning model, the processing module 203 includes: a construction module, an exploration module, and a determination submodule;
[0124] A construction module, configured to construct a multi-objective evolutionary solution model based on the importance of filters in the candidate first pruning model and the energy consumption;
[0125] An exploration module, configured to solve the multi-objective evolutionary solution model using a layer-by-layer pruning strategy, and explore an energy-efficient pruning solution for each layer of the convolutional neural network model to be pruned;
[0126] The determination submodule is configured to determine a low-energy pruning solution corresponding to each hardware device based on the energy-efficient pruning solution.
[0127] As an implementation method, the exploration module is specifically used to explore energy-efficient pruning solutions for each layer of a convolutional neural network model to be pruned:
[0128] Based on the multi-objective evolutionary solution model, construct a single-layer objective evolutionary solution model for each layer of the convolutional neural network model to be pruned;
[0129] The single-layer objective evolution solution model is used to explore energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned.
[0130] As an implementation method, in order to cope with the model deployment of newly introduced hardware devices, the multi-hardware energy consumption-oriented channel pruning device 200 further includes: a deployment module;
[0131] The deployment module is used to obtain the hardware characteristics of new hardware devices when new hardware devices are introduced;
[0132] Identifying a hardware device in an existing hardware device repository that has characteristics most similar to the hardware;
[0133] A hardware device with characteristics most similar to the hardware is used as a proxy to determine a pruning solution for the new hardware device.
[0134] In summary, the present application first combines a small amount of evaluation image data, uses a feature distribution difference evaluation model to rank the importance of the filters in the pruned convolutional neural network model, and deletes the filters with the lowest importance ranking to obtain a candidate first pruned model. Then, based on the actual measurement data, the energy consumption estimation model is used to determine the energy consumption of the candidate first pruned model, and the importance and energy consumption of the filters in the candidate first pruned model are weighed using a multi-objective evolutionary solution model to obtain a low-energy pruning solution corresponding to each hardware device. Finally, the low-energy pruning solution is used to prune the pruned convolutional neural network model to obtain a second pruning model corresponding to each hardware device. In this way, through the multi-objective evolutionary solution model, a pruning solution is provided for multiple hardware devices in one pruning process, thereby improving the pruning efficiency of the convolutional neural network model.
[0135] In addition, the present application also provides a multi-hardware energy consumption-oriented channel pruning device, including: a memory for storing a computer program; a processor for implementing the steps of the multi-hardware energy consumption-oriented channel pruning method as described in any one of the above items when executing the computer program.
[0136] In addition, the present application also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-hardware energy consumption-oriented channel pruning method as described in any of the above items are implemented.
[0137] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A channel pruning method oriented to multi-hardware energy consumption, characterized in that, the method includes: Combining the feature distribution of the original network model, using the feature distribution difference evaluation model to rank the importance of the filters in the convolutional neural network model to be pruned, and deleting the filters with the lowest importance ranking to obtain a candidate first pruning model; Based on the actual measurement data, using the energy consumption estimation model to determine the energy consumption of the candidate first pruning model; Using the multi-objective evolutionary solution model to balance the importance of the filters in the candidate first pruning model and the energy consumption to obtain a low-energy consumption pruning scheme corresponding to each hardware device; Using the low-energy consumption pruning scheme to prune the convolutional neural network model to be pruned to obtain a second pruning model corresponding to each hardware device.
2. The method according to claim 1, characterized in that, the combining the feature distribution of the original network model, using the feature distribution difference evaluation model to rank the importance of the filters in the convolutional neural network model to be pruned includes: Combining the evaluation image data to determine the feature distribution of the feature maps of each layer in the convolutional neural network model to be pruned; Combining the feature distribution, based on the feature distribution difference evaluation model, using the maximum mean difference function to evaluate the difference of the feature distribution to obtain a difference evaluation result value; Ranking the importance of the filters corresponding to the feature maps of the convolutional neural network model to be pruned based on the evaluation result value.
3. The method according to claim 1, characterized in that, the based on the actual measurement data, using the energy consumption estimation model to determine the energy consumption of the candidate first pruning model includes: Constructing a look-up table for each hardware device based on the actual measurement data; Based on the look-up table, using the energy consumption estimation model to determine the energy consumption of the candidate first pruning model.
4. The method according to claim 3, characterized in that, the using the energy consumption estimation model to determine the energy consumption of the candidate first pruning model includes: Using the energy consumption estimation model to determine the energy consumption value of each layer of the candidate first pruning model; Summing the energy consumption values to determine the energy consumption of the candidate first pruning model.
5. The method according to claim 1, characterized in that, the using the multi-objective evolutionary solution model to balance the importance of the filters in the candidate first pruning model and the energy consumption to obtain a low-energy consumption pruning scheme corresponding to each hardware device includes: Constructing a multi-objective evolutionary solution model based on the importance of the filters in the candidate first pruning model and the energy consumption; Adopting a layer-by-layer pruning strategy to solve the multi-objective evolutionary solution model to explore an energy-efficient pruning solution for each layer of the convolutional neural network model to be pruned; Determining a low-energy consumption pruning scheme corresponding to each hardware device based on the energy-efficient pruning solution.
6. The method according to claim 5, characterized in that, The multi-objective evolutionary solution model is solved using a layer-by-layer pruning strategy to explore energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned, including: Based on the multi-objective evolutionary solution model, a single-layer objective evolutionary solution model for each layer of the convolutional neural network model to be pruned is constructed; Using the single-layer objective evolutionary solution model, energy-efficient pruning solutions for each layer of the convolutional neural network model to be pruned are explored.
7. According to the method described in claim 1, wherein, the method further includes: When a new hardware device is introduced, obtain the hardware characteristics of the new hardware device; Identify a hardware device in the existing hardware device repository that is most similar to the hardware characteristics; Use the hardware device most similar to the hardware characteristics as an agent to determine the pruning scheme for the new hardware device.
8. A channel pruning device oriented to multi-hardware energy consumption, wherein, it includes: A deletion module, configured to combine the feature distribution of the original network model, use the feature distribution difference evaluation model to rank the importance of the filters in the convolutional neural network model to be pruned, and delete the filters with the lowest importance ranking to obtain a candidate first pruning model; A determination module, configured to determine the energy consumption of the candidate first pruning model using an energy consumption estimation model based on actual measurement data; A processing module, configured to use a multi-objective evolutionary solution model to balance the importance of the filters in the candidate first pruning model and the energy consumption to obtain low-energy pruning solutions corresponding to each hardware device; A pruning module, configured to prune the convolutional neural network model to be pruned using the low-energy pruning solution to obtain a second pruning model corresponding to each hardware device.
9. A channel pruning device oriented to multi-hardware energy consumption, wherein, it includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the multi-hardware energy consumption oriented channel pruning method according to any one of claims 1 to 7 when executing the computer program.
10. A readable storage medium, wherein, the readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-hardware energy consumption oriented channel pruning method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Adjustable hardware aware pruning and mapping framework based on ReRAM neural network accelerator
CN112598129A
FPGA-oriented deep convolutional neural network accelerator and design method
CN113487012A
Cache allocation method of distributed real-time system
CN115905043A
Class adaptive model pruning method and device, electronic equipment and storage medium
CN116796823A
Apparatus and a method for neural network compression
US20220083866A1
Cited By
Efficient hybrid expert model deployment system and method based on input dynamic pruning
CN121094108A