A Componentization Method of Convolutional Neural Network Based on Structured Pruning in an Edge-Cloud Collaboration Scenario
Through the channel control gate technology optimized by structured pruning and biological evolution algorithms, the customized compression problem of large-scale convolutional neural networks in edge-cloud collaborative scenarios is solved, efficient image perception tasks of edge devices are realized, and the adaptability and performance of the model are improved.
Patent Information
- Application Number
- CN202310840084.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-07-10
AI Technical Summary
In the edge-cloud collaboration scenario, it is difficult for the existing technology to customize and compress large convolutional neural networks into suitable sub-models according to the specific task requirements of edge devices, resulting in limited storage and computing capabilities of edge devices and inability to efficiently complete image perception tasks.
Through the structured pruning method, the channel control gate is used to quantify the contribution degree of the convolution channel, combine with biological evolution algorithms to optimize the linear combination coefficients, remove the irrelevant convolution channel, and generate customized sub-models to adapt to the specific task requirements of edge devices.
It realizes a significant reduction in model scale within the acceptable range of accuracy, improves the storage and computing efficiency of edge devices, and adapts to different tasks.
Smart Images

Figure CN117057409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition in the cloud-edge collaboration scenario. Specifically, it relates to a method for componentizing large convolutional neural networks based on structured pruning. Background Art
[0002] Convolutional neural networks (CNNs) have been widely used in image classification scenarios due to their powerful feature representation capabilities, such as scene classification of remote sensing images and scene classification techniques for autonomous driving vehicles. At the same time, cloud-edge collaboration is becoming a solution for these scenarios that require real-time performance and efficiency. This solution places tasks such as data collection, analysis, and processing at the edge, and transmits data to the cloud for monitoring and storage through network forwarding. This method can reduce system response latency, cloud pressure, and bandwidth costs, and has strong robustness and practicality. However, the high storage and high computing characteristics of neural network models conflict with the low storage space and low computing power characteristics of edge devices. To address the large-scale trend of CNN migration from the cloud to the edge, many model compression techniques have been developed to solve the above challenges.
[0003] In the field of model compression, many theoretical methods have been proposed, including pruning, tensor decomposition, and knowledge distillation, etc. Among them, pruning the model generally generates a compact sub-network by pruning weights (fine-grained) or pruning filters / channels (coarse-grained). Although the weight pruning method can better maintain the final inference accuracy of the network, it leads to irregular calculations and memory accesses, limiting the parallelism of hardware implementation. In contrast, structured pruning is to delete a group of weights rather than pruning individual weights. With the increasing trend of the depth of neural networks, directly deleting an entire convolutional kernel or channel can greatly compress the model, thus enabling better deployment on edge devices with poor storage capabilities. In addition, structured pruning provides effective processing in hardware by leveraging data parallel architectures, improves network compression, and reduces storage requirements, and is more widely used. Due to the various advantages of structured pruning, many studies have been conducted to further develop this method: by ablating certain neurons in the convolutional layer, researchers can observe a significant accuracy drop in specific categories, which reveals the differential process of neurons at different layers and also reveals that category-related information is mainly concentrated in a few neurons. Some researchers impose sparsity constraints on the scaling factors of the batch normalization (Batch Normalization) layers of the model and identify channels with lower scaling factors as channels with less information; there is also to associate the output channels of each layer of the model with a scalar control gate to learn the contribution degree of different convolutional channels to the final output of the neural network. A larger gate value means the importance of the convolutional channel, otherwise the convolutional channel is less important.
[0004] However, most of the work in the above methods mainly focuses on removing redundant layers and parameters in the model to reduce the model size, but ignores the actual need that in some scenarios, edge devices only need to perceive a very small part of the total number of input categories. In practical applications, such as in the fields of remote sensing image classification, industrial defect detection, and autonomous driving scene recognition, edge devices only need to classify or study specific categories, rather than considering all categories. In these scenarios, the task requirements of different edge devices vary. For a specific device, only a few categories may be the objects of its concern because these categories are closely related to its specific research or monitoring purposes. Our method will combine the specific task requirements of edge devices to customize a submodel with a greatly reduced model size, providing a solution for deploying models on some edge devices in the edge-cloud collaboration scenario. Summary of the Invention
[0005] Object of the Invention:
[0006] In some edge-cloud collaboration scenarios, such as in the fields of remote sensing image classification, industrial defect detection, and autonomous driving scene recognition, the cloud has more data sources to obtain large-scale datasets and can train a large over-parameterized model to distinguish all image scenarios. Edge devices only need to classify or study specific categories, rather than considering all categories. This means that the models on these devices only need to focus on the classification of specific categories. Therefore, we propose a general method for componentizing large convolutional neural networks based on structured pruning. Here, componentization refers to decomposing a large model into several submodels for different services. The cloud combines the specific task requirements of edge devices and deploys the submodels to applicable edge devices to perform tasks such as perception and inference. By removing convolutional channels irrelevant to specific task inference through structured pruning, it is ensured that the submodel deployed on the edge side has a greatly reduced model size while the accuracy slightly decreases, enabling edge devices with limited storage capacity and energy supply to more conveniently complete the task of image perception. Finally, through the collaborative work of edge devices and the cloud, the cloud can customize submodels according to the needs of edge devices, achieve the componentization effect of convolutional neural networks, and provide support for model training, optimization, and update. Edge devices can also more focusedly process the perception and inference tasks of specific categories, further strengthening the advantages of edge-cloud collaboration.
[0007] The technical solution of the present invention is as follows:
[0008] A method for componentizing a convolutional neural network based on structured pruning in an edge-cloud collaboration scenario, comprising the following steps:
[0009] Step 1: Obtain the channel control gate A at the image level;
[0010] Step 1.1: Use the training set to pre-train the over-parameterized complete model N_full, with the model depth being K;
[0011] Step 1.2: Initialize the control gate λ k ;
[0012] Step 1.3: Obtain an image x, use the model N_full to predict this image, and get the inference result f θ (x); and use the model after adding the control gate λ k to predict this image, and get the inference result f θ (x,A); The model after adding the control gate λ k refers to the model obtained by multiplying the control gate λ k channel by channel with the output of the k-th layer of the model in the complete model; the dimension of the control gate λ k is the number of convolutional channels of the k-th layer of the model, A = {λ1,λ2…λ K};
[0013] Step 1.4: Calculate the loss function with
[0014]
[0015] where γ is a hyperparameter, and the meaning of the crossEntropy() function is the cross-entropy loss function, and K is the depth of the model;
[0016] And perform gradient update on A with
[0017]
[0018] During the update process, ensure that each value of λ k is non-negative; after the update, calculate that the predicted class label of the model with the control gate added is j = argmax N_full(x,A); if i = j at this time, then the channel gate A has reached convergence in this iteration;
[0019] Step 1.5: Repeat Step 1.4 T times to complete the training of the control gate A; at this time, all λ k are composed layer by layer to obtain the image-level channel control gate A;
[0020] Step 2: Use the image-level channel control gate A of different images of the same category to establish the class-level control gate Y:
[0021] Step 2.1: Execute Step 1 one by one for all n images [x1,x2,x3…x n under the same target category to obtain the image-level control gate [A x1 ,A x2 ,A x3…A xn ;
[0022] Step 2.2: Add up the image-level channel control gates corresponding to each of the n images to obtain the class-level control gate for the target class
[0023] Step 3: Customize the sub-model:
[0024] Step 3.1: According to the customization requirements, use the result of Step 2 to obtain the class-level control gates [Y1, Y2, Y3... Y v ;
[0025] Step 3.2: Randomly generate s groups of different parameters [c1, c2, c3... c v for [Y1, Y2, Y3... Y v ;
[0026] Step 3.3: For a certain group [c1, c2, c3... c v , calculate S according to the formula
[0027] S fusion = c1 * Y1 + c2 * Y2 + c3 * Y3 +... + c v * Y v
[0028] Calculate S fusion ;
[0029] Step 3.4: According to the given pruning ratio, find the control gates λ with relatively low values corresponding to the corresponding ratio in S fusion , remove the convolutional channels corresponding to these control gate values to obtain the pruned sub-model according to this S k ; Calculate the inference accuracy of this sub-model on the training set, and calculate the fitness = 1.0 / (complete model inference accuracy - this sub-model inference accuracy); fusion
[0030] Step 3.5: Execute Steps 3.3 and 3.4 for the s groups of different parameters in Step 3.2, and obtain the fitness fitness for each group of parameters respectively;
[0031] Step 3.6: According to the fitness of each group in the s groups of parameters, select two groups of parameters with the highest fitness for crossover and mutation to obtain the individuals of the next generation, and return to Step 3.3 until the required number of iterations is reached;
[0032] Step 3.7: Execute Steps 3.2 - 3.6 T times, and obtain a group of parameters with the highest fitness each time; Select the group of parameters with the highest fitness from the T times of parameters;
[0033] Step 3.8: Calculate S using the parameters obtained in Step 3.7 fusion , and then, according to the given pruning ratio, find the control gates λ with relatively low values in the S fusion calculated in this step k , remove the convolutional channels corresponding to these control gate values, and the pruned sub-model obtained according to this S fusion is the customized sub-model.
[0034] In addition, the present invention also provides a computer-readable storage medium and a computer system based on the above method. The computer-readable storage medium stores computer-executable instructions, and the instructions are used to implement the above method when executed. The computer system includes one or more processors and a computer-readable storage medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0035] Advantageous Effects
[0036] The present invention provides a method for componentizing a convolutional neural network based on structured pruning in an edge-cloud collaboration scenario. For different types of inputs, the contribution degrees of different convolutional channels to the final inference are quantified in the form of channel control gates to obtain the control gates corresponding to all types. According to requirements, for a given combination of types, the best pruning structure is found by linearly combining their corresponding control gates. The process of exploring the linear combination coefficients is obtained by iterating through processes such as natural selection, crossover, and mutation in simulating the principles of biological evolution in biology. Finally, based on the combined control gate information, the original model is pruned to obtain a customized sub-model. Since the excellent performance of the convolutional neural network in the image classification scenario requires sufficient neural network depth to extract the features of the input image, and the algorithm proposed by us can enable a large convolutional neural network with parameter redundancy to be customized and compressed according to different edge requirements, so that the scale of the model is greatly reduced on the premise that the model performance degradation is within an acceptable range. Compared with traditional model compression methods, the method of the present invention is task-based and can disassemble the original model into components more flexibly for different tasks.
[0037] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0038] The above and / or additional aspects and advantages of the present invention will become apparent and be easily understood from the description of the embodiments in conjunction with the following drawings, where:
[0039] Figure 1 : Overall framework diagram of the method;
[0040] Figure 2 : Graph of experimental verification results.
[0041] Figure 3 : Graph of experimental verification results.
[0042] Figure 4 : Graph of experimental verification results.
[0043] Figure 5 : Graph of experimental verification results. Specific implementation manner
[0044] The present invention proposes a general method for componentizing large convolutional neural networks based on structured pruning, which can disassemble a large model into several sub-models for different services, and then deploy the sub-models to applicable edge devices to perform tasks such as perception and inference. The cloud is responsible for training a large model with a complex network structure using a large dataset to discriminate all image scenes, and customizing sub-models for different edge devices, achieving the effect of componentizing a large model. The scale of the sub-model is greatly reduced compared to the original model, ensuring that edge devices with limited storage capacity and energy supply can more conveniently complete the task of image perception, and the accuracy degradation is within an acceptable range.
[0045] In the present invention, inspired by the distillation-guided routing method, we introduce the concept of channel control gates to identify the contribution degree of different channels in the convolutional neural network to the output of the final model. In our method, the channel control gates are used to find the key channels, rather than retaining the parameter weights of the original model. We merge the control gates of all samples under the same category to obtain class-level control gates, and use a linear combination method to fuse different class-level control gates. The optimal parameters of the linear combination are obtained by iterating through a process simulating biological evolution. Through the fused control gates, we can remove the corresponding channels with low contribution degrees, thereby obtaining a customized sub-model. The overall framework is as Figure 1 shown.
[0046] We associate the output channels of each layer of the convolutional neural network with the control gates to identify the contribution degree of different channels in the convolutional neural network to the output of the final model. Without making any modifications to the weights of the pre-trained original convolutional network model, during the period when the model obtains an input and performs forward inference to obtain a result, a group of control gates λ k will be multiplied by the output of the k-th layer of the model channel by channel, and the dimension of the control gate λ k is the number of convolutional channels of the k-th layer of the model. Then, extracting the contribution degree of each convolutional channel during one inference process of the model can be simplified to optimizing A = {λ1, λ2…λ K}, where K is the depth of the model. To make the convolutional channels activated by different category inputs as sparse and non-overlapping as possible, and to meet the ultimate requirement of model componentization, λ should satisfy two conditions:
[0047] 1. The value of λ is non-negative. A negative λ value will cancel out the original output in the network.
[0048] 2. λ should be as sparse as possible, that is, the variance between the control gate values should be as large as possible, and more values should be close to 0. This is more in line with the view of the interpretability of neural networks.
[0049] Specifically, the optimization objective for the control gate A is:
[0050]
[0051] Here, f θ (x) = [p1, p2, p3... p m and f θ (x, A) = [p′1, p′2, p′3... p′ m are the inference results of the original model for the input x and the prediction results of the model after adding the control gate respectively. m is the total number of categories, and γ is a hyperparameter. The meaning of the crossEntropy() function is the cross-entropy loss function, which is used to measure the difference between the probability distribution of the model output and the actual label. Here, we use it to calculate the difference in the prediction performance of the model before and after adding the control gate. The significance of this optimization objective is to make the channel control gate values obtained by training as sparse as possible on the premise of ensuring that the model prediction results remain unchanged.
[0052] At first, set all λ k channel gate values to 1 uniformly. This means that the contribution degrees of all convolutional channels are the same and do not affect the inference performance of the model. After one forward propagation of the model, the gradient of the control gate can be calculated by the following formula:
[0053]
[0054] Use the gradient update method to update and optimize the control gate values. The control gate converges after several updates. The corresponding channels with larger channel control gate values contribute more to the output of the final neural network, and the corresponding channels with smaller or even zero channel control gate values contribute less or no contribution to the output of the final model.
[0055] The specific steps can be expressed as:
[0056] Step 1.1: Use the training set to pre-train the over-parameterized complete model N_full with a model depth of K; all possible image categories are included in the training set; this process can be carried out on the cloud side with sufficient computing power and a large training set;
[0057] Step 1.2: Initialize all λ k to 1;
[0058] Step 1.3: Input the image x, and the predicted class of the original neural network is i = argmax N_full(x); where N_full(x) gives the probabilities that the model output image x belongs to all classes, and argmax refers to the class label corresponding to the maximum probability value;
[0059] Step 1.4: Calculate the loss function with
[0060]
[0061] and update the gradient of A, ensuring that each value of λ
[0062]
[0063] is non - negative during the update process; after the update, calculate the predicted class label of the model with the control gate added as j = argmax N_full(x, A); if i = j at this time, then the channel gate A has converged in this iteration. k
[0064] Step 1.5: Repeat Step 1.4 for T times to complete the training of the control gate A. After Step 1.4, A has converged, but it has not reached a very sparse state, and multiple iterations are required to make A sparse.
[0065] Step 1.6: At this time, all λ k layer - by - layer form A.
[0066] The control gate A obtained here is the control gate trained for the image x, and we call it the image - level channel control gate. For example, if the model depth K = 3, the number of channels in each layer is 3, and finally, after A is optimized, λ1 = [1, 2, 3], λ2 = [2, 2, 4], λ3 = [9, 0, 3], and A = [λ1, λ2, λ3].
[0067] Through the above process, we train and obtain the control gate for a single input, that is, a single image - level channel control gate. Previous studies have found that inputs predicted to be of the same class usually activate a unique set of convolutional channels, which are different from those activated by other inputs and may even be mutually exclusive. Therefore, we can obtain the class - level channel control gate by merging the channel control gates of all images of the same class. The merging method here is to directly add up the control gate values of the corresponding convolutional layers. Such a class - level control gate can reflect the contribution degrees of different channels when the neural network infers inputs of the same class.
[0068] The specific steps can be expressed as follows:
[0069] Step 2.1: For all the images [x1, x2, x3…x n under the target category, execute the above image-level channel control gate algorithm one by one to obtain the image-level control gate for each image, denoted as
[0070]
[0071] Step 2.2: Sum up the image-level channel control gates corresponding to the n images respectively to obtain the class-level control gate of the target category
[0072] For example, for image x1, λ1 = [1, 2, 3], λ2 = [2, 2, 4], λ3 = [9, 0, 3], then For the image x2 of the same target category, λ4 = [3, 3, 3], λ5 = [2, 2, 2], λ6 = [1, 1, 1], then
[0073] To customize a sub-model that can classify images of a specified category, we propose to fuse different class-level channel control gates, and then perform pruning and compression according to the fused control gates. We simplify the fusion of channel gates into a problem of linearly combining control gates of different categories. The coefficients of the linear combination can be adjusted so that the fused channel gates can help obtain the best model pruning effect. However, the weight of each control gate is unknown, and a simple fully connected neural network can be used to solve the problem of linear combination, that is, the input is the channel control gates of v (v < total number of categories m) categories, and the output is the fused channel control gate. The loss function is defined as the gap between the performance of the sub-model obtained by pruning according to the fused control gate and the performance of the original model. However, since the structural pruning process cannot calculate gradients, traditional gradient update methods cannot be directly applied to this problem. In this case, we search for the best combination parameters by simulating processes such as natural selection, crossover, and mutation in the principle of biological evolution.
[0074] We represent the fusion of control gates by the following formula:
[0075] S fusion = c1*Y1 + c2*Y2 + c3*Y3 + … + c v *Y v
[0076] We use the coefficients [c1, c2, c3…c of the v-category channel gates for linear combination vis represented as an individual in the population. For the convenience of subsequent operations, we store these individuals in binary form. We first initialize a population and generate a set of random binary strings. Each individual binary string contains the coefficients corresponding to a set of control gates, i.e., [c1, c2, c3…c v . Then, decode the binary string into the actual coefficients c1, c2, c3…c v , and calculate the fitness of this set of coefficients. The calculation of fitness will be evaluated according to the performance of the sub-model obtained by the fused channel control gate S fusion . Here, we use the reciprocal of the difference between the inference accuracy of the sub-model and the original model on the training set as the fitness metric, that is, the higher the accuracy, the greater the fitness. Next, select the two optimal sets of parameters as the parent individuals according to the fitness for crossover and mutation operations to generate the next generation of individuals. In the crossover operation, we select two binary strings of the parent individuals and divide the crossover point at a randomly selected position. Then, exchange the binary string segments after the crossover point to form two offspring individuals. Such a crossover operation can help us combine the "genes" of different individuals' advantages and generate more potential solutions. After obtaining the offspring individuals, we perform mutation operations on some of the binary bits, that is, flip the values at randomly selected bits. This mutation operation increases the diversity of the offspring population and provides new possible solutions during the search process. Repeat the above process to gradually improve the fitness of the individuals until the maximum number of iterations is reached or the optimal solution is found.
[0077] Through the fused channel control gate S fusion , we can determine the contributions of each convolutional channel under the specified category combination, and remove some convolutional channels according to the information of the control gate, so as to help us obtain a sub-model with the best performance. Generally speaking, we need to first train an over-parameterized large model, and then delete redundant parameters by various methods without affecting the accuracy. Retaining the optimal architecture and weights was once considered the key to finally obtaining an efficient model. And previous studies have found that: the pruned network can achieve the same accuracy whether it inherits the weights in the original network or not. This shows that the essence of channel pruning is to find a good pruning structure, that is, the number of channels to be retained layer by layer in the neural network. The obtained sub-model does not necessarily need to inherit the weights of the large model, because such an approach may cause the pruned model to fall into a local minimum, even though these weights are considered important under the pruning criterion. Therefore, after we remove some channels according to the channel control gate S fusion to obtain the sub-model, we will fine-tune the sub-model with a small amount of training data to significantly improve its inference performance, and the subsequent experimental part also verifies this point.
[0078] The specific steps can be expressed as:
[0079] Step 3.1: We need to fuse v class-level control gates for [Y1, Y2, Y3... Y v , where v < the total number of classes m. Initialize the coefficients corresponding to each class-level control gate: c1 = 1, c2 = 1, c3 = 1... c v = 1;
[0080] Step 3.2: Randomly generate s groups of different parameters for [Y1, Y2, Y3... Y v , that is, s groups of [c1, c2, c3... c v ;
[0081] Step 3.3: For a certain group of [c1, c2, c3... c v , calculate S according to the formula
[0082] S fusion = c1 * Y1 + c2 * Y2 + c3 * Y3 +... + c v * Y v
[0083] Calculate S fusion ;
[0084] Step 3.4: According to the given pruning ratio, find the control gates λ fusion with relatively low values in S k , remove the convolutional channels corresponding to these control gate values, and obtain the pruned sub-model according to this S fusion . Calculate the inference accuracy of this sub-model on the training set. And calculate the fitness = 1.0 / (inference accuracy of the complete model - inference accuracy of this sub-model);
[0085] Step 3.5: Execute Steps 3.3 and 3.4 for the s groups of different parameters in Step 3.2, and obtain the fitness fitness for each group of parameters respectively;
[0086] Step 3.6: According to the fitness of each group in the s groups of parameters, select two groups of parameters with the highest fitness for crossover and mutation to obtain the individuals of the next generation:
[0087] For the two selected groups of parameters, convert them into two binary format strings, and divide the crossover point at a randomly selected position. Then, exchange the binary string segments after the crossover point to form two offspring individuals, that is, two new groups of parameters. This is the crossover; after obtaining the offspring individuals, we perform mutation operations on some of the binary bits, that is, flip the values at randomly selected bits from 0 to 1 or from 1 to 0. This is the mutation.
[0088] Generate s groups of new parameters as the new generation population by repeating the crossover and mutation operations, and return to Step 3.3 until the required number of iterations is reached.
[0089] Step 3.7: Execute Steps 3.2 - 3.6 for T times, and obtain a set of parameters with the highest fitness each time. Then select a set of parameters with the highest fitness from these T times.
[0090] Step 3.8: Calculate S using the parameters obtained in Step 3.7 fusion , and then according to the given pruning ratio, find the control gates λ with relatively low values of the corresponding ratio in the S fusion calculated in this step, and remove the convolution channels corresponding to these control gate values to obtain the pruned sub - model according to this S k , and this sub - model is the customized sub - model we want. fusion The above process is a general method completed in the cloud. Through the above process, we can componentize a large - scale convolutional neural network, which is convenient for us to customize sub - models according to the needs of edge devices.
[0091] To verify the effectiveness of the above, we use the VGG - 16 model to verify our method on CIFAR - 10 and CIFAR - 100. The two CIFAR datasets are respectively composed of natural images of 10 classes and 100 classes, with a resolution of 32×32. The number of training and test sets for both datasets is 50000 and 10000 respectively. We perform normalization and simple data augmentation on the input data.
[0092] In the class - level control gate training stage, we first input all samples of the same class to train the control gates. In each training, we re - initialize the control gate A and set all values to 1 to activate all convolution channels. During the training of the control gates, set the number of iterations T = 100 and the learning rate to 0.1. In addition, to make the channels activated by different class inputs as sparse and non - overlapping as possible, we adopt a mid - training decay strategy to adjust the hyperparameter γ. Specifically, at the 0.5 stage of the training cycle, γ is decayed by 100 times. Such a setting is because the control gate values have reached a high degree of sparsity before the 0.5 stage. At this time, decaying γ helps the model accuracy to quickly recover. Adopting such a strategy helps us obtain more sparse channel control gate values, that is, more of them are close to 0, which helps to clearly show the channel distribution structure contributing to the final inference, and this is helpful for our subsequent fusion of different class - level control gates.
[0093] When fusing v different class - level control gates, we adopt processes such as natural selection, crossover, and mutation in the principle of biological evolution to find the best coefficient combination [c1, c2, c3…c
[0094] When fusing v different class - level control gates, we adopt processes such as natural selection, crossover, and mutation in the principle of biological evolution to find the best coefficient combination [c1, c2, c3…c v. We represent the coefficients in binary encoding, and each coefficient c i (i ∈ [0, v])'s occupied number of bits is significant for the performance of the algorithm and the representation of the search space. More occupied bits can provide higher precision, but also increase the length of the binary encoding and the computational complexity. Considering the accuracy requirements of the algorithm and the search space, we set the occupied number of bits for each coefficient c i to be 7. For the population size s and the number of iterations T in the algorithm, generally speaking, a larger population size can increase the coverage of the search space, but will increase the computational cost, while more iterations can provide more search opportunities, but will also increase the algorithm execution time. After multiple experiments and evaluations, we found that when the number of iterations is 15, the optimal fitness of the population usually reaches convergence. For the setting of the population size s, we choose to adjust the population size at different stages of the algorithm. In the initial stage of the algorithm, we choose a larger population size to increase diversity and cover more solution spaces, while in the latter half of the algorithm, we choose to reduce the population size and search more concentratedly for the local optimal solutions in the solution space to accelerate convergence. Therefore, we set the number of iterations T to 20. In the first half of the iteration process, the population size is 20, and in the latter half of the iteration process, the population size is reduced to 10. To cooperate with the setting of the population size, we adopted an adaptive mutation rate strategy. In the initial stage of the algorithm, a higher mutation rate is selected to search the solution space more widely. In the latter half, as the algorithm progresses, the mutation rate is gradually reduced to search more concentratedly for the local optimal solutions in the solution space.
[0095] We selected any 3 classes and 5-class combinations from the CIFAR-10 dataset to customize sub-models (where pr = 0.85 when the number of target classes is 3, and pr = 0.75 when the number of target classes is 3) to verify our method. And we customized test sets for different class combinations respectively, and the test sets only contain data with target class labels. Figure 2 and Figure 3 show the performance of the corresponding sub-models under a part of the class combinations. The dots are the inference accuracies of the original complete model on the customized test sets.
[0096] When fusing the control gates of different categories of channels, similar to merging the control gates of all images in the same category, we attempt to sum and merge the control gates of different categories layer by layer. However, as shown by the cross-marked points in the figure, the control gates fused in this way prune the original model to obtain a customized model, and the inference performance is poor. Subsequently, we fine-tuned these models using a small amount of training data. As shown by the upper triangular marked points, the performance of the models has been greatly improved, just as previously stated: our method is to explore a good pruning structure, which does not mean that retaining the parameters of the original neural network can achieve good results. Retraining can greatly improve the performance of our models. However, there are still some sub-models with unsatisfactory inference performance under certain category combinations. We believe the reason for this phenomenon is that in a CNN, each channel of the convolutional kernel represents different feature information of the image. The image features of the same category are relatively similar, so the activated channels are also relatively concentrated. On the contrary, the image features of different categories are quite different, and directly adding and merging the channel gate values of different categories cannot effectively extract more important channel information.
[0097] We also used our method to fuse the control gates under these category combinations and pruned to obtain sub-models. The performance of the sub-models on the test set is shown by the square marked points. In most category combinations, the performance can already approach that of the original neural network. It is not excluded that there is still a performance gap in some category combinations. However, considering the significant compression of the scale of the original model, such performance differences are inevitable. It can be found that the fusion results obtained by our method have performance improvements in different category combinations compared to the results obtained by directly adding the channel control gates, proving that the algorithm is effective in solving this problem.
[0098] When pruning the original model, the pruning rate has a huge impact on the performance of the final sub-model. Under the method we proposed, we found that the more categories the sub-model needs to perceive, the more the accuracy drops under the same pruning ratio. We verified this conclusion on the CIFAR-10 dataset, as Figure 4As shown in the figure. To keep the accuracy degradation of the customized sub-model within 5%, we obtained a curve (green curve) showing the relationship between the model pruning ratio and the number of categories that the sub-model needs to perceive. The accuracy calculation method here is as follows: for different numbers of target categories, 30 different combinations are taken. For these 30 combinations, sub-models are customized respectively at the pruning rate of pr. When the average accuracy degradation of these sub-models compared with the original model is within 5% at the pruning rate pr, the pruning rate pr at this time meets the requirements and is used as a point on our curve. The part under this curve can make the accuracy of the customized sub-model degrade by less than 5% compared with the accuracy of the original neural network. Since the accuracy requirements are different in different scenarios, a curve (blue curve) that can keep the accuracy degradation of the sub-model within 10% is also obtained. It can be found that a looser accuracy requirement allows for a higher degree of model pruning. When specifically implementing this method, the corresponding pruning ratio can be used to prune the model during the training process so that the performance of the sub-model meets the corresponding task requirements.
[0099] To further verify the performance of the proposed method, we compared the relationship between the accuracy loss and the pruning ratio of VGGNet-16 on the CIFAR-10 and CIFAR 100 datasets respectively. We conducted experiments with the number of target categories being 5 and 8 respectively. For a given number of target categories, 20 different category combinations were selected. Sub-models were generated at different pruning ratios and tested on the customized test set. The change curve between the accuracy loss and the pruning ratio is as Figure 5 shown. At a lower pruning ratio, the accuracy of the sub-model generated by our method is better than that of the original model. As the pruning ratio increases, more parameters are pruned and the accuracy of the sub-model becomes lower. This proves that our method prunes the redundant convolutional layers in the CNN, retains the convolutional layers that contribute more to the final inference, and makes the performance of the sub-model even better than the original model after fine-tuning. This means that when applying this method, we can trade off a higher sub-model accuracy by reducing the degree of model compression. On the premise that the model scale is reduced by at least 50% compared with the original model, a higher model accuracy is achieved. In addition, it can be seen from this figure that the increase in the number of target categories leads to an earlier decrease in the accuracy of the sub-model, which proves that there is a relationship between the complexity of the task and the model parameters required by the sub-model. The sub-model needs more model parameters to perceive more categories, and the same ratio of pruning will cause a greater accuracy loss for the model that perceives more categories. When applying this method, the more categories the sub-model needs to perceive, the more appropriate pruning rate we need to reduce to ensure the performance of the sub-model.
[0100] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.
Claims
1. A method for componentizing a convolutional neural network based on structured pruning in an edge-cloud collaboration scenario, characterized in that: It includes the following steps: Step 1: Obtain the image-level channel control gate A; Step 1.1: Use the training set to pre-train the over-parameterized complete model N_full, and the depth of the model is K; Step 1.2: Initialize control gate λ k ; Step 1.3: Obtain an image x, and use the model N_full to predict the image to obtain the inference result f θ (x); and use the model after adding the control gate λ k to predict the image to obtain the inference result f θ (x, A); the model after adding the control gate λ k refers to the model obtained by multiplying the output of the k-th layer of the model channel by channel with the control gate λ k in the complete model; the dimension of the control gate λ k is the number of convolutional channels of the k-th layer of the model, A = {λ1, λ2…λ K}; Step 1.4: With Calculate the loss function, where γ is a hyperparameter, and the meaning of the crossEntropy() function is the cross-entropy loss function, and K is the depth of the model; And with Perform gradient update on A, ensuring that each value of λ k is non - negative during the update process; after the update, calculate the predicted class label of the model after adding the control gate as j = argmax N_full(x, A); if i = j at this time, the channel gate A has converged in this iteration; k Each value is non - negative; after the update, calculate the predicted class label of the model after adding the control gate as j = argmax N_full(x, A); if i = j at this time, the channel gate A has converged in this iteration; Step 1.5: Repeat Step 1.4 for T times to complete the training of control gate A; at this time, all λ k are combined layer by layer to obtain the image-level channel control gate A; Step 2: Use the image-level channel control gate A of different images of the same category to establish the class-level control gate Y; Step 2.1: For each of the n images [x1, x2, x3…x n in the same target category, perform Step 1 one by one to obtain the image-level control gates [A x1 , A x2 , A x3 …A xn ; Step 2.2: Add up the image-level channel control gates corresponding to each of the n images to obtain the class-level control gate for the target class Step 3: Customize the sub-model: Step 3.1: According to the customization requirements, use Step 2 to obtain v class-level control gates for different target categories [Y1, Y2, Y3…Y v ; Step 3.2: Randomly generate s sets of different parameters [c1, c2, c3... c v for [Y1, Y2, Y3... Y v ; Step 3.3: For a certain group [c1, c2, c3…c v , according to the formula S fusion = c1 * Y1 + c2 * Y2 + c3 * Y3 + … + c v * Y v Calculate S fusion ; Step 3.4: According to the given pruning ratio, find the control gates λ with relatively low values corresponding to the ratio in S fusion , and remove the convolutional channels corresponding to these control gate values to obtain the pruned sub-model according to this S k ; fusion Calculate the inference accuracy of the sub-model on the training set, and calculate the fitness = 1.0 / (complete model inference accuracy - sub-model inference accuracy); Step 3.5: Execute Steps 3.3 and 3.4 for s groups of different parameters in Step 3.2, and obtain the fitness fitness of each group of parameters respectively; Step 3.6: According to the fitness of each group in the s groups of parameters, select two groups of parameters with the highest fitness for crossover and mutation to obtain the individuals of the next generation, and return to Step 3.3 until the required number of iterations is reached; Step 3.7: Execute Steps 3.2 - 3.6 for T times, and obtain a group of parameters with the highest fitness each time; select a group of parameters with the highest fitness from the T times of parameters; Step 3.8: Calculate S using the parameters obtained in Step 3.7 fusion , and then, according to the given pruning ratio, find the control gates λ with relatively low values corresponding to the ratio in the S fusion calculated in this step k , remove the convolutional channels corresponding to these control gate values, and obtain the sub-model pruned according to this S fusion The pruned sub-model is the customized sub-model.
2. The method for componentizing a convolutional neural network based on structured pruning in an edge-cloud collaboration scenario according to claim 1, wherein: All possible image categories are included in the training set.
3. The componentization method of a convolutional neural network based on structured pruning in an edge-cloud collaboration scenario according to claim 1, characterized in that: The control gate λ satisfies the conditions: the value of λ is non-negative, and λ is as sparse as possible.
4. The method for componentizing a convolutional neural network based on structured pruning in an edge-cloud collaboration scenario according to claim 1, characterized in that: Perform Step 1.1 on the cloud side, and use the training set to pre-train the over-parameterized complete model N_full.
5. A computer-readable storage medium stores computer-executable instructions, and the instructions are used to implement the method according to any one of claims 1 to 4 when executed.
6. A computer system, comprising: One or more processors, the computer-readable storage medium according to claim 5, are used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Efficient deep convolutional neural network pruning method
CN113610227A
Neural network channel pruning method based on improved MetaPruning
CN116306880A