Student network acquisition method and system based on pruning and distillation fusion

By combining pruning and distillation techniques, decoupling knowledge distillation and adaptive pruning rate methods, the problem that students' network size cannot be quantified is solved, and the efficient compression and precise compression effect of deep convolutional neural networks is achieved.

CN120146100APending Publication Date: 2025-06-13CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510150875.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When compressing deep convolutional neural network models in the prior art, the size of the student model cannot be quantified and difficult to select, and there are problems with efficiency and accuracy in practical applications of pruning and distillation technologies.

Method used

Using a student network acquisition method based on pruning and distillation fusion, the initial lightweight convolutional neural network is trained by decoupling knowledge distillation, the target pruning rate and filter importance group are obtained, multiple rounds of pruning and distillation compression training are performed, and the lightweight convolutional neural network is updated until the preset training threshold is reached.

Benefits of technology

It realizes the streamlined structure of the student network, which can effectively compress the size of the deep convolutional neural network while maintaining the high accuracy of the model, and is suitable for resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146100A_ABST
    Figure CN120146100A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of convolutional neural networks, in particular to a student network acquisition method and system based on pruning and distillation fusion. Comprising the following steps: performing pruning distillation compression training on an initial lightweight convolutional neural network to obtain a student network; the pruning distillation compression training comprises the following steps: training an initial lightweight convolutional neural network through decoupling knowledge distillation to obtain a first lightweight convolutional neural network; obtaining a target pruning rate and a filter importance group; pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network; and updating the initial lightweight convolutional neural network based on the second lightweight convolutional neural network. According to the method, the training progress is accelerated by utilizing the characteristics of knowledge distillation, the problem of low convergence rate in the pruning process is solved, and meanwhile, the problem that the size of a student network model is difficult to determine in knowledge distillation is solved based on the target pruning rate and filter importance pruning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of convolutional neural networks, and in particular to a method and system for acquiring a student network based on pruning, distillation and fusion. Background Art

[0002] At present, in the field of artificial intelligence, the development of convolutional neural networks (CNN) is very rapid. Since the 1990s, convolutional neural networks have attracted widespread attention due to their powerful capabilities in image recognition and processing. With the advancement of technology, the structure and performance of convolutional neural network models have been significantly improved, and at the same time, the scale of their models has become larger. This not only increases the storage memory space occupied by the neural network model and the consumption of running computing resources, increasing its cost of use, but also makes the neural network model unsuitable for environments with limited computing resources or storage resources.

[0003] In order to solve the problem of the neural network model being too large, the field currently uses network compression technology to optimize the neural network model. Network compression technology aims to reduce the number of parameters and computational complexity of the neural network model while maintaining the original performance of the model as much as possible. Existing network compression technologies include pruning technology and distillation technology. Among them, pruning technology compresses the model size by removing part of the structure or parameters of the model, which has certain limitations. For example, unstructured pruning is difficult to accelerate model reasoning, the balance between pruning ratio and pruning accuracy is difficult to quantify and grasp, and the iterative pruning method of pruning technology is time-consuming, which is difficult to apply to large-scale and complex teacher models in practice. Distillation technology achieves model compression by transferring the knowledge of a large teacher model to a small student model. In actual application, the size of its student model cannot be quantified and designed, and it is difficult to select. If the student model is designed too large, the model compression effect is poor, and the large-scale and complex teacher model in the actual scene cannot be reduced. If the student model is designed too small, the model accuracy is difficult to maintain the original performance, and the complex network model scene in the teacher model cannot be accurately reproduced. Summary of the invention

[0004] The purpose of the present invention is to provide a student network acquisition method and system based on pruning and distillation fusion, which combines pruning technology and distillation technology to compress the neural network model and solve the above technical problems.

[0005] To achieve the above object, a first aspect of the present invention provides a method for obtaining a student network based on pruning and distillation fusion, including the following steps: obtaining an initial lightweight convolutional neural network based on a preset deep convolutional neural network; performing a round of pruning and distillation compression training on the initial lightweight convolutional neural network, and then obtaining a first target lightweight convolutional neural network based on the result of the pruning and distillation compression training as the model compression result. Wherein, the pruning and distillation compression training includes: training the initial lightweight convolutional neural network through decoupled knowledge distillation to obtain a first lightweight convolutional neural network; obtaining a target pruning rate and a filter importance group; pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network; updating the initial lightweight convolutional neural network based on the second lightweight convolutional neural network as the result of this round of pruning and distillation compression training.

[0006] The above method for obtaining a student network based on pruning and distillation fusion uses a deep convolutional neural network with a large scale and complex structure in the actual scenario as the teacher model. This method combines pruning technology and distillation technology. First, the initial lightweight convolutional neural network is trained through decoupled knowledge distillation to obtain a first lightweight convolutional neural network, and then the first lightweight convolutional neural network is pruned. The characteristics of decoupled knowledge distillation are used to accelerate the subsequent pruning speed and solve the problem of slow convergence speed in the pruning process. At the same time, pruning the first lightweight convolutional neural network through the target pruning rate and the filter importance group can adaptively adjust its model size according to the characteristics of the first lightweight convolutional neural network to obtain a second lightweight convolutional neural network. By updating the lightweight convolutional neural network through the above process, it is used as the student network of the deep convolutional neural network, which solves the problem that the size of the student network in the distillation process cannot be quantitatively designed and is difficult to select. On this basis, the student network obtained by this method can represent the original large-scale and complex deep convolutional neural network with a concise model structure, improving the effect of the lightweight convolutional neural network in simplifying the deep convolutional neural network. At the same time, the student network obtained by this method combines distillation technology and pruning technology to improve the model accuracy and enhance the accuracy of the lightweight convolutional neural network in reproducing the corresponding real complex scenario of the deep convolutional neural network.

[0007] In a possible implementation manner, the method for obtaining a student network based on pruning and distillation fusion further includes: performing several rounds of pruning and distillation compression training on the initial lightweight convolutional neural network until the number of training rounds reaches a preset training threshold; obtaining the target lightweight convolutional neural network based on the result of the last round of pruning and distillation compression training.

[0008] In this implementation, the lightweight convolutional neural network is subjected to multiple rounds of distillation training and pruning training, so that the structure of the obtained target lightweight convolutional neural network is further simplified, its storage space occupancy rate is further reduced, and the size of the obtained target lightweight convolutional neural network is further reduced compared with that of the deep convolutional neural network, that is, the effect of the student network simplifying the teacher model is improved; at the same time, the lightweight convolutional neural network learns the features of the deep convolutional neural network through distillation training in multiple rounds of training, so that the finally obtained target lightweight convolutional neural network can maintain the accuracy of the deep convolutional neural network and improve the performance of the student network.

[0009] The second aspect of the present invention provides a student network acquisition system based on pruning and distillation fusion, including a main control module, a model extraction module, and a compression training module, wherein: the model extraction module is used to obtain an initial lightweight convolutional neural network based on a preset deep convolutional neural network; the main control module is used to receive the initial lightweight convolutional neural network, so as to control the compression training module to perform a round of pruning and distillation compression training on the initial lightweight convolutional neural network, and then obtain a target lightweight convolutional neural network based on the result of the pruning and distillation compression training. The compression training module is used to perform the pruning and distillation compression training, including: training the initial lightweight convolutional neural network through decoupled knowledge distillation to obtain a first lightweight convolutional neural network; obtaining a target pruning rate and a filter importance group; pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network; updating the initial lightweight convolutional neural network based on the second lightweight convolutional neural network as the result of this round of pruning and distillation compression training.

[0010] The above-mentioned student network acquisition system based on pruning and distillation fusion uses the deep convolutional neural network with a large scale and complex structure in the actual scenario as the teacher model. Among them, the compression training module combines pruning technology and distillation technology. First, the initial lightweight convolutional neural network is trained by decoupled knowledge distillation to obtain the first lightweight convolutional neural network, and then the first lightweight convolutional neural network is pruned. Utilizing the characteristics of decoupled knowledge distillation speeds up the subsequent pruning speed and solves the problem of slow convergence speed during the pruning process. At the same time, the first lightweight convolutional neural network is pruned according to the target pruning rate and filter importance group, which can adaptively adjust its model size according to the characteristics of the first lightweight convolutional neural network to obtain the second lightweight convolutional neural network. By updating the lightweight convolutional neural network through the above process and using it as the student network of the deep convolutional neural network, the problem that the size of the student network cannot be quantitatively designed and is difficult to select during the distillation process is solved. On this basis, the student network obtained by this method can represent the originally large-scale and complex deep convolutional neural network with a streamlined model structure, that is, it improves the effect of the student network in simplifying the teacher model. At the same time, the student network obtained by this method combines distillation technology and pruning technology to improve the model accuracy and enhance the accuracy of the lightweight convolutional neural network in reproducing the corresponding real complex scenario of the deep convolutional neural network.

[0011] In a possible implementation manner, the main control module is further configured to control the compression training module to perform several rounds of pruning and distillation compression training on the initial lightweight convolutional neural network until the number of training rounds reaches a preset training threshold, and then obtain the target lightweight convolutional neural network based on the result of the last round of the pruning and distillation compression training.

[0012] In this implementation manner, multiple rounds of distillation training and pruning training are performed on the lightweight convolutional neural network, so that the structure of the obtained target lightweight convolutional neural network is further simplified, its storage space occupancy rate is further reduced, and the size of the obtained target lightweight convolutional neural network is further reduced compared with the deep convolutional neural network, that is, it improves the effect of the student network in simplifying the teacher model. At the same time, the lightweight convolutional neural network learns the features of the deep convolutional neural network through distillation training during multiple rounds of training, so that the finally obtained target lightweight convolutional neural network can maintain the accuracy of the deep convolutional neural network and improve the performance of the student network.

[0013] In a possible implementation manner, obtaining the target pruning rate and filter importance group includes: the compression training module obtains the first model accuracy rate based on the first lightweight convolutional neural network; if the compression training module determines that the first model accuracy rate is less than or equal to a preset accuracy rate threshold, then obtain an initial pruning rate, and then use the initial pruning rate as the target pruning rate.

[0014] In this implementation, since the lightweight convolutional neural network has weak expressive ability in the initial training stage of learning the knowledge of the deep convolutional network, its model parameters change rapidly and cannot accurately reflect the ability of the deep convolutional network; that is, the prediction accuracy of the lightweight convolutional neural network is relatively low in the initial training stage of learning the knowledge of the deep convolutional network. Therefore, when the above scheme determines that the accuracy of the first lightweight convolutional neural network is less than the accuracy threshold, it uses the initial pruning rate with a smaller pruning rate as the target pruning rate to prune the first lightweight convolutional neural network, ensuring that its filters will not be lost too much in the initial training stage, improving the accuracy of the lightweight convolutional neural network in the subsequent training process, and ultimately improving the accuracy of the student network in learning the knowledge of the teacher model.

[0015] In a possible implementation, if the compression training module determines that the accuracy of the first model is greater than the accuracy threshold, it obtains an adaptive pruning rate, and then uses the adaptive pruning rate as the target pruning rate; after pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain the second lightweight convolutional neural network, it further includes: the compression training module obtains the accuracy of the second model based on the second lightweight convolutional neural network, and then calculates the accuracy loss based on the accuracy of the first model and the accuracy of the second model; the compression training module updates the adaptive pruning rate based on the accuracy loss; if the compression training module determines that the accuracy loss is greater than or equal to the preset loss upper limit, it updates the initial lightweight convolutional neural network based on the first lightweight convolutional neural network as the result of this round of pruning and distillation compression training; if the compression training module determines that the accuracy loss is less than the loss upper limit, it updates the initial lightweight convolutional neural network based on the second lightweight convolutional neural network as the result of this round of pruning and distillation compression training.

[0016] In this implementation, when the accuracy of the first lightweight convolutional neural network is greater than the accuracy threshold, it is considered that its model parameters have tended to be stable, that is, the accuracy of the first lightweight convolutional neural network has tended to be stable. At this time, the adaptive pruning rate with a larger pruning rate is used as the target pruning rate to prune the first lightweight convolutional neural network, so that its unimportant filters can be removed during the pruning process, improving the efficiency of compressing the size of the neural network model, and ultimately making the size of the obtained student network structure more concise.

[0017] Further, the above solution quantifies and evaluates the accuracy loss of the lightweight convolutional neural network in this round of pruning by comparing the accuracy of the first model before pruning and the accuracy of the second model after pruning. When the accuracy loss is greater than the preset loss upper limit, it is considered that the accuracy loss in this round of pruning is too high and the pruning rate is too high, which does not meet the requirements for the lightweight convolutional neural network to learn the knowledge of the deep convolutional neural network. Therefore, it is considered that this round of pruning fails, and the initial lightweight convolutional neural network is updated based on the first lightweight convolutional neural network before this round of pruning to ensure that the result of this round of pruning does not affect the entire training process. Correspondingly, when the accuracy loss is less than the preset loss upper limit, it is considered that this round of pruning is effective, and thus the initial lightweight convolutional neural network is updated based on the second lightweight convolutional neural network after this round of pruning.

[0018] In addition, the adaptive pruning rate is updated based on the accuracy loss. When the accuracy loss is too high, the adaptive pruning rate is decreased to prevent the accuracy of the trained student network from being too low; when the accuracy loss is too low, the adaptive pruning rate is increased to improve the efficiency of compressing the size of the lightweight convolutional neural network model. Ultimately, this solution can balance the training accuracy of the student network and the model compression efficiency.

[0019] In a possible implementation manner, when the compression training module determines that the accuracy loss is greater than or equal to the loss upper limit, the updating of the adaptive pruning rate based on the accuracy loss includes: the compression training module calculates a decay pruning rate based on a preset decay factor and the adaptive pruning rate, and then updates the adaptive pruning rate based on the decay pruning rate. The formula is as shown in Equation 1 below:

[0020] p2 new =μ 1 *p2 old (Equation 1)

[0021] where μ 1 is the decay factor, p2 old is the adaptive pruning rate, and p2 new is the decay pruning rate.

[0022] In this implementation manner, when the accuracy loss is greater than or equal to the loss upper limit, it is considered that the accuracy loss is too high and the adaptive pruning rate needs to be decreased. Among them, the decay factor is a parameter between 0 and 1, which determines the decay amplitude of the adaptive pruning rate. Specifically, the closer the decay factor is to 0, the higher the decay amplitude of the adaptive pruning rate.

[0023] In a possible implementation, when the compression training module determines that the accuracy loss is less than the loss upper limit, updating the adaptive pruning rate based on the accuracy loss includes: the compression training module calculates a growth pruning rate based on the loss upper limit and the adaptive pruning rate, and then updates the adaptive pruning rate based on the growth pruning rate. The formula is shown as formula 2 below:

[0024]

[0025] where μ 2 is a preset growth factor, Δ max is the loss upper limit, Δ min is a preset loss lower limit, Δ is the accuracy loss, p2 old is the adaptive pruning rate, p2 new is the growth pruning rate.

[0026] In this implementation, when the accuracy loss is less than the loss upper limit, it is considered that the accuracy loss is low and the adaptive pruning rate needs to be increased. Specifically, the increase amplitude of the adaptive pruning rate is determined according to the loss upper limit, the loss lower limit, and the growth factor. Among them, the growth factor is a parameter greater than 1. The closer the growth factor is to 1, the lower the increase amplitude of the adaptive pruning rate.

[0027] In a possible implementation, training the initial lightweight convolutional neural network through decoupled knowledge distillation to obtain a first lightweight convolutional neural network includes: the model extraction module obtains target class knowledge and non-target class knowledge based on the deep convolutional neural network, and then constructs a loss function for the initial lightweight convolutional neural network based on the target class knowledge and the non-target class knowledge; the compression training module receives the loss function and then trains the initial lightweight convolutional neural network based on the loss function to obtain a first lightweight convolutional neural network.

[0028] In this implementation, the output of the deep neural network is divided into target class knowledge and non-target class knowledge. Among them, the target class knowledge is the output information of the categories that the lightweight convolutional neural network needs to focus on learning, and the non-target class knowledge is the output information that the lightweight convolutional neural network does not need to pay special attention to. By distinguishing the target class knowledge and the non-target class knowledge to construct the loss function, the lightweight convolutional neural network can learn the knowledge of the deep convolutional neural network more pertinently, and improve the accuracy of the first lightweight convolutional neural network obtained through decoupled knowledge distillation training.

[0029] In a possible implementation, the obtaining of the target pruning rate and the filter importance group includes: the compression training module obtains a filter set based on the first lightweight convolutional neural network; for any filter in the filter set, the compression training module obtains the filter norm and the filter independence based on the filter; the compression training module obtains the filter importance corresponding to the filter based on the filter norm and the filter independence, and then takes all the filter importances as the filter importance group.

[0030] In this implementation, the filter norm is determined by the height, width, number of channels of the filter convolution kernel, and the weight value at its position in the channel. The filter independence is calculated by performing the DBSCAN clustering algorithm on the filters after flattening and normalizing all the filters. Finally, the filter importance is comprehensively calculated based on the filter norm and the filter independence. Specifically, for any filter, the larger its filter norm and the stronger its filter independence, the more important the filter is considered. Brief Description of the Drawings

[0031] Figure 1 is a schematic flowchart of a method for obtaining a student network based on pruning and distillation fusion provided by an embodiment of the present invention;

[0032] Figure 2 is a schematic structural diagram of a system for obtaining a student network based on pruning and distillation fusion provided by an embodiment of the present invention;

[0033] Figure 3 is a schematic flowchart of another method for obtaining a student network based on pruning and distillation fusion provided by an embodiment of the present invention;

[0034] Wherein: 110, main control module; 120, model extraction module; 130, compression training module. Detailed Embodiments

[0035] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0036] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed descriptions of the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification, claims and above-mentioned drawings of this application are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification, claims or above-mentioned drawings of this application are used to distinguish different objects and are not used to describe a specific order.

[0037] It should be understood that although each step in the flowchart of the drawings is shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0038] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0039] At present, in order to solve the problem that the deep convolutional neural network model is too large and the structure is too complex, the field currently uses network compression technology to optimize the neural network model. Network compression technology aims to reduce the number of parameters and computational complexity of the neural network model while maintaining the original performance of the model as much as possible. Existing network compression technologies include pruning technology and distillation technology. Among them, pruning technology compresses the model size by removing part of the structure or parameters of the model, and there are certain limitations. For example, unstructured pruning is difficult to accelerate model reasoning, the balance between pruning ratio and pruning accuracy is difficult to quantify and grasp, and the iterative pruning method of pruning technology is time-consuming, and it is difficult to apply it to large-scale and complex teacher models in practice. Distillation technology realizes model compression by transferring the knowledge of large teacher models to small student models. In actual application, the size of its student model cannot be quantified and designed, and it is difficult to select. If the student model is designed too large, the model compression effect is poor, and the large-scale and complex teacher model in the actual scene cannot be reduced. If the student model is designed too small, the model accuracy is difficult to maintain the original performance, and the complex network model scene in the teacher model cannot be accurately reproduced.

[0040] In order to solve the above technical problems, see Figure 1 , an embodiment of the present invention provides a method for obtaining a student network based on pruning, distillation and fusion, comprising the following steps:

[0041] S110, obtaining an initial lightweight convolutional neural network based on a preset deep convolutional neural network.

[0042] The present invention uses a deep convolutional neural network (DCNN) as a teacher model. In the field of artificial intelligence technology, a deep convolutional neural network specifically refers to a deep learning model specifically used to process data with a grid structure, and specifically refers to a complex neural network with a large number of network layers and a relatively large scale, including a neural network model that partially processes image classification and image segmentation. Based on the deep convolutional neural network that needs to be compressed, an adaptive initial lightweight convolutional neural network is selected, wherein a lightweight convolutional neural network (LCNN) specifically refers to a neural network model with fewer layers and parameters, such as MobileNet, ShuffleNet and SqueezeNet; through the above-mentioned student network acquisition method based on pruning, distillation and fusion, the initial lightweight convolutional neural network can fully learn the knowledge of the deep convolutional neural network through distillation, and further reduce the number of layers and parameters through pruning, and generate a target lightweight convolutional neural network as a student network.

[0043] Specifically, the present invention uses VGG16 as the preset deep convolutional neural network of the teacher model. VGG16 has 13 convolutional layers and 3 fully connected layers. The dataset used in the present invention records the operation processes of electrical devices in three different aging states. Specifically, this dataset samples the current, voltage, and temperature of the device at a certain frequency to obtain the electrical characteristics of electrical devices in three different aging states during the corresponding operation processes. In terms of dataset construction, the present invention traverses the dataset with a window of length 24, inputs the current, voltage, and temperature as the eigenvalue of VGG16, and then obtains the corresponding aging state of the electrical device as the output; that is, input a time series data of 24*3 and output the corresponding aging state label. Since this dataset has recorded the operation processes of electrical devices in three different aging states, the trained VGG16 is a three-classification model, which is used to judge which of the three aging states the corresponding device belongs to according to the input electrical characteristic values of the device, and then judge the aging situation of the electrical device. After the present invention fully trains VGG16 using this dataset, a deep convolutional neural network with strong expressive ability is obtained as the teacher model of this embodiment.

[0044] Define the model size of the trained deep convolutional neural network as M t , and define the scaling factor λ∈[0,1] such that the model size of the initial lightweight convolutional neural network is M s =λ*M t , and use the initial lightweight convolutional neural network as the student network of this embodiment. Specifically, based on the preset scaling factor λ, the first 10 convolutional layers of the preset deep convolutional neural network are retained as the convolutional layers of the initial lightweight convolutional neural network, and two fully connected layers of the preset deep convolutional neural network are retained as the fully connected layers of the initial convolutional neural network. Among them, the obtained initial lightweight convolutional neural network includes two convolutional layers with 64 filters, two convolutional layers with 128 filters, three convolutional layers with 256 filters, and three convolutional layers with 512 filters. Finally, based on the three aging states of the electrical devices in the dataset, the output size of the fully connected layer of the initial convolutional neural network is set to 3, so that the initial convolutional neural network can judge the aging state corresponding to a specific electrical device according to the input feature data.

[0045] S120. Perform a round of pruning and distillation compression training on the initial lightweight convolutional neural network, and then obtain the first target lightweight convolutional neural network based on the results of the pruning and distillation compression training as the model compression result.

[0046] Among them, the pruning distillation compression training includes: S121. Training the initial lightweight convolutional neural network through decoupled knowledge distillation to obtain a first lightweight convolutional neural network; S122. Obtaining a target pruning rate and a filter importance group; S123. Pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network; S124. Updating the initial lightweight convolutional neural network based on the second lightweight convolutional neural network as the result of this round of pruning distillation compression training.

[0047] The student network obtained through this embodiment solves the problem that the size of the student network cannot be quantitatively designed and is difficult to select during the distillation process. On this basis, the student network obtained through this method can represent the originally large-scale and complex-structured VGG16 deep convolutional neural network with a streamlined model structure, improving the effect of the lightweight convolutional neural network in simplifying the memory occupancy and structural complexity of the deep convolutional neural network. Based on reducing the convolutional layer and fully connected layer of the model, the obtained first target lightweight convolutional neural network can maintain the same judgment accuracy as the VGG16 teacher model, judge the aging state corresponding to a specific electrical device according to the input current, voltage, and temperature characteristic data, and further make the first lightweight convolutional neural network obtained in this embodiment more suitable for being deployed in the resource-constrained local environment of electrical devices to monitor the aging condition of local electrical devices in real time and provide a more accurate analysis and judgment basis for the maintenance of local electrical devices.

[0048] In a possible embodiment, the method for obtaining a student network based on pruning distillation fusion further includes: performing several rounds of pruning distillation compression training on the initial lightweight convolutional neural network until the number of training rounds reaches a preset training threshold; obtaining the target lightweight convolutional neural network based on the result of the last round of pruning distillation compression training.

[0049] See Figure 2, the second aspect of the present invention provides a student network acquisition system based on pruning and distillation fusion, including a main control module 110, a model extraction module 120, and a compression training module 130, where: the model extraction module 120 is used to obtain an initial lightweight convolutional neural network based on a preset deep convolutional neural network; the main control module 110 is used to receive the initial lightweight convolutional neural network, so as to control the compression training module 130 to perform a round of pruning and distillation compression training on the initial lightweight convolutional neural network, and then obtain a target lightweight convolutional neural network based on the result of the pruning and distillation compression training. The compression training module 130 is used to perform the pruning and distillation compression training, including: training the initial lightweight convolutional neural network through decoupled knowledge distillation to obtain a first lightweight convolutional neural network; obtaining a target pruning rate and a filter importance group; pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network; updating the initial lightweight convolutional neural network based on the second lightweight convolutional neural network as the result of this round of pruning and distillation compression training.

[0050] The above-mentioned student network acquisition system based on pruning and distillation fusion uses the large-scale and complex deep convolutional neural network in the actual scenario as the teacher model. Among them, the compression training module 130 combines the pruning technology and the distillation technology. First, the initial lightweight convolutional neural network is trained through decoupled knowledge distillation to obtain a first lightweight convolutional neural network, and then the first lightweight convolutional neural network is pruned. The characteristics of decoupled knowledge distillation are used to accelerate the subsequent pruning speed and solve the problem of slow convergence speed in the pruning process; at the same time, the first lightweight convolutional neural network is pruned through the target pruning rate and the filter importance group, and its model size can be adaptively adjusted according to the characteristics of the first lightweight convolutional neural network to obtain a second lightweight convolutional neural network. By updating the lightweight convolutional neural network through the above process, it is used as the student network of the deep convolutional neural network, which solves the problem that the size of the student network in the distillation process cannot be quantitatively designed and is difficult to select. On this basis, the student network obtained by this method can use a streamlined model structure to represent the originally large-scale and complex deep convolutional neural network, that is, it improves the effect of the student network simplifying the teacher model; at the same time, the student network obtained by this method combines the distillation technology and the pruning technology to improve the model accuracy and improve the accuracy of the lightweight convolutional neural network to reproduce the corresponding real complex scenario of the deep convolutional neural network.

[0051] In a possible embodiment, the main control module 110 is further configured to control the compression training module 130 to perform several rounds of pruning and distillation compression training on the initial lightweight convolutional neural network until the number of training rounds reaches a preset training threshold, and then obtain the target lightweight convolutional neural network based on the result of the last round of pruning and distillation compression training.

[0052] In this embodiment, the lightweight convolutional neural network is subjected to multiple rounds of distillation training and pruning training, so that the structure of the obtained target lightweight convolutional neural network is further simplified, the storage space occupancy rate thereof is further reduced, and the size of the obtained target lightweight convolutional neural network is further reduced compared with that of the deep convolutional neural network, that is, the effect of the student network simplifying the teacher model is improved; at the same time, the lightweight convolutional neural network learns the features of the deep convolutional neural network through distillation training in multiple rounds of training, so that the finally obtained target lightweight convolutional neural network can maintain the accuracy of the deep convolutional neural network and improve the performance of the student network.

[0053] In a possible embodiment, the obtaining of the target pruning rate and the filter importance group includes: the compression training module 130 obtains a first model accuracy rate based on the first lightweight convolutional neural network; if the compression training module 130 determines that the first model accuracy rate is less than or equal to a preset accuracy rate threshold, an initial pruning rate is obtained, and then the initial pruning rate is used as the target pruning rate.

[0054] In this embodiment, since the expression ability of the lightweight convolutional neural network is weak in the initial training stage of learning the knowledge of the deep convolutional network, its model parameters change rapidly and cannot accurately reflect the ability of the deep convolutional neural network; that is, the prediction accuracy rate of the lightweight convolutional neural network is low in the initial training stage of learning the knowledge of the deep convolutional network. Therefore, when the above solution determines that the accuracy rate of the first lightweight convolutional neural network is less than the accuracy rate threshold, the initial pruning rate with a smaller pruning rate is used as the target pruning rate to prune the first lightweight convolutional neural network, ensuring that its filters will not be lost too much in the initial training stage, improving the accuracy of the lightweight convolutional neural network in the subsequent training process, and finally improving the accuracy rate of the student network learning the knowledge of the teacher model.

[0055] In a possible embodiment, if the compression training module 130 determines that the accuracy of the first model is greater than the accuracy threshold, an adaptive pruning rate is obtained, and then the adaptive pruning rate is used as the target pruning rate; after pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group, the following is further included: the compression training module 130 obtains the accuracy of the second model based on the second lightweight convolutional neural network, and then calculates the accuracy loss based on the accuracy of the first model and the accuracy of the second model; the compression training module 130 updates the adaptive pruning rate based on the accuracy loss; if the compression training module 130 determines that the accuracy loss is greater than or equal to the preset loss upper limit, the initial lightweight convolutional neural network is updated based on the first lightweight convolutional neural network as the result of the pruning distillation compression training in this round; if the compression training module 130 determines that the accuracy loss is less than the loss upper limit, the initial lightweight convolutional neural network is updated based on the second lightweight convolutional neural network as the result of the pruning distillation compression training in this round.

[0056] In this embodiment, when the accuracy of the first lightweight convolutional neural network is greater than the accuracy threshold, it is considered that its model parameters have tended to be stable, that is, the accuracy of the first lightweight convolutional neural network has tended to be stable. At this time, the adaptive pruning rate with a larger pruning rate is used as the target pruning rate to prune the first lightweight convolutional neural network, so that its unimportant filters can be removed during the pruning process, improving the efficiency of compressing the size of the neural network model, and finally making the size of the obtained student network structure more concise.

[0057] Furthermore, the above solution quantifies and evaluates the accuracy loss of the lightweight convolutional neural network in this round of pruning by comparing the accuracy of the first model before pruning and the accuracy of the second model after pruning. When the accuracy loss is greater than the preset loss upper limit, it is considered that the accuracy loss in this round of pruning is too high and the pruning rate is too high, which does not meet the requirements for the lightweight convolutional neural network to learn the knowledge of the deep convolutional neural network. Therefore, it is considered that this round of pruning fails, and the initial lightweight convolutional neural network is updated based on the first lightweight convolutional neural network before this round of pruning to ensure that the result of this round of pruning does not affect the entire training process. Correspondingly, when the accuracy loss is less than the preset loss upper limit, it is considered that this round of pruning is effective, and therefore the initial lightweight convolutional neural network is updated based on the second lightweight convolutional neural network after this round of pruning.

[0058] In addition, the adaptive pruning rate is updated based on the accuracy loss. When the accuracy loss is too high, the adaptive pruning rate is reduced to prevent the accuracy of the student network obtained by training from being too low; when the accuracy loss is too low, the adaptive pruning rate is increased to improve the efficiency of compressing the size of the lightweight convolutional neural network model, and finally making this solution able to balance the training accuracy and model compression efficiency of the student network.

[0059] In a possible embodiment, when the compression training module 130 determines that the accuracy loss is greater than or equal to the loss upper limit, updating the adaptive pruning rate based on the accuracy loss includes: the compression training module 130 calculates a decay pruning rate based on a preset decay factor and the adaptive pruning rate, and then updates the adaptive pruning rate based on the decay pruning rate. The formula is as shown in Formula 1 below:

[0060] p2 new = μ 1 * p2 old (Formula 1)

[0061] Where μ 1 is the decay factor, p2 old is the adaptive pruning rate, and p2 new is the decay pruning rate.

[0062] In this embodiment, when the accuracy loss is greater than or equal to the loss upper limit, it is considered that the accuracy loss is too high and the adaptive pruning rate needs to be reduced. Among them, the decay factor is a parameter between 0 and 1, which determines the decay amplitude of the adaptive pruning rate. Specifically, the closer the decay factor is to 0, the higher the decay amplitude of the adaptive pruning rate.

[0063] In a possible embodiment, when the compression training module 130 determines that the accuracy loss is less than the loss upper limit, updating the adaptive pruning rate based on the accuracy loss includes: the compression training module 130 calculates an increase pruning rate based on the loss upper limit and the adaptive pruning rate, and then updates the adaptive pruning rate based on the increase pruning rate. The formula is expressed as shown in Formula 2 below:

[0064]

[0065] Where μ 2 is a preset increase factor, Δ max is the loss upper limit, Δ min is a preset loss lower limit, Δ is the accuracy loss, p2 old is the adaptive pruning rate, and p2 new is the increase pruning rate.

[0066] In this embodiment, when the accuracy loss is less than the loss upper limit, it is considered that the accuracy loss is low and the adaptive pruning rate needs to be increased. Specifically, the increase amplitude of the adaptive pruning rate is determined according to the loss upper limit, the loss lower limit, and the increase factor. Among them, the increase factor is a parameter greater than 1, and the closer the increase factor is to 1, the lower the increase amplitude of the adaptive pruning rate.

[0067] In a possible embodiment, training the initial lightweight convolutional neural network through decoupled knowledge distillation to obtain a first lightweight convolutional neural network includes: the model extraction module 120 obtains target class knowledge and non-target class knowledge based on the deep convolutional neural network, and then constructs a loss function of the initial lightweight convolutional neural network based on the target class knowledge and the non-target class knowledge; the compression training module 130 receives the loss function and then trains the initial lightweight convolutional neural network based on the loss function to obtain a first lightweight convolutional neural network.

[0068] Specifically, the probability p corresponding to the target class knowledge t is as shown in Equation 3 below:

[0069]

[0070] The probability p corresponding to the non-target class knowledge nt is as shown in Equation 4 below:

[0071]

[0072] where represents the logits of the j-th class, represents the logits output of the target class knowledge, and C represents the total number of classes. From Equation 3 and Equation 4, the probability corresponding to a single non-target class knowledge can be obtained, as shown in Equation 5 below:

[0073]

[0074] Furthermore, the soft label loss KD of the initial lightweight convolutional neural network is defined as the KL divergence between the prediction probabilities of the deep convolutional neural network and the lightweight convolutional neural network, and its expression is as shown in Equation 6:

[0075]

[0076] where represents the deep convolutional neural network, represents the initial lightweight convolutional neural network, represents the probability of the target class knowledge of model χ. On this basis, the binary distribution probability b is defined, as well as the prediction probabilities for the target class knowledge or non-target class knowledge; at the same time, let p nt represent the non-target class knowledge probability, and the expression of the KD loss is as shown in Equation 7:

[0077]

[0078] Based on Equation 7, let

[0079] At the same time, let The expression of KD can be simplified to Equation 8:

[0080]

[0081] As can be seen from Equation 8, if the prediction accuracy of the deep convolutional neural network is very high, it will lead to the fact that the lightweight convolutional neural network has very limited learning of non-target class knowledge. Therefore, the KD loss is further rewritten as Equation 9:

[0082]

[0083] As can be seen from Equation 9, by controlling the magnitudes of parameter α and parameter β, the learning capabilities of the lightweight convolutional neural network for target class knowledge and non-target class knowledge can be balanced. Based on the derivation process from Equation 3 to Equation 9, the loss function of the initial lightweight convolutional neural network is finally obtained as shown in Equation 10:

[0084]

[0085] where is the soft label loss.

[0086] is the hard label loss.

[0087] In this embodiment, the output of the deep neural network is divided into target class knowledge and non-target class knowledge. Among them, the target class knowledge is the output information of the categories that the lightweight convolutional neural network needs to focus on learning, and the non-target class knowledge is the output information that the lightweight convolutional neural network does not need to pay special attention to. By distinguishing target class knowledge and non-target class knowledge to construct the loss function, the lightweight convolutional neural network can learn the knowledge of the deep convolutional neural network more targeted, and improve the accuracy of the first lightweight convolutional neural network obtained by decoupled knowledge distillation training.

[0088] In a possible embodiment, the obtaining of the target pruning rate and the filter importance group includes: the compression training module 130 obtains a filter set based on the first lightweight convolutional neural network; for any filter in the filter set, the compression training module 130 obtains the filter norm and the filter independence based on the filter; the compression training module 130 obtains the filter importance corresponding to the filter based on the filter norm and the filter independence, and then takes all the filter importances as the filter importance group.

[0089] Specifically, for any multi-channel convolution kernel, the i-th filter of the L-th layer is represented as Its filter norm is as shown in Equation 11 below:

[0090]

[0091] Among them, H represents the height of the convolutional kernel, W represents the width of the convolutional kernel, and C represents the number of channels of the convolutional kernel. represents the weight value of filter 0 at position (i, j) in channel c.

[0092] Flatten and normalize all filters in the filter set, and perform the DBSCAN clustering algorithm on the filters. Specifically, for any filter Select the farthest distance ε and the minimum number of points δ in the neighborhood, and then solve the ε-neighborhood of the filter, as shown in Equation 12 below:

[0093]

[0094] Among them, D represents the filter set; d() represents the Euclidean distance, that is

[0095] represents the number of elements in the set ; if then the filter is considered as the central filter, otherwise the filter is considered as the edge filter. For any filter Execute the above process until all filters in the filter set D are classified, obtaining the central filter set D1 and the edge filter set D2. For any filter in the central filter set D1, mark the filter as visited, and at the same time solve the ε-neighborhood of the filter, and then divide all elements in the ε-neighborhood into the same category set C k and mark them as visited. Specifically, if any element in the ε-neighborhood belongs to the central filter, further solve the sub-ε-neighborhood of the element, and then divide all filters in the sub-ε-neighborhood into a category set C k and mark them as visited. If any element in the ε-neighborhood belongs to the edge filter, divide it into the corresponding category set C k and mark it as visited. Finally, after all filters in the central filter set D1 are marked as visited, put the unmarked filters in the edge filter D2 into the edge set N, obtaining k category sets C k and an edge set N.

[0096] Solve the filter independence based on the k category sets C k and an edge set N. Specifically, when the filter is in the edge set N, the filter independence of the filter When the filter is in the category set Ck When it is, the original filter independence of the filter is as shown in Equation 13 below:

[0097]

[0098] where, it represents the class combination C k the number of filters in the filter; the center

[0099] where F k represents the filter in set k. It can be seen from Equation 13 that for any filter in the class set C k the closer it is to the center of the corresponding class set C k the more filters there are in the corresponding class set C k and the lower its filter independence.

[0100] After obtaining all the original filter independences, the original filter independences are normalized to obtain the filter independence whose expression is as shown in Equation 14 below:

[0101]

[0102] where, G1 max represents the maximum value of the original filter independence in the corresponding class set C k and G1 min represents the minimum value of the original filter independence in the corresponding class set C k

[0103] Finally, the filter importance expression is obtained as shown in Equation 15 below:

[0104]

[0105] It can be seen from Equation 15 that for any filter, the larger its filter norm, the stronger its filter independence, indicating that its filter importance is higher. Among them, the parameters α and β are used to balance the filter norm and the filter independence.

[0106] In this embodiment, the filter importance is calculated comprehensively through the filter norm and the filter independence. Specifically, for any filter, the larger its filter norm and the stronger its filter independence, the more important this filter is considered.

[0107] See Figure 3 , the embodiment of the present invention provides another method for obtaining a student network based on pruning and distillation fusion, including the following steps:

[0108] S210. Obtain an initial lightweight convolutional neural network based on a preset deep convolutional neural network.​

[0109] S220. Train the initial lightweight convolutional neural network through decoupled knowledge distillation to obtain the first lightweight convolutional neural network.

[0110] S230. Obtain the first model accuracy based on the first lightweight convolutional neural network.

[0111] S240. If the compression training module determines that the first model accuracy is less than or equal to the preset accuracy threshold, obtain the initial pruning rate, and then use the initial pruning rate as the target pruning rate.

[0112] S241. Obtain the filter importance group, and prune the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain the second lightweight convolutional neural network.

[0113] S260. Update the initial lightweight convolutional neural network based on the second lightweight convolutional neural network as the result of this round of pruning distillation compression training, and determine whether the number of training rounds reaches the preset training threshold. If the number of training rounds reaches the preset training threshold, execute step S270; if the number of training rounds does not reach the preset training threshold, return to execute step S220.

[0114] S270. Obtain the target lightweight convolutional neural network based on the result of the last round of pruning distillation compression training.

[0115] S250. If the compression training module determines that the first model accuracy is greater than the accuracy threshold, obtain the adaptive pruning rate, and then use the adaptive pruning rate as the target pruning rate.

[0116] S251. Obtain the filter importance group, and prune the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain the second lightweight convolutional neural network.

[0117] S252. Obtain the second model accuracy based on the second lightweight convolutional neural network, and then calculate the precision loss based on the first model accuracy and the second model accuracy.

[0118] S253. If the compression training module determines that the precision loss is greater than or equal to the preset loss upper limit, update the initial lightweight convolutional neural network based on the first lightweight convolutional neural network as the result of this round of pruning distillation compression training.

[0119] S254. If the compression training module determines that the precision loss is less than the loss upper limit, update the initial lightweight convolutional neural network based on the second lightweight convolutional neural network as the result of this round of pruning distillation compression training.

[0120] S255. Update the adaptive pruning rate based on the accuracy loss.

[0121] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0123] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as the protection scope of the present invention. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A student network acquisition method based on pruning, distillation and fusion, characterized in that: include: Obtaining an initial lightweight convolutional neural network based on a preset deep convolutional neural network; Performing a round of pruning distillation compression training on the initial lightweight convolutional neural network, and then obtaining a target lightweight convolutional neural network based on the results of the pruning distillation compression training; The pruning distillation compression training includes: Training the initial lightweight convolutional neural network by decoupling knowledge distillation to obtain a first lightweight convolutional neural network; Get the target pruning rate and filter importance group; Pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network; The initial lightweight convolutional neural network is updated based on the second lightweight convolutional neural network as the result of this round of pruning distillation compression training.

2. The method for obtaining a student network based on pruning, distillation and fusion according to claim 1, characterized in that: Also includes: Performing several rounds of pruning, distillation and compression training on the initial lightweight convolutional neural network until the number of training rounds reaches a preset training threshold; The target lightweight convolutional neural network is obtained based on the results of the last round of pruning, distillation and compression training.

3. A student network acquisition system based on pruning, distillation and fusion, characterized in that: It includes the main control module, model extraction module and compression training module, among which: The model extraction module is used to obtain an initial lightweight convolutional neural network based on a preset deep convolutional neural network; The main control module is used to receive the initial lightweight convolutional neural network, thereby controlling the compression training module to perform a round of pruning distillation compression training on the initial lightweight convolutional neural network, and then obtaining the target lightweight convolutional neural network based on the results of the pruning distillation compression training. The compression training module is used to perform the pruning distillation compression training, including: Training the initial lightweight convolutional neural network by decoupling knowledge distillation to obtain a first lightweight convolutional neural network; Get the target pruning rate and filter importance group; Pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network; The initial lightweight convolutional neural network is updated based on the second lightweight convolutional neural network as the result of this round of pruning distillation compression training.

4. According to the student network acquisition system based on pruning, distillation and fusion according to claim 3, it is characterized in that: Also includes: The main control module is also used to control the compression training module to perform several rounds of pruning distillation compression training on the initial lightweight convolutional neural network until the number of training rounds reaches a preset training threshold, and then obtain the target lightweight convolutional neural network based on the results of the last round of pruning distillation compression training.

5. According to claim 4, a student network acquisition system based on pruning, distillation and fusion, characterized in that: The obtaining of the target pruning rate and the filter importance group comprises: The compression training module obtains a first model accuracy based on the first lightweight convolutional neural network; If the compression training module determines that the accuracy of the first model is less than or equal to a preset accuracy threshold, an initial pruning rate is obtained, and then the initial pruning rate is used as the target pruning rate.

6. A student network acquisition system based on pruning, distillation and fusion according to claim 5, characterized in that: If the compression training module determines that the accuracy of the first model is greater than the accuracy threshold, an adaptive pruning rate is obtained, and the adaptive pruning rate is used as the target pruning rate; after pruning the first lightweight convolutional neural network based on the target pruning rate and the filter importance group to obtain a second lightweight convolutional neural network, the method further includes: The compression training module obtains a second model accuracy based on the second lightweight convolutional neural network, and further calculates the accuracy loss based on the first model accuracy and the second model accuracy; The compression training module updates the adaptive pruning rate based on the accuracy loss; If the compression training module determines that the accuracy loss is greater than or equal to the preset loss upper limit, the initial lightweight convolutional neural network is updated based on the first lightweight convolutional neural network as the result of the current round of pruning distillation compression training; If the compression training module determines that the accuracy loss is less than the loss upper limit, the initial lightweight convolutional neural network is updated based on the second lightweight convolutional neural network as the result of this round of pruning distillation compression training.

7. The student network acquisition system based on pruning, distillation and fusion according to claim 6 is characterized in that: When the compression training module determines that the precision loss is greater than or equal to the loss upper limit, updating the adaptive pruning rate based on the precision loss includes: The compression training module calculates the attenuated pruning rate based on the preset attenuation factor and the adaptive pruning rate, and then updates the adaptive pruning rate based on the attenuated pruning rate, which is expressed as follows: p2 new =μ1*p2 old ; Wherein, μ1 is the attenuation factor, p2 old is the adaptive pruning rate, p2 new is the attenuation pruning rate.

8. The student network acquisition system based on pruning, distillation and fusion according to claim 6 is characterized in that: When the compression training module determines that the precision loss is less than the loss upper limit, updating the adaptive pruning rate based on the precision loss includes: The compression training module calculates the growth pruning rate based on the loss upper limit and the adaptive pruning rate, and then updates the adaptive pruning rate based on the growth pruning rate, which is expressed as follows: Among them, μ2 is the preset growth factor, Δ max is the upper limit of loss, Δ min is the preset loss lower limit, Δ is the accuracy loss, p2 old is the adaptive pruning rate, p2 new is the growth pruning rate.

9. The student network acquisition system based on pruning, distillation and fusion according to claim 4 is characterized in that: The initial lightweight convolutional neural network is trained by decoupling knowledge distillation to obtain a first lightweight convolutional neural network, including: The model extraction module acquires target class knowledge and non-target class knowledge based on the deep convolutional neural network, and then constructs a loss function of the initial lightweight convolutional neural network based on the target class knowledge and the non-target class knowledge; The compression training module receives the loss function, and then trains the initial lightweight convolutional neural network based on the loss function to obtain a first lightweight convolutional neural network.

10. The student network acquisition system based on pruning, distillation and fusion according to claim 4, characterized in that: The obtaining of the target pruning rate and the filter importance group comprises: The compression training module obtains a filter set based on the first lightweight convolutional neural network; For any one filter in the filter set, the compression training module obtains a filter norm and filter independence based on the filter; The compression training module obtains the filter importance corresponding to the filter based on the filter norm and the filter independence, and then takes all the filter importances as the filter importance group.

Citation Information

Cited By

  • Deployment method, prediction method and system of electrical equipment state prediction network

    CN121561523A