Student model automatic generation system based on knowledge distillation
By introducing a method of dynamic update loss function of structural prediction network in the knowledge distillation system, the problem of difficult to automatically generate the structure size of student models in the prior art is solved, and the student models suitable for resource-constrained scenarios are generated, and the practicality of the model is improved.
Patent Information
- Application Number
- CN202510150870.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-13
AI Technical Summary
The existing knowledge distillation technology is difficult to automatically generate student models with structural sizes that meet practical application scenarios, making it impractical to deploy large-scale convolutional neural network models on devices with limited hardware and software resources.
By designing a student model automatic generation system based on knowledge distillation, the system includes an automatic generation module, a model acquisition module, a loss function building module, a structure prediction module and a knowledge distillation module. The system uses deep convolutional neural networks as the teacher model and introduces structural prediction networks to dynamically update the lightweight convolutional neural network loss function, allowing the student model to dynamically adjust its network structure during knowledge distillation training.
It realizes the student model with a structure size suitable for application in scenarios with limited hardware and software resources while maintaining the accuracy of the original complex teacher model, which improves the practicality of the student model.
Smart Images

Figure CN120146099A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network models, and particularly to an automatic generation system for a student model based on knowledge distillation. Background Art
[0002] With the continuous development of modern artificial intelligence technology, convolutional neural networks (CNNs) have achieved remarkable success in multiple fields such as image recognition, natural language processing, and speech recognition. The rapid development of convolutional neural networks benefits from their natural advantages in processing data with a grid structure such as pixel arrays of images. Through the organic combination of convolutional layers, pooling layers, and fully connected layers, convolutional neural networks can effectively extract and learn hierarchical features of data, thus demonstrating superior performance in the processing of various tasks.
[0003] However, with the continuous deepening and expansion of convolutional neural network models, the number of parameters and computational complexity of the models have also increased sharply. This has led to two main problems: one is that training large convolutional neural network models requires a large amount of computing resources and time, and the other is that it becomes impractical to deploy these large-scale convolutional neural network models on devices with limited hardware and software resources. To solve these problems, currently, the field is actively promoting the development of neural network compression technology. Neural network compression is a technology that tries to maintain its original performance while reducing the size and complexity of neural networks. Its main purpose is to reduce the computing resource requirements, storage space occupancy, and energy consumption of neural network models.
[0004] Among the existing network compression technologies, knowledge distillation is a relatively effective method. The concept of knowledge distillation was first proposed by Bucila et al. in 2006, and was further developed and popularized by Hinton et al. in 2015. The core idea of knowledge distillation is to transfer the knowledge of a large and complex teacher model to a smaller and simpler student model. In this way, the student model can maintain the original performance of the teacher model as much as possible while maintaining a small size and computing requirements. Existing knowledge distillation technology usually sets the target structure size of the student model before training, and then performs knowledge distillation training to enable the student model to learn the characteristics of the teacher model. In practical applications, how to design the target size of the student model needs to be based on a lot of experience and experiments, and there is no exact method, making it difficult to select the size of the student model. When the student model is designed to be too large, it cannot meet the actual application scenarios with limited hardware and software resources. When the student model is designed to be too small, it cannot accurately restore the characteristics of the large and complex teacher model. Once the target size of the student model is set, it cannot be changed during the knowledge distillation training process. Often, the practicality of the student model size can only be verified through actual applications after the training results are obtained, and retraining can be performed when its size does not meet the actual application scenario. Summary of the invention
[0005] The purpose of the present invention is to provide a student model automatic generation system based on knowledge distillation to solve the above technical problems, so that it can automatically generate a student model whose structure size meets the actual application scenario.
[0006] To achieve the above object, the present invention provides a student model automatic generation system based on knowledge distillation, including an automatic generation module, a model acquisition module, a loss function construction module, a structure prediction module, and a knowledge distillation module, wherein: the model acquisition module is used to obtain an initial lightweight convolutional neural network based on a preset deep convolutional neural network; the loss function construction module is used to construct a lightweight convolutional neural network loss function of the initial lightweight convolutional neural network based on the deep convolutional neural network; the structure prediction module is used to construct a structure prediction network based on the initial lightweight convolutional neural network; the automatic generation module is used to receive the initial lightweight convolutional neural network and the lightweight convolutional neural network loss function, and then control the knowledge distillation module to perform several rounds of knowledge distillation training on the initial lightweight convolutional neural network, so as to generate a student model based on the result of any round of the knowledge distillation training; the knowledge distillation module is used to perform the knowledge distillation training, including: training the initial lightweight convolutional neural network based on the lightweight convolutional neural network loss function to obtain a first lightweight convolutional neural network; updating the lightweight convolutional neural network loss function based on the structure prediction network and the first lightweight convolutional neural network; updating the initial lightweight convolutional neural network based on the first lightweight convolutional neural network as the result of this round of the knowledge distillation training.
[0007] The above student model automatic generation system based on knowledge distillation uses a deep convolutional neural network as the teacher model and a lightweight convolutional neural network as the student model. During the knowledge distillation training process, a structure prediction network corresponding to the actual application scenario is introduced, and the lightweight convolutional neural network loss function is dynamically updated through the structure prediction network and the first lightweight convolutional neural network, so that the lightweight convolutional neural network loss function can dynamically adjust the process of training the lightweight convolutional neural network. Finally, the lightweight convolutional neural network can dynamically adjust its network structure and change its structure size during the knowledge distillation training process, so that the trained student model can meet the hardware and software resource limitations of the actual application scenario on the premise of reflecting the accuracy of the original complex teacher model. At the same time, this system enables the automatic generation of a student model with a structure size suitable for application in scenarios with limited hardware and software resources and a model accuracy close to that of the complex teacher model without manually designing the target size of the student model based on a large amount of experience and experiments. As long as a structure prediction network corresponding to the actual application scenario is introduced, the structure prediction network can be used to dynamically adjust the knowledge distillation training process, improving the practicality of the student model.
[0008] In a possible implementation, in the knowledge distillation module, updating the lightweight convolutional neural network loss function based on the structure prediction network and the first lightweight convolutional neural network includes: obtaining a set of parameter tuples based on the first lightweight convolutional neural network; for any parameter tuple in the set of parameter tuples, obtaining its corresponding mask based on the structure prediction network; and updating the lightweight convolutional neural network loss function based on the set of parameter tuples and all the masks.
[0009] In this implementation, the structure of the first lightweight convolutional neural network is split into a set of parameter tuples; for any parameter tuple, its corresponding mask is obtained based on the structure prediction network. Specifically, each mask element takes a value of 0 or 1. When the mask value is 1, it means that the corresponding parameter tuple should be discarded from the lightweight convolutional neural network loss function; when the mask value is 0, it means that the corresponding parameter tuple should be retained in the lightweight convolutional neural network loss function. By the above process, updating the lightweight convolutional neural network loss function can discard unimportant structures and parameters in the lightweight convolutional neural network during the training process, thereby improving the sparsity of the finally generated student model, reducing the occupied space of the student model, and making it more suitable for applications in scenarios with limited hardware and software resources.
[0010] In a possible implementation, after the knowledge distillation module updates the lightweight convolutional neural network loss function based on the structure prediction network and the first lightweight convolutional neural network: the loss function construction module is further configured to obtain a structure prediction network loss function based on all the masks and the lightweight convolutional neural network loss function, so that the knowledge distillation module updates the structure prediction network based on the structure prediction network loss function.
[0011] In this implementation, updating the structure prediction network through the structure prediction network loss function can dynamically adjust the structure prediction network and the logic of generating masks according to the structure of each round of training, and further dynamically adjust the process of adjusting the lightweight convolutional neural network structure in each round of training, so that the entire training process can automatically explore the optimal structure of the lightweight convolutional neural network without manual intervention, and finally realize the automatic generation of a student model with a structure size suitable for applications in scenarios with limited hardware and software resources, while the model accuracy is close to that of a complex teacher model, improving the practicality of the student model.
[0012] In a possible implementation, the structure prediction network loss function is specifically:
[0013] where represents that when the input is x, the output is y, and the first lightweight convolutional neural network is The lightweight convolutional neural network loss function when the mask is w, and P(w) represents the model computation amount after updating the lightweight convolutional neural network loss function based on the set of parameter tuples and all the masks. γ is used to control the regularization strength.
[0014] In this implementation, updating the structure prediction network loss function based on the computation amount of the lightweight convolutional neural network can dynamically adjust the process of adjusting the structure of the lightweight convolutional neural network in each round of training according to the structure size of the lightweight convolutional neural network in each round of training, enabling the entire training process to automatically explore the optimal structure of the lightweight convolutional neural network, so that the finally generated student model will not be unable to meet the actual application scenarios with limited hardware and software resources due to an overly large structure, nor will it be unable to accurately restore the characteristics of a large and complex teacher model due to an overly small structure. Further, by minimizing the structure prediction network loss function, it can be ensured that, in the case of introducing a mask to update the lightweight convolutional neural network, the balance between the prediction accuracy and the structure size of the finally generated student network is achieved, making the accuracy of the student network model close to that of the complex teacher model while the structure size is suitable for application in the corresponding scenario.
[0015] In a possible implementation, the number of update rounds of the structure prediction network is preset in the knowledge distillation module; for any round of the knowledge distillation training, if the knowledge distillation module determines that it does not belong to the update rounds, the structure prediction network is not updated in this round of the knowledge distillation training.
[0016] In this implementation, considering that in the initial stage of the distillation training, the parameters of the lightweight convolutional neural network may not have converged yet, and at this time, when the structure prediction network is updated, the generated mask may be inaccurate. Therefore, for the entire process of performing several rounds of knowledge distillation training on the initial lightweight convolutional neural network, the starting round T start and the ending round T end of the structure prediction network training are set, and the rounds between T start and T end are used as the update rounds of the structure prediction network. By setting the training rounds of the structure prediction network, the structure size of the lightweight convolutional neural network remains unchanged in the initial stage and the final stage of the distillation training rounds, ensuring the stability of the structure adjustment of the lightweight convolutional neural network during the training process, thereby improving the accuracy of the finally generated student model's ability to express the knowledge of the teacher model. At the same time, it is ensured that the structure prediction network can generate an effective mask to guide the lightweight convolutional neural network to reasonably adjust its structure size and improve the rationality of the structure scale of the finally generated student model.
[0017] In a possible implementation, for any of the parameter tuples in the set of parameter tuples, the knowledge distillation module is further configured to: obtain the importance of the corresponding convolutional layer based on the first lightweight convolutional neural network; update the loss function of the lightweight convolutional neural network based on the set of parameter tuples, all the masks, all the convolutional layer importances, and a preset penalty threshold.
[0018] In this implementation, in addition to the set of parameter tuples and all the masks described above, all the convolutional layer importances and a preset penalty threshold are introduced to update the loss function of the lightweight convolutional neural network. For any parameter tuple, when the importance of the convolutional layer to which it belongs is less than the preset penalty threshold, it means that the parameter tuple should be discarded from the loss function of the lightweight convolutional neural network; when the importance of the convolutional layer to which it belongs is greater than or equal to the preset penalty threshold, it means that the corresponding parameter tuple should be retained in the loss function of the lightweight convolutional neural network. Through the above process, the accuracy of judging the important and unimportant structures of the lightweight convolutional neural network is further enhanced, thereby improving the accuracy of the final generated student model in expressing the knowledge of the teacher model; at the same time, through the above update process, the expression ability of the student model is concentrated in the convolutional layers with high importance, reducing the computational amount of the student model, improving the running efficiency of the student model, and making it more suitable for application in scenarios with limited hardware and software resources.
[0019] In a possible implementation, the process of updating the loss function of the lightweight convolutional neural network is specifically as follows:
[0020]
[0021] where L total represents the loss function of the lightweight convolutional neural network, represents the set of parameter tuples, [w]g represents the mask corresponding to the g-th parameter tuple, represents the importance of the convolutional layer corresponding to the g-th parameter tuple, and δ represents the penalty threshold.
[0022] In this implementation, λg represents the penalty coefficient. When λg takes the value of 0, it means that the corresponding g-th parameter tuple is not penalized and can be retained in the lightweight convolutional neural network; when λg takes the value of 1, it means that the corresponding g-th parameter tuple needs to be discarded from the lightweight convolutional neural network. Specifically, when the mask corresponding to the g-th parameter tuple takes the value of 1, λg takes the value of 0; when the importance of the convolutional layer corresponding to the g-th parameter tuple is greater than or equal to the preset penalty threshold, λg takes the value of 0. Through the above process, the accuracy of judging the important and unimportant structures of the lightweight convolutional neural network is further enhanced, thereby improving the accuracy of the final generated student model in expressing the knowledge of the teacher model; at the same time, through the above update process, the expression ability of the student model is concentrated in the convolutional layers with high importance, reducing the computational amount of the student model, improving the running efficiency of the student model, and making it more suitable for application in scenarios with limited hardware and software resources.
[0023] In a possible implementation, in the loss function construction module, the lightweight convolutional neural network loss function for constructing the initial lightweight convolutional neural network based on the deep convolutional neural network includes: obtaining a hard label loss function and a soft label loss function based on the deep convolutional neural network and the initial lightweight convolutional neural network, and then constructing the lightweight convolutional neural network loss function based on the hard label loss function and the soft label loss function.
[0024] In this implementation, the loss function of the lightweight convolutional neural network consists of a hard label loss function and a soft label loss function. Among them, the hard label loss function ensures that the learning objective of the lightweight convolutional neural network is consistent with that of the deep convolutional neural network, avoiding target deviation; the soft label loss function helps the lightweight convolutional neural network learn the complex knowledge features of the deep convolutional neural network, and finally makes the accuracy of the automatically generated student model close to that of the complex teacher model, improving the accuracy of the student model.
[0025] In a possible implementation, in the loss function construction module, training the initial lightweight convolutional neural network based on the lightweight convolutional neural network loss function to obtain the first lightweight convolutional neural network includes: obtaining the hard label loss function gradient based on the hard label loss function and obtaining the soft label loss function gradient based on the soft label loss function; constructing a lightweight convolutional neural network loss function gradient based on the hard label loss function gradient and the soft label loss function gradient, so that the knowledge distillation module trains the initial lightweight convolutional neural network based on the lightweight convolutional neural network loss function gradient through the gradient descent algorithm to obtain the first lightweight convolutional neural network.
[0026] In this implementation manner, the initial lightweight convolutional neural network is trained by the gradient descent algorithm, so that the loss function of the lightweight convolutional neural network gradually approaches the minimum value, and finally the accuracy of the automatically generated student network model is close to that of the complex teacher model, improving the accuracy of the student model.
[0027] In a possible implementation manner, the structure prediction module is used to construct a structure prediction network based on the initial lightweight convolutional neural network, including: obtaining an initial prediction network; obtaining a total data set based on the initial lightweight convolutional neural network; randomly sampling the total data set according to a preset sampling ratio to obtain a sub-data set; training the initial prediction network based on the sub-data set to obtain the structure prediction network.
[0028] In this implementation manner, a sub-data set is obtained through random sampling, and the initial prediction network is trained based on the sub-data set to obtain the structure prediction network, reducing the computational overhead and training time while ensuring the accuracy of the obtained structure prediction network. Description of the Drawings
[0029] Figure 1 is a structural diagram of a system for automatically generating a student model based on knowledge distillation provided by an embodiment of the present invention;
[0030] Wherein: 100, automatic generation module; 200, model acquisition module; 300, loss function construction module; 400, structure prediction module; 500, knowledge distillation module. Detailed Embodiments
[0031] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0032] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed descriptions of the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0033] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0034] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0035] Existing knowledge distillation techniques usually set the target structure size of the student model before training, and then perform the training of knowledge distillation to enable the student model to learn the features of the teacher model. In practical applications, how to design the target size of the student model requires a lot of experience and experiments, and is manually set by humans. There is no exact method, making it difficult to select the size of the student model. When the student model is designed too large, it cannot meet the actual application scenarios with limited hardware and software resources. When the student model is designed too small, it cannot accurately restore the characteristics of large and complex teacher models. Once the target size of the student model is set, it cannot be changed during the training process of knowledge distillation. Often, only after obtaining the training results can the practicality of the student model size be verified through actual applications, and retraining is required when its size does not meet the actual application scenarios.
[0036] To solve the above technical problems, refer to Figure 1 , an embodiment of the present invention provides a student model automatic generation system based on knowledge distillation, including an automatic generation module 100, a model acquisition module 200, a loss function construction module 300, a structure prediction module 400, and a knowledge distillation module 500. Among them:
[0037] The model acquisition module 200 is used to obtain an initial lightweight convolutional neural network based on a preset deep convolutional neural network.
[0038] The present invention uses a deep convolutional neural network (DCNN) as the teacher model. In the field of artificial intelligence technology, a deep convolutional neural network specifically refers to a deep learning model dedicated to processing data with a grid structure, especially a complex neural network with a relatively large number of network layers and a large scale, including some neural network models for image classification and image segmentation. Based on the deep convolutional neural network that needs to be model compressed, an appropriate initial lightweight convolutional neural network is selected. Among them, the lightweight convolutional neural network (LCNN) specifically refers to a neural network model with fewer layers and parameters, such as MobileNet, ShuffleNet, and SqueezeNet. Through the above-mentioned method for obtaining a student network based on pruning and distillation fusion, the initial lightweight convolutional neural network can fully learn the knowledge of the deep convolutional neural network through distillation, and further reduce the number of layers and parameters through pruning to generate a target lightweight convolutional neural network as the student network.
[0039] Specifically, the present invention uses VGG16 as the preset deep convolutional neural network of the teacher model. VGG16 has 13 convolutional layers and 3 fully connected layers. The dataset used in the present invention records the operation processes of electrical devices in three different aging states. Specifically, the dataset samples the current, voltage, and temperature of the device at a certain frequency to obtain the electrical characteristics of the electrical devices in three different aging states during the corresponding operation processes. In terms of dataset construction, the present invention traverses the dataset with a window of length 24, takes the current, voltage, and temperature as the eigenvalue inputs of VGG16, and then obtains the corresponding aging state of the electrical device as the output; that is, inputs a time series data of 24*3 and outputs the corresponding aging state label. Since the dataset has recorded the operation processes of electrical devices in three different aging states, the trained VGG16 is a three-classification model, which is used to judge which of the three aging states the corresponding device belongs to according to the input device electrical characteristic values, and then judge the aging situation of the electrical device. After the present invention fully trains VGG16 using this dataset, a deep convolutional neural network with strong expressive ability is obtained as the teacher model of this embodiment.
[0040] Define the model size of the trained deep convolutional neural network as M t and define the scaling factor λ ∈ [0,1] such that the model size of the initial lightweight convolutional neural network is M s = λ * M t, and use the initial lightweight convolutional neural network as the student network of this embodiment. Specifically, based on the preset scaling factor λ, this embodiment retains the first 10 convolutional layers of the preset deep convolutional neural network as the convolutional layers of the initial lightweight convolutional neural network, and retains the two fully connected layers of the preset deep convolutional neural network as the fully connected layers of the initial convolutional neural network. Among them, the obtained initial lightweight convolutional neural network includes two convolutional layers with 64 filters, two convolutional layers with 128 filters, three convolutional layers with 256 filters, and three convolutional layers with 512 filters. Finally, based on the three aging states of the electrical equipment in the dataset, the output size of the fully connected layer of the initial convolutional neural network is set to 3, so that the initial convolutional neural network can judge the aging state corresponding to a specific electrical equipment according to the input feature data.
[0041] The loss function construction module 300 is used to construct the lightweight convolutional neural network loss function of the initial lightweight convolutional neural network based on the deep convolutional neural network.
[0042] The structure prediction module 400 is used to construct a structure prediction network based on the initial lightweight convolutional neural network.
[0043] The automatic generation module 100 is used to receive the initial lightweight convolutional neural network and the lightweight convolutional neural network loss function, and then control the knowledge distillation module 500 to perform several rounds of knowledge distillation training on the initial lightweight convolutional neural network, so as to generate a student model based on the result of any round of the knowledge distillation training.
[0044] The knowledge distillation module 500 is used to perform the knowledge distillation training, including: training the initial lightweight convolutional neural network based on the lightweight convolutional neural network loss function to obtain the first lightweight convolutional neural network; updating the lightweight convolutional neural network loss function based on the structure prediction network and the first lightweight convolutional neural network; updating the initial lightweight convolutional neural network based on the first lightweight convolutional neural network as the result of this round of the knowledge distillation training.
[0045] The above-mentioned student model automatic generation system based on knowledge distillation uses the VGG16 deep convolutional neural network as the teacher model and the lightweight convolutional neural network as the student model. During the knowledge distillation training process, a structure prediction network corresponding to the actual application scenario is introduced, and the loss function of the lightweight convolutional neural network is dynamically updated through the structure prediction network and the first lightweight convolutional neural network, so that the loss function of the lightweight convolutional neural network can dynamically adjust the process of training the lightweight convolutional neural network. Finally, the lightweight convolutional neural network can dynamically adjust its network structure and change its structure size during the knowledge distillation training process, so that the trained student model can meet the hardware and software resource limitations of the actual deployment scenario of the electrical equipment while maintaining the accuracy of the original VGG16 model in judging the aging state of the electrical equipment. At the same time, this system makes it unnecessary to manually adjust the target size of the student model multiple times based on the local deployment scenario of the electrical equipment. As long as the structure prediction network corresponding to the local deployment scenario of the electrical equipment is introduced, the structure prediction network can dynamically adjust the knowledge distillation training process to automatically generate a student model whose structure size is suitable for the local deployment scenario of the electrical equipment with limited hardware and software resources, and the model accuracy is close to that of the complex VGG16 teacher model. This makes the student model obtained in this embodiment more suitable for deployment in the local environment of the electrical equipment with limited resources, and can judge the current aging state of the electrical equipment in real time according to the current electrical characteristics during the operation of the local electrical equipment, providing a more accurate analysis and judgment basis for the maintenance of the local electrical equipment.
[0046] In a possible embodiment, in the knowledge distillation module 500, updating the loss function of the lightweight convolutional neural network based on the structure prediction network and the first lightweight convolutional neural network includes: obtaining a set of parameter tuples based on the first lightweight convolutional neural network; for any parameter tuple in the set of parameter tuples, obtaining its corresponding mask based on the structure prediction network; and updating the loss function of the lightweight convolutional neural network based on the set of parameter tuples and all the masks.
[0047] Specifically, a parameter tuple refers to a combination of basic components of the learnable parameters included in the lightweight convolutional neural network. If all the learnable parameters in a parameter tuple are set to 0, the output of the components belonging to this parameter tuple has no influence on the final output result of the lightweight convolutional neural network. In particular, in the lightweight convolutional neural network, assuming the input feature image is In, the output feature image Out is as shown in Equation 1 below:
[0048]
[0049] where F represents the filter parameters; represents a convolution operation; μ represents the mean of the batch data; σ represents the variance of the batch data; γ and β are scaling factors used to adjust the output of the batch normalization layer, and Relu() is an activation operation, i.e., Relu(x) = max(0, x). Among the above parameters, F, γ, and β are learnable parameters. If F, γ, and β are all set to 0, the final output of the filter is 0, which has no impact on the subsequent convolution training operations of the lightweight convolutional neural network. Therefore, based on Equation 1, the i-th filter of the l-th layer is defined and the corresponding and is a tuple of parameters. The set of tuples of parameters for the entire lightweight convolutional neural network is where represents the number of tuples of parameters in, as shown in Equation 2 below:
[0050]
[0051] where C l represents the number of tuples of parameters of the l-th layer; [χ] g represents the g-th element of the tuple of parameters x.
[0052] In this embodiment, the structure of the first lightweight convolutional neural network is split into a set of tuples of parameters; for any tuple of parameters, its corresponding mask is obtained based on the structure prediction network. Specifically, each mask element takes a value of 0 or 1. When the mask value is 1, it means that the corresponding tuple of parameters should be discarded from the lightweight convolutional neural network loss function; when the mask value is 0, it means that the corresponding tuple of parameters should be retained in the lightweight convolutional neural network loss function. Through the above process, updating the lightweight convolutional neural network loss function can discard unimportant structures and parameters in the lightweight convolutional neural network during the training process, thereby improving the sparsity of the finally generated student model, reducing the occupied space of the student model, and making it more suitable for application in scenarios with limited hardware and software resources.
[0053] In a possible embodiment, after the knowledge distillation module 500 updates the lightweight convolutional neural network loss function based on the structure prediction network and the first lightweight convolutional neural network: the loss function construction module 300 is further configured to obtain a structure prediction network loss function based on all the masks and the lightweight convolutional neural network loss function, so that the knowledge distillation module 500 updates the structure prediction network based on the structure prediction network loss function.
[0054] In this embodiment, by updating the structure prediction network using the structure prediction network loss function, it is possible to dynamically adjust the structure prediction network and the logic of the generated mask according to the structure of each round of training, and further dynamically adjust the process of adjusting the lightweight convolutional neural network structure in each round of training, enabling the entire training process to automatically explore the optimal structure of the lightweight convolutional neural network without manual intervention. Finally, a student model with a structure size suitable for scenarios with limited hardware and software resources can be automatically generated, while the model accuracy is close to that of the complex teacher model, improving the practicality of the student model.
[0055] In a possible embodiment, the structure prediction network loss function is specifically as shown in Equation 3 below:
[0056]
[0057] Where, represents the lightweight convolutional neural network loss function when the input is x, the output is y, the first lightweight convolutional neural network is and the mask is w. P(w) represents the model computational amount after updating the lightweight convolutional neural network loss function based on the parameter tuple set and all the masks. γ is used to control the regularization strength.
[0058] In this embodiment, by updating the structure prediction network loss function based on the computational amount of the lightweight convolutional neural network, it is possible to dynamically adjust the process of adjusting the lightweight convolutional neural network structure in each round of training according to the structure size of the lightweight convolutional neural network in each round of training, enabling the entire training process to automatically explore the optimal structure of the lightweight convolutional neural network, so that the finally generated student model will not be unable to meet the actual application scenarios with limited hardware and software resources due to an overly large structure, nor will it be unable to accurately restore the characteristics of the large and complex teacher model due to an overly small structure. Further, by minimizing the structure prediction network loss function, it is possible to ensure the balance between the prediction accuracy and the computational amount of the finally generated student network in the case of introducing a mask to update the lightweight convolutional neural network, enabling the student network model accuracy to be close to that of the complex teacher model while the structure size is suitable for application in the corresponding scenario.
[0059] In a possible embodiment, the update rounds of the structure prediction network are preset in the knowledge distillation module 500; for any round of the knowledge distillation training, if the knowledge distillation module 500 determines that it does not belong to the update rounds, the structure prediction network is not updated in this round of the knowledge distillation training.
[0060] In this embodiment, considering that in the initial stage of distillation training, the parameters of the lightweight convolutional neural network may not have converged yet, and at this time, when the structure prediction network is updated, the generated mask may not be accurate. Therefore, for the entire process of performing several rounds of knowledge distillation training on the initial lightweight convolutional neural network, the round T at which the structure prediction network starts training is set. start and the round T at which the training ends end , and the rounds between T start and T end are used as the update rounds of the structure prediction network. By setting the training rounds of the structure prediction network, the structure size of the lightweight convolutional neural network remains unchanged at the initial stage and the end stage of the distillation training rounds, ensuring the stability of the structure adjustment of the lightweight convolutional neural network during the training process, thereby improving the accuracy of the final generated student model in expressing the knowledge of the teacher model. At the same time, it ensures that the structure prediction network obtains good prediction ability, enabling it to generate effective masks to guide the lightweight convolutional neural network to reasonably adjust its structure size and improve the rationality of the structure scale of the final generated student model.
[0061] In a possible embodiment, for any of the parameter tuples in the set of parameter tuples, the knowledge distillation module 500 is further configured to: obtain the importance of the corresponding convolutional layer based on the first lightweight convolutional neural network; update the loss function of the lightweight convolutional neural network based on the set of parameter tuples, all the masks, all the convolutional layer importances, and a preset penalty threshold.
[0062] In this embodiment, in addition to the set of parameter tuples and all the masks described above, all the convolutional layer importances and a preset penalty threshold are introduced to update the loss function of the lightweight convolutional neural network. For any parameter tuple, when the importance of the convolutional layer to which it belongs is less than the preset penalty threshold, it means that the parameter tuple should be discarded from the loss function of the lightweight convolutional neural network; when the importance of the convolutional layer to which it belongs is greater than or equal to the preset penalty threshold, it means that the corresponding parameter tuple should be retained in the loss function of the lightweight convolutional neural network. Through the above process, the accuracy of judging the important structure and unimportant structure of the lightweight convolutional neural network is further enhanced, thereby improving the accuracy of the final generated student model in expressing the knowledge of the teacher model; at the same time, through the above update process, the expression ability of the student model is concentrated in the convolutional layers with high importance, reducing the computational amount of the student model and improving the running efficiency of the student model, making it more suitable for application in scenarios with limited hardware and software resources.
[0063] In a possible embodiment, the process of updating the loss function of the lightweight convolutional neural network is specifically as shown in Equation 4 below:
[0064]
[0065] Among them, L total represents the lightweight convolutional neural network loss function, represents the set of parameter tuples, and [w]g represents the mask corresponding to the g-th parameter tuple, represents the importance of the convolutional layer corresponding to the g-th parameter tuple, and δ represents the penalty threshold.
[0066] Specifically, for a lightweight convolutional neural network with L convolutional layers, calculate the importance of each convolutional layer where represents the γ parameter of the i-th BN layer in the l-th convolutional layer. After obtaining the importance of all convolutional layers, set the penalty threshold δ, and the parameter tuples corresponding to the convolutional layer with importance less than the penalty threshold will be penalized.
[0067] In this embodiment, λg represents the penalty coefficient. When λg takes the value of 0, it means that the corresponding g-th parameter tuple is not penalized and can be retained in the lightweight convolutional neural network; when λg takes the value of 1, it means that the corresponding g-th parameter tuple needs to be discarded from the lightweight convolutional neural network. Specifically, when the mask corresponding to the g-th parameter tuple takes the value of 1, λg takes the value of 0; when the importance of the convolutional layer corresponding to the g-th parameter tuple is greater than or equal to the preset penalty threshold, λg takes the value of 0. Through the above process, the accuracy of judging the important and unimportant structures of the lightweight convolutional neural network is further enhanced, thereby improving the accuracy of the student model expressing the knowledge of the teacher model; at the same time, through the above update process, the expression ability of the student model is concentrated in the convolutional layers with high importance, reducing the computational amount of the student model, improving the running efficiency of the student model, and making it more suitable for application in scenarios with limited hardware and software resources.
[0068] In a possible embodiment, in the loss function construction module 300, the lightweight convolutional neural network loss function for constructing the initial lightweight convolutional neural network based on the deep convolutional neural network includes: obtaining a hard label loss function and a soft label loss function based on the deep convolutional neural network and the initial lightweight convolutional neural network, and then constructing the lightweight convolutional neural network loss function based on the hard label loss function and the soft label loss function.
[0069] Specifically, set the loss function of the lightweight convolutional neural network as Equation 5 below:
[0070] L total = L original + L distill (Equation 5)
[0071] where, L totalDenote the loss function of the lightweight convolutional neural network; \(L\) original Denote the hard-label loss; \(L\) distill Denote the soft-label loss. Further, the hard-label loss can be expressed as Equation 6 below:
[0072]
[0073] where \(N\) denotes the total number of classes; \(y\) i denotes the true label of the \(i\)-th class. If the sample belongs to the \(i\)-th class, its true value is 1, otherwise its true value is 0. \(p\) S (\(i\)) is the probability predicted by the lightweight convolutional neural network for the \(i\)-th class.
[0074] Further, let \(T(x)\) be the logits output of the deep convolutional neural network for the input \(x\), and \(S(x)\) be the logits output of the lightweight convolutional neural network. Then the corresponding softmax probability distributions \(p\) T and \(p\) S are shown in Equation 7 and Equation 8 respectively as follows:
[0075]
[0076] From Equation 7 and Equation 8, the expression of the soft-label loss can be obtained as Equation 9 below:
[0077]
[0078] where \(i\) is the class index, and \(\alpha\) is used to balance the soft-label loss function and the hard-label loss function.
[0079] In this embodiment, the loss function of the lightweight convolutional neural network consists of a hard-label loss function and a soft-label loss function. Among them, the hard-label loss function ensures that the learning objective of the lightweight convolutional neural network is consistent with that of the deep convolutional neural network, avoiding target deviation; the soft-label loss function helps the lightweight convolutional neural network learn the complex knowledge features of the deep convolutional neural network, and finally makes the accuracy of the automatically generated student model close to that of the complex teacher model, improving the accuracy of the student model.
[0080] In a possible embodiment, in the loss function construction module 300, training the initial lightweight convolutional neural network based on the lightweight convolutional neural network loss function to obtain a first lightweight convolutional neural network includes: obtaining a hard label loss function gradient based on the hard label loss function, and obtaining a soft label loss function gradient based on the soft label loss function; constructing a lightweight convolutional neural network loss function gradient based on the hard label loss function gradient and the soft label loss function gradient, so that the knowledge distillation module 500 trains the initial lightweight convolutional neural network based on the lightweight convolutional neural network loss function gradient through a gradient descent algorithm to obtain a first lightweight convolutional neural network.
[0081] Specifically, based on Equations 6 to 9, the hard label loss function gradient can be obtained as shown in Equation 10 below:
[0082]
[0083] Furthermore, based on Equations 6 to 9, the following Equation 11 can be obtained:
[0084]
[0085] where δ ik is the Kronecker function, and δ ik is 1 when i = k, and 0 otherwise. Thus, the expression 12 can be obtained:
[0086]
[0087] Furthermore, from Equation 11 and Equation 12, the soft label loss function gradient can be obtained as shown in Equation 13 below:
[0088]
[0089] From the hard label loss function gradient shown in Equation 10 and the soft label loss function gradient shown in Equation 13, the lightweight convolutional neural network loss function gradient can be obtained as shown in Equation 14 below:
[0090]
[0091] In this embodiment, the initial lightweight convolutional neural network is trained through a gradient descent algorithm, so that the loss function of the lightweight convolutional neural network gradually approaches the minimum value, and finally the accuracy of the automatically generated student network model is close to that of the complex teacher model, improving the accuracy of the student model.
[0092] In a possible embodiment, the structure prediction module 400 is configured to construct a structure prediction network based on the initial lightweight convolutional neural network, including: obtaining an initial prediction network; obtaining a total data set based on the initial lightweight convolutional neural network; randomly sampling the total data set based on a preset sampling ratio to obtain a sub-data set; and training the initial prediction network based on the sub-data set to obtain the structure prediction network.
[0093] Specifically, a sampling ratio ρ is set, and a sub-data set D is obtained by randomly sampling the total data set D 0 , such that |D 0 | = ρ * |D|. Use D 0 as the training data set of the structure prediction network.
[0094] In this embodiment, a sub-data set is obtained by random sampling, and the initial prediction network is trained based on the sub-data set to obtain the structure prediction network, reducing the computational overhead and training time while ensuring the accuracy of the obtained structure prediction network.
[0095] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0096] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0097] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as the protection scope of the present invention. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A student model automatic generation system based on knowledge distillation, characterized in that: It includes automatic generation module, model acquisition module, loss function construction module, structure prediction module and knowledge distillation module, among which: The model acquisition module is used to acquire an initial lightweight convolutional neural network based on a preset deep convolutional neural network; The loss function construction module is used to construct a lightweight convolutional neural network loss function of the initial lightweight convolutional neural network based on the deep convolutional neural network; The structure prediction module is used to construct a structure prediction network based on the initial lightweight convolutional neural network; The automatic generation module is used to receive the initial lightweight convolutional neural network and the lightweight convolutional neural network loss function, and then control the knowledge distillation module to perform several rounds of knowledge distillation training on the initial lightweight convolutional neural network, so as to generate a student model based on the result of any round of the knowledge distillation training; The knowledge distillation module is used to perform the knowledge distillation training, including: Training the initial lightweight convolutional neural network based on the lightweight convolutional neural network loss function to obtain a first lightweight convolutional neural network; Based on the structure prediction network and the first lightweight convolutional neural network, updating the lightweight convolutional neural network loss function; The initial lightweight convolutional neural network is updated based on the first lightweight convolutional neural network as a result of this round of knowledge distillation training.
2. According to the knowledge distillation-based student model automatic generation system of claim 1, it is characterized in that: In the knowledge distillation module, updating the lightweight convolutional neural network loss function based on the structure prediction network and the first lightweight convolutional neural network includes: Based on the first lightweight convolutional neural network, obtaining a parameter tuple set; For any parameter tuple in the parameter tuple set, obtaining a corresponding mask based on the structure prediction network; Based on the set of parameter tuples and all the masks, the lightweight convolutional neural network loss function is updated.
3. The student model automatic generation system based on knowledge distillation according to claim 2 is characterized in that: After the knowledge distillation module updates the lightweight convolutional neural network loss function based on the structure prediction network and the first lightweight convolutional neural network: The loss function construction module is also used to obtain the structure prediction network loss function based on all the masks and the lightweight convolutional neural network loss function, so that the knowledge distillation module updates the structure prediction network based on the structure prediction network loss function.
4. The student model automatic generation system based on knowledge distillation according to claim 3 is characterized in that: The structure prediction network loss function is specifically: in, Indicates that when the input is x and the output is y, the first lightweight convolutional neural network is The lightweight convolutional neural network loss function when the mask is w, P(w) represents the model operation amount after updating the lightweight convolutional neural network loss function based on the parameter tuple set and all the masks, γ is used to control the regularization strength.
5. The student model automatic generation system based on knowledge distillation according to claim 3 is characterized in that: The knowledge distillation module is preset with update rounds of the structure prediction network; for any round of the knowledge distillation training, if the knowledge distillation module determines that it does not belong to the update round, the structure prediction network will not be updated in this round of knowledge distillation training.
6. The student model automatic generation system based on knowledge distillation according to claim 2 is characterized in that: For any parameter tuple in the parameter tuple set, the knowledge distillation module is further configured to: Obtaining the importance of the corresponding convolutional layer based on the first lightweight convolutional neural network; Based on the parameter tuple set, all the masks, all the convolutional layer importances and a preset penalty threshold, the lightweight convolutional neural network loss function is updated.
7. The student model automatic generation system based on knowledge distillation according to claim 6 is characterized in that: The process of updating the lightweight convolutional neural network loss function is specifically as follows: Among them, L total represents the lightweight convolutional neural network loss function, represents the parameter tuple set, [w]g represents the mask corresponding to the g-th parameter tuple, represents the importance of the convolutional layer corresponding to the g-th parameter tuple, and δ represents the penalty threshold.
8. The student model automatic generation system based on knowledge distillation according to claim 1 is characterized in that: In the loss function construction module, the lightweight convolutional neural network loss function of the initial lightweight convolutional neural network is constructed based on the deep convolutional neural network, including: Based on the deep convolutional neural network and the initial lightweight convolutional neural network, a hard label loss function and a soft label loss function are obtained, and then based on the hard label loss function and the soft label loss function, the lightweight convolutional neural network loss function is constructed.
9. The student model automatic generation system based on knowledge distillation according to claim 8 is characterized in that: In the loss function construction module, the initial lightweight convolutional neural network is trained based on the lightweight convolutional neural network loss function to obtain a first lightweight convolutional neural network, including: Obtaining a hard label loss function gradient based on the hard label loss function, and obtaining a soft label loss function gradient based on the soft label loss function; Based on the hard label loss function gradient and the soft label loss function gradient, a lightweight convolutional neural network loss function gradient is constructed, so that the knowledge distillation module trains the initial lightweight convolutional neural network through a gradient descent algorithm based on the lightweight convolutional neural network loss function gradient, thereby obtaining a first lightweight convolutional neural network.
10. The student model automatic generation system based on knowledge distillation according to claim 1 is characterized in that: The structure prediction module is used to construct a structure prediction network based on the initial lightweight convolutional neural network, including: Get the initial prediction network; Acquire a total data set based on the initial lightweight convolutional neural network; Randomly sampling the total data set based on a preset sampling ratio to obtain a sub-data set; The initial prediction network is trained based on the sub-dataset to obtain the structure prediction network.