Model Training Method, Device, Equipment and Storage Medium Based on Knowledge Distillation
By calculating the representation similarity between the teacher network and the student network layer and constructing a similarity loss function, the problems of instability and non-convergence of the training process in knowledge distillation are solved, and the stability and convergence of model training are improved.
Patent Information
- Application Number
- CN202210816261.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-07-12
AI Technical Summary
The existing knowledge distillation method cannot effectively deal with the situation where the teacher network and student network are not matched very well, resulting in unstable and unconvergent model training process.
By obtaining the first model that satisfies the target condition and the second model that does not satisfies the target condition, calculate the representation similarity of the two network layers, construct the sum of the similarity loss function and the optimization loss function as the target loss function, and train the second model using the training set until the target loss function converges.
It improves the stability and convergence of the student network training process, and enhances the ability and scope of the second model to learn the first model.
Smart Images

Figure CN115062769B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a model training method, device, equipment and storage medium based on knowledge distillation. Background Art
[0002] At present, deep learning neural networks have been successfully applied to various computer vision applications, such as image classification, object detection, and semantic segmentation. Large-scale deep learning model training must be obtained from very large and highly redundant datasets. However, when the amount of data in the dataset is large, model training requires a large amount of time and storage space. Therefore, in order to shorten the training time and reduce resource occupancy, the use of knowledge distillation methods to compress large-scale deep learning networks has been widely used. The knowledge distillation method has high requirements for the matching degree between the teacher network and the student network. However, the current knowledge distillation method only optimizes the labels of the training sample set and cannot handle the situation where the matching degree between the teacher network and the student network is not high, resulting in unstable and non-convergent problems in the training process. Therefore, how to improve the training process of knowledge distillation to improve the stability and convergence of the student network training process has become an urgent problem to be solved. Summary of the Invention
[0003] Based on this, it is necessary to provide a model training method, device, equipment and storage medium based on knowledge distillation for the above technical problems to solve the problems of instability and non-convergence in the training process.
[0004] The first aspect of the embodiments of the present application provides a model training method based on knowledge distillation, and the method includes:
[0005] Obtain a first model that meets the target conditions and a second model that does not meet the target conditions. The first model includes M network layers, and the second model includes N network layers, where N and M are both integers greater than zero;
[0006] Update the initial loss function of the second model according to the optimization loss function constructed based on the output of the first model to obtain an updated second model;
[0007] Calculate the representation similarity between the M network layers in the first model and the N network layers in the updated second model respectively, and determine the target similarity through a preset selection condition;
[0008] Construct a similarity loss function according to the target similarity, and use the sum of the similarity loss function and the optimization loss function as the target loss function;
[0009] Train the second model using the training set until the target loss function converges to obtain a second model that meets the target conditions.
[0010] The second aspect of the embodiments of the present application provides a model training device based on knowledge distillation, the device includes:
[0011] An acquisition model module, configured to acquire a first model that meets the target condition and a second model that does not meet the target condition, the first model includes M network layers, the second model includes N network layers, and both N and M are integers greater than zero;
[0012] An update module, configured to update the initial loss function of the second model according to the optimization loss function constructed based on the output of the first model, and obtain the updated second model;
[0013] A target similarity determination module, configured to calculate the representation similarity between the M network layers in the first model and the N network layers in the updated second model respectively, and determine the target similarity through a preset selection condition;
[0014] A target loss function determination module, configured to construct a similarity loss function according to the target similarity, and use the sum of the similarity loss function and the optimization loss function as the target loss function;
[0015] A training module, configured to train the second model using a training set until the target loss function converges, and obtain a second model that meets the target condition.
[0016] In the third aspect, an embodiment of the present invention provides a computer device, the computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the model training method based on knowledge distillation as described in the first aspect.
[0017] In the fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the model training method based on knowledge distillation as described in the first aspect.
[0018] The beneficial effects of the present invention compared with the prior art are:
[0019] The present invention obtains a first model that meets the target conditions and a second model that does not meet the target conditions. The first model includes M network layers, and the second model includes N network layers. Both N and M are integers greater than zero. An optimized loss function is constructed based on the output of the first model to update the initial loss function of the second model, obtaining an updated second model. The representation similarities between the M network layers in the first model and the N network layers in the updated second model are calculated. Through a preset selection condition, a target similarity is determined. Based on the target similarity, a similarity loss function is constructed, and the sum of the similarity loss function and the optimized loss function is used as the target loss function. The second model is trained using a training set until the target loss function converges, obtaining a second model that meets the target conditions. Using the representation similarity between the network layers in the first model and the second model as part of the loss function in the second model makes full use of the information in the intermediate network layers, increasing the ability and scope of the second model to learn from the first model, thereby improving the stability and convergence during the training of the second model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 FIG. is a schematic diagram of an application environment of a model training method based on knowledge distillation provided by an embodiment of the present invention;
[0022] Figure 2 FIG. is a flowchart of a model training method based on knowledge distillation provided by an embodiment of the present invention;
[0023] Figure 3 FIG. is a schematic structural diagram of a model training device based on knowledge distillation provided by an embodiment of the present invention;
[0024] Figure 4 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0026] It should be understood that, as used in the specification of the present invention and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations.
[0027] It should also be understood that the term "and / or" as used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] As used in the specification of the present invention and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0029] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0030] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a particular feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0031] Embodiments of the present invention may acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use the knowledge to obtain the best results.
[0032] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0033] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0034] In order to illustrate the technical solution of the present invention, the following will be described by specific embodiments.
[0035] A model training method based on knowledge distillation provided by an embodiment of the present invention can be applied in an application environment such as Figure 1 wherein the client communicates with the server. The client includes but is not limited to computer devices such as a palm computer, a desktop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, and a personal digital assistant (PDA). The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0036] See Figure 2 , which is a schematic flowchart of a model training method based on knowledge distillation provided by an embodiment of the present invention. The above model training based on knowledge distillation can be applied to the server in Figure 1 The above server is connected to the corresponding client to provide model training services for the client. As shown in Figure 2 , the model training method based on knowledge distillation may include the following steps.
[0037] S201: Obtain a first model that meets the target conditions and a second model that does not meet the target conditions.
[0038] In step S201, the first model that meets the target condition and the second model that does not meet the target condition are deep learning convolutional neural network models. The first model that meets the target condition is obtained by training the first model using a training set. The first model includes M network layers, and the second model includes N network layers. Both N and M are integers greater than zero. The first model that meets the target condition is a trained deep learning model, and the second model that does not meet the target condition is an untrained deep learning model. The second model is used to learn the parameters in the first model to save training time. In this embodiment, the first model is the teacher model in knowledge distillation, and the second model is the student model in knowledge distillation. When obtaining the first model that meets the conditions, the sample data in the training set can be manually labeled first, and then the labeled sample data can be used to train the first model. This training process can be understood as a pre-training process. During the training process, the loss function of the first model can be used to calculate the loss value between the actual output of the first model and the labeled result, and backpropagation training can be performed according to the loss value until the loss converges to obtain the first model that meets the conditions. For example, the above first model is used for voiceprint recognition, and the actual output of the first model that meets the conditions can be a voiceprint feature vector.
[0039] It should be noted that the first model and the second model can be of the same type of neural network model, that is, the first model and the second model have the same network layer structure, or the first model and the second model can be of different types of neural network models, that is, the network layer structures of the first model and the second model are different. However, the number of network layers of the first model is less than that of the second model. That is, the model size and the number of model parameters of the first model are less than those of the second model. Since the second model has a large number of model parameters, it takes a long time to output the prediction result in actual applications. To improve the output speed of the prediction result and save computer resources, knowledge distillation can be performed on the large model to obtain a lightweight small model. Thus, in actual applications, the feature representation knowledge that the complex and strong-learning first network model has learned can be distilled out based on knowledge distillation of the model training and transmitted to the second model with fewer parameters and weaker learning ability.
[0040] It should be noted that in the embodiments of the present application, the sample data in the above training set can be the data obtained after preprocessing the recording data. For example, 40 hours of customer service recording data with voiceprint annotation can be amplified by adding noise, increasing the speech speed, and adding data perturbation to obtain a data set, and the data is divided according to the ratio of 8:2 for the training set and the test set. When dividing, the speaker information is fully considered to ensure that the speakers' voices in the training set and the test set are separated. The recording files in the training set are read to form a feature data combination of data labels (data-label). This feature data combination can be understood as the sample data in the training set.
[0041] S202: Update the initial loss function of the second model according to the optimized loss function constructed based on the output of the first model, and obtain the updated second model.
[0042] In step S202, the first model and the second model are network models with similar structures. The output result in the first model is used to guide the training of the second model, simplifying the training process of the second model. Among them, an optimized loss function is constructed according to the output of the first model, and the optimized loss function is used as the loss function in the training process of the second model to obtain the updated second model.
[0043] In this embodiment, the performance of the simple network is improved by using the output of a complex network with high performance and generalization ability and the real label data to train the simple network. Suppose there is a single or multiple models with complex networks and good performance, denoted as the first model, and there is a model with fewer network layers and low learning ability, denoted as the second model. Using the knowledge learned by training the first model as the training target of the second model, the updated second model is obtained. The performance of the updated second model is close to that of the first model. However, compared with the first model, the second model has fewer parameters and a shorter training time, thus achieving the compression and acceleration of the large model and improving the accuracy of the small model.
[0044] For example, the first model can be constructed based on the ResNet34 residual network; the second model can be constructed based on ResNet10. Since larger networks often face the problems of large and redundant deep learning network models and difficulty in meeting the real-time requirements of recognition speed, while small network models are very likely to have insufficient model feature representation ability due to fewer parameters, resulting in low model accuracy. The problem is that although it meets the real-time requirements of online applications, it cannot meet the accuracy requirements. Therefore, training the second model with the knowledge of the first model can enable the large network model to have a positive effect on the small network model, enabling the second model to obtain better fitting parameters and thus improving the accuracy of the second model.
[0045] It should be noted that since the number of network layers of the first model is greater than that of the second model, the knowledge distillation method can be adopted to determine the corresponding relationship between the network layers of the second network model and the network layers of the first model, so that the network layers of the second model learn to fit the corresponding network layers of the first model, that is, the network layers corresponding to the first network model are compressed into the network layers of the second model through knowledge distillation. Since the number of network layers of the first model is greater than that of the second model, the network layers of the first model correspond to the network layers of the second model at intervals. For example, when the first model includes 24 network layers and the second model includes 12 network layers, the first layer of the second model may correspond to the second layer of the first model, the second layer of the second model may correspond to the fourth layer of the first model, the third layer of the second model may correspond to the sixth layer of the first model, the fourth layer of the second model may correspond to the eighth layer of the first model, and so on.
[0046] Optionally, updating the initial loss function of the second model according to the optimization loss function constructed based on the output of the first model to obtain the updated second model includes:
[0047] Inputting the first training sample with the original label into the first model, outputting the new label data corresponding to the first training sample to obtain the second training sample;
[0048] Using the first training sample and the second training sample to train the second model respectively to obtain the first knowledge distillation loss function and the second knowledge distillation loss function;
[0049] Constructing an optimization loss function through the first knowledge distillation loss function and the second knowledge distillation loss function;
[0050] Updating the initial loss function of the second model according to the optimization loss function to obtain the updated second model.
[0051] In this embodiment, the true label of the original data is denoted as Hard - target, and the output probability of the first model is denoted as soft - target. Since the first model has many parameters and obtains more representation information, while the true label Hard - target contains very little representation information. For example, in a binary classification task, assume that the softmax output of the first model is [0.995, 0.005]. The negative - example output carries 0.005 sample information, which can show the similarity between the two classes, while the true label is only [1, 0], and the negative example being 0 does not contain useful information. By introducing the hyperparameter temperature T, a smoothing operation is performed on the soft - target. When T approaches 0, the largest value in the result will approach 1 and the other will be 0. When T is larger, the distributions of the two output results are flatter, and the gap between the two probability values is smaller, so that more retained similarity information is obtained. The output probability of the first model is used as the learning knowledge of the second model, thereby reducing the number of parameters and the computational complexity of the convolutional neural network and avoiding huge computational overhead.
[0052] It should be noted that it is also possible to use the output of the first model as the knowledge to train the second model, instead, using the intermediate - layer features of the first model input into the intermediate - layer features of the second model. This method allows the network layer of the second model to be more than that of the first model, but the number of neurons in the middle should be less than that of the first model.
[0053] Optionally, according to the optimization loss function constructed based on the output of the first model, update the initial loss function of the second model to obtain the updated second model, including:
[0054] Input the first training sample with the original label into the first model, and update the original label corresponding to the first training sample with the new label output by the first model to obtain the second training sample;
[0055] Use the first training sample and the second training sample to train the second model respectively to obtain the first knowledge - distillation loss function and the second knowledge - distillation loss function;
[0056] Construct an optimization loss function through the first knowledge - distillation loss function and the second knowledge - distillation loss function.
[0057] In this embodiment, when constructing the optimization loss function according to the first knowledge - distillation loss function and the second knowledge - distillation loss function, different parameters are randomly set for the first knowledge - distillation loss function and the second knowledge - distillation loss function, and the target parameters are obtained through the gradient - descent algorithm. In order to enable the second model to learn more guiding information from the first model, when setting the parameters, the parameters set for the first knowledge - distillation loss function are generally larger than the parameters corresponding to the second knowledge - distillation loss function.
[0058] It should be noted that during the gradient descent process, a suitable step size needs to be selected. A suitable step size can greatly reduce the number of iterations and even make it more convenient to converge to the global optimal solution. If the step size is too large during the gradient descent process, it may cause the gradient to directly skip the local minimum or even diverge. If the step size is too small, the convergence speed may be greatly reduced, and too much unnecessary time will be consumed during the convergence process.
[0059] S203: Calculate the representation similarity between the M network layers in the first model and the N network layers in the updated second model respectively, and determine the target representation similarity.
[0060] In step S203, the central kernel alignment algorithm is used to calculate the similarity between network layers, and the degree to which the second model learns to fit the first model is judged through the similarity. When the feature matrices output by two network layers with corresponding relationships are relatively similar, it can indicate that the parameters between the two network layers with corresponding relationships are relatively similar, that is, the network layers of the second model have successfully learned to fit the corresponding network layers of the first model. The representation similarity between the M network layers in the first model and the N network layers in the updated second model is determined as the target similarity through preset selection conditions.
[0061] In this embodiment, the linear CKA (centered kernel alignment) algorithm is used to calculate the similarity between different network layers. The linear CKA algorithm can determine the corresponding relationships between the hidden layers of neural networks trained with different random initializations and different widths. When the number of network layers in the first model and the second model is different, the relationship between each network layer in the first model and the second model is not one-to-one. It is necessary to find the corresponding relationship according to the similarity between different network layers in the first model and the second model, and this similarity is the target similarity.
[0062] Optionally, calculating the representation similarity between the M network layers in the first model and the N network layers in the updated second model respectively, and determining the target similarity through preset selection conditions includes:
[0063] Respectively obtain the feature matrices of each of the M network layers in the first model and the N network layers in the updated second model;
[0064] Calculate the representation similarity between the feature matrices of the M network layers in the first model and the feature matrices of the N network layers in the updated second model respectively, and obtain the representation similarity sequence corresponding to each network layer in the first model;
[0065] Determine the target representation similarity from the representation similarity sequence corresponding to each network layer in the first model.
[0066] In this embodiment, the first model and the updated second model are trained using the same training set. The data in the training set is input into the first model and the updated second model respectively, and the feature matrices output by each network layer of the second model and the feature matrices output by each network layer of the first model are obtained. The representation similarity between M network layers in the first model and N network layers in the updated second model is calculated to obtain a sequence of representation similarities corresponding to each network layer in the first model. The target similarity is determined from the sequence of representation similarities corresponding to each network layer in the first model through a preset selection condition.
[0067] It should be noted that when calculating the representation similarity, it can also be calculated layer by layer. For example, when the first layer of the second model corresponds to the third layer of the first model, the second layer corresponds to the sixth layer, the third layer corresponds to the ninth layer, and the fourth layer corresponds to the twelfth layer, the similarity between the feature matrix output by the first layer of the second model and the feature matrix output by the third layer of the first model can be calculated, the similarity between the feature matrix output by the second layer of the second model and the feature matrix output by the sixth layer of the first model can be calculated, the similarity between the feature matrix output by the third layer of the second model and the feature matrix output by the ninth layer of the first model can be calculated, and the similarity between the feature matrix output by the fourth layer of the second model and the feature matrix output by the twelfth layer of the first model can be calculated. The calculated similarity is the target similarity.
[0068] Optionally, determining the target similarity from the sequence of representation similarities corresponding to each network layer in the first model through a preset selection condition includes:
[0069] Obtain the maximum value of the representation similarity from the sequence of representation similarities corresponding to each network layer in the first model, and use the maximum value of the representation similarity as the target similarity.
[0070] In this embodiment, when the representation similarity is the largest, it is considered that the network layer in the second model learns more knowledge from the corresponding network layer in the first model, and features similar to those of the first model can be extracted.
[0071] S204: Construct a similarity loss function according to the target similarity, and use the sum of the similarity loss function and the optimization loss function as the target loss function.
[0072] In step S204, the similarity loss function is the difference value between the output features of the network layer learned after knowledge distillation of the second model and the output features of the corresponding network layer in the first model. The target loss function is calculated based on the similarity loss function and the optimization loss function. The target loss function is the sum of the differences between the output features of the network layers between the first model and the second model and the differences between the output features of the entire model.
[0073] In this embodiment, a similarity loss function is constructed based on the similarity between the first model and the second model. The similarity loss function is the difference value between the output features of the network layer learned after knowledge distillation of the second model and the output features of the corresponding network layer in the first model. When the similarity is greater, it is considered that the difference between the network layers of the first model and the second model is smaller. When the similarity is smaller, it is considered that the difference between the network layers of the first model and the second model is larger. Based on this, a similarity loss function is constructed. The weighted sum of the similarity loss function and the optimization loss function is used as the target loss function. Different loss functions can help the neural network model learn different knowledge. On the basis of the optimization loss function, the target loss function adds the similarity loss function, expanding the learning scope of the model.
[0074] Optionally, according to the target similarity, a similarity loss function is constructed, including:
[0075] Calculate the loss value of each network layer in the second model according to the target similarity;
[0076] Based on the loss value of each network layer in the second model, a similarity loss function is constructed.
[0077] In this embodiment, a similarity loss function is constructed based on the loss of each layer of the network in the second model. The value range of the representation similarity value in each network layer of the first model and the second model is (0, 1). When the size of the representation similarity is closer to 1, it is considered that the corresponding network layers of the first model and the second model are closer. When the size of the representation similarity is closer to 0, it is considered that the gap between the corresponding network layers of the first model and the second model is larger. Therefore, according to the target similarity, the corresponding network layer of each layer of the network in the second model in the first model is obtained, and the difference between the target similarity and 1 is used as the loss value of the network layer in the second model. Then, the similarity loss function in the second model is the sum of the loss values of each network layer in the second model.
[0078] S205: Use the training set to train the second model until the target loss function converges, and obtain the second model that meets the target conditions.
[0079] In step S205, when using the training set to train the second model, supervised training can be performed, and a corresponding threshold is set. When the difference between the training result and the supervision label is less than the threshold, it is considered that the target loss function converges, and the second model that meets the target conditions is obtained.
[0080] In this embodiment, when training the second model using the training set, for the target loss function in the second model, the training result and the label information in the training set are used to calculate the loss value through the target loss function, and it is determined whether the loss value meets the preset conditions. When the preset conditions are not met, the second model is updated by backpropagation according to the loss value to obtain the second model with updated model parameters. Then, the second model with updated model parameters is trained again based on the training set until the loss value meets the preset conditions, and the second model that has met the target conditions is obtained.
[0081] Optionally, training the second model using the training set until the target loss function converges to obtain the second model that meets the target conditions includes:
[0082] Construct sample pairs according to the positive and negative samples in the training set; the sample pairs include at least one positive sample and one negative sample;
[0083] Train the second model based on the sample pairs until the target loss function converges to obtain the second model that meets the target conditions.
[0084] In this embodiment, when training the second model using the training set, to accelerate the convergence speed, sample pairs can be used for training. The same positive sample and negative sample are selected for each sample pair. Since the magnitude of the sample backpropagation gradient is determined by the accumulation of the differences between each positive and negative sample pair, once the number of samples in the sample pair is too large, it will be difficult to find the training separation surface that satisfies all sample pairs in the feature space, and the convergence of training will deteriorate accordingly. At the same time, repeatedly calculating the distance difference between two samples as the backpropagation gradient each time will lead to redundancy in training and an increase in training time. Since the size comparison of a pair of positive and negative samples in different sample pairs may occur multiple times, increasing the number of positive and negative samples in the sample pair too much has little help for training. Therefore, one positive sample and one negative sample can be selected for each sample pair. Use the selected sample pairs to train the second model until the target loss function converges to obtain the second model that meets the target conditions.
[0085] In the present invention, a first model that meets the target conditions and a second model that does not meet the target conditions are obtained. The first model includes M network layers, and the second model includes N network layers, where both N and M are integers greater than zero. An optimized loss function is constructed based on the output of the first model to update the initial loss function of the second model, obtaining an updated second model. The representation similarities between the M network layers in the first model and the N network layers in the updated second model are calculated, and a target similarity is determined through a preset selection condition. Based on the target similarity, a similarity loss function is constructed, and the sum of the similarity loss function and the optimized loss function is used as the target loss function. The second model is trained using a training set until the target loss function converges, obtaining a second model that meets the target conditions. Using the representation similarity between the network layers in the first model and the second model as part of the loss function in the second model makes full use of the information in the intermediate network layers, increasing the ability and scope of the second model to learn from the first model. When training the second model, the stability and convergence of the second model can be improved.
[0086] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a model training device based on knowledge distillation provided by an embodiment of the present invention. In this embodiment, each unit included in the terminal is used to execute Figure 2 the corresponding steps in the corresponding embodiment. Specifically, please refer to Figure 2 and Figure 2 the relevant descriptions in the corresponding embodiments. For the sake of simplicity, only the parts related to this embodiment are shown.
[0087] See Figure 3 , the model training device 30 includes:
[0088] A model acquisition module 31, configured to acquire a first model that meets the target conditions and a second model that does not meet the target conditions. The first model includes M network layers, and the second model includes N network layers, where both N and M are integers greater than zero;
[0089] An update module 32, configured to update the initial loss function of the second model according to the optimized loss function constructed based on the output of the first model, obtaining an updated second model;
[0090] A target similarity determination module 33, configured to calculate the representation similarities between the M network layers in the first model and the N network layers in the updated second model respectively, and determine a target similarity through a preset selection condition;
[0091] A target loss function determination module 34, configured to construct a similarity loss function based on the target similarity, and use the sum of the similarity loss function and the optimized loss function as the target loss function;
[0092] The training module 35 is used to train the second model using the training set until the target loss function converges, obtaining a second model that meets the target conditions.
[0093] Optionally, the above update module 32 includes:
[0094] The second training sample acquisition unit is used to input the first training sample with the original label into the first model, and update the original label corresponding to the first training sample with the new label output by the first model to obtain a second training sample;
[0095] The knowledge distillation loss function acquisition unit is used to train the second model using the first training sample and the second training sample respectively to obtain a first knowledge distillation loss function and a second knowledge distillation loss function;
[0096] The optimized loss function acquisition unit is used to construct an optimized loss function through the first knowledge distillation loss function and the second knowledge distillation loss function;
[0097] The updated second model acquisition unit is used to update the initial loss function of the second model according to the optimized loss function to obtain an updated second model.
[0098] Optionally, the above optimized loss function acquisition unit includes:
[0099] The initial loss function acquisition subunit is used to set different initial parameters for the first knowledge distillation loss function and the second knowledge distillation loss function to obtain an initial distillation loss function;
[0100] The construction subunit is used to update the parameters of the initial distillation loss function using the gradient descent algorithm to obtain target parameters, and update the initial distillation loss function using the target parameters to obtain an optimized loss function.
[0101] Optionally, the above target similarity determination module 33 includes:
[0102] The feature matrix acquisition unit is used to respectively acquire the feature matrices of each network layer in M network layers of the first model and N network layers of the updated second model;
[0103] The representation similarity sequence acquisition unit is used to
[0104] calculate the representation similarities between the feature matrices of M network layers in the first model and the feature matrices of N network layers in the updated second model respectively, obtaining a representation similarity sequence corresponding to each network layer in the first model;
[0105] The target similarity acquisition unit is used to determine the target similarity from the representation similarity sequences corresponding to each network layer in the first model through a preset selection condition.
[0106] Optionally, the above-mentioned target similarity obtaining unit includes:
[0107] A target similarity determination subunit, configured to obtain the maximum value of the representation similarity from the representation similarity sequences corresponding to each network layer in the first model, and use the maximum value of the representation similarity as the target similarity.
[0108] Optionally, the above-mentioned target loss function determination module 34 includes:
[0109] A loss value determination unit for each network layer, configured to calculate the loss value of each network layer in the second model according to the target similarity;
[0110] A similarity loss function construction unit, configured to construct a similarity loss function based on the loss values of each network layer in the second model.
[0111] Optionally, the above-mentioned training module 35 includes:
[0112] A sample pair construction unit, configured to construct sample pairs according to the positive and negative samples in the training set; the sample pairs include at least one positive sample and one negative sample;
[0113] A second model obtaining unit that meets the target conditions, configured to train the second model based on the sample pairs until the target loss function converges, and obtain a second model that meets the target conditions.
[0114] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in any of the above-mentioned embodiments of the model training method based on knowledge distillation.
[0115] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 4 this is only an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include a network interface, a display screen, and an input device, etc.
[0116] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0117] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store the data that has been output or will be output.
[0118] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiment of the present invention, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiment can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0119] All or part of the processes in the above method embodiment of the present invention can also be completed by a computer program product. When the computer program product runs on a computer device, the computer device can be made to execute the steps in the above method embodiment.
[0120] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0121] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0122] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0123] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A model training method based on knowledge distillation, characterized in that, it includes: Obtain a first model that meets the target conditions and a second model that does not meet the target conditions. The first model includes M network layers, and the second model includes N network layers. Both N and M are integers greater than zero; Input the first training sample with the original label into the first model, and update the original label corresponding to the first training sample with the new label output by the first model to obtain a second training sample; Use the first training sample and the second training sample to train the second model respectively to obtain a first knowledge distillation loss function and a second knowledge distillation loss function; Set different initial parameters for the first knowledge distillation loss function and the second knowledge distillation loss function to obtain an initial distillation loss function; Use the gradient descent algorithm to update the parameters of the initial distillation loss function to obtain target parameters, and use the target parameters to update the initial distillation loss function to obtain an optimized loss function; According to the optimized loss function, update the initial loss function of the second model to obtain an updated second model; Use the same training set to train the first model and the updated second model. Input the data in the training set into the first model and the updated second model respectively, obtain the feature matrices output by each network layer of the second model, and the feature matrices output by each network layer of the first model, calculate the representation similarity between the M network layers in the first model and the N network layers in the updated second model respectively, and determine the target similarity through a preset selection condition; Construct a similarity loss function according to the target similarity, and use the sum of the similarity loss function and the optimized loss function as the target loss function; Use the training set to train the second model until the target loss function converges to obtain a second model that meets the target conditions. The training set is the data obtained after preprocessing the audio data.
2. The model training method based on knowledge distillation according to claim 1, characterized in that, The calculation of the representation similarity between the M network layers in the first model and the N network layers in the updated second model respectively, and determining the target similarity through a preset selection condition includes: Respectively obtain the feature matrices of each network layer among the M network layers in the first model and the N network layers in the updated second model; Calculate the representation similarity between the feature matrices of the M network layers in the first model and the feature matrices of the N network layers in the updated second model respectively to obtain a sequence of representation similarities corresponding to each network layer in the first model; Determine the target similarity from the sequence of representation similarities corresponding to each network layer in the first model through a preset selection condition.
3. The model training method based on knowledge distillation according to claim 2, characterized in that, The determination of the target similarity from the sequence of representation similarities corresponding to each network layer in the first model through a preset selection condition includes: Obtain the maximum representation similarity value from the representation similarity value sequences corresponding to each network layer in the first model, and use the maximum representation similarity value as the target similarity value.
4. The model training method based on knowledge distillation according to claim 1, wherein, constructing a similarity loss function according to the target similarity value includes: calculating the loss value of each network layer in the second model according to the target similarity value; constructing a similarity loss function based on the loss values of each network layer in the second model.
5. The model training method based on knowledge distillation according to claim 1, wherein, training the second model using the training set until the target loss function converges to obtain a second model that meets the target conditions, including: constructing sample pairs according to the positive and negative samples in the training set; the sample pairs include at least one positive sample and one negative sample; training the second model based on the sample pairs until the target loss function converges to obtain a second model that meets the target conditions.
6. A model training device based on knowledge distillation, wherein, the device includes: a model acquisition module, configured to acquire a first model that meets the target conditions and a second model that does not meet the target conditions, the first model includes M network layers, the second model includes N network layers, and both N and M are integers greater than zero; an update module, configured to input a first training sample with an original label into the first model, and update the original label corresponding to the first training sample with the new label output by the first model to obtain a second training sample; training the second model using the first training sample and the second training sample respectively to obtain a first knowledge distillation loss function and a second knowledge distillation loss function; setting different initial parameters for the first knowledge distillation loss function and the second knowledge distillation loss function to obtain an initial distillation loss function; updating the parameters of the initial distillation loss function using the gradient descent algorithm to obtain target parameters, and updating the initial distillation loss function using the target parameters to obtain an optimized loss function; updating the initial loss function of the second model according to the optimized loss function to obtain an updated second model; a target similarity value determination module, configured to train the first model and the updated second model using the same training set, input the data in the training set into the first model and the updated second model respectively, acquire the feature matrices output by each network layer of the second model and the feature matrices output by each network layer of the first model, calculate the representation similarity between the M network layers in the first model and the N network layers in the updated second model respectively, and determine the target similarity value through a preset selection condition; a target loss function determination module, configured to construct a similarity loss function according to the target similarity value, and use the sum of the similarity loss function and the optimized loss function as the target loss function. A training module, configured to train the second model using a training set until the target loss function converges, so as to obtain a second model that meets the target condition, where the training set is data obtained after preprocessing the recording data.
7. A computer device, characterized in that the computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the knowledge distillation-based model training method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the knowledge distillation-based model training method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Knowledge distillation method and device
CN109637546A
Language recognition method based on language model and text classification method and device
CN111554268A