Model Distillation Method, Device, Equipment and Storage Medium Based on Multi-Teacher Model

Through the multi-teacher model distillation method of dynamically selecting teacher models and assigning weights, the problem of poor performance of student models in the existing technology is solved, and the best performance of student models and user experience is improved.

CN114386604BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210044224.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-27
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

The existing model distillation method based on teacher models has resulted in poor performance of student models and the fixed weighting method cannot dynamically and adequately distillate effective knowledge to student models.

Method used

Using a model distillation method based on multi-teacher model, the teacher model is dynamically selected and the most appropriate weight ratio is assigned to multiple teacher models so that multiple teacher models can distillate effective knowledge to the student model. The method includes obtaining training sample data and hard labels, generating soft labels through multiple teacher models and student models, performing knowledge distillation learning, and updating teacher model selection strategies until the student model converges.

Benefits of technology

By dynamically selecting the teacher model and assigning weights, the student model obtained by distillation can be performed best, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114386604B_ABST
    Figure CN114386604B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence, and particularly to a model distillation method based on a multi-teacher model. The method includes: obtaining training sample data and corresponding hard labels, and identifying the training sample data through multiple teachers and a first student model to obtain multiple first soft labels and a second soft label; according to the hard label, the multiple first soft labels and the second soft label, performing knowledge distillation learning on the first student model to generate a second student model and obtaining a third soft label through the second student model; updating the teacher model selection strategy according to the first soft label and the third soft label to obtain an updated teacher model selection strategy, re-determining the corresponding teacher model based on the updated teacher model selection strategy, and performing knowledge distillation learning on the first student model according to the re-determined teacher model until the first student model converges to obtain a target student model. This can enable the performance of the distilled student model to reach the best and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a model distillation method based on a multi-teacher model, a model distillation device based on a multi-teacher model, a computer device and a storage medium. Background Art

[0002] The existing models have a huge number of parameters, which poses a huge challenge to technical personnel in fine-tuning and online deployment. For example, the BERT-base model has 110 million parameters and the BERT-large model has 340 million parameters. The massive number of parameters makes these models slow in fine-tuning and deployment, with high computational costs, causing great delays and capacity limitations for real-time applications. Therefore, model compression is of great significance.

[0003] As one of the three major methods of model compression, model distillation has been widely recognized and applied in academia and industry, and more and more distillation methods have been proposed and applied. The commonly used distillation method is based on a framework of a teacher model and a student model, which distills the knowledge learned by a single complex teacher model into a simple student model. To a certain extent, it maintains the reasoning speed while effectively improving the accuracy of the student model reasoning. However, it is not necessary to use a single complex teacher model to improve the accuracy of the student model. Summary of the invention

[0004] The present application provides a model distillation method based on a multi-teacher model, a model distillation device based on a multi-teacher model, a computer device and a storage medium, aiming to solve the problem of poor performance of the existing student model obtained by distillation based on the teacher model.

[0005] To achieve the above objectives, the present application provides a model distillation method based on a multi-teacher model, the method comprising:

[0006] Acquire training sample data and hard labels corresponding to the training sample data, identify the training sample data through multiple teacher models to obtain multiple first soft labels, and identify the training sample data through a first student model to obtain a second soft label;

[0007] Performing knowledge distillation learning on the first student model according to the hard label, the plurality of first soft labels and the second soft label to generate a second student model; wherein a model parameter of the first student model is different from a model parameter of the second student model;

[0008] Identify the training sample data using the second student model to obtain a third soft label;

[0009] updating the teacher model selection strategy according to the first soft label and the third soft label to obtain an updated teacher model selection strategy, wherein the teacher model selection strategy is used to select a teacher model;

[0010] Based on the updated teacher model selection strategy, the corresponding teacher model is re-determined, and knowledge distillation learning is performed on the first student model according to the re-determined teacher model until the first student model converges to obtain a target student model.

[0011] To achieve the above object, the present application also provides a model distillation device based on a multi-teacher model, and the model distillation device based on the multi-teacher model includes:

[0012] A first label generation module is used to obtain training sample data and hard labels corresponding to the training sample data, identify the training sample data through multiple teacher models to obtain multiple first soft labels, and identify the training sample data through a first student model to obtain a second soft label;

[0013] a model generation module, configured to perform knowledge distillation learning on the first student model according to the hard label, the plurality of first soft labels and the second soft label to generate a second student model; wherein the model parameters of the first student model and the model parameters of the second student model are different;

[0014] A second label generation module, used for identifying the training sample data through the second student model to obtain a third soft label;

[0015] a strategy updating module, configured to update the teacher model selection strategy according to the first soft label and the third soft label to obtain an updated teacher model selection strategy, wherein the teacher model selection strategy is used to select a teacher model;

[0016] A model determination module is used to redetermine the corresponding teacher model based on the updated teacher model selection strategy, and perform knowledge distillation learning on the first student model according to the redetermined teacher model until the first student model converges to obtain a target student model.

[0017] In addition, to achieve the above-mentioned purpose, the present application also provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement any one of the model distillation methods based on a multi-teacher model provided in the embodiments of the present application when executing the computer program.

[0018] In addition, to achieve the above-mentioned purpose, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor enables the processor to implement any one of the model distillation methods based on a multi-teacher model provided in the embodiments of the present application.

[0019] The model distillation method based on multiple teacher models, the model distillation device based on multiple teacher models, the equipment and the storage medium disclosed in the embodiments of the present application screen the teacher model through the performance of the student model obtained by distillation and dynamically select the teacher model, thereby enabling multiple teacher models to distill effective knowledge to the student model, so that the performance of the student model obtained by distillation reaches the best, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 It is a scenario diagram of a model distillation method based on a multi-teacher model provided in an embodiment of the present application;

[0022] Figure 2 It is a flow chart of a model distillation method based on a multi-teacher model provided in an embodiment of the present application;

[0023] Figure 3 is a schematic block diagram of a model distillation device based on a multi-teacher model provided in one embodiment of the present application;

[0024] Figure 4 It is a schematic block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0026] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor do they have to be executed in the order described. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions. In addition, although the functional modules are divided in the device schematic, in some cases, the module division can be different from that in the device schematic.

[0027] The term "and / or" as used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] The commonly used distillation method is based on a framework of a teacher model and a student model. The knowledge learned by a single complex teacher model is distilled into a simple student model, which effectively improves the accuracy of student model reasoning while maintaining the reasoning speed to a certain extent. However, it is not necessary to use a single complex teacher model to improve the accuracy of the student model. Choosing a more suitable teacher model with weaker performance for distillation may achieve better results than choosing a stronger teacher model.

[0029] Based on this idea, the distillation method based on multi-teacher model came into being. However, most of the current multi-teacher model distillation methods use fixed weights to distribute the influence of each teacher on the student, which cannot dynamically and fully distill effective knowledge to the student model, which may result in the performance of the distilled student model not being optimal.

[0030] To solve the above problems, the present application provides a model distillation method based on multiple teacher models, which can be applied in servers and, of course, can also be applied to terminal devices. It can dynamically select teacher models and assign the most appropriate weight ratios to multiple teacher models, thereby enabling multiple teacher models to distill effective knowledge to student models, so that the distilled student model performs best and improves user experience.

[0031] The terminal device may include a fixed terminal such as a mobile phone, a tablet computer, a personal digital assistant (PDA), etc. The server may be, for example, a separate server or a server cluster. However, for ease of understanding, the following embodiments will introduce the model distillation method based on the multi-teacher model applied to the server in detail.

[0032] For example, the model distillation method based on multiple teacher models proposed in the embodiment of the present application can be applied to application scenarios such as question answering (determining whether a question matches an answer) and sentence matching (whether two sentences express the same meaning). By continuously and dynamically selecting teacher models and determining the weight ratios corresponding to multiple teacher models, the student model obtained by distillation can be faster and more accurate in question answering and sentence matching.

[0033] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0034] like Figure 1 As shown, the model distillation method based on the multi-teacher model provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. The application environment includes a terminal device 110 and a server 120, wherein the terminal device 110 can communicate with the server 120 through a network. Specifically, the server 120 performs multiple knowledge distillation learning on multiple teacher models to obtain a target student model, and sends the target student model to the terminal device 110 so that the user can use the target student model through the terminal device 110. Among them, the server 120 can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal device 110 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application is not limited here.

[0035] See also Figure 2 , Figure 2 This is a schematic flow chart of a model distillation method based on a multi-teacher model provided in an embodiment of the present application. The model distillation method based on a multi-teacher model can be applied to a server, thereby enabling multiple teacher models to distill effective knowledge to a student model, so that the distilled student model performs optimally.

[0036] like Figure 2 As shown, the model distillation method based on the multi-teacher model includes steps S101 to S105.

[0037] S101, obtaining training sample data and hard labels corresponding to the training sample data, identifying the training sample data through multiple teacher models to obtain multiple first soft labels, and identifying the training sample data through a first student model to obtain a second soft label.

[0038] The training sample data is sample data used to train the parameters of the student model, which may be different sentence pairs or different pictures, etc. The teacher model is a complex model with excellent reasoning performance, and the student model is a concise and low-complexity model.

[0039] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0040] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0041] Specifically, the training sample data is identified by establishing a neural network with softmax as the loss function. The hard label corresponding to the training sample data is the real label corresponding to the training sample data. After being identified by the neural network with softmax as the loss function, it is represented as 0 or 1, 0 represents that the training sample data does not belong to a certain category, and 1 represents that the training sample data belongs to a certain category.

[0042] Specifically, the training sample data can be identified by all teacher models to obtain multiple first soft labels, and the training sample data can be identified by the first student model to obtain a second soft label. The first soft label is the probability corresponding to the selected teacher model for the training sample data belonging to a certain category, and its value is between 0 and 1. The first student model is a student model that has not undergone knowledge distillation, and the second soft label is the probability corresponding to the first student model for the training sample data belonging to a certain category.

[0043] It should be noted that each training sample data corresponds to a hard label, each teacher model corresponds to a first soft label for each training sample data, and the first student model corresponds to a second soft label for each training sample data.

[0044] For example, suppose there is a classification task about cars, and the training sample data needs to be classified, such as brand A cars, brand B cars, and bicycles, and the training sample data is brand B cars, then the corresponding hard label is [0, 1, 0]. If the training sample data is identified by the teacher model, the corresponding first soft label may be [0.09, 0.9, 0.01]. If the training sample data is identified by the first student model, the corresponding second soft label may be [0.4, 0.5, 0.1].

[0045] S102. Perform knowledge distillation learning on the first student model according to the hard label, the multiple first soft labels and the second soft label to generate a second student model; wherein the model parameters of the first student model are different from the model parameters of the second student model.

[0046] Among them, knowledge distillation refers to a method of using a teacher model with higher accuracy but complex structure to guide the training of a student model with lower accuracy but simpler structure. The model parameters may include hyperparameters, the number of model layers, the number of model parameters, etc.

[0047] Exemplarily, knowledge distillation learning can be performed on the first student model using the hard label, the multiple first soft labels, and the second soft label to initialize the model parameters of the first student model and generate the second student model.

[0048] In some embodiments, a first deviation value is determined based on the multiple first soft labels and the second soft label; a second deviation value is determined based on the hard label and the second soft label; a third deviation value is determined based on the first deviation value and the second deviation value, and the model parameters of the first student model are initialized according to the third deviation value to generate a second student model. The model parameters of the first student model are different from the model parameters of the second student model, the first deviation value is the average loss value of the multiple first soft labels and the second soft labels, the second deviation value is the loss value of the hard label and the second soft label, and the third deviation value is the comprehensive loss value adjusted according to the first deviation value and the second deviation value. In this way, the student model can be subjected to knowledge distillation through multiple teacher models, thereby initializing the model parameters of the student model.

[0049] Specifically, the loss function constructed by determining the first deviation value according to the multiple first soft labels and the second soft labels is:

[0050]

[0051] Where k is the total number of selected teacher models; y i,k,cSample x predicted by the k-th teacher model i The probability of belonging to the cth class, that is, the first soft label corresponding to the kth teacher model; p s (y i =c|x i θ s ) is the sample x predicted by the first student model j The probability of belonging to the cth class, that is, the second soft label corresponding to the first student model, where θ s are the model parameters of the student model.

[0052] Specifically, the loss function constructed by determining the second deviation value according to the hard label and the second soft label is:

[0053]

[0054] Among them, N[y i =c] is the predicted sample x i The corresponding hard tag.

[0055] In some embodiments, based on the reverse gradient propagation algorithm, the weight ratio corresponding to the first deviation value and the second deviation value is determined; and the third deviation value is determined according to the first deviation value and the second deviation value and the corresponding weight ratio. Among them, the reverse gradient propagation algorithm is a method for training an artificial neural network, which calculates the gradient of the loss function for all weights in the network. This gradient will be fed back to the optimization method to update the weights to minimize the loss function. In this way, the weight ratio can be adjusted to balance the two types of loss functions, thereby minimizing the comprehensive loss value, achieving the effect of initializing the model parameters of the first student model and generating the second student model.

[0056] Specifically, the loss function constructed by determining the third deviation value according to the first deviation value and the second deviation value is:

[0057] l KD =αl DL +(1-α)l CE

[0058] Among them, l KD is the third deviation value, l DL is the first deviation value, l CE is the second deviation value, α is the hyperparameter for balancing the two types of loss functions. By adjusting the value of α, the student model's attention to soft and hard labels during the distillation process can be adjusted, thereby minimizing the comprehensive loss value and achieving the effect of initializing the model parameters of the first student model to generate the second student model.

[0059] When the neural network uses only hard labels (i.e., α = 0), the information of the original data is lost, which reduces the difficulty of fitting the model to the data, making the model easier to fit and prone to overfitting, resulting in a decrease in the generalization ability of the model. When soft labels are used (i.e., α≠0), the model needs to learn more knowledge, such as learning the similarities and differences between two close probabilities, thereby enhancing the generalization ability of the model.

[0060] By minimizing the comprehensive loss value, the first student model can extract knowledge from all teachers on average, achieve the effect of initializing the model parameters of the first student model, and generate the second student model.

[0061] S103: Identify the training sample data using the second student model to obtain a third soft label.

[0062] The third soft label is the probability corresponding to the second student model that the training sample data belongs to a certain category.

[0063] Specifically, the training sample data can be identified by the second student model to obtain a third soft label, thereby quickly understanding the recognition accuracy of the second student model, and quickly determining the learning distillation of the second student model, so as to prepare for the subsequent update of the teacher model selection strategy.

[0064] S104. Update the teacher model selection strategy according to the first soft label and the third soft label to obtain an updated teacher model selection strategy, where the teacher model selection strategy is used to select a teacher model.

[0065] Among them, the teacher model selection strategy is used to select the strategy corresponding to the teacher model.

[0066] In some embodiments, a fourth deviation value is determined based on the hard label and the third soft label; a state vector parameter is generated based on the first soft label and the fourth deviation value; an update strategy is determined based on the state vector parameter corresponding to each of the teacher models; and the teacher model selection strategy is updated based on the update strategy to obtain an updated teacher model selection strategy. Thus, the teacher model selection strategy can be updated based on the performance of the teacher model and the student model with updated model parameters.

[0067] Among them, the fourth deviation value is the loss value of the hard label and the third soft label, and the state vector parameter includes the first soft label and the fourth deviation value, which is used to represent the state corresponding to each iteration in the update iteration process. Specifically, the state vector parameter can be represented by a word vector. The word vector is a common term in the field of natural language. In fact, it is a vector representation of a word. Since the computer cannot directly calculate the word that has not been processed, it is represented in the form of a word vector. The update strategy is used to update the teacher model selection strategy, which can specifically include a screened teacher model, etc.

[0068] Specifically, the loss function constructed by determining the fourth deviation value according to the hard label and the third soft label is:

[0069]

[0070] After the state vector parameters are generated, an update strategy is determined according to the state vector parameters corresponding to each of the teacher models; the teacher model selection strategy is updated based on the update strategy to obtain an updated teacher model selection strategy so as to reselect the corresponding teacher model.

[0071] In some embodiments, based on a threshold function, the scores of each teacher model in the teacher model selection strategy are calculated according to the state vector parameters corresponding to each teacher model; the teacher models in the teacher model selection strategy are screened according to the scores of each teacher model to obtain the screened teacher model; and the update strategy is determined based on the screened teacher model. Wherein, the threshold function is a sigmoid function of the trainable teacher model selection strategy parameter θ. The Sigmoid function is an S-type function used to map variables to between [0, 1], and the score is a value between [0, 1].

[0072] Among them, it can be specifically expressed by the formula:

[0073] π(s j , a j )=a j σ(AF(s j )+b)+(1-a j )(1-σ(AF(s j )+b))

[0074] Among them, π(s j , a j ) is used to represent whether to select or not select the jth teacher model. Specifically, when π(s j , a j )=1, select the jth teacher model; when π(s j , a j)=0, the jth teacher model is not selected. j is the jth state, a j is a constant of 0 or 1, which is determined according to the distribution of the state vector parameters; σ(AF(s j )+b) is the sigmoid function of the trainable teacher model selection strategy parameter θ, which is used to calculate the score of each teacher model in the teacher model selection strategy, where F(s j ) is the state vector parameter, A and B can be arbitrary parameters, which are determined through experiments or experience.

[0075] Specifically, the state vector parameters corresponding to each teacher model are mapped using a sigmoid function to obtain a corresponding score, and it is determined whether the score exceeds a preset score threshold; if the score exceeds the preset score threshold, the teacher model corresponding to the score is screened out to obtain a screened teacher model; if the score does not exceed the preset score threshold, the teacher model corresponding to the score is not screened out. The preset score threshold can be any value and is not specifically limited here.

[0076] For example, if the preset score threshold is 0.7, the score corresponding to model A is 0.8, the score corresponding to model B is 0.6, and the score corresponding to model C is 0.75, then the teacher models obtained after screening are model A and model C.

[0077] In some embodiments, the second student model is tested by a test set to obtain the accuracy of the second student model; a corresponding feedback strategy is determined according to the accuracy of the second student model; and the teacher model selection strategy is updated based on the feedback strategy and the update strategy to obtain an updated teacher model selection strategy.

[0078] The feedback strategies include three types, which can be expressed by the formula:

[0079]

[0080] Among them, γ is a model hyperparameter, which indicates the degree of attention of the two parts, which is determined by the user. D is the accuracy of the second student model.

[0081] Specifically, a corresponding feedback formula can be selected according to the actual situation to obtain a corresponding feedback strategy. The corresponding feedback strategy can also be determined according to the accuracy of the second student model. For example, when the accuracy of the second student model is less than the preset first accuracy threshold, select -1. CE This feedback strategy; when the accuracy of the second student model is not less than the preset first accuracy threshold and not greater than the preset second accuracy threshold, select -l CE-l DL This feedback strategy; when the accuracy of the second student model is greater than the preset first accuracy threshold, select γ*(-l CE -l DL )+(1-γ)*acc D This feedback strategy. Wherein, the first accuracy threshold is less than the second accuracy threshold, and the specific value is not specifically limited here.

[0082] Specifically, updating the teacher model selection strategy based on the feedback strategy and the update strategy can be expressed by the formula:

[0083]

[0084] Among them, θ is the teacher model selection strategy parameter, and updating the teacher model selection strategy is actually updating the teacher model selection strategy parameter; lr is the learning rate, which can be pre-set or adjusted according to the training process; is the vector differential operator and can be used to refer to the gradient operator.

[0085] Specifically, in the first iteration, according to the corresponding feedback strategy and update strategy, the teacher model selection strategy parameters can be updated by multiplying by a preset learning rate, thereby realizing the update of the teacher model selection strategy. In multiple iterations, the feedback strategy obtained in each iteration is accumulated, the gradient of each screened teacher model is calculated and accumulated, and finally multiplied by a preset learning rate, so that the teacher model selection strategy parameters can be dynamically updated after multiple iterations, thereby realizing the dynamic update of the teacher model selection strategy.

[0086] S105. Re-determine the corresponding teacher model based on the updated teacher model selection strategy, and perform knowledge distillation learning on the first student model according to the re-determined teacher model until the first student model converges to obtain a target student model.

[0087] The target student model is the best performing student model obtained by distilling the teacher model corresponding to the teacher model selection strategy after multiple iterations of updating the teacher model selection strategy. The convergence means that the model obtained by distillation no longer produces large fluctuations.

[0088] In some embodiments, knowledge distillation learning is performed on the first student model according to the re-determined teacher model to generate a third student model; a third deviation value corresponding to the third student model is obtained to determine whether the third deviation value meets the convergence condition; if the third deviation value meets the convergence condition, the third student model is used as the target student model.

[0089] Specifically, the training sample data is identified through the re-determined teacher model to re-obtain multiple first soft labels; based on the hard labels, the re-obtained multiple first soft labels and the second soft labels, the first student model is subjected to knowledge distillation learning to generate a third student model; wherein the model parameters of the third student model are different from those of the first student model and the second student model.

[0090] Similarly, a first deviation value is determined according to the plurality of first soft labels and the second soft label; a second deviation value is determined according to the hard label and the second soft label; a third deviation value is determined according to the first deviation value and the second deviation value, and the model parameters of the first student model are initialized according to the third deviation value to generate a third student model. Since the teacher model is re-determined, the third deviation value corresponding to the student model generated in each iteration is different.

[0091] Whether the third deviation value satisfies the convergence condition can be determined by determining whether the third deviation value is less than a preset deviation value; if the third deviation value is less than the preset deviation value, it is determined that the third deviation value satisfies the convergence condition, that is, the third student model can be used as the target student model; if the third deviation value is not less than the preset deviation value, it is necessary to re-update the teacher model selection strategy to adjust the model parameters of the student model obtained by distillation. The preset deviation value can be any value and is not specifically limited here, so that the comprehensive loss value can be controlled within a very small range, so that the student model obtained by distillation no longer has mutations.

[0092] Specifically, obtain the third deviation value corresponding to the third student model, and determine whether the third deviation value meets the convergence condition; if the third deviation value meets the convergence condition, use the third student model as the target student model; if the third deviation value does not meet the convergence condition, determine the loss degree of the third deviation value, and adjust the parameters in the third student model according to the loss degree.

[0093] For example, the greater the loss, the greater the adjustment of the parameters in the preset student model; the smaller the loss, the smaller the adjustment of the parameters in the preset student model. In this way, by adjusting the preset student model based on the loss value, a greater degree of adjustment can be made when the error degree of the student model is greater, thereby improving the convergence speed of the student model and the training efficiency. At the same time, it also makes the adjustment operation of the student model more accurate, thereby improving the accuracy of the student model training.

[0094] See also Figure 3 , Figure 3It is a schematic block diagram of a model distillation device based on a multi-teacher model provided in one embodiment of the present application. The model distillation device based on a multi-teacher model can be configured in a server to execute the aforementioned model distillation method based on a multi-teacher model.

[0095] like Figure 3 As shown, the model distillation device 200 based on the multi-teacher model includes: a first label generation module 201, a model generation module 202, a second label generation module 203, a strategy update module 204 and a model determination module 205.

[0096] A first label generation module 201 is used to obtain training sample data and hard labels corresponding to the training sample data, identify the training sample data through multiple teacher models to obtain multiple first soft labels, and identify the training sample data through a first student model to obtain a second soft label;

[0097] A model generation module 202, configured to perform knowledge distillation learning on the first student model according to the hard label, the plurality of first soft labels and the second soft label to generate a second student model; wherein the model parameters of the first student model and the model parameters of the second student model are different;

[0098] A second label generation module 203, configured to identify the training sample data by using the second student model to obtain a third soft label;

[0099] A strategy updating module 204 is used to update the teacher model selection strategy according to the first soft label and the third soft label to obtain an updated teacher model selection strategy, wherein the teacher model selection strategy is used to select a teacher model;

[0100] The model determination module 205 is used to redetermine the corresponding teacher model based on the updated teacher model selection strategy, and perform knowledge distillation learning on the first student model according to the redetermined teacher model until the first student model converges to obtain a target student model.

[0101] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0102] The method and apparatus of the present application can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer terminal devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices.

[0103] Exemplarily, the above method and apparatus may be implemented in the form of a computer program. The computer program may be implemented in the form of a computer program. Figure 4 Runs on the computer device shown.

[0104] See also Figure 4 , Figure 4 1 is a schematic diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.

[0105] like Figure 4 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0106] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any model distillation method based on a multi-teacher model.

[0107] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0108] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any model distillation method based on the multi-teacher model.

[0109] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that the structure of the computer device is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0110] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0111] In some embodiments, the processor is used to run a computer program stored in the memory to implement the following steps: obtaining training sample data and hard labels corresponding to the training sample data, identifying the training sample data through multiple teacher models to obtain multiple first soft labels, and identifying the training sample data through a first student model to obtain a second soft label;

[0112] According to the hard label, the multiple first soft labels and the second soft label, the first student model is subjected to knowledge distillation learning to generate a second student model; wherein the model parameters of the first student model are different from the model parameters of the second student model; the training sample data is identified by the second student model to obtain a third soft label; the teacher model selection strategy is updated according to the first soft label and the third soft label to obtain an updated teacher model selection strategy, and the teacher model selection strategy is used to select a teacher model; based on the updated teacher model selection strategy, the corresponding teacher model is re-determined, and knowledge distillation learning is performed on the first student model according to the re-determined teacher model until the first student model converges to obtain a target student model.

[0113] In some embodiments, the processor is also used to: determine a first deviation value based on the multiple first soft labels and the second soft labels; determine a second deviation value based on the hard label and the second soft label; determine a third deviation value based on the first deviation value and the second deviation value, and initialize the model parameters of the first student model according to the third deviation value to generate a second student model.

[0114] In some embodiments, the processor is also used to: determine the weight ratio corresponding to the first deviation value and the second deviation value based on the reverse gradient propagation algorithm; determine the third deviation value according to the first deviation value and the second deviation value and the corresponding weight ratio.

[0115] In some embodiments, the processor is also used to: determine a fourth deviation value based on the hard label and the third soft label; generate a state vector parameter corresponding to each of the teacher models based on multiple of the first soft labels and the fourth deviation value; determine an update strategy based on the state vector parameter corresponding to each of the teacher models; and update the teacher model selection strategy based on the update strategy to obtain an updated teacher model selection strategy.

[0116] In some embodiments, the processor is also used to: based on a threshold function, calculate the scores of each teacher model in the teacher model selection strategy according to the state vector parameters corresponding to each of the teacher models; filter the teacher models in the teacher model selection strategy according to the scores of the various teacher models to obtain the filtered teacher models; and determine the update strategy based on the filtered teacher models.

[0117] In some embodiments, the processor is also used to: test the second student model through a test set to obtain the accuracy of the second student model; determine the corresponding feedback strategy according to the accuracy of the second student model; update the teacher model selection strategy based on the feedback strategy and the update strategy to obtain an updated teacher model selection strategy.

[0118] In some embodiments, the processor is also used to: perform knowledge distillation learning on the first student model according to the re-determined teacher model to generate a third student model; obtain a third deviation value corresponding to the third student model, and determine whether the third deviation value meets the convergence condition; if the third deviation value meets the convergence condition, use the third student model as the target student model.

[0119] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and when the program instructions are executed, any model distillation method based on a multi-teacher model provided in the embodiment of the present application is implemented.

[0120] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.

[0121] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0122] The present invention refers to a new application model of computer technologies such as storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. of the blockchain language model. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, platform product service layer, and application service layer.

[0123] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A model distillation method based on a multi-teacher model, characterized in that, the method includes: Obtaining training sample data and the hard labels corresponding to the training sample data, identifying the training sample data through multiple teacher models to obtain multiple first soft labels, and identifying the training sample data through a first student model to obtain a second soft label, where the training sample data is a sentence pair or a picture; Performing knowledge distillation learning on the first student model according to the hard labels, the multiple first soft labels and the second soft label to generate a second student model; wherein, the model parameters of the first student model and the model parameters of the second student model are different; Identifying the training sample data through the second student model to obtain a third soft label; Updating the teacher model selection strategy according to the hard labels, the first soft labels and the third soft label to obtain an updated teacher model selection strategy, where the teacher model selection strategy is used to select a teacher model; Re-determining the corresponding teacher model based on the updated teacher model selection strategy, and performing knowledge distillation learning on the first student model according to the re-determined teacher model until the first student model converges to obtain a target student model; The performing knowledge distillation learning on the first student model according to the hard labels, the multiple first soft labels and the second soft label to generate a second student model includes: Determining a first deviation value according to the multiple first soft labels and the second soft label; Determining a second deviation value according to the hard label and the second soft label; Determining a third deviation value according to the first deviation value and the second deviation value, and initializing the model parameters of the first student model according to the third deviation value to generate a second student model; The updating the teacher model selection strategy according to the hard labels, the first soft labels and the third soft label to obtain an updated teacher model selection strategy includes: Determining a fourth deviation value according to the hard label and the third soft label; Generating state vector parameters corresponding to each teacher model according to the multiple first soft labels and the fourth deviation value; Determining an update strategy according to the state vector parameters corresponding to each teacher model; Updating the teacher model selection strategy based on the update strategy to obtain an updated teacher model selection strategy.

2. The method according to claim 1, characterized in that, the determining a third deviation value according to the first deviation value and the second deviation value includes: Based on the reverse gradient propagation algorithm, determining the weight ratio corresponding to the first deviation value and the second deviation value; Determining a third deviation value according to the first deviation value, the second deviation value and the corresponding weight ratio.

3. The method according to claim 1, characterized in that, the determining an update strategy according to the state vector parameters corresponding to each teacher model includes: Based on a threshold function, calculating the scores of each teacher model in the teacher model selection strategy according to the state vector parameters corresponding to each teacher model; Screen the teacher models in the teacher model selection strategy according to the scores of the respective teacher models to obtain the screened teacher models; Determine an update strategy based on the screened teacher models.

4. The method according to claim 1, wherein, after determining the update strategy according to the state vector parameters corresponding to each teacher model, the method further includes: Testing the second student model with a test set to obtain the accuracy rate of the second student model; Determining a corresponding feedback strategy according to the accuracy rate of the second student model; Updating the teacher model selection strategy based on the feedback strategy and the update strategy to obtain an updated teacher model selection strategy.

5. The method according to claim 1, wherein, The knowledge distillation learning of the first student model according to the re-determined teacher model until the first student model converges to obtain a target student model includes: Performing knowledge distillation learning on the first student model according to the re-determined teacher model to generate a third student model; Obtaining a third deviation value corresponding to the third student model and determining whether the third deviation value meets the convergence condition; If the third deviation value meets the convergence condition, then using the third student model as the target student model.

6. A model distillation device based on multiple teacher models, wherein, comprising: A first label generation module, configured to obtain training sample data and hard labels corresponding to the training sample data, identify the training sample data through multiple teacher models to obtain multiple first soft labels, and identify the training sample data through a first student model to obtain second soft labels, where the training sample data is a sentence pair or a picture; A model generation module, configured to perform knowledge distillation learning on the first student model according to the hard labels, the multiple first soft labels, and the second soft labels to generate a second student model; wherein, the model parameters of the first student model and the model parameters of the second student model are different; A second label generation module, configured to identify the training sample data through the second student model to obtain third soft labels; A strategy update module, configured to update the teacher model selection strategy according to the hard labels, the first soft labels, and the third soft labels to obtain an updated teacher model selection strategy, where the teacher model selection strategy is used to select teacher models; A model determination module, configured to re-determine a corresponding teacher model based on the updated teacher model selection strategy, and perform knowledge distillation learning on the first student model according to the re-determined teacher model until the first student model converges to obtain a target student model; The performing knowledge distillation learning on the first student model according to the hard labels, the multiple first soft labels, and the second soft labels to generate a second student model includes: Determining a first deviation value according to the multiple first soft labels and the second soft labels; Determining a second deviation value according to the hard label and the second soft label; Determine a third deviation value according to the first deviation value and the second deviation value, and initialize the model parameters of the first student model according to the third deviation value to generate a second student model; The updating the teacher model selection strategy according to the hard label, the first soft label, and the third soft label to obtain an updated teacher model selection strategy includes: Determine a fourth deviation value according to the hard label and the third soft label; Generate state vector parameters corresponding to each teacher model according to the plurality of first soft labels and the fourth deviation value; Determine an update strategy according to the state vector parameters corresponding to each teacher model; Update the teacher model selection strategy based on the update strategy to obtain an updated teacher model selection strategy.

7. A computer device, characterized in that, the computer device includes a memory and a processor; the memory is used for storing a computer program; the processor is configured to execute the computer program and, when executing the computer program, implement: the model distillation method based on multiple teacher models according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the model distillation method based on multiple teacher models according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Network training method and device for knowledge distillation, medium and electronic device

    CN110674880A

  • Model generation method, video screening method, related device and storage medium

    CN113392864A