Large model knowledge distillation method, device and equipment and storage medium

By using the confidence score of the teacher model and the probability score of the student model during knowledge distillation for double verification, and updating the teaching template of the teacher model, the problem of poor performance in the professional field is solved, and the model's ability to capture professional knowledge and iterative optimization performance is improved.

CN119962625APending Publication Date: 2025-05-09SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510004240.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the existing knowledge distillation methods, student models lack effective capture of professional knowledge, resulting in poor model performance.

Method used

By entering the prediction sample into the pre-trained teacher model, the confidence score of the teacher model is obtained and the prediction results of the target case with the confidence score greater than the preset threshold are input into the student model for knowledge distillation. At the same time, the teaching template of the teacher model is updated according to the probability score of the student model to achieve iterative optimization of the student model.

Benefits of technology

The performance of student models in professional fields is improved, the cases where the teacher models are confident in predicting results do not mislead students' models, the ability of students' models to capture professional knowledge is enhanced, and the model performance is optimized through feedback mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962625A_ABST
    Figure CN119962625A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of machine learning, and provides a large model knowledge distillation method, device and equipment and a storage medium, and the method comprises the steps: inputting a prediction sample into a pre-trained teacher model, and obtaining a prediction result of the teacher model for each sample case in the prediction sample and a confidence score of the prediction result; inputting the prediction result of the target case with the confidence score greater than a preset threshold value into a student model for knowledge distillation, finely adjusting the student model, and obtaining the prediction result of the student model for each target case and the probability score of the prediction result; and updating a teaching template of the teacher model according to the probability score, and returning and executing the step of inputting the prediction sample into the pre-trained teacher model so as to carry out iterative optimization on the student model until the student model meets a preset iteration termination condition. And on the basis of a double check and trust maximization mechanism of a confidence score and a probability score, the performance of the student model on professional knowledge in the field is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a large-model knowledge distillation method, device, equipment and storage medium. Background Art

[0002] In the field of natural language processing, large language models (LLMs) have become a key force in promoting the advancement of natural language processing technology. Through deep learning technology, large language models have achieved remarkable achievements in tasks such as text generation, dialogue systems, machine translation, etc. As the scale of the model continues to expand, challenges in deployment and application arise, mainly including high computing costs, which make it difficult to deploy and apply large language models on resource-constrained devices, and the model response speed is not fast enough to meet the needs of real-time processing.

[0003] In order to overcome the limitations of large language models in computing resources, model compression technology has emerged, among which knowledge distillation is a widely used method. The basic idea of ​​knowledge distillation is to extract knowledge from a complex and large-scale teacher model and transfer it to a simpler and lighter student model. This method can significantly reduce the model's demand for computing resources while maintaining or even improving model performance.

[0004] Existing knowledge distillation technology still has shortcomings in some aspects. For example, the student model does not perform well in some professional fields due to the lack of effective capture of domain expertise. When faced with complex or rare input samples, the student model may not be able to give accurate answers, resulting in poor model performance. Summary of the invention

[0005] The present invention provides a large-model knowledge distillation method, device, equipment and storage medium to solve the defect in the existing knowledge distillation method that the student model lacks effective capture of professional knowledge, resulting in poor model performance, and improve the performance of the student model in the professional field.

[0006] The present invention provides a large model knowledge distillation method, comprising: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a confidence score of the first prediction result; Inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation to fine-tune the student model, and obtaining the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; The teaching template of the teacher model is updated according to the probability score, and the step of inputting the predicted samples into the pre-trained teacher model is returned and executed to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0007] According to the large model knowledge distillation method provided by the present invention, the updating of the teaching template of the teacher model according to the probability score includes: Identifying, from among the target cases, an edge case having the smallest probability score; Inputting the edge case into the teacher model, and using the teacher model to generate explanation information of the edge case; The edge cases and the explanation information are input into the teacher model to update the teaching template of the teacher model.

[0008] According to the large model knowledge distillation method provided by the present invention, the prediction sample is input into the pre-trained teacher model, and the first prediction result of the teacher model for each sample case in the prediction sample and the confidence score of the first prediction result are obtained, including: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a prediction reason corresponding to the first prediction result; Based on a preset confidence assessment algorithm, a confidence assessment is performed on the first prediction result and the prediction reason to generate a confidence score for the first prediction result.

[0009] According to the large model knowledge distillation method provided by the present invention, after updating the teaching template of the teacher model according to the probability score, it also includes: Obtaining evaluation parameters of the student model and evaluation weights corresponding to the evaluation parameters; the evaluation parameters at least include model accuracy and model size; The performance of the student model is evaluated based on the evaluation parameters and the evaluation weights, and whether the performance of the student model satisfies a preset iteration termination condition is determined according to the evaluation result.

[0010] According to the large model knowledge distillation method provided by the present invention, before inputting the prediction sample into the pre-trained teacher model, the method further includes: Acquire a text dataset of a preset field, and annotate the text dataset based on a preset annotation algorithm to obtain an annotated text set; Inputting the annotated text set into the teacher model, and using the teacher model to convert the annotated text in the annotated text set into a first embedding vector; Based on the first embedding vector, case retrieval is performed on the training samples of the teacher model to obtain predicted samples.

[0011] According to the large model knowledge distillation method provided by the present invention, performing case retrieval on the training samples of the teacher model based on the first embedding vector to obtain prediction samples includes: Obtaining a set of embedding vectors corresponding to the training samples of the teacher model; Calculating the similarity between the first embedding vector and each second embedding vector in the embedding vector set; The training samples are retrieved according to the similarity, and a preset number of sample cases having the greatest similarity to the first embedding vector are selected from the training samples as prediction samples.

[0012] According to the large model knowledge distillation method provided by the present invention, after the student model satisfies a preset iteration termination condition, the method further includes: Deploy and apply the student model in a preset business application; During the application of the student model, feedback information of the student model is collected to monitor and optimize the student model.

[0013] The present invention also provides a large model knowledge distillation device, comprising the following modules: A prediction module, used for inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a confidence score of the first prediction result; A knowledge distillation module, used for inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation, so as to fine-tune the student model, and obtain the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; An iterative optimization module is used to update the teaching template of the teacher model according to the probability score, return to and execute the step of inputting the predicted sample into the pre-trained teacher model to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the large model knowledge distillation method as described in any one of the above are implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the large model knowledge distillation method as described in any one of the above.

[0016] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the large model knowledge distillation method as described in any one of the above.

[0017] The large-model knowledge distillation method, device, equipment and storage medium provided by the present invention perform a double verification of the prediction results of the same case during the knowledge distillation process through the confidence score of the teacher model and the probability score of the student model, verify the confidence score of the prediction result of each sample case in the prediction sample according to the teacher model, and input the prediction result of the target case with a confidence score greater than a preset threshold into the student model for knowledge distillation, thereby avoiding the teacher model from giving incorrect guidance to the student model for sample cases whose prediction results are not confident, and improving the student model's ability to effectively capture professional knowledge; verify the probability score of the prediction result of the target case according to the student model, update the teaching template of the teacher model according to the probability score, build a feedback mechanism for the teacher model, and realize iterative optimization of the student model, thereby improving the performance of the student model in the professional field. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is a schematic diagram of the process of the large model knowledge distillation method provided by the present invention.

[0020] Figure 2 It is a schematic diagram of the structure of the large-scale model knowledge distillation device provided by the present invention.

[0021] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] An embodiment of the present invention provides a large-model knowledge distillation method, which distills the knowledge of the teacher model into the student model based on double verification and trust maximization mechanism, and feeds back the edge cases with the lowest probability scores in the prediction results of the student model to the teacher model to update the teaching template, facilitate subsequent iterative optimization of the student model, and improve the performance of the student model in professional knowledge.

[0024] Specifically, Figure 1 is a flow chart of the large model knowledge distillation method provided by the present invention, such as Figure 1 As shown, the method comprises the following steps: Step 100, inputting the prediction sample into the pre-trained teacher model, obtaining the first prediction result of the teacher model for each sample case in the prediction sample, and the confidence score of the first prediction result; Step 200, inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation to fine-tune the student model, and obtaining the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; Step 300, updating the teaching template of the teacher model according to the probability score, returning to and executing the step of inputting the predicted sample into the pre-trained teacher model to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0025] The prediction sample is input into the pre-trained teacher model to obtain the first prediction result of the teacher model for each sample case in the prediction sample and the confidence score of the first prediction result. The teacher model predicts each sample case in the prediction sample, determines the category to which each sample case belongs and the probability of the category to which it belongs, thereby obtaining the probability distribution of the prediction result of the teacher model for the prediction sample.

[0026] The first prediction result of the target case whose confidence score is greater than the preset threshold is input into the student model, and the student model is subjected to knowledge distillation to fine-tune the student model, and the second prediction result of the student model for each target case and the probability score corresponding to the second prediction result are obtained.

[0027] Optionally, since the first prediction result of the teacher model for the sample case is obtained by classifying and predicting the sample case, the first prediction result is used to characterize the category to which the sample case belongs, and therefore, the confidence score is used to characterize the confidence that the sample case belongs to the category corresponding to the first prediction result. Similarly, the second prediction result of the student model for the sample case is obtained by classifying and predicting the target case, and the second prediction result is used to characterize the category to which the target case belongs, and therefore, the probability score is used to characterize the probability that the target case belongs to the category corresponding to the second prediction result.

[0028] For the student model, the teacher model only passes the prediction results of the target cases whose confidence scores exceed the preset threshold to the student model, and eliminates the sample cases whose confidence scores do not exceed the preset threshold, thereby avoiding passing wrong guidance to the student model.

[0029] Optionally, for sample cases whose confidence scores are less than or equal to a preset threshold, the teacher model is not confident enough in its prediction results and needs to be re-prompted. For target cases whose confidence scores are greater than the preset threshold, the teacher model is confident enough in the prediction results, so the prediction results of the teacher model for the target cases are input into the student model as soft labels for the student model to supervise the learning of the student model.

[0030] During the knowledge distillation process, the student model's prediction results for the target case will continue to approach the teacher model's prediction results for the target case, and learn the knowledge of the teacher model by minimizing the difference between its prediction results and the teacher model's soft labels.

[0031] Optionally, during the knowledge distillation process, appropriate parameters of the teacher model and the student model can be adjusted and optimized according to actual needs to improve the efficiency and effectiveness of knowledge transfer.

[0032] Furthermore, in the process of knowledge distillation, the second prediction result of the student model for each target case and the probability score of the target category corresponding to each second prediction result are obtained, and the teaching template of the teacher model is updated according to the probability score. Then, return to and execute step 100, so as to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0033] Optionally, the preset iteration termination condition includes the number of iterations reaching a preset number, or the performance of the student model meets the accuracy requirement.

[0034] Optionally, the prediction results of the teacher model for the target case are input into the student model, the student model is fine-tuned, and the prediction results of the student model for each target case are obtained. At the same time, the student model calculates a probability score for the prediction result of each target case.

[0035] Optionally, the teaching template of the teacher model is updated according to the probability score of the prediction result of the student model for the target case, so as to facilitate the subsequent knowledge distillation of the student model. In one embodiment, the sample case with the largest difference in prediction results between the student model and the teacher model, or the sample case with a difference in prediction results between the student model and the teacher model greater than a preset value, is determined according to the probability score, and the teaching template of the teacher model is updated based on this. In another embodiment, the sample case with the lowest probability score is obtained, and the teaching template of the teacher model is updated based on this.

[0036] It should be noted that the teaching template of the teacher model can be the soft label output by the teacher model, that is, the prediction result of the teacher model. In some embodiments, the role played by the teacher model in the knowledge distillation process can be used as a teaching template to provide a learning example or reference for the student model through the output and behavior of the teacher model. The teaching template here is not static, but dynamically generated based on the processing and response of the teacher model to the input data.

[0037] The role of the teaching template of the teacher model in knowledge distillation is mainly reflected in that it provides learning goals and guidance for the student model through the output prediction results (such as soft labels), helping the student model to learn and imitate the complex features and decision-making capabilities of the teacher model.

[0038] In this embodiment, the prediction results of the same case are doubly verified in the knowledge distillation process through the confidence score of the teacher model and the probability score of the student model. The confidence score of the prediction result of each sample case in the prediction sample is verified according to the teacher model, and the prediction result of the target case with a confidence score greater than a preset threshold is input into the student model for knowledge distillation, thereby avoiding the teacher model from giving incorrect guidance to the student model for sample cases whose prediction results are not confident, and improving the ability of the student model to effectively capture professional knowledge; the probability score of the prediction result of the target case is verified according to the student model, and the teaching template of the teacher model is updated according to the probability score, and a feedback mechanism for the teacher model is constructed to realize iterative optimization of the student model, thereby improving the performance of the student model in the professional field.

[0039] In one embodiment, when updating the teaching template of the teacher model according to the probability score, the teaching template of the teacher model is specifically updated according to the target case with the lowest probability score. Based on this, in step 300, updating the teaching template of the teacher model according to the probability score may also include: Step 301, identifying the edge case with the smallest probability score from each of the target cases; Step 302, inputting the edge case into the teacher model, and using the teacher model to generate explanation information of the edge case; Step 303: input the edge case and the explanation information into the teacher model to update the teaching template of the teacher model.

[0040] According to the probability score of the target category corresponding to the prediction result of each target case by the student model, the edge case with the smallest probability score is identified from each target case, and the edge case is input into the teacher model. The teacher model is used to generate explanation information of the edge case, which is used to explain why the confidence of the prediction result of the student model for the edge case is low. The identified edge case and its explanation information are input into the teacher model, and the teaching template of the teacher model is updated.

[0041] In some related technologies, a confidence check is performed between the teacher model and the student model. When the student model has low confidence in the prediction results, it seeks help from the teacher model to improve the accuracy of the final output prediction results. This method improves the interactivity and reliability of the model. When faced with complex or rare inputs, the student model may not be able to give an accurate answer, especially in the absence of sufficient feedback mechanisms. When the student model encounters uncertain situations, it is difficult to generate explanations or bases for the prediction results in a systematic way, which affects the transparency of the model and user trust.

[0042] Based on the interactivity between the teacher model and the student model, in this embodiment, in view of the limitations of the student model in dealing with uncertain or rare situations, edge cases are identified and corresponding explanatory information is generated to explain the behavior of the model. A feedback mechanism of the student model to the teacher model is established, and the teaching template of the teacher model is updated, thereby enhancing the generalization ability of the student model when facing novel or complex situations.

[0043] In order to improve the understanding and adaptability of the student model to the knowledge in the professional field, the context alignment technology is used to obtain the prediction sample. In step 100, before the prediction sample is input into the pre-trained teacher model, the following steps may also be included: Step 001, obtaining a text dataset of a preset field, and annotating the text dataset based on a preset annotation algorithm to obtain an annotated text set; Step 002, inputting the annotated text set into the teacher model, and using the teacher model to convert the annotated text in the annotated text set into a first embedding vector; Step 003: Perform case retrieval on the training samples of the teacher model based on the first embedding vector to obtain predicted samples.

[0044] Optionally, the text dataset can be represented as , is the n data points associated with the text data in the text dataset D. The annotation algorithm is shown in the following formula 1: ; (1)

[0045] In formula 1, is the original text data, For text data The annotation results of It is a labeling function. The labeling of text data includes but is not limited to binary classification of categories in preset fields and marking of technical causal relationships.

[0046] Furthermore, the annotated text set is input into the teacher model, and the annotated text in the annotated text set is converted into an embedding vector using the teacher model to obtain a first embedding vector. Based on the first embedding vector, case retrieval is performed on the training samples of the teacher model to obtain a predicted sample.

[0047] Furthermore, case retrieval of training samples of the teacher model based on the first embedding vector is implemented based on the similarity between the embedding vectors. Therefore, step 003 also includes: Step 0031, obtaining a set of embedding vectors corresponding to the training samples of the teacher model; Step 0032, calculating the similarity between the first embedding vector and each second embedding vector in the embedding vector set; Step 0033: perform case retrieval on the training samples according to the similarity, and select a preset number of sample cases with the greatest similarity to the first embedding vector from the training samples as prediction samples.

[0048] When performing case retrieval based on the first embedding vector, first obtain the embedding vector set corresponding to the training samples of the teacher model, then calculate the similarity between the first embedding vector corresponding to the annotated text set and each second embedding vector in the embedding vector set, sort the sample cases in the training samples according to the similarity, and select a preset number of sample cases with the greatest similarity to the first embedding vector to form a prediction sample for fine-tuning the student model, thereby improving the student model's understanding of the professional knowledge in the preset field.

[0049] When performing case retrieval on the training samples of the teacher model based on the first embedding vector, the similarity between the first embedding vector and each second embedding vector in the embedding vector set is calculated, and the second embedding vectors in the embedding vector set are sorted according to the similarity, that is, the sample cases in the training samples are sorted, and a preset number of sample cases with the greatest similarity to the first embedding vector are selected according to the sorting order to form a prediction sample.

[0050] Optionally, the teacher model includes an embedding layer, which can convert the input context text into an embedding vector, and the specific implementation method is shown in Formula 2: ; (2)

[0051] in, is the context-annotated text data, Is the context text The corresponding embedding vector.

[0052] In one embodiment, the similarity between the first embedding vector and each second embedding vector in the embedding vector set is calculated using cosine similarity, and the specific calculation formula is as follows: ; (3)

[0053] In formula 3, is the first embedding vector, is the second embedding vector. According to the calculation method shown in Formula 3, the cosine similarity between the embedding vector output by the teacher model and the corresponding embedding vector in the embedding vector set can be calculated, that is, the cosine similarity between the embedding vector output by the teacher model and each embedding vector corresponding to the training sample can be calculated.

[0054] Sort by similarity, perform case retrieval according to Formula 4, and select the most similar K cases as prediction samples to assist the teacher model in generating more accurate outputs.

[0055] ; (4)

[0056] in, is the embedding vector corresponding to the annotated text set output by the teacher model, that is, the embedding vector of the query context, and E is the set of embedding vectors corresponding to the training samples of the teacher model, which includes multiple embedding vectors.

[0057] It is understandable that the training samples of the teacher model may include text data corresponding to professional knowledge in multiple fields, and fine-tuning the student model based on samples in a specific field can help improve the student model's effective capture of professional knowledge in a specific field, thereby improving the student model's performance in that specific field.

[0058] In one embodiment, when the teacher model generates a prediction result for a prediction sample, it also generates a prediction reason, and the confidence score of the teacher model for the prediction result is generated based on the prediction result and the prediction reason. Therefore, step 200 further includes: Step 201, inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a prediction reason corresponding to the first prediction result; Step 202: Based on a preset confidence assessment algorithm, a confidence assessment is performed on the first prediction result and the prediction reason to generate a confidence score for the first prediction result.

[0059] The prediction sample is input into the pre-trained teacher model, and the first prediction result of the teacher model for each sample case in the prediction sample and the prediction reason corresponding to the first prediction result are obtained. Based on the preset confidence assessment algorithm, the first prediction result and the prediction reason are confidence assessed to generate a confidence score for the first prediction result.

[0060] For example, the confidence score is generated as shown in Formula 5-6: ; (5)

[0061] ; (6)

[0062] in, is the first prediction result of the teacher model for the sample case, is the reason for prediction, Represents one of the K sample cases in the prediction sample, That is the evaluation function corresponding to the confidence evaluation algorithm, Represents the teacher model.

[0063] The double check in the knowledge distillation process includes the confidence judgment of the teacher model. If the confidence score of the teacher model's prediction result for the sample case is lower than the preset threshold, the teacher model is not confident enough about the prediction result and needs to be prompted again. The threshold judgment of the confidence score of the teacher model can be implemented as shown in the following formula 7: ; (7)

[0064] In Formula 7, It is the preset threshold of the confidence score, which can also be called the confidence threshold.

[0065] Furthermore, based on the confidence threshold judgment of the teacher model, the first prediction result of the target case with a confidence score greater than the preset threshold is input into the student model to fine-tune the student model. Specifically, the student model calculates each possible category and generates a probability score: ; (8)

[0066] The teacher model also includes a hidden layer. In Formula 8, is the category label, t is the text representation, Yes Category The weight of is the hidden layer representation of t, that is, the text feature of t in the hidden layer, M is the total number of categories of all possible categories, represents the weight of category m.

[0067] Based on Formula 8, when calculating whether a sample case belongs to a category When the probability score is calculated, the text representation of the sample case is first calculated in the category Then calculate the weighted value of all categories, and represent the text of the sample case in the category The ratio of the weighted value under the category to the sum of its weighted values ​​under all categories is used as a sample case belonging to the category The probability score of .

[0068] In one embodiment, the student model calculates the probability score of the target case belonging to each category, and selects the category with the largest probability score as the prediction result for the target case, thereby obtaining a second prediction result and the probability score corresponding to the second prediction result. Then, based on the probability scores of each target case, the edge case with the lowest probability score is selected, and the teacher model is used to generate explanation information for the edge case, and the edge case and its explanation information are input into the teacher model, and the teaching template of the teacher model is updated.

[0069] Furthermore, in step 300, after updating the teaching template of the teacher model according to the probability score, the following steps may also be included: Step 310, obtaining evaluation parameters of the student model and evaluation weights corresponding to the evaluation parameters; the evaluation parameters at least include model accuracy and model size; Step 320, evaluating the performance of the student model based on the evaluation parameters and the evaluation weights, and determining whether the performance of the student model meets a preset iteration termination condition according to the evaluation result.

[0070] After updating the teaching template of the teacher model, the evaluation parameters of the student model and the evaluation weights corresponding to the evaluation parameters are obtained. The evaluation parameters include at least the model accuracy and the model size. Different parameters in the evaluation parameters have different corresponding weights.

[0071] Further, based on the evaluation parameters and the evaluation weights, the performance of the student model is evaluated, and it is determined whether the performance of the student model meets the preset iteration termination condition according to the evaluation result. Optionally, if the performance of the student model meets the preset iteration termination condition, the fine-tuning and training of the student model are completed, otherwise, if the performance of the student model does not meet the preset iteration termination condition, return to and execute step 100, repeat the above process, and iteratively optimize the student model until the performance of the student model meets the preset iteration termination condition.

[0072] In one embodiment, taking the model accuracy and model size as evaluation parameters, the performance of the student model is evaluated according to the model evaluation method shown in the following formula 9 to obtain the model score: : ; (9)

[0073] In Formula 9, represents the model accuracy, Indicates the model size, that is, the model magnitude or model scale, Represents the evaluation weight of the model accuracy, An evaluation weight representing the model size.

[0074] Optional, and It is a configurable weight parameter that can be configured according to the performance requirements of the student model. The evaluation parameters of the student model can also include recall rate and F1 score.

[0075] Furthermore, after the student model satisfies the preset iteration termination condition, the following steps may also be included: Step 401, deploying and applying the student model in a preset business application; Step 402, during the application process of the student model, feedback information of the student model is collected to monitor and optimize the student model.

[0076] The student model is deployed and applied in the preset business application, and during the application of the student model, feedback information of the student model is collected, and the student model is monitored and optimized according to the collected feedback information. The monitoring of the student model includes but is not limited to the monitoring of the performance of the student model, and the feedback information of the student model can be generated by automatic evaluation of the monitoring of the student model, or can be generated by user feedback of the business system, which is not specifically limited.

[0077] In this embodiment, the student model is fine-tuned by selecting prediction samples through context alignment, which enhances the student model's effective capture of professional knowledge in the domain, improves the student model's understanding and adaptability to professional knowledge, and thus improves the student model's performance on professional knowledge in the domain.

[0078] Furthermore, based on the double verification and trust maximization mechanism, by checking the confidence of the prediction results of the teacher model, the prediction results with higher confidence of the teacher model are screened out to guide the student model, which can achieve efficient distillation of professional knowledge. And by checking the probability score of the prediction results of the student model, the edge cases with the lowest probability score are screened out to update the teaching template of the teacher model. A feedback mechanism between the student model and the teacher model is established, which can dynamically adjust the response of the model, which can not only achieve iterative optimization of the student model, but also improve the interactivity and reliability of the model.

[0079] In addition, in view of the limitations of the student model in dealing with uncertain or rare situations, edge cases are identified and corresponding reasons are generated to explain the model behavior, which are used to update the teaching template, thereby enhancing the generalization ability of the student model when facing novel or complex situations.

[0080] The large model knowledge distillation device provided by the present invention is described below. The large model knowledge distillation device described below and the large model knowledge distillation method described above can be referenced to each other.

[0081] Reference Figure 2 , the large model knowledge distillation device provided by the embodiment of the present invention includes: The prediction module 10 is used to input the prediction sample into the pre-trained teacher model, obtain the first prediction result of the teacher model for each sample case in the prediction sample, and the confidence score of the first prediction result; A knowledge distillation module 20 is used to input the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation, so as to fine-tune the student model, and obtain the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; The iterative optimization module 30 is used to update the teaching template of the teacher model according to the probability score, return to and execute the step of inputting the predicted sample into the pre-trained teacher model to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0082] In one embodiment, the iterative optimization module 30 is further used to: Identifying, from among the target cases, an edge case having the smallest probability score; Inputting the edge case into the teacher model, and using the teacher model to generate explanation information of the edge case; The edge cases and the explanation information are input into the teacher model to update the teaching template of the teacher model.

[0083] In one embodiment, the prediction module 10 is further configured to: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a prediction reason corresponding to the first prediction result; Based on a preset confidence assessment algorithm, a confidence assessment is performed on the first prediction result and the prediction reason to generate a confidence score for the first prediction result.

[0084] In one embodiment, the large model knowledge distillation device further includes a model evaluation module, which is used to: Obtaining evaluation parameters of the student model and evaluation weights corresponding to the evaluation parameters; the evaluation parameters at least include model accuracy and model size; The performance of the student model is evaluated based on the evaluation parameters and the evaluation weights, and whether the performance of the student model satisfies a preset iteration termination condition is determined according to the evaluation result.

[0085] In one embodiment, the large model knowledge distillation device further includes a context alignment module, which is used to: Acquire a text dataset of a preset field, and annotate the text dataset based on a preset annotation algorithm to obtain an annotated text set; Inputting the annotated text set into the teacher model, and using the teacher model to convert the annotated text in the annotated text set into a first embedding vector; Based on the first embedding vector, case retrieval is performed on the training samples of the teacher model to obtain predicted samples.

[0086] In one embodiment, the context alignment module is further used to: Obtaining a set of embedding vectors corresponding to the training samples of the teacher model; Calculating the similarity between the first embedding vector and each second embedding vector in the embedding vector set; The training samples are retrieved according to the similarity, and a preset number of sample cases having the greatest similarity to the first embedding vector are selected from the training samples as prediction samples.

[0087] In one embodiment, the large model knowledge distillation device further includes a deployment application module for: Deploy and apply the student model in a preset business application; During the application of the student model, feedback information of the student model is collected to monitor and optimize the student model.

[0088] Figure 3An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communications interface 320 and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the steps of the large model knowledge distillation method, for example, including: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a confidence score of the first prediction result; Inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation to fine-tune the student model, and obtaining the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; The teaching template of the teacher model is updated according to the probability score, and the step of inputting the predicted samples into the pre-trained teacher model is returned and executed to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0089] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0090] On the other hand, the present invention further provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the large model knowledge distillation method provided in the above embodiments, for example, including: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a confidence score of the first prediction result; Inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation to fine-tune the student model, and obtaining the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; The teaching template of the teacher model is updated according to the probability score, and the step of inputting the predicted samples into the pre-trained teacher model is returned and executed to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0091] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the steps of the large model knowledge distillation method provided in the above embodiments, for example, including: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a confidence score of the first prediction result; Inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation to fine-tune the student model, and obtaining the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; The teaching template of the teacher model is updated according to the probability score, and the step of inputting the predicted samples into the pre-trained teacher model is returned and executed to iteratively optimize the student model until the student model meets the preset iteration termination condition.

[0092] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model knowledge distillation method, characterized in that: include: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a confidence score of the first prediction result; Inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation to fine-tune the student model, and obtaining the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; The teaching template of the teacher model is updated according to the probability score, and the step of inputting the predicted samples into the pre-trained teacher model is returned and executed to iteratively optimize the student model until the student model meets the preset iteration termination condition.

2. The large model knowledge distillation method according to claim 1, characterized in that: The updating of the teaching template of the teacher model according to the probability score comprises: Identifying, from among the target cases, an edge case having the smallest probability score; Inputting the edge case into the teacher model, and using the teacher model to generate explanation information of the edge case; The edge cases and the explanation information are input into the teacher model to update the teaching template of the teacher model.

3. The large model knowledge distillation method according to claim 1, characterized in that: The step of inputting the prediction sample into the pre-trained teacher model, obtaining the first prediction result of the teacher model for each sample case in the prediction sample, and the confidence score of the first prediction result, includes: Inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a prediction reason corresponding to the first prediction result; Based on a preset confidence assessment algorithm, a confidence assessment is performed on the first prediction result and the prediction reason to generate a confidence score for the first prediction result.

4. The large model knowledge distillation method according to claim 1, characterized in that: After the teaching template of the teacher model is updated according to the probability score, it also includes: Obtaining evaluation parameters of the student model and evaluation weights corresponding to the evaluation parameters; the evaluation parameters at least include model accuracy and model size; The performance of the student model is evaluated based on the evaluation parameters and the evaluation weights, and whether the performance of the student model satisfies a preset iteration termination condition is determined according to the evaluation result.

5. The large model knowledge distillation method according to claim 1, characterized in that: Before inputting the prediction sample into the pre-trained teacher model, the method further includes: Acquire a text dataset of a preset field, and annotate the text dataset based on a preset annotation algorithm to obtain an annotated text set; Inputting the annotated text set into the teacher model, and using the teacher model to convert the annotated text in the annotated text set into a first embedding vector; Based on the first embedding vector, case retrieval is performed on the training samples of the teacher model to obtain predicted samples.

6. The large model knowledge distillation method according to claim 5, characterized in that: The performing case retrieval on the training samples of the teacher model based on the first embedding vector to obtain a predicted sample includes: Obtaining a set of embedding vectors corresponding to the training samples of the teacher model; Calculating the similarity between the first embedding vector and each second embedding vector in the embedding vector set; The training samples are retrieved according to the similarity, and a preset number of sample cases having the greatest similarity to the first embedding vector are selected from the training samples as prediction samples.

7. The large model knowledge distillation method according to claim 1, characterized in that: After the student model satisfies a preset iteration termination condition, the method further includes: Deploy and apply the student model in a preset business application; During the application of the student model, feedback information of the student model is collected to monitor and optimize the student model.

8. A large model knowledge distillation device, characterized in that: include: A prediction module, used for inputting the prediction sample into a pre-trained teacher model, obtaining a first prediction result of the teacher model for each sample case in the prediction sample, and a confidence score of the first prediction result; A knowledge distillation module, used for inputting the first prediction result of the target case whose confidence score is greater than a preset threshold into the student model for knowledge distillation, so as to fine-tune the student model, and obtain the second prediction result of the student model for each of the target cases, and the probability score corresponding to the second prediction result; An iterative optimization module is used to update the teaching template of the teacher model according to the probability score, return to and execute the step of inputting the predicted sample into the pre-trained teacher model to iteratively optimize the student model until the student model meets the preset iteration termination condition.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the large model knowledge distillation method as described in any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the large model knowledge distillation method as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Model updating method and device, computer equipment and storage medium

    CN121918854A

  • Model update method, apparatus, computer equipment and storage medium

    CN121918854B