Large model dynamic knowledge distillation method and system
By dynamically adjusting the distillation parameters, the prediction uncertainty and category fuzzy problems of the static knowledge distillation method under dynamic data changes are solved, and the stability and robustness of the model are improved, and efficient knowledge transfer is adapted to complex question-and-answer scenarios.
Patent Information
- Application Number
- CN202510726044.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
AI Technical Summary
The existing static knowledge distillation method cannot adapt to the dynamic changes in data distribution, resulting in large fluctuations in prediction confidence in the Q&A scenarios and blurred boundaries in category distinctions, making it difficult to meet the high requirements of model output stability.
By obtaining the Q&A pair dataset, using the student model to generate the first and second soft labels, calculate the predicted uncertain values and distinguishing ability, dynamically adjust the distillation intensity weight and distillation temperature values, iteratively train the student model until the performance indicators reach the preset threshold, and realize dynamic knowledge distillation.
It improves the stability of model output, reduces noise interference from fuzzy samples, improves the generalization ability of complex semantics, and enhances the robustness of the model for multi-level policy interpretation and cross-departmental collaborative Q&A.
Smart Images

Figure CN120235216A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dynamic knowledge distillation of large models, and particularly to a method and system for dynamic knowledge distillation of large models. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, intelligent systems based on large language models have gradually become an important means to improve service efficiency. However, the professionalism and complexity of the field pose special requirements for the knowledge understanding ability of the model, and traditional knowledge distillation methods face significant challenges in dealing with dynamic scenarios.
[0003] Currently, the construction of large models in the field mainly relies on the transfer learning technology of general language models, and the capabilities of large-scale pre-trained models are transferred to lightweight student models through knowledge distillation. However, existing distillation methods are mostly static knowledge distillation and cannot adapt to the characteristics of dynamic changes in data distribution. In the question-and-answer scenario, the input questions often involve complex semantics such as multi-level policy clause interpretation and cross-departmental business collaboration, resulting in large fluctuations in the prediction confidence of the student model during the knowledge transfer process and blurred category discrimination boundaries, often leading to stagnation or overfitting of the model performance in the later stage of distillation, and it is difficult to meet the high requirements of the scenario for the stability of model output. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method and system for dynamic knowledge distillation of large models for the dynamic knowledge distillation of large models in the field, solve the problems of prediction fluctuations and category fuzziness caused by data dynamics, improve the stability of the model, and suppress overfitting.
[0005] To achieve the above purpose, the embodiments of the present invention provide a method for dynamic knowledge distillation of large models, including: obtaining a question-and-answer pair dataset in the relevant field, randomly collecting input question-and-answer pairs in the question-and-answer pair dataset, and dividing the input question-and-answer pairs into training input question-and-answer pairs and fine-tuning input question-and-answer pairs; loading a pre-trained language model as the student model, and using the student model to encode the training input question-and-answer pairs to generate the first soft label; performing distillation, and using the student model again to encode the training input question-and-answer pairs to generate the second soft label; calculating the prediction uncertainty value and discrimination ability of the student model according to the first soft label and the second soft label; dynamically adjusting the distillation intensity weight and distillation temperature value according to the prediction uncertainty value and the discrimination ability; iteratively executing the process of training the student model until the performance index of the student model reaches a preset threshold, and completing the dynamic knowledge distillation.
[0006] Optionally, calculating the prediction uncertainty value and discrimination ability of the student model includes: inputting the training input Q&A pairs into the student model for prediction to generate a probability distribution of each possible output category, and this probability distribution is the second soft label of the current round; quantifying the discreteness of the second soft label through probability variance to obtain the prediction uncertainty value; setting a first variance threshold, if the prediction uncertainty value is less than or equal to the first variance threshold, it is determined that the prediction uncertainty value of the current student model is low; if the prediction uncertainty value is greater than the first variance threshold, it is determined that the prediction uncertainty value of the current student model is high.
[0007] Optionally, calculating the prediction uncertainty value and discrimination ability of the student model further includes: retrieving the first soft label, quantifying the distribution difference between the first soft label and the second soft label through relative entropy to obtain the discrimination ability value; setting a first discrimination ability threshold, if the discrimination ability value is less than or equal to the first discrimination ability threshold, it is determined that the discrimination ability of the current student model is strong, if the discrimination ability value is greater than the first discrimination ability threshold, it is determined that the discrimination ability of the current student model is weak.
[0008] Optionally, dynamically adjusting the distillation intensity weight and distillation temperature value according to the prediction uncertainty value and the discrimination ability includes: if the prediction uncertainty value of the current student model is low, increasing the distillation intensity weight, if the prediction uncertainty value of the current student model is high, decreasing the distillation intensity weight; if the discrimination ability of the current student model is strong, decreasing the temperature value, if the discrimination ability of the current student model is weak, increasing the distillation temperature value.
[0009] Optionally, dynamically adjusting the distillation intensity weight and distillation temperature value according to the prediction uncertainty value and the discrimination ability further includes: using the adjusted distillation intensity weight and distillation temperature value to calculate the knowledge distillation loss and the task loss to obtain the total loss function; updating the parameters of the student model according to the total loss function.
[0010] Optionally, iteratively executing the process of training the student model until the performance index of the student model reaches a preset threshold to complete dynamic knowledge distillation includes: evaluating the model ability, training efficiency, and business adaptability of the trained student model.
[0011] Optionally, evaluating the model ability, training efficiency, and business adaptability of the trained student model includes: evaluating the model ability includes calculating the relative entropy output by the current student model, evaluating the training efficiency includes calculating the parameter update variance of the current student model and the convergence speed of the current student model, evaluating the business adaptability includes calculating the entity coverage rate and multi-round dialogue coherence of the current student model, summarizing the above evaluations to obtain the comprehensive score of the student model, and setting the training termination condition of the student model according to the comprehensive score of the student model.
[0012] Optionally, setting the termination condition for training the student model according to the comprehensive score of the student model includes: if any one or more individual indicators among the model ability, training efficiency, and business adaptability of the current student model are lower than the preset indicator threshold for three consecutive rounds, then fine-tuning is triggered.
[0013] Optionally, the fine-tuning of the student model includes: inputting the fine-tuning input Q&A pairs into the trained student model to generate fine-tuning prediction outputs; calculating the total loss of the fine-tuning prediction outputs, and adjusting the learning rate of the trained student model according to the total loss of the fine-tuning prediction outputs; re-evaluating the student model after the learning rate adjustment until the set termination condition for training the student model is met.
[0014] On the other hand, the present invention provides a large model dynamic knowledge distillation system for implementing the large model dynamic knowledge distillation method. The system includes a control module. The control module includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the large model dynamic knowledge distillation method.
[0015] In the above technical solution, the same student model is reused to generate soft labels, eliminating the need to introduce a teacher model, reducing the occupation of computing resources. By predicting the uncertainty value to dynamically adjust the distillation weight, the knowledge transfer efficiency for high-confidence samples is enhanced, and the noise interference for fuzzy samples is reduced, significantly improving the output stability. Secondly, by quantifying the distribution difference between model iterations through KL divergence and dynamically adjusting the distillation temperature parameter, the confusion rate of category boundaries decreases in interpretation tasks. At the same time, the high-temperature mechanism captures potential cross-domain associations, enhancing the generalization ability of complex semantics. Thirdly, a scenario-customized evaluation system is constructed, integrating business indicators such as entity coverage rate and multi-round dialogue coherence. Combining the dynamic termination condition and the fine-tuning mechanism, it avoids the overfitting risk of traditional static distillation. This solution breaks through the dependence on fixed data distribution in static distillation, and through the parameter dynamic linkage mechanism, it adapts to the dynamic evolution characteristics of data. While ensuring high-precision knowledge transfer, it significantly enhances the robustness of the model in complex scenarios such as multi-level policy interpretation and cross-department collaborative Q&A, providing a reliable technical support for intelligent services.
[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation section. Description of the Drawings
[0017] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings: Figure 1It is a flowchart of the dynamic knowledge distillation method for large models.
[0018] Figure 2 It is a schematic diagram of model parameter update. Specific implementation manners
[0019] The following combines the attached Figure 1 - attached Figure 2 The specific implementation manners of the embodiments of the present invention will be described in detail. It should be understood that the specific implementation manners described here are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0020] It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solution of this application all comply with the relevant regulations of national laws and regulations. In the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0021] The inventors of this application found in the process of implementing the present invention that in the prior art, the static knowledge distillation method relies on fixed parameters, which is difficult to adapt to dynamic updates and cross-department data heterogeneity, resulting in large fluctuations in prediction confidence and incomplete entity coverage; there is a lack of a semantic boundary adaptive mechanism, making it difficult to accurately analyze multi-level policy associations.
[0022] Embodiment 1 Refer to Figure 1 - Figure 2 , which is the first embodiment of the present invention. This embodiment provides a dynamic knowledge distillation method for large models, including: S100: Obtain a domain question-and-answer pair dataset, randomly collect input question-and-answer pairs in the question-and-answer pair dataset, and divide the input question-and-answer pairs into training input question-and-answer pairs and fine-tuning input question-and-answer pairs.
[0023] For example, this solution can be used in the government affairs field, and data can be randomly collected from a public question-and-answer library or directly from a preset question-dialog training dataset. Before each round of iteration, 20% of the data is randomly sampled from the training dataset. Here, 20% of the data includes input question-and-answer pairs in the form of single-round conversations and their corresponding soft labels.
[0024] Furthermore, divide the input question-and-answer pairs into training input question-and-answer pairs and fine-tuning input question-and-answer pairs, and perform preprocessing. The training input question-and-answer pairs are 15% of the data in the random sampling, and the fine-tuning input question-and-answer pairs are 5% of the data in the random sampling.
[0025] Preferably, perform random masking on the training input question-and-answer pairs. The random masking accounts for 15% of the training input question-and-answer pairs to obtain masked samples.
[0026] S200: Load the pre-trained language model as the student model, and use the student model to encode the training input Q&A pairs to generate the first soft label.
[0027] For example, fully load the pre-trained language model (e.g., Qwen1.5-14B) as the student model, and do not freeze any network layers. Convert the input Q&A pairs (questions and answer labels) into a format that can be processed by the pre-trained language model (such as tokenization, vectorization, etc.). Pass the input Q&A pairs (such as the pre-trained Qwen1.5-14B), that is, the current student model (at round t-1), through the model to encode the input sentences and make a preliminary prediction on the input data to generate the first soft label (soft-label). Store the generated first soft label. At this time, no distillation is performed, which serves as a reference for the next iteration, that is, as the iterative reference for subsequent dynamic knowledge distillation.
[0028] Preferably, fully load large models such as Qwen1.5-14B, retain the general semantic understanding ability in the pre-training stage, and there is no need to train from scratch; generate the first soft label (prediction distribution at round t-1) to provide a comparison benchmark for each subsequent iteration; do not freeze the network layers, allowing the model to simultaneously adjust the underlying semantic encoding and high-level task adaptation ability during distillation.
[0029] Preferably, before generating the first soft label, add a masked task branch, introduce a dynamic masking mechanism based on attention weights, preferentially mask high-attention regions, input the masked samples into the student model to obtain self-supervised pseudo-labels, and combine them with the first soft label to form a dual-path soft label generation. Different from traditional static masking, this solution dynamically selects the masking positions.
[0030] S300: Perform distillation, and again use the student model to encode the training input Q&A pairs to generate the second soft label.
[0031] For example, load the input Q&A pairs, the input format is single-round dialogue, and again use the student model (at round t) to encode the training input Q&A pairs to generate the second soft label.
[0032] Preferably, generate the second soft label (prediction distribution at round t), which reflects the latest state of the current model. Comparing it with the first soft label can quantify the iterative effect. Reusing the same student model to generate the second soft label eliminates the need to introduce a teacher model and reduces the consumption of computing resources.
[0033] S400: According to the first soft label and the second soft label, calculate the prediction uncertainty value and discrimination ability of the student model.
[0034] In a preferred embodiment of the present invention, calculating the prediction uncertainty value and discrimination ability of the student model includes: inputting the training input question-and-answer pair into the student model for prediction to generate a probability distribution of each possible output category, and this probability distribution is the second soft label of the current round; quantifying the discreteness of the second soft label through probability variance to obtain the prediction uncertainty value; setting a first variance threshold, if the prediction uncertainty value is less than or equal to the first variance threshold, it is determined that the prediction uncertainty value of the current student model is low; if the prediction uncertainty value is greater than the first variance threshold, it is determined that the prediction uncertainty value of the current student model is high.
[0035] Further, the prediction uncertainty value is calculated as follows:
[0036] Where, represents the prediction uncertainty value; represents the variance of the probability distribution, measuring the uncertainty of the student model about the prediction result. The larger this variance, the higher the uncertainty of the student model about the prediction result, and the smaller this variance, the higher the certainty of the student model about the prediction result. represents the student model in the t round of iteration for the input sample i generates the probability distribution.
[0037] It should be noted that the first variance threshold is not a fixed threshold and is set according to the dynamic sorting or ratio of the variances. For example, in each iteration, calculate the of all samples, and take the 40% quantile of the distribution as the first variance threshold.
[0038] Further, calculating the prediction uncertainty value and discrimination ability of the student model further includes: retrieving the first soft label, and quantifying the distribution difference between the first soft label and the second soft label through relative entropy (calculating the KL divergence, Kullback-Leibler Divergence) to obtain the discrimination ability value; setting a first discrimination ability threshold, if the discrimination ability is less than or equal to the first discrimination ability threshold, it is determined that the current student model has strong discrimination ability, and if the discrimination ability is greater than the first discrimination ability threshold, it is determined that the current student model has weak discrimination ability.
[0039] Further, the discrimination ability is calculated as follows:
[0040] Where, represents the discrimination ability, represents the KL divergence, represents the soft label distribution in the t-th round, represents the soft label distribution in the (t - 1)-th round.
[0041] Furthermore, the first discrimination ability threshold is set according to the KL divergence distribution statistics. Calculate the KL divergence values for a large number of samples, and statistically analyze the distribution of the KL divergence values, such as the mean and standard deviation. Set the discrimination ability threshold according to the KL divergence distribution statistics, which can be set to 0.35.
[0042] Preferably, the prediction confidence is measured by the variance of probability, which directly reflects the model's mastery of domain knowledge and avoids the misguidance of hard labels; the variance threshold is dynamically adjusted based on the quantile (such as the 40% quantile) to adapt to different data distributions and training stages, improving robustness. Use the KL divergence to compare the soft label distributions of the previous and current rounds to quantify the effectiveness (discrimination ability) of the model parameter update, which can effectively guide the adjustment of the subsequent distillation intensity.
[0043] Preferably, the first soft label can be replaced with a dual-path soft label in this step. The dual-path soft label fusion technology automatically balances the semantic association between the masked samples and the original samples, which can significantly improve the context understanding ability of polysemous words in the document and reduce the risk of intention misjudgment.
[0044] S500: Dynamically adjust the distillation intensity weight and the distillation temperature value according to the predicted uncertainty value and the discrimination ability.
[0045] Specifically, if the predicted uncertainty value of the current student model is low, increase the distillation intensity weight; if the predicted uncertainty value of the current student model is high, decrease the distillation intensity weight; if the discrimination ability of the current student model is strong, decrease the temperature value; if the discrimination ability of the current student model is weak, increase the distillation temperature value.
[0046] Furthermore, the adjustment formula for the distillation intensity weight is as follows:
[0047] where, represents the distillation weight, k is a hyperparameter used to control the sensitivity of the predicted uncertainty value, represents the predicted uncertainty value.
[0048] Furthermore, the adjustment formula for the distillation temperature value is as follows:
[0049] where, represents the distillation temperature value, represents the initial set base temperature value, represents the adjustment coefficient, represents the discrimination ability.
[0050] Further, using the adjusted distillation intensity weight and distillation temperature value, calculate the knowledge distillation loss and the task loss to obtain the total loss function; the calculation formula of the total loss function is as follows:
[0051] Among them, represents the total loss function, represents the distillation weight, represents the knowledge distillation loss, represents the distillation temperature value, represents the task loss.
[0052] Preferably, if double-path soft labels are used, a masking loss needs to be added to the total loss function, and the total loss function is as follows:
[0053] Among them, represents the weight of the masking loss, represents the masking loss.
[0054] Further, update the parameters of the student model according to the total loss function, perform backpropagation on to calculate the total gradient of the parameters of the student model. The calculation formula of the total gradient is as follows:
[0055] Among them, represents the total loss function, represents the distillation weight, represents the knowledge distillation loss, represents the distillation temperature value, represents the task loss, represents the parameter gradient.
[0056] Similarly, if double-path labels are used, the calculation formula of the total gradient is:
[0057] Among them, represents the gradient of the parameter , represents the weight of the masking loss, represents the masking loss.
[0058] Preferably, if double-path soft labels are used, through the triple mechanisms of dynamic masking position selection, dual-channel soft label fusion, and gradient collaborative control, the knowledge transfer efficiency and model robustness of the question-answering system in the government affairs field are significantly improved.
[0059] Further, update the student model parameters according to the total gradient, and the parameter update formula is as follows:
[0060] where represents the learning rate, represents the parameters of the student model gradient, represents the total loss function, represents the parameters of the student model after the t-th round of iterative update, represents the parameters of the student model after the end of the (t - 1)-th round of iteration.
[0061] Preferably, when the predicted uncertainty value is low, increase the weight to force the model to learn high-confidence samples; when the predicted uncertainty value is high, decrease the weight to avoid noise interference; when the discrimination ability is strong, decrease the temperature (sharpen the probability distribution) to enhance the model's decision-making confidence; when the discrimination ability is weak, increase the temperature (smooth the distribution) to encourage exploration of potential patterns. The total loss function combines the task loss (such as cross-entropy) and the distillation loss, and dynamically weights and balances the domain knowledge transfer and task performance optimization.
[0062] S600: Iteratively execute the process of training the student model until the performance index of the student model reaches a preset threshold, and complete the dynamic knowledge distillation.
[0063] Specifically, the process of training the student model is S200 to S500.
[0064] Further, evaluate the model ability, training efficiency, and business adaptability of the trained student model. Evaluating the model ability includes calculating the relative entropy output by the current student model, evaluating the training efficiency includes calculating the parameter update variance of the current student model and the convergence speed of the current student model, and evaluating the business adaptability includes calculating the entity coverage rate and multi-turn dialogue coherence of the current student model.
[0065] Further, the calculation formula of the entity coverage rate is as follows:
[0066] It should be noted that the higher the entity coverage rate, the stronger the adaptability of the current student model to domain knowledge.
[0067] Further, the multi-turn dialogue coherence measures the context logical consistency in multi-turn dialogues through manual evaluation or automated evaluation (such as BLEU, ROUGE). The higher the coherence, the stronger the practicality of the current student model in the dialogue scenario. The calculations of other evaluations belong to the conventional technical means in this field and will not be described here.
[0068] Further, summarize the above evaluations to obtain a comprehensive model score, and set the training termination condition for the student model according to the comprehensive model score.
[0069]
[0070] Among them, 、 、 are all weight coefficients.
[0071] Preferably, , =0.5, =0.3, =0.2.
[0072] Further, setting the training termination condition for the student model according to the comprehensive score of the student model includes: the number of iterations reaches the set maximum value, or the comprehensive model score reaches the set score threshold and the model ability, training efficiency, and business adaptability of the current student model all reach the preset threshold, and meeting either of the two is sufficient. At this time, it can be considered that the performance of the student model reaches a satisfactory level and the training terminates.
[0073] Preferably, if any one or more of the single - item indicators of the model ability, training efficiency, and business adaptability of the current student model are lower than the preset indicator threshold for multiple consecutive rounds (preferably 3 rounds), then trigger fine - tuning.
[0074] It should be noted that the single - item indicators are the model ability indicators, training efficiency indicators, and business adaptability indicators corresponding to the above - mentioned scheme. An indicator threshold is set for each indicator, and the indicator threshold is set according to specific requirements.
[0075] Further, setting the training termination condition for the student model according to the comprehensive score of the student model also includes: input the fine - tuning input Q&A pairs into the trained student model to generate a fine - tuning prediction output, calculate the total loss of the fine - tuning prediction output, and adjust the learning rate of the trained student model according to the total loss of the fine - tuning prediction output; re - evaluate the student model after the learning rate adjustment until the set training termination condition for the student model is met and then end the adjustment.
[0076] Further, preset a fine - tuning loss threshold. If the total loss of the fine - tuning prediction output is less than or equal to the preset fine - tuning loss threshold, then magnify the current learning rate by 10% to accelerate the model convergence. If the total loss of the fine - tuning prediction output is greater than the preset fine - tuning loss threshold, then halve the current learning rate to make the student model update more stably.
[0077] Preferably, by integrating model ability, training efficiency, and business adaptability, single - metric deviation is avoided; termination is triggered based on comprehensive scoring or non - compliance with metrics in multiple consecutive rounds to prevent overfitting or ineffective iteration and save computing resources; the learning rate is dynamically adjusted according to the loss of the fine - tuning set (accelerating convergence when the loss is low and updating stably when the loss is high), enhancing the robustness of the final student model.
[0078] The present invention also provides a large - model dynamic knowledge distillation system for implementing the large - model dynamic knowledge distillation method. The system includes a control module. The control module includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the large - model dynamic knowledge distillation method.
[0079] An embodiment of the present invention provides a storage medium with a program stored thereon. When the program is executed by a processor, the large - model dynamic knowledge distillation method is implemented.
[0080] An embodiment of the present invention provides a processor for running a program. When the program runs, the large - model dynamic knowledge distillation method is executed.
[0081] An embodiment of the present invention provides a device. The device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, large - model dynamic knowledge distillation is implemented. The device herein can be a server, a PC, a PAD, a mobile phone, etc.
[0082] The present application also provides a computer program product that is suitable for executing large - model dynamic knowledge distillation when executed on a data - processing device.
[0083] Those skilled in the art should understand that the embodiments of the present application can provide a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk memory, CD - ROM, optical memory, etc.) containing computer - usable program code.
[0084] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0087] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0088] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0089] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0090] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0091] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for dynamic knowledge distillation of large models, characterized in that, Including: Obtain a domain Q&A pair dataset, randomly collect input Q&A pairs from the Q&A pair dataset, and divide the input Q&A pairs into training input Q&A pairs and fine-tuning input Q&A pairs; Load a pre-trained language model as the student model, and use the student model to encode the training input Q&A pairs to generate a first soft label; Perform distillation, and use the student model again to encode the training input Q&A pairs to generate a second soft label; According to the first soft label and the second soft label, calculate the prediction uncertainty value and discrimination ability of the student model; According to the prediction uncertainty value and the discrimination ability, dynamically adjust the distillation intensity weight and the distillation temperature value; Iteratively execute the process of training the student model until the performance index of the student model reaches a preset threshold, and complete dynamic knowledge distillation.
2. The large model dynamic knowledge distillation method according to claim 1, wherein The calculating the prediction uncertainty value and discrimination ability of the student model includes: Input the training input Q&A pairs into the student model for prediction to generate a probability distribution of each possible output category, and this probability distribution is the second soft label of the current round; Quantify the discreteness of the second soft label through probability variance to obtain the prediction uncertainty value; Set a first variance threshold. If the prediction uncertainty value is less than or equal to the first variance threshold, it is determined that the prediction uncertainty value of the current student model is low; if the prediction uncertainty value is greater than the first variance threshold, it is determined that the prediction uncertainty value of the current student model is high.
3. The large model dynamic knowledge distillation method according to claim 2, characterized in that, The calculating the prediction uncertainty value and discrimination ability of the student model further includes: Retrieve the first soft label, and quantify the distribution difference between the first soft label and the second soft label through relative entropy to obtain the discrimination ability value; Set a first discrimination ability threshold. If the discrimination ability value is less than or equal to the first discrimination ability threshold, it is determined that the current student model has strong discrimination ability; if the discrimination ability value is greater than the first discrimination ability threshold, it is determined that the current student model has weak discrimination ability.
4. The large model dynamic knowledge distillation method according to claim 3, characterized in that The dynamically adjusting the distillation intensity weight and the distillation temperature value according to the prediction uncertainty value and the discrimination ability includes: If the prediction uncertainty value of the current student model is low, increase the distillation intensity weight; if the prediction uncertainty value of the current student model is high, decrease the distillation intensity weight; If the current student model has strong discrimination ability, lower the temperature value; if the current student model has weak discrimination ability, increase the distillation temperature value.
5. The large model dynamic knowledge distillation method according to claim 1, wherein The dynamically adjusting the distillation intensity weight and the distillation temperature value according to the prediction uncertainty value and the discrimination ability further includes: Use the adjusted distillation intensity weight and distillation temperature value to calculate the knowledge distillation loss and the task loss to obtain the total loss function; Update the parameters of the student model according to the total loss function.
6. The large model dynamic knowledge distillation method according to claim 1, wherein The iteratively executing the process of training the student model until the performance index of the student model reaches a preset threshold and completing dynamic knowledge distillation includes: evaluating the model ability, training efficiency, and business adaptability of the trained student model.
7. The large model dynamic knowledge distillation method according to claim 6, wherein The evaluating the model ability, training efficiency, and business adaptability of the trained student model includes: Evaluating the model ability includes calculating the relative entropy output by the current student model, Evaluating the training efficiency includes calculating the parameter update variance of the current student model and the convergence speed of the current student model. Evaluating the business adaptability includes calculating the entity coverage rate of the current student model and the coherence of multi-round conversations. Summarize the above evaluations to obtain a comprehensive score of the student model, and set the training termination condition of the student model according to the comprehensive score of the student model.
8. The large model dynamic knowledge distillation method according to claim 7, characterized in that, Setting the training termination condition of the student model according to the comprehensive score of the student model includes: if any one or more single indicators of the model ability, training efficiency, and business adaptability of the current student model are lower than the preset indicator threshold for multiple consecutive rounds, then trigger fine-tuning.
9. The large model dynamic knowledge distillation method according to claim 8, wherein The fine-tuning of the student model includes: Input the fine-tuning input Q&A pairs into the trained student model to generate fine-tuning prediction outputs. Calculate the total loss of the fine-tuning prediction outputs, and adjust the learning rate of the trained student model according to the total loss of the fine-tuning prediction outputs. Re-evaluate the student model after the learning rate adjustment until the set training termination condition of the student model is met.
10. A large model dynamic knowledge distillation system, characterized in that The system includes a control module, and the control module includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the large model dynamic knowledge distillation method according to any one of claims 1-9.
Citation Information
Cited By
Knowledge distillation method and device based on AI large model and electronic equipment
CN120654774A
Question and answer model training method and question and answer task processing method
CN120744074A
Question and answer model training method and question and answer task processing method
CN120744074B
Lightweight target detection knowledge distillation method
CN121119044A